Monday, November 25, 2013

Afternoon mental health break: Missing dialogue from Star Trek

from ath716:

Missing dialogue from Star Trek:

Kirk: "Scottie I need full power!"
Scottie: "We'll need 6 weeks in spacedock, captain."
Kirk: "You've got 30 seconds."
[30 second later]
Kirk: "Scottie, do we have full power?"
Scottie: "Did we spend 6 weeks in spacedock?"
Kirk: "No."
Scottie: "Then you don't have full power, do you?"

Thursday, November 21, 2013

Links for 11-21-2013

Video: Hybrid cloud and software-defined storage

I had a clip during a Red Hat Storage Server (Gluster) webcast a while back. I talk about why distributed software-only storage is especially important in hybrid cloud environments.

Wednesday, November 20, 2013

Red Hat Open Hybrid Cloud: My ASEAN presentation and text

This presentation on Open Hybrid Clouds is a version of the one I gave at Red Hat Forum in Asia in October 2013. So let's get started.













You want to scale like Google, you want to manage data like Facebook, and you want to have the agility of Amazon.com and Amazon Web Services. Or maybe you're thinking "Well, no, because those big Internet giants are really unique. They're nothing like my enterprise, or my organization."

One of the things I want to convince you of is that while those companies certainly are unique, and certainly are not typical, they also in many ways shine a spotlight on where IT is going, and therefore where the IT in your organization is will be going.


 Let's first talk about scale. However you define computing scale in your own organization, is probably increasing and this isn't just about the number of servers. It's about the exposed of mobile, which is one of the "third platform technologies," that market researchers IDC talk about.

IT is also getting more and more central to businesses of more and more firms. It touches more customers, more processes, and more data and his in just in the traditional IT centric industries, like financial services. We also see the role of IT increasing in mining, in agriculture, in heavy manufacturing, indeed, just about everywhere.

Now, we move on to talk about managing data. Here's a quote from McKinsey and Company, the management consultants, a couple years ago, saying, The use of big data will become a key basis of competition and growth for individual firms and that all companies, therefore, need to take big data seriously." I could argue that we've maybe done ourselves something of a disservice, as an industry with this big data term, because it tends to focus the attention on the volume of data or maybe the real-timeness of data, or the variety of data.

While that's certainly a trend, I'd argue that perhaps the bigger trend is just the pervasive use of data in many forms. MIT's Andrew McAfee told an interesting anecdote at the MIT Sloan CIO Symposium earlier this year when he talked about how the publishing industry was increasingly moving from a culture of lunches which is to say that people got together over lunch and discussed their hunches about what books were going to sell well and which weren't, and instead is being replaced by a culture where there is more systematic use of data and of course that's driven by Amazon.com and the other Internet businesses.

The third characteristic is agility, flexibility, and this takes a number of different forms. What I've shown here is really the developer-centric view of that where developer needs to be able to very quickly code and deploy and have their users and have their businesses reap the benefits of that code.
The resources devoted to given application need to be very rapidly scale up and down. And the applications themselves need to be developed more quickly given that they are both more familiar and are also being more directly connected to revenue than ever before. So agility, flexibility, ability to spin up and down resources for promotions or whatever the business needs are becoming more important than ever.

And what that all comes together to say is that the old way doesn't get you there any longer.
There is business demand for these different innovations. There are business demands for these new ways of doing business. And traditional IT infrastructure doesn't get you there. What you are left with is what you can call an IT services gap or an IT innovation gap; it doesn't really matter that much what you call it. But the bottom line is, you can't just keep doing things in your traditional way and hope to meet up with these new business demands for innovation.

The way the public cloud providers and Internet giants have done it is with open source, open standards and with the open innovation model that's made possible by and is driven by open source. If you look at the big public clouds, they are all running Linux. They are all making extensive use of other types of open source software to greater or lesser degrees. They are making contributions back into a variety of open source projects that touch in areas that are of particular interest to them.

But even if we look at what's running on those public clouds at Amazon Web Services for example, there is far more Linux than there is Windows. And this just reflects that the open source way is where a lot of the new innovation is happening out there. It's not happening in proprietary software world as much as it is happening in the open source world.

Let's talk OpenStack as an example of how this open innovation is happening. A huge number of companies is contributing to it, and in fact I'd argue that one of the main reasons why OpenStack is such an interesting project is not because of some specific interesting technologies within OpenStack today (although those are certainly there), but rather because such a community of developers and partners and others have gathered around OpenStack and that massive participation is what's made OpenStack so interesting.

But it's not just about what's happening at the Internet giants and with Web 2.0 companies. This is some research the IDC conducted last year of a range of organizations, about what their feelings about having an open cloud were, or open cloud technologies, and 72 percent said it was either an absolute requirement or at the very least a key factor. This importance of openness and clouds is not unique to just one space. It spreads across the entire industry.
 I am also going to argue that transformed IT requires a hybrid approach. Now by hybrid here I don't mean the narrow cloud-bursting "move workloads from public clouds to private clouds minute to minute" as was sometimes talked about when that term was first introduced.

Rather I mean hybrid as this combination of different technology platforms, some of which are on premise, some of which are in various types of public cloud providers and you have a choice of which of those different technology platforms and which of those different types of ownership and payment models you want to choose on an individual project basis.

And then IT really manages this entire portfolio that is running in different places. So let's bring this all together. IT is becoming more strategic, because they have to, to deliver these services and capabilities that the business needs. They are hybrid so they are going to be managing services that are being deployed in many different places, and they are achieving that with open source, open standards-based technologies. Now I've been using IT as a single term.

Now I am going to discuss breaking that apart a little bit and talk about developers, IT Operations, and the business. These different constituencies are in fact working more closely together, I'd argue, than they have historically but I think it's still useful to look at the kind of different requirements and mandates for each of those parts of the overall business.

So let's first of all talk about the developer. The developer's role has arguably changed more than anyone's in recent years. My former analyst colleague Stephen O'Grady came out with a book last year called The New Kingmakers in which he made the argument that developers were increasingly driving decision making within an IT organization.

I think this argument can be overstated, but nonetheless it's an important dynamic as we move forward and as new applications and new services are important to more and more different types of businesses.


So what does the developer care about? Well, the developer has all the applications to write, so the developer wants to be productive and part of being productive is having access to the tools that they need to do their jobs. These include getting access to the new innovations that are happening in open source. There's very little in the way of language development these days that isn't open source, so they need access to all that. They also want access to all of their familiar tools as well. They also want to be able to focus on the details that matter.

A lot of the mechanics of getting a server and setting it up and deploying it are plumbing a developer needs to have done so they can develop applications, but it's not actually something that developers care about. They also need ways to integrate this new application development they're doing with existing software and processes within their business.
How do we do this? Platform as a Service is one important ingredient here. Here is a quote from Gartner: "The use of Platform as a Service technology will enable IT organizations to become more agile and more responsive to the business needs."

If we look at this slide, it breaks down the tasks needed to write and deploy an application or at least some of them. The point here is that there are a lot fewer of those steps with a PaaS then there is with a physical server. In fact, even a virtualized server doesn't necessarily help all that much; PaaS is really about getting the infrastructure out of the way of the developer.

 In terms of integrating with existing applications and processes, this is an interesting announcement that Red Hat made back in October, JBoss xPaaS for OpenShift.

What you see here is you've got the OpenShift PaaS on the bottom here. You have the JEE application server on top there; lots of people today are using OpenShift to develop Java applications. And then on top you have the integration PaaS, the BPM PaaS, and the mobile PaaS.

What we're doing here is using various elements of JBoss Middleware portfolio, such as messaging and business rules and so forth, and bringing them into the PaaS environment. What this does is it really enables applications developed with OpenShift to be integrate with existing business processes, with existing enterprise applications, and so forth. We really view Platform as a Service as not something that's just about developing certain styles of new applications but something that can be used as a fundamental part of enterprise application development.

 If we shift our focus to IT operations now, what are they trying to do? I could actually summarize almost every thing that's in the right hand side of this slide with one word and that would be automation.

If I had two words, I'd probably use something like relentless automation because if we talk about moving to cloud environments that is really one of the very key pieces, the need to delegate out to end users for self service but then automate that provisioning.

Automate things like scaling. Automate things like security remediation and so forth. Ultimately, this is all about creating standardization because standardization not only reduces the amount of work that's needed to keep environments up and running, but it also creates a much more consistent experience and experience that can be much more federated and while kept under a centralized policy control.

There is another interesting aspect with how infrastructures are evolving for the cloud. They're evolving to take into account what are largely new styles of workloads. If you look at the existing sometimes called "systems of record" workloads, you have individual instances, which are carefully maintained. They may potentially conflict with other services and you may need to keep them isolated in that respect.

This is typically a complex setup and once you have the thing up and running, you don't want to breathe on it too much, if you would. You tend to keep these production workloads around for a long time, whereas if you look at cloud style workloads, they're much smaller, they're much more modular, they tend to be replicated a lot. That's where that automation comes in and they're just overall much lighter weight. Bill Baker who used to be at Microsoft sometimes refers to traditional workloads as pets. If a pet gets sick, you take it to the vet and try to make it better. New-style workloads, on the other hand are cattle. If the cow gets sick, well, you get a new cow.

That is the model of these two workloads, the stateful and the stateless.

We even announced at Red Hat a particular product offering that explicitly recognizes these two different styles of workloads as well as the hybrid workloads that combine the two. So you your traditional enterprise virtualization for the traditional enterprise workloads. You've got Red Hat Enterprise Virtualization there.

Your cloud-style workloads are a great match for Red Hat Enterprise Linux OpenStack Platform.

Then, with Red Hat CloudForms, where you have a cloud management platform, you get common policy and that single pane of glass management on top of both those styles.


Finally, you have the business imperative.

Business basically cares about delivery of revenue producing services. The business is also in many cases interested in keeping some abstraction between underlying infrastructure and the application development on top of that. This matters in a number of different environments, but it's perhaps most obvious in a certain types of government procurement situations because often agencies want to run their own infrastructure and contract out to do application development for example and there are some real benefits from a procurement standpoint in keeping those two layers separate from each other and ultimately what we are moving towards here is having a set of services that the business users can consume wherever and whenever they need to.
 This is another analyst quote, this time from Forrester Research: "Management in the future of Cloud shifts to providing measuring and using portfolio services, focus and building enterprise libraries of reusable application images and workloads for public clouds and hybrid consumption." This really is a fundamentally new way of doing things that's perhaps one of the most radical changes that we have seen in how IT operates within an organization.
Much of this is being driven by open source. I talked about open source in the context of public clouds earlier but you look at say Red Hat's portfolio here, all of our products are open source. They're all driven by open source innovation, whether we are talking hybrid cloud management with CloudForms, Platform-as-a-Service with OpenShift, Infrastructure-as-a-Service with OpenStack, distributed storage with Red Hat Storage Server, enterprise virtualization with Red Hat Enterprise Virtualization, JBoss Middleware and of course Red Hat Enterprise Linux which forms the foundation for so much of this.
And that's Red Hat Open Hybrid Cloud. It's a rich portfolio of upstream communities as well as commercial products. It all ties in with each other so that, I know it's a cliche but the whole is really greater than the sum of the parts with Red Hat's Cloud Portfolio.

Links for 11-20-2013

Monday, November 18, 2013

Links for 11-18-2013

Parsing USGS river conditions in JSON

Over the summer, I wrote a post about writing an application for OpenShift by Red Hat to display current USGS river levels on an interactive map. That post went into quite a bit of detail about populating the MongoDB database, using MongoDB's geo features, displaying the map, and so forth. However, in the interests of keeping the post to a manageable length, I only lightly touched on the topic of retrieving and interpreting the data from the USGS in the first place. That will be the topic of this post. It's a good case study because this particular data set is relatively complex.

Historically, it could be fairly difficult to retrieve USGS data. As the site itself notes: "Most data from the USGS Water Data for the Nation site are currently downloaded as tab-delimited (rdb) data files. While this approach works, it uses 20th century approaches rather than 21st century approaches." Fortunately, for our purposes here, the current conditions data can now be retrieved using the USGS instantaneous values web service in either XML or JSON (Javascript Object Notation) format. This post will describe the procedure for using JSON.

Determining the syntactically correct URL to use with the USGS Instantaneous Values REST web service

The first step is to determine the right URL or URLs to use with the web service. The USGS provides a nice tool that allows you to figure this out interactively. You should note that the USGS won't let you pull data associated with all their gauges at one time; you have to specify at least one "major filter." For my purposes, the best approach seemed to be to create a list of two letter lowercase state abbreviations (plus DC and Puerto Rico) and create URLs for each individual state. I also observed that the service didn't always return a result--timing out instead. Therefore, having an update failure of only a subset of the total data would be preferred in any case.

Retrieving the data

The below code shows how to retrieve current conditions in this way in Python 2.7 using urllib2:

statelist = ["al","ak","az","ar","ca","co","ct","de","dc","fl","ga","hi","id","il","in","ia","ks","ky","la","me","md","ma","mi","mn","ms","mo","mt","ne","nv","nh","nj","nm","ny","nc","nd","oh","ok","or","pa","ri","sc","sd","tn","tx","ut","vt","va","wa","wv","wi","wy","pr"]
for i in statelist: 
    requesturl = "http://waterservices.usgs.gov/nwis/iv/?format=json,1.1&stateCd=" + i+"&parameterCd=00060,00065&siteType=ST" 
    req = urllib2.Request(requesturl) 
    opener = urllib2.build_opener() 
    f = opener.open(req) 
    entry = json.loads(f.read())

Each request returns all the gauges in the state of type "ST" (stream).

For each iteration (state), we now need to parse the contents of variable entry. Before doing so, however, let's look at what the USGS has sent to our computer.

The JSON

If you just type an appropriate URL into your browser, you'll get back a big block of text. For example, the URL:

http://waterservices.usgs.gov/nwis/iv/?format=json,1.1&stateCd=ma&parameterCd=00060,00065&siteType=ST

returns a string that starts with the following fragment:

{"name":"ns1:timeSeriesResponseType","declaredType":"org.cuahsi.waterml.TimeSeriesResponseType","scope":"javax.xml.bind.JAXBElement$GlobalScope","value":{"queryInfo":{"creationTime":null,"queryURL":"http://waterservices.usgs.gov/nwis/iv/","criteria":{"locationParam":"[]","variableParam":"[00060, 00065]","timeParam":null,"parameter":[],"methodCalled":null},"note":[{"value":"[ma]","type":null,"href":null,"title":"filter:stateCd","show":null},{"value":"[ST]","type":null,"href":null,"title":"filter:siteType","show":null},{"value":"[mode=LATEST, modifiedSince=null]","type":null,"href":null,"title":"filter:timeRange","show":null},{"value":"methodIds=[ALL]","type":null,"href":null,"title":"filter:methodId","show":null},{"value":"2013-11-18T17:59:12.444Z","type":null,"href":null,"title":"requestDT","show":null},{"value":"2172fbb0-507b-11e3-9141-6cae8b6642ea","type":null,"href":null,"title":"requestId","show":null},{"value":"Provisional data are subject to revision. Go to http://waterdata.usgs.gov/nwis/help/?provisional for more information.","type":null,"href":null,"title":"disclaimer","show":null},{"value":"sdas01","type":null,"href":null,"title":"server","show":null}],"extension":null},"timeSeries":[{"sourceInfo":{"siteName":"NORTH NASHUA RIVER AT FITCHBURG, MA","siteCode":[{"value":"01094400","network":"NWIS","siteID":null,"agencyCode":"USGS","agencyName":null,"default":null}],"timeZoneInfo":{"defaultTimeZone":{"zoneOffset":"-05:00","zoneAbbreviation":"EST"},"daylightSavingsTimeZone":{"zoneOffset":"-04:00","zoneAbbreviation":"EDT"},"siteUsesDaylightSavingsTime":true},"geoLocation":{"geogLocation":

Not so easy to parse as you can see. The problem is that the returned JSON is deeply nested, which makes it very hard to interpret just by eyeballing it. I found that using a JSON viewer really helped. There are a lot of different ones out there but I used this one at chris.photobooks.com. (I'm including some screenshots below, but I'd encourage you to retrieve your own JSON from the USGS using your browser, paste it into this or another JSON viewer, and play along at home.

I'm not going to go through how to access every field you may be interested in, but I'll hit a few that demonstrate interesting points about the data structure.

Interpreting the data


The first point to notice is that the returned data has a singe top-level object and you actually have to go two levels deep under [value][timeSeries] before you get to the key to the data structures that contain the current conditions data for each of the gauge locations.

So let's say we want to iterate over this data and get to the gauge number. Here's what that looks like:

Screen Shot 2013 11 18 at 1 15 03 PM

count = int (len(entry['value']['timeSeries']) - 1)
while count >= 0:
#We construct an array of the relevant values associated with a guage number
#Note that gage height and discharge are in separate entries
#Right here we're just filling out the "permanent" values
#Gauge Number. This will be the dictionary index
agaugenum = entry['value']['timeSeries'][count]['sourceInfo']['siteCode'][0]['value']

Per the screenshot that drills down to the individual gauge level, we're first iterating on the data structure that holds the individual gauge conditions (['value']['timeSeries'][count]). Then we drill down another four levels until we get to the 'value' field holding the gauge number. This is why this data is very hard to parse just by looking at the text.

Screen Shot 2013 11 18 at 1 24 12 PM

There's also a wrinkle to be aware of when working with this data. I've been talking as if the individual "records" were for gauges. But, actually, they're not. They're for values. 

The value can be found here:

entry['value']['timeSeries'][count]['values'][0]['value'][0]['value']

But what that value means is defined here:

entry['value']['timeSeries'][count]['variable']['variableCode'][0]['variableID']

What this means in practice for stream data is that most--but not all gauges--have two records. One is for the current river height (in feet) and one is for current river flow (in cfs). A variableID of 45807202 corresponds to height and a variableID of 45807197 corresponds to flow. 

I should probably mention that my code assumes that these are the only types of variables and that their units don't vary from site to site. I haven't seen any exceptions but for any sort of critical application, I'd probably use additional data encoded in the JSON rather than just assuming it's all consistent and invariable. (My application also just updates the current conditions rather than looking for changes to any of the parameters associated with the gauge itself such as location.)

The same basic process that I've described here can be used to make use of data from a wide range of sources. For example, you can read about the structure associated with tweets on the Twitter developer site. It's easy to get started playing around. And when you get to the point where you want to run an application with a database or other components, give OpenShift Online by Red Hat a try if you're not already doing so. It's free and easy to get up and running with the language of your choice. 

Friday, November 15, 2013

Links for 11-15-2013

AWS refreshed take on the enterprise and a thought on innovation

I didn't make it out to AWS re:Invent this year but I've been following along virtually using this thing called the Internet. A few observations.

Even Amazon now accepts it's a hybrid IT world. In Amazon Web Services grudgingly accepts hybrid cloud, Nancy Gohring quotes Andy Jassy, senior vice president of AWS:

We understand that enterprises have a number of on premise data centers that they aren’t ready to retire yet,” he said. “What they really want is the ability to use those on premise data centers easily with AWS. That’s what we’re spending considerable resources enabling.

It's worth noting that statements like this are in no way inconsistent with an Amazon vision that, in the long term, most computing will happen on public clouds rather than private IT resources. But it does represent something of a concession for the nearer term.

At Red Hat, we've been very focused on hybrid approaches to cloud. On example of this is our

Screen Shot 2013 11 15 at 11 13 59 AM

 OpenShift platform-as-a-service. We just announced enhancements to our Online offering. Alternatively, organizations can implement a PaaS on-premise using OpenShift Enterprise. And, for managing hybrid environments, we offer Red Hat CloudForms; the recently announced CloudForms 3 added support for both OpenStack and for additional AWS features.

The acknowledgement of hybrid also represents something of a shift in how Amazon is positioning itself in the context of (traditional) enterprise cloud adoption. It's less about going all-in on public cloud today and more about evolving to a cloud computing architecture. That said, many of the new features that Amazon announced at the event fall squarely into the "removing enterprise objections to the public cloud" bucket. Traditional enterprises aren't the only organizations that care about security, reliability, and performance of course. But something like fine-grained access control seems particularly relevant to the must-haves list of a large IT organization.

Amazon's Workspaces virtual desktop service is also something that's likely to be primarily of interest to traditional enterprises. On the one hand, it struck me as a bit of an odd announcement. Virtual Desktop Infrastructure (VDI) is an idea that's been around for a while and everyone and their brother has taken a crack at it. There are some big VDI deployments out there but it's also never truly hit the mainstream. Higher costs than promised. Too many gotchas. More incremental alternatives. On the other hand, a lot of the VDI that is out there is already handled through some sort of managed service. Furthermore, to the degree that traditional fat client desktops increasingly revolve around running some legacy applications, perhaps that's a pretty good reason to run them using an ahgressively-priced cloud service.

As was the case last year though, the customers on stage are still dominated by the usual suspects. Of course, there was the mandatory appearance by Adrian Cockcroft of Netflix. (I don't think you're allowed to run a cloud conference without Adrian in attendance.) But, even Netflix aside, the speaker list was heavily dominated by startups with greenfield Web-oriented products like Narrative. Especially given Amazon's interest in promoting a message that it's not just for this sort of company, their high profile at the show shows that public cloud use at scale is still far more about Silicon Valley than about Main Street. 

A final thought.

In his keynote, Amazon CTO made a big deal about innovation and listening to customers. I was struck though by his point of comparison, which was to traditional proprietary software companies. What I find interesting about this is that Vogels was effectively arguing that Amazon does a better job of proprietary software development than the competition. Yet, although Amazon makes extensive use of open source software, this approach doesn't really tap into the open source software development model and the innovation that stems from that. So how different is Amazon really from Oracle?

Thursday, November 14, 2013

Links for 11-14-2013

Monday, November 11, 2013

Pro Tip for speaker notes

Going back pretty much as far as I remember, putting together speaker notes for a deck of slides (or, worse, trying to cajole someone else into putting together a set of speaker notes for their slides) was one of those perpetually incomplete tasks. It just doesn't get done. But this can be a real problem for decks designed for others to deliver--especially when the deck itself isn't especially self-documenting. It also means that "the slides" attendees want after conferences aren't, in fact, any more useful than other conference tchotchkes.

I've come across what, for me, is a very effective solution.

Regular readers will probably have noticed that I publish transcripts to go along with my podcasts. My reasoning is that a lot of folks, including yours truly, find it generally easier to skim through text than it is to find the dedicated time to listen to a podcast. I use a service called CastingWords, which is  essentially a front-end for Amazon Mechanical Turk. It's both high quality and cost effective--$1.50 per minute plus or minus depending upon how quickly you want your transcript.

For speaker notes, I now just record myself "giving" a presentation--often in slightly abbreviated form. I have that recording transcribed, do some light editing, and then cut & paste into the slide notes. One could also put the slides and associated notes into a separate text document. 

This approach is a lot less effort--and certainly less mental effort--than sitting down and typing up a full set of speaker notes. The notes are probably more detailed doing it this way as well.

Thursday, November 07, 2013

Gartner: OpenStack seen as mitigating proprietary clouds

I like this line in a CIO article by Joab Jackson:

In a recently published survey of its clients, IT research firm Gartner found that OpenStack is increasingly being viewed by organizations as an option to mitigate against being locked into proprietary cloud services.

Learning about DevOps from Science Fiction

I saw @geekygirldawn give this presentation earlier this year at LinuxCon. It's a fun look at what you can learn about working together. This slideshare doesn't come close to replicating seeing the presentation delivered of course but it's definitely worth a flip through.

Here's a version with speaker notes.

Arthur's Seat

Arthur's Seat by ghaff
Arthur's Seat, a photo by ghaff on Flickr.

I hiked up to Arthur's Seat when I was in Edinburgh for CloudOpen Europe. Both Arthur's Seat and Castle Rock--where Edinburgh Castle sits--are eroded plugs of rock from the vents of an ancient volcano.

Wednesday, November 06, 2013

My Red Hat Open Hybrid Cloud keynote from ASEAN

This video's a slightly abridged and amended version of the presentation that I did at the Red Hat Forum in ASEAN during October. I talk trends, cloud openness, hybrid IT, and developers. If you're interested in a high-level view of Red Hat's cloud strategy without a lot of product details, check this out.

Tuesday, November 05, 2013

Links for 11-05-2013

CloudForms 3.0, OpenStack, and Amazon Web Services. Oh my!

 Sadly, I couldn't make it to Hong Kong for the OpenStack Summit, but there's lots of interesting news coming out of the event. I'm going to focus here on the management piece, partly because there's still a lot of confusion about the distinction between infrastructure, like OpenStack, and the management of that infrastructure. (And, to be sure, the boundaries of where various functionality lives are still shifting a bit in these still relatively nascent days of cloud computing.)

CloudForms 3.0, which Red Hat announced today, is software that manages clouds--including OpenStack--and is therefore complementary to OpenStack and other infrastructure platforms, such as VMware vSphere, Red Hat Enterprise Virtualization, and Amazon AWS. To use Gartner's terminology, CloudForms 3.0 is a cloud management platform which Gartner defines thusly:

Cloud management platforms are integrated products that provide for the management of public, private and hybrid cloud environments. The minimum requirements to be included in this category are products that incorporate self-service interfaces, provision system images, enable metering and billing, and provide for some degree of workload optimization through established policies. More-advanced offerings may also integrate with external enterprise management systems, include service catalogs, support the configuration of storage and network resources, allow for enhanced resource management via service governors and provide advanced monitoring for improved “guest” performance and availability.

OPST PressConf FINAL odp

The cloud management capabilities provided by CloudForms 3.0 include the following:

Seamless Self-service Portals that provide users with role-delegated, automated self-provisioning of catalog-driven IT services, with requisite request approvals and integration with enterprise service catalogs. Cloud Lifecycle Service Management spans from provisioning to retirement, with automatic aging, tracking, and monitoring of services.

Advanced Chargeback, Quotas and Metering, with detailed usage tracking by configurable classifications and support for multiple rates tables (fixed cost, allocation and usage) and reservation based chargeback.

Continuous Discovery and Insight from automatic, agent-free discovery of OpenStack instances and relationships and capacity and utilization, along with configuration tracking and drift comparison.

Unified Operations Management, offering multi-site federation that provides “single pane of glass” visibility across cloud and virtual infrastructures, including runtime operations, service configuration, utilization, events, timelines, reports and customizable dashboard mash-ups.

And these cloud management capabilities are hybrid capabilities that can span heterogeneous infrastructure--whether it's heterogeneous with respect to the technology or vendor of whether it's heterogeneous in the sense of on-premise and public. This type of choice and flexibility is essential for IT organizations.

In fact, Red Hat's CIO Lee Congdon made this very point in a webinar we were doing together just this morning. (Build the cloud your developers want and your business needs) Providers change. Your calculus about what gets operated where changes. But through that, you need to maintain management continuity and control. (Lee also talks DevOps, application integration, JBoss technologies like A-MQ and PicketLink, and why Red Hat IT actually accelerated its OpenShift deployment relative to initial plans; definitely check it out!)

In addition to support for OpenStack, the new CloudForms release also beefs up its management of Amazon AWS. For example, it now handles Amazon Virtual Private Cloud (VPC), a logically isolated section of the AWS Cloud containing user-defined virtual networks. Users can select authorized VPC/subnets to provision and run their workloads, which are managed by CloudForm's policies. This is important because a lot of organizations are viewing VPCs as the way to get the degree of workload isolation that gives them the comfort level top use public clouds. 

CloudForms 3.0 can also provision Amazon Machine Instances (AMIs) in a policy controlled manner, through enterprise defined self-service portals and service catalogs. State policies are enforced on source AMI’s as well as instance placement into destination regions and availability zones. Identity and group affiliations also condition what can or cannot be provisioned per enterprise regulatory, business and IT rules.

Stepping back from the details of today's announcement though, I think there's a bigger point to be made. Namely, that successful cloud deployments aren't going to be about buying a singular product and installing it. Rather, they're about taking a broad view of IT that considers the broader goals of the business and how the role of IT will morph within that context. As Forrester analyst James Staten wrote on his blog earlier this year: "Once you have a handle on the current state of your hybrid cloud environment, you should then shift your focus to managing and maintaining this hybrid architecture going forward - because you are just going to become more hybrid from here. This calls for a good enterprise architectural approach. "

That's good advice and you're going to need hybrid cloud management to get there.

Monday, November 04, 2013

Friday, November 01, 2013

Why it's not about containers "winning" or "losing"

About a month ago, I cranked out a long blog post about containers and how they came about. In a nutshell, containers virtualize an OS; the applications running in each container believe that they have full, unshared access to their very own copy of that OS. This is analogous to what virtual machines do when they virtualize at a lower level, the hardware. In the case of containers, it’s the OS that does the virtualization and maintains the illusion.

I dusted off some old research notes I had written and was writing about containers again because they have become something of a hot topic, especially in the context of platform-as-a-service (PaaS). This hotness in turn has led to, alas, predictable stories and claims that containers will kill the currently dominant approach to workload separation, namely, hardware-based virtualization like Linux' KVM or VMware's ESX.

That won't happen. Though containers are certainly an important technology. Let me take you through, in the form of a Q&A, why I believe both those statements to be true. (The original post and the research notes it links to give a lot more background, which I won't repeat here.)

What advantages do containers have over virtual machines in particular?

Probably the biggest advantage for most purposes is density. With hardware-based virtualization (henceforth just "virtualization"), each guest instance runs a full operating system copy. With OS virtualization (aka "containers"), there's just one operating system instance for the whole physical server; only a modest subset of the host OS is duplicated for the individual containers. 

Furthermore, containers have a single kernel with direct visibility into all the workloads running on the entire server. This eliminates some overhead and indirection associated with the fact that the contents of the guests in a virtualized server are a level of abstraction away from the controlling hypervisor (which is itself an operating system). Suffice it to say that, all things being equal, containers can provide significantly higher guest densities on a given piece of hardware. 

Starting up new instances and resizing instances is much faster as well. Again, with containers, you're not starting up or reconfiguring an entire operating system copy; you're effectively just fiddling with resource groups within an OS.

So, if containers are so gosh darn great, why aren't they everywhere?

Because they didn't do as good a job of addressing enterprise problems c. 2001 as the virtualization alternative did. (They did and continue to do a very good job of solving hosting provider problems which is one of the reasons that containers are so widely used in that environment.

It is indeed true that everyone, including enterprises, was desperate to "do more with less" as the cliche goes in the post-dot-com hangover. So any product that would allow multiple workloads to run on a single physical server was very welcome. (Windows workloads such as Exchange were notorious for demanding that they have an entire server to themselves even when they used just a small fraction of the capacity. 15 percent utilization was considered good pre-virtualization.)

But, with enterprises in particular, these weren't cookie-cutter workloads. Multiple OSs, lots of different versions, lots of different libraries and other customizations. Not every workload was a unique little flower. But lots of them were. And indeed, some early virtualization deployments were as much about supporting multiple operating system versions on a single machine as they were about consolidation as such.

As a result, virtualization was a better match for enterprise workloads than containers. And the density thing wasn't that big a deal. Getting from 15 percent to, say, 50 percent utilization looked like a big win to most enterprises and getting to the next level of optimization wasn't nearly the priority it was for service providers. (And enhancements to both software and the introduction of hardware assists by the chip makers incrementally improved performance in any case.)

Why didn't enterprises just adopt both containers and virtualization?

This gets into speculation but organizations have a limited bandwidth to adopt new technologies. And virtualization was a big consumer of that bandwidth throughout much of the 2000s (or naughts or whatever we're calling that decade). Indeed, it remains a big consumer of enterprise IT bandwidth today. And, for the reasons stated above, containers just didn't offer that much of an incremental win for enterprises.

Virtualization also added a lot of services of interest to enterprises that made use of the hypervisor (live migration, disaster recovery, storage snapshotting, and so forth) that are arguably a less natural fit for container-based approaches.

So what changed? Why are we talking about containers again?

The cloud changed. By which I mean two things in particular.

The first is that, as I wrote in my previous piece, cloud-style workloads tend towards scale-out, stateless, and loosely coupled. They also tend to run on more homogeneous environments (alongside existing applications under hybrid cloud management) and use languages (Java, Python, Ruby, etc.) that are largely abstracted from the underlying operating system. You typically don't have or want to have a highly disparate set of underlying OS images because that makes management harder. You also tend to have a large number of smaller and shorter-lived application instances. These are all a good match for containers.

The second is that PaaS amplifies all this in that it explicitly  abstracts away the underlying infrastructure and enables the rapid creation and deployment of applications with auto-scaling. This is a great match for containers, both because of the high densities and the rapid resource reallocation they enable (and, indeed, require). 

In short, it's not so much the case that the containers concept has changed but, rather, that the nature of workloads is changing to be better aligned with what containers do well.( There's also been significant open source innovation in the containers space as there has also been with KVM and oVirt in virtualization.)

So why doesn't virtualization go away?

The flip answer is "because nothing goes away." 

The less flip answer is "horses for courses." In principle, containers could be adapted to do (almost) anything subject to running on a common kernel. But I suspect that, if you have a requirement for Infrastructure-as-a-Service with a fair bit of heterogeneity in the guests, virtualization will continue to be just a more natural fit. The exact manner in which IaaS and PaaS evolve and converge over time will certainly change the details of how we do multi-tenancy.  History suggests strongly though that we'll continue to manage workloads in multiple ways. 

Everything old is new again

Oftentimes it seems that there are few genuinely new concepts in the IT biz. Rather, the crank turns again and this time the opportunities and the pitfalls are a bit better understood. Or there's a better match between the technology and current market needs. Or the necessary user and development ecosystem come together better. I'd argue all of these are true with containers.

Controlling Clouds: Beyond Safety presentation

I gave this presentation at CloudOpen Europe in Edinburgh earlier this month. Here's the abstract:

As an industry, we’ve mostly moved on from naive notions about cloud computing being inherently “safe” or “risky.” However, more sophisticated discussions require both greater nuance and greater rigor. This presentation takes attendees through frameworks for evaluating and mitigating potential issues in hybrid cloud environments, discusses key risk factors to consider, and describes some of the relevant standards and provider certifications. This is a broad and sometimes complex topic. However, it’s very manageable if individual risk factors are considered systematically and specifically. This session will give IT professionals tools and knowledge to help them make informed decisions.

I intend to record a variation of this presentation at some point, but I wanted to make this available in the meantime.

Running Google+ Hangouts: Tips and Tricks

Goggle+ Hangouts seem to be a "thing" these days. I'm a bit ambivalent about them myself. I confess to finding low to middling quality video of talking heads tends to detract from content rather than adding to it. That said, a lot of people seem to really like video and Hangouts are a really easy way to create some. Therefore, on the premise that anything worth doing is worth doing right, I thought I'd share some tips and tricks to doing Hangouts right. 

Where noted, some of these tips are lifted from a post by my Red Hat colleague Rich Bowen. Others are based on my own experiences.

Your studio

I use the "studio" term somewhat ironically because most people participating in a Hangout won't have a studio, nor is having one expected by the audience. But a little care will go a long way in the quality of the final product. 

Use a microphone and headphones. The mic doesn't need to be anything especially fancy, but there is a world of difference between using your laptop's internal microphone and just about any external microphone including even something like a gamer's headset. 

If you get into doing Hangouts seriously, you may consider getting a good-quality external webcam like the Logitech C920 although this definitely falls into the "optimizations" bucket as opposed to something that's really essential.

Shut the door. Avoid barking dogs and children.

Take some care with the lighting. Again, it's not essential that you setup a bunch of fancy studio lights. But avoid being backlit against a window or otherwise being especially poorly lit.

Similarly, you don't need to get all fancy with the background but people really don't want to see your unmade bed. In some cases, it may make sense to get an appropriate backdrop printed in order to provide a consistent visual identity for your Hangouts. (As with other social media use, however, be careful not to do something that violates your company's guidelines about using their brand.)

Hangout limits (via Rich)

Hangouts have a 10 attendee limit. We found that out the hard way the first time. So your audience doesn't actually join the hangout. Instead, they watch a live stream on Youtube, where there's no limit. Don't invite people to the hangout unless they are actually going to be talking. 

You can schedule a Hangout, but you cannot schedule an On Air Hangout. Don't get confused and schedule a Hangout and then have to change it later on. That'll just lose you half your audience.

[I'll just add that, in general, I find the whole UX associated with Hangouts to be incredibly confusing and obtuse. It will probably take you a while to get the hang of it.]

Mute the mic (via Rich)

Camera switches on voice - that is, the person talking gets the camera. So if you're not presenting, mute your microphone so that it doesn't suddenly switch to a shot of you cleaning your ears when you happen to rustle a paper. 

[You can also "mute" the video, which may or may not make sense depending upon the format of the Hangout. It probably doesn't make sense if it's an interactive discussion. But, if it's in the form of sequential presentations, your video is probably just a distraction if you're not speaking and, furthermore, you don't really want to spend the entire preso studiously looking at the webcam and trying to appear interested.]

Test, test, test (via Rich)

Presenting in a Google Hangout requires plugins, and the controls aren't immediately obvious. Create a test On Air event and have the presenter join and share their slides. It can be weird talking to your screen for an hour without seeing the audience, and it can be doubly flustering if things don't work quite how you thought they would. Give them the option of doing a full runthrough if they really want to, but at a minumum ensure that they can join the test hangout and get their slides shared, and that their microphone works well, there isn't a terrible echo in the room they're in, and so on.

YouTube stream URL and more UX confusion (via Rich)

The event that you've created in Google Plus, and the actual "On Air" event, are not connected to one another in any way. You need to create an event, and then put the URL of the On Air event in the description. Unfortunately, you can't schedule the On Air event ahead of time. So you create the On Air event (i.e., YouTube stream) and then edit the description 20-30 minutes before you're scheduled to start.

Start the On Air event well in advance and share a window with details about the hangout and when it will start. You can edit that bit off when you're done, so that it's not in the final video.

What I'm going to do next time is have placeholder page for the weeks leading up to the event. Then, the day of the event, create the On Air event and redirect that URL to the event. Of course, to do that, you'll need to have a server where you can redirect URLs, or a URL shortener that you can update, or something like that. Failing that, just update the Google Event when you have the YouTube URL, or perhaps publicize a blog post or forum post which you then update with that URL.

Lower Third Tool

Especially for panel-style discussions, the Lower Third Tool is a handy way to clearly identify the participants.

Q&A (via Rich)

Although the Hangout has a Q&A tool, that's only for people that are in the hangout. I'm still looking for a good Q&A tool, but for now we're using IRC. I recommend that you create a new IRC channel, rather than using your existing one. This removes the off-topic chatter, and makes it easier to have a transcript and an attendee list.

For the many people who are not comfortable with IRC, give them a web link, like https://webchat.freenode.net/?channels=rdohangout which takes them directly to the web chat applet [in this case, for our RDO--upstream Red Hat Enterprise Linux OpenStack Platform--hangouts].

Editing the video (via Rich)

Once the event is over, Youtube has a simple edit tool that will let you trim off the 30 minutes of dead air at the beginning of the recording. This takes a LONG time, so make sure it's done before you start publicizing the URL of the video. 

Thursday, October 31, 2013

Links for 10-31-2013

Friday, October 18, 2013

Links for 10-18-2013