- Thanks, NSA, you're killing the cloud | Cloud Computing - InfoWorld - "Personally, I don't see much of a connection between the NSA and cloud computing, but those on the fence regarding cloud computing will cite this as another reason to kick the can further down the road. Thanks for nothing, NSA."
- Prepare for change! This is not your father’s database industry — Tech News and Analysis - This this overstates Hadoop vis-a-vis other options but still a good read.
- 50 Top Sources Of Free eLearning Courses
- Google Maps Help Find City Quizzes, Dylan Trivia, Garage Sales - Games with Google Maps
- 18 Useful Internet Marketing Statistics that You Can't Ignore - RT @samanthastone: This just in...nurtured leads make 47% larger purchases than non-nurtured leads. That stat and more can be found at
- Photos: MIT Class of 2017 Hacks Harvard Class of 2017 Website | BostInno - MIT’s incoming class hacked the Harvard 2017 website Sunday night by replacing all the pictures with Mitt Romney’s mug, according to a Reddit thread.
Thursday, June 27, 2013
Links for 06-27-2013
IT risk isn't just about security
Because it mirrors a point I often feel compelled to make when discussing security and cloud computing, I wanted to highlight a couple of paragraphs from Cloud Computing: Assessing the Risks by Jared Carstensen, Bernard Golden and J.P. Morgenthal as excerpted in Tech Target.
It's important to understand that this risk limitation [whereby service providers shift the primary responsibility for risk to consumers] is not unique to Cloud Computing. Outsource providers (e.g. firms that take over operating a company's IT data centre) also limit their financial responsibility in the event of an outage. Therefore, it is important not to regard this risk limitation as a complete restriction on using a Cloud provider, unless, that is, a company regards any risk limitation by a service provider as unacceptable. In that case, the company should continue to operate its own computing environment and forego use of an external Cloud provider.
The important point from this discussion is that when Cloud Computing security is raised as an issue, other issues are often being addressed. It's important to distinguish what type of issue is of concern, as that will change the method of evaluating the issue, the demarcation of the trust boundary and the appropriate actions to be taken by the Cloud user.
One of the reasons why I think this point is important is that discussing overall IT governance discussions solely in terms of security (whether we're talking public clouds, private clouds, or--increasingly--some manner of hybrid IT) is far too narrow a framing. This narrow framing, in turn, often leads to thinking about the issue in narrow technical terms such as multi-tenant security features, encryption and key management, and physical facility protection.
These are important matters certainly. But they're also matters that public cloud providers (like other types of outsources) can reasonably argue they have well in-hand using well-established procedures and processes. The more difficult answers about where workloads should run come down to broader questions--and those answers may well change over time.
(I covered some of these broader issues in a presentation at the Red Hat Summit in June. I'm hoping to get a version of that presentation Beyond Safety: Controlling Clouds posted over the next month or so.)
Wednesday, June 26, 2013
Links for 06-26-2013
- Has the Time Arrived for Cloud Insurance? - Leverhawk
- Coursera Fantasy: Teacher Authority and Student Initiative in a MOOC
- Coursera Fantasy: Peer Feedback: The Good, the Bad and the Ugly :-)
- udacity and coursera – python editor, peer reviewing and complaining | Down Home Country Coding With Scott Selikoff and Jeanne Boyarsky
- The Problems with Peer Grading in Coursera | Inside Higher Ed
- How the Design of Soda Cans Have Changed Over Time
- 66 job interview questions for data scientists - Data Science Central
- OpenShift Origin Community Day (Boston) Deep Dive into Cartridges with Jhon Honce - YouTube - RT @pythondj: Just posted: @OpenShift Origin Community Day (Boston) "Deep Dive into Cartridges" by Jhon Honce via @redhat
- What were the most ridiculous startup ideas that eventually became successful? - Quora - RT @tsimonite: Anti-pitches for tech startups: "Google: World's 20th search engine"; "iphone: It won't have cut and paste"
- From Red Shoes To Red Hat | Rishidot Research - RT @krishnan: I am joining Red Hat in July as Director, OpenShift Strategy
Monday, June 24, 2013
Links for 06-24-2013
- David Simon | We are shocked, shocked…
- Aho/Ullman Foundations of Computer Science
- Heirs of Infocom: Where interactive fiction authors and games stand today | Ars Technica
- You Don't Have To Like Edward Snowden - “@dbfarber: You Don’t Have To Like Edward Snowden via @buzzfeedpol” << Smart piece.
- Cloud Platform Blog: Enabling Google App Engine to run in the Private Cloud with CapeDwarf - RT @DanJuengst: Code on! Google and Red Hat collaborating to help folks run GAE apps on #OpenShift PaaS and JBoss. #RedHat #gae
- Photo by ghaff • Instagram - Bloody Mary time at SFO United Club.
- ThinkGeek :: Tac Bac - Tactical Canned Bacon - RT @mccrory: Greatest invention EVER! Tactical #Bacon !!! Yes, that's not a mistake, TACTICAL BACON!
- PKN Gibsons #6 Mueller - YouTube - RT @pythondj: Sweet! my @discovertotems geospatial app presentation from @pechakucha Is up on YouTube hosted @openshift
- Photo by pythondj • Instagram - RT @pythondj: Rockin @Mongodb in the @openshift booth @ #mongonyc today with fotois @ New York Marriott Marquis
- PayPal now builds products on its internal PaaS — Tech News and Analysis - RT @krishnan: PayPal now builds products on its internal PaaS by @gigaom <- Anyone still thinking PaaS is for hobbyists? :-)
- Twitter / adrianco: Just shared a photo #throughglass ... - RT @adrianco: Just shared a photo #throughglass
- John McAfee releases NSFW video on how to uninstall security code • The Register - RT @valleyhack: This McAfee dude is turning drug use into performance art
- NVD3.js :: re-usable charts for d3.js
- Vega: A Visualization Grammar
- Rickshaw: A JavaScript toolkit for creating interactive time-series graphs
- REST web services with Python, MongoDB, and Spatial data in the Cloud - Part 2 | OpenShift by Red Hat
- Generalized Set of Color Schemes
Friday, June 21, 2013
The GigaOm Structure 2013 zeitgeist
I like GigaOm Structure. I find it gives a good sense of the current zeitgeist in cloud computing and related areas. What's being talked about and what isn't? What new or reimagined techs are emerging as memes?
GigaOm's own writers (among others) covered the event in considerable depth and I won't attempt to recreate their reporting here. Rather, I wanted to hit on some general themes I noticed and a few points that particularly struck me. So with no further and in no particular order, here we go.
OpenStack was omnipresent. Other on-premise IaaS? Not so much.
Mind you. There was still a bit of commentary about how OpenStack was still in relatively early days. Maybe. Ryan Granard of PayPal told the audience that his company runs 20 percent of its production infrastructure on OpenStack. As readers probably know, there was much ado about PayPal's adoption of OpenStack--they were and are heavy VMware users--a while back. One of Grandard's points though was that PayPal has a strategy of deliberately making several bets as a way of getting velocity while still having a robust infrastructure.
Notably absent from this Structure was the once ubiquitous Marten Mickos of Eucalyptus. Nor did I hear much mention of CloudStack--though Citrix did have a sponsor workshop, which I didn't attend.
Speaking of PayPal. A nice endorsement for PaaS and OpenShift.
Granard also articulated, as well as I've heard it from anyone, why PaaS is such a big deal for organizations. As reported by GigaOm's Jordan Novet:
Companies big and little have been jumping aboard the concept of on-premise PaaS, to some degree because security, regulatory compliance and cloud vendor lock-in fears remain part of the conversation about running on public infrastructure.
How is PayPal going about this? It’s been running Red Hat’s OpenShift on-premise PaaS to build out products such as PayPal Here — the company’s answer to Square — as well as a developer sandbox.
With that tool, Granard said, a developer chooses a product to work on “and in minutes, we have you up and running in a fully connected container” with infrastructure resources immediately allocated.
The real money quote for me though was that PaaS lets PayPal "enable developers and get out of the way."
x86 vs. ARM. Come back next year.
By which I mean that there were a few threads on this topic. Especially in the vein of whether ARM will make a meaningful dent in the server world. But no clear resolution.
To the degree that there was something of a consensus among folks I spoke with, it largely parallels my opinion and goes something like the following: x86 is the clear incumbent on the server. ARM is the clear incumbent in new-style mobile (tablets, cell phones, etc.). There's considerable inertia to that default condition for reasons of ecosystem and other things. For either architecture to make a major dent (narrow use cases aside) outside of its home base will require it to develop a 10x advantage--which most people don't think is going to happen.
What's that SDN stuff anyway?
There was some skepticism. For example, Arne Josefsberg CTO of ServiceNow said that “The conversations today sound almost exactly like the conversations we had three to four years ago." Indeed, the "hot or not" panel he sat on declared SDN a loser technology.
But that seemed to be a minority opinion. Session after session returned to the idea that the networking component of infrastructures need the same sort of rewiring in software that you can do with compute and storage if the whole dynamic IT process is going to be realized. While I think it's fair to observe that there were still a lot of open questions about how we're going to get there (and what exactly it will look like), the consensus was squarely behind SDN--at least as a concept.
Private, private, hybrid
This topic really deserves a separate post, especially given that I was on a panel about private clouds at the ODCA Forecast event preceding structure. Suffice it to say that it's a complicated topic for a variety of reasons:
- Depending upon the specific requirements, there are strong economic reasons to choose private over public or vice versa.
- Different organizations have strong pre-dispositions for in-sourcing vs. out-sourcing
- Existing applications can't be ignored
- Regulation is a factor that may or may not be "fixed" (from the perspective of public clouds.
The bottom line is that there are plenty of arguments and cherry-picked examples available to bolster "your" side. That said, there was widespread agreement that, for at least the next n years (where n is much less agreed upon), the cloud world will be hybrid.
Podcast: Geo, mobile, and more on OpenShift
Here are links to some of the posts and other topics covered in the podcast:
OpenShift
Spatial MongoDB in OpenShift, be the next FourSquare - Part 1 | OpenShift by Red Hat
REST web services with Python, MongoDB, and Spatial data in the Cloud - Part 2 | OpenShift
Kinvey
Appcelerator
Adopt-a-Hydrant
Shameless plug: I'll be on a panel with GigaOm analysts talking about PaaS Tuesday, June 25: Red Hat-Flexibility with PaaS: How to Keep Your Options Open — GigaOM Pro.
Listen to MP3 (0:17:57)
Listen to OGG (0:17:57)
Monday, June 17, 2013
Links for 06-17-2013
- Layman's Introduction to Random Forests - Edwin Chen's Blog
- Red Hat Summit: Open source trends, cloud outlook, innovation and more
- High-Performance Blending: Nutrition
- Lawsuit Filed To Prove Happy Birthday Is In The Public Domain; Demands Warner Pay Back Millions Of License Fees | Techdirt - RT @webmink: Lawsuit Filed To Prove Happy Birthday Is In The Public Domain; Demands Warner Pay Back Millions Of License Fees:
- Spatial MongoDB in OpenShift, be the next FourSquare - Part 1 | OpenShift by Red Hat
- Using Node.JS, MongoDB, Express for your spatial Web Service - and it's Free! | OpenShift by Red Hat
- Twitter / fbijlsma: @ghaff a lot of demand for ... - “@fbijlsma: @ghaff a lot of demand for Gordons book ” << We have a few copies left at #redhat booth
- Red Hat opens OpenShift PaaS cloud for business | ZDNet - RT @sjvn: Red Hat opens OpenShift PaaS cloud for business #RedHat #Linux #cloud by @sjvn
- Red Hat | Red Hat Unveils Fully-Supported Public PaaS Offering, OpenShift Online - OpenShift Online PaaS now generally available w paid support. (Free tier still available too.)
- OpenShift Origin Community Day (Boston) Writing Cartridges V2 by Jh... - RT @pythondj: Slides are up from @OpenShift Origin Community Day (Boston) Check Out: Writing Cartridges @slideshare #cloud #paas
Podcast: The past, present, and future of Linux systems management
Listen to MP3 (0:24:09)
Listen to OGG (0:24:09)
[Transcript]
Wednesday, June 05, 2013
Stop data silos
Apropos of the post I put up earlier today, this sponsored post on GigaOm notes that:
Data silos and integration challenges, more than security, are the biggest barriers to cloud adoption. Siloed data in cloud apps and data centers is costing companies millions annually due to inconsistency, inaccuracy and inefficiency across the business. And the enterprise software market is crossing the threshold of another transformation, now that cloud computing has shifted the center of gravity for data.
The MIT Media Lab's Sandy Pentland noted something similar at the MIT Sloan CIO Symposium a few weeks back when he said that "About 20% of big data is getting the data out of the silos and transforming it."
That's why at Red Hat we're so interested in the idea of open hybrid cloud--which includes open hybrid storage. (Starting with Red Hat Storage based on GlusterFS.) This isn't to minimize the important, and really difficult, role of organizational change in breaking down the silos. But there's at least technology available to help break down silos rather than aid in their creation.
Links for 06-05-2013
- As Data Floods In, Massive Open Online Courses Evolve | MIT Technology Review - "It’s unclear whether the laundry lists of refinements that result from A/B testing will add up to a grand theory of learning and teaching that challenges tradition. Ng says he doesn’t think a grand theory is needed for MOOCs to succeed. “I read Piaget and Montessori, and they both seem compelling, but educators generally have no way to choose what really works,” he says. “Today, education is an anecdotal science, but I think we can turn education into a data-driven science, where you do what you know works.”"
- A comparison of approaches to large-scale data analysis
- New Government Documents Show the Sean Parker Wedding Is the Perfect Parable for Silicon Valley Excess - Alexis C. Madrigal - The Atlantic - Just... wow.
- Refactoring Coursera | Mike Caulfield
- The Ed Techie: You can stop worrying about MOOCs now
- MOOC as Courseware: Coursera's Big Announcement in Context |e-Literate
- Coursera Jumps the Shark | HESA
- Business as usual | Music for Deckchairs
- Economic Effects of State Bans on Direct Manufacturer Sales to Car Buyers
- Zipf, Power-law, Pareto - a ranking tutorial
- Your favorite tech companies/products get sassy new slogans courtesy of Reddit thread
- Why Big Data Is Not Truth - NYTimes.com
- The Lorenz Cipher and how Bletchley Park broke it
Why we should probably retire "Big Data"
Data is hugely important. Long has been of course. It's often argued that applications are the longest-lived IT asset. Arguably the data that those long-lived applications create and access sticks around for at least as long. Data is only going to become more important.
Is there a lot of hype around data today? Sure. Raw data isn't information. And information doesn't necessarily lead to useful action if organizations (or their customers and users) aren't willing to change behavior based on new information. Just because pricing mechanisms can be used to reduce traffic congestion based on sensor data--just one of the ideas discussed under the "Smart Cities" term--doesn't mean that even basic congestion pricing is necessarily politically viable.
Furthermore, I strongly suspect that the lots of data hammer will turn out to be a rather unsuitable tool for certain (perhaps many) classes of problems--the enthusiasms of the "End of Theory" crowd notwithstanding. For example, it's unclear to what degree more data will really help companies design better products or target their ads better. (To be clear, there's long been lots of market research in the consumer space. It's just that it hasn't been terribly effective in creating the killer ad or the killer product.)
But it would be a mistake to think that today's data trends are 1990s-style Data Warehousing with just a fresh coat of paint. Whatever the jokes about the local weatherman, weather forecasting has improved. Self-driving cars will happen, though they may take a while to come into mainstream use. DNA sequencing is now commonplace--although, in a common theme, we're still in the early days of figuring out what we can (and should) do with the information obtained. And we're well on the way to sensors of all sorts becoming pervasive.
Which makes the "Big Data" term somewhat unfortunate in my view. I realize that may seem a bit of a contradiction given what I wrote above. Let me explain.
My first problem is that "Big Data" is too narrow. This is true even if we use the term in the broader sense of data that is atypical in some respect--not necessarily in its volume. The four Vs is a common shorthand. (For a less precise, but possibly more accurate description, I like the "Difficult Data" term that I heard from the University of Washington's Bill Howe.)
But an emerging culture of data doesn't have to be about big or even difficult. Discussions about data at the MIT Sloan CIO Symposium last month included big data examples, but it was also, in no small part, about cultures of data and the breaking down of silos. Just as with IT broadly and cloud computing, data and storage have to increasingly be based on a hybrid model in which data can be accessed when and where it is needed and not locked up in one department or even one organization. Governments and others are increasingly making open data available as well.
It's also worth remembering Nate Silver made headlines for calling the last US presidential election correctly, not because he did big data stuff or even because he applied particularly innovative or sophisticated analysis to polling data, but mostly because he used data and not his gut.
The second issue I have with "Big Data" isn't really that term's fault. Rather, it's that "Big Data," today, is so frequently conflated with Hadoop.
Based on Google's MapReduce concept, Hadoop divides data into many small chunks, each of which may be executed or re-executed on any node in a cluster of servers. Per Wkipedia: "A MapReduce program comprises a Map() procedure that performs filtering and sorting (such as sorting students by first name into queues, one queue for each name) and a Reduce() procedure that performs a summary operation (such as counting the number of students in each queue, yielding name frequencies)."
Hadoop also provides a distributed file system that stores data on the compute nodes, providing very high aggregate bandwidth across the cluster. (The standard file system is HDFS, but other filesystems, such as Gluster, can be substituted for higher scalability or other desirable characteristics.
Hadoop is often a useful tool. If you can split up data sets and work on them with some degree of autonomy, Hadoop workloads can scale very large. It also allows data to be operated in-situ without being loaded and transformed into a database, which can greatly decrease overhead for certain types of jobs. (This presentation by Sam Madden at MIT CSAIL offers some benchmarks as well as some pros and cons of Hadoop relative to RDBMS systems.)
However, data can be processed and analyzed using a wide variety of tools, including NoSQL databases of various kinds, "NewSQL" databases, and even traditional RDBMs like PostgreSQL (which can still scale sufficiently to handle a great many types of data analysis and transformation tasks). In fact, we even see something of a trend with some of the new-style databases adding back in traditional RDBMS features that had been assumed to be unnecessary.
Even high volume data doesn't begin and end with Hadoop. As Dan Woods writes for CITO Research: "The Obama campaign did have Hadoop running in the background, doing the noble work of aggregating huge amounts of data, but the biggest win came from good old SQL on a Vertica data warehouse and from providing access to data to dozens of analytics staffers who could follow their own curiosity and distill and analyze data as they needed."
Hadoop is an important tool in the kit when the amount of data is large. But there are lots of other options for that kit bad too. And never forget that it's not just about the bigness of the data but whether you share it with those who need it and whether you do anything with it.
Monday, June 03, 2013
Links for 06-03-2013
- A valley of ashes by Emily Esfahani Smith - The New Criterion - "In The Living Moment: Modernism in a Broken World, the literary critic Jeffrey Hart traces the efforts of a small but influential group of poets and novelists who sought to create a new cultural order following the chaotic aftermath of World War I. Their efforts came together in a new movement whose legacy is still with us today—literary modernism. The cultural fallout of the war—its devastation—was immense. The traditional order of nineteenth-century Europe had been blown to bits. “The First World War inaugurated the manufacture of mass death that the Second brought to a pitiless consummation,” in the words of the historian John Keegan."
- (3) Fraser Cain - Google+ - Tips and Tricks for Hangouts on Air updated Sept. 13, 2012 …
- The Flawed Ambitions of Better Place | MIT Technology Review - “@mlamonica: Flawed ambitions - how EV battery-swapping startup Better Place misjudged consumers and automakers #EVs”
Friday, May 31, 2013
Links for 05-31-2013
- How Land Rover, England's Ugliest Station Wagon, Became One of the World's Most Luxurious Brands | Adweek
- If We Still Used Punch Cards - YouTube - RT @LivingComputers: In this video @GHaff uses concise visuals to demonstrate punched card obsolescence:
- Twitter / ghaff: Hmm. Bing's running this Bing ... - Hmm. Bing's running this Bing vs. Google web search contest. Were smoked in my (blind) trial.
- The 5th Tenet of Open Hybrid Cloud: Start With an IaaS Private Cloud | tentenet.net - An under-the-covers look at #redhat technologies for hybrid IaaS by @bryanwche
- Mary Meeker’s 2013 internet trends: all the slides plus highlights - Quartz
- Doc Searls Weblog · Let’s help Airbnb rebuild the bridge it just burned - Not to defend Airbnb here, but I suspect we're going to see a lot of stress points as sharing services try to move out of the margins to something more mainstream.
- Forecast 2013 Registration - RT @opendatacenter: Enterprise tends to turn to private #cloud, but will it stay that way? @ghaff will discuss on a panel at #Forecast13:
- Why Do Not Track is destined to fail (DNT is DOA) | getwired.com - "Rather than driving efforts like DNT, which fundamentally cannot occur (in the manner users think those words mean “do not track”), we’d do a lot better as an industry to drive standards that delineate what types of information a specific site or tracking engine like Google Analytics or Adobe’s Omniture products can collect on you. But even if you throw those back at users, they’ll be overwhelmed. Perhaps the best angle is to reinforce that no activity on the Internet is totally anonymous, and no matter how hard you try, you cannot ever completely prevent being tracked."
- On the Internet, Nobody Knows You're a Dog
- NetAppVoice: Securing The Cloud: Why You Need Cast-Iron Guarantees - Forbes
- Data Science eBook by Analyticbridge - 2nd Edition - Data Science Central
- Memo to this year’s YC class: It’s damn hard to build an enterprise company | PandoDaily - RT @utollwi: Memo to this year’s YC class: It’s damn hard to build an enterprise company #SaaS
- http://maps.bpl.org's photos. - RT @flickrock: @ghaff Check out that Flick gallery using Flickrock!
- Flickr: Norman B. Leventhal Map Center at the BPL's Photostream - Cool map resource on Flickr from the Boston Public Library
- Google IO — Benedict Evans - "My main impression of Google IO was not so much any specific announcement as the overwhelming sense of ambition and self-confidence."
Thursday, May 23, 2013
Data in, hunches out
"Going with your gut is out." That line, from Russell Reynolds Associates managing director Shawn Banerji, neatly summed up a big chunk of yesterday's MIT Sloan CIO Symposium. There was discussion of the computing side of things as well. I especially liked EMC CTO John Roese's description of public clouds evolving to a "chaotic" (as in heterogeneous/hybrid) pool of resources in which special purpose clouds would have specific functions. But the majority of the day revolved around data.
Not necessarily "big data" by the way. One panelist—Jack Norris of MapR—even remarked that the "big data term is probably short lived." But, rather, the pervasive use of data, in whatever form and at whatever speed, to drive decision-making. As Annabelle Bexiga, the CIO of financial services firm TIAA-CREF put it: "Big data is just a richer set of data. [It's a] natural evolution of where data is going."
Erik Brynjolfsson (above) is the Director of the MIT Center for Digital Business. He spoke to how the data explosion was a revolution of technology but was also (and required) a revolution in management.
He offered the example of publishing, described as historically a "culture of lunches"—which is to say a culture of hunches and people networks. But Amazon brought a culture of numbers to the industry. And things haven't been the same since.
A 2011 paper, which Brynjolfsson co-authored with Lorin Hitt and Heekyung Hellen Kim found that data-driven decision making at firms resulted in 5 percent higher productivity than at firms which weren't so data oriented.
Brynjolfsson also discussed what might be called ambient data, data collected almost incidentally from sources such as cell phone records. He offered the example of streetbump in Boston, which uses an iPhone app to find potholes. At the same time, he observed that the app did best at finding potholes in the Back Bay, Beacon Hill, and other relatively upscale Boston locales. Why? Because that's where iPhones predominate. As Sloan's Andrew McAfee would elaborate on in his closing keynote (paraphrasing science fiction author William Gibson), "the future is already here. It's just unevenly distributed."
The panel which Brynjolfsson moderated also touched on some of the privacy issues associated with this ambient data. The MIT Media Lab's Sandy Pentland told how a big data commons, created by French telco Orange, was used to reduce commuting times in the Ivory Coast by 10 percent by rearranging bus routes using location-based data from mobile phones. Pentland went on to note, with more of a bit of understatement, that this sort of thing is politically controversial in places like the UK.
For his closing keynote, Andrew McAfee (above) took as his jumping off point a fascinating graph that appears in Ian Morris' Why the West Rules—For Now. At the risk of offering up spoilers, the central thesis of the book is that, viewed from the perspective of today, the level of worldwide social development prior to the industrial revolution is effectively in the noise.
(Although McAfee didn't get into this, the answer to the book title's question is basically that the West was better positioned to create and take advantage of the industrial revolution when the factors making it possible came together. It's a great read. If nothing else, it's a good history of the world from the perspective of the Western and Eastern core.)
The industrial revolution had such an impact because it overcame the limitations of human muscles. (I suppose, given farm animals, it would be more accurate to say it overcame the limits of human and domesticated mammal muscles generally.) In any case, though, McAfee's thesis that that today we're starting to overcome the limitations of our individual minds.
He laid out four elements to this:
Cyborgs—as in new combinations of people and machines.
Open—which will define successful organizations in a variety of ways. (If you want to delve more deeply into this thought, I point you to a presentation I gave at ProductCamp Boston a few weeks back.)
Data-driven—because for the first time we have data-driven visibility in all sectors of the economy.
Evolving—for which McAfee offered the example of the car rental industry which evolved only incrementally since it was founded after the Second World War but has seen the introduction of radically new services made possible by the Web and mobile phones from Zipcar to Lyft.
So it's more than data. But data—along with the compute needed to operate on it and the networks needed to move things around and tie them together—is a common thread. Big challenges lie in gaining access to the right data, even within single organizations. Cutting across data silos was also a theme heard more than once throughout the day. Asking the right questions matters too. As McKinsey's Michael Chui summed up that thought: "Be data driven. But don't suck at it."
Tuesday, May 21, 2013
Links for 05-21-2013
- Photographer Shares His Lightning Quick Lightroom Workflow
- Python 3 Metaprogramming - YouTube
- Has Intel finally landed that elusive Atom deal? | ITworld - RT @apatrizio: My debut blog with ITworld on the chips biz: rumors of an #Atom-powered #Samsung #Galaxy tab.
- New Orleans Buck
- (403) http://blogs.forrester.com/james_staten/13-05-16-hybrid_cloud_future_too_late - RT @Staten7: Hybrid Cloud? You are already hybrid. The question: What are you doing about it? #Forr blog: @stefanried
- Twitter / krishnan: Best definition of PaaS I have ... - RT @sravish: HILARIOUS (and perhaps true?) -> “@krishnan: Best definition of PaaS I have come across - ”
- Making Big Data Technologies Work in the Enterprise - “A lot of the architectures and products that technology managers may have been accustomed to for traditional transactional activity don’t map well to a big-data world,” says Gordon Haff, corporate technology evangelist at Red Hat and the author of Computing Next, a book on cloud computing. “You very much need to think about an architecture in the context of big data.”
- Ten years on: How did that cloud strategy pan out? • The Register
- www.martinlamonica.com/wp-content/uploads/2013/05/Making-Big-Data-Technologies-Work-in-the-Enterprise.pdf - Nice piece on big data techs for enterprise by @mlamonica (w quotes from me)
- Netflix, Reed Hastings Survive Missteps to Join Silicon Valley's Elite - Businessweek
- 'Weeds' Creator Kohan Dishes on How Netflix Drives Hollywood Insane - Businessweek
- The Serious Superficiality of The Great Gatsby : The New Yorker - "Baz Luhrmann’s “The Great Gatsby” is lurid, shallow, glamorous, trashy, tasteless, seductive, sentimental, aloof, and artificial. It’s an excellent adaptation, in other words, of F. Scott Fitzgerald’s melodramatic American classic. Luhrmann, as expected, has turned “Gatsby” into a theme-park ride. But he’s done it in exactly the right way. He hasn’t tried to make the novel more respectable, intellectual, or realistic. Instead, he’s taken “The Great Gatsby” very seriously just as it is."
- Twitter / Caterina: LinkedIn says I could apply ... - RT @Caterina: LinkedIn says I could apply to be Sr. Product Manager of Mobile for Flickr! Plus some other jobs. Woo!
Thursday, May 16, 2013
My first MOOC (Massively Open Online Course)
I recently completed my first Massively Online Online Course (MOOC), a term that presumably is at least passingly inspired by MMORGs, an online gaming genre that's most popularly represented by World of Warcraft.
The class was on Gamification and was well-taught by Wharton prof Kevin Werbach. But my focus here isn't to review or critique this particular class but, rather, to offer more general reactions to the instructional method. To reflect on what these courses seem to do well or at least handle relatively naturally, and what they struggle at. It's just a sample size of one—well, 1.5 actually as I'm currently taking a Data Science course that offers some additional insights—but I think it nonetheless exposes certain patterns.
I also encourage those interested in the topic to read Nathan Heller's "Is College Moving Online?" in the New Yorker, a thorough examination of the state of MOOCs and their potential effects on education—both for good and ill.
The format
The format for Gamification—which it seems is fairly typical—is built around a series of lectures. These consist of fairly typical Powerpoint slides with video of the instructor superimposed or off to the side in a small window. The production values are generally high and the combination of slideware and video is engaging.
There's a syllabus with links to various articles and other (free) materials. Prof. Werbach actually has a book on the topic of Gamification but it wasn't required for the course. None of the Coursera courses I've taken a look at had much if any in the way of stuff to buy.
The course then had a series of multiple-choice quizzes and a multiple-choice final exam—plus three written assignments of increasing length and scoring weight. My current Data Science course likewise has a series of lectures. But, in this case, the score comes from a series of programming and other assignments that relate to lecture topics although they are more hands-on and practical.
Type of schedule
Gamification, like the course I'm currently taking, comes from Coursera, which has a large course catalog from a wide range of schools. It was started by two Stanford professors and has received $16 million in funding from Kleiner Perkins Caufield & Byers.
Coursera's model, like that of the non-profit edX, is to offer classes in more of less real-time—by which I mean, an eight week course has a fixed start date followed by an end date about eight weeks later. Depending on the class, there may be more or less flexibility in how closely students have to hew to a weekly schedule for assignments, quizzes, and the like. But, fundamentally, the class is on a calendar and you can't dip in and out on the basis of work or family obligations, travel schedules, and inclination.
The downsides to this approach are obvious. There are a couple of classes I've considered but opted against because they overlapped periods when I would't have been able to devote much time to them.
On the other hand, after taking a class, I better appreciate why one might want to run a class to a schedule. Discussion boards, assignments (especially peer-graded ones—more on those in a bit), the availability of staff to answer course questions or address problems, and just getting forced into the "flow" of a class all require or at least greatly benefit from a schedule and associated incentive structure. (Education has a lot in common with a gamification system.)
In a related vein, I also better understand why you might not want to break a course into overly granular chunks, i.e. one or two week classes. I'm not just talking about any particular administrative overheads associated with putting a course in a catalog, but a broader set of transaction costs borne by all the participating actors such as figuring out how the class is run and understanding or dealing with prerequisites or tools. In many cases, it seems these would collectively just add too much overhead to a short course unless that course were more intensive than most people with a full-time job could undertake.
Does the fact that these Coursera courses so reflect the form and content of traditional university classes simply reflect tradition and the fact that much of the content was originally developed for such classes? Perhaps to a degree. On the other hand, whatever issues higher education may have today, it's also likely that not everything about traditional class instruction is wrong.
Another form of MOOC largely goes self-paced, even if related lectures still largely parallel the content of a semester of classes. Machine-graded quizzes can still exist, as can other types of computer- or self-evaluated assignments. Udacity, another VC-funded MOOC, follows this model today.
Other types of online instruction that increasingly diverge from a true MOOC model include iTunes University and the many instructional videos on YouTube. More focused sites, such as Code Academy, teach specific skills.
I think both more- and less-structured forms will have their place. Many of us have a limited ability to schedule courses that follow a fixed schedule and find the flexibility of watch-when-you-can attractive. At the same time, I appreciate the relative discipline and other potential benefits provided by a more formal course structure.
Grading and certification
Coursera, like edX, follows another aspect of most traditional university courses; it grades you. In Coursera's case, this takes the form of a Certificate of Completion based on your performance in various assignments and quizzes as determined by the individual class—70 percent in the case of Gamification. (Coursera is also starting to offer a "Verified" version for certain classes.)
I suspect that, at this point, a Coursera certificate is more of a gamification element, i.e. a motivator, than something that's especially useful outside of Coursera. However, it's also true that Coursera's business model will ultimately depend on being able to sell the ability to gain meaningful certifications—which means that they need to be able to grade.
Schools have, of course, been grading forever. But remember what the "M" in MOOC stands for? It's "massive," indicating that it's not at all unusual to have 50,000 people sign up for a MOOC. (Although far fewer will complete it.) It's obviously not practical for a professor and a few TAs to grade at that scale. In the future, I wouldn't be surprised to see hybrid models in which something "MOOC-like" is augmented with one-on-one and one-to-few interaction with professors and TAs for a fee—indeed we're starting to see examples of such—but let's stick to the subject at hand for today, given that what's actually being paid for starts to become a complicated question.
A couple observations about the state of grading in MOOCs.
Multiple choice works well, subject to the limitations of multiple choice. Given well-written questions, there's no more ambiguity than with any other multiple-choice exam. It's easy and instantaneously graded by computer. And modest feedback can be made available with respect to the answers.
Computers also do well at grading freeform but well-bounded and unambiguous answers, such as a numerical solution with only one correct response. It's "42" or it's wrong.
However, start talking more open-ended problems, even in quantitative topics like programming, and automated graders can start failing answers in unexpected ways. Without going into all the gory details, the autograder for the first assignment in my current data science course was highly sensitive to, for example, how the programmer chose to parse text fields and to any quirks in the output's format. While not fatal flaws, it's indicative of how automated grading challenges magnify exponentially the more creativity is allowed in solving an open-ended problem.
Also problematic is peer review. On the one hand, this offers a way for massive scale evaluation of freeform text in a way that it's hard for imagine computers tackling anytime soon. On the other hand, across a huge class with students of all ages, skills, language abilities, and interest, it's not hard to imagine that the evaluations can be… quirky—even given a relatively detailed scoring rubric.
I didn't personally have a huge problem with my results. But it was obvious from the discussion boards that a lot of people took low evaluations made without comment and evaluations made with obvious disregard for the scoring instructions very personally. And it's worth observing that, given a large class, statistics suggests that some will have simple "bad luck" with those they draw to evaluate their written assignments—even given multiple graders, multiple assignments, and algorithms to screen bad actors.
At the current stage of MOOCs, I personally find it easy enough to shrug my shoulders and get on with things. I have lots of diploma and certificate things gathering dust somewhere. But say a class has a written assignment contributing say, 33 percent of the grade. To the degree this grade has meaningful consequences for the student (for a resume or for tuition reimbursement if MOOCs start charging in some form), that's a problem. It's also fair to say that, real world significance aside, grades can have a motivating factor for many as well.
This hasn't been intended as an exhaustive look at the "grading issue" but it's been evident to me it's something that will have to be at least improved on—not that students are ever wholly happy about their grades—as MOOCs look to start collecting money from people and offering meaningful certifications.
Overall
I got a lot out of this course and it looks as if there's a lot of quality content out there—more certainly than I'll have time to dig into. I'm also happy to see both for-profit and non-profit initiatives probing at ways to make higher-education better and more efficient.
What I don't have a real opinion on is what the effect of Coursera, edX, and their ilk will be. The effect will certainly be uneven although one wonders if some aspects of MOOCs can replace elements of classes even at elite institutions. (Though I'd note that we've had the technology to replace 500 student freshman lectures for at least a decade.) I do suspect, or at least hope, that MOOCs—or at least MOOC methodologies—can replace low-value-add broadcast education in many situations.
One of the reasons that this particular crystal ball is cloudy is that higher education is often this odd hybrid of credentials, socialization, and learning that can be impersonal, highly personalized, solitary, involving lots of peer interaction, or some combination all of those. MOOCs clearly don't address all those modes but it arguably can do a subset rather well.

