Friday, August 23, 2019

Hugh Brock on Red Hat Research

Hugh Brock is Research Director at Red Hat. In this podcast, Hugh discusses how open source makes the way Red Hat approaches research different from the way it's done at other companies. He also talks about how the research program got started and, in particular, the role that Boston University has played.

Show notes:


Podcast:


Transcript:


-->
Gordon:  I thought I would get started off by talking about research programs in general at companies. There's a long history. Research labs, research organizations in corporations have often taken different forms. Sometimes there's a lot of fundamental research. Sometimes it's pre‑product development in a sense.
Can you take us through what some of the thinking was in Red Hat forming a research program and how you think about research, development, collaboration, intellectual property? That should take us a few minutes.
Hugh:  I think the way we're doing this at Red Hat is really exciting and also quite different from the way industry research has traditionally been done. The way companies...even the way companies have traditionally worked with universities. What we do at Red Hat Research is try to connect our engineers with researchers in universities via graduate student PhD projects.
The reason we work that way...well, there's two things, really. The reason we're able to work that way is because we're an open‑source company. The universities are happy to talk to us because they know that we're not interested in their intellectual property.
The reason it is an advantage for us to work that way is that it gives our engineers a chance to broaden their horizons, and it helps the universities focus what they're doing on something that is achievable in this now very fast moving world of IT and computer science research.
We think we have a winning model there. If you look at the way companies have traditionally done research, there's kind of two models. One model is you pay a researcher at a university to develop a project for you. I went to visit...Oh, I forget the guy's name. The research director at Mitsubishi Electric over here in Cambridge when I was first getting into this job.
One of the first things we learned in our conversation was that we have completely different jobs actually. Dick's whole enterprise is going to MIT and figuring out what he needs to pay them to do, which is a good model. It's good for MIT. It's good for Dick. He gets what he needs. All of that research then goes into a patent vault. Mitsubishi gets to use it but it doesn't get opened up until the patents expire.
Our model is completely different. What we're trying to do is, not pay researchers to do stuff that we want. It is to help them do what they want and get it into the open faster and more effectively than they could do it without our help.
So far, it's turned out to be a winning model. We hope that it will grow as we're able to contact more people and connect more people with our engineers.
Gordon:  One of the things that seems interesting in terms of academia, in industry, and open‑source, working together is...There are, of course, different objectives. There's also often different time scales in the way things happen.
In a way, I find academia interesting here because on the one hand, academia can often work in things that to someone sitting in a company, can look often very pie in the sky, very speculative. Maybe we will see something related to this in 15 years.
On the other hand, there's this incredible pace of change that you alluded to in IT and in tech where some project like Kubernetes goes from something internal at Google to something everybody is using in the course of just a few years. At least from teaching classes and so forth, things don't get revised that quickly at universities.
Obviously, there's also a tension of, you don't want universities chasing the toolkit of the day too quickly either.
Hugh:  Yeah. This is absolutely right. It's a real tension for universities. When we launched our relationship with Boston University, which would have been going on three years ago ‑‑ it'll be three years this December that we launched the Red Hat‑BU Collaboratory ‑‑ the case that my opposite number over at BU made to our VP of product, Paul Cormier was not only that Red Hat should do this because Red Hat will get interesting research out of it, but also that Red Hat should do this because we can help the universities do a better job of getting interesting research into the public.
My partner at BU, Dr. Orran Krieger, is a professor of computer engineering there. The case that he made to Paul was that industry is going so fast that academia needs to be pushed to keep up with them or academia will, in fact, become irrelevant. At least, in the computer science and computer engineering space. There's already a danger of that happening right now.
To some extent, there's a danger of it across the sciences with some of AI developments that we're seeing. The point is that we really are in a unique position at Red Hat to push research in a direction that's going to keep up with what we need in industry, and make a better, and closer partnership that works. That isn't parasitic or whatever, but actually serves the interest of both industry and university.
Gordon:  We'll get into some of that interesting AI work in a couple minutes. Before we dive down quite that deeply, let's talk about some of the threads that really came together in this whole program. Obviously, there's the universities, including universities in the Boston area.
Though, certainly not limited to that. The original Massachusetts Open Cloud work, which you just mentioned. Then, a lot of the interesting medical research that's happening in the Boston area. Probably many of my listeners know this, but Boston is known for having among the world class hospitals.
A lot of them are big research hospitals that are doing a lot of work. Can you talk about how all of this came together?
Hugh:  Yeah, it's actually a really interesting story. I mentioned the Red Hat Boston University Collaboratory. Red Hat Collaboratory at Boston University, I guess is the official name. This became the foundation of our research program at Red Hat. The way it came to being is a very interesting story.
Dr. Krieger over at BU wrote and received an NSF grant with his partner Peter Desnoyers about six years ago now, I want to say, five, six years. The grant was to study the feasibility of developing what's called an Open Cloud Exchange, which to put it very concisely is a bazaar to the public cloud's cathedral.
If Amazon is operating a single owner monolithic set of services that they control from top to bottom, the Open Cloud that Orran wanted to study is a marketplace where anybody can play. This is an interesting concept. There's a number of things that you need in order to make it real. The first thing you need is a data center.
It turns out that the five major research universities in the Boston area, Harvard, MIT, BU, Northeastern, and UMass system collaborated a few years ago on a data center called the Mass Green High Performance Computing Center MGHPCC, which is a very large data center in western Mass next to a hydro‑electric dam. Power's cheap.
They built this with the help of the commonwealth of Massachusetts so that they could move all of their research IT infrastructure out there.
With that done, it now is possible all of a sudden to start thinking about, OK, how can we build an Open Cloud that allows all five of these universities and the state and industries, small businesses, manufacturing to all participate in what amounts to a marketplace for computing services that looks like a cloud, but is ultimately much more efficient because there is no single owner.
It's more price efficient. This was the thing that Orran wanted to study. He got a grant. He got some computers, got a network infrastructure, set the thing the up, and set it up on the OpenStack while we Red Hat are a major OpenStack vendor. Orran kept pinging us saying, "Hey, do you guys want to help with this? This is really interesting."
He was trying to do some stuff with some of the OpenStack services that we at the time didn't think made sense. For a while, we put him off, like, "No, we're not really interested."
Eventually, he got through. The thing he was able to offer to Paul that no one else really had before was an opening for doing research in a practical way. Research that would draw our engineering team into partnership. That was a huge deal.
That was basis of this Red Hat BU Collaboratory, which is a million‑dollar a year partnership where we fund...Really, we fund basic research at BU, but we do it on a way that's connected to do what we want to do as an engineering company.
Fast forward a couple of years, we've established the Collaboratory. We're working with Orran at BU. Orran teaches what's called a Cloud Computing course. It's a project‑based course in Cloud Computing. This was the first year he had done it.
He brings in lots of industry partners as mentors to lead projects. We had a bunch of projects over there. One of the projects that we found in the Cloud Computing course was a little app that somebody at Boston's Children's Hospital had written to put a UI around medical image processing codes.
Codes that process MRI, for example, or any x‑rays or CT scans, or whatever, and do stuff with them. Not really AI, exactly, but advanced image processing. This project as it turns out was staffed by the wife of one our researchers at BU. Even funnier, the sponsoring researcher at BU turns out to be Orran's wife. He didn't know this at the time, which is hysterical.
We didn't realize that they were both working on opposite ends of the same project for a long time. They never talk apparently. Anyway, we got into this project and we realize that we could take this thing and put it to OpenShift fairly easily. That actually would be not only a great demo project for us, but a really nice contribution that would make back to this thing.
It is developed into a thing called the ChRIS project. That's how that got started. We're continuing to contribute to it, partly as because it's a great demo project for what you can do with OpenShift [Container Platform] on our tooling. Also, because we've been able to demonstrate some of the more advanced research results that we've found at BU.
Things like multi‑party computing, we have integrating that into ChRIS at that beginning part of this year. There are more things going to be coming there this year as well. That's been a really interesting story and it was a lot of fun.
Gordon:  The multi‑party computing and homomorphic encryption and differential privacy probably should be the topic for another podcast because they're really interesting. Essentially, ways that you can share data in a machine learning, AI context among institutions, and to have third‑party computing without compromising privacy.
That's actually a really interesting topic that...stay tuned, going to cover that in later episode in more detail. In addition to the MPC and chRIS, what are some other interesting projects that are going with research right now?
Hugh:  At BU in particular, we have two parallel tracks. One of them is the privacy‑preserving AI that you just mentioned. All the range of technologies around how to do machine learning in a privacy‑preserving way. The other kind of major research thrust that we have is all around how industry is going, and the world, is going to deal with the end of Dennard scaling.
Dennard scaling is the idea that you can continue to increase the number of transistors per square millimeter on a chip by doing various tricks to make that possible. It is a subject of a lot of debate exactly how long we will be able to continue to do Dennard scaling, but nobody is arguing that we'll be able to do it forever. It seems clear that we're approaching a point of diminishing returns.
What that means is that we are going to start looking again at specialty devices. Processors that are built for a particular purpose, the GPU is the most obvious example, but there are many of them. Many different types. The OS, which is our primary product here at Red Hat is going to have to start understanding how to deal with these things.
You can imagine a typical computer system in the cloud in three or four years is going to have not just a whole bunch of general purpose CPUs, but also GPUs, FPGAs which are programmable processors. You can program the architecture on the fly. Other devices we haven't even really thought of yet. All of these things are going to live together in the same machine.
We have a number of really interesting research projects going on right now that are all around the different aspects of that problem. Everything from can we build a Linux unikernel that really works to what do we need to do to create an open source tool chain for FPGAs. Any of a wide range of other projects along these lines. Partitioning hypervisors is another kind of key piece.
We've been very fortunate that BU turns out to have really strong Operating Systems department. We didn't know that when we made that partnership so that worked out quite well for us. We think we're going to do some groundbreaking stuff there. The unikernel project, in particular.
The idea of the unikernel is you take your app and you build it with the kernel it's going to run on so that you’ve basically built a bootable app. There's all kinds of interesting reasons why you might want to do that. It turns out that our kernel folks are really interested in this with the PhD who's working on it right now.
They're basically telling him, "If you can get this last stage of thing that you're working on right now to work, then we want to start looking at how we can actually use this in practice."
This is gone really quickly from pie‑in‑the‑sky idea Orran and Ali Raza who is the PhD, come in and say, "Hey, we want to look at unikernels and whether we can make a unikernel out of Linux and we think it's going to be really hard. We doubt it'll work," to in the space of not even, 18 months, Ali talking to our kernel engineers about the details of how we could make this real.
It is going very quickly. I think everybody's astonished that it's happened that quickly. It's an example of lots of different threads coming together at the same time. We've been really fortunate to be in the middle of it, but that's what we try to do at Red Hat Research. We weren't pulling these threads together, and then they would never intersect.
Gordon:  Probably topics for at least a couple more podcasts there. We have probably all the detail we can get into today. In closing out, where can people go to learn more about this? Many of these sounds interesting?
Hugh:  The best place to look for anything that we're doing is our website, which is research.redhat.com. We try to maintain there a list of all of the active projects and what the status is at all of the universities that we work with.
We also post there details of the events that we sponsor so colloquiums, workshops, things like that as well as the quarterly research review magazine that we produce to go into detail on these projects that we do.
DevConf is an annual conference that was launched in our Brno office in the Czech Republic. It's been running there now for, I think, 12 years. We did the first one here in the US last summer [2018] at the Boston University Student Center, the GSU, George Sherman Union.
We'll be doing it again August 15th‑17th this summer. It should be really good, I think. All of our interns will be presenting something there, particularly all of the PhD projects that I just mentioned. We think it's going to be a lot of fun.
Gordon:  Well, thank you, Hugh. Anything you'd like to close with?
Hugh:  I just want to thank you for reaching out and making this happen, Gordon. Thanks for listening everybody. If you're interested, again, in participating or just in knowing what's going on, check out research.redhat.com.

Wednesday, August 21, 2019

Trust, Enarx and TEEs, and open source security

Today, the Linux Foundation announced the intent to form the Confidential Computing Consortium, a community dedicated to defining and accelerating the adoption of confidential computing. As it so happens, I recorded this podcast at devconf.us last week with Red Hat security experts Mike Bursell and Nathaniel McCallum in which we discuss Red Hat Enarx, a project for providing hardware independence for securing applications using Trusted Execution Environments (TEE). It’s one of the projects that will be contributed to this consortium. We also cover broader issues of trust and open source security.

This episode is part of Innovate @Open, a new podcast that focuses on open source with a particular focus on how collaboration and openness are leading to new inventions and innovations.

Show notes:
Listen to podcast:
  • Trust, Enarx and TEEs, and the nature of open source security [15:51 MP3]
Transcript:

-->
Gordon Haff:  You're listening to "Innovate @Open." Stories from the cutting edge of technology innovation rooted in open‑source software and collaborative processes. I'm your host, Gordon Haff.
[music]
Gordon: What I have today is a podcast I recorded last week at devconf.us with Mike Bursell and Nathaniel McCallum. In that podcast we talked about the nature of trust and specifically about a new project called Enarx, which is an application deployment system that lets applications run within trusted execution environments.
This was particularly timely because today, August 21st, the Linux Foundation announced the intent to form the confidential computing consortium.
The basic idea here, is that as companies move their workloads to a bunch of different environments, hybrid computing environments, they need protection controls for sensitive IP and workload data and they're increasingly seeking greater assurances and more transparency of those controls.
The challenge is that current approaches in cloud computing address data at rest and in transit. Encrypting data in use is considered the third and possibly the most challenging step to providing a fully encrypted life cycle for sensitive data. Let's kick things off by having Mike tell us about trust.
Mike Bursell:  When you run any process, or you run any application, any program on a computer, it's an exercise in trust. You are trusting that all the layers below what you've written, assuming you've written it right in the first place, are things you can trust to do what they say they're going to do on the card.
I've got to trust my middleware, I've got to trust the firmware, I've got to trust the BIOS, I've got to trust the CPU or the hardware. The OS, the hypervisor, the kernel, all the different pieces of the software stack. I've got to trust them to do things like, not steal my data, not change my data. Not divert my data to somebody who shouldn't be seeing it. So that's a lot of pieces.
If you're looking at a standard stack of 10, 12 pieces, just think about all the different libraries you're using, all the different parts of the kernel, all those different bits. How can you trust that? That's a real difficulty.
It's the reason that people don't run sensitive workloads or keep really sensitive data on the public cloud, generally. Do you want to put your really sensitive data, your research, algorithms, programs on a public cloud service provider where they could look at it?
Or even, if you've got sysadmins, how much do you trust all of your sysadmins? Because a sysadmin, if they have root, could look at anything on any of those systems, even your internal systems.
Do you want your CEO’s payroll data to be in there? What about legal data about companies you are acquiring?
All of these things are difficult, and they are something that concerns a lot of people in the enterprise, in government, throughout the world.
We wanted to look at this. It just turns out there's some new set of technologies coming out right now called trusted execution environments. They are CPU and chipset technologies, which allow you to run programs, applications, in such a way that even the hypervisor, even root, even the kernel can't look into what you're doing.
That's great. Fantastic. They are publishing information from AMD, from Intel, from IBM, all the stuff coming out. But they're all different. They all handle the problem in a different way.
We, at Red Hat, started thinking about this and came up with some ideas. We decided that we wanted to make it easier for you to use these things.
Gordon:  That sounds really interesting. We are going into a little bit more about what this means, about what we have to trust, what we don't need to trust any longer. Nathaniel, could you walk our listeners through in a little more in detail, how this whole thing works?
Nathaniel McCallum: One of the things that we are concerned about is that a lot of our existing technologies require you, essentially, to write your application to the technology. It should be no surprise to the listeners here that Red Hat is very much against lock-in.
We want it to be possible for you to write your applications using the standard APIs that you already use, in the languages you already use, with the frameworks that you already use, and to be able to deploy these applications inside any hardware technology possible.
This is the goal of the Enarx project. One of the things we realized early on was that there's a new technology called WebAssembly which is being used in browsers all around the world. Literally, every single browser supports WebAssembly. It's being looked to very much as a sort of future to JavaScript.
The thing that's really interesting about WebAssembly to us, is that the capabilities that WebAssembly can deliver in conjunction with the WebAssembly system API. It is almost exactly the same set of functions that you can actually do inside these hardware environments.
It also means that you get to write an application in your own language with your own tooling. You can compile it to WebAssembly. Then Enarx will aid you in securely delivering that all the way into a cloud provider, and to be able to execute that remotely. The way that we do this, is we take your application as inputs, we perform an attestation process with the remote hardware.
We validate that the remote hardware is in fact, the hardware that it claims to be, using cryptographic techniques. The end result of that is not only an increased level of trust in the hardware that we're speaking to. It's also a session key, which we can then use to deliver encrypted code and data into this environment that we have just asked for cryptographic attestation on.
The end result is that you get to write your own application the way you want to write it. To get to deploy it in Enarx where you see fit, and you don't have to make your application depend upon specific hardware technologies.
Gordon:  In the show notes, I'm going to link to some information about this project. Could you take us through fairly a fairly high level how this works?
Nathaniel:  Basically, the way it works is that, once we've completed our attestation, we now have a session key that proves cryptographically that we are talking to our remote party. We can do so in a way that is encrypted using all of our standard cryptographic technologies.
We deliver the WebAssembly code that you have produced as part of your application, directly to our secure execution environment on the remote host. At that point, it is then just in time compiled for the actual CPU that you are going to run on. Then everything will be executing in that environment on the native processor.
We're also going to take care to enforce additional security measures. For example, if you persist any data, you'll be able to at some point, we don't currently implement this, but at some point will be available to read and write to a file system. The host will only see encrypted block devices.
The same thing is going to happen for networking. We're not going to allow unencrypted networking, but we will allow you to do TLS, for example, to communicate out. The end result is you just get to write your application. Then when you deploy it in this way, you have very strong assurances that there's a whole class of attacks against your application, that won't be able to get off the ground.
Mike:  What we're doing is we're basically...Remember, I talked to the beginning about how you've all got all of these layers you need to trust. We're removing the need for you to trust most of those layers, because the only things you need to trust are the chip vendor, and the firmware they provided. Which is all cryptographically signed, you can check that.
The Enarx code and the application you've written yourself and of course the Enarx code is going to be open. It's one of the kind of weird things about security. In order to be really confidential and to be closed to everyone else, you need to do your actual implementation and your coding and your design in the open. It's generally accepted these days that open source provides for better security overall.
We at Red Hat, of course, want everything to be open source. Enarx is completely open, will always be open. We're using open source technologies all over the place. We're using Rust as the main language. It's very well regarded for security, and for knowledge of what happens when things go wrong.
If you have faults in in your application, you know what's going to happen. It's not just going to start spilling digital as a place that could be used by a malicious host to work out what you're doing, for instance. We're developing in the open, so that you can do stuff in a closed way, with your sense of data, which you should control data and algorithms.
Gordon:  Take one of the things that we've seen over the last number of years and you being in security, Michael and Nathaniel, is that. I think so many people came to open source or looked at open source from the perspective of, "Oh, you never published the schematics for your alarm system if you're a bank."
No matter how good you think your security is, bring over that analog into the open‑source world and marry into the cryptography world, of course, that doesn't really apply.
Mike:  It's difficult because people assume that… That will be good analogy, that you've got your schematic of your bank vault, and a key. Cryptography is very different from that. You absolutely should never be using a cryptographic algorithm which is not known, and open, and peer‑reviewed. A really good cryptographic algorithm, the only thing that needs to be secret is the keys you're using.
They say that any fool can create a cryptographic algorithm that they can't break. I've certainly created cryptographic algorithms that I couldn't break and other people showed me how I'd gone wrong. It's one of these little things you learn to do as your apprenticeship in moving into security is doing this, so you understand how it goes wrong.
Cryptography is very much not like that. You need peer review, you need academics. Security is different, in some ways, to other parts of open source, in that there's this well‑known dictum that with enough eyes, all bugs are shallow, which is a great dictum of course.
But it's not quite that easy for security, because the number of people who have expertise in security, it's small. You need to ensure that their eyes are being applied. It's not good enough to have lots of un‑expert eyes looking at security. You need expert eyes, looking at security.
That's one of the reasons that companies like Red Hat ‑‑ but there are many others, Microsoft these days, Intel, IBM ‑‑ are spending a lot of time getting their security experts looking at cryptography and open‑source cryptography. Because it benefits the entire community and the whole ecosystem. What I call the commonwealth of what we are as an open‑source.
Nathaniel:  Just to add to what Mike was saying, this notion that with enough eyes, all bugs are shallow, this is predicated on the ratio of eyes to the amount of code.
When you are dealing with secure code, one of the things that you want to ensure is that because there are a limited number of eyes, we also need to try to limit the amount of code as much as possible.
This is why a project like an Enarx is all about reducing the trusted computing base. We want there to be a lot less code that you have to trust, which means that we need less eyes to review it to make sure it's secure.
Gordon:  We've certainly seen recently low levels of hardware in the light, that you get into these very complex pieces of engineering and it gets harder and harder to predict or to figure out every possible security exploit.
Let’s go maybe a little bit far afield as we wind this down.
Coming back to trust, what are some of the other areas in this low level, the software stack, in terms of trusted execution environments, in terms of firmware, in terms, perhaps, of CPUs themselves, where work is being done, or you think that there are possibilities to increase the security?
Mike:  Let me start with one, which is TPMs. TPMs have been around for quite a long time and people have not been using them. Partly because within the open‑source world, there was a great concern, 10, 15 years ago now, I guess in the early 2000s, that they were going to be used for DRM.
DRM has long been anathema to much of the open‑source community. They never got taken up usually within Linux and the open‑source community. There's a new version of TPM 2.0, which is much improved, and people are beginning to realize there's great benefit in using them.
The thing about a TPM is it's a hardware root of trust. It's really good for that if you need to be building up levels of trust because you can't do everything in Enarx, yet. There are times you need to build up trust, and it's a very good building block for those sorts of things. That's one example. Nathaniel, have you got some others?
Nathaniel:  There's been a variety of technologies, even besides the TPM. Unfortunately, none of them have really gone very well. Most of them have been hard to use, they've been hard to enable on the system, and they've been driven by a lot of concerns like DRM. Concerns that don't put the user first. This is why one of the key principles of the Enarx project is to make sure that we always put the user first.
Mike:  For years now, we've understood about encrypting data at rest when it's stored, encrypting data in transport when it's going over the network. We're now moving into a world where we need to encrypt data and algorithms in process.
That's what TEEs are for and that's what Enarx aims to make it easy for you to do as a developer.
Gordon:  Thank you for listening to this episode of Innovate @Open. For future episodes, subscribe to Innovate @Open on your favorite podcast app.

Saturday, January 12, 2019

Keeping my herbs alive: An indoor watering system

IMG 2815As I was again reminded in a recent Twitter thread, I like having fresh herbs. But you often don’t need a lot of them, which in turn means that it’s nice to have some pots of them growing at home so you can snip off just a little bit rather than buying 20 times what you need at the store. 

The problem is that 1.) I travel and 2.) I forget to water plants. One common low-tech method for automatically watering is to fill a wine bottle (or soda bottle, etc.) upside down in the soil. This works reasonably for a few days but my issue was the 2-week plant-killing excursions. So I looked around for commercially-available solutions.

There were a few systems online but at least one of them had to sit up above the plants, which wasn’t really feasible for my setup. And none of them seemed to have very good reviews. So I decided to see what I could come up with myself. I already had a small hydroponics system in the same room which turned my thoughts to supplying the water from an aquarium pump in a bucket of water on the ground hooked up to a timer. This is indeed what I ended up doing but getting to a system that actually worked took a fair bit of fiddling and experimentation.

My first thought was to “borrow” the aquarium pump I already had to keep my hydroponics system topped off in the summer and use some T-connectors to fan out the tubing so that I could water multiple plants. To make a long story short, the pump I had wasn’t powerful enough to lift water and force it through a network of T-connectors. Elevating the bucket on a stool helped. Sort of. But if I lifted it too high, the water started siphoning when the pump turned off. Furthermore, as I discovered when I experimented with different tubing sizes, in a network of tubing like this, it’s hard to get relatively even flows out of the different lines.

What ended up working the best had three basic elements:

A more powerful aquarium pump (head of about 8 feet)

A manifold intended to split the output of an aquarium air pump

A digital timer with 1 second resolution

The manifold was really the key thing here because it splits the flow pretty evenly. Furthermore, each of the outputs has a little flow control valve that lets you tweak the flows so you can account for different line lengths or different amounts of water for different pots. You just have to experiment to see how long you want to run the pump for. For me, it’s about 15 seconds which is probably a little on the short side; I sometimes end up watering a little every week or two. 

Now, in practice, actually building this was more complicated because of getting all the tubing sizes right. The big culprit was the pump. In order to get an aquarium/pond pump with sufficient head (pressure) you need to get one capable of vastly more flow than needed for this application. And because the pump is designed for that higher flow, its output size is fairly large so you need to adapt the tubing down to the significantly smaller diameter that’s appropriate for this application (and the manifold). I got it almost right; I had to use some epoxy paste in one place because I couldn’t find an adapter that was quite what I needed.

(There’s also a lot of inconsistency in how sizes are advertised, e.g. both the check valve and the manifold are supposed to be 3/8” but the tubing fits easily on one and was a very tight squeeze on the other. You may need to fiddle around.)

Finally, I added a check valve though it probably isn’t needed.

In addition to the parts listed above, here’s what you need:

  • Short piece of 3/4” tubing (I had some black pond corrugated pond tubing from earlier experiments)
  • You may need some epoxy paste if a connection is a bit loose connecting this tubing to the adapter
  • Multi-hose adapter
  • 3/8” tubing (probably ID)
  • Aquarium airline tubing (I believe it’s 3/16” diameter)

Finally I was trying to figure out what I could stick the pots in so that I didn’t need to worry if there was some overflow. I was coming up more or less blank. The options either weren’t really the right shape or they were a lot deeper than I needed or wanted (or both). Then I found the perfect thing: a 26” water heater drain pan. Just cover up the hole for the drain with duct tape. To make it even more perfect, I happened to  have a round 26” table up in my attic.

Now my herbs are happy and I’m happy. 

Friday, January 11, 2019

More 2008 redux: Open APIs

Given the discussion going on around API openness these days, I thought I’d resurrect yet some more text from an "Open Source vs. the Cloud" research note that I wrote in 2008. (See also "Open Source vs. the Cloud Redux” and this twitter thread.)

At the same time, to focus on source code is to focus on a specific type of openness and freedom that was important historically—but may not be as important going forward. Indeed, in the case of Web services running on massive server farms and cooperating over a network with all manner of other code, services, and data, the value of code is questionable. After all, you can hardly just load it up on a server and do anything useful with it anyway. One needs all those servers and interlocking pieces. Also, the ability to view, modify, and redistribute source code is only one of many rights or protections to consider in a Cloud Computing world. For example, consider these other things that might matter more:

...

Open APIs. Open Source as we know it today evolved largely in the context of Unix-like operating systems and the programs that ran directly on top of them using “libc” and other system libraries. While we may run monolithic programs over the network, much of the action in Web 2.0 has been in services such as Facebook, Flickr, Google Maps, and Salesforce.com that expose application programming interfaces (API) at a higher level. This allows developers considerable freedom to extend these platforms. Thus, whether a platform or application is Open Source or not, given public APIs, it can be extended and consumed in ways that are very analogous to Open Source. At the same time, the predictability and transparency of the terms of service for APIs—especially in the case of consumer-oriented services—raise their own issues.

Thursday, January 10, 2019

The cloud vs. open source redux

If you’re reading this, you’re probably aware that there is a fracas going on around open source licensing. Quite a bit has been written on the topic and I won’t rehash the specific details here; they have been well covered by:

However, to net l’affaire out, long-simmering issues associated with building businesses on the back of open source software are boiling over. In particular, cloud providers like Amazon Web Services (AWS) not only rely on vast amounts of open source software to run their infrastructure but they’re increasingly offering cloud services that directly compete with the companies that created much of that open source software in the first place. Furthermore, there’s a widespread (largely justified) perception, that some of these providers in particular are taking from the open source commons far more than they’re giving back.

Some thoughts.

This is not a new concern

As an industry analyst, I wrote a research paper titled “The Cloud vs. Open Source” in 2008. That’s just two years after AWS debuted. Much of the paper cautions against getting too fixated on source code when thinking about user freedoms and openness generally. This remains true today and is the subject of an entire chapter of my book How Open Source Ate Software that I published last year.

However, I also argued that throwing up roadblocks to making use of open source software was ultimately unproductive.

Today, Open Source is widely embraced by all manner of technology companies because they’ve found that, for many purposes, Open Source is a great way to engage with developer and user communities—and even with competitors. Therefore, the concern that, left to their own devices, companies will wholesale strip-mine Open Source projects and “take it all private” seems anachronistic. That’s not to say that everyone will always contribute as much code without copyleft as with it, but the suggestion that copyleft is all that’s holding the whole Open Source process together just doesn’t square with the facts.

Was I just wrong?

Now, at this point, you might turn around and say: “But wholesale strip-ming is exactly what’s happening. We need even stronger protections if the commons is not to be ruthlessly exploited!"

One problem is that all the evidence suggests this doesn’t work. Permissive licenses like Apache, MIT, and BSD have gained in popularity over time. There’s a reason for this. Much of modern open source’s success isn’t about the ability to view source code. It’s about its collaborative development model. And the Eclipse Foundation's Ian Skerrett argues that "projects use a permissive license to get as many users and adopters, to encourage potential contributions. They aren't worried about trying to force anyone. You can't force anyone to contribute to your project; you can only limit your community through a restrictive license."

Another data point is the AGPL. At the time the new version of the copyleft GPL came out (GPLv3 in 2007), the Affero General Public License was introduced as a new GPL variant. Copyleft basically says that if you distribute software, you have to make the source code available. This includes any changes you made. The rub is that, under the GPL’s terms, “distributing” basically means shipping software on a disc or offering it for download. This creates what some saw as a loophole because offering the software as a cloud service isn’t distribution as traditionally defined.

(Is this starting to sound familiar? I told you none of this was new.)

Enter the AGPL, which was just like the GPL except the definition of distribution was broadened to include offering software as a service.

However, the AGPL hasn’t been much used. Ironically, one of the users was MongoDB, which is one of the current companies that have relicensed their software to prevent its use by cloud providers. In general, lots of companies are nervous that the AGPL could potentially interact with internal code that they don’t want to make publicly available. So its often on the license no fly list for application development.

Which is all to say that the overall direction in open source has been away from restrictive licenses. Leaving aside whether an even more restrictive license could still be reasonably considered “open source,” there just seems very little appetite for such a creature.

Words matter

In my view, a lot of the heat around licenses like the Commons Clause comes about because the companies involved seem to be, on the one hand, trying to gain the perceived value of a proprietary license while also getting credit for still being open source. “Open core” arguably plays the same parlor trick.

It still might have been news if one or more of these companies simply relicensed some or all of their software to a license that was unabashedly proprietary even if it retained some aspects of open source. But I suspect it would have been much less of a tempest.

Whether or not doing so would have been a good idea is a separate question. But it’s their software. Their business challenges are real. Own it. If you’re not going to have an open source development model, I’m not sure why you particularly even care if it’s technically open source or not.

Can we make cloud providers do better?

While the software vendors are taking heat from one side. Cloud providers are taking it from another. There is indeed a widespread view that most cloud providers are takers rather than givers. AWS, as the #1 cloud provider, takes particular heat. It’s mostly deserved. Although Adrian Cockcroft’s team has arguably moved the needle in making AWS play better with open source communities, much more could be done.

However, publicly shaming Amazon will not be a very effective strategy to drive change. If you can't sell the business value of participating in open source, you've pretty much lost the battle. Shaming might net you some contributions for the PR value but certainly no real commitment. Pinning your hopes on Jeff Bezos’ altruism is not a winning move.

Instead, as Linux Foundation Executive Director Jim Zemlin  told me during an interview at the Open Source Leadership Summit last year: 

The epiphany that many companies have had over the last three to four years, in particular, has been, "Wow. If I have processes where I can bring code in, modify it for my purposes, and then, most importantly, share those changes back, those changes will be maintained over time.

"When I build my next project or a product, I should say, that project will be in line with, in a much more effective way, the products that I'm building.

"To get the value, it's not just consumed, it is to share back and that there's not some moral obligation, although I would argue that that's also important. There's an actual incredibly large business benefit to that as well." The industry has gotten that, and that's a big change.

In closing

None of this is to dismiss the underlying challenges that these changes came in response too. There will always be challenges at the level of the individual company trying to build a business no matter what the product. But there are more macro dynamics here as well. 

This shift of computing towards public clouds recreates a new type of vertically integrated stack. One-time chief technology officer of Sun Microsystems, Greg Papadopoulos, one suspects hyperbolically and with an eye towards something IBM founder Thomas J. Watson probably never said, suggested that “the world only needs five computers,” which is to say there would be “more or less, five hyperscale, pan-global broadband computing services giants” each on the order of a Google.  

Some cloud giants have indeed made significant contributions to open source projects. For example, Google originally created Kubernetes, the leading open source project for managing software containers, based on the infrastructure it had built for its internal use. Facebook has open sourced both software and hardware projects.

But, for the most part, these dominant companies use open source to create what are largely proprietary platforms far more than they reinvest to perpetuate ongoing development in the commons. And they’re sufficiently large and well-resourced that they mostly don’t depend on cooperative invention at this point.

It’s easy to dismiss free-riding as a problem given that organizations are missing out in some ways if they do so. However, to the degree that large tech companies, both cloud providers and others such as Apple, take far more from the open source commons than they contribute back, this at least raises concerns about open source sustainability.

The last section is based in part on content from How Open Source Ate Software (Apress 2018)

Wednesday, January 02, 2019

2019 New Year updates

A new year. A few updates.

In 2018, for the second year in a row, I published a book. How Open Source Ate Software with Apress. Buy early, buy often as the saying goes. If I do anything along these lines in 2019, it will probably be a mini-book about the lessons we can take away from Intel’s Itanium processor that I followed closely while an analyst. I’ve written an outline and we’ll have to see if I get the energy to do something more.

In part because of my open source book, this blog has been a bit inactive of late. I’ve also been publishing in a number of other places, including opensource.com, The Enterprisers Project, and TechTarget. I plan to continue doing so but I’m going to try to get back into the swing of jotting down quick thoughts here—related to both professional interests and otherwise. (I had planned to kick off a separate travel and food blog last year but I pretty much got no further than registering a domain. I’m just going to post here as the mood strikes.)

One decision I made over the holidays was that I’m going to pull the plug on my Cloudy Chat podcast. It’s gotten pretty irregular and unfocused and I don’t really see that changing. Furthermore, while my work focus remains fairly broad, it’s shifted towards emerging technology topics, which “cloud” really isn’t any longer even if some of its newer aspects are. This isn’t to say that I won’t do the occasional published interview. But I think a podcast implies a schedule and topic focus that I don’t really see being in the cards for the immediate future.

What may take its place is some more work with video. I haven’t fleshed out what that means exactly but it’s an idea I’ve noodled on over time. I’m hoping that taking “I really ought to do a new podcast” off my omnipresent mental to-dos will free up some cycles to create short educational videos related to emerging tech areas.

Contact and social media information is unchanged. I’m on twitter as @ghaff. My photos are on Flickr. (I’m hopeful for Flickr’s revitalization post-Yahoo.) I’m on LinkedIn but, if I don’t know you, you send a form invite, and our connection isn’t obvious to me, I’ll probably ignore it. I’m still on Facebook but I’m pretty selective about friend requests; if I ignore you don’t take it personally if you’re a casual and/or professional acquaintance.

I occasionally do product and book reviews. Feel free to inquire if you have something you think might interest me. But I’m don’t do a lot and I emphatically don’t write about things that I haven’t gotten hands-on with. I also sometimes do interviews at events. (See above re: podcasts however.) I mostly limit these to discussions about open source projects and other non-commercial topics. Otherwise, there’s just too much opportunity for conflicts of interest. 

I’m always open to speaking at/attending events on professional topics of interest. Some presentations are up on Slideshare. I do have a busy travel schedule though so I have to prioritize where I spend my time. (Somehow, the first few months of the year ended up especially crazy.)

To be done: Updating my overall website. It’s still a good source for links but I never really liked the design I used and I’ve held off on updating it recently as a result. 

Tuesday, August 07, 2018

What was new at Serverlessconf?

Cue the obligatory “There still servers in serverless” and “There is no cloud; it’s just someone else’s computer.” Like many, I’m not a fan of the term but it’s seemingly here to stay. I’ll get over it, just like I did with private cloud. In any case, the more interesting question is what’s going on with serverless—which is what took me to Serverlessconf in cool, gray San Francisco last week.

This is the point where it makes sense to introduce serverless for anyone who may have heard of the term but hasn’t studied it closely. And that turns out to also be a good segue for discussing the conference as a whole. As recently as last December, I was on a panel and a member of the audience asked us about the difference between serverless and Function-as-a-Service. We were able to offer the beginnings of an answer at the time—not that everyone was quite singing from the same hymnal—but my sense is that most everyone is in alignment now, with a caveat that I’ll get to presently.

So what are serverless and FaaS? For that I’ll turn to a recent blog post by my Red Hat colleague William Markito Oliveira who wrote in a recent blog post:

Functions-as-a-Service (FaaS) is an event-driven computing execution model that runs in stateless containers and those functions manage server-side logic and state through the use of services. Serverless is the architectural pattern that describe applications that combine FaaS and those hosted (managed) services. MartinFowler.com has a great article that provides more details and the origin of the terms.

In other words, FaaS is one component of serverless. But an application written using serverless patterns will also generally use a variety of standard building blocks to provide common services across many applications. Databases, authentication, and proxies are examples of services that many different applications require. These will be managed by an operations team; from a developer’s perspective, it doesn’t matter if that team works for the same company and runs them on-premise or if the team is employed elsewhere, such as a public cloud provider.

This brings us to my caveat. While most everyone agrees on what makes for a serverless architectural pattern, there’s far less unanimity on the degree to which this pattern mostly applies to public clouds only, fits with various hybrid application development and deployment models, and over what timeframe the assumed shift to increased serverless usage takes place.

Serverlessconf itself rotates fairly far on the public cloud angle. But that really shouldn’t be surprising. The organizers of the show are a training organization focused on public clouds. Of course, they’re not the only ones to see serverless as part and parcel of the rich set of public cloud services that complement FaaS on public clouds. The finer pricing granularity of many of these cloud services (including FaaS) relative to per-hour or even per-minute on Amazon Web Services EC2, for example, is also seen as a feature that doesn’t really translate to an on-premise environment. (That said, there was far more emphasis at this conference on developer productivity advantages than on pricing models; I want to say this represents at least a shift of degree from this conference in New York last year.)

At the same time, a number of talks expressed a pragmatic recognition that serverless and FaaS isn’t for everyone, at least today. Amiram Shachar disused “Shipping Containers as Functions.” Yochay Kiriaty pointed out “Mistakes and Anti-Patterns in Serverless (or when NOT to use Serverless).” Kiriaty noted that serverless best fits with a specific async and event-driven programming model. Other characteristics that are mostly needed to benefit from serverless include: stateless logic, idempotence of functions, one task per function, and functions that finish quickly and avoid recursion.

More generally, Erica Windisch argued that 12 factor provided guidelines. But serverless enforces them. 12 factor is most associated with early hosted platform-as-a-service, so this may give you some sense of the type of applications that are primary serverless targets today.

A tweet from Andrew Clay Shafer that I assume was inspired by this or another talk stated: “Serverless is a particularly opinionated PaaS.” This is worth pondering, if only because I think there are a lot of hard and fast lines being drawn around what are essentially architectural patterns: serverless, containers, PaaS, VMs, IaaS, CaaS (containers-as-a-service) whereas it’s more of a continuum with adjacencies that blend into each other and combine features. A topic for another day.

Wednesday, July 25, 2018

Google Next: Enterprise, ML/AI, open source, and hybrid

Things I learned at Google Next (or at least I think I did). Think of these as preliminary observations based on sessions I watched, people I spoke with, or tweets I read/replied to.

One thing about the show reminded me of an AWS re:Invent from four or so years ago. Earlier re:Invents mostly trotted the usual suspects up on stage: Netflix, SmugMug, other startups that probably are no longer with us. Suddenly, we had NASDAQ and other well-known large enterprise customers who demanded mission-critical out of their infrastructure. This year at Google Next saw a similar transformation. There were young companies doing cool stuff; Indonesia's Go-Jek made a particular splash. But there were also plenty of speakers from companies like Nielsen and Target (which is leaving AWS in favor of Google). [ADDED: At least that was the image projected from the main tent stage. As a colleague correctly noted to me, the show floor told a much more startup-centric story relative to where AWS and Microsoft Azure are today.]

I'm not sure I heard any direct, or even oblique, references to competitors but I think it was pretty clear where Google thinks its differentiation lies. The most prominent area was ML and AI which was omnipresent. There were some explicit announcements of ML services, mostly aimed at making ML more accessible to the millions of developers who are not data scientists. But elements of AI/ML were pervasive whether as part of Google Maps or GSuite. The second area was open source. Open source as a central strategy was strongly reflected in the small Community day I was invited to (and gave a lightning talk at) on Monday but it was also front-and-center on the opening day keynotes as well.

Google announced that Cloud Functions was GA and took the covers off Knative. (Knative essentially helps create a common building block for serverless on top of Kubernetes across hybrid clouds. My colleague William Markito Oliveira has a nice piece up that discusses FaaS, serverless, and Knative in more detail.) Google sort of soft-pedaled serverless (which I'll use as the general term even if I don't like it) though. They announced Knative on Day 1 in press release. But serverless only got a short segment in the Day 2 keynote. I'm not sure what to make of it. One theory I heard, which I sort of like, is that given AWS' FaaS time-to-market and mindshare lead with Lambda, Google is taking advantage of Kubernetes' container mindshare to enter the market through that door rather than take on Lambda directly.

To expand on the previous point slightly, here's a quote from the Day 2 keynote: "Containers are the universal platform for cloud." I'm pretty sure neither Microsoft nor AWS would make that statement. And I think I'm not over-analyzing to say that Google views serverless as part-and-parcel of a broader container infrastructure and cloud-native app dev environment as opposed to a discrete technology. For what it's worth, I agree with this view. I think there's too much drawing of hard lines between these different approaches to writing and running services going on--but that's a topic for another day.

Finally, the hybrid thing. Google announced an on-prem version of its Google Kubernetes Engine (GKE). First of all, I think I should be getting royalties for some of their messaging; it sounds a lot like various things I've written over the years. But I digress. It's a good story and one with which I obviously agree. There's clearly an appetite for being able to run workloads portably across different environments. But I'd just observe that this is very new territory for Google. Enterprise customers bring a lot of quirks, integration needs, and customization requests to their in-house infrastructure. Heck, if they are happy with a fully standardized offering, they probably should be looking at just using a public cloud. So, strategically this makes sense. But it's not really in Google's wheelhouse and they may find this sort of offering less amenable to the sort of technical solution they're accustomed to creating.

More to come but these are some observations after the first couple of days.

Wednesday, July 11, 2018

Links for 07-11-2018