Wednesday, October 19, 2005

Wikipedia mea culpa

I've picked at Wikipedia quality issues a few times of late, such as in this post. Andrew Orlowski has a nice piece in The Register on Wikipedia quality in which he notes
Encouraging signs from the Wikipedia project, where co-founder and überpedian Jimmy Wales has acknowledged there are real quality problems with the online work.

Criticism of the project from within the inner sanctum has been very rare so far, although fellow co-founder Larry Sanger, who is no longer associated with the project, pleaded with the management to improve its content by befriending, and not alienating, established sources of expertise. (i.e., people who know what they're talking about.)

Meanwhile, over at Smalltalk Tidbits, Industry Rants, James Robertson argues that it's really hard to identify true expertise, especially where topics are controversial. For example, historians still argue about the origins of World War I much less recent U.S. presidential elections.

I can't argue with that. Nonetheless, many of Wikipedia's flaws that I've seen aren't about divining hard-to-understand causes and effects but about pretty basic matters of fact--and a lot of really bad writing.

Tuesday, October 18, 2005

Interruptions

"Continuous Partial Attention," the state of constant interruption that characterizes many of our lives get written about a fair bit in one form or another. See here, for example. Good article on the phenomenon in this New York Times article (with data. Perish the thought!)
Yet while interruptions are annoying, Mark's study also revealed their flip side: they are often crucial to office work. Sure, the high-tech workers grumbled and moaned about disruptions, and they all claimed that they preferred to work in long, luxurious stretches. But they grudgingly admitted that many of their daily distractions were essential to their jobs. When someone forwards you an urgent e-mail message, it's often something you really do need to see; if a cellphone call breaks through while you're desperately trying to solve a problem, it might be the call that saves your hide. In the language of computer sociology, our jobs today are "interrupt driven." Distractions are not just a plague on our work - sometimes they are our work. To be cut off from other workers is to be cut off from everything.

For a small cadre of computer engineers and academics, this realization has begun to raise an enticing possibility: perhaps we can find an ideal middle ground. If high-tech work distractions are inevitable, then maybe we can re-engineer them so we receive all of their benefits but few of their downsides. Is there such a thing as a perfect interruption?

Via this Joel on Software post

Good questions about Video iPods

I'm always a bit hesitant to comment on the usage models underpinning a lot of electronic gadgets when it's pretty clear that I'm not the target demographic. I own a couple of different iPod flavors (a Shuffle and a standard pre-photo 40GB one), but don't have a whole lot of interest in showing off photos on the little screen—much less video. However, I don't typically show snapshots around either or take a lot of video with my digital camera, so I've sort of shrugged and figured I'm just not the audience.

That may be true, but this article from MSNBC about the video iPod mirrors a lot of my thoughts. For example:
You can listen to your iPod at work, at home, at the gym, in the car. Basically, that means everywhere. But watching your iPod might not be as convenient. It will be tough to watch at work or while jogging. And in the car? Passengers, yes. Drivers, no.

Good questions.

[UPDATE: For the opposing view, here's a piece from Information Week.]

Thursday, September 29, 2005

The Post-Modern Equivalent of Brass Candlesticks

In this day and age, not many folks still have brass candlesticks to polish. But we apparently have a replacement activity: Apple nano polishing

via: Make. .

Amen

The highlight of an interview with Jeff Jones of IBM Analyst Relations:
What are things you would change?

I would rewrite PowerPoint to allow no more than 10 charts in any presentation. I would rewrite Notes' calendar feature to disallow the creation of meeting invitations that lack at least five sentences of explanation as to the purpose of the meeting. I would also remove the recurring meetings feature of Notes' calendar.

I don't mind the recurring feature so much. (Although when I worked for a vendor, I did accumulate a lot of pointless weekly meetings over time so I know where he's coming from.) But "no more than 10 charts." Yes!

via helzerman.com

Knowing Cultural Context

Tokyo-based blogger Sean Kinsell eviscerates this WaPost story purporting to find a new girlyness among Japanese men. For a primer on how important local knowledge is to understanding any supposed trend, this post is hard to beat.

Worthwhile reading. I don't have any good links handy but I can't count the number of times that I've read something which drew broad brush conclusions about behaviors or trends in a city or place. (Of course, local reporters do this too.)

via Virginia Postrel

Friday, September 23, 2005

Opting In and Opting Out

We're starting to hear the "Opt Out" defense a lot in some recent copyright cases, first with the Internet Library (which I discussed here, here, and here) and now with Google's tussle with the Author's Guild. As Tim at O'Reilly Radar comments:
Google's opt-out position is exactly the right one. If we were to wait for publishers to opt in, only current, in print works would get into the index.

I think he's almost certainly right. Because inertia and fear of losing control of their IP, most publishers would probably think it simpler and safer to do nothing. However, as I commented earlier in the Internet Library situation, it hard to see the legal significance of providing an opt-out mechanism. The sort of archiving that Google is doing may or may not fall under fair use (I don't really have an opinion on that, but let publishers opt-out is just a courtesy. Here's what an expert has to say (The Patry Copyright Blog--a great, if very detailed and technical source, for copyright discussions):
The legal issue remains the same, however: whether copying of an entire work without authorization is an infringement where the ultimate user is able to see only a few sentences of the original. Since fair use is an unconsented to use, the fact that publishers object doesn't matter, regardless of the chutzpadik way Google may have handled the issue (The Second Circuit is divided on whether bad faith is a fair use factor). And whether Google is actually an advertising behemoth that doesn't want its own service to be used to investigate itself, whether it is, therefore, a false great white knight in the culture wars (if there are any) shouldn't matter either.

Tuesday, September 20, 2005

More Wikipedia Weakness

It's often tempting to give Wikipedia a pass. After all, many of its most egregious errors get fixed over time--at least if the topic isn't controversial and the article has had enough time to "settle down" from breaking news or changing facts. But, as I've commented before here, here, and here, when Wikipedia is bad, it can be pretty bad.

Case in point, the other day I happened to run across an article on Data General AViiON servers, a topic with which I have a more than passing acquaintance. I was the product manager for the first AViiONs and handled marketing for many products of successive generations. Now this is the type of article on which one would probably be inclined to trust Wikipedia to get things more or less right. After all, it's an uncontroversial technical topic of the dead (but not too distant) past.

Don't trust those instincts. Apparently the article is also sufficiently off the beaten path that it hasn't had a chance to benefit from the Wikipedia "community" because it's rife with howling factual errors. It has SCO writing DG/UX (Data General's flavor of Unix) for example, contains basic misconceptions about Non-Uniform Memory Access (NUMA) architectures (an important piece of the AViiON line), and its storyline about DG's historical competition with DEC is pretty far off-base. But my intent here isn't to belabor the details--many will probably be fixed eventually, perhaps by me. And, I've seen trade press articles almost as bad. But this serves as yet another cautionary tale. Even where Wikipedia "should" be good, it often is not. Exercise appropriate caution.

Monday, September 19, 2005

DVD Wars

I'm not an expert on this stuff, but this writeup from Engadget seems to do a nice job of cutting to the basics of the Blu-ray vs, HD DVD fight (with a nice precis of DVD history into the bargain). Bottom line?
Still with us? No? Blu-ray discs are more expensive, but hold more data; there, that's all.

So now that you know why Blu-ray discs cost more and why Sony/Philips and Toshiba are all harshing on one another so much, we can get to the really important stuff: the numbers, and who’s supporting who.

Friday, September 16, 2005

Disk To Flash

Yes, it's been a while since I've posted. My bad. Lots of travel in various permutations of business and pleasure and the organizational wreckage that comes from such. (Intel Developer Forum in San Francisco, hiking around Mt. Shasta and Point Reyes, Sun's "Galaxy" launch in NY, etc.) But I'm more or less dug out now.

To me, one of the more interesting aspects of the iPod nano announcement was that it's replacing the mini in Apple's product line--and thereby, in a single stroke, shifting a huge chunk of portable MP3 player volumes from disk to flash. (The iPod mini uses a hard disk while the nano uses flash memory.) Does this presage even more shifts from spinning steel to silicon?

The short answer is probably yes. The longer one is a bit more complicated.

Certainly it's hard to see the shift as anything but a complete one for many, many years to come. This post by Illuminata colleague Tom Deane gives some of the reasons. Price per bit is one big issue. Limited read/write cycles are another. Just as tape continues to exist alongside disks, disks will continue to exist alongside flash.

However, flash is well on its way to eclipsing disk for most really portable applications. Advantages like size, ruggedness, and lower power trump somewhat higher per-bit costs for the most part. The nano's a classic case study of the tradeoffs. Apple's taken the somewhat daring move of replacing its wildly popular mini with a device that's more expensive per capacity today. It's betting that the svelteness and longer batter life of the nano make up for giving up storage. One can reasonably argue with the timing; I would probably disagree, but it's a plausible argument that Apple could have waited until some price crossover point was reached. But there can be little dispute that flash continues to take over disk territory in these handheld (and smaller) devices.

That's because Moore's Law is outpacing usage models. Human hearing isn't getting better. We're not better able to watch movies on a small LCD screen. Useful digital photo resolution ahsn't plateaued yet but it's probably getting close. In other words, we're rapidly moving towards points where additional capacity in many types of devices becomes less and less important relative to other characteristics such as lightness of weight. As flash memories rapidly head into the multi-GB range, the size and number of MP3 files aren't increasing apace.

The replacement of disk by flash is one implication of these intersecting curves. Another is that silicon technology will increasingly not be a factor in how many functions can be crammed into a single device. Which isn't to say that everything will be. There are certainly user interface issues (which Apple can probably solve as well as anyone) as well as more subtle and less rationalist issues of style and brand.

Thursday, August 18, 2005

Scheduling Problems

Fellow analyst Stephen O'Grady asks "Why is Scheduling Still So Damn Hard?"


Think about how you schedule meetings with folks outside your own calendar system:
* Step 1: Determine your own availability
* Step 2: Communicate that availability to an external party; typically means cut and pasting or manually writing some openings into an email
* Step 3: If you're lucky, some of these work, and you receive a reply which requires you to create a new calendar entry
* Step 4: If you weren't lucky in Step 3, the available slots didn't work, and the external party has proposed some alternatives so you're back to Step 1. Rinse, lather, repeat.

Certainly one root of the problem is the lack of appropriate protocols and mechanisms to selectively open up our calendars beyond the firewall. Yet just another example of the horrible state of collaboration, Microsoft Office 2003 dinosaur ads notwithstanding. Yes, it would be nice to do the same sort of group scheduling we can do with Outlook/Exchange with folks at other companies--or indeed could do with proprietary office automation products like Data General's CEO, fifteen years ago. Yes, that would be a good start. But we should also set our sights higher because the Outlook way of doing things isn't all that scalable either. The ultimate goal should be to have the system able to make intelligent decisions for us rather than just present us with a bunch of out-of-context data.

The real problem here is that computers are so bloody literal. They can schedule a block of time around already scheduled blocks of time, but that's about it. But scheduling is more complex than that. This isn't new. Consider the following from The Digital Deli, a marvelous look at PC culture circa 1984.
Next [in the list of computer applications to avoid] we come to the computerized electronic calendar. It doesn't let you make dates; you have to make "events." Can you imagine saying, "We fell in love on our first event"?

Worse, it suffers from that picky literal-mindedness of machine-think. For me, as for most folks, time unravels in a drinks-with-Chris-late-next-week sort of way. Electronic calendars are not that loose. To them, "late next week" means nothing. "Noonish" means nothing. They don't know about "happy hours," and they've never even met Chris. But "07-08-83, 6:00 P" they understand. No wonder we don't get along.

The electronic calendar also wants me to tell it just how long the "event" will last. The program divides the day into fifteen-minute blocks and has to know how many of those will be filled by my event with Chris. Now, I usually know roughly (within a half-hour, say) how things will go, but one has to be flexible on this sort of thing. Maybe an old mutual friend stops by the table, or we suddenly decide to go see a movie. Or there's a full moon out and it's a warm night ... You get the idea. Well, in electronic date-books, it's not enough to write "Chris" at 5:00. They want to know where you'll be at 5:15, 5:30, 5:45, 6:00. Sometimes I get the feeling this program was designed by somebody's mother.

I want a way to describe to the computer that I'd prefer to not have three meetings in a row, that I prefer not to have a call at 5pm but I can if there's no other choice, that I've got three hours blocked off for writing but I can take a call during that slot if it's important enough, and so forth. I'm not so naive to believe that we can easily get a computer to actually grok all our preferences and options, but we should at least have such considerations in mind as we tackle the nearer term mechanical problems.

Losing Money on Volume

A few weeks ago I suggested that
at least relative to the Harry Potter books of the world, there's a bit less discounting competition on the Long Tail which should at least partially offset higher costs associated with lower volumes. With purely digital goods, the Long Tail comes even closer to pure gravy.

It turns out that for the blockbuster hits, discounting can hit profitability hard. That's because a "big box" reseller like Best Buy uses hot new DVD releases as loss leaders to bring customers into the store. It may, in fact, sell new DVDs below cost for a time. This practice presumably works out for Best Buy (otherwise one supposes they wouldn't do it) because enough people who come into Best Buy to buy the latest Star Wars DVD also buy another non-sale DVD, a CD, or an ink cartridge. However, in the process, Best Buy pretty much destroys (or at least greatly reduces) the profitability of new blockbuster hits for everyone else in the process. Stores, both online and bricks-and-mortar, that drag less other business along in the wake of the bestseller can end up making relatively little money on their highest volume titles as a result. By contrast, smaller titles may have somewhat higher stocking and inventory costs, but tend not to be exposed to loss leader discount pressures from the big retailers.

Wednesday, August 17, 2005

Code From Books

Dan Bricklin's latest Software Licensing podcast is with Tim O'Reilly, founder and CEO of O'Reilly media. A chunk of it deals with a topic that was once very relevant to me That's the question of what's allowable use for the code in a programming book.
Tim O'Reilly's short answer on his Web site is this:
You can use and redistribute example code from our books for any non-commercial purpose (and most commercial purposes) as long as you acknowledge their source and authorship. The source of the code should be noted in any documentation as well as in the program code itself (as a comment).

He then goes into a bit more detail. The bottom line is that using the code as part of a larger project is generally OK so long as proper credit is given. Competitive uses-for example, publishing a CD with the code snippets or including them in another book are not. It's got some elements of the original BSD license (with advertising clause) with the important caveat that O'Reilly will frown upon uses that directly undercut the market value of its own (or the author's) products or services.

O'Reilly's position seems quite commonsensical. It also somewhat codifies what's long been common practice. My personal interest is that I once developed and sold a DOS file manager, Directory Freedom, which (to quote the docs):
originally grew out of a variety of programs which owe their "look and feel" to Michael Mefford's DR and CO utilities in PC Magazine Volume 6, #17 and #21. DF was most directly adapted from Peter Esherick's DC (Directory Control) version1.05B. Peter helped get DF started by making the source code for DC available to me and has also shared some fixes which he has made in subsequent revisions of his program.

This type of reuse and adaptation was fairly common on a small scale in the days before today's Open Source hit in a big way. (I'm talking roughly the mid-eighties to mid-nineties here.) Code in magazines like PC Magazine, Dr. Dobbs, and PC Techniques was certainly copyrighted, but it basically existed to sell the magazines, as opposed to being commercial software in its own right. In practice, it was generally assumed by just about everyone that most uses of the code were proper, even if good manners suggested giving credit where credit is due.

One interesting aspect is that this code, while not licensed as Open Source, can in practice be used more flexibly than true Open Source licensed under a "viral" license like the GPL which requires that derivative works be likelwise licensed under the GPL. In fact, my Directory Freedom could not have been a shareware product had the PC Magazine utilities been explicitly under the GPL rather than the implicit loosey-goosey de facto BSDish terms under which they were actually published and used.

Tuesday, August 16, 2005

NPR On Demand?

Podcasting and related technologies have been called the "End of Radio." I haven't been buying; at least for the most part. Given that podcasts have to be consumed in real time, a person can consume far less audio content than is the case with inherently skimmable written blogs. As a result, I've previously argued that:
as professional broadcasters like the BBC start putting content on the air, (e.g. "In Our Time") many--probably most--people will largely devote their limited audio-listen minutes to professionally-produced broadcasts. Call this podcasting if you like, but it's really just on demand radio as you can record more crudely today with software like Replay Radio.

Conspicuously absent from such on demand radio has been NPR which, like the BBC, would seem to have many programs tailor-made for the purpose-both highly topical and less so. (I'd argue that the current state of the technology in which several manual steps are needed to sync a program, for most people programs that can be listened to a week or a month after broadcast are preferable.)

Well, it looks as if NPR may not be totally clueless after all. It's apparently decided not to renew its contract with Audible, the maker of lame, proprietary audiobooks and the like. (In all fairness, Audible was long about the only game in town.)
As we formulate a more comprehensive strategy, we chose not to renew our agreement with Audible when it recently expired. We are now developing a new strategy for making NPR content downloadable and portable. Once the plan is finalized, we will announce it publicly.

via O'Reilly Radar

Dan Bricklin on Software patents

Dan Bricklin weighs in on software patents in this post.

Dan has some history here. As the creator of VisiCalc, he's previously written about why he didn't patent that software. His latest rejoinder is in reaction to:
Russ Krojec's blog entry "What if VisiCalc was Patented?" Russ argues that since VisiCalc wasn't patented competitors found it "...safer to copy the currently winning formula and avoid having to innovate. In this case, the lack of patents brought innovation to a standstill and we are all running spreadsheet programs that still operate like 25 year old software."

Dan correctly points out that Russ' position just doesn't square with historical facts--and demonstrates thereby the value of having an understanding of history. Back in the DOS days, I clearly remember any number of programs that were perhaps inspired by the spreadsheet metaphor but deliberately took different paths. There's some discussion of different approaches and programs here. A number of the alternatives were more explicitly multi-dimensional or iterative or capable of solving more flexibly-designed equations than conventional spreadsheets, but none were particular successes. Products that were very direct VisiCalc successors (basically Lotus 1-2-3 and then Excel) ruled instead, but it wasn't because there weren't alternatives. How come?

One reason that I'm pretty confident in giving is that, as the PC became more widely used outside of hobbyists and specialists, this forced a certain regularization of applications, for lack of a better term. Suddenly, the computer unsavvy needed to use these apps. A whole ecosystem of specialized training classes, books, and support systems sprung up. This tended to marginalize mainstream software that broke with established models. It was hard enough to teach people new command codes (which were highly irregular in the days before Windows), much less radically different mental models for how softrware worked.

Another reason is a bit more philosophical and I'm correspondingly less sure about it. But, perhaps the spreadsheet was just a metaphor and model that really worked and connnected with people--and the alternatives were just more copmplex variations on a theme that generally detracted rather than improved. In fact, most of the "innovation" around spreadsheet replacements has since been replicated in mainstream spreadsheets (e.g. for multi-dimensional, think pivot tables). And a tiny percentage of spreadsheeters use any of these capabilities-at least on a day-in, day-out basis. Furthermore, like word processors, the spreadsheet had a familiar physical analog--the accountant's ruled sheet.

I'm not convinced that software patents are bad in toto, however flawed the current system may be. But to say that more patents would have spurred greater invention in desktop productivity software just doesn't have a historical basis. And, indeed, at least with the reality of today's overly broad patents, it seems likely that just the opposite would have been the case with every remotely-related product litigated.

Thursday, August 11, 2005

The Web 2.0 Debate

Overheated blog conversations often get more wrapped up in the terminology than the "thing" itself. I've written about past transgressions like folksonomies. Tim Bray tackles the debate around Web2.0. Spot on. There's meat here but lose the freekin' buzzwords.

The Four Seasons and Generic Luxury

I like reading reading View of the Wing, even if (fortunately from my perspective) I'm not a serious enough road warrior to appreciate or take advantage of a lot of my advice. Sometimes I have to laugh out loud though:
Clearly this is a Four Seasons, but which one? In their zeal to determine what Thomas Jefferson’s bedroom might have looked like if he had had electricity and modern plumbing, Four Seasons has stumbled into the sort of routinized design philosophy embraced by mid-market chains with out the wherewithal to spend tens of millions building a hotel. I can only assume that the Four Seasons interior design team was let go in a corporate downsizing and their last cruel act was to commit the company to a 20 year supply of fake chesterfield TV cabinets fitted with mini-bars.

To the list of really good Four Seasons properties, I'd add The Olympic in Seattle. (Especially nice after a week of climbing or backpacking. The last time I was there, I drove up with this dust-covered rental car. I'm sure it took all the doorman's poise to not back away and avoid messing up his nice uniform.

Risks of Blogging

There's a good interview with Sun's Tim Bray, Director of Web Technologies, and Simon Phipps, Chief Technology Evangelist (Love that title!). Tim says in response to a question about the risks of blogging:
Yes, there are potential risks. But at the moment, a year into this, I would say that we are seeing almost all reward and no downside, so whatever potential risks there are, none of them have come forth yet. As one of our smart legal staff pointed out when we were starting to work on the policies, if we had come to a lawyer 15 years ago and asked him what he thought about e-mails he would have been horrified at the idea!

Many in the computer industry have been using email for longer than that; when I joined Data General almost 20 years ago, CEO (Comprehensive Electronic Office--a minicomputer-based integrated email/word processing/calendaring package) was already an ingrained part of the culture. However, in my previous job in the oil drilling business, there was no email and every memo-which is to say pretty much all written communications--had to be signed off by the appropriate level of authority. So, I certainly agree with the sense of Tim's comments. There's been a lot of loosening up just about everywhere. The convenience of email is too great and it basically doesn't work if it has to always go through channels.

That said, a lot of these discussions take place in the context of the technology industry and the Coasts. This SF Chronicle article, for example, is generally very upbeat--but most of the cases that it discusses are in the Valley where companies are often looser than elsewhere. I'm not about to argue a "blogs are risky" meme, but counsel that the standards of the tech elite don't necessarily applky elsewhere.

Thursday, August 04, 2005

The Long Tail For Quants

As regular readers know, I studiously avoid a lot of what I consider to be "blogosphere" hype - including the term blogosphere and such favorites as folksonomies and tags. However, the "Long Tail" remains a particularly powerful concept for me, in part because it ties directly into the business models of the likes of Netflix and Amazon rather than just a generic "linkiness is good." Yes, the partilly ego-boo-driven reviews on IMDB, and Amaon, and Netflix are part of what make the Long Tail possible - as Irving Wlawdawsky-Berger describes in the comments to this post - but it's the sales themselves that are the Long Tail.

That makes data like this which quantizes the Long Tail particularly valuable. Individual data points that are arrived at in indirect ways are always a bit suspect, but there's an increasing body of research that indicates a Long Tail at online retailers like Amazon - in this case meaning titles ranking below the top 100K (or roughly comparable to large brick-and-mortar inventory - of in the vicinity of a quarter to a third of total sales. The latest data corrects earlier figures which indicated a possibly (much) higher number, but even a quarter of Amazon sales is still a substantial number if they can be delivered with modest incremental per-unit cost. (I also suspect that, at least relative to the Harry Potter books of the world, thee's a bit less discounting competition on the Long Tail which should at least partially offset higher costs associated with lower volumes. With purely digital goods, the Long Tail comes even closer to pure gravy.)

Wednesday, August 03, 2005

10 Years That Changed the World

It's a bit ironic that, these days, Wired remains one of the few magazines that I still receive in dead tree form. I like the mix of story and item content and length. And its sense of style helps keep me from just going online. This month's cover story, 10 Years That Changed the World, is written by Kevin Kelly who helped found Wired and was its first executive editor. It's one of the best single pieces that I've read about the years since Netscape's IPO, in part because Kelly crystallizes certain underpinning concepts and dynamics of the Internet and the Web with considerable clarity.

For example, on the way that the Web turned the creation of content on its head:
Problem was, content was expensive to produce, and 5,000 channels of it would be 5,000 times as costly. No company was rich enough, no industry large enough, to carry off such an enterprise. The great telecom companies, which were supposed to wire up the digital revolution, were paralyzed by the uncertainties of funding the Net...Netscape's public offering took off, and in a blink a world of DIY possibilities was born. Suddenly it became clear that ordinary people could create material anyone with a connection could view. The burgeoning online audience no longer needed ABC for content. Netscape's stock peaked at $75 on its first day of trading, and the world gasped in awe. Was this insanity, or the start of something new?

And he has some marvelous turns of phrase:
But if we have learned anything in the past decade, it is the plausibility of the impossible.

For an industry so accustomed to hype that almost every announcement is and was a bew paradigm and a world-changing event, this is why the mainstream Internet and the Web caught many of us unawares. So much seemed impossible. For lack of a better hook, the Netscape IPO was a "Day the Universe Changed" to use the title of James Burke's old BBC series.

Tuesday, July 26, 2005

Wikipedia Data and Anecdotes

I like both data (that is generalized data from lots of people) and anecdotes (specific data from a few). Her's a bit of both about a topic that I follow with interest-Wikipedia. Why the interest? Well, for one thing, I find it a personally useful resource. For another, I consider it a bit of a bellwether of interactive collaboration.

Here's the home page for the stats. Diving down a bit, it looks like there might be at least some suggestion that the rate of new articles generation is slowing (at least in English) although I'd hesitate to draw any real conclusions before seeing more months of data.

On the anecdotal side, Tim Bray points some of the usual problems:
Dave Winer's right, the Wikipedia's article on RSS is a crock. Dave's gripe is that it's "highly political", mine is that it's just wrong: for example, the introductory bit suggests that full-content feeds are impossible. Also, it's badly-organized. Dave's problem is going to be harder to address because RSS itself is highly political; but at least the political narrative should be coherent. Anyhow, it would be nice if someone level-headed were to take responsibility for it. I currently ride herd on two or three other articles and that's all my Wikipedia cycles. It's not as hard as you might think, and here's why: the kinds of people who want to put stupid, irrelevant, badly-written junk in the Wikipedia in my experience are easily discouraged. Just hang in, keep on fixing things they break and explaining why in a calm tone of voice on the Discussion page, and pretty soon they go away.

I'm not sure I fully share Tim's "it will work out" faith. That said, I think it reinforces the view of Wikipedia as a valuable resource-but not a totally dependable one.

Wednesday, July 20, 2005

Some More on Digital Archives

Although not directly on the point to the current Internet Archive case, this piecewritten by Adam Mathes, Graduate School of Library and Information Science, University of Illinois Urbana-Champaign has some interesting discussion about archiving software programs for preservations purposes--as well as current exemptions to the DMCA that aid in that effort.
In addition to processing issues, this brings up some of the legal issues involved in the collection. The mere act of copying these digital works, especially for the eventual purpose of enabling access on a different hardware platform, should arguably be considered a fair use. However, if these disks have "copy protection" schemes, even outdated ones that can be bypassed, care must be used to make sure the collection does not run afoul of the Digital Millennium Copyright Act (DMCA). Although recently Archive.org was given an exemption for particularly this reason it presents a considerable barrier and must be dealt with. (1) It may be helpful to amass multiple copies of the works in many formats to further bolster the legal backing to shift and archive the materials.

See also this reference:"Internet Archive Gets DMCA Exemption To Help Archive Vintage Software." Internet Archive. 2003. February 23, 2004.

Friday, July 15, 2005

The Internet Archive continued

In response to my last post, fellow analyst James Governor speculates that perhaps libraries like the Library of Congress might not be one possible analogy.

If, for purposes of argument, we consider the Internet Archive a library, that could well grant them some exemptions to the rights conferred to content creators by copyright law. Consider this from Section 108 of the U.S. Copyright Act:
§ 108. Limitations on exclusive rights: Reproduction by libraries and archives
(a) Except as otherwise provided in this title and notwithstanding the provisions of section 106, it is not an infringement of copyright for a library or archives, or any of its employees acting within the scope of their employment, to reproduce no more than one copy or phonorecord of a work, except as provided in subsections (b) and (c), or to distribute such copy or phonorecord, under the conditions specified by this section, if—
(1) the reproduction or distribution is made without any purpose of direct or indirect commercial advantage;
(2) the collections of the library or archives are
(i) open to the public, or
(ii) available not only to researchers affiliated with the library or archives or with the institution of which it is a part, but also to other persons doing research in a specialized field; and
(3) the reproduction or distribution of the work includes a notice of copyright that appears on the copy or phonorecord that is reproduced under the provisions of this section, or includes a legend stating that the work may be protected by copyright if no such notice can be found on the copy or phonorecord that is reproduced under the provisions of this section.

See also here and here for various pointers about copyright law as it applies to libraries. (By the way, nothing in there about the content creator having to give permission or being able to withdraw permission--count another strike against robots.txt having any significance in this case.

However, as with many things digital, I'm still a bit suspicious of physical world analogs. Not just because of the different nature of the media, but also the nature of the institution. We can all agree that the Library of Congress is a Library and that Widener at Harvard is a library and even little Thayer Library in my town of Lancaster is a library. And it's perhaps not too much of a stretch to see the Internet Archive as a form of library. But what if I were to declare my own little web site a library and compile Dilbert cartoons there? (I picked this as an example of content that's posted publicly but only for a limited time.) My guess is that Scott Adams might not approve. Yet, what makes the Internet Archive different in any fundamental way?

Thursday, July 14, 2005

Thoughts on The Wayback Machine Kerfuffle

The Internet Archive a.k.a. Wayback Machine is being sued by a firm called Healthcare Advocates for storing copies of old web pages. (See Good Morning Silicon Valley, for example.) These archived pages are causing the company heartburn in a separate trademank dispute so it's unhappy. Further, for some reason, the pages were allegedly stored in spite of being flagged with a "robots.txt" file to not be archived, cached, spidered, etc.

The case has generated the predictable throwing up of hands in disgust throughtout the online world. As Good Morning Silicon Valley's John Paczkowski succinctly puts it: "Uh, you published that information to a public medium ..." Now I'm certainly sympathetic with the Internet Archive here. At some level, the archiving and caching of publicly-displayed web pages seems almost part of the fabric of the Web and the way it works. However, I'm less convinced than some others that this is Much Ado About Nothing. I preface the following comments and observations with a standard "I Am Not a Lawyer"--and would welcome any on point case law that might be relevant here.

I think we can all stipulate that web pages and such are copyrighted material and freely displaying them to the public doesn't reduce or eliminate that copyright in any way.

I do agree with John that the robots.txt angle seems a wit wacky.
Why? The robots.txt protocol is purely advisory. It has no legal bearing whatsoever. "Robots.txt is a voluntary mechanism," said Martijn Koster, a Dutch software engineer and the author of a comprehensive tutorial on the robots.txt convention (robotstxt.org). "It is designed to let Web site owners communicate their wishes to cooperating robots. Robots can ignore robots.txt."

Ignoring robots.txt may be bad manners, but it's hard to see the legal significance. (There are perhaps analogs in physical trespass laws--posting your property and the like--but my understanding is that the details of such as typically goverened by explicit state and local laws.)

However--and here I perhaps stray into less charted territory--what exactly gives the permission to copy and archive web sites anyway? Certainly, there's no explicit permission like a negative robots.txt file that affirmatively gives the right to replicate, store, transmit, archive, etc. web pages. I suppose the theory is that there is some sort of implicit permission based on custom and social contract. Which seems a rather loosey-goosey state of affairs.

I can't think of any really good analogs here. Yes, I can record TV and radio--but only for my personal use. It's quite well established I can't put those recordings on a server for all to access. Usenet postings might be the most analagous situation; they're now archived as Google Groups and in more fragmentary form elsewhere. However, as far as I know, the legal status of Usenet and other types of online postings doesn't have much case law underpinning it. Furthermore, I think one could easily argue that such postings have a more explicit element of transmission of content out into the world--with the full knowledge that said content will be forwarded and stored for at least some interval--than Web pages which reside on a controlled site.

Nor can I see the exemplary historical service that the Internet Archive is providing with its activities having any bearing. "Preservation of the past" may be a social good, but it's got little to do with copyright law. After all, Abandonware has the same legal status as any other warez in the absence of the copyright owner's explicit permission to release it into the wild.

From where I sit, robots.txt certainly seems like a red herring in this case--given the lack of laws compelling its observence. But there's a much larger issue of caching and archive that seems to rest on very sandy foundations.

Monday, July 11, 2005

Podcasting Redux

Podcasting continues to be a beloved trend of the plugged-in elite. I've commented rather dismissively about it before. Since then I've spent more time checking out the various podcasting options--both software and content. Have I revised my opinion? Not really, I still think that there are some fundamental reasons why podcasting won't have the impact of text-based RSS. Which is not to say that podcasting doesn't have merits within a limited scope.

Chris Anderson at The Long Tail gives three reasons why podcasts aren't a big deal (yet).
  1. They don't have internal permalinks to section and subjects, so they don't get much link-love.

  2. They aren't searchable. How hard would it be for some service to run podcasts through a quick-n-dirty voice recognition program to autogenerate transcripts? They don't need to be exactly right; 80% accurate search is better than the 0% we've got now.

  3. They're meant to be consumed linearly, and pretty much at the (agonizingly slow and amateurish) pace they were created. Who, aside from trapped commuters, has time for that?

These comport with my impressions, but it's the linear consumption that's the real killer. David Winer, who wrote the RSS 2.0 specification, described the web as a "skimming" medium on this Steve Gillmor podcast and you can't really skim audio feeds effectively. As a result, you end up selecting a few favorite programs that you might listen to during audio-friendly periods--which is to say, typically driving in the car. And, guess what, as professional broadcasters like the BBC start putting content on the air, (e.g. "In Our Time") many--probably most--people will largely devote their limited audio-listen minutes to professionally-produced broadcasts. Call this podcasting if you like, but it's really just on demand radio as you can record more crudely today with software like Replay Radio. (By the way, NPR, get with the program!)

So when are podcasts good? I can think of a few things, both based on my personal experiences and things I've read about.

Business uses--for example, a weekly "broadcast" to a sales force. This is sort of a special case of the "content to listen to while commuting/driving.
Certainly, interesting interviews with the sort of specialists which don't make it onto mainstream broadcasts in any depth are interesting. Steve Gillmore and Dan Bricklin's podcasts are good examples of this. Talks at conferences are another good example--as at IT Conversations.

Thus, I certainly don't argue that podcasting is "bad" or useless. I like on demand listening and RSS syndication provides a handy mechanism to more easily (if hardly automagically) get updated audio content from favored sources to my car. But I'll continue to argue that it remains a largely peripheral trend rather than the "end of radio."