Showing posts with label Google Book Search. Show all posts
Showing posts with label Google Book Search. Show all posts

Friday, May 27, 2011

The Google Books Mess

There were a couple of tell-tale signs last week that Google may be having some pain and problems with its vastly ambitious Google Books project. First, was the news that Google was pulling the plug on its corresponding, open-ended, plan to scan and database masses of historic newspaper archives. Second a report that Google was diverting all its programmers from its eBookstore and perhaps not vigorously pursuing plans selling eBooks.

The problem that Google has, is that there was huge momentum within the company towards its grandiose plan for a comprehensive universal digital library and this vision, with its accompanying class action settlement [ASA or amended settlement agreement] was decisively stopped in March by the opinion of Judge Chin (USDC SDNY)


While the digitization of books and the creation of a
universal digital library would benefit many, the ASA would
simply go too far. It would permit this class action - - which
was brought against defendant Google Inc. ("Google") to challenge
its scanning of books and display of "snippets" for on-line
searching - - to implement a forward-looking business arrangement
that would grant Google significant rights to exploit entire
books, without permission of the copyright owners. Indeed, the
ASA would give Google a significant advantage over competitors,
rewarding it for engaging in wholesale copying of copyrighted
works without permission, while releasing claims well beyond
those presented in the case. (Opinion 22 March 2011)

Chin's decision is styled an opinion, and it might yet be appealed or revised, but most observers would tell you that it has pretty well stopped the Google project in its tracks.

Google has got a lot of figuring out to do:

  1. Google is not out of its legal woes, although such a rich and powerful company can probably stall or out-manoeuvre the authors and publishers who are parties to the original suite in the USA. Yet Google will need some resolution to the case or it risks enormous damages for breach of copyright ($3.6 trillion according to one scholar).
  2. Google will not find it straightforward to avoid legal actions in other jurisdictions. It has ongoing legal woes in France, and if some French publishers win substantial damages, many others will charge through these same gates.
  3. Google is continuing to scan without permission millions of works which are not out of copyright on behalf of its library partners. So the liabilities grow.
  4. Google will be required to deliver digital library services to some of its core collaborating libraries. The libraries of Michigan and Stanford in particular. To the extent that these services depend on copyright works digitized without permission Google remains at significant risk.
  5. There will be increasing concern about advantages that may accrue to Google from the works that it has already scanned and databased, and which it may use in ways impervious and invisible to external actors. Perhaps Google will gain enormous advantage in the fields of search, automated translation and semantic technologies through private access to vast amounts of unregistered, unlicensed, copyright material. That putative advantage creates legal risks for Google from competitors and regulators.
  6. Without a recognized and legitimized settlement Google cannot deliver services of general public benefit, and at some point Google loses good will. Without a settlement Google cannot even be generous.
  7. Google has plenty of agreements with publishers and authors for the distribution, display and potential licensing of millions of copyright works. So it could be an active participant in the eBooks market, but it has been strangely hesitant and stuttering in recent years about its commercial activities. Almost certainly because Google's lawyers are anxious about the way such commercial exploitation may play against the unresolved matters in dispute. If Google carries on havering it will lose its opportunity in the digital books market, much as it appears to be losing its opportunity in the market for digital music.
I am not sure that Google has an easy way of stepping out of this mess. But it needs to find, or create through disruptive action, some solution.

The original goal of a universal library designed, built and maintained by a single technical player was hubristic and naive, driven by the enthusiasm and commitment of the founders (Page in particular who felt that he owed a debt to his alma mater, the University of Michigan). Google's best hope now would be to distance its involvement from the prospect of private gain and to place all works not public domain, and not explicitly licensed to Google, in the sole care and control of the public academic institutions from which the original works were taken, and to renounce any commercial advantage through its involvement in converting 'orphan' works. Google will have to pay the authors and publishers something (if only to cover some of the legal bills, that will otherwise be pursued to the bitter end on a contingency basis by the other side), it can afford to finance the first blocks of a Rights Registry, but it should be more open and more public, more consultative, in part foundation funded, than the original design. Google does not need and should not look for special advantages on rights and forward-looking business models. If Google were to do that it could help to promote the cause of orphan works legislation in a disinterested manner. Google needs to get legitimate, beyond all shadow of doubt, fast.

Google often likes to play the 'open' card, but it has been far too closed and 'private' over its books project. It needs to rethink the game-plan and its style of involvement. That way it will retain the good will of the library community and the reading public. By being highly generous and public spirited it looks after the interests of its shareholders also. Page is now CEO and he may need to bite on the books bullet and own up to a change of course, only be being much more open and generous can Google hope to make something like the Google Books project a reality.

Tuesday, November 30, 2010

Google goes into Culture Commerce

The rumour mill has it that Google will launch a Chrome netbook, cloud-based, computer before Christmas or early in the New Year. When you put this rumour alongside the others coming from the Googleplex you get an interesting picture

  1. Is Google going to buy a big package of movie rights? Is that why it has hired the former Netflix executive George Kynci?
  2. Google is possibly quite close to signing a deal with the major music labels for its cloud-based music-streaming service.
  3. For a couple of years, Google has been apparently on the brink of releasing a digital books service in collaboration with book publishers. Most recently Dan Clancy told us that Google Editions will be launching very soon ("très bientôt")
The rumours about the Chrome netbook suggest that its really all about the web, cloud-based productivity and web browsing, but if its launch is accompanied by, or closely followed by, a Google distribution and e-commerce solution for books, films and music, the market place for publishers and entertainment companies may change very fast. Google will be a formidable competitor if it becomes an information publisher and an e-commerce platform for film, music and books. Competitor primarily for Apple and Amazon, Google may well be seen as more of a 'friend', because more collaborative and more open than either Amazon or Apple, by the big incumbent publishers and media groups. Knowing, as we do, the way Google works (quiet launches, 'beta' services, and something of a scatter gun approach) I think its unlikely that Google will launch a fully fledged, cloud-based, Chrome-machine, with a multi-channel, multi-media dashboard in place in the first quarter of next year. It is surely more likely that this hardware platform and this constellation of media services will each emerge in their own good time. But if the plans come off and Google has these publishing partnerships in good order, it is highly likely that Google will be selling a lot of consumer products next Christmas. And I do not meant via Groupon; a commercial solution that can stream all kinds of media stuff from your locker in the cloud to Android and Chrome platforms, will be a dazzling consumer attraction

Saturday, July 24, 2010

In Praise of Not-Reading

Reading is, in these days, an over-rated activity.

Most of what is most important about books is now about not-reading them. I was reminded of this deep but counter-intuitive truth by a blogger (An American Editor) recently complaining that his To Be Read pile (TBR Pile) was getting unmanageable because it was full of ebooks that often cost nothing and were without physical presence.

Which brings us to the special problem of ebooks. Yes, ebooks are a special problem because they take up virtually no space — just a bunch of bits and bytes, digits if you will, on a disk that can store gigabytes of digits. And so that TBR pool steadily grows. I looked this morning and I have more than 300 TBR ebooks, and that pile keeps growing. Acquiring Books for the TBR Pile: The Special Case of eBooks
American Editor is here admitting to a very old-fashioned mistake. He has not caught up with the twenty-first century. Books are now not really for reading -- or to be more accurate -- they are only occasionally, under the most special circumstances, for reading. Publishers are partly to blame for this (culpable, since all publishers, especially of newspapers and magazines, know that their profits are entirely dependent on selling stuff that the customers do not read) and digital book experts would be much more on the button if they spent less time fretting about 'reading'. And part of the problem is that the digital experts operate with a vastly over-simple model of what reading is. The conventional wisdom is that proper reading (sometimes called 'long form reading' -- a ghastly phrase for a dubious concept) is the measurable phase in which you open all the pages of the book and look at them, the hours and minutes through which a book, conveyor-like, passes, between the moment that you bought it and the moment that you shelve it in your personal library, never to be looked at again. Incidentally, 're-reading' is a much more interesting concept than mere 'reading', but we note that in passing and may return to the topic on another occasion (you will have observed that you can do that with writing as with reading). This Taylorean model of, conveyor-like, reading predicates that in serious reading our eyes scan more or less consecutively the whole book from page 1 to page umpteen. Efficiently and quickly. The time and motion expert holding a stop-watch, just as Google analytics calibrates our use of the Google library. As though reading a book might not actually comprise understanding it, or failing to enjoy it, or realising pretty instantly that it was not worth reading at all. Under any circumstances.

Of course, reading has always been, but is becoming steadily more, episodic; very little of our reading is like this conventional model of continuous reading, and most of us who now work in intellectual or bureaucratic activities which involve web-based reading, spend a lot of time, yes reading, in ways which are not at all like the way you first read and enjoyed Babar, P.G. Wodehouse or Jane Austen. You see, we spend a lot of our time and energy deciding what not to read. And these decisions matter. Possibly even as much as enjoying Babar, or re-reading Jane Austen.

Our understanding of digital books would be much better if we spent less time wondering about how we might read them, and a lot more time thinking about the ways in which we may use them without necessarily, or even at all, reading them. For certainly, and beyond all doubt, when there are 20 million books in Google Books Search we will not seriously, continuously, read more than the tiniest fraction of them. There are a lot of things that we need to do with books and it is not at all clear to me that we have a framework in which these activities can take place with digital books, half as effectively as with print books. For example, we need to be able to:
  • search them (that activity appears to be brilliantly covered by the already mentioned Google Books Search)
  • provide access to them (possibly well covered, in the USA, by the afore-mentioned)
  • buy them
  • listen to them
  • lend them to a friend or a colleague
  • translate them (well)
  • quote from them
  • (ideally) cite them when we quote them
  • non-consumptively compute them (we none of us know quite what that might involve)
These are all important points, but I will admit to playing a rhetorical trick with this list, my bullet points, and the bold face. The key point about the list is the recurrent 'them'. There are so many things that we need to do with books aside from, and apart from reading them. The key thing about digital books is that we need them. We need digital books to be the 'object' of all these newly digital verbs and activities. Digital books need to be as versatile and as 'real' as physical books in all these ways, even though they are now becoming entirely virtual and insubstantial. The big challenge that Google, Apple and Amazon have yet to project is that books themselves are becoming networked. And the Google, Apple and Amazon models of network usage will inevitably fail if they are not truly book-centric.

May I recommend (unreservedly, though I have forgotten most of it, and disagreed with much of it) Pierre Bayard's How to Talk About Books You Have Not Read. Which, in case you mistakenly decide not to read it, has many reviews here.

Wednesday, July 14, 2010

Google Books Search over the Summer

Judge Chin is still considering his decision in the case of Google ..... His ruling may come this week, next week, or in a few months. Only he and his team have a good idea of that. Meanwhile Professor Pam Samuelson has produced a very thorough, balanced, somewhat critical review of the proposed Settlement and of Google's efforts in a 60 page paper for the Minnesota Law Review. If you haven't been following GBS too closely, this is an excellent place to get an insightful review and summary of what has been going on. If, like several hundred lawyers and digital library experts you have been following GBS too closely for years, you will already have read her piece and it will have reminded you of stuff that you had forgotten. Her conclusion:

The future of public access to the cultural heritage of humankind embodied in books is too important to leave in the hands of one company and one registry that will have a de facto monopoly over a huge corpus of digital books and rights in them.
Google has yet to accept that its creation of this substantial public good brings with it public trust responsibilities that go well beyond its corporate slogan about not being evil. Google Books Search and the Future of Books in Cyberspace


I have been a 'qualified' supporter of Google Books Search from the beginning. The qualifications are coming more to the fore. Whatever Judge Chin decides, we can be sure that Google Books Search is going to be mired in legal complexities for years to come. The international ramifications of the venture are hopeless and will sap energy and innovation. Google Books Search, if it is approved, will work badly and too patchily for European literature and libraries, and it will be especially rough and unsatisfactory for British literature, libraries and universities. It will be a mess of conflicting and irresolvable copyright regimes for years. Google itself seems to find it hard to innovate or roll out new services. A clean and direct implementation of Google Editions has been 'promised' for this summer, or this year, but it has been promised before. Several times. No doubt part of the reason that it is being held up is that its roll out may have unpredictable or unwelcome legal consequences (or unwelcome splash-back from the court of public opinion). Google Editions when it comes should be a very useful and popular service, but Google have to get it out of the door before it can properly grow and bed itself into the array of digital books that is now mushrooming.

Pamela Samuelson points to the lack of substance in Google's mantra 'we will not be evil'; but its arguable that Google has failed in a more fundamental and troubling way. It has failed to sacrifice the idols of its founders; it has failed in corporate governance. Page and Brin met and worked together in a project for digital libraries. The Google Books Search proposition was clearly motivated in part by Page's promise to digitize the libraries of his alma mater the University of Michigan. The two big leaps in the Google Books enterprise, were first to dream of digitizing millions of books in one universal searchable index (the original project, defended by an appeal to 'fair use' and the transformative effect of a large database of books) and then secondly to aim for a commercial settlement to the 'class action' suite, through which Google, the authors and publishers would effectively enclose, exploit and privatize millions of copyrights for which they cannot claim ownership. I suspect that the Google Books project, and especially the Library component, has always been too close to the goals and aspirations of Google's young founders. The big and aggressive steps that the company has taken to stake out its claims have been part of the founding DNA, the dreams that brought Brin and Page together. The third 'founder', Erik Schmidt joined the company in 2001. At about that time the initial steps for the Google Books enterprise may have been taken, perhaps Schmidt may have been too much the 'new boy' to question the goals of the original founders. Schmidt should have spotted that there were copyright problems, he should have noticed that there were at least issues of politesse involved in digitizing and then using for profit, stuff that did not belong to Google or to the Universities with which they worked. I bet that he has since then wished that the aims of the Google Books undertaking were more clearly understood within and without the company. And more cautiously and generously drawn. At some point Google has to take a much more humble view of its role, and at that point things might start to work in its favour.

Wednesday, June 23, 2010

Nominalism, Realism and Digital Books

There is a quasi-philosophical disagreement underlying the steady digitisation of literature. A radical disagreement about what digital books really are. In a strange manner this dispute parallels the controversy between nominalists and realists in medieval scholastic philosophy about the status of universals (properties, numbers, virtues etc). Texts in the twenty-first century take the place of properties in the fourteenth. Are books more than texts, are texts more than digital file formats? Are these abstract concepts: "red", "thirteen", "chastity" real entities or are they simply instances and constructs based on our experience of coloured objects, groups of cakes and the people we meet? The nominalists denied the reality of these abstractions and the realists retaliated. Blood was shed. Now we find the digerati divided over the question whether a book is really more than a text; since the ebook nominalists, finding meaning in sentences and texts and not much else, would be be happy with books digitised as texts (preferably in the ePub standard) and the realists say that a book is much more than its text and that the pagination matters, the layout matters, the entirety of the book matters, the references and the citations to the book matter, and of course the illustrations matter; therefore in pursuit of realism, digital systems should virtualise the whole book, not just its text. While Project Gutenberg is at one end of the spectrum (nominalists carefully proof-reading and hunkered down in ASCII or XML), Google with its Book Search digitisation project is at the realist edge -- some would say 'hyper-realist' in its acceptance of blank endpapers and leather bindings, all part of the 'real book' as represented in a Google database. Google probably would, if it could, encode the sensory aura of historic books, the vinegar that comes from cholera-touched books.

But the modern predicament over books as texts, or books as virtual objects, is complicated by a dimension of uncertainty over the appropriateness of treating books as a collective whole as parts of a library and a literature, or of digitizing them one at a time as individual atoms; perhaps, in some cases, with unique and unusual bibliographic or structural properties. Digital nominalists are governed by a standard of simplicity and hold that a text is a text, is a text. But some nominalists are atomic, whilst others favour a more holistic and uniform approach, in the interests of creating a library or a reading platform, in which all books can be searched and individual books isolated as readable downloads. Correspondingly, on the 'realist' side sits Google with its holistic and scalable method, Google's whole strategy for digitizing books has been based on an assumption that all books should be accessed, searched and distributed through a single canonical library. Amazon, which has in most respects taken a 'nominalist' approach to the distribution of eBooks, it doesn't do ePub but its proprietary format 'AZW' is just another ASCII encoding standard, has also embraced a 'holistic' attitude. Amazon offers its customers global searching of the Amazon archive and encourages users to build up a collection, a mini-library of eBooks on their Kindle. Amazon, just as much as Google, would like to have a scalable and totalitarian solution to the whole of published literature. All Kindle titles are atoms in the same collective, distinguished by the fact that they can be sampled or acquired from Amazon, one book at a time as the consumer dictates and purchases.

Perhaps a diagram will aid the explanation of this digital predicament for computerized books:
















What approach to digital books heads to the top left quadrant of this matrix? Why, apps of course. If we think of books as apps, they do have a reality and concreteness which exceeds the flatness of the mere ASCII text, but the book as app is also highly individual. Perhaps a paradigm of this approach is the Atomic Antelope Alice App which has caused such a stir. The inventors of the Alice book found an intriguing way of applying 'physics', acceleration and gravity, to the Teniel illustrations in the Alice book. This is obviously a very special and un-generalisable treatment of a classic work, but as an app it is a brilliant proposition. Apps can afford to be sui generis since they stand on their own, and if this gives cataloguers and librarians a headache, too bad. The Exact Editions book and magazine apps are also in this segment of the diagram. It is, I would suggest, the potential inventiveness and the unpredictable future of the book as an application that has the most intriguing potential for the future of digital books (and libraries). If digital books do something completely novel and free-standing, something unprecedented in the world of print books, it will be because they are also software applications and can in that way assume a digital reality which exceeds our expectations of the traditional text.



Thursday, June 17, 2010

Stone, scissors, paper: and the Digital Book Race

It sometimes seems that Google, Apple and Amazon are engaged in a three-way fight over the digital books space. They each have a very important area of strength: Google dominates search (and search-based advertising), Apple leads and designs the very best consumer devices, Amazon has amazing strength in consumer transactions (logistics). And the three fighters are subtly trying to manoeuvre the struggle into the terrain where their particular strengths dominate. So Google is building a massive library that will interact and benefit from Google's supremacy in search, and a lot of stuff should be free and any hardware device can access its service. Amazon is trying to build individual consumer accounts fed by their logistic strength and transactional breadth, obviously not limited to books, physical and digital, and Amazon care more about managing transactions in the consumer's account than they do about owning the device. Amazon only cares about consumers, so -- unlike Google -- they do not offer anything much to libraries. So Kindle books are readable on the iPad. Apple is trying to establish a superior hardware platform in which their range of interoperating devices cannot be matched: desktop, phone and tablet format working together. If Apple owns the superior hardware platform they will control vital pinch-points through their app store (so far limited to Apple hardware for apps and books). They are all fighting on several fronts at once. So Apple, is erecting a system of apps and phone-based demographics which will be immune to Google search. Google will be blocked in Apple's mobile domain, and Apple has built an e-commerce system that certainly rivals Amazon's, though it does not have anything like the breadth of the Amazon offering (yet). Amazon attempted to get into Apple's patch with a hardware device, the Kindle, which now appears to be completely outclassed by the iPad. Google does not want to be boxed-out by the Apple iOS, so it has launched its Android system which it hopes will attract the collective ingenuity of all the consumer electronics companies who are concerned that Apple might eat their lunch. With retaliatory ingenuity Apple is building an advertising system iAds which may seriously limit the Google advertising dominance. So the battle is three way and in deadly earnest.

Is it a game of stone, scissors and paper? Perhaps one being played in several dimensions. If so, Apple is the 'paper', they were always going to be on top of the rather dull 'Kindle' from Amazon: stone-coloured if not yet quite sunk. But Apple appears to be threatened by Google's incredible scissor-like sharpness in search. Perhaps the analogy breaks down with Amazon contra Google. Is there a way a battleground on which Amazon is beating Google? Amazon does appear to have the beating of Google in one area: third party, cloud-based, web services -- Amazon S3 etc. Perhaps this is the area in which Amazon might ultimately have the beating of Google's over-centralised approach to building a universal library. Amazon needs to build a digital book service which is more collaborative and de-centralised and which attracts libraries, authors and publishers as effectively as they have attracted book buyers.

Is there something that could change the dynamics of this three-way tussle? Lots of things. One development that would certainly change matters would be if Facebook entered the fray. Suppose that Facebook, live up to its name and did a deal with Amazon? We might then see an Amazon that really could outsmart Google in the relationships game. Cloud-based book networks.

Do you think that the Google Book Search judgement will come tomorrow? One of these days....

Tuesday, February 23, 2010

GBS and the Judgement of Solomon

The Google Books Search Settlement came to court last week for a Fairness Hearing. There is now a full transcript, but like many who follow the case closely, but not too closely (it could easily become an obsession), I have mainly relied on Twitter comments and the excellent blog of Professor James Grimmelmann, The Laboratorium, as a way of keeping in touch with what is going on. Grimmelmann and his students have also produced a fascinating, colourful and impressionistic report of the presentations and the behaviour of the actors on the day. This is highly recommended if you have an interest in how the case is developing. The whole process leaves me in considerable admiration for the American legal system -- astonishment even; though one knows that it is very possible that a perverse decision may be formulated in the end. And then dragged out and mangled with a decade-long process of delay and appeal all the way to the supreme court.

Admiration, that the process involves a Fairness Hearing -- a hearing where all parties are invited to present arguments for and against the 'fairness' of the proposed settlement. Fairness really is at issue and fairness should speak. Admiration for the remarkable skill and ingenuity of the critics and the proponents of the settlement. Admiration, really, that the Judge is clearly looking for a solution. At several points in the process he asks for help -- he especially wanted suggestions from the critics as to how the Settlement could be fixed. Grimmelmann notes in his commentary that critics who could not come up with ideas as to how to fix the problems they were focussing on, lost ground.

The Judge (Denny Chin) was not asking for outside help, but like Solomon he clearly needs it. Perhaps his case would work in a Biblical framework? The dispute has unnerving parallels with the one brought to Solomon, but there are also many additional complications: Rule 23 and the arcane process of American class actions, identical factual predicates, Firefighters and all (no, I don't know what Firefighters is about, but a lot hangs from its precedent). The crucial point is that this is once again a dispute about a child who should have a long and healthy future and there is a danger that it may be smothered or torn apart in his chambers. The orphan books should thrive! But there are too many jealous 'foster parents' and the judge will need a masterly stroke if he is to separate the shameful pretenders from the true mother. Is there scope for the judge to put the settlors to a Solomonic test? Two moments in the argument struck me as particularly crucial in this regard:

First: BONI (for the Authors' Guild) on orphans in dialogue with the Judge:
THE COURT: I think I agree with Mr. Katz and the government that if you give an opt-in, you would eliminate a lot of the objections.
MR. BONI: We would eliminate a lot of objections but we wouldn't have a settlement, and here's why. Number one, and most importantly for us, we will not -- we as class representatives --THE COURT: Well, I would assume -- before I said I would surmise. But I would surmise that Google wants the orphan books and that's what this is about -- (Transcript p138)

Second: DURIE (for Google) in dialogue with the Judge:
THE COURT: If Google had been digitizing entire books and not just making portions available but making the entire portions available and indeed selling them, would that be
something that Google would have tried to defend?
MS. DURIE: Selling the work, no. Making the entire work available, that is a more complicated question, in the following respect. We were giving an entire copy of the book to the library.....(Transcript p150)

Boni says that there wouldn't be a settlement if it had to be opt-in (presumably because Google would not work on that basis. But are we sure about that? They are working on an opt-in basis from now on). Durie, speaking for Google, concedes and volunteers that Google have been quite willing to give away entire copies of books in copyright (books that it did in no sense own). Google is not ungenerous. I think these positions conceal an axis on which Judge Chin may be able to turn the case with a judgment worthy of Solomon. Notice his comment to Boni: 'make it opt-in and you would eliminate a lot of the objections'. There are bluffs to be called. The parties should be forced to live with a purely opt-in solution, which incidentally keeps copyright the right way up, will keep Ursula le Guin, and the French and German governments happy; or (and at this point Judge Chin needs to stroke the handle of his sword, even test the mettle of the blade with his forefinger) Google must be much more generous with the copyrights it has opted from the orphans. Generous to the public domain and non-exclusive to its competitors.

With a crafty swipe of his rapier, Judge Chin should be able to pierce the settling parties apart on this issue. And put the whole thing back together in a more pleasing fashion. I am not sure whether or not copyright will still be the right way up. But we may hope!

Friday, February 05, 2010

What Will Google Do Next?

Is the DoJ filing, the nail in the coffin of the Google Books Settlement? James Grimmelmann, who has been following these things as closely as anybody, thinks that it is increasingly unlikely that the Settlement will be approved (which is how I interpret this tweet: "I'm updating my priors on the chances of the settlement's passage downward substantially." -- James, poor guy, has been reading too many briefs these last few months).

So, if the Settlement is not approved, what happens next? In fact, even if the Settlement is approved since it will be ensnared in appeals and delays for years yet to come we might well wish to know: what happens next? The most visible part of the grand Google library project, university-based subscriptions to most of 20th Century literature and published knowledge, modest royalty streams flowing to orphaned works, public access terminals in libraries etc, is stalled. The Books Right Registry may not come to pass. There are three good things that could and should nevertheless happen when Google finally washes its hands of the Settlement and shrugs off its law suits:

  1. There should be a way of delivering the original index-ing service that was the primary goal when the whole exercise began. Negotiating for that to emerge, may be a way of letting the Author's Guild and the APA off the nasty hook that they have constructed for themselves and Google. It would also, of course, be a way for Google to satisfy many of the obligations it has by now built up to its library partners.
  2. Google may finally get round to delivering Google Editions, about which there has been some talk, and more than a little rumour. But if it really is going to appear in the first half of 2010 it was time that it had a bit more visibility.
  3. Google should give greater prominence to Google Scholar which has been for too long a neglected but steadily useful aspect of the Google service
Whatever Judge Chin decides, Google has 10 million+ books scanned, databased, interpreted and searchable in its servers. It has had a good deal of encouragement and collaboration from the publishing industry and I can't see the publishers really wanting to pursue their original copyright infringement suit to the bitter end. Those books are going to be put to some use at some point. I bet there is intense internal debate at the Googleplex about what to do next. The critical mass of 10 million books will be part of the answer.

Friday, January 29, 2010

Too Many App Stores?

As is generally known the Apple App store handles payments for developers and some distribution functions in exchange for a 30% cut on the revenues obtained from the sale of those apps. 70% is passed along to the publishers and developers who make the apps. The Android App store is run along similar lines (30% to Google for handling the transaction, collecting the monies and maintaining the store front: 70% to the developer). Amazon last week announced its own App store, along with an invitation to developers to produce Apps for the Kindle platform. Co-incidentally it announced a new and improved deal for authors and publishers, whereby under certain not too onerous conditions the author/publisher will get 70% from the Amazon sale. It is fishy the way that these deals seem to cluster around a 30% cut. Is there something about the investment and infrastructure needed to set up and run an App store that dictates a 30:70 deal? Who is going to be first to blink, and move to 20:80?

Along with its brilliant iPad, Apple just announced an iBook application within the iTunes store. Edward Nawotka, an industry pundit wonders whether this presages Apple offering a direct route to publication for authors. Amazon has been cutting direct deals with authors (sidelining the publishers). Google is in the class-action settlement of the century in its effort to become the digital publisher or republisher of millions of out of print but in copyright books. It would seem that in there effort to establish dominant positions in the distribution of digital books these three great companies are moving back up the publishing chain in an effort to secure greater security of supply.

As well as trying to buy into a dominant supply position (it would be fascinating to see the details of the exclusive deal that Amazon struck with Rosetta for McEwan's backlist. What guarantees or minimums are in that package?), these great companies are also trying to muscle into the other guy's distribution channel. Both Google and Amazon have built apps for the iPhone which allow users of the Apple device to access resources hosted/published for Kindle or by Google Books. Somehow it is barely conceivable that Apple will produce apps for the Kindle or the Android app stores.

What does this tell us? It tells us that building an app for the other guy's store is a sign that you have either lost the market, know that you are going to lose the battle in the long run, or are not really concerned to establish a dominant software or hardware platform for books in the first place. Apple can afford to ignore (in fact can afford to welcome) the Kindle and Google Apps, because these book reading systems will only take off on the iPhone/iPad platform if the users are able to purchase media content directly to the device through iTunes and the app store. Any such transactions are a direct win for Apple at the expense of rival platforms. If Google or Amazon were to support Apple's in-app purchasing they have lost the market and 'lost' the customer relation. My hunch is that Amazon really doesn't care too much about the market for 'soft reading systems' and does not care at all in the long run about the hardware market. They care about selling digital books and built Kindle as the first stage of the rocket that would take them to being everyone's digital library. They would be very happy if the Kindle became a mere brand, a virtual personal library system; if need be on Apple's hardware and O/S. They really do not want to lose the digital books market, especially not to Google. There is therefore scope for an alliance of sorts between Amazon and Apple, if Apple wanted it. Apple has the whip hand in these matters with its clearly superior hardware and software package, but it could offer Amazon a pretty exciting prospect as the digital books back-end to the iTunes content management system.

Monday, November 16, 2009

Google Book Settlement -- ReDux

I go away for two days in the mountains and come back to find the Google Books Settlement II, a 173 page renegotiated version of the original deal (red-lined). Catch up with Grimmelmann.

I am not sure that I will read the new contract; getting through the first version was quite a strain! Here is one view of why it matters and why it doesnt.

The Google Book Settlement matters because something pretty much like what is envisaged in the Settlement is going to happen. And that is a good thing. Google has 10 million books scanned and databased in their servers. They are revved up and no doubt waiting to go. A large part of the most valuable human knowledge of the twentieth century will be accessible and will be being digitally read in American universities from some time in 2010 (that is a very good thing; a very bad thing is that they will NOT be available in the rest of the world, and the legal technicians have not much clue about how that broader accessibility can happen). Even if the judge were to reject the Settlement, even if the DOJ were to file and insist upon some swingeing limitations to the scope of the agreement, most books will be Googled from now on. 10 million books have already been databased in the way pioneered by the Google Books system, many of them at the express request of their publishers. Some aspects of the Google project have been very controversial, which is why we have a court case and why the legal hostilities may meander on for years yet. The outcome is predictably messy but the change has occurred.

The publishing paradigm has shifted and most books will now be accessible and searchable in various ways via Google (and perhaps, let us hope, via alternative search engines). Five years ago it was by no means obvious that all books would be digitised en masse, in their entirety (even many of the bindings), full text searchable, that they would be page-rendered, that they would be straightforwardly citeable in something like the ways print scholars have cited books for centuries (volume, chapter, page), that they would be readable in much the same way as ordinary web pages, that illustrations and indexes should be in place (though for many of the 20th Century books this will not be the case -- as a direct result of the Settlement orphan illustrations are more orphan than the texts). None of this was settled in 2004.

So there has been a revolutionary shift in the publishing paradigm. But now for the other shoe: I am not confident that the Google 'victory' in the case of the Settlement, will be seen that way in the future. Thomas Kuhn who coined the usage of 'paradigm shift' to explain the way in which with a scientific revolution a period of upheaval with its paradigm shift, then led to a period of 'normal science' when investigations proceeded under the shelter of the new paradigm. I am not sure that Google will now find itself in a period of normal science or calmer waters. Google has made a tremendous step forward, with its vision of the comprehensive digital database of published books, but the suspicion is growing that this is not a terminus. The Googled library/bookshop of all published literature is not a finished product and it is highly probable that it will fairly soon be overtaken by other models and by other paradigm shifting changes in the technology. I have the very strong intuition (it is merely that and I can not prove it) that digital books will soon be used in ways that surprise us greatly and have very little in common with the current operation of the Google Books service. We have barely started on the path of understanding how digital books can be used; and as an early entrant to the field, Google has every chance of finding itself out-innovated. Some Twitter-type of disruptive service will doubtless come along soon to show us how computation really should work with digital editions.

Google will surely be making the incumbent's mistake if they suppose that the Settlement really settles, solves or finalizes the direction that digital book technology should now take. To quote the Scottish sage "I hae me doots".

Wednesday, October 28, 2009

Will the App Store make a Good Book Store?

There are a few reasons for thinking that the iPhone's App Store may become the next and best digital book store. These are a few of the reasons that occur to me:

  • iTunes is already the most important digital music store and the AppStore is inheriting a lot of the momentum and the kharma of the iTunes e-commerce system
  • the iPhone App Store is already a pretty good App Store and seems to be building Apple a possibly dominant position in the race for mobile Apps. Robert Scoble has some perceptive observations on this.
  • Apple is rumoured to be building and close to launching a tablet computer, which will share the iPhones touch interface and the e-commerce system that supports the iPhone and the iTouch. The possibly mythical Apple tablet was last seen bounding through the Australian outback looking for media content, but when when this wallabook/kangoozine finally lands it will be a gorgeous display for newspapers and books. So Apple in producing a tablet is trying to make their hardware the best for books (and newspapers, films and albums!).
  • Users like reading stuff off their iPhones and when there is a tablet the chances are that they will like that even more.
  • The Apple system despite its creaky approval process, and the very weird rules that Apple imposes on its developers, is in some respects (and surprisingly) more open than the Amazon or the Google systems. Amazon for its Kindle and Google (for Google Books, or Google Editions) require that the books they distribute or will sell reside on their servers and in their 'format'. Amazon and Google already know what digital books are. Apple is not so sure. The architectural potential with Apple is more open: any publisher or author or inventor can throw an App with some new software and display potential at the Apple system (paying their small fee) and see if it catches on (Vooks or Enhanced Editions can be experimented with in the Apple media space). Google Editions and Amazon's Kindle have no such open-ness, no inventor-driven potential, and the same goes for Sony and Barnes and Noble.
On the other hand, we can find some reasons for not too readily buying into the Apple-flavoured vision of the App Store as bookstore (or even Vookstore or Nookstore).
  • I think it was Tim O'Reilly who said that the book as App does not scale well. Which I took to mean that whilst we can envisage having one or two or several Books as Apps on our phone, it is not likely that we can manage libraries this way. That may be correct, but individuals, unlike institutions, merely accumulate libraries. We buy books one at a time and if we buy enough of them within our iPhone ways will be found for managing those collections. I have been impressed by the way in which Apps can be found within the App Store. Even though it seems to be lamentably lacking in shopper-oriented convenience and friendliness. Users have been finding and buying the Exact Editions Apps for the Spectator and Opera magazine, though there has been very little explicit advertising or promotion for them (yet). The offerings within the App store are being found. Traffic is being generated. With a bit of merchandising skill from Apple, it is conceivable that millions of individual book Apps could be found and purchased within the App store by the tens of millions of users who have iTunes accounts. Scaling may not be such a big problem.
  • Perhaps a more serious issue with the one book per App model for the Apple Bookstore is that books need to be open and to work with each other in ways which Apps do not. Apps are self-contained applications and do not provide much scope for interoperation and interaction in the ways that digital books really need to do. But is this merely a short-term problem? Apple are developing their mobile developer environment and interoperating Apps are bound to come.
The Appstore/Bookstore could work out rather well. Especially if Apple resists the temptation to over-control the environment. If Steve Jobs really is sitting on the final specification of the iTabloid as he scans the latest field reports from Wooloongabba, Woomera, Bullaroo, Geelong and Gulgong, my advice is that he should ditch the pink, opt for matt black, OK the slightly larger form factor, the bigger memory, better battery (please! a better battery) and sign off. It will not be in my Christmas stocking but I am ready to stand in line in January, or February, or whenever....

Tuesday, October 20, 2009

Bookserver and the Right Architecture for Books

The pace of change in the books space is hotting up. Two weeks ago Amazon announced that its Kindle will be internationally available. Google last week at Frankfurt announced its Google Editions proposition (or perhaps we should say they re-announced it). In three weeks Google has an appointment with Judge Chin on November 9, to re-present its much discussed Google Books Search Settlement. Techcrunch had a piece on 24 Android phones (some of which are admittedly merely rumoured, but most of which will be touted for reading books). Yesterday Barnes and Noble presented their new Nook, eReader. And the day before yesterday the Internet Archive (with several collaborators) announced its BookServer project.

The BookServer proposition seems to be very much a work in progress. Thinking on the hoof and probably a fair bit of smoke and mirrors (see Peter Brantley's Web of Books presentation and Roy Tennant's blog) . But one aspect of it feels a great deal more right than the Google Books Search proposition in its various forms. The architecture is essentially and deliberately open and multipolar:

As the audience for digital books grows, we can evolve from an environment of single devices connected to single sources into a distributed system where readers can find books from sources across the Web to read on whatever device they have. Publishers are creating digital versions of their popular books, and the library community is creating digital archives of their printed collections. BookServer is an open system to find, buy, or borrow these books, just like we use an open system to find Web sites. (Internet Archive's BookServer page)
This is a good central position around which the Internet Archive can build its coalition. But it would seem that they may end up with some unlikely allies. They are following the path of, and working with, Lexcycle's Stanza (now Amazon owned) in their strategy of orchestrating, coalescing access to formatted ePUB files, and one wonders whether the BookServer backers will fill their obvious lack of full text search through an alliance with Microsoft's Bing. There are indeed some interesting challenges for Microsoft and Amazon as they contemplate whether to address Google's potentially massive lead in proprietary book aggregation by making a more Open alliance with the Internet Archive and other champions of free and open. If the Internet Archive can maintain its footing in a genuinely open and independent position (which includes encouraging Google to spider and search, as well as Bing), it has a good chance of establishing the crucial principles that it articulates. It has a good chance of being more open to innovation than Google.

Friday, October 16, 2009

Google is Going for It

It looks as though Google has decided that it is going to get the Google Books Settlement that it wants (and I suspect that a large part of the American public want it too). Google is just going to push it through as close as it can to the original proposition. There were, after all, two ways of interpreting the decisive and late intervention from the DoJ. According to Michael Cairns it gives a blueprint for how the original Settlement could be renegotiated to pass muster. According to Pamela Samuelson, Google Book Settlement 1.0 is History, it would be astonishing if the Settlement in its current form were to be approved: a new balance and a new settlement is needed which takes account of the deep issues identified in the DoJ brief.

Normally, I would back Professor Samuelson, one of America's most distinguished IT lawyers, against a self-confessed 'know-it-all' mere publishing consultant. But on this occasion I suspect that Cairns has called it right (he may know more about how publishers go about arm wrestling with regulators). The Settlement will be nudged to accommodate the 'objections' raised by the DoJ but it will be broadly as agreed. Here are three tell-tale signs:

  1. The timescale for resolution is short. The revised Settlement will be presented to the Judge in on November 9 and he has said that he wants matters to be concluded speedily (within a few months). So Google cannot be planning major changes. No extensive process of consultation and commentary is in view.
  2. There was the rather extraordinary New York Times opinion column, A Library to Last Forever, in which Sergey Brin explained the great advantages of the Googe Book Search project, exactly as though it were to be executed as envisaged in the Settlement. He breezily acknowledges that there have been concerns about 'competition' and 'copyright' but effectively dismisses these concerns as misunderstandings (I bet the DoJ lawyers were surprised to see their analysis so airily brushed aside).
  3. Google has this week announced its Google Editions project, which is intended to make its books database resource, title by title, available to all readers everywhere in every format of ebook reader. So Google at one step is embracing and enlisting all the ebook platforms which might otherwise be a form of competitive counterbalance to the Google Books Library. I wonder what the DoJ competition experts think of the decisive way in which Google has also pre-announced, long before it is operative, how the discount system is going to work between Google, the publishers and the various retailers envisaged as operating Google Editions? And with Google selling direct as well? And with Google running the search engine for all parties? Is there no potential for anti-trust concerns in all of this?
Google Editions has nothing directly to do with the Google Books Settlement. Google has taken off its left shoe and is banging the table: "Forget about orphan books, the books of rights owners are also on our servers with their full agreement and hey we are going to be selling them as well soon, moreover in distribution channels not mentioned at all in the Settlement. Wake up please!" Google could have announced the plan last year; or next year (it is in any case not going to be ready until the first half of next year); or the year after. Google Editions, so far, only appears to be a distribution channel for books sourced from publishers or books in the public domain. As such it is quite independent of the Settlement. Yep Google Editions is orthogonal from Google Editions. Nothing to do with that controversy. But one wonders why Google should announce it now, drawing attention to the fact that Google may, one way and another, become the primary source for books in all digital and electronic formats? Three weeks before they present the revised deal to the judge (presumably with comments from DoJ).

My guess is that Google is pretty confident that the DoJ really do not want to veto the deal (do you want to stop a deal that enables ten million titles to reach the 30 million Americans who are 'print-impaired'?). The judge probably will not want to veto the deal. Not when push comes to shove. So Google had better hang in for the deal to be the way it wants. After all, Google has digitised the 10 million titles. Google has, we presume, clear vision of the way this is going to work. When the deal is agreed, Google not the Federal Government has to deliver the service. Note who is doing the heavy lifting. Note who is paying the lawyers.

Google has a great database system (oxymoron alert), Google has the ability to deliver a fabulous service, Google has been at the forefront and it has sustained its initiative, and if turns out that it has a de facto, and to a degree, a de jure monopoly, once the dust settles that will be challenged. A monopoly of books will be defeated (ultimately by competition as much as by the courts). Google's confidence is justified and is the other face of its boldness in setting the whole scheme going, and if it misapplies its momentum (which it probably will), it will come unstuck.

There is plenty of scope for Google critics: Angela Merkel, Pamela Samuelson, Peter Brantley and Bob Darnton. And we need them. But Google is getting done something which will bear a lot of fruit. So this settlement is going to go through and stopping Google Books Search dead in its tracks is not going to happen. Sure it will be appealed..... all the way (paid for by those of Google's competitors that want to slow the juggernaut down). But more to the point the Google Books Search settlement is not the last innovative step in this march. Time to move forward. Time to compete with Google through innovation, not through writs.

Saturday, October 10, 2009

Healthy Media

Joe Wikert's Publishing 2020 blog is often thought-provoking, mostly about ebooks. He has been a great fan of the Kindle, but today he fires a salvo in the direction of the existing generation of ebook readers with the Kindle bang in the centre of target:

The problem with these devices is that they encourage quick print-to-e content conversion and nothing more. In fact, they even discourage some of the simplest ways of enhancing print-to-e conversions. Embedded links are a great example. If you're a Kindle owner how often do you click on those links? More specifically, how often do you groan as you click on those links, knowing that the browsing experience ahead is painful at best? The irony is that although the Kindle was the first to include wireless functionality, that feature is really only good for one thing: buying content from Amazon. Every other time I've used the "experimental" browser I've been disappointed. That's because, at its heart, the Kindle is a reader and it doesn't encourage any other use. How the Kindle Prevents eContent from Evolving
Well put. It is lame to have live links in a digital text if the system does not support good browsing. Part of the trouble with much of the current design and thought about ebooks, is that too often publishers and technologists assume that the only thing that matters with a book is the reading of it (perhaps abetted by the 'buying' of it). The Kindle has been designed and the Amazon digital e-commerce system has been built as though the only thing that really matters with books is the reading of them. One after the other. Books are much more multi-functional than the buying/reading/moving on, modulus would suggest. And digital books need to be more than digital representations of the printed object (though they need to be that at least).

Where does this take us? Arguably books, magazines and newspapers in their digital form need to be at least as good as print books. At least as useful as print books. But also much more open and much more inter-related. They need to be on the web and of the web, because that is where we increasingly do our reading. And we don't only read books and newspapers. Our digital reading experience is inevitably becoming more various, more public, more connected and more horizontal (embracing other media types as well as all forms of print media). Taking digital books as seriously as they need to be taken, is a matter of enabling them to be open to all forms of cognitive action (reading, referring, learning, analysing, interpreting, sharing, comparing, etc) but also making them inherently more open (to search, to citation, to annotation, to quotation...). My thoughts were brought in this direction by a fascinating blog by the Philosopher of Information Luciano Floridi. Floridi writes about the ways in which ICT (digital technologies) are revolutionising health care and medicine:
Behind the success of ICT-based medicine and well-being lie two phenomena and two trends.

The first phenomenon may be labelled “the transparent body”. By measuring, monitoring and managing our bodies ever more deeply, accurately and non-invasively, ICT have made us more easily explorable, have increased the scope of possible interactions from without and from within our bodies (e.g. nanotechnology), and made the boundaries between body and environment increasingly porous (e.g. fMRI). We were black boxes, we are quickly becoming white boxes through which anyone can see.

The second phenomenon is that of “the shared body”. “My” body can now be easily seen as a “type” of body, thus easing the shift from “my health conditions” to “health conditions I share with others”. And it is more and more natural to consider oneself not only the source of information (what I tell the doctor) or the owner of information about oneself (my Google health profile), but also a channel to transfer DNA information and corresponding biological features between past and future generations (see 23andme). Arsenic and e-Health
Books as bodies? Or bodies as books? The metaphor can work in both directions. As our books, newspapers and magazines become digital they should become transparent. Digital editions will be shared, even integrated in other contexts, and as publishers and editors we need to understand and extrapolate the way in which the information they contain can flourish in other information systems. This is not a matter of abandoning the old formats but of reinvigorating them in a new technology matrix in which they become more porous. A Kindle which traps books in its hermetically sealed account is missing the main chance. A Google Books which abandons all pictures and illustrations is stunting the digital possibilities of the books it fillets. And newspaper owners or author's agents who think that they can close those digital systems down are pointing us in the wrong direction.

Wednesday, September 23, 2009

What Should Google Do Next?

The DoJ, late on Friday, produced a brief ("A Statement of Interest of the United States") which stopped the Google Books Search Settlement in its tracks. Pam Samuelson thinks "This is the most significant development since the settlement itself was announced." (HuffPo). She also predicts "Now that the DOJ has weighed in so forcefully, however, it would be astonishing if Judge Chin approved the settlement in its current form."

I guess that Google, The Authors Guild and the APA take the same view because they have now asked the Judge to delay matters whilst they re-negotiate. The judge will surely agree to the delay, but the renegotiation is not going to be easy. The DoJ's brief can be read in two ways. Pam Samuelson reads it as pretty much a rejection. That is the way it struck me. But there are also elements of the rhetoric which suggest that the DoJ (speaking for "the United States") would like to see something positive emerge from the brouhaha. There are sentences like this "Because a properly structured settlement agreement in this case offers the potential for important societal benefits, the United States does not want the opportunity or momentum to be lost." (see Page 4) Michael Cairns has picked up on this at Personanondata, and he reckons that this is an invitation to the parties to join with the DoJ and amend the settlement forthwith "this agreement will be approved with many of the changes DoJ has specified" (Deal Done).

The trouble is that the three principal difficulties that the DoJ identifies are profound and go to the heart of the deal. Are the members of the class adequately identified and their interests adequately represented? Are there fundamental difficulties in relation to copyright which is at the core of the dispute? Is there a looming anti-trust problem? The DoJ has not prescribed a solution to any of these tricky issues and if it were to prescribe a solution it would set the bar very high. I dont think they will be in the smoke-filled room with the Google lawyers and the Authors' Guild hammering out royalty rates, distributor discounts, and pricing algorithms.

One doubts that the Google lawyers are looking forward to the next few months. Few months? This is beginning to look like an interminable process. How could it be terminated? Since we think that the DoJ is right to like features of the Settlement we wonder whether the Google lawyers can craft a simple and interim agreement which gives them and all parties something in the short-term, whilst the long term negotiations on the big Settlement proceed.

An intriguing possible interim solution would be one that separated the issue of search from the issue of full text distribution. A partial service which enabled full text indexing and snippet search results of the kind envisaged in the original Google Print project. Ideally, Google should propose that this corpus of material would also be open to any other search engine (ie Google does not leverage its first mover status -- though inevitably it still has that!). Such an Interim Settlement should permit anybody to deliver search services, up to and including limited display snippets over 'books' as defined in the draft Settlement.

This specific interim compromise may not fly, but one hopes that the parties can identify some provisional elements of the project which can soon see the light of day. Holding everything up whilst the legal complications ramify is an unattractive prospect.

Friday, September 18, 2009

Google Fast Flip and the Future of Magazines/Newspapers

We took a pot-shot at Fast Flip the other day. There are a few more lessons to be drawn. The Search Engine Journal take particularly struck me. My issue was really with their headline: Google Labs Rolls Out Fast Flip, Google Book Preview for News
Whilst one can "kind of" see what SEJ are getting at: images of the publication; full page views; a linear arrangement for navigation (but note horizontal rather than the predominantly vertical scroll in Books Search); the same database squirrelling away in the background to serve pages, searches and deliver links. The overall feel is certainly more like Google Books than it is like Google News. But what had struck me when I first sampled it was the way in which Fast Flip as an interface differs from that supplied for the books collection. The intention is also very different. Fast Flip is really about skimming and Google Books is a proto-reading system. There really are some big differences between Fast Flip and the way Book Search works.

  1. Google Books Search does not present us with arrays of parallel book content, fast flipping between books, the emphasis in Books Search is to drill down into the book. In effect to read the book. If Google had launched Books Search with the manifesto: "In the age of the internet the atomic unit of consumption is the page and the Google library will allow you to fast flip between pages in different books using the relevance tags inserted by Google." there would have been uproar in the halls of learning.
  2. Some magazines are included in Book Search (though periodicals, including magazines) are excluded from the Settlement) and they are treated pretty much in the same way as books. But again with horizontal layout and thumbnails as an option (primarily, I guess, because of the prevalence of double page spreads). But these magazines are the print magazines not the web services.
  3. Google Books Search allows the user to preview the actual image of the book's page. But Fast Flip is entirely predicated on the fact that many newspapers have built websites which repurpose their issues to web pages. Google Books Search is deeply 'type-based'. Fast Flip is snapping images from web pages. Quite a different matter and much less predictable as web pages are often dynamic. In its approach to books Google presupposes that what matters is the actual printed text and its fixed pagination.
  4. Google in the light of their probable/maybe settlement of the Google Books Search cases will be committed to selling individual books. Selling them in large packages to libraries, but also selling them individually. But Fast Flip has been launched in a barrage of Google comment that presupposes (along with much conventional wisdom) that newspapers and magazines cannot be sold through subscriptions on the web. Or only in exceptional circumstances. But Google has a business plan in which millions/billions of licenses to individual titles within Google Book Search are going to be sold to individual consumers (and held within individual accounts until the expiry of copyright or account holder)? If so, digital magazines should also be saleable to subscribers. As we know, at Exact Editions, that they can be.
Google Books Search has plenty of problems (not all of them legal), but in its conception and its goals it is much, much better than Fast Flip. Newspapers and magazines would do much better to pay full attention to the way in which the Google Books Service is shaping up. Fast Flip is an experiment, a jeu, from Google Labs. Google Books is a mammoth, a juggernaut, a tsunami for publishers. Not only for book publishers. Magazines and newspapers need to figure out how they can make money selling digital editions of their publications which sit alongside that juggernaut and which are database driven web services providing paid for and (on many occasions) free access to the content which would otherwise have been printed. If digital books are to be sold online to libraries and individuals why should not newspapers and magazines be licensed to subscribers in very similar fashion?

Tuesday, August 11, 2009

Lessig on the Google Books Settlement

Lawrence Lessig contributed a 40 min discussion to the Berkman Center's seminar "Alternative Approaches to Open Digital Libraries in the Shadow of the Google Book Search Settlement”. (In the 'shadow' of the Google Settlement -- doesnt this make it sound a bit ominous?)

He opens with a comparison between Tiger/Kitten and Tiger/Tiger. Google has to be the Tiger. So although not explicitly anti-Google, his rather mournful assessment of the Google project is moving away from it. Watch out for the claws. He recognises that the GBS Settlement may represent progress and have some positive results, there are even so a lot of downsides: "We need a framework to encourage experimentation"; "We should not trust our culture to kittens that turn into tigers"; There is a tendency in the extraordinarily complex settlement agreement "against the ecology of free access which we have had since the invention of printing".

Lessig's position is not hard and fast, and tries to avoid being anti-Google. There is something rather soft, touchy-feely, about his extreme example of what is happening to books: it is far-fetched, in my view, to suppose that books will be as ham-strung with temporary permissions as documentary films. It is not clear what his recommendation really amount to. The 'appropriate or the best ecology of access' is a vague idea.

But Lessig is putting his finger on some of the tender issues in the Google project. There is a worrying tendency for the Google Books project to dissappear in a vastly complicated and centralised network of permissions, concessions, exceptions, pettifogging access restrictions, content omissions and database-driven implementation decisions which may yet stifle the project. Or, at the very least, cramp its style. With Google Book Search, code is very much becoming and making law, but not in ways that Lessig can welcome. Something looser, more rounded, more democratic and multi-polar is needed. The ultimate and inevitable failure of Google's project as it is currently shaped is that it is not putting books in the centre of its intentions. Books are not being given room to breathe.

Thursday, June 11, 2009

Is Google Making the Celera Mistake?

Celera was the company founded by Craig Venter, and funded by Perkin Elmer, which played a large part in sequencing the human genome and was hoping to make a massively profitable business out of selling subscriptions to genome databases. The business plan unravelled within a year or two of the publication of the first human genome. With hindsight, the opponents of Celera were right. Science is making and will make much greater progress with open data sets.

Here are some reaons for thinking that Google will be making the same sort of mistake as Celera if it pursues the business model outlined in its pending settlement with the AAP and the Author's Guild:

  1. The task and the cost of curating the data cannot be separated from the responsibility and the expertise of those who generate it. Celera's hope for massive private value in its private databases was undermined by the preference for publicly funded research to go its own sweet way into the public arena. Does Google really want to manage and control, assume the responsibility for all those who write books and how they can be distributed? Does Google and the Books Right Registry really think that Authors want their activities to be regulated in this fashion?
  2. Genomic databases are extraordinarily valuable, it does not follow that you can sell them as big ticket items. Is there a massive market out there for closed subscription databases to millions of books sold to institutions? Celera did make some sales of its promised proprietary databases, but it was never believable that there was available funding to support a market for billions of dollars per annum on genomic databases. Those chimerical numbers were needed to support the astronomical market cap Celera briefly touched. Google may not have such sky expectations of its digital library subscription revenues, but I wonder how well the expectations that it does have, match with the funding currently available to the public library system and educational institutions?
  3. PE was very good at building automated sequencing systems and selling them to researchers. Very, very good. It turned out to be not nearly so good at building a business to manage, curate and exploit genome databases that would be licensed to scientists and researchers. Such different activities do not mix, and your customers are likely to suspect a conflict of interest, and this is one reason why Celera was spun-out from Perkin Elmer. Google is very good, six times "very good" at managing search-sensitive advertising and large scale intentional databases drawn from web use. Are Google's customers going to be happy working with a system in which their reading attention, and referential record is always being calibrated and used to influence their buying pattern and subscription budget?
  4. Hubris. Almost certainly in the case of Perkin Elmer, but they did have the sense to pull back. With Google it is hard to say..... hubris and ambition are sometimes confused, or mistaken, the one for the other.
There are plenty of differences between these two situations. Nor am I suggesting that all literary copyrights should be put into the public domain (nor indeed should all genomic data be treated as public). Differences and contrasts abound, but Eric Schmidt should put Sulston and Ferry's book The Common Thread on his summer reading list.

Monday, April 27, 2009

Cloud Computing and Content Services

Richard Wallis who blogs as Panlibus for Tallis (the Solihull UK-based library automation specialist which seem to be going places after 40 years of quietly honing their LMS in the black country), has an interesting post on the way that library suppliers are moving into the cloud. Of course that is going to happen, and it will be interesting to see how OCLC, "the 500 lb" gorilla impacts the traditional library automation market. Especially with Google "the 15 million book" gorilla, hobbled by obvious metadata shortcomings, lurking in the background. Following on from Richard's post, I started to wonder how the content aggregators are going to react to the opportunity and challenge of selling content services within a cloud computing framework.

A recent survey paper from Berkeley summarises the hardware innovations of Cloud Computing as follows:

1. The illusion of infinite computing resources available on demand, thereby eliminating the need for Cloud Computing users to plan far ahead for provisioning.
2. The elimination of an up-front commitment by Cloud users, thereby allowing companies to start small and increase hardware resources only when there is an increase in their needs.
3. The ability to pay for use of computing resources on a short-term basis as needed (e.g., processors by the hour and storage by the day) and release them as needed, thereby rewarding conservation by letting machines and storage go when they are no longer useful. Above the Clouds: A Berkeley View of Cloud Computing (p 1)

A service model and a charging system of this kind would be very attractive to content users if it could achieve the radical cost savings typical of cloud computing. Service subscribers would jump at the promise of a service which gives them the 'illusion' of access to infinite information, (this is where the Google library of at least 10 million books comes in) and which eliminates the need for upfront commitments. The first and second capabilities are straightforward, but there does not appear to be a rationale for 3. The problem is that Information is not rivalrous. From the supplier's point of view, whether as creator or as intermediary, there is nothing saved when users do not use a resource. The intrinsic value of copyright resources does not increase because less use is made of them. Paradoxically the value of a scientific resource such as Science Direct actually increases if its usage is widespread. From the information suppliers standpoint the 'trick' is to maintain the illusion that a resource is effectively available wherever it is needed, even though a great deal is being charged for it and the barriers to entry and easy use are high. Paying for the resource on a short-trem or intermittent basis is unlikely to appeal to the rights holder. I suspect that the Books Rights Registry will be slow to sanction Google in the introduction of an hourly 'pay as you go' access model to its main collection.

A solution to this conundrum will emerge, and I suspect that it will evolve in the direction that publishers and rights holders want their information to be accessible, searchable, citeable, and to a limited extent viewable for free. But they do not want to give it all away. The necessity and the attraction of charging for data and content will be limited to services which are in some extent premium, whether by virtue of extreme topicality, of outstanding readable quality, or of additional value services. The Exact Editions service already deploys a cloud-based content management system it will be interesting to see how our partner publishers evolve solutions for end-user pricing and public access (all our public-facing services are now searchable without need for a subscription or controlled access). In fact every page can be viewed at thumbnail size without the need to register an account. Perhaps we need to evolve towards a stage where every citation at least delivers a thumbnail view even of 'closed' pages.

Saturday, April 25, 2009

On Thinking Beyond the Google Settlement

We enjoyed the London Book Fair earlier this week. There was considerable interest in the project that Bloomsbury have announced to create 'shelves' of content for public libraries, using the Exact Editions platform. There was also much interest in the fact that Exact Editions is able to directly support the iPhone user experience.

The fair was slightly quieter than last year and some publishers are feeling the squeeze of recession in reduced consumer purchases, but there was also a great deal of optimism and excitement about the books industry and in particular about the potential for new digital markets. I am sure that book publishers are in much better shape than newspaper or magazine publishers to adapt to the new challenge of digital publishing. If there were a Trade Show like the LBF for newspapers or magazines this year in London it would be so gloomy, or probably postponed.

There was certainly some discussion of Google Books Search, of the looming, probable, approval of the Settlement in the dispute with the Authors' Guild and the Association of American Publishers. There were some meetings where Google was discussed. A literary agent, Piers Blofeld, penned an angry diatribe against Google Book Search (see p 12 of the Bookseller Daily for 21 April): I am sure that I was not the only reader of this piece to be thinking "Canute". But overall there seemed to be a quiet air of business as usual, with the book trade; publishers, authors and agents, not quite understanding what is about to hit it. What is about to hit it, or them?

The London Book Fair is much smaller than the Frankfurt Book Fair, but its focus is overwhelmingly on British or English language books. The Main Hall at Earls Court accommodates several hundred stands from leading British, European and American publishers. Those publishers and the booksellers, agents, librarians and authors spend three days discussing, negotiating, dealing, buying, selling, promoting, praising, discounting, rubbishing, remaindering and very occasionally reading tiny little bits of the 300,000 odd books published or about to be published in the English language. The show is overwhelmingly concerned with new books and major sellers from the back list. The book business pretty much is the business that is on display at the London Book Fair and it has a focus on this year and next year's books (2 years worth of books from the UK and US market takes us to 300,000 or so new books).

If we think of this rather large and hangar-like hall being occupied by the books that are currently the focus of the commercial market for books, we can also imagine a skyscraper of 30 or perhaps 40 stories being built above the Earls Court Stadium. The stacked stories of this skyscraper will each contain another 300,000 mostly older books, but this time all of them ordered, regimented and deployed in total silence and precise obedience with no noisy haggling or discordant trading. Such a skyscraper would be a serious obstacle on the flight path for planes approaching Heathrow, but its towering shadow does give us an idea of the relative scale of the Google Books Search project as set against the current (this year, last year) output of publishers in the English language. The 10,000,000+ books that Google will have in its arsenal when the Google Book Search library goes live in a year of two will completely dwarf the current activity. The 40 odd stories of the Google Books skyscaper will not need the traditional tools and mechanisms of the book trade. The transactions, accessibility, searchability, and reading of these millions of books will all be a matter of database and web-driven activity. Commercial arrangements will be settled by the Books Rights Registry or the publishers' agreements with Google and the commercial transactions and access rules will be executed by Google or its contracted distributors. There will be very little need for human intervention, except at the periphery. When authors, agents and publishers decide to put things into the system, or, at the consumer edge, when readers, searchers, librarians or consumers decide that they wish to have some form of access to the repository. Of course Google will also not need a skyscraper at all. The few hundred terabytes, possibly by then one or two petabytes, that may be needed for the Google nearly-complete libary in 2012 will comfortably fit in the confines of the whirling, bladed and racked systems, housed in a single standard freight container. We should add a few more trailers to cope with the bandwidth of a billion users, but it is all fitting nicely in the underground loading bay that they have at Earls Court. The efficiency and reliability of the Google system does not require large physical infrastructure. Push on a couple of years, and by 2014 I think one can be sure that Google will have most of the world's published literature in the Google database. How will new books then be working in relation to the 50, 60, 70, 80 stories high skyscraper of previously published but now completely databased and universally accessible digital books?

So how indeed is the traditional world of books going to cope with the fact that most of the world's published literature will be available, purchasable, readable and eminently usable as a database system within a few years? The new books which are then being published will still need the care, design, attention and promotion of publishers and editors, but will readers be expecting to buy new books in volume form when everything else is usable and being used as part of a database system. Will Google be totally dominating the market for new books, as it apparently will be monopolising the market for 'orphan' books?

We may wonder. One suspects that there is far too much that goes on in the world of writing, authorship and publishing, for the talent, the style and the colour to disappear into the smooth and virtual maw of a Googlised library. Google, in my view, will not end up owning the books business, and its monopolistic trajectory will stall or run into natural limits. But in walking the aisles beneath this 40 story skyscraper, rising inexorably above Earls Court, full of scanned and indexed titles that are in many cases neglected and orphaned, one does wonder whether there can possibly be a future for proprietary file formats, Kindle or Sony readers and the non-Google searchable distribution networks that publishers are building and commissioning for themselves? What does Mr Bezos think he is achieving by locking users into a DRM which is Google inaccessible? Elsevier, Springer and Wiley are in enough trouble with their own non-standard, pre-Google Book Search, content management and access systems for it to be doubtful that the world needs another 175 variations on the same theme. There is quite a lot going on in the world of digital books that is completely irrelevant to the GBS seismic shift. The Guardian journalist writing about eBooks at the Fair managed to avoid mentioning Google or GBS at all. If the digital book is not in a format that can be searched by Google (or similar search engines) you might as well forget it. Google will not be selling access to everything, it will not be allowed to build such a monopoly, but it will be searching everything that is published. That is the big win for Google from the Settlement.

There will be one hell of a row when it is fully appreciated that this wonderful Google system is only to be properly and fully deployed in the USA. But leaving that on one side, the London Book Fair of 2012 will take place very much in the shadow of the decision Judge Denny Chin in the Southern District of New York this summer. We all need to be thinking post-Google Settlement.