Showing posts with label orphan copyrights. Show all posts
Showing posts with label orphan copyrights. Show all posts

Friday, May 27, 2011

The Google Books Mess

There were a couple of tell-tale signs last week that Google may be having some pain and problems with its vastly ambitious Google Books project. First, was the news that Google was pulling the plug on its corresponding, open-ended, plan to scan and database masses of historic newspaper archives. Second a report that Google was diverting all its programmers from its eBookstore and perhaps not vigorously pursuing plans selling eBooks.

The problem that Google has, is that there was huge momentum within the company towards its grandiose plan for a comprehensive universal digital library and this vision, with its accompanying class action settlement [ASA or amended settlement agreement] was decisively stopped in March by the opinion of Judge Chin (USDC SDNY)


While the digitization of books and the creation of a
universal digital library would benefit many, the ASA would
simply go too far. It would permit this class action - - which
was brought against defendant Google Inc. ("Google") to challenge
its scanning of books and display of "snippets" for on-line
searching - - to implement a forward-looking business arrangement
that would grant Google significant rights to exploit entire
books, without permission of the copyright owners. Indeed, the
ASA would give Google a significant advantage over competitors,
rewarding it for engaging in wholesale copying of copyrighted
works without permission, while releasing claims well beyond
those presented in the case. (Opinion 22 March 2011)

Chin's decision is styled an opinion, and it might yet be appealed or revised, but most observers would tell you that it has pretty well stopped the Google project in its tracks.

Google has got a lot of figuring out to do:

  1. Google is not out of its legal woes, although such a rich and powerful company can probably stall or out-manoeuvre the authors and publishers who are parties to the original suite in the USA. Yet Google will need some resolution to the case or it risks enormous damages for breach of copyright ($3.6 trillion according to one scholar).
  2. Google will not find it straightforward to avoid legal actions in other jurisdictions. It has ongoing legal woes in France, and if some French publishers win substantial damages, many others will charge through these same gates.
  3. Google is continuing to scan without permission millions of works which are not out of copyright on behalf of its library partners. So the liabilities grow.
  4. Google will be required to deliver digital library services to some of its core collaborating libraries. The libraries of Michigan and Stanford in particular. To the extent that these services depend on copyright works digitized without permission Google remains at significant risk.
  5. There will be increasing concern about advantages that may accrue to Google from the works that it has already scanned and databased, and which it may use in ways impervious and invisible to external actors. Perhaps Google will gain enormous advantage in the fields of search, automated translation and semantic technologies through private access to vast amounts of unregistered, unlicensed, copyright material. That putative advantage creates legal risks for Google from competitors and regulators.
  6. Without a recognized and legitimized settlement Google cannot deliver services of general public benefit, and at some point Google loses good will. Without a settlement Google cannot even be generous.
  7. Google has plenty of agreements with publishers and authors for the distribution, display and potential licensing of millions of copyright works. So it could be an active participant in the eBooks market, but it has been strangely hesitant and stuttering in recent years about its commercial activities. Almost certainly because Google's lawyers are anxious about the way such commercial exploitation may play against the unresolved matters in dispute. If Google carries on havering it will lose its opportunity in the digital books market, much as it appears to be losing its opportunity in the market for digital music.
I am not sure that Google has an easy way of stepping out of this mess. But it needs to find, or create through disruptive action, some solution.

The original goal of a universal library designed, built and maintained by a single technical player was hubristic and naive, driven by the enthusiasm and commitment of the founders (Page in particular who felt that he owed a debt to his alma mater, the University of Michigan). Google's best hope now would be to distance its involvement from the prospect of private gain and to place all works not public domain, and not explicitly licensed to Google, in the sole care and control of the public academic institutions from which the original works were taken, and to renounce any commercial advantage through its involvement in converting 'orphan' works. Google will have to pay the authors and publishers something (if only to cover some of the legal bills, that will otherwise be pursued to the bitter end on a contingency basis by the other side), it can afford to finance the first blocks of a Rights Registry, but it should be more open and more public, more consultative, in part foundation funded, than the original design. Google does not need and should not look for special advantages on rights and forward-looking business models. If Google were to do that it could help to promote the cause of orphan works legislation in a disinterested manner. Google needs to get legitimate, beyond all shadow of doubt, fast.

Google often likes to play the 'open' card, but it has been far too closed and 'private' over its books project. It needs to rethink the game-plan and its style of involvement. That way it will retain the good will of the library community and the reading public. By being highly generous and public spirited it looks after the interests of its shareholders also. Page is now CEO and he may need to bite on the books bullet and own up to a change of course, only be being much more open and generous can Google hope to make something like the Google Books project a reality.

Wednesday, January 13, 2010

What do Literary Agents Do?

The other day I was talking (mostly listening) to someone who works in trade publishing. That is the kind of publishing in which books are 'sold' for advances of £25K, or maybe £250K, or very occasionally more than a million smackeroos. This friend/acquaintance was explaining that the house she works for controls very few of the electronic/digital rights in the books they publish and she was frustrated that agents seem highly reluctant to grant any rights, or even to experiment with digital propositions.

This got me to thinking. What is the point of an agent who does not do deals? We do not hear much about agents doing digital deals. Are they just sitting on their authors' rights and not exploiting them at all? Or are there soon going to be a rash of direct deals by agents with the likes of Amazon, Sony, Apple, Plastic Logic, Google etc? Perhaps there will be some deals: a couple of months ago Amazon flew a dozen top literary agents to Seattle for frank discussions. A few days ago Amazon announced that they had done an exclusive deal with Paul Coelho, exclusive for all his e-books in Portuguese. I wonder how much Amazon had to guarantee or pay as an advance for the exclusive rights? But all the e-books rights for Portuguese Paul Coelho, (why only Portuguese?), does not sound like such a big deal (oh yes, I know he is Brazilian, so it is a fairly big deal).

I suspect that exclusivity is the key issue here. Agents are used to handling and dealing in exclusive rights, and they are working with the hypothesis that digital rights are going to be like the exclusive rights that they have learned to carve out of the traditional book-publishing contract. Identify and separate the rights and sell each of them for as much as possible to one counter-party. But are digital rights like this? Does exclusivity really cut it in the innovative market for digital books? It has always seemed to me that copyright owners would be better off, and publishers would also be in a stronger position, if digital deals were almost always non-exclusive. Why do an exclusive eBooks deal with one supplier if there are 15 different players in the market, each with their own 'installed base'? Why do a five year exclusive with Amazon if the market for digital is going to end up with Apple, or Google or someone else?

If you look at the couple of dozen eBook reading platforms that were announced, re-announced, released or previewed at last week's CES (Consumer Electronics Show), it would appear to be quite possible that the market for digital rights is going to become extremely diverse and based on many different types of non-exclusive exploitation. Are agents capable of handling this kind of fast moving market? Is your typical literary agent capable of identifying and negotiating deals with dozens or scores of technology partners? How many literary agents were at CES in Las Vegas last week? Not too many, and few literary agents are comfortable in evaluating technology propositions.

Perhaps agents should get used to the idea of granting all digital rights to the book publishers they deal with on a non-exclusive basis, retaining the right to do non-exclusive deals themselves in certain circumstances. That way publishers and agents will all be working for the trade authors they represent. Just at the moment, it appears that a degree of paralysis and ignorance is ensuring that as few deals as possible are taking place. We are seeing the emergence of a new class of 'neglected exploitation' rights, somewhat analogous to the 'orphan copyrights' which lie at the core of the Google Books Search Settlement.

Tuesday, August 11, 2009

Lessig on the Google Books Settlement

Lawrence Lessig contributed a 40 min discussion to the Berkman Center's seminar "Alternative Approaches to Open Digital Libraries in the Shadow of the Google Book Search Settlement”. (In the 'shadow' of the Google Settlement -- doesnt this make it sound a bit ominous?)

He opens with a comparison between Tiger/Kitten and Tiger/Tiger. Google has to be the Tiger. So although not explicitly anti-Google, his rather mournful assessment of the Google project is moving away from it. Watch out for the claws. He recognises that the GBS Settlement may represent progress and have some positive results, there are even so a lot of downsides: "We need a framework to encourage experimentation"; "We should not trust our culture to kittens that turn into tigers"; There is a tendency in the extraordinarily complex settlement agreement "against the ecology of free access which we have had since the invention of printing".

Lessig's position is not hard and fast, and tries to avoid being anti-Google. There is something rather soft, touchy-feely, about his extreme example of what is happening to books: it is far-fetched, in my view, to suppose that books will be as ham-strung with temporary permissions as documentary films. It is not clear what his recommendation really amount to. The 'appropriate or the best ecology of access' is a vague idea.

But Lessig is putting his finger on some of the tender issues in the Google project. There is a worrying tendency for the Google Books project to dissappear in a vastly complicated and centralised network of permissions, concessions, exceptions, pettifogging access restrictions, content omissions and database-driven implementation decisions which may yet stifle the project. Or, at the very least, cramp its style. With Google Book Search, code is very much becoming and making law, but not in ways that Lessig can welcome. Something looser, more rounded, more democratic and multi-polar is needed. The ultimate and inevitable failure of Google's project as it is currently shaped is that it is not putting books in the centre of its intentions. Books are not being given room to breathe.

Wednesday, April 15, 2009

The Trouble with Orphans

In the old dispensation books (and magazines and newspapers) used to be published and then gradually disappear. A few copies of any particular print run would be kept in archival conditions in important libraries, but by and large they gradually mouldered away. In fact they biodegraded into mulch. Something similar happened to the 'copyrights': to the intellectual property that the publications crystalised. After a few decades, and with the exception of a very few masterpieces or works of genius, the intellectual property that they represented was of negligible value or interest, and they would sooner or later fall into the public domain, probably before the physical book biodegraded. At that point the 'IP' did not matter, or rather it mattered only to the public domain.

In the last 10 years there has been a growing tumult about 'orphan' copyrights. Or 'orphan works'. In the eyes of some of the key critics, eg James Grimmelmann, the real problem with the Google Book Search proposition and the Settlement that Google is reaching with Authors and Publishers is all about the orphans that are being swept up into the maw of Google's 7 million, and counting scanned digital books.

But the funny thing about 'orphan works' is that the very category is defined by the technology which makes it possible to replace or preserve printed books by digital books which could last forever. The books and photographs are no more 'orphan' than they ever were, it is just that they look like they should be imortal rather than biodegraded, so who can speak up for them on that? Suddenly old and mostly forgotten copyrights seem to have some possible value, because the digital books could last for ever, and who knows but some of them (a few) will surely have considerable hidden value? To many of the critics it does not seem right that this value should accrue or crystalise to Google (and to the Authors Guild and some Publishers), rather than to anybody else.

If you buy the idea that the 'orphan' status of a book (or some other piece of intellectual property) is more a function of the new technology than of the old which was around when the object was born or conceived, there is an interesting corollary. As technology improves the orphans become more valuable. The 'orphans' may indeed become a lot more valuable when computation advances again, as it will. Especially if the orphans can be used to construct something else, something that we dont yet understand. Who is to say what value they have? There is a lot in the Google Books Settlement about 'non-consumptive' research (which roughly means 'reading by computers and software'). Who knows how valuable that could become?

We may get a glimpse of this when Wolfram's intriguing Alpha project is unveiled. From Rudy Rucker's recent blog about what Alpha portends, it certainly sounds to me as though Wolfram and his team have been doing some pretty sensitive 'non-consumptive' reading of key reference books:

I asked him how he is handling the daunting task of finding out all the possible scientific models. “There’s only so many linear feet of reference books that exist in the world,” remarked Wolfram. “Nowadays when I go into a library I look at the reference shelves and try and estimate how many of them we’ve picked up. I think we’re close to ninety percent by now. Right now my office is mounded with books with bookmarks for things we still need to implement, and one by one the bookmarks and the books are going away.” www.hplusmagazine.com
When the database representation of what a book is about gets to be that powerful and expressive, non-consumptive reading is arguably more useful and valuable than the old fashioned human kind of reading. Orphan copyrights, in a clever enough computer environment, have much more value than their publishers or authors could have imagined....

Monday, March 16, 2009

Google Books Search: What is Good for Google is Good for the USA

There was an important Conference on the Legal and Publishing impact of the Google Books Settlement at Columbia Law School on Friday. Several attendees, led by Peter Brantley, were actively Twittering the event, see #gbslaw for the Twitter-stream. There are one, two very useful reflective summaries posted by Peter Hirtle (lawyer at Cornell Library).

Apparently one of the recurring themes in the conference was this mantra "What is good for Google is good for the USA." I am sure that it was said in jest/irony, but that must nevertheless have made the Google participants unhappy. Even if ironic, the comparison is wounding. Just now being compared to General Motors is nearly as bad as being compared to AIG, and is frankly worse than being compared to Microsoft (which would also be very unfair and unwelcome to Google, but the comparisons are coming...). The mantra is especially unfortunate, since it is far too close to the bone: the whole way the Google Book Search settlement is working out is far too US-centric, as though Detroit was the market, and the accessibility of digital books in the rest of the world was not a matter of importance to the US or to Google. General Motors has been building inefficient and slipshod cars which had limited appeal in the rest of the world and failed the ultimate tests of quality engineering and sustainability. Could Google fall into a similar trap of building too much, too wastefully, for local demand and national circumstance without full attention to all the factors which build quality, openness and sustainability? Apparently some anxieties on this score were raised at the meeting. Somewhere in the Twittering I saw someone questioning how the US would feel if another country adopted a similar approach a private enclosure and database representation of all the books in the English language held by French libraries (the French or even more probably the Chinese Union Database Library? It will probably happen). Can you imagine the uproar? Senator Conyers would have most unfavoured nation legislation in train within a twinkling...

A lot of the books from these dusty stacks in Michigan and California are foreign published. Through the group of libraries in the US with which it is collaborating Google will catch in its net of NotYet OutOfCopyright but OutOfPrint titles a vast swathe of books originally published by British, French and German publishers. Google has apparently spent $7 million in the last two month on press advertisements in over a hundred countries to advise authors and publishers of the rights that they may have in the Settlement to the use of their books in the US market ($7 million on print ads for the legal notice, few text database projects have had a total investment this large). But the authors of those books are also readers and if the eventual legal and technological effect of the Settlement is to make the access to those books much less viable in the countries in which they were written or published?

Spare a thought for Google: not only it is it being compared to General Motors, they now also have to deliver on the very substantial obligations which the Settlement imposes on them, in particular to roll out commercial services to libraries and to individuals (to reiterate: these obligations are only to deliver services to the US market). This is going to keep Google very busy. Many critics of the Settlement have pointed out that it creates an enormous (millions of books) private preserve for Google, from books which look more like they belong to the public domain, either because they are orphan, or because they close to orphan. This monopolistic position is seen as an obstacle to competition. Of course it is in one way a matter of enormous advantage for Google.

But there is another way of looking at the situation. Google is now under the obligation, the heavy public expectation of delivering services from this massive collection. I believe that it will be under a very heavy public expectation and moral obligation to deliver, or find some legal way to enable, similar services to overseas markets. Google has assumed an onerous obligation to curate and deliver services for a large class of legacy titles. Inevitably it has been taking short-cuts, there is a weird absence of metadata, it has missed some quality goals, the books are not always exciting, many of them are out of print for good reason, I suspect that the difficulty and the importance of this legacy task will in itself make it impractical for Google to be the innovator in the book space that it might like to become. It is much easier to deliver an innovative and truly revolutionary social service for book readers when you are not curating 10 million titles. Hirtle concludes his excellent notes with this:


Yet while there may be great disappointment with the process used to generate the settlement, I also detected no incipient revolution against the settlement itself. No one was calling for rights holders to register and submit comments to the court (as they can do until 5 May). No one was saying the court should reject it and tell the parties to start over. Yes, the class may be too large and the mechanism too crude, but we created this problem when we abandoned formalities, lengthened copyrights, and started treating every copyrighted item in the world like it was a Disney movie. Given this procrustean bed we have made for ourselves, the settlement may be our only way out. Yes, Congress should create a compulsory license authorizing the use of out-of-print books - but don't hold your breadth waiting for that. In the interim, the settlement may be the best we can hope for - even though it has the potential to radically alter all of our worlds. (Hirtle: Library Law Blog)

Google will proabably get its way, for the most part, with the Settlement, but it may also find the bed it has made for itself, with the aid of Publishers and the Author's Guild, somewhat procrustean. The tasks it faces are Herculean. It will surely get a lot of attention from lawyers (within and without the business). There will be worries about monoploy and anti-trust but there will be plenty of competition.

Monday, January 19, 2009

Information Infrastructure Projects

Today Obama is inaugurated and steering his economic recovery plan through Congress will be a key target for his first 100 days. Governments throughout the developed world are looking for significant infrastructure projects to stimulate employment and to stimulate demand. That is the good old Keynesian solution to a recession which has come right back into fashion. Why has there been so little attention given to public infrastructure investments in the information field? Universal broad band seems to be the only plausible candidate so far ($6 Billion is the projected cost of this in the US plan, less than %1 of the total package). Most of the ideas that we read about from Obama, or from think tanks, are that employment will be created by investing in mass transit, or in efficient power transmission, social housing, or neglected infrastructure projects (old bridges and roads etc). Surely we should also be thinking about infrastructure projects for the information sector? But whether the projects are physical or digital, it has to be recognised that there are at least two problems with pump-priming investments. In the blogged comment of the Nobel Laureate Gary Becker:

The activities stimulated by the [Obama] package to a large extent would draw labor and capital away from other productive activities. In addition, the government programs were unlikely to be as well planned as the displaced private uses of these resources. (Becker -- Infrastructure in a Stimulus Package)

So large-scale publicly funded information infrastructure investments that are to be part of a stimulus package must not suck up highly skilled labour which is in short supply, and if they are to be part of a timely intervention they must not involve substantial leadup and planning time. Its a no-no to suggest putting billions into the creation of vast libraries of Open Source software, after all most competent programmers are gainfully employed. As for the requirement of effective planning, it will be much easier to identify a promising project that can be scaled up or replicated than to start something completey de novo. There are such projects in the information field, and rapid scaling up looks most promising for initiatives which involve large amounts of data collection. Data collection is often relatively unskilled, or with skills that can be easily taught. Of course, it is also important that the information being collected is of real use. Here are three suggestions:
  • Open Street Map -- has been making great progress has global ambitions and is run and led by enthusiastic amateurs who appear to be capable of organising relatively large scale social endeavours (mapping parties/conferences). Governments could fund volunteer groups who would generate large amounts of geo-data, perhaps with particular emphasis on ecologically relevant mapping data, and culturally relevant historical data which is neglected by commercial cartographers.
  • Google Book Search has made a very substantial step towards solving the problem of 'Orphan Copyrights' in the USA, but there are nevertheless plenty of gaps in the US data, and the Google collaboration with university libraries is much less advanced in other countries. Although using this data to full-effect reguires the Google engine (or something like it from Open Source or another company), the accumulation of the raw scans is a relatively low-tech business. It needs little more than a careful and well planned logistic pathway and an investment in scanning machines. If governments in Europe and the rest of the developed world were now to invest in large-scale scanning of the published literature, they could reach effective agreements for Open Access or appropriately regulated Approved Access to enormous bodies of printed literature. This would have the additional benefit of putting in the public domain, material which ought to be of general benefit; the Google effort would also be, and seem to be, much less monopolistic if there was a large body of government funded scanning as a counterweight. Much of the technical expertise necessary for scanning the world's out of print literature already exists in libraries and, with the example of Google Book Search in front of them, the world's libraries could be geared up to scan, within a few years, several times the 7 million books already scanned by Google.
  • Ecological audits. As we tackle global climate change it will become increasingly necessary to have better information about animal and plant species and the environmental impacts of climate change on living systems. Amateur efforts like the British Bird Survey are pioneers in this work. If Governments invest now to create better and ongoing records of natural diversity these records will be very useful in the global challenge of moderating and putting a break on climate change.

Friday, December 19, 2008

Institutional Licenses

We started offering Institutional, or Site Licenses a year ago. This was not a part of our original business model, but the Exact Editions founders know that market from previous experience. We have, even so, been surprised at the strength of interest for consumer magazines. To judge by the range of libraries subscribing to our services (government departments, international organisations and schools as well as universities) there is a large global market for subscriber content delivered via IP to private networks. It is a 'rule of thumb' in STM publishing that there are 2/3,000 universities world wide that constitute the market for periodical literature. The market for digital books and consumer magazines in libraries world wide is potentially two or three orders of magnitude higher (think schools and public libraries).

There is a large market, but there must be real pressure on the budgets to acquire information. At the high end, STM academic databases are very, very expensive (individual universities are often spending millions on their scientific periodical budgets). It seems likely that new entrants in the books and digital magazines space will thrive if they keep their prices down (by which I mean significantly less than $1,000 per title, per institution). Digital books may need to be priced in the region of $100 pa, per site, if they are to achieve widespread adoption.

Perhaps much lower? When Google Book Search starts selling its massive archive of mainly 'out of print but in copyright' books it will be fascinating to see how that is priced. On a per title basis the collection will need to be priced at much less than $1 a title if individual universities and colleges are to subscribe for access (there are already 7 million titles in the aggregate).

It is at least possible that the advent of the Google system will lead to a two-tier market. A huge pile of mainly little-used books in the Google mass, and slivers of high value and highly current content which will be marketed and promoted direct by the publishers, or perhaps by Google in a 'premium stream'. That will pose problems of channel contention and regulation. It may not be too long before authors, publishers and even Google are saying that books in the 'slush pile' should be free. One suspects that this is what Google may have wanted as an outcome from the beginning.

PS 'Slush pile' is here merely a technical term. Slush piles contain many books of outstanding quality. The Google collection will be very valuable and useful, though, of course, with plenty of rubbish in it!

Sunday, November 09, 2008

Regulating the Google Settlement

While it is a very good thing that Google and the authors and publishers are not going to be involved in years of fruitless and expensive litigation, there may be some awkward consequences. The draft settlement stops the head-on dispute, but the compromise does appear to have some rough edges. Signing off on this settlement is going to be a tricky problem: no judge will want to be blamed for approving a system which violates public trust or creates a de facto monopoly. A lot in the settlement is indicative and provisional and interim (they don't quite say "If this doesnt work both parties agree that we will whistle up something else", but they come damned close to doing so on more than one occasion). Who, at this stage, knows how the various business models will work out (take a look at Georgia Harper's speculations on pricing "bins" here)? Will a poorly crafted and hastily approved settlement create as many problems as it solves?

But one of the clear things is that there is going to be a Books Rights Registry. This doesnt wait for the judge. It is already whirling into action and authors and publishers are addressing it. This agency is something that the books world needs and it has precedents and cousins in the many 'collection societies' that look after dispersed copyright interests (eg in music, graphic art, xerography etc). So we have a new 'Rights Society' one which serves the interests of authors and publishers in the management and exploitation of digital texts (so far only in the US, but the same model will doubtless be rolled out in other jurisdictions -- think about it: we just called up 150 or more digital collection agencies in different jurisdictions and languages). Google is paying $34.5 million for the creation of the first Books Rights Registry (whose ongoing operation will be funded by a levy from the rights managed) and it would seem highly likely that Google is already building it. That Google is doing this is in many ways a good thing -- what an appaling prospect if the publishers were to try and build such a system! But there are dangers and ironies in a situation where Google as the commercial fox, the first and prime exploiter of the distribution opportunities flowing from the settlement, is also designing the chicken wire and building the coop in which the hens will be housed. It is a bit odd for a commercial operator to building its own regulator. Yes, I know that the 8 directors of the Registry are all appointed by the publishers and the authors (4 each). But directors decide the issues that havent already been decided, its the architect and the plumbers who get the building to function. Odd, but possibly unavoidable in these strange circumstances.

Google, unlike the publishers, the authors or their agents, is capable of rapidly and elegantly building a databases system which maps and regulates the incredibly complex real world of copyright exceptions (I recall Frances Haugen's comment about Google's management of 'amazingly complicated' viewability restrictions). Google's code already understands much of the bizarre detail of the world of rights and Google also understands how these rights might need to be exploited (or 'circumnavigated' the international ramifications are quite mind-boggling) so their system is more than likely going to work. This is certainly an area in which code will become law.

But all this makes me wonder whether the judge who signs off on the settlement will really devote a small portion of his/her time to the 300 odd pages in the settlement documentation. A lot of that documentation will and should evolve in the light of experience. She should really be looking very carefully at the API which the Rights system will incorporate and the principles which underline the API. Devising the principles which should govern this API and crystalising the objectives that the Rights Registry should foster is a matter on which the judge can make a real impact. These are matters of principle and public good, barely touched on in the public documents about the settlement, where we need judicial oversight. Perhaps she will spend some of her time looking at the Android constitution and I hope she will require that the commercial exploitation of literary rights is as open and at least as un-Google biased as Google has promised to make the Android playing field. Some of the Android slogans work rather well for our vision of digital books: 'Books without borders', 'Books can easily embed the web', 'Books are created equal', 'Books can run in parallel'. Digital books should do all of that and if they run on Android devices as well, who knows all books may soon be 'available' anywhere for everybody. Nearly all available for free search..... but that is not yet enough.

Wednesday, October 29, 2008

Google Book Search Settlement

The rumours of a settlement were correct. The publishers and Google have settled their law suit (the agreement is here). That is a very good thing and it opens up a new chapter in the development of publishing and digital libraries. The main thing is that digital libraries are going to be incredibly important and a large part of our cloud-based knowledge systems: furthermore they will be run along the lines of the Google Book Search system (database-driven, page-oriented, url-guaranteed, access-managed and to a considerable extent free). There are plenty of interesting blog comments: the Laboratorium, Peter Suber, Medialoper, PersonaNonData, and TechDirt.

There are winners and losers in this settlement, and I agree with the Medialoper view "the only entities that don’t seem to have fared so well are parties who weren’t involved in the suits"; by and large the settling parties look like the winners (Google, publishers, authors and libraries). But I wonder whether there is not an element of a 'winner's curse' about to descend on Google. Some parts of the settlement outline a fantastically complicated and ingenious business model for our future access to digital books. Very specific mechanisms for the pricing of books and the regulation of access, access to content within books, and access from within institutions to digital resources. If you read the stuff about 'Pricing Bins' and 'Pricing Algorithms' (pp49-50) you will get a good flavour of the extraordinarily detailed prescriptions.

A lot of this setup and this detail really needs to be established by innovation, by experiment and by markets, not by a court approved Settlement to a private dispute. Google may find itself subject to a lot more regulation and attention whilst it attempts to make these business models work (many business models or modes of exploitation are encompassed in the agreement). The settlement appears to be highly transparent and open, but it is not so transparent how the split between authors and publishers and other rights holders is intended to work. That may be a rather crucial consideration which may now be the subject of discussion between the Authors Guild and the AAP!

Tuesday, March 11, 2008

Orphan Copyrights

At the recent Tools of Change conference I met up with Peter Brantley, the Executive Director for the Digital Library Federation. We had a very brief conversation about Orphan Copyrights and I suggested to Peter that he produce some thoughts on the subject for the Exact Editions blog. He has produced some specific and intriguing suggestions in an essay which should attract discussion. Since his essay is in part a call-for-action which may involve lots of parties -- the posting is now up on his Shimenawa blog.

I particularly like the features of Brantley's proposals that amount to an Adoption Agency for Orphan-ed copyrights (improved access, digitization, take-down provision, escrow account, and reconciliation of public and private interests). This is a set of important issues which is best tackled by the broadest possible range of interests.

Thursday, October 25, 2007

Libraries working with Google Book Search, Or Not

At the weekend there was an interesting article in the NYTimes about the increasingly wary reaction of libraries to the Google Book Search proposition. Major research libraries are looking for a more open distribution model, without Google proprietary restrictions, and supporting the OCA (Open Content Alliance); and more are realising that they can do their own thing.

Interesting comments on this article from Michael Cairns at PersonaNonData, and from Peter Brantley at O'Reilly. Interestingly different, but they both highlight the idea of Digital Interlibrary Loan.

But I am not sure that the concept of Digital Interlibrary Loan really holds up. Well it works fine if digital libraries are composed of Books-as-files, since you can of course loan and track a PDF file; but if digital libraries are databases of searchable books and manuscript collections, where the book lives by virtue of being searched with and linked to other books, the concept of an interlibrary loan is redundant. Consider this question: how are you going to find this rare out of print book which might be available to you through digital interlibrary loan? Before you can borrow a book you need to know that it exists. So you are going to search for it in the complete library catalogue which provides full text searching as part of the catalogue, and then offers you Google-style snippets of the content. That is roughly the way things are going to work in state of the art libraries in 2010. The catalogue you are searching is in a library on another continent. And yet the book looks really good so you want to have it on interlibrary loan.....

But, but hold on a minute, you have been searching it and snippeting it and its already 'on' the server where you are searching the catalogue, so having it available to read digitally is just a matter of being able to access, search and read every page. Its just a matter of access and of lifting up the snippeted grid that stands between you and the book in all its veridical, full text, scanned image, glory. There is nothing to be loaned, its just a matter of providing access. Once books are searchable through the web, the idea that they need to be loaned is otiose. Before we get to digital interlibrary loans we are going to have campus to campus digital walk-in access......

Wednesday, October 17, 2007

Google and Copyright

Google have just announced a new set of Content ID tools which will help copyright owners protect their content on the YouTube platform. More detail is given here. These new policies and the copyright ID platform may enable Google to shrug off or negotiate a way out of the onerous suits it faces from Viacom and the UK's Premier League. But it also seems to back away from the idea of establishing a 'fair use' of video clippings or quotations -- a contentious issue which is at the heart of the YouTube success. The new Google approach appears to give copyright owners total control over the distribution of their video content.

Google will have to make similar proposals to the owners and custodians of literary copyrights. We can expect a comparable "highly complicated technology platform -- [with] content identification tools" to be in preparation for the Google Book Search platform (they already have much of it in place already) . It would be hard to go before a judge saying that literary copyrights are going to be treated differently from video copyrights. I predict this is going to lead Google to handing a lot more power to its Publisher partners and less leeway to its Library partners in the construction of the Google Book Search 'library'.

Some of the Google statements are quite striking and humble:

No matter how accurate the tools get, it is important to remember that no technology can tell legal from infringing material without the cooperation of the content owners themselves.....The best we can do is cooperate with copyright holders to identify videos that include their content and offer them choices about sharing that content. As copyright holders make their preferences clear to us up front, we'll do our best to automate that choice while balancing the rights of users, other copyright holders, and our community as a whole. [See videoID-about]
It is especially tricky to see how one can automate the choices of copyright holders whilst balancing the rights of users....As John Batelle wonders its not at all clear what happens to fair use. But book publishers will certainly welcome the idea that they might be given more control 'up front'. Its what they have been asking for all along.

Trouble is that literary copyrights can be a lot more confused and complicated even than video copyrights. All serious literary publishing requires that scope be given to 'fair use'.


Thursday, September 13, 2007

Google Book Search Project or the Human Library Search Project?

Siva Vaidhyanathan discusses how the Google Book Project threatens copyright in a short podcast, posted at First Monday (hat tip to IF Book blog). Siva says that Google is 99% certain to lose its copyright case in the courts (although he also mentions that a settlement is quite possible, so I guess he means that Google will lose or settle in a way which loses the key issue). Siva is a lawyer and he is 99% sure that Google will lose. I wonder how the case looks to the Google lawyers?

But Professor Vaidhyanathan is not a Luddite, he is very much in favour of the project of a global, universal, web-based library; but not as a private venture. He draws the comparison with the Human Genome Project. Making the universal digital library, through which all out-of-copyright information could be accessed, is worthy of national and international support.

We’re willing to do these sorts of big projects in the sciences. Look at how individual states are rallying billions of dollars to fund stem cell research right now. Look at the ways the United States government, the French government, the Japanese government rallied billions of dollars for the Human Genome Project out of concern that all that essential information was going to be privatized and served in an inefficient and unwieldy way.

So those are the models that I would like to see us pursue. What saddens me about Google’s initiative, is that it’s let so many people off the hook. Essentially we’ve seen so many people say, “Great now we don’t have to do the digital library projects we were planning to do.”....... transcript

A very interesting idea, and it would need drivers like Jim Watson, John Sulston, the NIH and the Wellcome Trust to make it happen (and the Health component will be a big part of the public justification). Digital magazines will be part of such a global library and some of their archives will be freely accessible. All published magazines should be searchable through the web and that will happen because it obviously needs to happen and because it will enormously increase their value when it does. As Siva says the next five years are going to be interesting.

Thursday, March 08, 2007

Google vs Microsoft II

There have been many informative follow-ups to the Rubin speech mentioned here 2 days ago. Surely the most insightful analysis comes from Danny Sullivan. If one follows some of the links in his posting one can see how amazingly thorough Danny is. The bravest posting was mildly critical of Microsoft from the Microsoft-employed blogger Don Dodge. Tim O'Reilly speaks eloquently in Google's defence. He usually does on Book Search. So does Lawrence Lessig. Andrew Grabois's follow-through in the O'Reilly posting, comments on the spuriously inflated percentage of orphan copyrights, seems to reduce the force of O'Reilly and Lessig's argument. The Grabois points took me over to the PersonaNonData blog, for Michael Cairns's own comments. The pnd blog also alerted me to the mournful post from Peter Brantley. That is the contribution that would most worry me if I were the Google Book Search product manager (eek!). But today's PaidContent note on Google's Chinese library efforts leads one to the conclusion that Google is now so far in, it just has to get it right.

Google surely needs to settle with the publishers, listen to its friendly critics in the library community and work with other players in the digital book space. That is the only way to get it right and the only way to

organize the world's information and make it universally accessible and useful
Whenever Google plans a new information service, someone in the Googleplex should intone 100 times that it is the world's information, not Google's. Google is great and it needs to stay that way.