Showing posts with label copyright. Show all posts
Showing posts with label copyright. Show all posts

Thursday, June 11, 2009

Is Google Making the Celera Mistake?

Celera was the company founded by Craig Venter, and funded by Perkin Elmer, which played a large part in sequencing the human genome and was hoping to make a massively profitable business out of selling subscriptions to genome databases. The business plan unravelled within a year or two of the publication of the first human genome. With hindsight, the opponents of Celera were right. Science is making and will make much greater progress with open data sets.

Here are some reaons for thinking that Google will be making the same sort of mistake as Celera if it pursues the business model outlined in its pending settlement with the AAP and the Author's Guild:

  1. The task and the cost of curating the data cannot be separated from the responsibility and the expertise of those who generate it. Celera's hope for massive private value in its private databases was undermined by the preference for publicly funded research to go its own sweet way into the public arena. Does Google really want to manage and control, assume the responsibility for all those who write books and how they can be distributed? Does Google and the Books Right Registry really think that Authors want their activities to be regulated in this fashion?
  2. Genomic databases are extraordinarily valuable, it does not follow that you can sell them as big ticket items. Is there a massive market out there for closed subscription databases to millions of books sold to institutions? Celera did make some sales of its promised proprietary databases, but it was never believable that there was available funding to support a market for billions of dollars per annum on genomic databases. Those chimerical numbers were needed to support the astronomical market cap Celera briefly touched. Google may not have such sky expectations of its digital library subscription revenues, but I wonder how well the expectations that it does have, match with the funding currently available to the public library system and educational institutions?
  3. PE was very good at building automated sequencing systems and selling them to researchers. Very, very good. It turned out to be not nearly so good at building a business to manage, curate and exploit genome databases that would be licensed to scientists and researchers. Such different activities do not mix, and your customers are likely to suspect a conflict of interest, and this is one reason why Celera was spun-out from Perkin Elmer. Google is very good, six times "very good" at managing search-sensitive advertising and large scale intentional databases drawn from web use. Are Google's customers going to be happy working with a system in which their reading attention, and referential record is always being calibrated and used to influence their buying pattern and subscription budget?
  4. Hubris. Almost certainly in the case of Perkin Elmer, but they did have the sense to pull back. With Google it is hard to say..... hubris and ambition are sometimes confused, or mistaken, the one for the other.
There are plenty of differences between these two situations. Nor am I suggesting that all literary copyrights should be put into the public domain (nor indeed should all genomic data be treated as public). Differences and contrasts abound, but Eric Schmidt should put Sulston and Ferry's book The Common Thread on his summer reading list.

Sunday, March 29, 2009

A Democratic Quality to Digitization

Robert Darnton used this interesting phrase in his recent NPR comments on the Google Book Search project and the Settlement (a 7 minute interview here).

There is a democratic quality to digitization but if those supplying it are simply trying to maximise profits the whole thing could turn sour. (The Infinite Shelf .. On the Media, 27 March 2009)
Is it true that there is a democratic quality to digitization? I think there may be a profound truth there, and getting at it, may do something to reduce or quieten Darnton's worries. He is right to be worried. If Google were to become the predominant and monopolistic supplier of books (and other print, digital print, resources) through the web, that would be a disaster. But that is a big if because digitization of our print heritage is a broadly democractic shift. It is a democractic shift in much the same way as the invention and adoption of print led to, or was one of the necessary preconditions of the democratic thrust of the Enlightenment (see Darnton's original post on Google & the Future of Books). Darnton rightly points out that the democracy of the Enlightenment was partial and restricted in its reach by privilege
Far from functioning like an egalitarian agora, the Republic of Letters suffered from the same disease that ate through all societies in the eighteenth century: privilege. Privileges were not limited to aristocrats. In France, they applied to everything in the world of letters, including printing and the book trade, which were dominated by exclusive guilds, and the books themselves, which could not appear legally without a royal privilege and a censor's approbation, printed in full in their text. (NYRB 12 February 2009)
The web is putting the final nail in the coffin which restricted the privileges of print (bolstered by the legal privilege of copyright), initially to men (rather than women), the rich and then the wealthy, the formally educated and which even now in our own time excludes some. The democratization implicit in digitization works in two ways. It works for the universality and openness of distribution because it is now a fact that digital copies and digital access are available at marginal cost for everyone. A lot more stuff will be free, partly because advertising which accompanies or supports it can generate profits, but also because it really is dirt cheap to provide free web access. So cheap that to anyone who provides digital services, providing some services free, some access to content for free, is a no-brainer. Digital access is strikingly open and democractic in its thrust because it actually (and obviously) costs more to exclude someone or anyone from access to a web resource than to enable it for everyone. 'Open' is simply, for the supplier, the lowest cost access model on the web. Authenticating, selling, registering for or targetting access costs more. But the democractic bias of digitization works also at the point of creating digital resources. It is much easier to create and if necessary re-create digital resources than to look after them in any other way.

Moving from the democratic thrust of free access from digitization, a digital process is like a printer's press in that it enables us to originate digital masters. Digitization as a method of data capture, a means for transforming cultural objects to web presence, is also becoming more feasible and more necessary. Digitization as a process is democratic because it is repeatable and reliable and affordable. Digitization is also likely to be of higher quality if it is various and competitive (Google's problems with quality of capture are notorious). Digitization as a transformative process, relying on software, computers and scanning instruments, is becoming easier and cheaper at something close to Moore's law. Even Google's massive digitization project is now much easier and cheaper than it was when they started. Since digitizing books (films, works of art, music etc) is becoming more affordable and easier every year we should have more of it. We will probably soon have consumer-targetted, hand-held, intelligent scanners.

The real danger in the Google Book Search service and the Settlement is that libraries and publishers should start to think that digitization is best left to the uniquely specialised Google. To prevent a monopoly we need a choice of services which digitize books and print resources and serve them openly (or as commercial services) to audiences through the web.

I think Darnton is right, there is a democratic thrust to digitization and it is in all our interests that there should be lots of alternatives to the digitization engine that Google has created with the help of the New York Public library, the Oxford, Harvard, Michigan and Stanford University Libraries (of course many more univerisities are now in the Google ship). Surely, the Google Books Library, (for it is rapidly becoming that), needs to be watched so that it does not become an engine for monopoly pricing, but the best safeguard against this is to create and sustain alternatives. Having played a part in kicking off the Google initiative, Harvard can help the next and better proposition that comes along. Darnton as Harvard's librarian should be there to support it.

Monday, March 16, 2009

Google Books Search: What is Good for Google is Good for the USA

There was an important Conference on the Legal and Publishing impact of the Google Books Settlement at Columbia Law School on Friday. Several attendees, led by Peter Brantley, were actively Twittering the event, see #gbslaw for the Twitter-stream. There are one, two very useful reflective summaries posted by Peter Hirtle (lawyer at Cornell Library).

Apparently one of the recurring themes in the conference was this mantra "What is good for Google is good for the USA." I am sure that it was said in jest/irony, but that must nevertheless have made the Google participants unhappy. Even if ironic, the comparison is wounding. Just now being compared to General Motors is nearly as bad as being compared to AIG, and is frankly worse than being compared to Microsoft (which would also be very unfair and unwelcome to Google, but the comparisons are coming...). The mantra is especially unfortunate, since it is far too close to the bone: the whole way the Google Book Search settlement is working out is far too US-centric, as though Detroit was the market, and the accessibility of digital books in the rest of the world was not a matter of importance to the US or to Google. General Motors has been building inefficient and slipshod cars which had limited appeal in the rest of the world and failed the ultimate tests of quality engineering and sustainability. Could Google fall into a similar trap of building too much, too wastefully, for local demand and national circumstance without full attention to all the factors which build quality, openness and sustainability? Apparently some anxieties on this score were raised at the meeting. Somewhere in the Twittering I saw someone questioning how the US would feel if another country adopted a similar approach a private enclosure and database representation of all the books in the English language held by French libraries (the French or even more probably the Chinese Union Database Library? It will probably happen). Can you imagine the uproar? Senator Conyers would have most unfavoured nation legislation in train within a twinkling...

A lot of the books from these dusty stacks in Michigan and California are foreign published. Through the group of libraries in the US with which it is collaborating Google will catch in its net of NotYet OutOfCopyright but OutOfPrint titles a vast swathe of books originally published by British, French and German publishers. Google has apparently spent $7 million in the last two month on press advertisements in over a hundred countries to advise authors and publishers of the rights that they may have in the Settlement to the use of their books in the US market ($7 million on print ads for the legal notice, few text database projects have had a total investment this large). But the authors of those books are also readers and if the eventual legal and technological effect of the Settlement is to make the access to those books much less viable in the countries in which they were written or published?

Spare a thought for Google: not only it is it being compared to General Motors, they now also have to deliver on the very substantial obligations which the Settlement imposes on them, in particular to roll out commercial services to libraries and to individuals (to reiterate: these obligations are only to deliver services to the US market). This is going to keep Google very busy. Many critics of the Settlement have pointed out that it creates an enormous (millions of books) private preserve for Google, from books which look more like they belong to the public domain, either because they are orphan, or because they close to orphan. This monopolistic position is seen as an obstacle to competition. Of course it is in one way a matter of enormous advantage for Google.

But there is another way of looking at the situation. Google is now under the obligation, the heavy public expectation of delivering services from this massive collection. I believe that it will be under a very heavy public expectation and moral obligation to deliver, or find some legal way to enable, similar services to overseas markets. Google has assumed an onerous obligation to curate and deliver services for a large class of legacy titles. Inevitably it has been taking short-cuts, there is a weird absence of metadata, it has missed some quality goals, the books are not always exciting, many of them are out of print for good reason, I suspect that the difficulty and the importance of this legacy task will in itself make it impractical for Google to be the innovator in the book space that it might like to become. It is much easier to deliver an innovative and truly revolutionary social service for book readers when you are not curating 10 million titles. Hirtle concludes his excellent notes with this:


Yet while there may be great disappointment with the process used to generate the settlement, I also detected no incipient revolution against the settlement itself. No one was calling for rights holders to register and submit comments to the court (as they can do until 5 May). No one was saying the court should reject it and tell the parties to start over. Yes, the class may be too large and the mechanism too crude, but we created this problem when we abandoned formalities, lengthened copyrights, and started treating every copyrighted item in the world like it was a Disney movie. Given this procrustean bed we have made for ourselves, the settlement may be our only way out. Yes, Congress should create a compulsory license authorizing the use of out-of-print books - but don't hold your breadth waiting for that. In the interim, the settlement may be the best we can hope for - even though it has the potential to radically alter all of our worlds. (Hirtle: Library Law Blog)

Google will proabably get its way, for the most part, with the Settlement, but it may also find the bed it has made for itself, with the aid of Publishers and the Author's Guild, somewhat procrustean. The tasks it faces are Herculean. It will surely get a lot of attention from lawyers (within and without the business). There will be worries about monoploy and anti-trust but there will be plenty of competition.

Sunday, February 01, 2009

Google Book Search and the Tragedy of the Anti-Commons

Michael Heller, a property lawyer at Columbia University, has coined the term the 'tragedy of the anti-commons'. This is a twist on the more familiar idea of 'the tragedy of the commons' -- which is thought to be the cause of such ecological disasters as the implosion of fisheries, perhaps even the nearing apocalypse of global heating. Heller's insight is that too much private ownership can be as much of a problem as too little: “When too many owners control a single resource, cooperation breaks down, wealth disappears and everybody loses.” He gives plenty of examples in his book The Gridlock Economy -- the book's argument is forcibly stated in its subtitle: How Too Much Ownership Wrecks Markets, Stops Innovation, and Costs Lives.

There is a good chance that the Google Books Settlement is going to show us all how this tragedy of the anti-commons works out in the world of books. The Google project, which is backed by the American publishers and American Authors's representatives should be (in my view will be) a wonderful resource for American universities, schools, public libraries and through them for American consumers. By 2011, if the Settlement is approved, at least 5 million out of print but not yet out of copyright [OOPnotYOOC] titles will be available to readers in the US market. This resource will have little opportunity to work so well for authors, readers and consumers in the rest of the world. The books will by and large not be available in the rest of the world (perhaps in American embassies?).

Google is already serving a very different and vastly narrower view of Google Book Search to the rest of the world (even to Canada and Mexico). Books which are public domain and wholly visible and readable in the US are not visible and readable elsewhere. And this copyright caution about territorial rights is unlikely to change, because the Settlement, when it is approved, is only going to be approved and agreed for the US market. Google has been persuaded (or has volunteered?) to accept the territorial restrictions and complications inherent in the market of copyright books. In my view, Google will not risk starting court actions in other jurisdictions, for the very simple reason that they might be lost, or worse still settled on a different basis from the US dispute. Google will be bound to leave the ex-US position of its wonderful aggregate of unloved (mostly 'orphan') copyrights in a national limbo. The orphans will remain unloved outside the 50 states.

The complexity of the rights situations of these millions of titles is effectively unmanageable and un-negotiable, which is pretty much what Michael Heller means by a tragedy of the anti-commons. By developing and growing an intricate and incredibly complex system of rights for different legal regimes and market territories the publishing industry has produced a system where a negotiated and innovative new service is probably impossible. It would take something like a new Berne convention on copyright to make this a level plane for all jurisdictions.

One might say that this hopeless and impenetrable thicket of rights which are largely historical and dormant is a problem for the rest of the world and for scholars outside the US. It is not a problem for the US, or for Google. Well maybe..... but it is also possible that this lack of international and global relevance will undermine the authority and the prestige of a US-centric resource. I wonder whether US scholars will accept a situation in which citations and references cannot be made and verified in a global context?

There is another dimension in which the impenetrable complexity of the rights position OOPnotYOOC titles: illustrations and photographs in these titles are in effect excluded from the scope of active exploitation by Google. Interestingly enough, Children's Book Illustrations are to be treated differently. They are defined as 'inserts' and therefore fall within the scope of the settlement and will presumably be in the searchable and readable services that Google produces. But, in the place of ordinary illustrations and photographs in books which are not 'Children's Books' we should expect gaps or blanks, such as one already finds in the Google Book Search service. Eric Rumsey thinks that I may be on my own in reading the Google Settlement this way, but some apparently well-informed, anonymous, commenter makes a similar point in a comment on the Martyn Daniels blog. Why should illustrations in Children's Books be treated differently from those in other books? I suspect that the publishers and the Authors Guild felt that they could negotiate with certainty on these rights (as also on quotation rights, rights in poetry etc) but they knew that they could not negotatiate for the owners of artistic rights.

Will it matter that Google Book Search, when it is marketed as a commercial subscription service for libraries and universities cannot be accessed or read in the world at large? Will it matter that many of the photographs and illustrations in millions of the OOPnotYOOC titles will not be there? Yes, it will matter, and that it matters will be another instance of the tragedy of the anti-commons.

Wednesday, October 15, 2008

Copyright Czar and Copyright U-turn

The outgoing US President has signed a bill which creates a US Copyright Czar. See the PC Magazine report.

Bush signed the Prioritizing Resources and Organization for Intellectual Property (PRO-IP) Act, a measure that will create several new government enforcement positions.
If it turns out that it is President Obama who is in fact charged with appointing this Czar, there must be a small chance that Laurance Lessig will fill the post. That would be an ironic turn of events.

As a copyright loyalist, one has to recognise that it is often the supposed advocates and defenders of copyright who are its worst enemies. Copyright would be stronger if it were more permissive and less onerous. Witness the ludicrous decision of the German courts who have decided that Google is infringing copyrights when it includes thumbnail images in search results. Whatever the technicalities of the German law on this point, we can be certain that incredibly useful services such as Google image search will not be derailed. At a certain point technology simply plows on and works its way around obstacles of this kind.

Wednesday, August 06, 2008

More Mygazines

The Press Gazette has more about the Mygazines site and its possible business model. They reproduce a lengthy but empty email from the creator of the web site, supposedly 'John Smith', but that may well be an alias for whoever has built the service. Here is an extract from John Smith's email:

The true future of the industry lies in the final stages of our site concept. We can easily transition to the final revenue model quickly with the co-operation of the publishers. We cannot however reveal the full concept at this time as we are saving that discussion for the publishing industry directly.........
As per our press release: We have every intention of working with the industry to provide not only revenue streams that are vast, but also an answer for the Publishers in general. Our method will increase current revenue, halt and reverse advertising revenue lost to the internet, and overcome the lack of the ability for magazines to stay current.
There is more in that bombastic and questionable tone.

There was something fishy about the Mygazines claim that these magazines (hundreds of complete magazines) were being uploaded by end users who aimed to share 'their' magazines with others. One tell-tale sign, most of the magazines were uploaded in their entirety with Contents Pages clearly identified. If a magazine sharing site, built by the community, was for real it would be chock full of magazines which had been partially uploaded or badly annotated in the upload process. Mygazines content looked far too perfect. Much more probable that it was the work of one or two bodies toiling away with a guilotine and feeding the scanned results into the database system that lies at the heart of the service. There is no 'safe harbor/user generated content' defence for doing that.

It all looks like a shameless and pointless ripoff operation, with no under-lying business proposition which could possibly appeal to publishers. What a pity that the ingenuity and effort that has been put into building this 'service' was not applied to a more worthwhile and sensible project.

Wednesday, July 23, 2008

How to build an audience quickly.....

What is Mygazines? ..... (answer from their web site)..."Mygazines is your free place to browse, share, archive and customize unlimited magazine articles uploaded by you, the Mygazines community."

The service appears to be in part a YouTube for magazines, and an aspirant social network. The technology is adequate to impressive: the magazines are rendered in Flash, and the user is able to mark. 'share', comment upon, pages and magazine articles. The system is largely automated, except that some of the 'cleverer' bits are 'user generated'. So in uploading a magazine as a scanned PDF (I havent done this) the user is expected to tell the system which pages are front, covers, contents pages, and where articles begin. I am distinctly impressed that users will in fact do this (we have always assumed that publishers would not be reliable about marking this information on PDF files). Also, while the site is very new, maybe only a week old, and likely to disappear very soon (for reasons we will come to), they already have 800 titles including many of the mainstream US and Canadian titles (think Time, MacWorld, Wired, Business Week, Maclean's.

The system also does the trick of intelligently OCRing the scanned uploads (we were blogging about this last week), searching across issues works pretty well, so all in all I can see this getting a lot of usage. Except that it appears to have been done without the permission of the publishers and in apparent disregard of the laws of copyright. Mashable, Joho and gHacks comment favourably on the experiment. But we are all rather surprised at the legal presumption; in Dave Weinberger's words: "I don’t know what they think they’re going to do about the obvious copyright issues."

So I fear that the founders of mygazines are very quickly building an audience of ...... lawyers. Scores of lawyers from all the big magazine publishers. Is there some killer twist that we havent thought about? Some unrevealed aspect of the business model which will make the publishers look favourably on the development. I have my doubts....

Wednesday, July 09, 2008

Charkin Blog in book form

The Charkin Blog went silent 8 months ago, and there have been rumours that it will soon appear in print. The rumour was confirmed when we received a request for permission to include a snapshot, originally taken from our web site, so that the thumbnail (of a Berkshire Publishing reference work) could appear in the book publication. The email requesting permission was very polite and of course we promptly granted permission. This is how the project was described.

"In September 2008 Pan Macmillan will publish Charkin Blog: the Archive, by Richard Charkin, an edited print-on-demand version of the blog he published at http://charkinblog.macmillan.com/default.aspx while chairman of the company."
and they asked for blanket permission in all territories. But there was no mention of digital rights, so does this mean that there will not be a digital edition? I hope not, since I am a great believer that anything that is worth printing is worth having in digital format. On the other hand I am more of a believer of exact editions where the digital editon exactly matches the print edition: can an exact edition go in the other direction? How can one compensate for all the missing links, the immediacy of navigation etc? We intend to buy the book to find out.

The Charkin blog was a very good read while it lasted, it will be interesting to see if it can work in volume form. Of course, Macmillan as a large publisher would take the permissions issue very seriously, but can you imagine how many permissions emails they will have had to generate? It is a reminder that blogs just could not exist if every blog re-usage required permission. Publishing and blogging on the web thrives because the reins are a little bit looser. Publishers who insist that copyright issues must all remain 'opt in' (ask before you use) rather than 'opt out' (if you object I will take down) are living without the web.

Thursday, July 03, 2008

Copyrights and Back Issues

An important decision on a dispute that has been rumbling for years in the US

Back-to-back rulings by federal appellate courts in Atlanta and New York favoring the National Geographic Society will allow magazine and newspaper publishers to transfer their published archives to computer discs and sell them commercially without infringing on freelance contributors' copyrights. ....... the 11th Circuit majority determined that because National Geographic's digital library reproduced complete magazine issues "exactly as they are presented in the print version," publishers retained the privilege of reproducing them under federal copyright laws without renegotiating contracts with their writers and photographers. (see report at Law.com)

The British courts may not follow the American courts in this decision, but we have always felt that it is common sense that a magazine which is an exact and faithful replica of a print edition should be treated in the same way as the print issue from the point of view of licensing and copyrights (we especially like the phrasing of the court "exactly as they are presented in the print version," could the judge have taken out a subscription to Exact Editions before he coined his phrase?).

Photographers and picture agencies will feel that profitable exploitation of digital editions should have some beneficial consequences for photographers and illustrators. That should happen, provided publishers are able to develop effective digital publishing strategies. Not being able to include photographs or illustrations for rights reasons is not a practical way of developing a digital service. Without an effective digital publication publishers will be less able to afford fees for photos and illustrations. Lets hope that the British courts and the British picture agencies take notice of these specifically american rulings. Movement along these lines is in the interests of all the rights holders.

Monday, June 30, 2008

Google Book Search is it Rudderless?

Some librarians are complaining that they have been used by Google (hat tip to David Rothman) and they worry that Google is now losing interest in the library market. Google certainly seems to have backed away from publishers (no longer attending the main trade fairs, not making a concerted pitch towards them). So is the Google Book Search project losing its direction? Here are three guesses about that:

  1. Google has made tremendous progress with the data capture project. There are no public aggregate statistics, but Michigan passed the 1m books target earlier this year, so I would estimate that Google Book Search has over 4 million titles contributed by the libraries, plus perhaps 1 million from publishers (Springer will have over 30,000 now). (If anyone has any good data on this please add as a comment). So in this sense Google Book Search is working very well as a powerful data-service, but no one at Google has a good idea about how to drive the books operation as a commercial service. Text-driven advertising is not going to monetise most of the books in the collection. GBS is a computer science project which is working really well but it is hard to see how it can become a pay-for-itself proposition. I think this is why Microsoft pulled out of its 'shadow Google Book Search' play, a month ago. It didnt see the point of being second best at something which might not have a commercial justification even if they were 'first best'. Microsoft doesnt believe in fundamental computer science engineering the way that Google does. The GBS project is not losing its direction, it was just a 'moon shot' with a long time to come to fruition. Come back in 10 years time. By then the computer science on handling a 50 million volume text database will be part-done. Google is not being slow or neglecting anybody. Its just a huge project.
  2. Google is waiting until the legal mess around the status of in copyright titles is cleared up before putting a clear commercial direction on the Google Book Search service. So GBS is not so much rudderless as in 'legal limbo'. The direction will be resolved as part of a settlement with the publishers and authors and this settlement will give Google a big head start in providing a commercial book service, sanctioned by the publishers. Peter Brantley is worried that this may be where we are. But I am not convinced, because I suspect that Google is more interested in prolonging and delaying the legal issues than it is in reaching a settlement. Google gains by prolonging the dispute, because its hard to negotiate what it wants, and in the end technology will 'prise open' the copyright position that publishers (and agents) will never agree to surrendering. Publishers and 'old fashioned' authors and agents want to maintain the requirement that texts may only be copied with explicit permission. Google doesn't think like that and takes the view that texts like any physical object can be digitised, and that the digital object can be computed without permission, (though accepting that secondary commercial exploitation may need explicit permission). So Google is not expecting a legal victory, or a negotiated agreement anytime soon. If we think that the Google Book Search project is all about delivering books in the largest possible numbers, in the best possible format, to the greatest number of human readers, they had better get on and settle the disputes and start rolling out the commercial services before Amazon has walked off with all the commercial advantage using its Kindle. Google is just being too slow to get commercial because of legal hassles.
  3. Finally, there is the possibility that the Google project really is 'rudderless' and they would have been better off taking on board explicit bibliographic and librarianship skills from the begining and they they can still do this and need to change tack in order to do so. They would need to re-orient and declare open some of their proprietary positions, perhaps they could co-opt Brewster Kahle, but an 'open source' revision to their project might have some benefits. Having a complete input from librarians and using the objective of creating a free open library of all no-longer-in-copyright material would have been a worthy target for Google and perhaps they will revert to operating in this way, if they decide that the legal obstacles to a fully commercial service of the kind that they are building are perhaps too fraught and tricky for them. Google Book Search is somewhat rudderless, because they have not defined the appropriate goal for their massive enterprise.
I tend to alternate between (1) and (3), but that may mean that I am wrong about (2) also. It could be that Google is well advanced with plans for a commercial version of Google Book Search and will launch a 'pay per view' implementation next week. Who knows?

Wednesday, May 07, 2008

Amazingly Compilcated Viewability Restrictions

One hesitates to recommend a 50 minute podcast. But this chat at Talis's The Library 2.0 Gang had some interesting comments. The focus of the discussion was on the recently release Google Book Search Viewability API, and there seemed to be fairly general agreement that it was a step in the right direction but not yet enough.

Google needs to loosen up a bit and open up some more to enable some really interesting literary mashups to take hold. There were some particularly interesting contributions from Frances Haugen, a Google Book Search Product Manager. She spoke passionately and idealistically about the aims of the Google Book Search project. She agreed that an API which allowed some server-side interactions would be a good idea. But in passing she noted that there were legal issues and limitations. I was particularly struck by her comment that the Google rules on access limitations on international viewability are 'amazingly complicated'.

Google's lawyers are being strict on the extent to which works which may not be public domain in other countries can be accessed/viewed outside the US (but the majority almost certainly are in most places). It is not surprising that such a set of house rules limits the extent to which a useful API can be defined. The problem is not so much copyright, as the differing terms of copyrights in different jurisdictions and the penumbra of uncertainty about who has what.

Google Book Search will work better for Google if they can outsource the business of establishing who has clear title in a text and where. That could mean negotiating with publishers before digitising the text. It may come to that, and Google Book Search will be more comprehensive and more accessible when it does so.

Monday, April 21, 2008

Hate to be OUP, CUP and Sage

Dorothy Salo who pens an insightful blog, Caveat Lector (what a brilliant name for a blog) has the headline "Hate to be Georgia State". She is referring to the controversy surrounding the action against that University, for blatant copyright infringement, taken by the two biggest University Presses, OUP, CUP and Sage, one of the most respected publishers of high level social science. The action is backed, perhaps underwritten, by the AAP.

If you read the formal Complaint for Declaratory Judgement and Injunctive Relief, you may feel a tremor of concern for Georgia State's officers and librarians. As I read the documents prepared by the publishers' lawyers the case looks pretty strong. But of course that is what lawyers do. They prepare suits that make a strong case. We have not yet heard the Georgia State side of the case. I suspect that the dispute, even if it is won or more probably settled out of court, will in the long run be as bad for publishers as it appears to be for the university. No publisher wants to sue good customers and Georgia State is certainly a good customer for many publishers. It can not be a happy sight for two prestigious British university presses to be suing an American State University. Questions will be asked.

This story was broken in the New York Times a week ago, and their reporter elicited an interestingly specific comment from CUP's Editorial Director, a well respected publisher:

Frank Smith, editorial director for academic books at Cambridge University Press, said that for electronic use in a course, Cambridge typically charges 17 cents a page for each student, and generally grants permission for use of as much as 20 percent of a book.
If you do the math on 17c per page, per student, you get an astronomical price for a course pack in a popular subject. A thousand page e-reserve for a course taken by 100 students is going to cost the university $17,000 per annum. This pricing policy, per student per page, per annum, may once have been appropriate (when universities produced local print anthologies in lieu of buying books) but it is inherently unreasonable for the kind of ambient access that web-based teaching requires. The libraries for the most part already have the books and publishers absolutely need to find ways of encouraging and facilitating reasonable access to those books and periodicals. Suing the university is a terrible idea, nor is it sensible to price to the limit on what students might be recommended to read. E-reserves, digital course packs, are good ideas and publishers should not be pricing in such a way as to prevent them from working.

Tuesday, April 01, 2008

Section 108 Copyright report

A band of experts has spent a lot of time constructing a thoughtful report on possible reforms to the US law of copyright. The excellent Open Access News blog gives you the essential links.

I have not read it all; it is a lot of reading -- 150 dense pages. And it is all recommendations for the Copyright Office to consider before possibly asking Congress to change the law. But whoever was charged with finalising the report for publication did not think very carefully about how it would be read. Perhaps the publishers on the panel were not closely enough involved in the final stages. Under the normal default settings for Preview/PDF on my Mac, the light blue box in which the recommendations appear completely occludes the text. It is wearisome to have to cut and paste the text out of its blue boxes in order to read it.

It would be mischievous to suggest that they did not want their recommendations to be read (so they put it with a blue background which disagrees with monitors and photocopying systems). The specific recommendation on this page will give scant encouragement to Google in its arguments with Publishers and others on the Library project. The recommendation rather explicitly disowns the Google Book Search modus operandi and the Google grab of copies which have not been explicity sought from rights holders.


















Here is the text of the invisible recommendation that will be disagreeable to Google -- all three sub clauses are each enough to give GBS some heartburn.

. Section 108 should be amended to allow a library or archives to authorize
outside contractors to perform at least some activities permitted under
section 108 on its behalf, provided certain conditions are met, such as:
a. The contractor is acting solely as the provider of a service for which
compensation is made by the library or archives, and not for any
other direct or indirect commercial benefit.
b. The contractor is contractually prohibited from retaining copies
other than as necessary to perform the contracted-for service.
c. The agreement between the library or archives and the contractor
preserves a meaningful ability on the part of the rights holder to
obtain redress from the contractor for infringement by the contrac-
tor.
(Section 108 Study Group Report p. iv of the Executive Summary)


If this is the way copyright experts are thinking, it would seem to be very clear that Google needs to find a way of backing out of its aggressive stance on rights in library copies.

Wednesday, March 19, 2008

All you can Eat Music

Apple through its iTunes and Nokia through its "Comes with music deal" are preparing to offer unlimited access to the major music companies catalogues through monthly subscription plans. See today's report in the Financial Times. Such services would be pitched at the $7-8 per month level, with Apple apparently needing a large slice of the revenue (this we understand is Apple's style).

I wonder if this will work for music? I wonder if it would work for those of us with non-mainstream interests? Surely premium music would elude this framework? One can al least think about such a scheme in the case of music since the 4 majors control a large part of the recorded music pie. Book publishing is incredibly much more fragmented (also by language), so it is scarcely conceivable that a technology platform could negotiate a global rights pie with umpteen different major print publishers.

Mind you its an idea which might be attractive to some. Would such an over-arching subscription scheme be one way for Google to negotiate a settlement with the publishers and author's societies which are opposing its Google Book Search project? Google would then need to start charging subscriptions for full access to the in-copyright resources in its database.

If this is to be the Google digital books charging model, I would guess that the other players in the market have some years in which to test alternative approaches. Which is what we are doing with the Open Searching/Subscription Content Reading offering that we this week launch for Berkshire Publishing.

Thursday, February 07, 2008

Clipping from Books

The Exact Editions platform supports a *Clipper* which helps you to cut a selection from one of the JPEGs which show the detail of a magazine/book. The Clipper tool also provides information on the publication and a link back to the source. The Clipper was designed for magazine columns and it now works well with books, especially if you need to blog a short quotation:



The limitation of no more than 12% of a page, per clipping, remains. It is there as a marker for traditional views on the permitted extent for 'fair use', or 'fair dealing'. One supposes that as publishers realise the advantages of sanctioning and enabling fair use through Clippings, especially those which provide a citation, there will be pressure for this limit of 12% to be raised.

Wednesday, February 06, 2008

University of Michigan has 1 million Google Book Searchable books

What an amazing achievement:

Here is the millionth book.
Paul Courant's blog about the milestone.

In a very few years all 7.5 million bound volumes (that must include magazines and newspapers) will be searchable, by anyone, anywhere. That is right the University of Michigan will allow searching of its collections by anyone (not reading of entire volumes or even pages, for reasons of copyright, but searching). It can hardly be imagined what potential this has for scholarship (especially in the humanities). Michigan's reputation will soar (rightly). Universities are highly competitive and international competition is getting more urgent, this is a knowledge race. Michigan will be at the head of a chasing pack.

We can be sure that this is putting competitive heat on universities and their planners everywhere. The book publishing industry and the magazine industry are well behind in rising to meet this challenge, whereas the scientific periodical publishers have matters well in hand (not quite as well in hand as Google).

Friday, December 14, 2007

The Optimal Term of Copyright

Rufus Pollock, a researcher at Cambridge University, has produced a rather brilliant but difficult, because formal and mathematical proof (whilst yet embodying many empirical propositions), that the term of copyright should be much lower than it is now, and that it should in general become shorter as technology advances. He suggests that the optimum term of copyright should be about 15 years from creation. The argument is complex, but he provides some neat informal guidance vis:

........consider the situation with respect to books, music, or film. Today, a man could spend a lifetime simply reading the greats of the nineteenth century, watching the classic movies of Hollywood’s (and Europe’s) golden age or listening to music recorded before 1965. This does not mean new work isn’t valuable but it surely means it is less valuable from a welfare point of view than it was when these media had first sprung into existence. Furthermore, if we increase protection we not only restrict access to works of the future but also to those of the past.
As a result the optimal level of protection must be lower than it was initially in fact it must fall gradually over time as our store of the creative work of past generations gradually accumulates to its long-term level. Forever Minus a Day?...

Pollock does not consider the related question: how will the efficient and optimal pricing of information services respond to changes in technology which reduce barriers to access? As more information becomes available how should a commercial information supplier i.e. a publisher, price his subscription services? How much should be given away and what to charge for that which is sold?

Wednesday, October 17, 2007

Google and Copyright

Google have just announced a new set of Content ID tools which will help copyright owners protect their content on the YouTube platform. More detail is given here. These new policies and the copyright ID platform may enable Google to shrug off or negotiate a way out of the onerous suits it faces from Viacom and the UK's Premier League. But it also seems to back away from the idea of establishing a 'fair use' of video clippings or quotations -- a contentious issue which is at the heart of the YouTube success. The new Google approach appears to give copyright owners total control over the distribution of their video content.

Google will have to make similar proposals to the owners and custodians of literary copyrights. We can expect a comparable "highly complicated technology platform -- [with] content identification tools" to be in preparation for the Google Book Search platform (they already have much of it in place already) . It would be hard to go before a judge saying that literary copyrights are going to be treated differently from video copyrights. I predict this is going to lead Google to handing a lot more power to its Publisher partners and less leeway to its Library partners in the construction of the Google Book Search 'library'.

Some of the Google statements are quite striking and humble:

No matter how accurate the tools get, it is important to remember that no technology can tell legal from infringing material without the cooperation of the content owners themselves.....The best we can do is cooperate with copyright holders to identify videos that include their content and offer them choices about sharing that content. As copyright holders make their preferences clear to us up front, we'll do our best to automate that choice while balancing the rights of users, other copyright holders, and our community as a whole. [See videoID-about]
It is especially tricky to see how one can automate the choices of copyright holders whilst balancing the rights of users....As John Batelle wonders its not at all clear what happens to fair use. But book publishers will certainly welcome the idea that they might be given more control 'up front'. Its what they have been asking for all along.

Trouble is that literary copyrights can be a lot more confused and complicated even than video copyrights. All serious literary publishing requires that scope be given to 'fair use'.


Friday, October 05, 2007

Radiohead and the Future of Print

OK, I know that is a mildly ridiculous headline. But hear me out. The Oxford-based band Radiohead have made a move which is giving the music industry the jitters. Radiohead are launching a new album without the help of the majors and they are asking their fans to pay what they want to pay for it ("its up to you") if they download the music digitally. They are also selling an expensive package of physical goods CD/DVD vinyl disks etc. for £40/$80. Michael Arrington thinks this marks a turning point in the inevitable march of music towards free. Jeff Gomez (at the Print is Dead blog) notes the control with which Radiohead have managed this publication process themselves:

So with one fell swoop Radiohead shatters half-a-dozen rock-star rituals, and further makes the existence of record labels a questionable thing in a digital age.
Jeff does not ask whether print publishers are similarly vulnerable. But the question hangs in the air (it is a Print is Dead blog, right?). On the other hand, maybe print publishers are in a better position. After all, in a curious way the Radiohead exercise is lavishing particular attention on the packaging and the physical product. There is even a book in the package as well as the CDs and the vinyl. One can see that £40 package becoming a collectors item. It is possible that as we embrace the digital, the quality and the value of print magazines and books will actually increase (though as a luxury item) whilst the digital versions become the most popular and evanescent form in which the works are enjoyed.

Monday, October 01, 2007

Widgets and Namespaces

Having just had four days holiday without web-access, one realises that things move too quickly right now. Here is some stuff that I hope to catch up with:

Tim O'Reilly posts about Adobe opening up Share, a generalisable document widget system. A kind of YouTube for documents. Looks interesting and one more copyright challenge for publishers and authors to think about. Yet another reason for keeping close control of those PDF files before they get shared in ways that were not possible a few years ago! But, I wonder whether Adobe have positioned this quite right: {I only raise the question} - perhaps the 'Share' concept is missing the revolutionary point about the YouTube analogy. YouTube was viral because it was very easy to share videos that way, but I reckon that the key step forward with YouTube (and similar services) is that they have shown how it is possible, useful, viral and creative to QUOTE videos. Quoting is much more productive and creative than another potentially abusive sharing technology......The problem of standards, of 'fair use' and techniques of digital quotation through the web (which is one step beyond citation and mere linking) has not yet been solved.

The Exact Editions/Berkshire announcement drew an insightful and appreciative response from Outsell:

There are services which offer similar opportunities, Amazon’s Search Inside! being a prime example. However, the functionality is limited when compared to Exact Editions, both for the publisher and the end user - through Amazon, users can only search inside one book at a time, for instance, and can never look at every page of a single title. This move from Berkshire may indicate that book publishers are becoming less cautious about exposing their content on the web, and more likely to start experimenting in earnest with ways in which the networked environment can not only help to boost sales, but can also deliver valuable new functionality around existing content.
I suppose Kate Worlock's way of putting this point makes it clear that the Exact Editions service is also doing what the new Adobe system is doing, but our system makes it easy for the publisher to control and brand the content in the network environment, and with our clipper the quotation carries the attribution/citation with the quotation. Her conclusion is the essence: "Not only boost sales but deliver new functionality....." I will remember to reuse that phrase (with proper attribution to Kate Worlock of course).

This looked interesting on harmonising meta-data: Lorcan Dempsey blog, on why we need a Strunk and White for namespaces.

Finally, just before I took my break, Richard Charkin said goodbye to Macmillan and set sail for Bloomsbury. The trade press reports it here. Richard is such a talented academic and STM publisher that I will lay long odds that Bloomsbury will now make some forays in that direction. STM publishing has become way too congested, predictable and costive. Time for a shakeup and some innovation. Bloomsbury could do that.