Showing posts with label Google Books. Show all posts
Showing posts with label Google Books. Show all posts

Wednesday, September 30, 2009

Banned Books Week and Fighting the Last War Over Again

This week the ALA (and many libraries) are celebrating Banned Books Week. But as a recent Wall Street Journal editorial mentioned, "banning" is not the same as challenging, and it is challenging that the ALA seems to be condemning. And most challenges are not successful. The small number of challenges (and the even smaller number of successful ones) tells me that we have won this battle.

Why are we still devoting so much attention to this issue? We have limited resources, and the world has a limited attention span. Why not spend our time on a battle that we are still fighting?

At this year's Access meeting in Canada, author Cory Doctorow suggested librarians use what influence we have to promote rational copyright laws-- you know, the kind that protect the greater good of society (as copyright was originally intended), instead of using copyright to shore up old business models with increasingly draconian penalties.

The "orphan books" being fought over in the Google Book Settlement are mostly the result of continually extending copyright coverage beyond normal commercial viability. Libraries need to be making the case that publishers are not the only party at the table here-- someone has to stand up for the rights of society to it's own cultural legacy.

The old saying was that generals were always preparing to fight the last war over again, rather than adapting strategies and tactics to a changing world. We should not be guilty of the same mistake. The battle for librarians today is over laws that restrict access to information-- not through banning or burning, but through one-sided legislation.

Powered by ScribeFire.

Saturday, February 07, 2009

E-Books Go Mobile

The use of mobile phones as ebook readers in common in Japan, and is growing in the US and elsewhere. A number of publishers are making the leap (last month, for example, Books on Board announced their catalog of 20,000 books would be available for the iPhone).

Now comes word that Google is entering this market. Google has
launched a mobile phone version of Google Book Search that could could eventually grow to include the 1.5 million public domain books scanned as part of their digitization project.

The books currently exist as scanned images-- these mobile versions will be text created through optical character recognition. Where the computers produce only garbled text, readers can click on the sport to retrieve that part of the scanned image.

Not only does this open up smart phones to the vast public domain resources harvested through Google's digitization project, but this also shows that OCR technology has improved to the point where Google (at least) thinks it is ready for prime time.

Scanned images are just the first phase of bringing books into the digital world. Ebooks need to exist as digital text, and human-based projects like Project Gutenberg are probably proceeding too slowly. OCR is vital to the next phase of mass-digitization. We'll soon see if Google's timing is right.


Monday, October 13, 2008

Google Books Libraries Establish HathiTrust Repository


LISNews reports that the University Libraries involved in the Google Books project have established a central repository for the 2 million digital books scanned thus far.

Called the
Hathi Trust, and with lead participants the Universities of Michigan and Indiana, the repository will serve as a backup should Google go out of business, or lose interest in the project (as Microsoft did with its Live Books project). "Hathi" is the Hindi word for elephant, who are famously good at remembering things.

A large
scale search feature is planned for the repository, as are a number of other intriguing features, including an API to allow partner libraries to integrate the collection into their local systems, access mechanisms for the disabled, the ability to publish virtual collections, the ability to add (or "ingest") non-Google content, and a public discovery interface.

As full-text search looks to replace the traditional library search methods over the next few years, it's great to have a non-corporate source for the search data. Congrats to the HathiTrust team!


Friday, May 23, 2008

Microsoft Ends Book and Article Digitization, Shuts Down Live Search Books

Just saw the announcement that after digitizing 750,000 books, Microsoft is pulling the plug on Live Search Books and Live Search Academic, and it says it's leaving the field to libraries and publishers-- but doesn't mention a rather large competitor still in the game (Google).

Whether this is more about Microsoft's faltering Live efforts, or really shows that there's not enough money to be made here for private industry, is hard to say. I expect Google will answer this question for us over the next few years.


Wednesday, April 16, 2008

"The Million Books Problem" and Silent Movies

On my commute home last night I listened to an Educause podcast from Scott Kirsner's keynote at NERCOMP 2008 entitled "What Innovators Can Learn From Hollywood".

Since there was a lot of traffic, I had plenty of time to think about what he was saying-- basically how the many technical changes in the movie business have succeeded (when they have succeeded) in the face of stiff resistance from people who were comfortable with the established tools (indeed who were often geniuses in their use).


Advocates of new technologies usually underestimated how long change would take (
Technicolor was introduced in 1917 and took decades to succeed in the market). On the other hand, sometimes even the imagination of technology boosters falls short. When The Jazz Singer" introduced the concept of talking pictures, it was thought of as a niche technology for musicals. Dramas, comedies, and other films worked fine as silent films-- a whole generation of actors and film makers had created an expressive and often beautiful body of work without muddying up the visual with sounds. Why would anyone need to add a soundtrack?

I'd never thought of it before, but the real revolutionary thing was not so much the invention of the capability of making talking pictures as it was that the market quickly decided it only wanted talking films. And in a year or two, that's all that was being produced.

Which leads me to consider the library business today. We're on the verge of an age of pervasive, free access to the digitized contents of the million books being processed by Google, the Open Content Alliance, and others. In this kind of world, how many libraries need to duplicate this access in print? Will electronic access be enough? Will print become the niche market? This "Million Books Problem" is getting a
lot of attention in library circles today.

The key question is will this technology take decades to become the norm, or just a few years? I've been thinking we'd have a comfortable number of years to adapt to changing demands, but what if we don't?