Showing posts with label digitization. Show all posts
Showing posts with label digitization. Show all posts

Sunday, March 01, 2009

recent reading

Some reports and posts that caught my attention recently:

The Andrew W. Mellon Foundation released a progress report from the DuraSpace project, a joint project of the DSpace Foundation and the Fedora Commons.

"MetaTools - Investigating Metadata Generation Tools" from JISC.

Merrilee Proffitt from RLG/OCLC posted on the "Legal and Ethical Implications of Large-Scale Digitization of Manuscript Collections" symposium at UNC-Chapel Hill. Posting Part 1 and Posting Part 2.

Andrew Richard Albanese published an article for Library Journal called "Institutional Repositories: Thinking Beyond the Box." It's a very balanced presentation of a number of points of view on the failure and success of IRs.

Wednesday, December 17, 2008

ICDL adding European collections

From an article on Forbes.com, the International Children's Digital Library (ICDL) announced a partnership with the Taliaferro Family Fund to increase the number of European children's titles in the collection. The Elias Project will target three collections in Europe: the Norwegian Children's Book Institute in Oslo, Norway, the International Youth Library in Munich, Germany , and the National Center for Children's Books in Paris, France.

After reading the article, I check in at the ICDL site, which I hadn't visited in a few months, and noticed two other news announcements: ICDL and the Google Book project will be sharing public domain children's book titles; and ICDL has launched an iPhone app with full access to the collection, a new titles features, and an offline mode and an airplane mode. It's great to see such a worthwhile project making such advances in collection building and in adding new services.

(I didn't see a press release about the European project on the ICDL site. I saw the press release on some other sites, so I assume it's meant to be out there.)

Thursday, November 20, 2008

Europeana

The prototype site for Europeana, the European digital library funded by the EC, is set to launch today, November 20, 2008.

The initial collection of 2 million items comes from museums, libraries, archives, and audio-visual collections, and includes paintings, maps, videos and newspapers. The interface is in French, English, and German, with more languages planned. Highlights include the Magna Carta from Britain, the Vermeer painting “Girl With a Pearl Earring” from the Mauritshuis Museum in The Hague, a copy of Dante’s “Divine Comedy'" and newsreel footage of the fall of the Berlin Wall. Content is free of copyright restrictions.

New York Times and EU Business articles provide more details. The goal is to add 10 million items into the digital library by 2010, at a price of 350 million to 400 million euros.

Sunday, November 16, 2008

arl guide to the google book search settlement

The Association of Research Libraries has created "A Guide for the Perplexed: Libraries & the Google Library Project Settlement," a 23-page document intended to help libraries understand the impact of the proposed Google Book Search settlement.

google book search session at dlf

I was going to spend some time transforming my notes from Dan Clancy's session on Google Book Search from the DLF Fall 2008 Forum into more coherent prose, but for the sake of timeliness, I'm going to post them as is.

  • 20% of the content in Google Book Search is in the public domain, 5% is in print, and the rest is in an unknown “twilight zone” -- unknown status and/or out-of-print.
  • 7 million books scanned, over 1 million are public domain, 4-5 million are in snippet view.
  • Early scanning was not performed at an impressive rate, and it took way longer than expected to set up.
  • Priorities are working search quality, and exposure to google.com.
  • Search is definitely not solved and “done,” and is harder given the big distribution of relatively successful hits.
  • They are working to improve the quality of scanning and the algorithm to process the books and improve usability. They admit that they still have work to do, especially with the re-processing of older scans.
  • Data to support Long Tail model is right.
  • Creating open APIs, including one to determine the status of a book, and a syndicated viewer that can be embedded.
  • Trying to identify the status of orphans, and release a database of determinations. But institutions need to use determinations to guide their decisions, not just follow them because “Google said so.”
  • On the proposed settlement agreement: Google thought they would benefit users more to settle than to litigate.
  • The class is defined as anyone in the U.S. with a copyright interest in a book, in U.S. use. (no journals or music)
  • For all books in copyright, Google is allowed to scan, index, and provide varying access models dependent upon the status of the book -- if in print or out-of-print. Rights holders can opt out.
  • 4 access models: consumer digital purchase (in the cloud, not downloads – downloads are not specifically included in agreement); free preview of up to 20% of book; institutional subscription for the entire database (site license with authentication, can be linked into course reserves and course management systems); public access terminals for public libraries or higher ed that do not want to subscribe (1 access point in each public library building, some # by FTE for high ed institutions) which allows printing (for 5 years or $3 million underwriting of payments to rights holders).
  • Books Rights Registry to record rights, handle payments to rights holders. It can operate on behalf of other content providers, not just Google.
  • Plan to open up government documents, because they feel that the rights registry organization will deal with the issue of possible in-copyright content included in gov docs, which kept them from opening gov docs before.
  • Admits that publishers and authors do not always agree if publishers have the rights for digital distribution of books. Some authors are adamant that they did not assign rights, some publishers are adamant that even if not explicit, it's allowed. The settlement supposedly allows sharing between authors and publishers to cover this.
  • What is “Non-consumptive research”? OCR application research. Image processing research. Textual analysis research. Search development research. Use of the corpus as a test corpus for technology research, not research using the content. 2 institutions will run data centers for access to the research corpus, with financial support from Google to set up the centers.
  • What about their selling books back to the libraries that contributed them via subscriptions? They will take the partnership and amount of scanning into account and provide a subsidy toward a subscription. Stanford and Michigan will likely be getting theirs free. Institutions can get a free limited set of their own books for the length of the copyright of the books. They can already do whatever they want with their public domain books.
  • They will not necessarily be collecting rights information/determinations from other projects for the registry. In building the registry, they are including licensed metadata (from libraries, OCLC, publishers, etc), so they cannot publicly share all the data that will make up the registry. But they will make public the status of book that are identified/claimed as in copyright.
  • If Google goes away or becomes “evil Google,” there is lots of language in contracts and settlement for an out.
  • The settlement is U.S. only because the class in the suit was U.S. only. Non-U.S. terms are really challenging because many countries have no concept of class-action, and there is a wide variation of laws.
  • A notice period begins January 5. Mid 2009 is the earliest time this could be approved by the court.

Tuesday, October 28, 2008

google book search settlement agreement announced

Today it was announced that Google has reached a settlement in the lawsuit filed by the Authors Guild, the Association of American Publisher, and a group of individual authors.

Some of the details are available at Google. The changes that I am the most interested in are these:

"Until now, we've only been able to show a few snippets of text for most of the in-copyright books we've scanned through our Library Project. Since the vast majority of these books are out of print, to actually read them you'd have to hunt them down at a library or a used bookstore. This agreement will allow us to make many of these out-of-print books available for preview, reading and purchase in the U.S.. Helping to ensure the ongoing accessibility of out-of-print books is one of the primary reasons we began this project in the first place, and we couldn't be happier that we and our author, library and publishing partners will now be able to protect mankind's cultural history in this manner."

...

"The agreement will also create an independent, not-for-profit Book Rights Registry to represent authors, publishers and other rightsholders. In essence, the Registry will help locate rightsholders and ensure that they receive the money their works earn under this agreement. You can visit the settlement administration site, the Authors Guild or the AAP to learn more about this important initiative."
I'm all for more access to these books and for rightsholders to get their due, but what does it mean to assign a value to them?

They also plan to offer subscriptions: "We'll also be offering libraries, universities and other organizations the ability to purchase institutional subscriptions, which will give users access to the complete text of millions of titles while compensating authors and publishers for the service." I have mixed feelings -- the subscription model is not an unusual one, and libraries have certainly provided digitized materials from their collections for paid subscription services before, i.e., with ProQuest. I wonder if the partners will get any share in the compensation for providing the content for the service?

I'm currently at an Open Content Alliance meeting and I'm looking forward to what I am sure will be many discussions among the attendees today.

EDIT: There's now a joint press release from the University of Michigan, the University of California, and Stanford University, a FAQ from the American Association of Publishers, a Google rightsholders site, a Google blog post, in addition to the site above and the press release.

Wednesday, October 15, 2008

First Monday article on Google Books and OCA

The newest issue of First Monday (volume 13, number 10, 6 October 2008) has an interesting article by KalevLeetaru -- "Mass book digitization: The deeper story of Google Books and the Open Content Alliance."
The article compares what is publicly known about the Google Book and OCA projects.

From the conclusions:

While on their surface, the Google Books and Open Content Alliance projects may appear very different, they in fact share many similarities:

  • Both operate as a black box outsourcing agent. The participating library transports books to the facility to be scanned and fetches them when they are done. The library provides or assists with housing for the facility, but its personnel are not permitted to operate the scanning units, which must be staffed by personnel from either Google or OCA.

  • Neither publishes official technical reports. Google engineers have published in the literature on specific components of their project, which offer crucial insights into the processes they use, while talks from senior leadership have yielded additional information. OCA has largely been absent from the literature and few speeches have unveiled substantial technical details. Both projects have chosen not to issue exhaustive technical reports outlining their infrastructure: Google due to trade secret concerns and OCA due to a lack of available time.

  • Both digitize in–copyright works. Google Books scans both out–of–copyright books and those for which copyright protection is still in force. OCA scans out–of–copyright books and only scans in–copyright books when permission has been secured to do so. Both initiatives maintain partnerships with publishers to acquire substantial in–copyright digital content.

  • Both use manual page turning and digital camera capture. Large teams of humans are used to manually turn pages in front of a pair of digital cameras that snap color photographs of the pages.

  • Both permit libraries to redistribute materials digitized from their collections. While redistribution rights vary for other entities, both the Google Books and OCA initiatives permit the library providing a work for digitization to host its own copy of that digitized work for selected personal use distribution.

  • Both permit unlimited personal use of out–of–copyright works. While redistribution rights vary for other entities, both the Google Books and OCA initiatives permit the library providing a work for digitization to host its own copy of that digitized work for selected personal use distribution.

  • Both enforce some restrictions on redistribution or commercial use. Google Books enforces a blanket prohibition on the commercial use of its materials, while at least one of OCA’s scanning partners does the same. Google requires users to contact it about redistribution or bulk downloading requests, while OCA permits any of its member institutions to restrict the redistribution of their material.

From the section on "Transparency"
A common comparison of the Google Books and Open Content Alliance projects revolves around the shroud of secrecy that underlies the Google Books operation. However, one may argue that such secrecy does not necessarily diminish the usefulness of access digitization projects, since the underlying technology and processes do not matter, only the final result. This is in contrast to preservation scanning, in which it may be argued that transparency is an essential attribute, since it is important to understand the technologies being used so as to understand the faithfulness of the resulting product. When it comes down to it, does it necessarily matter what particular piece of software or algorithm was used to perform bitonal thresholding on a page scan? When the intent of a project is simply to generate useable digital surrogates of printed works, the project may be considered a success if the files it offers provide digital access to those materials.
To me, that paragraph gets at the key issue in discussing and comparing the projects -- are books being scanned in a consistent way and being made accessible through at least one portal, enforcing current rights restrictions? Yes? Then both these projects are, at a basic level, successful and provide a useful service.

Yes, there are issues to quibble with for both projects. More technical transparency is desirable for both projects. Both have controlled workflows that limit what can be contributed to the projects in different ways. There are aspects of the Google workflow that Google contractually requires its partners to keep secret. That's their right to include in their contracts, and a potential partner's decision to make if they find it objectionable and therefore choose not to participate. Each documents and enforces rights in different ways and to different extents -- we should be looking to standards in that area. Each sets different requirements for allowing reuse. If only there could be agreement.

One note on preservation. Neither projects are preservation projects -- they're access projects. Even if there were something we could point to and say "that's a preservation-quality digital surrogate" -- if such a concept as "preservation-quality" exists -- neither project aims for that. Both projects do, however, allow the participating libraries to preserve the files created through the projects. These files should and must be preserved because they can be used to provide digital modes of access, and, in some cases, they may be the only surrogates ever made if the condition of a book has deteriorated. Look at the HathiTrust for more on the topic of preserving the output of mass digitization projects.

And one note about the Google project providing "free" digitization for its participants. Yes, Google is underwriting the cost of digitization. But each partner library is bearing the cost of staffing and supplies for project management, checkout/checkin, shelving, barcoding, cataloging, and conservation activities, not to mention storage and management of the files. The overall cost is definitely reduced, but not free.

Wednesday, September 17, 2008

smithsonian digitization initiative

There's an announcement on CNN that the Smitshonian plans to put its 137 million object collection online. The new Smithsonian Secretary G. Wayne Clough said in an interview that they do not yet know how long it will take or how much it will cost to digitize the full 137 million-object collection and will do it as money becomes available. A team will prioritize which artifacts are digitized first. They plan to focus on making the collections usable for the K-12 audience.

When I was at the Smithsonian yesterday for David Weinberger's talk, this seemed to be a buzzing topic of discussion among audience members; one Smithsonian employee even mentioned it in a question to Weinberger, expressing a certain level of surprise.

Monday, September 08, 2008

google newspaper digitization

Google is digitizing newspapers.

Not only will you be able to search these newspapers, you'll also be able to browse through them exactly as they were printed -- photographs, headlines, articles, advertisements and all.

This effort expands on the contributions of others who've already begun digitizing historical newspapers. In 2006, we started working with publications like the New York Times and the Washington Post to index existing digital archives and make them searchable via the Google News Archive. Now, this effort will enable us to help you find an even greater range of material from newspapers large and small, in conjunction with partners such as ProQuest and Heritage, who've joined in this initiative. One of our partners, the Quebec Chronicle-Telegraph, is actually the oldest newspaper in North America—history buffs, take note: it has been publishing continuously for more than 244 years.

You’ll be able to explore this historical treasure trove by searching the Google News Archive or by using the timeline feature after searching Google News. Not every search will trigger this new content, but you can start by trying queries like [Nixon space shuttle] or [Titanic located]. Stories we've scanned under this initiative will appear alongside already-digitized material from publications like the New York Times as well as from archive aggregators, and are marked "Google News Archive." Over time, as we scan more articles and our index grows, we'll also start blending these archives into our main search results so that when you search Google.com, you'll be searching the full text of these newspapers as well.
It's interesting that they're working directly with publishers and with aggregators such as ProQuest to digitize and improve discoverability of back files. That's good news, but do they also plan to work with major newspaper open access projects such as the National Digital Newspaper Program? Are they digitizing any collections in addition to publisher collections?

When I last looked at the Google news archive in September 2006 I found that way too much of the content was pay-per-view, made you pay even if your institution had licensed subscription access, and didn't work with OpenURL resolvers. I don't see that any of that has changed. I hope it will.

Wednesday, September 03, 2008

HathiTrust

The University of Michigan has announced that their MBooks initiative has grown into a shared repository effort called the HathiTrust (pronounced hah-TEE).

HathiTrust was originally a collaboration of the thirteen universities of the Committee on Institutional Cooperation (CIC) to establish a repository for those universities to archive and share their digitized collections. All content to date has been supplied by the University of Michigan and the University of Wisconsin, and Indiana University and Purdue University will soon be contributing their digital materials. 20% of its current content is open access and 80% is restricted. Don't look for a single search interface yet -- it's planned. As they say: "Good, useful, technology takes time.... and the strength and insight born of collaborative work."

The new HathiTrust initiative has been funded for an initial five-year period beginning January 2008, and is now open to other institutions. Partners will be charged a one-time start-up fee based on the number of volumes added to the repository, in addition to an annual fee for the curation of the data. They already support both open access and dark archive materials, and will also do so for new partners.

Their July 2008 monthly report gives a good sense of their activities. It is interesting to note that the only initial ingest workflow supported is the Google partner workflow. That's not too surprising since this work is based on the MBooks project developed in support of Google content workflows. That in and of itself ensures that there many institutions who'll be considering partnership.

This announcement is exceptionally exciting. I look forward to its development as a service.

Wednesday, August 27, 2008

dead sea scrolls

When I was growing up, my mother had a small selection of books displayed between decorative bookends on her coffee table -- a set of 4 art history overview volumes with high quality color reproductions on glossy paper, and a book on the Dead Sea Scrolls. I was fascinated by the volume on ancient art and the book on the scrolls because of their sheer antiquity. I don't remember there being many illustrations in the book, but the story of the discovery of the scrolls was a very engaging one. I don't remember every asking my Mom why she that volume on display, or, if I did, what her answer was.

The New York Times today reports on the project to digitize the Scrolls. It's interesting to read that they plan to create new digital images, as well as digitizing the infrared images created of the scrolls in the 1950s.

Tangentially, there was an article in The Australian a couple of week ago about the conservation and multi-spectral imaging of scrolls from the Villa dei Papyri at Herculaneum.

Tuesday, August 26, 2008

Executive Director of OCA named

A press release went out tonight naming Maura Marx -- founder of the Digital Library Program at the Boston Public Library -- as the first Executive Director of the Open Content Alliance.

“Maura's background in working both inside and outside the library system will help her communicate with a broad public audience the shape of the new public library services in this digital age." said Brewster Kahle, Digital Librarian of the Internet Archive. “Her dynamic style, deep-seated commitment to open principles, and demonstrated success at implementing partnerships and initiatives in the digital space will be a powerful combination in taking the OCA to the next level.”
I met Maura at a meeting this spring, and I know that she's an excellent choice!

Wednesday, August 06, 2008

time for links and nothing more

I'm really swamped these days, and only have time to post some links to things that caught me eye during the past week:

vi.sualize.us seems like a really interesting social bookmarking tool for images. Perhaps they'll learn what delicious learned and give up the tortured . 's.

William Patry stopped blogging
. I'm not surprised if folks thought his personal blog was the word of Google. It's sad that he also decided to erase his archives, but I understand that he didn't want his past postings to live on and continue to be misunderstood.

It seems that Google is making some of its machine-translation technologies and translation management tools available to human translators, at least as a beta. I'm working with a project that requires translation into 7 languages. Managing this process is very challenging, and I've seen some very bad tools for the process.

Following the Digitization and the Humanities Symposium, Jennifer Schaffner and Merilee Profitt wrote a brief report, The Impact of Digitizing Special Collections on Teaching and Scholarship: Reflections on a Symposium about Digitization and the Humanities. The report acts as a summary of the symposium, and also gives some calls to action, especially about metrics for success.

Duke has launched its Open Library Environment Project with Mellon support. Its focus on back-end open tools is worhtwhile, but I'm not sure I know how this will be integrated with other activities in the community.

Tuesday, July 29, 2008

on NYRB article about Google Books

Jean-Claude Guédon and Boudewijn Walraven submitted letters to the New York Review of Books which have been published as "Who Will Digitize the World's Books?" They are commenting on Robert Darnton's "The Library in the New Age", and he responds to their letters.

Thursday, July 17, 2008

international copyright law and digitization

In one of those great synchronicities, I've encountered two publications on international copyright law and digitization, both of which are worth reading.

The first is an Information World Review article entitled "Scan and Deliver" about how issues of copyright clearance have affected the British Library's digitization program. (I keep hearing Adam Ant's "Stand and Deliver" in my head)

The second is the International Study on the Impact of Copyright Law on Digital Preservation just released by the Library of Congress. The report is a joint effort of the Library of Congress National Digital Information Infrastructure and Preservation Program, the Joint Information Systems Committee, the Open Access to Knowledge (OAK) Law Project, and the SURFfoundation.

Tuesday, May 27, 2008

more on the end of Windows Live Search Books

A roundup of commentary:

ars technica (includes comments from Brewster Kahle)

shimenawa/Peter Brantley

TeleRead

New York Times

Friday, May 23, 2008

Windows Live Search Books going away

Peter Brantley forwarded a Microsoft message from the Live Search blog to the DLF community. Excerpt:

"Today we informed our partners that we are ending the Live Search Books and Live Search Academic projects and that both sites will be taken down next week. Books and scholarly publications will continue to be integrated into our Search results, but not through separate indexes.

"This also means that we are winding down our digitization initiatives, including our library scanning and our in-copyright book programs. We recognize that this decision comes as disappointing news to our partners, the publishing and academic communities, and Live Search users.
...

"Based on our experience, we foresee that the best way for a search engine to make book content available will be by crawling content repositories created by book publishers and libraries."
I never really used Live Search Books, but I have colleagues who said very good things about it, some of whom thought it was better than Google Books in terms of search success and consistency and delivery UI. I hope that the output of the digitization by the partners can be re-purposed into other services. Their message encourages partners to continue working with the Internet Archive, so I feel hopeful that the equipment and processes put into place for this project will also continue to produce output even without Microsoft's involvement.

Thursday, May 22, 2008

more on oclc and google

Here's an article in Information Today on the OCLC/Google agreement.

This focuses more on the addition of Google Book links into existing WorldCat records, and the creation of new records for volumes not currently in WorldCat. I wonder if this means adding 856 field links to digital surrogates, or creating separate digital resource records? All OCLC member institutions will be able to add records into their catalogs.

This article doesn't mention one aspect of the agreement -- that the agreement now allows Google Book partners to share records from their catalogs that have an OCLC provenance with Google (or let OCLC do it for them). There has always been a lot of discussion about what rights member institutions had vis-a-vis sharing OCLC-sourced records that represent their holdings, and not everyone agrees with OCLC's assertions of its rights. Given Google's need to know something about the volumes that it's digitizing to provide access, it seems unavoidable that sharing of some OCLC-sourced metadata between Google Book participants and Google has already happened. Now it's a recognized need and activity.

Wednesday, May 21, 2008

oclc and google cooperation

On Monday a press release was issued about cooperation between OCLC and Google. Excerpted:

OCLC and Google Inc. have signed an agreement to exchange data that will facilitate the discovery of library collections through Google search services.

Under terms of the agreement, OCLC member libraries participating in the Google Book Search™ program, which makes the full text of more than one million books searchable, may share their WorldCat-derived MARC records with Google to better facilitate discovery of library collections through Google.

Google will link from Google Book Search to WorldCat.org, which will drive traffic to library OPACs and other library services. Google will share data and links to digitized books with OCLC, which will make it possible for OCLC to represent the digitized collections of OCLC member libraries in WorldCat.

...

WorldCat metadata will be made available to Google directly from OCLC or through member libraries participating in the Google Book Search program.

Google recently released an API that provides links to books in Google Book Search using ISBNs, LCCNs and OCLC numbers. This API allows WorldCat.org users to link to some books that Google has scanned through a “Get It” link. The link works both ways. If a user finds a book in Google Book Search, a link can often be tracked back to local libraries through WorldCat.org.

The new agreement enables OCLC to create MARC records describing the Google digitized books from OCLC member libraries and to link to them. These linking arrangements should help drive more traffic to libraries, both online and in person.

There are a couple of big wins here for different communities.

For WorldCat users, there is direct access to Google Book Search volumes. For Google Book users, there is improved access to physical volumes.

For Google Book participant libraries, there are better mechanisms for getting metadata about collections to Google.

Even more importantly -- and I am being hopeful here and reading something into this that may not be there -- this is a potential way to get representation of volumes digitized as part of the Google project not only into WordCat but into the OCLC/DLF Registry of Digital Masters. For folks unfamiliar with that project, it's a registry of digitized volumes -- which meet certain digitization standards and are publicly accessible -- that can be used as a tool by libraries and users to determine if volumes have already been digitized and are available. It's a slowly growing registry where a devoted group of participants have been working to develop the standards for describing digital masters and the work flows for adding records. This is a service that is poised to become essential.

Tuesday, April 01, 2008

convert paper to ipaper

When I first saw this on BoingBoing, I and others were certain it was an April Fool's joke: "Free bulk-scanning, OCR and web-publishing service launched by Scribd."

Then someone from Scribd posted in the comments that this was, indeed, a real offer with a very badly timed annoucement.

Part of me still suspects it's a prank, but this could be a real, limited-time prototype service offer.

My thoughts about iPaper are still as undecided as when I first wrote about it.