Showing posts with label texts. Show all posts
Showing posts with label texts. Show all posts

Tuesday, June 16, 2009

milestones for the National Digital Newspaper Program

Today there was an exciting press event at the Newseum for the National Digital Newspaper Program, sponsored by the Library of Congress and the National Endowment for the Humanities. There was a great live demo, a video on digital production for the project from the University of Kentucky, and some nice speechmaking. The event promoted the milestone where the project surpassed 1,000,000 pages available at the Chronicling America site, the addition of seven new state partners, and the addition of images of illustrated newspaper supplements to the LoC Flickr Commons set (with more to come every month).

So far the AP has an article available, and there were representatives of other news outlets at the event. Check out the press release. Roy Tennant has a post that includes some of the technical specs supplied by my colleague Ed Summers. Ed and Dan Krech have done some great work to update the underlying application, improving the ingest and search functionality, adding the functionality that allows the site to be crawled, and exposing the data as RDF for a multitude of possibilities.

Edit: Here's the Washington Post article, and the official LoC blog posting.

Thursday, April 02, 2009

new flip book beta

From Peter Brantley on the OCA blog -- A new beta version of the Flipbook bookreader has been released open source under GNU license. The source code is available from the Open Library site.

Saturday, February 21, 2009

Catalogue of Digitized Medieval Manuscripts

A team at UCLA has launched the Catalogue of Digitized Medieval Manuscripts, a centralized online archive of holdings worldwide.

The Catalogue first began to take form in Christopher Baswell's talk at the MLA conference in December, 2005. Generous support by the Center for Medieval and Renaissance Studies at the University of California, Los Angeles, has enabled Professors Matthew Fisher and Christopher Baswell to develop this site, and make it publicly available in its current form through the CMRS web site. An additional grant from the UCHRI (University of California Humanities Research Institute) made possible additional data entry, and substantive refinements to the back-end technologies in place.
...
Eventually, the site will have a collaborative layer of some sort, so that scholars can share their expertise with other researchers and with libraries, which do not always have the most accurate information for each manuscript, according to Mr. Fisher. He’d like the catalog to provide a general set of digital tools, too, so that similar databases can be built in other fields.
To date the project has located over 5,000 digitized manuscripts, and over 1,o00 have been cataloged for inclusion. An article in the Chronicle of Higher Ed provides background on the project.

Sunday, February 08, 2009

Yiddish books online

In October 2008 at an Open Content Alliance meeting, I saw a presentation about the National Yiddish Book Center. It has just been announced that over ten thousand Yiddish texts -- estimated as over half of all the published works in Yiddish currently in existence -- are now available online through a joint venture with the Internet Archive. From the press release:

The National Yiddish Book Center is proud to offer online access to the full texts of nearly 11,000 out-of-print Yiddish titles. You can browse, read, download or print any or all of these books, free of charge. These titles were scanned under the auspices of our Steven Spielberg Digital Yiddish Library, and have been made available online through the Internet Archive.


Original, used copies and new, print-on-demand hardcover reprints of most titles in our collection are available at nominal cost.

Some of rights issues are apparently unclear, but it seems so important to make this collection available -- works written in an at-risk language that were at one point systematically destroyed -- that any potential legal risk is worthwhile in my mind.

A brief announcement appeared in the New York Times.

Thursday, January 15, 2009

oclc summary of proposed google book settlement

Ricky Erway from OCLC has distilled the proposed Google Book Settlement, its appendices, and the three library registry agreements from 320 pages to a 4 1/2 page summary. It's an excellent overview of the proposal.

Saturday, January 03, 2009

on electronic texts

I just read an article at Information Today by Nicholas Tomaiuolo, an instruction librarian at Central Connecticut State University, entitled "U-Content: Project Gutenberg, Me, and You." He outlines the requirements and steps for preparing an etext for Project Gutenberg.

At one point in the article, there is a discussion about the requirements for full text, not just a PDF created from page images. The author wrote this from the point of one unfamiliar with PG's requirements, illustrating the process one might follow to create an acceptable PG submission -- images to PDF, and images to OCR to corrected plain text -- I found myself thinking quite a bit about the often heard statement (not in this article, mind you) that PDF is the ultimate format for texts.

I'm in no way denigrating PDF. PDFs is an absolutely required format for texts. PDF is highly portable and shareable and readable, and, if the source files are good enough, clearly printable. But it's not innately analyzable or easily repurposed. That requires full text.

I am not unfamiliar with what it takes to create an accurate plain text transcription of a text. When Gutenberg was in its early days, we were really talking about transcriptions, as in people typing in text. OCR has greatly streamlined that process, but the proofreading required is non-trivial. Want to work with a highly formatted text, or one with tables or formulae or figures? Challenging. Adding layers of structural and semantic markup to plain text, as with TEI, is time consuming. Rich markup, including identifying dates or names or geographical places, or providing normalized versions of said dates and names is a large undertaking. A full text with structural and sematic markup can be repurposed into many formats, including ebooks and PDF.

And you do want ebooks. Some months ago I had the great opportunity to demonstrate the prototype World Digital Library site at the National Book Festival. There is no greater focus group than thousands of people who love to read! The two top requests were that the books should be downloadable as ebooks and that all the text content be available as full text in all seven project languages. These were not academics or librarians (although there were some of the former and many of the latter who stopped by), but parents and commuters and researchers and genealogists.

Both are daunting requests when you do not have full text available to work from. There will be PDFs. The others are goals to strive for.

Friday, November 14, 2008

roman de la rose digital library

Johns Hopkins University and the Bibliothèque nationale de France have announced that the Roman de la Rose Digital Library available at http://romandelarose.org/. The goal is to bring together digital surrogates of all the approximately 270 extant manuscript copies of the Roman de la Rose. By the end of 2009 they expect to have 150 versions included in the resource. There is an associated blog available at http://romandelarose.blogspot.com/.

I am particularly interested in the pageturner and image browser that they used -- the FSI Viewer, a Flash-based tool. It seems to work with TIF, JPG, FPX, and PDF (but not JPEG2000?), and converts files to multi-resolution TIFs. It's a very intuitive interface.

Friday, October 17, 2008

digital book access at John Hopkins

Jonathan Rochkind has posted a great description of digital book access features that he's put into production in the link resolver and OPAC at Johns Hopkins. They're remarkable in the sense that he's taken advantage of so many different service APIs (Google Books, IA, OCLC, Amazon, HathiTrust) to provide functionality with conditional options to provide as much collection coverage as possible.