Thursday, January 15, 2009

oclc summary of proposed google book settlement

Ricky Erway from OCLC has distilled the proposed Google Book Settlement, its appendices, and the three library registry agreements from 320 pages to a 4 1/2 page summary. It's an excellent overview of the proposal.

d-lib article on some LC tool development

My colleague Justin Littman has just published an excellent article in the January/February 2009 issue of D-Lib Magazine: "A Set of Transfer-Related Services."

"The Office of Strategic Initiative's (OSI) Repository Development Team (RDT) is developing a portfolio of services and components to address the challenges posed by scaling transfer processes. While the portfolio is expanding, the focus of this article will be on two core services, the Inventory Service and the Workflow Service. Before proceeding to examine these services, it will be useful to further delineate the transfer problem space. After examining these services, their role in mitigating preservation risks will be considered."

Monday, January 12, 2009

presidential records and donation reform

On January 7, 2009, the U.S. House of Representatives approved H.R. 35, the "Presidential Records Act Amendments of 2009," and H.R. 36, the "Presidential Library Donation Reform Act of 2009." These were chosen by the House leadership as the first pieces of substantive legislation passed in 2009 as a symbol of government transparency.

The Presidential Records Act Amendments restores meaningful public access to presidential records by nullifying a 2001 Bush executive order, and the Presidential Library Donation Reform Act requires the disclosure of big donors to presidential libraries. The Senate still has to pass its versions of the bills before they can go to soon-to-be-President Obama to be signed, which he has apparently indicated that he would.

The National Coalition for History provides a good overview of the Records Reform Act. The House Speaker's site provides an overview of both.

Sunday, January 11, 2009

I want a Palm Pre

Pretty much everyone who knows me knows my loyalty to Palm. I've had one since 1997. I've been syncing it with enterprise calendar systems so I have my personal and work calendars going back to February 1996. I currently have a Centro, and even though I have to manually key in all my work events because we use such an old version of Groupwise that I can't seem to find a sync that works, I am still devoted to my Palm.

I cannot count the number of friends and colleagues who have iPhones and have done their best to convert me. The Urbanspoon app and its clever use of the accelerometer to randomize recommendations by shaking the phone almost had me. My answer is always that if they can promise me that I can port over everything I have on my Centro -- 13 years of calendar, hundreds of contacts, notes, to-do lists, and ebooks -- then I'll consider it.

I am now waiting with baited breath for the Palm Pre. It's such a step forward in the interface (a card stack metaphor) and operating system (its WebOS is Linux) and browser (based on WebKit). It doesn't use the Palm desktop anymore, which while be a paradigm switch for me, but they have promised data migration tools for Centro users.

I know that some apps I use won't work anymore, and Raymond's concern about no backwards compatibility of the OS is a real one. But they need to move on and I need to move on, and I'm glad it looks like it will be to another Palm. You can't always be fully backwards compatible. Hey, I only complained a little when I discovered I couldn't run FileMake Mobile on my Centro, didn't I?

Reviews at Gizmodo, ars technica, PC World, and a PC World FAQ.

Saturday, January 03, 2009

on electronic texts

I just read an article at Information Today by Nicholas Tomaiuolo, an instruction librarian at Central Connecticut State University, entitled "U-Content: Project Gutenberg, Me, and You." He outlines the requirements and steps for preparing an etext for Project Gutenberg.

At one point in the article, there is a discussion about the requirements for full text, not just a PDF created from page images. The author wrote this from the point of one unfamiliar with PG's requirements, illustrating the process one might follow to create an acceptable PG submission -- images to PDF, and images to OCR to corrected plain text -- I found myself thinking quite a bit about the often heard statement (not in this article, mind you) that PDF is the ultimate format for texts.

I'm in no way denigrating PDF. PDFs is an absolutely required format for texts. PDF is highly portable and shareable and readable, and, if the source files are good enough, clearly printable. But it's not innately analyzable or easily repurposed. That requires full text.

I am not unfamiliar with what it takes to create an accurate plain text transcription of a text. When Gutenberg was in its early days, we were really talking about transcriptions, as in people typing in text. OCR has greatly streamlined that process, but the proofreading required is non-trivial. Want to work with a highly formatted text, or one with tables or formulae or figures? Challenging. Adding layers of structural and semantic markup to plain text, as with TEI, is time consuming. Rich markup, including identifying dates or names or geographical places, or providing normalized versions of said dates and names is a large undertaking. A full text with structural and sematic markup can be repurposed into many formats, including ebooks and PDF.

And you do want ebooks. Some months ago I had the great opportunity to demonstrate the prototype World Digital Library site at the National Book Festival. There is no greater focus group than thousands of people who love to read! The two top requests were that the books should be downloadable as ebooks and that all the text content be available as full text in all seven project languages. These were not academics or librarians (although there were some of the former and many of the latter who stopped by), but parents and commuters and researchers and genealogists.

Both are daunting requests when you do not have full text available to work from. There will be PDFs. The others are goals to strive for.

Friday, January 02, 2009

flickr commons developments

I am not part of the Library of Congress Flickr project team, and I in no way speak for them.

There is a lengthy discussion on Wired and at Found History about The Commons on Flickr, given that Yahoo laid off key staff member George Oates who shepherded the project. It doesn't seem that the project is in any imminent danger, but a dedicated group has stepped forward to evangelize, innovate, and curate thematic sets from the across the collection. I'm thrilled to see a community of use developing around The Commons, but it's sad that this is what precipitated its full coalescence.

DCC obsolescent data and files challenge

There's still time to send a message to Chris Rusbridge at the Digital Curation Center to enter his personal data recovery challenge:

I will do my best to recover the first half dozen interesting files that I’m told about… of course, what I really mean is that I’ll try and get the community to help recover the data. That’s you!

OK, I define interesting, and it won’t necessarily be clear in advance. The first one of a kind might be interesting, the second one would not. Data from some application common of its time may be more interesting than something hand-coded by you, for you. Data might be more interesting (to me) than text. Something quite simple locked onto strange obsolete media might be interesting, but then again it might be so intractable it stops being interesting. We may even pay someone to extract files from your media, if it’s sufficiently interesting (and if we can find someone equipped to do it).

The only reference for this sort of activity that I know of is (Ross & Gow, 1999, see below), commissioned by the Digital Archiving Working Group.

What about the small print? Well, this is a bit of fun with a learning outcome, but I can’t accept liability for what happens. You have to send me your data, of course, and you are going to have to accept the risk that it might all go wrong. If it’s your only copy, and you don’t (or can’t) take a copy, it might get lost or destroyed in the process. You’ll need to accept that risk; if you don't like it, don't send it. I might not be able to recover anything at all, for many reasons. I’ll send you back any data I can recover, but can’t guarantee to send back any media.

The point of this is to tell the stories of recovering data, so don’t send me anything if you don’t want the story told. I don’t mind keeping your identity private (in fact good practice says that’s the default, although I will ask you if you mind being identified). You can ask for your data to be kept private, but if possible I’d like the right to publish extracts of the data, to illustrate the story.
His deadline is Twelfth Night, January 5, 2009. Read his post for the rest of the details.

I've been thinking a lot lately about data migration and recovery, and I think this is a great way to illustrate the challenges and solutions -- with real data and the stories behind its creation, loss, and potential recovery.

make

I am definitely a proponent of the self-made. Cooking, clothing, jewelry, assisted forays into small electronic projects, etc. At various times I have designed and made costumes, painted, made paper, etched prints, made lampwork glass beads, and learned the basics of Chinese brush painting. I haven't had as much time for that in the last couple of years, but that doesn't mean that I don't still strive to make.

I'm a big fan of Make, and so are a lot of others. In that vein, William Turkel has posted his list of books for humanist makers.

interview with Vint Cerf

Siva Vaidhyanathan has posted his interview with Vint Cerf about developments in search technology and Google.

Thursday, January 01, 2009

resolutions

I don't know if these are resolutions or goals ...

I have been getting back to writing in the past few weeks, but I will write more about what we're working on and getting the word out.

I will be a more hard-nosed project manager vis-a-vis deadlines. I am too often too understanding of delays.

I will explore Washington D.C. more. I lived in metro D.C. for part of my childhood and have been in Virginia for 7 years, but still have a disconnected sense of the city. Perhaps that's in part an artifact of traveling by Metro and not having a real sense of the city's topography.

I will meet more people in other divisions at LoC. This is harder than it sounds.

I will eat lunch at my desk less often. I will make more of an active effort to organize social opportunities. (Why don't I remember this every time I think about switching jobs -- it gets harder and harder to build a new social network every time.)

I have faith that our house in Charlottesville will finally sell.

EDIT: There is another long term goal that I'm already working on: I've signed up for a Japanese class through the Federal employee graduate school. This is the year to begin transforming my ragtag knowledge of Japanese into usable language skills.

Tuesday, December 30, 2008

archiving the bush administration

There is a great article on ars technica today about the major processing effort that will be required at the National Archives when the Bush administration leaves office. The ars technica piece references a New York Times article on the topic from this past weekend.

This section really strikes home:

The contingency plan will entail "ingesting" the Bush White House's data into a separate system before integrating it with the ordinary archive. As the plan explains, "the current PERL [Presidential Electronic Records Library] system architecture was not scalable to actually support the volume of records that are expected from the current Presidential administration."

It's not just size that matters, though: the Archives will also need to process reams of information locked in some quaint proprietary formats. The RMS index, for example, "consists of an implementation of a customized older version of Documentum running on Oracle, with image files (including copies of scanned records) incorporated as objects in the database." The photos are stored in a "proprietary photo management software called MerlinOne, running on Microsoft SQL as the database engine," and it has apparently taken several months to extract the images and metadata for relinkage outside the Merlin format.

First, the use of quotation marks should remind us all that "ingest" means absolutely nothing to someone who is not a repository manager.

I have participated in some discussions about a potential data migration project at work. I recently saw an inventory of media formats -- not file formats, but media formats -- that the project would need to encompass, and it is lengthy. The only source I can think of for hardware to read some of the formats is EBay. That doesn't even take into account the files themselves. It's interesting how quickly a format becomes obsolete, and how many customized systems federal agencies use.

Monday, December 29, 2008

blogging has fallen by the wayside

Between a month almost solely dedicated to a single high-stress project and a lot of other writing commitments -- revising a paper for a conference, drafting a conference proposal and a co-authored conference proposal, and a writing chapter for a book -- I find that I haven't made time to blog. I promise to make time soon.

best metrics for comparing hardware?

Recently I spent 4 weeks on a project where we were considering hardware options for a large amount of storage for a data migration project. We ended up with 4 different proposals -- three from vendors and one to be built in-house. One of the tasks that I worked on was a matrix to compare the 4 potential solutions.

There were the easy metrics -- the amount of raw and usable storage, number of racks/tiles required, electrical and cooling requirements, cost, etc. Comparing supportability was trickier but doable, with 24*7 versus 12*5 phone support, availability of on-site technicians, warranty terms, support contract costs, etc. Where it became more difficult was identifying metrics to compare performance. Ratio of processors to storage? Location of processing nodes in the architecture? I/O rates? Time to read all data? And how do you best calculate those last two with four quite architecturally different proposals? We ended up with metrics that not everyone agreed upon, in part because there was a requirement that not everyone agreed upon.

I'm curious how other folks have gone about doing this. I'd be interested in hearing from anyone who is willing to share their strategies.

Wednesday, December 17, 2008

ICDL adding European collections

From an article on Forbes.com, the International Children's Digital Library (ICDL) announced a partnership with the Taliaferro Family Fund to increase the number of European children's titles in the collection. The Elias Project will target three collections in Europe: the Norwegian Children's Book Institute in Oslo, Norway, the International Youth Library in Munich, Germany , and the National Center for Children's Books in Paris, France.

After reading the article, I check in at the ICDL site, which I hadn't visited in a few months, and noticed two other news announcements: ICDL and the Google Book project will be sharing public domain children's book titles; and ICDL has launched an iPhone app with full access to the collection, a new titles features, and an offline mode and an airplane mode. It's great to see such a worthwhile project making such advances in collection building and in adding new services.

(I didn't see a press release about the European project on the ICDL site. I saw the press release on some other sites, so I assume it's meant to be out there.)

Tuesday, December 16, 2008

letter to santa

Nik Honeysett has posted a great letter to Santa on the Musematic blog.

If enough of us ask for that image format, will Santa grant our wish?

Friday, December 12, 2008

interview with Paul LeClerc in New York Times

Paul LeClerc, director of the New York Public Library, answered questions online at the New York Times that have been made available in three parts: part one, part two, part three. Topics include budgets, branch closures and renovations, ebooks, and preservation efforts.

In part two, he briefly mentioned their participation in the Library of Congress National Digital Newspaper Project and its public access content web site Chronicling America. It was nice to see this project mentioned get a media mention in the context of preserving and providing access to often ephemeral newspapers.

Library of Congress releases report on flickr pilot

The Library of Congress has released its report on its Flickr Commons pilot, where approximately 5,000 images were uploaded for a crowdsourcing metadata experiment. A full report and a summary report are available, both PDFs.

The photos have drawn more than 10 million views, 7,166 comments and more than 67,000 tags. When Flickr commenters provide updated place and personal names, dates, and event identification, staff from the Library's Prints and Photographs Division verify the information and have so far updated more than 500 records in their catalog -- with many more in the queue -- citing the Flickr Commons Project as the source of the new information.

Thursday, December 11, 2008

creative commons wants feedback on licenses

Creative Commons is conducting a study to collect feedback on the term “noncommercial” and how it should be covered in its licenses. The hope is that what’s learned from the survey can improve the licenses that allow or restrict noncommercial uses. The questionnaire has to be completed by this Sunday, December 14, 2008. Everyone who has taken advantage of CC licenses as a creator or a user should take some time to answer the questions.

world war II collection at the national archive and footnote

The US National Archives and the historical document website Footnote.com have collaborated on the digitization of a large collection of documents from the US involvement in World War II, which are now available on the footnote.com web site. There is an ars technica article on the collection and interface.

Like the ars technica writer, I had a lot of difficulty finding anything that I hoped to find. My grandfather, father, and uncle all served in WWII. My grandfather died in a friendly fire incident where allied planes accidentally sunk a ship carrying prisoners of war to be returned. I found nothing. There was nothing in the documents nor in the photos. Although I did find out that a man with almost the same name as my uncle (same middle initial but different middle name) was listed as missing when his plane was shot down in 1943. Still, it's a lot of useful content that I'm glad to see digitized and OCR'ed.

I was disappointed I wasn't surprised. I found the navigation to be a bit puzzling. I found I had to have multiple tabs open to easily go back to search. Not just the image but the entire image viewer screen had to come into focus when I selected something to view.

The ars technica writer said that his view of the site included the disclaimer "All Free (for a limited time)," and commented that "... it would be nice to think that a service based on government records of a significant American experience would be free indefinitely." The original press release describing the collaboration is worth reviewing, because it addresses that point in the ars technica article. The agreement allows Footnote.com non-exclusive access, and "After an interval of five years, all images digitized through this agreement will be available at no charge through the National Archives web site." So, Footnote can charge for it for now, but it will all revert to the National Archives for free and open access.

I don't see that disclaimer when using my Library of Congress computer because we have full access -- I wonder how long it will be fully accessible for those without subscriptions?

Sunday, December 07, 2008

laine farley named cdl director

I just saw the press release naming Laine Farley as the new Director of the California Digital Library. I am thrilled for Laine, who's been serving as the interim Director for over 2 years. I have worked with her on Aquifer and at least one other collaborative community project, and I know what an experienced and capable person she is.