Tuesday, March 24, 2009

Ada Lovelace Day

I have had the pleasure in my life of working with a number of strong (and strong-willed) women who have seen me through various stages of my career. On the occasion of Ada Lovelace Day, I'd like to write about a colleague who I have known for many years, although we only had the opportunity to work in the same place for 4 weeks: Caroline Arms.

Caroline joined the Library of Congress in 1995 to work on the American Memory project. While the initial focus of the project was digitization and access, she saw the underlying issue that was created by such an effort: preservation. There was a profound lack of awareness in the library world about digital preservation at the time.

Caroline thought long and hard about the life cycle of digital objects, focusing in particular on one of the most vital areas that have consequences for all preservation efforts: standards for metadata and file formats. Preservation is always easier if good choices are made about digital formats. Curators should make collection decisions knowing which formats will and won’t be easily sustainable. For an object to be useful long into the future, its formats should be carefully selected and the specifications and characteristics of its formats must be documented.

Caroline and LC colleague Carl Fleischhauer’s exhaustive format research led to their creation of the Digital Formats web site, the first definitive inventory of information about current and emerging digital formats. The site is an essential resource for the international digital preservation community. Caroline also made a concerted effort to promote the use of formats with open standards, and to shepherd file formats through the standards review process.

She was also involved with the development of the Open Archives Initiative Protocol for Metadata Harvesting. I first met Caroline working on a collaborative OAI harvesting project, and I owe much of my expertise to her mentoring.

It was a great loss to LC that Caroline retired in June 2008. She did not retire from the community, however, and is participating in a LC group looking at metadata even now.

Thank you, Caroline.

Friday, March 20, 2009

ennui

I have been suffering through a state of ennui of late. Low energy, not feeling like cooking, a short attention span for reading, lack of interest in TV shows I usually enjoy, and a strong desire to work at home, curled up on the sofa with cats. Not even getting a great deal on some fabulous shoes to wear once sandal season returns, or going to a farmer's market we'd never visited before and finding -- wonder of wonders -- bacon sage ravioli, has cheered me up much. Tasks that usually give me a strange sense of accomplishment, like a successful presentation for a group at work or being caught up on the laundry or finally depositing a stack of checks for small denominations that I kept accumulating to take to the bank all at once, have done little to enhance my mood for long.

It's partly the cold, March, why-isn't-it-really-spring-yet doldrums. It's just past the one year anniversary of putting our house in Charlottesville on the market. I have also spent the lion's share of my time doing almost nothing but writing. In the past few weeks I have written a chapter for a book, revised a conference paper, written two conference proposals, and written 3 lengthy technical documents for one project alone, not to mention sending countless emails. All that writing and spreading myself across Twitter, Facebook, and this blog has had the effect of cutting down on posting overall.

Successfully (I hope) completing my first Japanese course next week will help. And two projects I've been working on have launches next month -- getting those out the door and having the chance to talk about them them will almost certainly improve my outlook.

Sunday, March 01, 2009

copyright registries

I attended a great presentation by Siva Vaidhyanathan and James Grimmelmann at Georgetown University last Friday on the Google Book Search settlement. The question that I most wanted to raise during the discussion period (why did the facilitator never call on me?) was about their opinions on the proposed registry. This seems to me to be one of the topics most in need of clarification in the settlement.

I chatted with both of them afterwards. I worry about a potential lack of transparency of the registry's contents and its mode of operation. I have heard Dan Clancy from Google say that it will not be made fully publicly available.

While there a student from the University of Michigan School of Information mentioned Michigan's IMLS grant supported effort to create a Copyright Review Management System to increase the reliability of copyright status determinations of books published in the United States from 1923 to 1963. Last week Lorcan Dempsey was blogging about the OCLC Copyright Registry Evidence Initiative. Stanford has a Copyright Renewal Database. John Mark Ockerbloom at the University of Pennsylvania researched periodicals renewals in addition to posting scans from many volunteer institutions (including Carnegie Mellon's and Project Gutenberg's extensive work) in his Catalog of Copyright Entries. The U.S. Copyright Office has records from 1978 onward online.

So, where does a Library (or anyone, for that matter?) go to research the copyright status of a published work? One of these places? All of these places? And where might the ownership status of orphan works someday be researched and recorded and made public? What will be the most authoritative source? Will there be open resources and less open resources? This looks like an area where there might be too much competition, almost a splintering of attention that calls out for a sense of coordination in the community.

recent reading

Some reports and posts that caught my attention recently:

The Andrew W. Mellon Foundation released a progress report from the DuraSpace project, a joint project of the DSpace Foundation and the Fedora Commons.

"MetaTools - Investigating Metadata Generation Tools" from JISC.

Merrilee Proffitt from RLG/OCLC posted on the "Legal and Ethical Implications of Large-Scale Digitization of Manuscript Collections" symposium at UNC-Chapel Hill. Posting Part 1 and Posting Part 2.

Andrew Richard Albanese published an article for Library Journal called "Institutional Repositories: Thinking Beyond the Box." It's a very balanced presentation of a number of points of view on the failure and success of IRs.

eReading

Via TeleRead, I found an essay about eReading devices by Jennifer Chapelle on treocentral. The piece, "Centro, iPhone, and that Other Reading Device (Kindle 2)," briefly describes her experiences with a Centro and an iPhone, focusing on the new Kindle 2.

Overall, she liked it. But she's not throwing away her other devices.

If you've ever been interested in getting an eReader type of device, I can definitely recommend the Kindle 2. It's not the cheapest gadget, but it does have a lot of features, and don't forget that 3G Sprint radio inside. If you want an eReader that is thin, lightweight, fast, looks great, has a built-in dictionary and a battery saving sleep-mode with some cool portraits, the Kindle 2 from Amazon is a great choice.

And if you don't care about those eReaders like the Kindle and the Sony device, just stick with your Treo or Centro. Those are great little eBook readers! And we know all the other great stuff you can do on them like talking on the phone, texting, writing documents, listening to music, taking photos, surfing the internet on decent looking web browsers, playing games, etc. My Centro and Treo Pro will be staying right by my side, Kindle or no Kindle.

I saw an interview with Jeff Bezos on Charlie Rose last week, which was primarily a discussion of the Kindle 2. My take-away is that the killer feature for the Kindle is the wireless purchasing of books that does not require a PC. Bezos is also a huge fan of the ability to bookmark your location in a text on your Kindle, and when you pick up another of your Kindles, the devices will sync up and you will find the same bookmark. Interesting, but I'm not sure I understand yet why you would have more than one. One at home and one at work? One downstairs and one upstairs? It's already portable. The functionality that they are working on where you can sync between your Kindle and a reader app on a cell phone and back interests me more. His example was reading on a cell phone while waiting in line at the grocery store, and having your Kindle aware of your new bookmark once you get home. That use case works better for me.

His statement that he wants to deliver "Every book ever in print in any language" gives me pause. That feels potentially monopolistic for the eBook distribution sector. Well, at least for their proprietery AZW ebooks. But if theirs becomes the most successful pipeline for eBooks, will other creators and distributors of other formats be able to compete? I can only assume the open access eBook realm will not fade away.

I found myself looking at the Sony eReader a week ago. The touchscreen and non-touchscreen versions boths have some different usability issues. The touchscreen is the better of the two, and supports annotation. It supports more files formats that the Kindle. It requires a PC has no wireless features. And it runs on MonteVista Linux, which a member of my family worked on a couple of years ago.

For now at least I plan to continue to read books on my Centro. I have about 3 dozen books, some recent, some classics. And I haven't divested myself of my nearly 3,000 dead tree books. Or my library cards.

Sunday, February 22, 2009

Caldwell collection

The Cooper-Hewitt Library is celebrating the release of Shedding Light on New York: Edward F. Caldwell Collection. The collection contains more than 50,000 images consisting of approximately 37,000 black & white photographs and 13,000 original design drawings of lighting fixtures and other fine metal objects that they produced from the late 19th to the mid-20th centuries.

Caldwell & Co. was America’s premier producer of lighting and other metal objects during the turn of the 20th century through the 1940s, and the archives are currently stored in the Cooper-Hewitt National Design Museum Library in New York City. Notable clients of Caldwell lighting fixtures included the Rockefellers, the Carnegies, and the Roosevelts, and the company was also commissioned for famous landmarks such as the Grand Central Terminal, Radio City Music Hall, and the Waldorf-Astoria in New York City. Caldwell & Co. manufactured unique and intricate lighting fixtures in their Manhattan factory, such as chandeliers, electrified lamps and wall scones, which were then shipped to prominent residences all over the United States.
New York Public Library also has Caldwell & Co records.

Saturday, February 21, 2009

Catalogue of Digitized Medieval Manuscripts

A team at UCLA has launched the Catalogue of Digitized Medieval Manuscripts, a centralized online archive of holdings worldwide.

The Catalogue first began to take form in Christopher Baswell's talk at the MLA conference in December, 2005. Generous support by the Center for Medieval and Renaissance Studies at the University of California, Los Angeles, has enabled Professors Matthew Fisher and Christopher Baswell to develop this site, and make it publicly available in its current form through the CMRS web site. An additional grant from the UCHRI (University of California Humanities Research Institute) made possible additional data entry, and substantive refinements to the back-end technologies in place.
...
Eventually, the site will have a collaborative layer of some sort, so that scholars can share their expertise with other researchers and with libraries, which do not always have the most accurate information for each manuscript, according to Mr. Fisher. He’d like the catalog to provide a general set of digital tools, too, so that similar databases can be built in other fields.
To date the project has located over 5,000 digitized manuscripts, and over 1,o00 have been cataloged for inclusion. An article in the Chronicle of Higher Ed provides background on the project.

Friday, February 20, 2009

FDsys federal content management system

Via Open Access News, the FDsys (Federal Digital System) of the US Government Printing Office (GPO) has entered its public beta. FDsys is an advanced digital system that will enable GPO to manage Government information in a digital form, and enable GPO to manage information from all three branches of the U.S. Government.

For more detail, see Joab Jackson's article about it in Government Computer News, February 5, 2009. There are five major releases planned over the next three years.

Duke Library's Trident Metadata Tool

The Duke University Library is blogging about its Trident Metadata Tool development. Their February 13 post is the first on their architecture.

good week for open source releases

The Indiana University Library has released open source software to create a digital music library system.

Indiana University today announces the release of open source software to create a digital music library system. The software, called Variations, provides online access to streaming audio and scanned score images in support of teaching, learning, and research.

Variations enables institutions such as college and university libraries and music schools to digitize audio and score materials from their own collections, provide those materials to their students and faculty in an interactive online environment, and respect intellectual property rights.

A key feature of the system for faculty and students is the ability to create bookmarks and playlists for use in studying or in preparing classroom presentations, allowing easy access later on to specific audio time points or segments. A key feature for libraries is a flexible access control and authentication system, which allows libraries to set up access rules based on their own local institutional policies.

This software is the culmination of nearly fifteen years of development and use of digital music library systems at Indiana University. Creation of the current Variations software platform was originally funded by the National Science Foundation. In 2005, the Institute of Museum and Library Services awarded Indiana University a National Leadership Grant to extend this highly successful system to the nationwide library community. Beyond IU, the software is currently being used at the Ohio State University, University of Maryland, New England Conservatory of Music, and the Philadelphia area Tri-College Consortium (Haverford, Swarthmore, and Bryn Mawr).

This open source release of Variations complements IU’s earlier release of the open source Variations Audio Timeliner, which lets users identify relationships in passages of music, annotate their findings, and play back the results with simple point-and-click navigation. This tool is also included as a feature of the complete Variations system.

Indiana University plans to offer a free one-hour Variations webinar at 4:00 PM EST on March 4, 2009 for institutions and individuals interested in learning more about the system. To register, e-mail mnotess@indiana.edu.

The Indiana University Digital Library Program created Variations in collaboration with faculty and students in IU’s Jacobs School of Music. The IU Digital Library Program is a collaborative effort of the Indiana University Libraries and the Indiana University Office of the Vice President for Information Technology.

For more information on the Variations open source release, see: http://variations.sourceforge.net/

The Washington Times released some Django open source tools (has a newspaper even released open source software before?):

The Washington Times has always focused on content. After careful review, we determined that the best way to have the top tools to produce and publish that content is to release the source code of our in-house tools and encourage collaboration.

The source code is released under the permissive Apache License, version 2.0. The initial tools released are:

  • django-projectmgr, a source code repository manager and issue tracking application. It allows threaded discussion of bugs and features, separation of bugs, features and tasks and easy creation of source code repositories for either public or private consumption.

  • django-supertagging, an interface to the Open Calais service for semantic markup.

  • django-massmedia, a multi-media management application. It can create galleries with multiple media types within, allows mass uploads with an archive file, and has a plugin for fckeditor for embedding the objects from a rich text editor.

  • django-clickpass, an interface to the clickpass.com OpenID service that allows users to create an account with a Google, Yahoo!, Facebook, Hotmail or AIM account.

The opensource.washingtontimes.com web site will be hosting the code and issue tracking software, using django-projectmgr.

Sunday, February 08, 2009

Yiddish books online

In October 2008 at an Open Content Alliance meeting, I saw a presentation about the National Yiddish Book Center. It has just been announced that over ten thousand Yiddish texts -- estimated as over half of all the published works in Yiddish currently in existence -- are now available online through a joint venture with the Internet Archive. From the press release:

The National Yiddish Book Center is proud to offer online access to the full texts of nearly 11,000 out-of-print Yiddish titles. You can browse, read, download or print any or all of these books, free of charge. These titles were scanned under the auspices of our Steven Spielberg Digital Yiddish Library, and have been made available online through the Internet Archive.


Original, used copies and new, print-on-demand hardcover reprints of most titles in our collection are available at nominal cost.

Some of rights issues are apparently unclear, but it seems so important to make this collection available -- works written in an at-risk language that were at one point systematically destroyed -- that any potential legal risk is worthwhile in my mind.

A brief announcement appeared in the New York Times.

Thursday, February 05, 2009

wikipedia loves art

The Smithsonian American Art Museum has announced its participation is a really interesting initiative -- help illustrate Wikipedia articles with your images of art from their collection. Photograph items from the collection, following some guidelines, upload your images to flickr, and your images will likely be used to illustrate a Wikipedia article.

Over the next month we are participating in Wikipedia Loves Art, a scavenger hunt and free content photography contest among 15 museums and cultural institutions worldwide. The project, in conjunction with Flickr, is aimed at illustrating Wikipedia articles. The event is planned to run for the whole month of February 2009.

We're inviting you to come into the museum and shoot photos of our artworks based on various themes. You can shoot on your own or form a small team (10 people, tops). The photogs or teams with the most points will win prizes.

The details about participation are available at Wikipedia.

This is part of the larger Wikipedia Loves Art project where a number of museums are participating in a scavenger hunt. The only thing that is not clear to me is what the Wikipedia articles that these images will illustrate are about. Scholarly topics? Topics related to art and art history? Articles about the museums or the specific works of art? I am curious.

DCC paper on interoperability

The Digital Curation Centre has released a short briefing paper on interoperability. Its a good, brief primer on the basic issues.

JHOVE2 requirements available

The latest version of the JHOVE2 Functional Requirements have been posted. I'm still interested in what isn't documented yet, e.g., the final list of formats that will be supported.

Sunday, January 25, 2009

National Film Board of Canada puts archives online

The National Film Board of Canada (NFB) has opened up its archives - more than 500 films, clips and trailers are now available on their new Screening Room web site. They're freely available for online viewing (there are costs for public broadcast and educational use), with more to be added regularly.

the burden of twitter

Steven Levy has written an essay for Wired about the guilt that one can feel for not participating enough in ones social network. Following tweets but not twittering, not blogging often enough, or not updating ones Facebook status. It's a brief but interesting read on privacy and a weird sense of duty to keep those public lines of communication open.

Nicholas Carr has posted a very interesting reaction to Levy's essay.

There's an arrogance to sharing the details of one's life in public with strangers - it's the arrogance of power, the assumption that such details somehow deserve to be broadly aired. And as for the people, those strangers, on the receiving end of the disclosures, they suffer, through their desire to hear the details, to hungrily listen in, a kind of debasement. At the risk of going too far, I'd argue that there's a certain sadomasochistic quality to the exchange (it's a variation on the exchange that takes place between celebrity and fan). And I'm pretty sure that Levy's remorse comes from his realization, conscious or not, that he is, in a very subtle but nonetheless real way, displaying an undeserved and unappetizing arrogance while also contributing to the debasement of others.
This seems a bit strong to me, but not entirely off base. Arrogance of power? Debasement? Sadomasochistic? OK, that may be true for some who participate in social networking, just the same as for some participants in a real life communities. There is something a bit egotistical in assuming that others will follow your tweets/blog/delicious tags/flickr set/facebook. There is something a bit creepy that, if you don't require approval, complete strangers read your tweets where you might be discussing where you are at any given time. I like to think that most use social networking to actually keep in touch, not to obsessively stalk one another.

There's that public sharing expectations thing again. I know, I think about this a lot. People I do not know read my blog, see many (but not all) of my flickr images, and join my delicious network to see most (but again, not all) of my bookmarks. I have made a conscious decision to share these things. I had to struggle with getting over the creepiness factor. It was well over a decade ago that a woman from China, upon being introduced to me at a conference reception, exclaimed "Oh! I know who you are -- you have an interest in folk art and you like armadillos!" She had come across my personal web page (remember those?) while researching the conference speakers.

There's no turning back. There's only self-selecting your level of exposure.

Folger Library launches online image collections

The Folger Shakespeare Library just expanded access to its Digital Image Collection by offering over 20,000 images online. The collection includes books, theater memorabilia, manuscripts, art, and 218 of the Folger’s pre-1640 quarto editions of the works of William Shakespeare.

Online use is through the Luna Insight Browser -- you have to add an exception to your popup blocker or the software will not function properly. To access their Shakespeare Quartos collection and to get full functionality (saving searches, exporting html pages) you have to install the free Insight Java client.

They have a "how-to" page and search tips available.

Library of Congress SourceForge release

Last month the Library of Congress had a soft launch of an open source software release. We officially announced the release in the January 2009 issue of the Library of Congress Digital Preservation
Newsletter
. This is the first software that the Library has formally released as open source.

The tools are available through SourceForge under the “Library of Congress Transfer Tools” project. The project includes tools for use with BagIt specification, a hierarchical file packaging format for the exchange of digital content jointly developed by the Library of Congress and the California Digital Library.

Three tools developed by the Library's Repository Development Group are available now. Parallel Retriever implements a simple Python-based wrapper around wget and rsync to optimize the transfer of content between locations through parallelization. It supports rsync, HTTP, and FTP transfers. Bag Validator is a Python script that validates a Bag, checking for missing files, extra files, and duplicate files. VerifyIt is a shell script that verifies file checksums within a Bag manifest using parallel processes.

The Library plans to release additional tools as part of a suite of solutions and software development resources as they are completed over time. There are already more tools in the pipeline.

Friday, January 23, 2009

mobile is the new black

There's a new WorldCat Mobile pilot service.

NYPL has announced its NYPL Mobile beta.

The DC Public Library launched an iPhone app.

Stanford has a new version of an iStanford iPhone app that ties into its student services system.

The International Children's Digital Library launched an iPhone app last November.

technology transition at the white house

There was a great Washington Post article yesterday about how White House technology is "in the Dark Ages." I laughed bemusedly over my toast and read the article aloud at the breakfast table.

The White House is not being singled out. I work for a Federal Agency. I know folks who work at numerous other Federal Agencies, some of whom have worked at said agencies for decades. Federal agencies have many, many rules about hardware and software security, and every agency has to interpret and enforce those rules themselves. Security levels of content muddy the waters. This can cause a certain amount of confusion as to what is and isn't allowed. Someone told me that their agency (not the White House) hasn't yet approved Firefox. News that the White House counsel's office approved use of Gmail accounts for some press office activities has been forwarded to many Federal IT units, I'm sure.

Edit, 25 January: Wired has posted a Wired/Tired overview of White House tech, and a list of recent technology projects from various agencies. Nice to see the shout out for the LoC Flickr project.