Showing posts with label LoC. Show all posts
Showing posts with label LoC. Show all posts

Tuesday, June 30, 2009

LoC on iTunes

The Library of Congress now has content on iTunes U. iTunes U is the area of the iTunes Store which offers open educational audio and video content from universities and other educational institutions. The Library’s initial iTunes U content includes historical videos such as original Edison films and a series of 1904 films from the Westinghouse Works, as well as event videos such as author talks from the National Book Festival, the "Books and Beyond" series, discussions with curators, and lectures from the Kluge Center. The audio content includes Library podcast series such as "Music and the Brain," slave narratives from the American Folklife Center, and interviews with authors from the National Book Festival. The collection also includes Library-produced classroom and educational materials, such as courses from the Catalogers’ Learning Workshop.

You must be running iTunes to be able to view the LoC content.

Saturday, June 27, 2009

new BIL on SourceForge and update to BagIt spec

This week saw a couple of events around the BagIt specification and tools.

A revision of the BagIt specification went out this week. You will note that it is still 0.96 -- the revisions were only in language to clarify some questions that had been received. There are some discussions going on about 0.97 - join the Digital Curation Google group. I'd like to see some more activity there!

Version 3.0 of BIL, the BagIt Library for Java, was released on SourceForge this week. It's available as binary and source code.

Plus, there was the BagIt video ...

BagIt video

The first in a planned series of digital preservation videos is available on the digitalpreservation.gov site -- an introduction to BagIt! Brian Vargas did a great job as "the talent" -- e.g., the narrator -- but folks should know that Brian was not selected just for his acting experience: he wrote many of our transfer tools (like the transfer scripts on SourceForge) and is a co-author of the BagIt specification.

The video premiered this week at the annual NDIIPP Partner's Meeting to great acclaim. It's aimed at a general audience.

EDIT: The NDIIPP site has added a great new page on the Transfer Tools with a link to the video.

Tuesday, June 16, 2009

milestones for the National Digital Newspaper Program

Today there was an exciting press event at the Newseum for the National Digital Newspaper Program, sponsored by the Library of Congress and the National Endowment for the Humanities. There was a great live demo, a video on digital production for the project from the University of Kentucky, and some nice speechmaking. The event promoted the milestone where the project surpassed 1,000,000 pages available at the Chronicling America site, the addition of seven new state partners, and the addition of images of illustrated newspaper supplements to the LoC Flickr Commons set (with more to come every month).

So far the AP has an article available, and there were representatives of other news outlets at the event. Check out the press release. Roy Tennant has a post that includes some of the technical specs supplied by my colleague Ed Summers. Ed and Dan Krech have done some great work to update the underlying application, improving the ingest and search functionality, adding the functionality that allows the site to be crawled, and exposing the data as RDF for a multitude of possibilities.

Edit: Here's the Washington Post article, and the official LoC blog posting.

Tuesday, April 21, 2009

World Digital Library Launch

The World Digital Library is now available.

The site is launching with 1,170 objects from 26 partner institutions. WDL focuses on significant primary materials reflecting the cultural heritage of all UNESCO member countries, including manuscripts, maps, rare books, recordings, films, prints, photographs, architectural drawings, and other types of primary sources from varying time periods. The project will continue to add content to the site, and will enlist new partners from the widest possible range of institutions and countries.

The site is available in seven different languages: Arabic, Chinese, English, French, Russian, Spanish, and Portuguese. The content is not translated -- the items appear in their original language. The metadata and all the site navigation is translated to make it possible to search and browse the site in any of the languages. The metadata came from partner institutions or was created by catalogers at the Library of Congress, and much of the translation was provided by Lingotek.

The site was built using the Django Python framework, nginx, Lucene/Solr, and a mySQL database. The zooming in the imageviewer and pageturner is Seadragon Ajax. There is heavy use of Javascript, jquery, JSON and underlying XML. Check out the image carousels and timeline tool! The project also developed a cataloging tool to manage the metadata and cataloging process and interact with the Lingotek translation system via their API.

Friday, March 27, 2009

New LC multimedia collection sharing initiatives

This is news ... The Library of Congress will begin sharing content from its vast video and audio collections on the YouTube and Apple iTunes web services as part of a continuing initiative to make its incomparable treasures more widely accessible to a broad audience. The new Library of Congress channels on each of the popular services will launch within the next few weeks.

...

The General Services Administration today also announced agreements with Flickr, YouTube, Vimeo and blip.tv that will allow other federal agencies to participate in new media while meeting legal requirements and the unique needs of government. GSA plans to negotiate agreements with other providers, and the Library will explore these new media services when they are appropriate to its mission and as resources permit.

Read the Press Release.

Sunday, January 25, 2009

Library of Congress SourceForge release

Last month the Library of Congress had a soft launch of an open source software release. We officially announced the release in the January 2009 issue of the Library of Congress Digital Preservation
Newsletter
. This is the first software that the Library has formally released as open source.

The tools are available through SourceForge under the “Library of Congress Transfer Tools” project. The project includes tools for use with BagIt specification, a hierarchical file packaging format for the exchange of digital content jointly developed by the Library of Congress and the California Digital Library.

Three tools developed by the Library's Repository Development Group are available now. Parallel Retriever implements a simple Python-based wrapper around wget and rsync to optimize the transfer of content between locations through parallelization. It supports rsync, HTTP, and FTP transfers. Bag Validator is a Python script that validates a Bag, checking for missing files, extra files, and duplicate files. VerifyIt is a shell script that verifies file checksums within a Bag manifest using parallel processes.

The Library plans to release additional tools as part of a suite of solutions and software development resources as they are completed over time. There are already more tools in the pipeline.

Thursday, January 15, 2009

d-lib article on some LC tool development

My colleague Justin Littman has just published an excellent article in the January/February 2009 issue of D-Lib Magazine: "A Set of Transfer-Related Services."

"The Office of Strategic Initiative's (OSI) Repository Development Team (RDT) is developing a portfolio of services and components to address the challenges posed by scaling transfer processes. While the portfolio is expanding, the focus of this article will be on two core services, the Inventory Service and the Workflow Service. Before proceeding to examine these services, it will be useful to further delineate the transfer problem space. After examining these services, their role in mitigating preservation risks will be considered."

Friday, December 12, 2008

Library of Congress releases report on flickr pilot

The Library of Congress has released its report on its Flickr Commons pilot, where approximately 5,000 images were uploaded for a crowdsourcing metadata experiment. A full report and a summary report are available, both PDFs.

The photos have drawn more than 10 million views, 7,166 comments and more than 67,000 tags. When Flickr commenters provide updated place and personal names, dates, and event identification, staff from the Library's Prints and Photographs Division verify the information and have so far updated more than 500 records in their catalog -- with many more in the queue -- citing the Flickr Commons Project as the source of the new information.

Saturday, August 30, 2008

web archiving

The Library of Congress has a phenomenal Web Capture team, staffed with very dedicated people who take a lot of effort to identify web sites that best document an event, crawl and capture sites through partner Internet Archive, work with cataloging to get the sites described to enhance discoverability, do quality control to make sure the archived sites will run correctly, and then make the archived sites live for public access. This process can take a very long time to ensure that a site is fully captured, preserved, and accessible.

The Web Capture team is, as they have with previous years, documenting the 2008 elections. They don't crawl sites without permission, and they always send requests. A colleague at another library sent me a link to a post and series of comments on Wonkette that were a reaction to a LoC request to capture the site. The post itself is fine. It is more than a bit surreal to get such a request from LoC -- they're going to collect what I write? -- and making fun of it is OK.

Some of the comments, however, are another story.

The reaction to the notice of the request includes strings of profanity, vulgarity, and various exhortations to "archive this, LoC!" Some comment that it's possibly a fake request similar to a Nigerian scam, some liken it to FBI wiretapping, and one comment says that it's a waste of taxpayer dollars to have federal employees reading websites in order to identify what should be archived. One comment conjectures that by "capture," we mean print out the site and store it in a box next to the Ark of the Covenant. Some of the comments are obviously humorous and some are serious, and it's hard to tell with others.

I have a sense of humor, especially about political topics. Of course it's funny to the Wonkette participants that whatever is said, whether profound or mundane or profane, the Library of Congress will crawl it. I remember my own reaction when I was approached about submitting my email to the MCN archives covering the period when I was on its board, which contained such highlights as "The membership brochure is at the printer" and "Don't faint when you see how much the conference hotel wants to charge us for internet access." But for some reason this really struck a nerve because some of the commenters were so "f--- you, Library of Congress." That saddened and angered me.

It's a huge effort to collect ever-changing interactive born-digital resources compared to print materials, but we and many others libraries do it because it's an equally important form of publishing. Libraries collect whatever is relevant regardless of their form of publication. Sites like these are important because they reflect what's really being said and what people really think about the political process. What about that isn't worth collecting?

I'll cop to being a bit overly sensitive on this, but only because I place very high value on such collecting activities.

Tuesday, July 15, 2008

ndiipp partners meeting

Last week I attended the three-day 2008 meeting for the partners in the Library of Congress National Digital Information Infrastructure and Preservation Program (NDIIPP). Yesterday a colleague who couldn't attend asked me what stood out for me in the program. I didn't take a lot of notes -- I kept forgetting to because I just wanted to listen -- but I see some patterns in the cryptic, poorly-keyed memo on my Centro. (Note to organizers -- get more wireless connections next time. I didn't bother with my laptop because there was very little chance of getting on the network)

Private LOCKKSS Networks were everywhere. MetaArchive, Arizona State Library and Archives PeDALS, Data-PASS, ETD preservation, and, of course, journal content. It's interesting to see the LOCKSS distributed and self-replicating architecture being used for all types of content.

Distributed and/or replicated storage overall was definitely a trend. iRODS was mentioned in several sessions, I learned more about Dataverse, and I attended a meeting with the FACIT partners.

The packaging and transfer of files between institutions was discussed quite a bit. I was pleased to see the positive reaction to the BagIt package standard that LoC has been working on, which has been put into use with some NDIIPP partners including CDL and Stanford. I was really intrigued with a presentation that Tom Habing did on the ECHO DEPository project. I've seen it presented before, but something really clicked this time when I saw their Hub and Spoke architecture and listened to him talk about packaging between systems and services.

What really stuck with me was something that Micah Altman from Harvard said. He was discussing selection for digital preservation and declared that we need to "select the selectors" in identifying what should be preserved, because we can't save everything. If we identify key researchers and tie preservation to their research, we're assured to capture at least some vital resources. But there are so many disciplines that no one institution can identify what should be preserved, so the corollary need is for many, many institutions to involves themselves in selection and preservation, so there is more preservation coverage for the future. I was glad to hear selection described as a necessary activity.

Friday, June 06, 2008

BagIt

A press release went out this week on digitalpreservation.gov about the BagIt format specification. BagIt is a lightweight specification for the description of data packages meant for transfer between institutions. The need for such a standard was initially identified in working with NDIIPP partners such as CDL (John Kunze of CDL played a major role in the format development and testing and is one of the principal authors). Other partners have expressed interest, and we are moving forward with a prototype submission web app that will take advantage of BagIt.

My colleague Ed Summers has posted a fantastic overview of the specification to which I can add nothing except reinforcing the kudos deserved by everyone involved.

There is an official Internet-Draft and comments are welcome.

Thursday, January 17, 2008

LC and flickr

One couldn't turn around online today without seeing a post on the Library of Congress flickr experiment today: Read Write Web, Shifted Librarian, eFoundations, Catalogablog, Digital Koans, etc, and of course on the LOC blog.

In case you've been under a rock, LC has put up 3,000 images of photographs from their Prints and Photographs division for which no rights restrictions are known to exist, and are asking folks to tag them to help in enriching the metadata. The pilot project is in a new area of flickr called "The Commons." LC has a GREAT FAQ on participation.

Three things stand out for me:

They have made use of a new rights statement "no known copyright restrictions," which they may or may not make available for other images uploaded to flickr.

They are interested in hearing from other cultural heritage institutions to gauge interest in other projects becoming part of The Commons.

This is a pilot, and they do not know how long they might keep the images on flickr. So, review and tag these while you can. They have other candidate collections in mind.