Showing posts with label preservation. Show all posts
Showing posts with label preservation. Show all posts

Friday, December 31, 2010

2010 in Review

I am mortified to find that I only posted three times in 2010. I'd like to be able to say that it was for some glorious reason, but, to be honest, I just haven't made time. I've tweeted (and re-tweeted) quite a bit. I went out on the road and spoke at a number of conferences. I had one article come out that I wrote in 2009. But in 2010, I just didn't make much time to write.

I could make a public resolution, but that's risky...since the proof of failure or success would be right here. So no resolution.

There were a number of topics that caught my attention this year.

Twitter donated their archive to the Library of Congress this year. It has been startling to me just how much public outcry there was. It's not unlike a journal -- a very public journal, aggregated from millions of people. Given the Library's collections of personal papers and man-on-the-street collections, Twitter seems perfectly in keeping with the Library's other analog and digitial collections. And, the Library archives web sites. In one sense, archiving Twitter is archiving another part of the web.

I got more involved in web archiving this year. I've been involved in web archiving before - I started up an initiative to archive course web sites in 2000. It's been gratifying to become involved again, and see how much has been saved and will be saved. Not just at institutions like the Library of Congress or other national libraries or research universities (check out the institutions that are part of the IIPC), but through personal, volunteer efforts. I'm looking at you, Archive Team.

Archives are acquiring increasing numbers of born-digital collections. I've been thrilled to see the increased interest in the use of digital forensics tools in the appraisal and processing and accessing of such collections. But there are challenges. Archives are looking at vintage media, which often requires vintage hardware and software. The collection at the Library's Package Campus is something to behold, but I shudder at what it will take to keep the equipment operational. To understand some of the challenges, a couple of key reports came out this year, on Preserving Virtual Worlds and Digital Forensics in Cultural Heritage.

In that same vein, I've always been interested in computing history. I am going to resolve to return to reading more on that subject.

I've been thinking a lot about documenting computing history in aid of digital preservation. There are multiple initiatives to document and verify file formats. There is at least one initiative to document carrier media. There are archives of manuals and media. I am thinking a lot about what other sorts of documentation are needed - operating systems, application software, hardware of all types... I heard these challenges subtly woven through many presentations and discussions at our storage architecture meeting this year.

I've been thinking a lot about standards this year. That comes from working with an initiative to collect content from the wild, as published. How do we collect things as they are, but minimize the grief required to deal with variety when ingesting them into a managed environment? That I wish I had a great answer for. But that's what 2011 will in part be about.

I almost forgot about personal digital archiving! We started an initiative at the Library in 2010, with Personal Archiving Day and a Personal Digital Archiving booth at the National Book Festival. I loved working at both events. There's a lot more to be done about public awareness and promoting best practices.

Saturday, June 27, 2009

BagIt video

The first in a planned series of digital preservation videos is available on the digitalpreservation.gov site -- an introduction to BagIt! Brian Vargas did a great job as "the talent" -- e.g., the narrator -- but folks should know that Brian was not selected just for his acting experience: he wrote many of our transfer tools (like the transfer scripts on SourceForge) and is a co-author of the BagIt specification.

The video premiered this week at the annual NDIIPP Partner's Meeting to great acclaim. It's aimed at a general audience.

EDIT: The NDIIPP site has added a great new page on the Transfer Tools with a link to the video.

Friday, June 26, 2009

Chesapeake Project Legal Information Archive

I came across a very interesting resource today -- the Chesapeake Project Legal Information Archive -- and the just-released results of a study they did on archiving legal resources on the web:

The Chesapeake Project Legal Information Archive has released a comprehensive report evaluating its digital preservation efforts during the project's two-year pilot phase.

The project evaluation reveals that nearly 14 percent — or approximately one in seven — of the online publications archived between March 2007 and March 2009 have already disappeared from their original locations on the Web but, due to the project's efforts, remain accessible via permanent archive URLs. A similar analysis in 2008 showed that slightly more than 8 percent of archived titles had disappeared from their original URLs, demonstrating a dramatic increase in "link rot," or inactive URLs, among archived content over the past year.

During the two-year pilot phase, the libraries participating in the project archived more than 4,300 digital objects and tracked more than 177,000 visits to www.legalinfoarchive.org, the home of The Chesapeake Project's digital archive collections. Users of the project's Web site visited from educational, government, and military institutions in the United States, as well as from countries abroad throughout the Americas, Europe, the Middle East, Asia, Africa, Australia, and the Pacific Islands.

Not too surprisingly, the second highest class of domain to where resource loss is found is .edu, after .info. Academic institutions are not always very conscientious about preserving access to their content, and with their academic term structure and the movement of faculty between institutions, web content on .edu sites is highly variable in its longevity. I don't see a characterization of how old the resources are that they harvested -- that can be very difficult to identify -- but it is a high percentage of bitrot, and there was quite an increase from the end of the first year to the end of the second year.

Download the PDF of their report.

Sunday, April 05, 2009

DigCCurr 2009

I was in Chapel Hill the first week of April for the DigCCurr 2009 conference and to attend a meeting to brainstorm about personal digital collection preservation. I thought the conference was very good, better than the first one in 2007. I saw many excellent presentations, had some great conversations, and got a good response to my presentation on LC's work with file transfer and inventory tools. As with the last conference, I walked out thinking that I should have been an archivist.

I strongly recommend the proceeding form DigCCurr 2009. They're available as a free download from Lulu, or you can buy a POD version. You can also look up the very active twittering history at #digccurr.

I found it strangely hard to write up my notes from this meeting. I think it's because I'm still struggling with some aspects of the digital preservation problem space.

I absolutely agree that the activities of traditional archival practice have a place in the preservation of digital records. Where I found myself disagreeing with some presenters is in the balance between collecting and saving what we can versus an appraisal process to select what we will collect/save. In collection development practices for general collections, there is the often-held discussion about never knowing what might prove useful in the future, so it is a disservice to be too selective now. I guess that I have taken that point of view to heart, and I want to see our institutions cast as open a net as possible for digital collections. If we don't grab it when we can, there will be nothing to select.

I also found myself bristling occasionally over the implied scope of the term "digital collections" as I most often heard that phrase used at the meeting. There was very much a focus on electronic records and the digital realm of personal papers. Of course there were some great discussions around multimedia, web sites, audio/video, and image collections, but what I pretty much never heard anybody mention was born-digital scholarship and teaching and learning materials.

My first web site preservation project was at the Harvard Design School in the late 1990s, where, while developing courseware software, I realized that we were losing the history of what we taught and the products of the courses as we overwrote sites every term. Part of an institution's records are its lists of course offerings, course syllabi and reading lists, and, for some courses, the projects that the students created and put online in the course site. This was particularly true at at graduate school with programs in architecture, landscape architecture, and urban planning where the studio courses produced important site-specific work and case studies that was often lost after every term. I felt so strongly about this that I launched a course site preservation project that would have involved retrieving sites off server archives. We were looking at using METS (in its early days) to map the sites. But, as often happens, I ended up leaving before the project got very far along and no one felt nearly as devoted to the project as I did and it didn't go very far.

At UVA we launched a project called "Sustaining Digital Scholarship" to preserve born-digital scholarship, primarily in the humanities and social sciences. We instituted a technical assessment process and were working on documenting and migrating some major digital scholarly resources with varying strategies. That project is still going on in a limited way. It can take a lot of resources to assess and document a large digital archive.

That said, I was excited by some of the tools that I saw. ACE from the University of Maryland. MOPSEUS from Greece. The PARSE.Insight draft preservation roadmap. CASPAR for representation information. PLATO and Hoppla from Austria. LANL's ReMember Framework for OAI-ORE. CDL's Pairtree directory structure. Prometheus and MediaPedia from Australia. All very much worth looking into.

There was also a thread in this meeting on the use of digital forensics, transitioning some tools and practices from legal digital forensics into archival digital forensics. This interested me very much and I intend to read up in this area.

Tuesday, March 24, 2009

Ada Lovelace Day

I have had the pleasure in my life of working with a number of strong (and strong-willed) women who have seen me through various stages of my career. On the occasion of Ada Lovelace Day, I'd like to write about a colleague who I have known for many years, although we only had the opportunity to work in the same place for 4 weeks: Caroline Arms.

Caroline joined the Library of Congress in 1995 to work on the American Memory project. While the initial focus of the project was digitization and access, she saw the underlying issue that was created by such an effort: preservation. There was a profound lack of awareness in the library world about digital preservation at the time.

Caroline thought long and hard about the life cycle of digital objects, focusing in particular on one of the most vital areas that have consequences for all preservation efforts: standards for metadata and file formats. Preservation is always easier if good choices are made about digital formats. Curators should make collection decisions knowing which formats will and won’t be easily sustainable. For an object to be useful long into the future, its formats should be carefully selected and the specifications and characteristics of its formats must be documented.

Caroline and LC colleague Carl Fleischhauer’s exhaustive format research led to their creation of the Digital Formats web site, the first definitive inventory of information about current and emerging digital formats. The site is an essential resource for the international digital preservation community. Caroline also made a concerted effort to promote the use of formats with open standards, and to shepherd file formats through the standards review process.

She was also involved with the development of the Open Archives Initiative Protocol for Metadata Harvesting. I first met Caroline working on a collaborative OAI harvesting project, and I owe much of my expertise to her mentoring.

It was a great loss to LC that Caroline retired in June 2008. She did not retire from the community, however, and is participating in a LC group looking at metadata even now.

Thank you, Caroline.

Friday, January 02, 2009

DCC obsolescent data and files challenge

There's still time to send a message to Chris Rusbridge at the Digital Curation Center to enter his personal data recovery challenge:

I will do my best to recover the first half dozen interesting files that I’m told about… of course, what I really mean is that I’ll try and get the community to help recover the data. That’s you!

OK, I define interesting, and it won’t necessarily be clear in advance. The first one of a kind might be interesting, the second one would not. Data from some application common of its time may be more interesting than something hand-coded by you, for you. Data might be more interesting (to me) than text. Something quite simple locked onto strange obsolete media might be interesting, but then again it might be so intractable it stops being interesting. We may even pay someone to extract files from your media, if it’s sufficiently interesting (and if we can find someone equipped to do it).

The only reference for this sort of activity that I know of is (Ross & Gow, 1999, see below), commissioned by the Digital Archiving Working Group.

What about the small print? Well, this is a bit of fun with a learning outcome, but I can’t accept liability for what happens. You have to send me your data, of course, and you are going to have to accept the risk that it might all go wrong. If it’s your only copy, and you don’t (or can’t) take a copy, it might get lost or destroyed in the process. You’ll need to accept that risk; if you don't like it, don't send it. I might not be able to recover anything at all, for many reasons. I’ll send you back any data I can recover, but can’t guarantee to send back any media.

The point of this is to tell the stories of recovering data, so don’t send me anything if you don’t want the story told. I don’t mind keeping your identity private (in fact good practice says that’s the default, although I will ask you if you mind being identified). You can ask for your data to be kept private, but if possible I’d like the right to publish extracts of the data, to illustrate the story.
His deadline is Twelfth Night, January 5, 2009. Read his post for the rest of the details.

I've been thinking a lot lately about data migration and recovery, and I think this is a great way to illustrate the challenges and solutions -- with real data and the stories behind its creation, loss, and potential recovery.

Tuesday, December 30, 2008

archiving the bush administration

There is a great article on ars technica today about the major processing effort that will be required at the National Archives when the Bush administration leaves office. The ars technica piece references a New York Times article on the topic from this past weekend.

This section really strikes home:

The contingency plan will entail "ingesting" the Bush White House's data into a separate system before integrating it with the ordinary archive. As the plan explains, "the current PERL [Presidential Electronic Records Library] system architecture was not scalable to actually support the volume of records that are expected from the current Presidential administration."

It's not just size that matters, though: the Archives will also need to process reams of information locked in some quaint proprietary formats. The RMS index, for example, "consists of an implementation of a customized older version of Documentum running on Oracle, with image files (including copies of scanned records) incorporated as objects in the database." The photos are stored in a "proprietary photo management software called MerlinOne, running on Microsoft SQL as the database engine," and it has apparently taken several months to extract the images and metadata for relinkage outside the Merlin format.

First, the use of quotation marks should remind us all that "ingest" means absolutely nothing to someone who is not a repository manager.

I have participated in some discussions about a potential data migration project at work. I recently saw an inventory of media formats -- not file formats, but media formats -- that the project would need to encompass, and it is lengthy. The only source I can think of for hardware to read some of the formats is EBay. That doesn't even take into account the files themselves. It's interesting how quickly a format becomes obsolete, and how many customized systems federal agencies use.

Sunday, November 16, 2008

oxford university preserv2 project session at dlf

My brief, unstructured notes from a presentation by Sally Rumsey from Oxford University on the Preserv2 project, at the DLF Fall 2008 Forum.

  • Effort focused primarily on IRs with research data, but applies to all types of repositories.
  • “Boxed up -- but innovative at the time – repository environment.” 6 different repositories at Oxford, multiple Fedora & Eprints, some sharing the same storage.
  • A desire for decoupled services -- distributed storage and transportable data. Scalability issue.
  • Implemented British National Archives seamless flow approach to preservation. There are three categories of activities: Characterization (inventory and format identification). Preservation planning and technology watch (using the PRONOM technology watch service). Preservation Actions (file migration, rendering tools, etc, based on repository policy).
  • Smart Storage. Ability to address objects in open storage through a repository layer or directly in the storage system. DROID was used as the tool to verify new items as they are stored, and continually check files in storage. They are implementing a scheduler.
  • Oxford looking at Sun Honeycomb server architecture.

Friday, October 31, 2008

JISC Digital Preservation Policies Study

JISC has released a two-part study of digital preservation policies: Digital Preservation Policies Study and Digital Preservation Policies Study, Part 2: Appendices—Mappings of Core University Strategies and Analysis of Their Links to Digital Preservation. The study aims to provide an outline model for digital preservation policies and to analyse the role that digital preservation can play in supporting and delivering key strategies for higher ed institutions.

Wednesday, October 08, 2008

DCC Curation Lifecycle Model

Via the Digital Curation Blog, I came across the DCC Curation Lifecycle Model. This is a very interesting high-level overview of the life cycle stages in digital curation efforts. There's an introductory article available.

The model proposes a generic set of sequential activities -- creating or receiving content, appraisal, ingest, preservation events, storage, etc. There are some decisions points at the appraisal and preservation event stages about next steps -- refusal, reappraisal, migration, etc. A colleague and I sat together and looked it over this afternoon. We were both looking at it from a perspective of a digital collections repository and not an IR, and the model was designed primarily with IRs in mind, so our thoughts are coming from a different place in terms of what we wanted to see additionally taken into account in the visualization.

There's a "transform" activity -- definitely something that takes place potentially multiple times in a data life cycle. In the visualization this appears sequentially after "store" and "access, use and reuse." This is an activity that's hard to include in a visualization of a sequence because it can take place at so many points, but it feels like it should be earlier in the sequence, perhaps before those two steps.

The next ring is labeled with the activities "curate" and "preserve" with arrows. Does the placement of the terms and arrows mean anything in relation to the outermost ring? Are "ingest," "preservation activity" and "store" part of "preserve" and the rest part of "curate?" Or does this more simply represent ongoing activities?

The center of the model is the data. It's surrounded by a ring for descriptive and presentation information. It's an activity of central importance and is directly related to the data as is shown, but we weren't sure how its placement related to the sequence of tasks in the visualization.

"Preservation planning" is the next ring out. Planning and implementation are a central, ongoing activity. We also weren't sure when this ongoing activity meshed with the sequence.

"Community watch and participation" is the last remaining inner ring. It's also on ongoing activity. What actions might the outcomes of this activity affect?

Overall, this is a good model for planning. It's challenging to create a visualization for complex processes and dependencies and this covers a lot of ground. And of course it's meant to be generic and high-level, to be made more concrete by an institution that makes use of it. It certainly stimulated our thinking in terms of how we might model our data life cycle and the dependencies between the various tasks.

NOTE: Sarah Higgins, who created the model, has provided excellent responses to my thoughts and questions in the comments to this post. Please read them!

Tuesday, July 15, 2008

ndiipp partners meeting

Last week I attended the three-day 2008 meeting for the partners in the Library of Congress National Digital Information Infrastructure and Preservation Program (NDIIPP). Yesterday a colleague who couldn't attend asked me what stood out for me in the program. I didn't take a lot of notes -- I kept forgetting to because I just wanted to listen -- but I see some patterns in the cryptic, poorly-keyed memo on my Centro. (Note to organizers -- get more wireless connections next time. I didn't bother with my laptop because there was very little chance of getting on the network)

Private LOCKKSS Networks were everywhere. MetaArchive, Arizona State Library and Archives PeDALS, Data-PASS, ETD preservation, and, of course, journal content. It's interesting to see the LOCKSS distributed and self-replicating architecture being used for all types of content.

Distributed and/or replicated storage overall was definitely a trend. iRODS was mentioned in several sessions, I learned more about Dataverse, and I attended a meeting with the FACIT partners.

The packaging and transfer of files between institutions was discussed quite a bit. I was pleased to see the positive reaction to the BagIt package standard that LoC has been working on, which has been put into use with some NDIIPP partners including CDL and Stanford. I was really intrigued with a presentation that Tom Habing did on the ECHO DEPository project. I've seen it presented before, but something really clicked this time when I saw their Hub and Spoke architecture and listened to him talk about packaging between systems and services.

What really stuck with me was something that Micah Altman from Harvard said. He was discussing selection for digital preservation and declared that we need to "select the selectors" in identifying what should be preserved, because we can't save everything. If we identify key researchers and tie preservation to their research, we're assured to capture at least some vital resources. But there are so many disciplines that no one institution can identify what should be preserved, so the corollary need is for many, many institutions to involves themselves in selection and preservation, so there is more preservation coverage for the future. I was glad to hear selection described as a necessary activity.

Friday, April 04, 2008

PREMIS 2.0

The PREMIS Editorial Committee has released PREMIS Data Dictionary for Preservation Metadata, version 2.0, a revision of the May 2005 report. A draft XML schema -- still undergoing a month of review before its final release -- is also available.

Audiovisual Research Collections and Their Preservation

TAPE (Training for Audiovisual Preservation in Europe) has published Audiovisual Research Collections and Their Preservation. This report looks at the requirements for access and re-use, focusing on the potential of digitization for creating distributed content-based archives.

Tuesday, March 25, 2008

document migration

There's a very thoughtful post on the Digital Curation blog about the use of Open Office as a migration tool.

Conversions Plus has personally saved my hide when I needed access to older file formats, but it's not meant as a preservation tool. What tools do people use for conversion? How do they scale?

Monday, March 03, 2008

Warrick

I just came across Warrick, a neat research project that takes advantage of cached web crawls to restore lost web site. Warrick is a utility for reconstructing or recovering a website when a back-up is not available. It searches the Internet Archive, Google, Live Search, and Yahoo for stored pages and images and will save them to your filesystem. It's not guaranteed and it is a research project and not a production service, but it could help when there's no other option. It falls under the general category of research that they call "Lazy Preservation," a phrase that I can see some loving and some hating.

When I followed the link to the about page and I saw that it was a research project at Old Dominion University, I immediately suspected that it was one of Michael Nelson's students, and it was. I briefly blogged Joan Smith's mod-oai work last year. Michael is always working on something interesting, especially his current work on OAI-ORE.

Friday, February 15, 2008

why bind?

I love an effective visualization, and my colleague Holly has a very effective visualization to make the case for her binding budget.

Tuesday, January 15, 2008

D-Lib article on CRATE

The January/February 2008 issue of D-Lib includes an article by Joan Smith and Michael Nelson on their proposed CRATE mod-oai utility to produce preservation metadata for web resources. I saw this work presented at Open Repositories 2007 and I thought it was interesting then. It's a good article on an interesting proposal.

Friday, January 11, 2008

Report on LC/SDSC Data Transfer and Storage Tests

The Library of Congress has released the report Data Center for Library of Congress Digital Holdings: A Pilot Project; Final Report. I've only had the chance to glance through it, but it looks to be a very clear and straightforward report on technical, procedural, and cultural issues encountered in this digital image collection data transfer experiment, and how they resolved those issues (or not).

http://www.digitalpreservation.gov/pdf/SDSC_LC_data-storage_report_2.pdf

Tuesday, December 04, 2007

Alliance for Permanent Access

I haven't been able to track down much on the new Alliance for Permanent Access. There's this press release. Did anyone attend the Second International Conference on Permanent Access to the Records of Science held in Brussels on November 15, 2007 where the Alliance was launched?

Friday, October 05, 2007

digital lives project

Digital Koans posted about the Digital Lives research project, "focusing on personal digital collections and their relationship with research repositories."

For centuries, individuals have used physical artifacts as personal memory devices and reference aids. Over time these have ranged from personal journals and correspondence, to photographs and photographic albums, to whole personal libraries of manuscripts, sound and video recordings, books, serials, clippings and off-prints. These personal collections and archives are often of immense importance to individuals, their descendants, and to research in a broad range of Arts and Humanities subjects including literary criticism, history, and history of science.

...

These personal collections support histories of cultural practice by documenting creative processes, writing, reading, communication, social networks, and the production and dissemination of knowledge. They provide scholars with more nuanced contexts for understanding wider scientific and cultural developments.

As we move from cultural memory based on physical artifacts, to a hybrid digital and physical environment, and then increasingly shift towards new forms of digital memory, many fundamental new issues arise for research institutions such as the British Library that will be the custodians of and provide research access to digital archives and personal collections created by individuals in the 21st century.

I very much look forward to seeing the results of this work, as university archives and institutional repositories increasingly have to cope with not only managing and preserving deposited personal digital materials, but have to potentially describe, organize, and make such collections usable.

While not the focus of their study, anyone who has ever supported teaching with images knows a tangential area of this problem space intimately. Faculty develop their own collections of teaching images -- their own analog photography, purchased slides, digital photography, images found on the open web, images from colleagues, etc. We have licensed images and surrogates of our own physical collections. They want to use materials from their own collections and our repositories together in their teaching. What is the relationship between their image collections and our repositories and teaching tools? Do we integrate their collections into ours? Do we have a role in digital curation and preservation of their data used in teaching and research, which happen to be images? We struggle with the legal and resource allocation issues every day.