Showing posts with label digital library systems. Show all posts
Showing posts with label digital library systems. Show all posts

Thursday, January 15, 2009

d-lib article on some LC tool development

My colleague Justin Littman has just published an excellent article in the January/February 2009 issue of D-Lib Magazine: "A Set of Transfer-Related Services."

"The Office of Strategic Initiative's (OSI) Repository Development Team (RDT) is developing a portfolio of services and components to address the challenges posed by scaling transfer processes. While the portfolio is expanding, the focus of this article will be on two core services, the Inventory Service and the Workflow Service. Before proceeding to examine these services, it will be useful to further delineate the transfer problem space. After examining these services, their role in mitigating preservation risks will be considered."

Sunday, November 30, 2008

costs of operating a data center

In recent weeks I've been in a number of meetings about storage architecture. In comparing potential solutions for a particular project's needs, there has been a lot of discussion about requirements for space, cooling, power, etc., which are some of the metrics for comparison of the proposals. Today I came across a posting by James Hamilton from Microsoft's Data Center Futures Team, on misconceptions about the cost of power in large-scale data centers.

There is an error in the table for the two amortization entries: the 180-month amortization of the facility says 3 years in the notes column when it should say 15, and the 36-month amortization of the servers says 15 years when it should say 3. The right numbers are used in the calculation -- it's just the explanatory notes that are switched.

There are some very interesting calculations that I plan to forward to other folks. One commenter points out that the figures and graph don't take staffing into account, which, according to Hamilton, is because that cost is a very small percentage of the cost. At Microsoft's scale that's probably true. I guess we're a relatively small- to medium-scale data center, because we are definitely taking that into consideration. And, while the cost of power may be a lower percentage of overall cost than previously considered, power capacity is still definitely a factor when you're considering adding a number of servers and switches.

Sunday, November 16, 2008

oxford university preserv2 project session at dlf

My brief, unstructured notes from a presentation by Sally Rumsey from Oxford University on the Preserv2 project, at the DLF Fall 2008 Forum.

  • Effort focused primarily on IRs with research data, but applies to all types of repositories.
  • “Boxed up -- but innovative at the time – repository environment.” 6 different repositories at Oxford, multiple Fedora & Eprints, some sharing the same storage.
  • A desire for decoupled services -- distributed storage and transportable data. Scalability issue.
  • Implemented British National Archives seamless flow approach to preservation. There are three categories of activities: Characterization (inventory and format identification). Preservation planning and technology watch (using the PRONOM technology watch service). Preservation Actions (file migration, rendering tools, etc, based on repository policy).
  • Smart Storage. Ability to address objects in open storage through a repository layer or directly in the storage system. DROID was used as the tool to verify new items as they are stored, and continually check files in storage. They are implementing a scheduler.
  • Oxford looking at Sun Honeycomb server architecture.

university of texas repo atom use session at dlf

My somewhat unstructured notes from a presentation by Peter Keane from UT Austin on the use of Atom and Atom/Pub in their DASe repository, at at the DLF Fall 2008 Forum.

  • DASe project: lightweight repository, 100+ collections, 1.2 million files, 3 million metadata records.
  • DASe has replaced their image reserves system. Home grown (“built instead of borrowed”), originaly prototyped 2004/2005.
  • They didn’t originally plan to build a repository, they were building an image slideshow and ended up with a repository, too.
  • It’s a data first application. Data comes from spreadsheets, FM, Flickr, iPhoto, file headers, etc. System includes a variety of different collection-based data models. Needed to map to/from standard schemas. Accepted as is, no normalization or enrichment at all.
  • SynOA: Syndicated Oriented Architecture. Importance recognized in being RESTful. DASe is a Rest framework.
  • Use the Atom publishing protocol to represent collections and items and searches. Used internally between services, including upload and ingest (uses http get, post, etc). Everything is Atom with a UI (Smarty PHP templates) on top of it.
  • Working on a Blackboard integration.
  • Interesting use of Google spreadsheets – create Google spreadsheet for whatever they have a name/value pairs, automatically outputs atom, can ingest from feed.
  • No fielded search across collections, only within a single collection. They could map across data models to a common standard, but haven’t. (corrected as per comment below)
  • Repositories were considered a door to libraries, all trying to create a better door. This is not the right concept, instead should be exposing in a standard way to any and all services.
  • Loves REST; used the term “RESTafarian.”

Friday, November 14, 2008

omeka 0.10

Omeka v 0.10 has been released. Omeka 0.10b incorporates many requested changes: an unqualified Dublin Core metadata schema and fully extensible element sets to accommodate interoperability with digital repository software and collections management systems; elegant reworkings of the theme and plugin APIs to make add-on development more intuitive and more powerful; a new, even more user friendly look for the administrative interface; and a new and improved Exhibit Builder.

Greenstone release

Greenstone v2.81 has been released. Improvements include handling filenames that include non-ASCII characters, accent folding switched on by default for Lucene, and character based segmentation for CJK languages. There are many other significant additions, including the Fedora Librarian Interface (analogous to GLI, but working with a Fedora repository).

Friday, October 17, 2008

digital book access at John Hopkins

Jonathan Rochkind has posted a great description of digital book access features that he's put into production in the link resolver and OPAC at Johns Hopkins. They're remarkable in the sense that he's taken advantage of so many different service APIs (Google Books, IA, OCLC, Amazon, HathiTrust) to provide functionality with conditional options to provide as much collection coverage as possible.

Monday, September 01, 2008

kete 1.1

I've blogged about Kete before - version 1.1 has been released.

From the announcement:

Kete 1.1 is now available with a giant helping of new features and improvements. This is also the first release where you can grab Kete from our code repository's new home at Github.com. See http://kete.net.nz/site/topics/show/25-downloads for details or browse the code online at http://github.com/kete/kete/.

For those who haven't seen Kete in action, Kete is open source software that enables communities, whether the community is a town or a company, to collaboratively build their own digital libraries, archives and repositories. Kete combines features from Knowledge and Content Management Systems as well as collaboration tools such as wikis, blogs, tags, and online forums to make it easy to add and relate content on a Kete site. You could create a service like Google's Knol for your community using Kete.

An in-depth list of features and issues resolved can be found at http://kete.net.nz/documentation/topics/show/182-kete-11-features-and-bug-fixes , but here are some highlights:

Friday, August 15, 2008

Red Island Repository Institute

This week the Red Island Repository Institute took place, with a week-long immersion in all things Fedora. The instructors were Sandy Payette, Richard Green, and Matt Zumwalt. It would be hard to think of people who are a better choice to teach the institute besides these three!

Powerpoints from presentations are online, and they provide a great overview of Fedora.

Wednesday, August 06, 2008

time for links and nothing more

I'm really swamped these days, and only have time to post some links to things that caught me eye during the past week:

vi.sualize.us seems like a really interesting social bookmarking tool for images. Perhaps they'll learn what delicious learned and give up the tortured . 's.

William Patry stopped blogging
. I'm not surprised if folks thought his personal blog was the word of Google. It's sad that he also decided to erase his archives, but I understand that he didn't want his past postings to live on and continue to be misunderstood.

It seems that Google is making some of its machine-translation technologies and translation management tools available to human translators, at least as a beta. I'm working with a project that requires translation into 7 languages. Managing this process is very challenging, and I've seen some very bad tools for the process.

Following the Digitization and the Humanities Symposium, Jennifer Schaffner and Merilee Profitt wrote a brief report, The Impact of Digitizing Special Collections on Teaching and Scholarship: Reflections on a Symposium about Digitization and the Humanities. The report acts as a summary of the symposium, and also gives some calls to action, especially about metrics for success.

Duke has launched its Open Library Environment Project with Mellon support. Its focus on back-end open tools is worhtwhile, but I'm not sure I know how this will be integrated with other activities in the community.

Wednesday, July 30, 2008

Hooray -- Fedora 3.0 released

Hooray -- the highly anticipated (at least by me) formal release of Fedora 3.0 is available.

Excerpted from the press release:

Fedora 3.0 features the Content Model Architecture (CMA), an integrated structure for persisting and delivering the essential characteristics of digital objects in Fedora. The software is available at
<http://www.fedora-commons.org/> and at <http://sourceforge.net/projects/fedora-commons>. The Fedora CMA plays a central role in the Fedora architecture, in many ways forms the over-arching conceptual framework for future development of Fedora Repositories. Fedora 3.0 features include:
  • Content Model Architecture - Provides a model-driven approach for persisting and delivering the essential characteristics of digital content in Fedora
  • Fedora REST API - A new API that exposes a subset of the Access and Management API using a RESTful Web interface contributed by MediaShelf
  • Mulgara Support - Fedora supports the Mulgara 2.0 Semantic Triplestore replacing Kowari -Migration Utility - Provides an update utility to convert existing collections for Content Model Architecture compatibility
  • Relational Index Simplification - The Fedora schema was simplified making changes easier without having to reload the database and significantly increasing scalability
  • Dynamic Behaviors - Objects may be added or removed dynamically from the system moving system checks into run-time errors
  • Error Reporting - Provides improved run-time error details
  • Multiple Owner as a CSV String - Enables using a CSV string as ownerID and in XACML policies
  • Java 6 Compatibility - Fedora may be optionally compiled using Java 6 while retaining support for Java Enterprise Edition 1.5 deployments
  • Relationships API - API-M has been extended to enable adding, removing, and discovering RDF relations between Fedora objects
  • Revised Fedora Object XML Schemas - The new schemas are simpler, supporting the CMA and removing Disseminators
  • Atom Support - Fedora objects can now be imported and exported in the Atom format
  • Messaging Support - Integrates JMS messaging for sending notification of important events
  • Validation Framework - Provides system operators a way to validate all or part of their repository, based on content models
  • 3.0-Compatible Service Releases - New versions of the OAI Provider and GSearch services are compatible with Fedora 3.0. The GSearch release also enables messaging support for GSearch, which allows for more robust and seamless integration with the Fedora repository.
I have been waiting for the CMS for some time -- this update to the architecture will greatly improve the flexibility of a Fedora implementation by removing the tight bindings between objects and disseminators and allowing for easier disseminator updating. The validation support is also key if one is interested in working with workflow engine to automatically process tasks and validate production outcomes. I am intrigued by the Atom support -- Dan commented on BagIt/SWORD as a possible repository SIP in one of our discussions at RepoCamp. This could become a very real experiment.

Tuesday, July 29, 2008

Fedora and DSpace collaboration

A press release hit the streets today about a formal collaboration between the Fedora Commons and the DSpace Federation. Excerpt from the press release:

The decision to collaborate came out of meetings held this spring where members of DSpace and Fedora Commons communities discussed multiple dimensions of cooperation and collaboration between the two organizations. Ideas included leveraging the power and reach of open source knowledge communities by using the same services and standards in the future. The organizations will also explore opportunities to provide new capabilities for accessing and preserving digital content, developing common web services, and enabling interoperability across repositories.

In the spirit of advancing open source software, Fedora Commons and DSpace will look at ways to leverage and incubate ideas, community and culture to:

1. Provide the best technology and services to open source repository framework communities.

2. Evaluate and synchronize, where possible, both organizations'technology roadmaps to enable convergence and interoperability of key architectural components.

3. Demonstrate how the DSpace and Fedora open source repository frameworks offer a unique value proposition compared to proprietary solutions.

The announcement came on the heels of an event sponsored by the Joint Information Systems Committee's (JISC) Common Repository Interface Group (CRIG) held at the Library of Congress. The event, known as "RepoCamp," was a forum where developers gathered to discuss innovative approaches to improving interoperability and web-orientation for digital repositories. Sandy Payette, Executive Director of Fedora Commons, and Michele Kimpton, Executive Director of the DSpace Foundation, reiterated their commitment to collaboration and encouraged input and participation from both communities as work gets underway.
The full press release is available. Theres a great photo of Sandy and Michelle in a ceremonial handshake at LoC. Sandy and Michelle led a brief discussion about this last Friday at RepoCamp, and was exciting to watch this initiative launch.

Monday, July 28, 2008

RepoCamp

Last Friday I spent all day at RepoCamp. There were at least 40 participants from I don't know how many institutions! Major kudos to David Flanders from the JISC Common Repository Interface Group (CRIG) who did a fabulous job organizing the event, and to Ed Summers who facilitated the LoC side.

There was a lot of great discussions around SWORD and OAI-ORE. I was happy to have the opportunity to talk about BagIt with a group who hadn't encountered it yet, and we had some really interesting discussions. Talking through use cases beyond our initial simple use case -- files from Institution A are transferred to Institution B and stored for preservation with no active access -- there is an obvious need for BagIt profiles that specify what is contained in a Bag and how it's organized for other uses -- like potentially as a SIP for ingest into a repository. Folks were also really interested in the idea of "Holey Bags" where the manifest is a list of URIs for retrieving files. Ideas were batted around about crawls that start out with a minimal manifest of URIs, capture those files, generate checksums, follow links from those files to capture more files and checksums, ending up with a Bag generated on-the-fly from that crawl so you have captured the files and record of the URIs where the files were found. A really interesting suggestion was the use of an OAI-ORE Resource Map to instantiate such a capture. Or for that matter, serve as the fetch file. Bags of course can simply be files on disk, but when it's a Bag of web resources you might want more structure than just a list of locations the files came from. After listening to the discussions I'm convinced that the work we're doing (I should say Ed is doing) with a web app for a Bag deposit service that uses SWORD is going in the right direction. I think were developing some real traction with BagIt.

They video recorded all the elevator pitches and the reports to the whole group. I don't know if, where, or when they'll be available. I am not a fan of seeing myself on video.

It was nice to see some of the Fedora team -- Sandy Payette, Dan Davis, Eddie Shin -- and some UVA colleagues. Its only been 3 or so months but it feels like I left so long ago. I was pleased to meet Ben O'Steen from Oxford (we were following his Fedora IR work when I was at UVA) but I didn't get the chance to really sit down and talk with him. I had so many other interesting conversations that I need to follow up on ...

Tuesday, July 22, 2008

what is a repository?

Yesterday a colleague was chatting with me about what make up a repository. Have we been overthinking what is needed? Can we simplify the tools we use? Recombine lightweight tools in a new way?

This was very timely because I'd seen a posting that JISC's Information Environment team is experimenting with IdeaScale to have a discussion about defining repositories to feed into JISC work on repository architecture.

First -- about IdeaScale:

It begins with an idea posted to your IdeaScale community by a user. Each idea can be expanded through comments by the community. The ultimate measure of an idea is determined by a voting system. Any idea can be voted to the top or buried back down to the bottom. It combines the "wisdom of the crowds" concept with Web 2.0 models like Digg.
I think it's interesting that JISC is trying this approach -- have discussants set out a series of statements about repositories, allow comments, and let members of the community sign up to vote +1 or -1 on the positions.

I have a love/hate relationship with the word "repository." It's next to impossible to define or describe, but I haven't been able to come up with anything better. I'm not sure this activity will produce any solid definitions, but it is generating a very interesting public discussion.

Tuesday, June 17, 2008

Fedora and DSpace meeting

There was a developer meeting last week with representatives of the DSpace and Fedora communities, and notes from that meeting are available. I am pleased that there was a focus on identifying and solving the needs of the community. I saw a lot of mention of work flows. There was a recognition of the need to get the communities TOGETHER to have discussions about these common needs. There was a brief mention of the upcoming RepoCamp as one such potential venue for discussion.

I also loved the final sentence: "Michele, Sandy, Brad, and Thorny will have a beer together to summarize ideas generated in this meeting and conceive of next steps." That's friendly collaboration at its finest.

Fedora 3.0 Beta 2 released

The second beta release of Fedora 3.0 is now available for testing. This release completes most of the features planned for the general 3.0 release this fall. I'm personally most excited by the new Content Model Architecture (CMA), an integrated structure for persisting and delivering the essential characteristics of digital objects in Fedora, which replaces the previous architecture for binding objects, behaviors and mechanisms. Formal instantiation of content models opens up interesting new avenues for work flows and validation. That this new architecture also allows for easier modification of mechanisms is a huge operational boon.

Read the announcement for more details on the release.

Friday, June 13, 2008

save the date for RepoCamp!

RepoCamp will be a one-day event where folks who are interested in managing and creating digital repository software and their contents can gather and share ideas. RepoCamp will take place on July 25, 2008, at the Library of Congress in Washington D.C. PLEASE NOTE: Due to space constraints, there will only be room for 30-ish people.

For those of you who are familiar with the concept, this will be a "barcamp," where the sessions are proposed and scheduled by the attendees the day of the event. Having participated in beCamp the past couple of years, I'm very excited to see this opportunity for the repository community.

Visit the wiki for more information: http://barcamp.pbwiki.com/RepoCamp

Thursday, May 08, 2008

Kete

I recently become aware of an interesting content management system called Kete, which was developed for the Kete Horowhenua site in New Zealand. It's a repository and discovery service that supports uploading and metadata creation through a web interface. It supports the inclusion of:

  • Images
  • Audio recordings
  • Video recordings
  • Documents
  • URLs for web resources
Metadata can be "locked" so only the creator can edit it, or be open for any to edit. They are collecting some amazing biographical details for their Anzac (veterans) collection through the community. Every "topic" (a subject, a place, a person) can have its own discussion.

Kete Horowhenua was developed with Ruby on Rails, utilizes Zebra z39.50 full text indexing engine developed by IndexData, is fully compatible with Koha, and will be released under a GNU General Public License (GPL). The Kete software is available for download. They are in the process of building a release of the code without the Horowhenua project customizations that can be deployed using a web based wizard that supports customization. They are looking for funding to support this work, as they admit that they underestimated the development needs. They are even accepting PayPal donations to help the work along!

It's an interesting looking site and tool, but I don't know how much is specific to the Horowhenua version and what will be in the generalized version. The browse UI could use a little refinement (I couldn;t figure out how to sort, or if you can sort), but there's a lot of promise here.

Tuesday, April 15, 2008

Blacklight MATC nomination

You can still comment in support of Project Blacklight's nomination for the Mellon Foundation "Mellon Award for Technology Collaboration" (MATC). Anyone who'd like to say something positive about Blacklight, please visit the site and comment:

http://matc.mellon.org/nominate/university-of-virginia/project-blacklight

Saturday, April 12, 2008

Project Blacklight MATC nomination

Project Blacklight is one of the many worthy nominees for the Mellon Foundation "Mellon Award for Technology Collaboration" (MATC). Folks can submit comments in support of nominations, and I am encouraging anyone who'd like to say something positive about Blacklight to please visit the site and comment before 5 PM Eastern Time on Monday April 14.

http://matc.mellon.org/nominate/university-of-virginia/project-blacklight