Showing posts with label digital library services. Show all posts
Showing posts with label digital library services. Show all posts

Tuesday, April 21, 2009

World Digital Library Launch

The World Digital Library is now available.

The site is launching with 1,170 objects from 26 partner institutions. WDL focuses on significant primary materials reflecting the cultural heritage of all UNESCO member countries, including manuscripts, maps, rare books, recordings, films, prints, photographs, architectural drawings, and other types of primary sources from varying time periods. The project will continue to add content to the site, and will enlist new partners from the widest possible range of institutions and countries.

The site is available in seven different languages: Arabic, Chinese, English, French, Russian, Spanish, and Portuguese. The content is not translated -- the items appear in their original language. The metadata and all the site navigation is translated to make it possible to search and browse the site in any of the languages. The metadata came from partner institutions or was created by catalogers at the Library of Congress, and much of the translation was provided by Lingotek.

The site was built using the Django Python framework, nginx, Lucene/Solr, and a mySQL database. The zooming in the imageviewer and pageturner is Seadragon Ajax. There is heavy use of Javascript, jquery, JSON and underlying XML. Check out the image carousels and timeline tool! The project also developed a cataloging tool to manage the metadata and cataloging process and interact with the Lingotek translation system via their API.

Thursday, December 11, 2008

world war II collection at the national archive and footnote

The US National Archives and the historical document website Footnote.com have collaborated on the digitization of a large collection of documents from the US involvement in World War II, which are now available on the footnote.com web site. There is an ars technica article on the collection and interface.

Like the ars technica writer, I had a lot of difficulty finding anything that I hoped to find. My grandfather, father, and uncle all served in WWII. My grandfather died in a friendly fire incident where allied planes accidentally sunk a ship carrying prisoners of war to be returned. I found nothing. There was nothing in the documents nor in the photos. Although I did find out that a man with almost the same name as my uncle (same middle initial but different middle name) was listed as missing when his plane was shot down in 1943. Still, it's a lot of useful content that I'm glad to see digitized and OCR'ed.

I was disappointed I wasn't surprised. I found the navigation to be a bit puzzling. I found I had to have multiple tabs open to easily go back to search. Not just the image but the entire image viewer screen had to come into focus when I selected something to view.

The ars technica writer said that his view of the site included the disclaimer "All Free (for a limited time)," and commented that "... it would be nice to think that a service based on government records of a significant American experience would be free indefinitely." The original press release describing the collaboration is worth reviewing, because it addresses that point in the ars technica article. The agreement allows Footnote.com non-exclusive access, and "After an interval of five years, all images digitized through this agreement will be available at no charge through the National Archives web site." So, Footnote can charge for it for now, but it will all revert to the National Archives for free and open access.

I don't see that disclaimer when using my Library of Congress computer because we have full access -- I wonder how long it will be fully accessible for those without subscriptions?

Tuesday, November 25, 2008

commentary on why google must die

John Dvorak write an essay in PC Magazine entitled "Why Google Must Die." It's a pithy article on search engine optimization (SEO) and the SEO tricks that are in play to work best with Google or get around a Google feature. This is an an essay that I never would have noticed had it not been referenced in a posting by Stephen Abram that I very much took notice of, also entitled "Why Google Must Die."

His post is a response to the often-heard suggestion that OPACs, federated search, and web site search engines should be "just like Google." He asks what should be implemented first:

1. Should I start manipulating the search results of library users based on the needs of advertisers who pay for position?
2. Should I track your users' searches and offer different search results or ads based on their private searches?
3. Should I open library OPACs and searches to 'search engine optimization' (SEO) techniques that allow special interest groups, commercial interests, politicians (as we've certainly seen with the geotagged searches in the US election this year), racist organizations (as in the classic MLK example), or whatever to change results?
4. Should I geotag all searches, using Google Maps, coming from colleges, universities or high schools because I can ultimately charge more for clicks coming from younger searchers? Should I build services like Google Scholar to attract young ad viewers and train or accredit librarians and educators in the use of same?
5. Should I allow the algoritim to override the end-user's Boolean search if it meets an advertiser's goal?
6. "Evil," says Google CEO Eric Schmidt, "is what Sergey says is evil." (Wired). Is that who you want making your personal and institutional values decisions?
There's more to the post. I admire a forthright post like this that pushes back on the assertion that doing things the Google way is automatically better.

I used to have lengthy discussions with a library administrator in a past job who wanted image searching to be just like Google images, because searches on a single word like "horse" would always produce images of horses at the top of the results. It was a lot of effort to explain that this was somewhat artificial, due to the sheer number of images and, that, in the absence of descriptive metadata, that having the string "horse" in the file name would ensure that they would be near the top of the list and that Google didn't actually recognize that it was an image of a horse. Sorry, we really did need to expend effort on descriptive metadata.

Monday, November 24, 2008

europeana a victim of its own success

There are a number of article documenting the wild success and consequent server failure of the Europeana digital library: Times Online, PublicTechnology.net, and Yahoo Tech. The development site that documents the project is still available.

This is a cautionary tale for those of us who are working on the World Digital Library project, which is set to launch in April 2009. We know that there is a potentially high level of interest in a multi-lingual international digital collection site -- albeit one with a much initial smaller collection -- and seeing this confirms for us that making plans for mirroring is a necessity.

Monday, November 17, 2008

facebook repository deposit app using sword

From Stuart Lewis' blog comes word of a Facebook app -- SWORDAPP -- for depositing content into SWORD-enabled repositories. It's meant to encourage social deposit: notice of your deposits goes out in your Facebook newsfeed, and you can receive news of your friend's deposits. You have to already be eligible to authenticate and deposit into a repository somewhere.

This is an interesting use of SWORD in an application that lives in a really different context. He's looking for testers and feedback.

Wednesday, September 03, 2008

HathiTrust

The University of Michigan has announced that their MBooks initiative has grown into a shared repository effort called the HathiTrust (pronounced hah-TEE).

HathiTrust was originally a collaboration of the thirteen universities of the Committee on Institutional Cooperation (CIC) to establish a repository for those universities to archive and share their digitized collections. All content to date has been supplied by the University of Michigan and the University of Wisconsin, and Indiana University and Purdue University will soon be contributing their digital materials. 20% of its current content is open access and 80% is restricted. Don't look for a single search interface yet -- it's planned. As they say: "Good, useful, technology takes time.... and the strength and insight born of collaborative work."

The new HathiTrust initiative has been funded for an initial five-year period beginning January 2008, and is now open to other institutions. Partners will be charged a one-time start-up fee based on the number of volumes added to the repository, in addition to an annual fee for the curation of the data. They already support both open access and dark archive materials, and will also do so for new partners.

Their July 2008 monthly report gives a good sense of their activities. It is interesting to note that the only initial ingest workflow supported is the Google partner workflow. That's not too surprising since this work is based on the MBooks project developed in support of Google content workflows. That in and of itself ensures that there many institutions who'll be considering partnership.

This announcement is exceptionally exciting. I look forward to its development as a service.

Wednesday, August 20, 2008

Ithaka's 2006 Studies of Key Stakeholders in the Digital Transformation in Higher Education

Ithaka has recently released the full findings from their 2006 surveys of the behavior and attitudes of faculty members and academic librarians. The faculty study focuses on the relationship between faculty and the library, faculty perceptions and uses of electronic resources, the transition from print to electronic journals, faculty publishing preferences, e-books, digital repositories, and the preservation of scholarly journals. The librarian survey complements the faculty study, exposing the similarities and differences between faculty and librarian views of key topics.

This is an extended quotes describe two very interesting key perceptual differences:

Over the course of these three surveys, we have tested three “roles” of the library – purchaser, archive and gateway. We have attempted to track how the importance of these three different roles has changed over time. Most highly rated among these roles is that of library as purchaser – faculty don’t want to have to pay for scholarly resources, a finding which holds across disciplines and has remained stable over time. There is slightly more variation by discipline in views on the importance of the library’s preservation function, but valuation of this role is also uniformly high and has remained static over time. The importance of the role of the library as a gateway for locating information, however, varies more widely and has fallen over time.

The declining importance assigned to the gateway role is cause for concern in general, and especially when considered by discipline. The importance to faculty of this role has decreased across all disciplines since 2003, most significantly among scientists. While almost 80% of humanists rate this role as very important, barely over 50% of scientists do so. Beyond the differences between these general disciplinary groups, there also exist substantial variations by individual discipline, as demonstrated by the perceptions of economists. Between 2003 and 2006, the percentage of economists indicating they found the library’s gateway role to be very important dropped almost fifteen percentage points. In 2006, the percentage of economists who believed this gateway role to be very important was actually below the average level of scientists, falling to 48%.

The decreasing importance of this gateway role to faculty is logical, given the increasing prominence of non-library discovery tools such as Google in the last several years. Since 2003, the number of scholars across disciplines who report starting their research at non-library discovery tools, either a general purpose search engine or a specific electronic resource, has increased, and the number who report starting in directly library-related venues, either the library building or the library OPAC, has decreased. Despite the rising popularity of tools like Google, overall, general purpose search engines still slightly trail the OPAC as a starting point for research, and are well behind specific electronic research resources. This overall picture, however, hides a number of variations by discipline; scientists typically prefer non-library resources, while humanists are more enthusiastic users of the library.

The declining importance of this role to faculty stands in stark contrast to the perceptions of librarians, as shown by our 2006 librarian survey. Although the importance of the library’s role as a gateway to faculty is decreasing, rather dramatically in certain fields, over 90% of librarians list this role as very important, and almost as many – only 5 percentage points less – expect it to remain very important in 5 years. Obviously there is a mismatch in perception here.

Librarians at all sizes of institutions see this gateway role as among their primary goals; this, along with the licensing of electronic resources and maintaining a catalog of their resources, are by far the roles most broadly considered important. They expect most of the roles of the library to rise in importance, or at least hold steady, over the next five years, with some notable exceptions to be found in roles focused on nondigital materials, such as roles relating to traditional print preservation and the maintenance of a local print journal collection, which are expected to decline in importance. There are some variations by institution size. Several roles, most notably the development and maintenance of special collections and several more technical tasks such as the management of datasets, are significantly more important at larger libraries than smaller ones. And unlike smaller libraries, larger libraries view licensing as their single most important activity, with less emphasis put on the gateway and catalog roles. This may be a sign that leading-edge libraries are beginning to change their priorities to match those of faculty and students. Still, the mismatch in views on the gateway function is a cause for further reflection: if librarians view this function as critical, but faculty in certain disciplines find it to be declining in importance, how can libraries, individually or collectively, strategically realign the services that support the gateway function?

... and

Perceptions of a decline in dependence are probably unavoidable as services are increasingly provided remotely, and in some ways these shifting faculty attitudes can be viewed as a sign of library success. One can argue that the library is serving faculty well, providing them with a less mediated research workflow and greater ability to perform their work more quickly and effectively. In the process, however, they may be making their own role less visible. This indicates a challenge facing libraries in the near future – as faculty needs are increasingly met without the direct intermediation of the library, the importance of the library decreases. Libraries must consider ways which they can offer new and innovative services to maintain, or in some cases recapture, the attention and support of faculty.

Read their full white paper, or review the raw data from the faculty survey or the librarian study.

Wednesday, July 02, 2008

Mbooks Collections

The BLT announced a new feature in MBooks at Michigan -- Collections Pages.

I love that I can browse collections that others have made public -- Perry has a great Gothic Literature collection going (no public domain Castle of Otranto yet?) -- and that you can view the full collection or limit your view to the full-text items. I also love that I can copy items into my own collections to help populate them. I couldn't find almost any MBooks items in Mirlyn to build my first attempt at a collection. Not too many public domain texts on voodoo are available yet...

Monday, June 23, 2008

Strategies for Sustaining Digital Libraries now in print

I am pleased to announce that the volume Strategies for Sustaining Digital Libraries, edited by Katherine Skinner and Martin Halbert is now in print. I'm pleased because 1), it's a Emory University Digital Library publication, and 2), I have a chapter in the book.

This collection of essays on sustaining digital libraries is a report of early findings from member of the community who have developed ongoing services and collections intended to be sustained over time in ways consistent with the long-held practices of print-based libraries. The essays address a variety of issues in advancing innovative information services to an ongoing programmatic mode of sustaining digital libraries for the long haul.

It's available in print and as a free PDF version, under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 License. For more information, please visit the Emory web site.

Thursday, May 22, 2008

more on oclc and google

Here's an article in Information Today on the OCLC/Google agreement.

This focuses more on the addition of Google Book links into existing WorldCat records, and the creation of new records for volumes not currently in WorldCat. I wonder if this means adding 856 field links to digital surrogates, or creating separate digital resource records? All OCLC member institutions will be able to add records into their catalogs.

This article doesn't mention one aspect of the agreement -- that the agreement now allows Google Book partners to share records from their catalogs that have an OCLC provenance with Google (or let OCLC do it for them). There has always been a lot of discussion about what rights member institutions had vis-a-vis sharing OCLC-sourced records that represent their holdings, and not everyone agrees with OCLC's assertions of its rights. Given Google's need to know something about the volumes that it's digitizing to provide access, it seems unavoidable that sharing of some OCLC-sourced metadata between Google Book participants and Google has already happened. Now it's a recognized need and activity.

Wednesday, May 21, 2008

oclc and google cooperation

On Monday a press release was issued about cooperation between OCLC and Google. Excerpted:

OCLC and Google Inc. have signed an agreement to exchange data that will facilitate the discovery of library collections through Google search services.

Under terms of the agreement, OCLC member libraries participating in the Google Book Search™ program, which makes the full text of more than one million books searchable, may share their WorldCat-derived MARC records with Google to better facilitate discovery of library collections through Google.

Google will link from Google Book Search to WorldCat.org, which will drive traffic to library OPACs and other library services. Google will share data and links to digitized books with OCLC, which will make it possible for OCLC to represent the digitized collections of OCLC member libraries in WorldCat.

...

WorldCat metadata will be made available to Google directly from OCLC or through member libraries participating in the Google Book Search program.

Google recently released an API that provides links to books in Google Book Search using ISBNs, LCCNs and OCLC numbers. This API allows WorldCat.org users to link to some books that Google has scanned through a “Get It” link. The link works both ways. If a user finds a book in Google Book Search, a link can often be tracked back to local libraries through WorldCat.org.

The new agreement enables OCLC to create MARC records describing the Google digitized books from OCLC member libraries and to link to them. These linking arrangements should help drive more traffic to libraries, both online and in person.

There are a couple of big wins here for different communities.

For WorldCat users, there is direct access to Google Book Search volumes. For Google Book users, there is improved access to physical volumes.

For Google Book participant libraries, there are better mechanisms for getting metadata about collections to Google.

Even more importantly -- and I am being hopeful here and reading something into this that may not be there -- this is a potential way to get representation of volumes digitized as part of the Google project not only into WordCat but into the OCLC/DLF Registry of Digital Masters. For folks unfamiliar with that project, it's a registry of digitized volumes -- which meet certain digitization standards and are publicly accessible -- that can be used as a tool by libraries and users to determine if volumes have already been digitized and are available. It's a slowly growing registry where a devoted group of participants have been working to develop the standards for describing digital masters and the work flows for adding records. This is a service that is poised to become essential.

Friday, May 02, 2008

OhioLINK EAD repository and tools

OhioLINK has launched a Finding Aid Creation Tool and Repository for use by any institution in Ohio, which does not require membership in OhioLINK to use.

The OhioLINK Finding Aid Repository takes advantage of XTF in a clean and simple way. It's not clear what repository is behind it. The EAD FACTORy tool is available for use by Ohio institutions for EAD authoring. It doesn't say when or if the code for the tool will be released.

Tuesday, April 15, 2008

OR08 repository case studies

The Repositories Support Project has released two dozen UK, European, and North American digital repository case studies that were prepared for the Open Repositories 2008 conference. I thought it was a great idea when they solicited these for the conference, and I am very happy that my UVA case history is included.

Friday, March 14, 2008

Google Books Viewability API

Google has released a new API that supports links to volumes in Google Book Search. Web developers can use the Books Viewability API to quickly find out a book's viewability on Google Book Search and, in an automated fashion, embed a link to that book in Google Book Search on their own sites.

We'd already created our own service for our Virgo catalog for digitized UVA volumes that are available as full-text in GBS. We were thinking about how we'd potentially port that to Blacklight -- now we can look at the API in addition to what we'd done ourselves.

Monday, March 10, 2008

Texas Digital Library Repository

Via DigitalKoans, the Texas Digital Library Repository has launched with content from four of its partner institutions. There's a program that's making some real headway: an IR for ETDs (multi-institutional, even!), journal hosting, and real progress with inter-institutional Shibboleth – on top of what they’re already doing with digitization and online collections.

Wednesday, December 05, 2007

developing a service vision for a repository

Dorothea rightly challenged me for not including a service vision in my post on repository goals and vision. I do have something like a vision, but I wouldn't say that it's quite where it needs to be yet. That said, I said I would post it, so I am.

What are the services needed around a repository?

  • Identification and acquisition of valuable content
    • You can't wait for content to come to you – research what’s going on in the departments and at the University, and initiate a dialog.
    • Digital collections must also come from the Library and other University units – University Archives, Museums, etc.
  • Consulting Services
    • Advise on intellectual property and contract/licensing issues for scholarly output.
    • Assistance in preparing files for deposit, creating or converting metadata, and in the actual deposit process.
  • Access
    • Easy-to-use discovery interface with full-text searching and browse.
    • Instruction for community on how to find and use and cite content.
    • Make the content shareable via Open Archives Initiative (OAI).
  • Promotion and Marketing
    • Build awareness of the high cost of scholarly journals, and that we are buying back our own institutional scholarship.
    • Promote the value of building sustainable digital collections – preservation is more than just backing up files.
    • Promote the goals of the Open Access movement, including managed, free online access and a focus on improved visibility and impact.
    • Show faculty that they can build personal and community archives.
    • Market repository building services that will enable the institution to build a body of digital content.
    • Market the repository as a content resource and a venue that increases the visibility of the institution.

Tuesday, December 04, 2007

goals and vision for a repository

Last week I had the opportunity to have a lengthy conversation with some folks about our Repository. In doing so I was able to get at some really simplified statements about our activities.

Why a Repository?

  • A growing body of the scholarly communications and research produced in our institutions exists solely in digital form.
  • Valuable assets -- secondary or gray scholarship such as proceedings, white papers, presentations, working papers, and datasets -- are being lost or not reproduced.
  • Numerous online digital collections and databases produced through research activity are not formally managed and are at risk.
  • An institutional repository is needed as a trusted system to permanently archive, steward, and manage access to the intellectual work – both research and teaching – of a university.
  • Open Access, Open Access, Open Access and Preservation, Preservation, Preservation.
What's the vision for a Repository?
  • A new scholarly publishing paradigm: an outlet for the open distribution of scholarly output as part of the open access movement.
  • A trusted digital repository for collections.
  • A cumulative and perpetual archive for an institution.
What does success look like?
  • Improved open access and visibility of digital scholarship and collections.
  • Participation from a variety of units, departments, and disciplines at the institution.
  • Usable process and standards for adding content.
  • Content is actively added.
  • Content is used: searched and cited and downloaded.
  • There is a wide variety of content types.
  • Simple counts are NOT a metric.
I really appreciate having the chance to formulate ideas like these that have nothing to do with the technology but everything to do with why we're doing what we do. I want to work this up into something more formal to share broadly.

Wednesday, October 31, 2007

LibX and OpenURL Referrer browser extensions

We launched our UVA Library LibX plugin for Firefox in June 2007, and its gotten some rave reviews from staff. Now that UVA has approved the rollout of Vista and IE 7 on its computers, we're testing the beta IE version of LibX. I understand we've supplied some feedback on installation and running on Vista.

When I saw the recent announcement of the availability of OCLC's OpenURL Referrer for IE, I paused a bit when considering who to send the annoucement to. LibX is the tool we promote with our users, it recognizes DOIs, ISBNs, ISSNs, and PubMed IDs, and supports COinS and OCLC xISBN, and works with our resolver and our catalog Virgo. We have our resolver working with Google Scholar.

In the end, I didn't forward the annoucement because we're trying to promote the use of LibX and I didn't want to dilute that message for our staff and users. The OpenURL Referrer is a very cool tool and a great use of the OCLC Resolver Registry so users don't have to know anything except the name of their institution to set it up. I'm just not sure if we need both, at least not right now.

I need to ask if we know how much use our LibX toolbar is getting.

Friday, October 12, 2007

what is publishing?

I'm in the process of making a transition in my organization, shifting into a newly created position as Head of Digital Publishing Services.

The first question that everyone asks is "What will you be doing?" The second is "What is publishing in a Library?"

We partner with faculty who are selecting content, organizing it, describing it, analyzing and identifying and creating intellectual relationships, and presenting and interpreting that content in new ways as born-digital scholarship. These projects are more frequently being considered in the promotion and tenure process. Is that the Library supporting a publishing activity? Most certainly.

We digitize collections, describe them, organize them, present them online, and promote them to our community for their use in teaching in research. Is that a publishing activity? It can be argued either way (and has been) -- I'm on the side that leans toward yes.

We provide production support and hosting for peer-reviewed electronic journals. No one would argue that we're participating an a publishing activity.

We're evaluating an Institutional Repository. Not a publishing activity per se, but a way to preserve the publishing output of our community. That's a service related to our stewardship role.

As a Library we're already very active participants in publishing activities. We have our Scholars' Lab, Research Computing Lab, and Digital Media Lab serving as the loci for our faculty collaborations. My role will be formalizing what our publishing services are, what our work flows should be, and how we can sustain and expand our consulting services related to scholarly communication.

Tuesday, September 18, 2007

new york times open access

The story of the day seems to be about the NY Times opening up its archives. So far I've seen postings at boing boing, if:book, open access news, o'reilly radar, and teleread.

So why am I bothering to blog this? Because this made me think about something I blogged about some months ago -- Google News Archive Search. One of the things that galled me at the time was how much of what they indexed was behind a pay firewall. Now, the NY Times is opening almost all their content up (save for 1923-1986), making this a more useful service, at least for resources from one newspaper. If only there wasn't so much other for-fee public domain newspaper content controlled through ProQuest Archiver. I still hope for an OpenURL Resolver service so authorized users can get to authorized resources at ProQuest Historical Newspapers instead.