DCC paper on interoperability
The Digital Curation Centre has released a short briefing paper on interoperability. Its a good, brief primer on the basic issues.
What I'm thinking about digital libraries and other things
The Digital Curation Centre has released a short briefing paper on interoperability. Its a good, brief primer on the basic issues.
Posted by
Leslie Johnston
at
3:58 PM
0
comments
Labels: standards
The latest version of the JHOVE2 Functional Requirements have been posted. I'm still interested in what isn't documented yet, e.g., the final list of formats that will be supported.
Posted by
Leslie Johnston
at
3:53 PM
0
comments
Labels: standards
The National Film Board of Canada (NFB) has opened up its archives - more than 500 films, clips and trailers are now available on their new Screening Room web site. They're freely available for online viewing (there are costs for public broadcast and educational use), with more to be added regularly.
Posted by
Leslie Johnston
at
9:50 PM
0
comments
Labels: collections, video
Steven Levy has written an essay for Wired about the guilt that one can feel for not participating enough in ones social network. Following tweets but not twittering, not blogging often enough, or not updating ones Facebook status. It's a brief but interesting read on privacy and a weird sense of duty to keep those public lines of communication open.
Nicholas Carr has posted a very interesting reaction to Levy's essay.
There's an arrogance to sharing the details of one's life in public with strangers - it's the arrogance of power, the assumption that such details somehow deserve to be broadly aired. And as for the people, those strangers, on the receiving end of the disclosures, they suffer, through their desire to hear the details, to hungrily listen in, a kind of debasement. At the risk of going too far, I'd argue that there's a certain sadomasochistic quality to the exchange (it's a variation on the exchange that takes place between celebrity and fan). And I'm pretty sure that Levy's remorse comes from his realization, conscious or not, that he is, in a very subtle but nonetheless real way, displaying an undeserved and unappetizing arrogance while also contributing to the debasement of others.This seems a bit strong to me, but not entirely off base. Arrogance of power? Debasement? Sadomasochistic? OK, that may be true for some who participate in social networking, just the same as for some participants in a real life communities. There is something a bit egotistical in assuming that others will follow your tweets/blog/delicious tags/flickr set/facebook. There is something a bit creepy that, if you don't require approval, complete strangers read your tweets where you might be discussing where you are at any given time. I like to think that most use social networking to actually keep in touch, not to obsessively stalk one another.
Posted by
Leslie Johnston
at
8:59 PM
2
comments
Labels: digital life, privacy
The Folger Shakespeare Library just expanded access to its Digital Image Collection by offering over 20,000 images online. The collection includes books, theater memorabilia, manuscripts, art, and 218 of the Folger’s pre-1640 quarto editions of the works of William Shakespeare.
Online use is through the Luna Insight Browser -- you have to add an exception to your popup blocker or the software will not function properly. To access their Shakespeare Quartos collection and to get full functionality (saving searches, exporting html pages) you have to install the free Insight Java client.
They have a "how-to" page and search tips available.
Posted by
Leslie Johnston
at
8:34 PM
0
comments
Labels: collections
Last month the Library of Congress had a soft launch of an open source software release. We officially announced the release in the January 2009 issue of the Library of Congress Digital Preservation
Newsletter. This is the first software that the Library has formally released as open source.
The tools are available through SourceForge under the “Library of Congress Transfer Tools” project. The project includes tools for use with BagIt specification, a hierarchical file packaging format for the exchange of digital content jointly developed by the Library of Congress and the California Digital Library.
Three tools developed by the Library's Repository Development Group are available now. Parallel Retriever implements a simple Python-based wrapper around wget and rsync to optimize the transfer of content between locations through parallelization. It supports rsync, HTTP, and FTP transfers. Bag Validator is a Python script that validates a Bag, checking for missing files, extra files, and duplicate files. VerifyIt is a shell script that verifies file checksums within a Bag manifest using parallel processes.
The Library plans to release additional tools as part of a suite of solutions and software development resources as they are completed over time. There are already more tools in the pipeline.
Posted by
Leslie Johnston
at
8:03 PM
0
comments
Labels: LoC, open source, tools
There's a new WorldCat Mobile pilot service.
NYPL has announced its NYPL Mobile beta.
The DC Public Library launched an iPhone app.
Stanford has a new version of an iStanford iPhone app that ties into its student services system.
The International Children's Digital Library launched an iPhone app last November.
Posted by
Leslie Johnston
at
10:55 PM
0
comments
Labels: mobile technology
There was a great Washington Post article yesterday about how White House technology is "in the Dark Ages." I laughed bemusedly over my toast and read the article aloud at the breakfast table.
The White House is not being singled out. I work for a Federal Agency. I know folks who work at numerous other Federal Agencies, some of whom have worked at said agencies for decades. Federal agencies have many, many rules about hardware and software security, and every agency has to interpret and enforce those rules themselves. Security levels of content muddy the waters. This can cause a certain amount of confusion as to what is and isn't allowed. Someone told me that their agency (not the White House) hasn't yet approved Firefox. News that the White House counsel's office approved use of Gmail accounts for some press office activities has been forwarded to many Federal IT units, I'm sure.
Edit, 25 January: Wired has posted a Wired/Tired overview of White House tech, and a list of recent technology projects from various agencies. Nice to see the shout out for the LoC Flickr project.
Posted by
Leslie Johnston
at
10:26 PM
1 comments
Labels: digital life
Ricky Erway from OCLC has distilled the proposed Google Book Settlement, its appendices, and the three library registry agreements from 320 pages to a 4 1/2 page summary. It's an excellent overview of the proposal.
Posted by
Leslie Johnston
at
3:28 PM
0
comments
Labels: intellectual property, texts
My colleague Justin Littman has just published an excellent article in the January/February 2009 issue of D-Lib Magazine: "A Set of Transfer-Related Services."
"The Office of Strategic Initiative's (OSI) Repository Development Team (RDT) is developing a portfolio of services and components to address the challenges posed by scaling transfer processes. While the portfolio is expanding, the focus of this article will be on two core services, the Inventory Service and the Workflow Service. Before proceeding to examine these services, it will be useful to further delineate the transfer problem space. After examining these services, their role in mitigating preservation risks will be considered."
Posted by
Leslie Johnston
at
3:09 PM
0
comments
Labels: digital library systems, LoC, tools
On January 7, 2009, the U.S. House of Representatives approved H.R. 35, the "Presidential Records Act Amendments of 2009," and H.R. 36, the "Presidential Library Donation Reform Act of 2009." These were chosen by the House leadership as the first pieces of substantive legislation passed in 2009 as a symbol of government transparency.
The Presidential Records Act Amendments restores meaningful public access to presidential records by nullifying a 2001 Bush executive order, and the Presidential Library Donation Reform Act requires the disclosure of big donors to presidential libraries. The Senate still has to pass its versions of the bills before they can go to soon-to-be-President Obama to be signed, which he has apparently indicated that he would.
The National Coalition for History provides a good overview of the Records Reform Act. The House Speaker's site provides an overview of both.
Posted by
Leslie Johnston
at
3:09 PM
0
comments
Labels: intellectual property, privacy
Pretty much everyone who knows me knows my loyalty to Palm. I've had one since 1997. I've been syncing it with enterprise calendar systems so I have my personal and work calendars going back to February 1996. I currently have a Centro, and even though I have to manually key in all my work events because we use such an old version of Groupwise that I can't seem to find a sync that works, I am still devoted to my Palm.
I cannot count the number of friends and colleagues who have iPhones and have done their best to convert me. The Urbanspoon app and its clever use of the accelerometer to randomize recommendations by shaking the phone almost had me. My answer is always that if they can promise me that I can port over everything I have on my Centro -- 13 years of calendar, hundreds of contacts, notes, to-do lists, and ebooks -- then I'll consider it.
I am now waiting with baited breath for the Palm Pre. It's such a step forward in the interface (a card stack metaphor) and operating system (its WebOS is Linux) and browser (based on WebKit). It doesn't use the Palm desktop anymore, which while be a paradigm switch for me, but they have promised data migration tools for Centro users.
I know that some apps I use won't work anymore, and Raymond's concern about no backwards compatibility of the OS is a real one. But they need to move on and I need to move on, and I'm glad it looks like it will be to another Palm. You can't always be fully backwards compatible. Hey, I only complained a little when I discovered I couldn't run FileMake Mobile on my Centro, didn't I?
Reviews at Gizmodo, ars technica, PC World, and a PC World FAQ.
Posted by
Leslie Johnston
at
11:40 AM
0
comments
Labels: digital life, mobile technology
I just read an article at Information Today by Nicholas Tomaiuolo, an instruction librarian at Central Connecticut State University, entitled "U-Content: Project Gutenberg, Me, and You." He outlines the requirements and steps for preparing an etext for Project Gutenberg.
At one point in the article, there is a discussion about the requirements for full text, not just a PDF created from page images. The author wrote this from the point of one unfamiliar with PG's requirements, illustrating the process one might follow to create an acceptable PG submission -- images to PDF, and images to OCR to corrected plain text -- I found myself thinking quite a bit about the often heard statement (not in this article, mind you) that PDF is the ultimate format for texts.
I'm in no way denigrating PDF. PDFs is an absolutely required format for texts. PDF is highly portable and shareable and readable, and, if the source files are good enough, clearly printable. But it's not innately analyzable or easily repurposed. That requires full text.
I am not unfamiliar with what it takes to create an accurate plain text transcription of a text. When Gutenberg was in its early days, we were really talking about transcriptions, as in people typing in text. OCR has greatly streamlined that process, but the proofreading required is non-trivial. Want to work with a highly formatted text, or one with tables or formulae or figures? Challenging. Adding layers of structural and semantic markup to plain text, as with TEI, is time consuming. Rich markup, including identifying dates or names or geographical places, or providing normalized versions of said dates and names is a large undertaking. A full text with structural and sematic markup can be repurposed into many formats, including ebooks and PDF.
And you do want ebooks. Some months ago I had the great opportunity to demonstrate the prototype World Digital Library site at the National Book Festival. There is no greater focus group than thousands of people who love to read! The two top requests were that the books should be downloadable as ebooks and that all the text content be available as full text in all seven project languages. These were not academics or librarians (although there were some of the former and many of the latter who stopped by), but parents and commuters and researchers and genealogists.
Both are daunting requests when you do not have full text available to work from. There will be PDFs. The others are goals to strive for.
Posted by
Leslie Johnston
at
6:25 PM
0
comments
Labels: texts
I am not part of the Library of Congress Flickr project team, and I in no way speak for them.
There is a lengthy discussion on Wired and at Found History about The Commons on Flickr, given that Yahoo laid off key staff member George Oates who shepherded the project. It doesn't seem that the project is in any imminent danger, but a dedicated group has stepped forward to evangelize, innovate, and curate thematic sets from the across the collection. I'm thrilled to see a community of use developing around The Commons, but it's sad that this is what precipitated its full coalescence.
Posted by
Leslie Johnston
at
11:49 AM
1 comments
Labels: images
There's still time to send a message to Chris Rusbridge at the Digital Curation Center to enter his personal data recovery challenge:
I will do my best to recover the first half dozen interesting files that I’m told about… of course, what I really mean is that I’ll try and get the community to help recover the data. That’s you!His deadline is Twelfth Night, January 5, 2009. Read his post for the rest of the details.
OK, I define interesting, and it won’t necessarily be clear in advance. The first one of a kind might be interesting, the second one would not. Data from some application common of its time may be more interesting than something hand-coded by you, for you. Data might be more interesting (to me) than text. Something quite simple locked onto strange obsolete media might be interesting, but then again it might be so intractable it stops being interesting. We may even pay someone to extract files from your media, if it’s sufficiently interesting (and if we can find someone equipped to do it).
The only reference for this sort of activity that I know of is (Ross & Gow, 1999, see below), commissioned by the Digital Archiving Working Group.
What about the small print? Well, this is a bit of fun with a learning outcome, but I can’t accept liability for what happens. You have to send me your data, of course, and you are going to have to accept the risk that it might all go wrong. If it’s your only copy, and you don’t (or can’t) take a copy, it might get lost or destroyed in the process. You’ll need to accept that risk; if you don't like it, don't send it. I might not be able to recover anything at all, for many reasons. I’ll send you back any data I can recover, but can’t guarantee to send back any media.
The point of this is to tell the stories of recovering data, so don’t send me anything if you don’t want the story told. I don’t mind keeping your identity private (in fact good practice says that’s the default, although I will ask you if you mind being identified). You can ask for your data to be kept private, but if possible I’d like the right to publish extracts of the data, to illustrate the story.
Posted by
Leslie Johnston
at
11:27 AM
0
comments
Labels: preservation
I am definitely a proponent of the self-made. Cooking, clothing, jewelry, assisted forays into small electronic projects, etc. At various times I have designed and made costumes, painted, made paper, etched prints, made lampwork glass beads, and learned the basics of Chinese brush painting. I haven't had as much time for that in the last couple of years, but that doesn't mean that I don't still strive to make.
I'm a big fan of Make, and so are a lot of others. In that vein, William Turkel has posted his list of books for humanist makers.
Posted by
Leslie Johnston
at
11:18 AM
1 comments
Labels: digital life
Siva Vaidhyanathan has posted his interview with Vint Cerf about developments in search technology and Google.
Posted by
Leslie Johnston
at
10:54 AM
0
comments
Labels: digital life, search
I don't know if these are resolutions or goals ...
I have been getting back to writing in the past few weeks, but I will write more about what we're working on and getting the word out.
I will be a more hard-nosed project manager vis-a-vis deadlines. I am too often too understanding of delays.
I will explore Washington D.C. more. I lived in metro D.C. for part of my childhood and have been in Virginia for 7 years, but still have a disconnected sense of the city. Perhaps that's in part an artifact of traveling by Metro and not having a real sense of the city's topography.
I will meet more people in other divisions at LoC. This is harder than it sounds.
I will eat lunch at my desk less often. I will make more of an active effort to organize social opportunities. (Why don't I remember this every time I think about switching jobs -- it gets harder and harder to build a new social network every time.)
I have faith that our house in Charlottesville will finally sell.
EDIT: There is another long term goal that I'm already working on: I've signed up for a Japanese class through the Federal employee graduate school. This is the year to begin transforming my ragtag knowledge of Japanese into usable language skills.
Posted by
Leslie Johnston
at
2:21 PM
0
comments
Labels: digital life
There is a great article on ars technica today about the major processing effort that will be required at the National Archives when the Bush administration leaves office. The ars technica piece references a New York Times article on the topic from this past weekend.
This section really strikes home:
The contingency plan will entail "ingesting" the Bush White House's data into a separate system before integrating it with the ordinary archive. As the plan explains, "the current PERL [Presidential Electronic Records Library] system architecture was not scalable to actually support the volume of records that are expected from the current Presidential administration."First, the use of quotation marks should remind us all that "ingest" means absolutely nothing to someone who is not a repository manager.It's not just size that matters, though: the Archives will also need to process reams of information locked in some quaint proprietary formats. The RMS index, for example, "consists of an implementation of a customized older version of Documentum running on Oracle, with image files (including copies of scanned records) incorporated as objects in the database." The photos are stored in a "proprietary photo management software called MerlinOne, running on Microsoft SQL as the database engine," and it has apparently taken several months to extract the images and metadata for relinkage outside the Merlin format.
Posted by
Leslie Johnston
at
11:00 AM
0
comments
Labels: preservation
Between a month almost solely dedicated to a single high-stress project and a lot of other writing commitments -- revising a paper for a conference, drafting a conference proposal and a co-authored conference proposal, and a writing chapter for a book -- I find that I haven't made time to blog. I promise to make time soon.
Posted by
Leslie Johnston
at
2:42 PM
0
comments
Labels: blogging
Recently I spent 4 weeks on a project where we were considering hardware options for a large amount of storage for a data migration project. We ended up with 4 different proposals -- three from vendors and one to be built in-house. One of the tasks that I worked on was a matrix to compare the 4 potential solutions.
There were the easy metrics -- the amount of raw and usable storage, number of racks/tiles required, electrical and cooling requirements, cost, etc. Comparing supportability was trickier but doable, with 24*7 versus 12*5 phone support, availability of on-site technicians, warranty terms, support contract costs, etc. Where it became more difficult was identifying metrics to compare performance. Ratio of processors to storage? Location of processing nodes in the architecture? I/O rates? Time to read all data? And how do you best calculate those last two with four quite architecturally different proposals? We ended up with metrics that not everyone agreed upon, in part because there was a requirement that not everyone agreed upon.
I'm curious how other folks have gone about doing this. I'd be interested in hearing from anyone who is willing to share their strategies.
Posted by
Leslie Johnston
at
12:12 PM
1 comments
Labels: hardware, project management
From an article on Forbes.com, the International Children's Digital Library (ICDL) announced a partnership with the Taliaferro Family Fund to increase the number of European children's titles in the collection. The Elias Project will target three collections in
After reading the article, I check in at the ICDL site, which I hadn't visited in a few months, and noticed two other news announcements: ICDL and the Google Book project will be sharing public domain children's book titles; and ICDL has launched an iPhone app with full access to the collection, a new titles features, and an offline mode and an airplane mode. It's great to see such a worthwhile project making such advances in collection building and in adding new services.
(I didn't see a press release about the European project on the ICDL site. I saw the press release on some other sites, so I assume it's meant to be out there.)
Posted by
Leslie Johnston
at
9:45 PM
0
comments
Labels: collections, digitization
Nik Honeysett has posted a great letter to Santa on the Musematic blog.
If enough of us ask for that image format, will Santa grant our wish?
Posted by
Leslie Johnston
at
3:29 PM
0
comments
Labels: humor
Paul LeClerc, director of the New York Public Library, answered questions online at the New York Times that have been made available in three parts: part one, part two, part three. Topics include budgets, branch closures and renovations, ebooks, and preservation efforts.
In part two, he briefly mentioned their participation in the Library of Congress National Digital Newspaper Project and its public access content web site Chronicling America. It was nice to see this project mentioned get a media mention in the context of preserving and providing access to often ephemeral newspapers.
Posted by
Leslie Johnston
at
2:26 PM
0
comments
Labels: libraries
The Library of Congress has released its report on its Flickr Commons pilot, where approximately 5,000 images were uploaded for a crowdsourcing metadata experiment. A full report and a summary report are available, both PDFs.
The photos have drawn more than 10 million views, 7,166 comments and more than 67,000 tags. When Flickr commenters provide updated place and personal names, dates, and event identification, staff from the Library's Prints and Photographs Division verify the information and have so far updated more than 500 records in their catalog -- with many more in the queue -- citing the Flickr Commons Project as the source of the new information.
Posted by
Leslie Johnston
at
11:16 AM
1 comments
Labels: collections, images, library 2.0, LoC
Creative Commons is conducting a study to collect feedback on the term “noncommercial” and how it should be covered in its licenses. The hope is that what’s learned from the survey can improve the licenses that allow or restrict noncommercial uses. The questionnaire has to be completed by this Sunday, December 14, 2008. Everyone who has taken advantage of CC licenses as a creator or a user should take some time to answer the questions.
Posted by
Leslie Johnston
at
5:16 PM
0
comments
Labels: intellectual property
The US National Archives and the historical document website Footnote.com have collaborated on the digitization of a large collection of documents from the US involvement in World War II, which are now available on the footnote.com web site. There is an ars technica article on the collection and interface.
Like the ars technica writer, I had a lot of difficulty finding anything that I hoped to find. My grandfather, father, and uncle all served in WWII. My grandfather died in a friendly fire incident where allied planes accidentally sunk a ship carrying prisoners of war to be returned. I found nothing. There was nothing in the documents nor in the photos. Although I did find out that a man with almost the same name as my uncle (same middle initial but different middle name) was listed as missing when his plane was shot down in 1943. Still, it's a lot of useful content that I'm glad to see digitized and OCR'ed.
I was disappointed I wasn't surprised. I found the navigation to be a bit puzzling. I found I had to have multiple tabs open to easily go back to search. Not just the image but the entire image viewer screen had to come into focus when I selected something to view.
The ars technica writer said that his view of the site included the disclaimer "All Free (for a limited time)," and commented that "... it would be nice to think that a service based on government records of a significant American experience would be free indefinitely." The original press release describing the collaboration is worth reviewing, because it addresses that point in the ars technica article. The agreement allows Footnote.com non-exclusive access, and "After an interval of five years, all images digitized through this agreement will be available at no charge through the National Archives web site." So, Footnote can charge for it for now, but it will all revert to the National Archives for free and open access.
I don't see that disclaimer when using my Library of Congress computer because we have full access -- I wonder how long it will be fully accessible for those without subscriptions?
Posted by
Leslie Johnston
at
3:29 PM
0
comments
Labels: digital library services
I just saw the press release naming Laine Farley as the new Director of the California Digital Library. I am thrilled for Laine, who's been serving as the interim Director for over 2 years. I have worked with her on Aquifer and at least one other collaborative community project, and I know what an experienced and capable person she is.
Posted by
Leslie Johnston
at
8:32 PM
0
comments
In recent weeks I've been in a number of meetings about storage architecture. In comparing potential solutions for a particular project's needs, there has been a lot of discussion about requirements for space, cooling, power, etc., which are some of the metrics for comparison of the proposals. Today I came across a posting by James Hamilton from Microsoft's Data Center Futures Team, on misconceptions about the cost of power in large-scale data centers.
There is an error in the table for the two amortization entries: the 180-month amortization of the facility says 3 years in the notes column when it should say 15, and the 36-month amortization of the servers says 15 years when it should say 3. The right numbers are used in the calculation -- it's just the explanatory notes that are switched.
There are some very interesting calculations that I plan to forward to other folks. One commenter points out that the figures and graph don't take staffing into account, which, according to Hamilton, is because that cost is a very small percentage of the cost. At Microsoft's scale that's probably true. I guess we're a relatively small- to medium-scale data center, because we are definitely taking that into consideration. And, while the cost of power may be a lower percentage of overall cost than previously considered, power capacity is still definitely a factor when you're considering adding a number of servers and switches.
Posted by
Leslie Johnston
at
10:10 AM
0
comments
Labels: digital library systems
Jamie Boyle's book The Public Domain: Enclosing the Commons of the Mind has been published by Yale University Press, and is also available for free download under a Creative Commons license.
I've seen Jamie Boyle speak two or three times, and I consider him a very important voice in the discussions on the public domain, intellectual property, patents, the economics of same, and their place in technology and culture.
Read this book.
Posted by
Leslie Johnston
at
10:27 AM
0
comments
Labels: intellectual property
The release of LuSql has been announced on a few email lists:
LuSql is a high-performance, simple tool for indexing data held in a DBMS into a Lucene index. It can use any JDBC-aware SQL database.
It includes a tutorial with a series of increasingly complex use cases, showing how article metadata held in a series of MySql tables can be indexed and how file system files containing full-text can also be indexed.
It has been tested extensively, including using 6.4 million metadata and full-text records to produce a 86GB index in 13.5 hours.
It is licensed with the Apache 2.0 license.
Posted by
Leslie Johnston
at
1:13 PM
0
comments
Labels: search
John Dvorak write an essay in PC Magazine entitled "Why Google Must Die." It's a pithy article on search engine optimization (SEO) and the SEO tricks that are in play to work best with Google or get around a Google feature. This is an an essay that I never would have noticed had it not been referenced in a posting by Stephen Abram that I very much took notice of, also entitled "Why Google Must Die."
His post is a response to the often-heard suggestion that OPACs, federated search, and web site search engines should be "just like Google." He asks what should be implemented first:
1. Should I start manipulating the search results of library users based on the needs of advertisers who pay for position?There's more to the post. I admire a forthright post like this that pushes back on the assertion that doing things the Google way is automatically better.
2. Should I track your users' searches and offer different search results or ads based on their private searches?
3. Should I open library OPACs and searches to 'search engine optimization' (SEO) techniques that allow special interest groups, commercial interests, politicians (as we've certainly seen with the geotagged searches in the US election this year), racist organizations (as in the classic MLK example), or whatever to change results?
4. Should I geotag all searches, using Google Maps, coming from colleges, universities or high schools because I can ultimately charge more for clicks coming from younger searchers? Should I build services like Google Scholar to attract young ad viewers and train or accredit librarians and educators in the use of same?
5. Should I allow the algoritim to override the end-user's Boolean search if it meets an advertiser's goal?
6. "Evil," says Google CEO Eric Schmidt, "is what Sergey says is evil." (Wired). Is that who you want making your personal and institutional values decisions?
Posted by
Leslie Johnston
at
2:45 PM
2
comments
Labels: digital library services, search
The Library of Congress is hosting the SearchCampDC barcamp on Tuesday, December 2. From the wiki:
The goal of SearchCampDC is to bring together people working on and using search releated technologies in and around Washington DC. We all rely on tools like Google in our daily lives, but as IT professionals we often need to build or integrate search technologies into applications, and the enterprise.Search technologies typically rely on a wide variety of techniques such as fulltext indexing, geocoding, named entity recognition, data mining, natural language parsing, machine learning, distributed processing. Perhaps you've used a piece of technology, and care share how well it worked? Or perhaps you help develop a search-related tool that you'd like to get feedback on? Or maybe you've got a search-itch to scratch, and want to learn more about how to scratch it? If so, please join us for SearchCampDC.
The idea for this event came about from a happy coincidence that several Lucene, Solr, OpenLayers and TileCache developers were going to be in the DC area at the same time. The idea is to provide them (and hopefully others like you) a space for short & sweet presentations about stuff they are working on, and also to provide a collaborative space for people to try things out, share challenges, ideas etc.
Posted by
Leslie Johnston
at
10:34 AM
0
comments
Labels: conferences, search, searchcampdc
There are a number of article documenting the wild success and consequent server failure of the Europeana digital library: Times Online, PublicTechnology.net, and Yahoo Tech. The development site that documents the project is still available.
This is a cautionary tale for those of us who are working on the World Digital Library project, which is set to launch in April 2009. We know that there is a potentially high level of interest in a multi-lingual international digital collection site -- albeit one with a much initial smaller collection -- and seeing this confirms for us that making plans for mirroring is a necessity.
Posted by
Leslie Johnston
at
4:52 PM
0
comments
Labels: digital library services
The prototype site for Europeana, the European digital library funded by the EC, is set to launch today, November 20, 2008.
The initial collection of 2 million items comes from museums, libraries, archives, and audio-visual collections, and includes paintings, maps, videos and newspapers. The interface is in French, English, and German, with more languages planned. Highlights include the Magna Carta from Britain, the Vermeer painting “Girl With a Pearl Earring” from the Mauritshuis Museum in The Hague, a copy of Dante’s “Divine Comedy'" and newsreel footage of the fall of the Berlin Wall. Content is free of copyright restrictions.
New York Times and EU Business articles provide more details. The goal is to add 10 million items into the digital library by 2010, at a price of 350 million to 400 million euros.
Posted by
Leslie Johnston
at
7:26 AM
0
comments
Labels: collections, digitization
Google has launched a hosted collection of newly-digitized images includes photos and etchings produced and owned by LIFE Magazine.
Apparently only a small percentage of these images have been published; the remainder come from their photo archive. Google is digitizing them: 20 percent of the collection is online, and hey are working toward he goal of having all 10 million photos online.
There are some great Civil War, WWI, WWII and Vietnam war images, portraits of Queen Victoria and Czar Nicholas, civil rights-era documentation, and early images of Disneyland. There are photographs by Mathew Brady, Alexander Gardner, Margaret Bourke-White, Dorothea Lange, Alfred Eisenstaedt, Carl Mydans, and Larry Burrows, among others.
Posted by
Leslie Johnston
at
3:28 PM
0
comments
Labels: collections, images
I love the idea of a flash mob volunteer effort to catalog the book and videotape collection at St. John's Church in Beverly Farms, Massachusetts. There's more from one of the volunteers.
I'd like to see this effort replicated at small historical societies, house museums, and other organizations that are often supported only by a small cadre of volunteers.
Posted by
Leslie Johnston
at
10:13 AM
0
comments
Labels: digital life, metadata
From Stuart Lewis' blog comes word of a Facebook app -- SWORDAPP -- for depositing content into SWORD-enabled repositories. It's meant to encourage social deposit: notice of your deposits goes out in your Facebook newsfeed, and you can receive news of your friend's deposits. You have to already be eligible to authenticate and deposit into a repository somewhere.
This is an interesting use of SWORD in an application that lives in a really different context. He's looking for testers and feedback.
Posted by
Leslie Johnston
at
5:49 PM
0
comments
Labels: collections, digital library services
The Association of Research Libraries has created "A Guide for the Perplexed: Libraries & the Google Library Project Settlement," a 23-page document intended to help libraries understand the impact of the proposed Google Book Search settlement.
Posted by
Leslie Johnston
at
10:58 PM
0
comments
Labels: digitization, intellectual property, publishing
My brief, unstructured notes from a presentation by Sally Rumsey from Oxford University on the Preserv2 project, at the DLF Fall 2008 Forum.
Posted by
Leslie Johnston
at
1:44 PM
0
comments
Labels: digital library systems, preservation
My somewhat unstructured notes from a presentation by Roberta Fox from Harvard, at at the DLF Fall 2008 Forum.
Posted by
Leslie Johnston
at
1:35 PM
0
comments
Labels: collections, search
My somewhat unstructured notes from a presentation by Tito Sierra and Jason Casden from NCSU on the Course Views tools, at at the DLF Fall 2008 Forum.
Posted by
Leslie Johnston
at
1:26 PM
2
comments
My brief, unstructured notes from a presentation by Ryan Chute from Los Alamos National Labs on the Djatoka JPEG 2000 Image Server, at at the DLF Fall 2008 Forum.
Posted by
Leslie Johnston
at
1:21 PM
0
comments
My somewhat unstructured notes from a presentation by Peter Keane from UT Austin on the use of Atom and Atom/Pub in their DASe repository, at at the DLF Fall 2008 Forum.
Posted by
Leslie Johnston
at
1:12 PM
1 comments
Labels: digital library systems
My somewhat unstructured notes from a presentation by Virginia Rutledge, an attorney from Creative Commons, at the DLF Fall 2008 Forum.
Posted by
Leslie Johnston
at
12:43 PM
0
comments
Labels: intellectual property
I was going to spend some time transforming my notes from Dan Clancy's session on Google Book Search from the DLF Fall 2008 Forum into more coherent prose, but for the sake of timeliness, I'm going to post them as is.
Posted by
Leslie Johnston
at
12:24 PM
0
comments
Labels: digitization, intellectual property, publishing
Omeka v 0.10 has been released. Omeka 0.10b incorporates many requested changes: an unqualified Dublin Core metadata schema and fully extensible element sets to accommodate interoperability with digital repository software and collections management systems; elegant reworkings of the theme and plugin APIs to make add-on development more intuitive and more powerful; a new, even more user friendly look for the administrative interface; and a new and improved Exhibit Builder.
Posted by
Leslie Johnston
at
4:04 PM
0
comments
Labels: digital library systems
I came across a link to a flickr set that someone has created with covers and illustrations from Scholastic Book Services books from the 1960s and 1970s. I _loved_ it when the Scholastic book order forms were distributed, and I always ordered something like a dozen books every time. This set includes a number of books that I know I owned, and even a very few that I _still_ own. This person has a great collection.
EDIT: Here's another flickr set.
Posted by
Leslie Johnston
at
3:33 PM
1 comments
Greenstone v2.81 has been released. Improvements include handling filenames that include non-ASCII characters, accent folding switched on by default for Lucene, and character based segmentation for CJK languages. There are many other significant additions, including the Fedora Librarian Interface (analogous to GLI, but working with a Fedora repository).
Posted by
Leslie Johnston
at
2:56 PM
0
comments
Labels: digital library systems
BBC News reported on a release of a collaboration between Google Earth and the Rome Reborn project. Ancient Rome is the first historical city to be added to Google Earth. The model contains more than 6,700 buildings, with more than 250 place marks linking to key sites in a variety of languages.
Posted by
Leslie Johnston
at
2:47 PM
1 comments
Labels: visualization
Johns Hopkins University and the Bibliothèque nationale de France have announced that the Roman de la Rose Digital Library available at http://romandelarose.org/. The goal is to bring together digital surrogates of all the approximately 270 extant manuscript copies of the Roman de la Rose. By the end of 2009 they expect to have 150 versions included in the resource. There is an associated blog available at http://romandelarose.blogspot.com/.
I am particularly interested in the pageturner and image browser that they used -- the FSI Viewer, a Flash-based tool. It seems to work with TIF, JPG, FPX, and PDF (but not JPEG2000?), and converts files to multi-resolution TIFs. It's a very intuitive interface.
Posted by
Leslie Johnston
at
7:15 AM
0
comments
Labels: collections, texts, tools
Via BoingBoing, the UI for Photoshop recreated with real objects, created by the agency Bates 141 in Jakarta for Software Asli. Follow the links to the image and to the "making of" flickr set.
Posted by
Leslie Johnston
at
11:27 AM
0
comments
Labels: humor
The U.S. District Court for the Middle District of Pennsylvania has issued an opinion in the case United States v. Crist that a hash value analysis in a criminal investigation counts as a Fourth Amendment "search." Read a synopsis at ars technica.
Posted by
Leslie Johnston
at
1:49 PM
0
comments
Labels: privacy
JISC has released a two-part study of digital preservation policies: Digital Preservation Policies Study and Digital Preservation Policies Study, Part 2: Appendices—Mappings of Core University Strategies and Analysis of Their Links to Digital Preservation. The study aims to provide an outline model for digital preservation policies and to analyse the role that digital preservation can play in supporting and delivering key strategies for higher ed institutions.
Posted by
Leslie Johnston
at
1:43 PM
0
comments
Labels: preservation
An interesting new book -- The Tower and The Cloud: Higher Education in the Age of Cloud Computing -- has been published by Educause. The term "cloud computing" is usually used to refer to applications that run on remote systems in "the cloud" rather than on desktop computers or to the storage of files remotely rather than locally, but the book defines the term more broadly, including open-source software and social-networking tools. The full book is available online as a free PDF.
Posted by
Leslie Johnston
at
1:35 PM
0
comments
Labels: cyberinfrastructure
Thanks to Amanda for pointing this out -- I am addictively following the twitter production of War of the Worlds, an homage to the Orson Welles radio production. How this came about is described at the Ask a Wizard blog.
Posted by
Leslie Johnston
at
1:04 PM
0
comments
Labels: digital life, mobile technology
Today it was announced that Google has reached a settlement in the lawsuit filed by the Authors Guild, the Association of American Publisher, and a group of individual authors.
Some of the details are available at Google. The changes that I am the most interested in are these:
"Until now, we've only been able to show a few snippets of text for most of the in-copyright books we've scanned through our Library Project. Since the vast majority of these books are out of print, to actually read them you'd have to hunt them down at a library or a used bookstore. This agreement will allow us to make many of these out-of-print books available for preview, reading and purchase in the U.S.. Helping to ensure the ongoing accessibility of out-of-print books is one of the primary reasons we began this project in the first place, and we couldn't be happier that we and our author, library and publishing partners will now be able to protect mankind's cultural history in this manner."I'm all for more access to these books and for rightsholders to get their due, but what does it mean to assign a value to them?
...
"The agreement will also create an independent, not-for-profit Book Rights Registry to represent authors, publishers and other rightsholders. In essence, the Registry will help locate rightsholders and ensure that they receive the money their works earn under this agreement. You can visit the settlement administration site, the Authors Guild or the AAP to learn more about this important initiative."
Posted by
Leslie Johnston
at
9:40 AM
0
comments
Labels: digitization, intellectual property
Some argue that search engines such are copyright violators because they scrawl, index and keep an archive of web sites. That copied archive -- or cache -- is, according to this argument, an unauthorized copy. Found via TechDirt, the Pennsylvania Eastern District Court held that a Web site operator's failure to deploy a robots.txt file containing instructions not to copy and cache Web site content gave rise to an implied license to index that site.
In Parker v. Yahoo!, Inc., 2008 U.S. Dist. LEXIS 74512 (E.D. Pa. Sep. 26, 2008), the court found that the plaintiff's acknowledgment that he deliberately chose not to deploy a robots.txt file on the site containing his work was conclusive on the issue of implied license. In so ruling the court followed Field v. Google, a similar copyright infringement action brought by an author who failed to deploy a robots.txt file and whose works were copied and cached by the Google search engine.
The court further ruled, though, that a nonexclusive implied license may be terminated. Parker may have terminated the implied license by the institution of the litigation, and he alleged that the search engines failed to remove copies of his works from their cache even after the litigation was instituted. If proved, "the continued use over Parker's objection might constitute direct infringement." That issue will likely be resolved at a later date.
For an analysis, see the New Media and Technology Law Blog.
The same plaintiff's earlier Parker v. Google, Inc., No. 06-3074 (3d Cir. July 10, 2007) is also a search engine copyright infringement case.
Posted by
Leslie Johnston
at
10:25 AM
0
comments
This struck me as hilarious -- Someone noted that a box of Cascadian Farms frozen broccoli had teeny, tiny faces worked into the image on the label:
http://bread-and-honey.blogspot.com/2008/10/wtf-broccoli.html
A comment in another blog said this (unsubstantiated):
"They've been putting tiny faces of employees, family and friends on the labels since at least 1995, which was when someone first showed me this on the labels of Cascadian Farms jams when I was first working at Fresh Fields (later bought by Whole Foods). CF has since been bought by General Mills, but it seems the tiny faces continue."
A bulletin board thread from 2007 claimed that the creamed corn packaging had a little hidden baby's face. Strange. That thread described it as a version of an "easter egg" in a video game or DVD ... an undocumented feature that you have to really try to find. That's pretty apt.
Posted by
Leslie Johnston
at
4:17 PM
0
comments
Labels: humor
On a long drive recently, my partner Bruce and I were reminiscing about places we used to eat at in Los Angeles. He grew up there and has a longer list than I (maybe for another post). So many places we used to patronize as recently as the early 1990s are now gone, or, as the L.A. Time Machines site puts it, "extinct." I decided that I would try to write down the places I remember frequenting that are no longer open. Then I started semi-obsessively researching them.
Posted by
Leslie Johnston
at
5:54 PM
6
comments
Jonathan Rochkind has posted a great description of digital book access features that he's put into production in the link resolver and OPAC at Johns Hopkins. They're remarkable in the sense that he's taken advantage of so many different service APIs (Google Books, IA, OCLC, Amazon, HathiTrust) to provide functionality with conditional options to provide as much collection coverage as possible.
Posted by
Leslie Johnston
at
11:12 AM
0
comments
Labels: digital library systems, search, texts
I've just read an interesting paper from a presentation at the recent CIDOC meeting: Nicholas Crofts, “Digital Assets and Digital Burdens: Obstacles to the Dream of Universal Access,” 2008 Annual Conference of CIDOC (Athens, September 15-18, 2008).
The premise is that technology is not the issue keeping our institutions from reaching a goal of universal access -- it's a number of post-technical issues, including varied intellectual property barriers, institutions' desires to protect their digital assets, and collection documentation that is not well-suited to sharing.
From the section on "Suitability of Documentation":
... but while this technical revolution has taken place, there has not been a corresponding revolution in documentation practice. The way that documentation is prepared and maintained and the sort of documentation that is produced are still heavily influenced by pre-Internet assumptions. The documentation found in museums – the raw material for diffusion – is often ill-suited for publication.From the conclusion:
While making cultural material freely available is part of their mission, and therefore a goal that they are obliged to support, it may still come into conflict with other factors, notably commercial interests: the need to maintain a high-profile and to protect an effective brand image. If museums are to cooperate successfully and make digital resources widely available on collaborative platforms, they will either need to find ways of avoiding institutional anonymity, or agree to put aside their institutional identity to one side.It's a frank and interesting paper. I think there has been progress in documentation practice -- look at the CCO and the Aquifer Shareable Metadata efforts, and the earlier Categories for the Description of Works of Art -- but it's true that this hasn't yet taken hold in a widespread way.
Posted by
Leslie Johnston
at
9:35 AM
0
comments
The newest issue of First Monday (volume 13, number 10, 6 October 2008) has an interesting article by KalevLeetaru -- "Mass book digitization: The deeper story of Google Books and the Open Content Alliance."
The article compares what is publicly known about the Google Book and OCA projects.
From the conclusions:
From the section on "Transparency"While on their surface, the Google Books and Open Content Alliance projects may appear very different, they in fact share many similarities:
Both operate as a black box outsourcing agent. The participating library transports books to the facility to be scanned and fetches them when they are done. The library provides or assists with housing for the facility, but its personnel are not permitted to operate the scanning units, which must be staffed by personnel from either Google or OCA.
Neither publishes official technical reports. Google engineers have published in the literature on specific components of their project, which offer crucial insights into the processes they use, while talks from senior leadership have yielded additional information. OCA has largely been absent from the literature and few speeches have unveiled substantial technical details. Both projects have chosen not to issue exhaustive technical reports outlining their infrastructure: Google due to trade secret concerns and OCA due to a lack of available time.
Both digitize in–copyright works. Google Books scans both out–of–copyright books and those for which copyright protection is still in force. OCA scans out–of–copyright books and only scans in–copyright books when permission has been secured to do so. Both initiatives maintain partnerships with publishers to acquire substantial in–copyright digital content.
Both use manual page turning and digital camera capture. Large teams of humans are used to manually turn pages in front of a pair of digital cameras that snap color photographs of the pages.
Both permit libraries to redistribute materials digitized from their collections. While redistribution rights vary for other entities, both the Google Books and OCA initiatives permit the library providing a work for digitization to host its own copy of that digitized work for selected personal use distribution.
Both permit unlimited personal use of out–of–copyright works. While redistribution rights vary for other entities, both the Google Books and OCA initiatives permit the library providing a work for digitization to host its own copy of that digitized work for selected personal use distribution.
Both enforce some restrictions on redistribution or commercial use. Google Books enforces a blanket prohibition on the commercial use of its materials, while at least one of OCA’s scanning partners does the same. Google requires users to contact it about redistribution or bulk downloading requests, while OCA permits any of its member institutions to restrict the redistribution of their material.
A common comparison of the Google Books and Open Content Alliance projects revolves around the shroud of secrecy that underlies the Google Books operation. However, one may argue that such secrecy does not necessarily diminish the usefulness of access digitization projects, since the underlying technology and processes do not matter, only the final result. This is in contrast to preservation scanning, in which it may be argued that transparency is an essential attribute, since it is important to understand the technologies being used so as to understand the faithfulness of the resulting product. When it comes down to it, does it necessarily matter what particular piece of software or algorithm was used to perform bitonal thresholding on a page scan? When the intent of a project is simply to generate useable digital surrogates of printed works, the project may be considered a success if the files it offers provide digital access to those materials.To me, that paragraph gets at the key issue in discussing and comparing the projects -- are books being scanned in a consistent way and being made accessible through at least one portal, enforcing current rights restrictions? Yes? Then both these projects are, at a basic level, successful and provide a useful service.
Posted by
Leslie Johnston
at
12:36 PM
0
comments
Labels: digitization, intellectual property
Via TeleRead, the 2008 Frankfurt Book Fair conducted a survey on how digitization will shape the future of publishing. The summary results are available in a press release.
These are the top four challenges facing the industry identified through the survey:
• copyright – 28 per cent
• digital rights management – 22 per cent
• standard format (such as epub) – 21 per cent
• retail price maintenance – 16 per cent
Not knowing what the details of these concerns really are in their survey results, as generalizations the first three are an interesting overlap with challenges facing digital collection building in libraries. What are appropriate terms for copyright and licensing for libraries? How do we identify/document copyright (and other rights) status? How do we manage access and provide for fair use with varying DRM scenarios? What standards will enhance preservation and ongoing access?
Posted by
Leslie Johnston
at
4:12 PM
0
comments
Labels: publishing
Via the Digital Curation Blog, I came across the DCC Curation Lifecycle Model. This is a very interesting high-level overview of the life cycle stages in digital curation efforts. There's an introductory article available.
The model proposes a generic set of sequential activities -- creating or receiving content, appraisal, ingest, preservation events, storage, etc. There are some decisions points at the appraisal and preservation event stages about next steps -- refusal, reappraisal, migration, etc. A colleague and I sat together and looked it over this afternoon. We were both looking at it from a perspective of a digital collections repository and not an IR, and the model was designed primarily with IRs in mind, so our thoughts are coming from a different place in terms of what we wanted to see additionally taken into account in the visualization.
There's a "transform" activity -- definitely something that takes place potentially multiple times in a data life cycle. In the visualization this appears sequentially after "store" and "access, use and reuse." This is an activity that's hard to include in a visualization of a sequence because it can take place at so many points, but it feels like it should be earlier in the sequence, perhaps before those two steps.
The next ring is labeled with the activities "curate" and "preserve" with arrows. Does the placement of the terms and arrows mean anything in relation to the outermost ring? Are "ingest," "preservation activity" and "store" part of "preserve" and the rest part of "curate?" Or does this more simply represent ongoing activities?
The center of the model is the data. It's surrounded by a ring for descriptive and presentation information. It's an activity of central importance and is directly related to the data as is shown, but we weren't sure how its placement related to the sequence of tasks in the visualization.
"Preservation planning" is the next ring out. Planning and implementation are a central, ongoing activity. We also weren't sure when this ongoing activity meshed with the sequence.
"Community watch and participation" is the last remaining inner ring. It's also on ongoing activity. What actions might the outcomes of this activity affect?
Overall, this is a good model for planning. It's challenging to create a visualization for complex processes and dependencies and this covers a lot of ground. And of course it's meant to be generic and high-level, to be made more concrete by an institution that makes use of it. It certainly stimulated our thinking in terms of how we might model our data life cycle and the dependencies between the various tasks.
NOTE: Sarah Higgins, who created the model, has provided excellent responses to my thoughts and questions in the comments to this post. Please read them!
Posted by
Leslie Johnston
at
3:27 PM
2
comments
Labels: preservation, visualization
A cyclorama was the cutting-edge multimedia installation of its time in the 1870-90s. A massive 360 degree painting in the round, it was often accompanied by narration, music, and a light show to heighten the illusion. Today I went to see the conserved, restored, and reinstalled Gettysburg Cyclorama at the new visitors' center. The center opened in April, but the Cyclorama only reopened 10 days ago.
I'm glad I went. True, you only get to spend 15 minutes in the Cyclorama gallery and you have to sit through a short movie about the battle first because the museum, movie, and cyclorama are on one ticket. The new museum is very nicely designed and installed (and extensive), the movie is not too long and very well-done, and the tickets are reasonably priced.
The painting (by Paul Philippoteaux, 1884) was installed in the new facility with its diorama foreground illusions recreated. They run a 15-minute narrated sound and light show to recreate Pickett's Charge (dawn over the battlefield is amazing), then they bring up the lights for a few minutes so you can see the entire painting clearly. In some spots the diorama leads seamlessly into the painting. It's still an amazing illusion and it takes your breath away.
The Cyclorama painting was previously housed in a Richard Neutra-designed building at Getttysburg. The Neutra building is scheduled for demolition in December 2008, but there is litigation to attempt to stop it. That will be a difficult case -- battlefield restoration versus Modern architecture preservation.
Posted by
Leslie Johnston
at
10:03 PM
1 comments
Labels: museums
The Federal Agencies Digitization Guidelines Initiative site went live on September 30, 2008. The initiative represents a collaborative effort between U.S. government agencies to establish a common set of guidelines for digitizing historical materials. Participants include the Defense Visual Information Directorate, the Library of Congress, the National Agricultural Library, the National Archives and Records Administration, the National Gallery of Art, the National Library of Medicine, the National Technical Information Service, the National Transportation Library, the Smithsonian Institution, the U.S. Geological Survey, the U.S. Government Printing Office, and The Voice of America.
The Still Image Working Group is focusing its efforts on books, manuscripts, maps, and photographic prints and negatives. There are draft "Digital Imaging Framework" and "TIFF Image Metadata" documents available. The Audio-Visual Working Group effort will cover sound and video recordings and will consider the inclusion of motion picture film as the project proceeds. That group is still at the document drafting stage.
Posted by
Leslie Johnston
at
2:32 PM
0
comments
Labels: standards
From Boing Boing:
American Memory is a new and compelling DVD coming from extended Skinny Puppy posse members William Morrison and Justin Bennett later this year. It took me a while to figure out exactly what was going on (and exactly who was responsible), but that didn't detract from this hypnotic and ultimately forceful piece.
The voice in the clip on the DVD's trailer is that of former slave Alice Gaston, interviewed in her eighties for the Library of Congress in 1941. The actress is lip-synching to her dialogue. Videomaker William Morrison explains that the whole project works this way, using audio from the American Memory Archive along with new and processed footage. And, of course, Skinny Puppy music.
According to Morrison: "The theoretical context of the project is that some time in the very distance future, long after America is gone, some artists scouring the backwater of whatever the net has become discover the American Memory Archive. They have no context for it's meaning but are intrigued by the sights and sounds. They create surreal impressions of the material they find and broadcast it back through time. A quantum radio channel beamed into the sub conscious minds of the 21st century."
A few different permutations of the band will be playing a show on December 4 at the Gramercy in NYC.
Posted by
Leslie Johnston
at
4:17 PM
0
comments
Labels: digital life
This is a personal weblog. The opinions expressed here are my own and not those of my employer.
This work is licensed under a Creative Commons Attribution-NonCommercial 2.5 License.