Thursday, October 02, 2008

interesting re-use of American Memory content

From Boing Boing:

American Memory is a new and compelling DVD coming from extended Skinny Puppy posse members William Morrison and Justin Bennett later this year. It took me a while to figure out exactly what was going on (and exactly who was responsible), but that didn't detract from this hypnotic and ultimately forceful piece.

The voice in the clip on the DVD's trailer is that of former slave Alice Gaston, interviewed in her eighties for the Library of Congress in 1941. The actress is lip-synching to her dialogue. Videomaker William Morrison explains that the whole project works this way, using audio from the American Memory Archive along with new and processed footage. And, of course, Skinny Puppy music.

According to Morrison: "The theoretical context of the project is that some time in the very distance future, long after America is gone, some artists scouring the backwater of whatever the net has become discover the American Memory Archive. They have no context for it's meaning but are intrigued by the sights and sounds. They create surreal impressions of the material they find and broadcast it back through time. A quantum radio channel beamed into the sub conscious minds of the 21st century."

A few different permutations of the band will be playing a show on December 4 at the Gramercy in NYC.

Wednesday, September 24, 2008

new version of getty introduction to metadata

The third edition of the Getty Introduction to Metadata -- edited by Murtha Baca, with essays by Tony Gill, Anne J. Gilliland, Maureen Whalen, and Mary Woodley -- is now available online and in hard copy. This is a very useful overview and it's nice to see it updated.

Thursday, September 18, 2008

generational myths

Siva Vaidhyanathan has a great article in The Chronicle Review entitled "Generational Myth." Siva and I first met through our online discussion of this topic -- I very strongly agree with him on this issue.

Lorcan Dempsey posted about a couple of blog posts by Andy Powell and Dave White about their takes on this issue. Dave's proposed "Resident" and "Visitor" categories and his acknowledgment of the spectra of behaviors that these categories represent is a well-considered take on how libraries might better understand styles of learning of distance students in particular. I'm obviously not a fan of human categorization -- people are notoriously hard to pigeonhole. But I think these are actually more akin to personas than categories, like those that you'd develop as an exercise when designing a new online service. Not unerringly accurate, but not without usefulness. It's certainly supplements the often simplistic thinking about our users as "faculty" or "graduate students" or "undergraduates" or "the public."

I also strongly recommend Janna Brancolini's blog Generation Underrated, her response to Mark Bauerlein's The Dumbest Generation: How the Digital Age Stupefies Young Americans and Jeopardizes Our Future (Or, Don't Trust Anyone Under 30). Check out what someone under 30 has to say, who also happens to be the daughter of a digital librarian.

grapes need a eula?

From Serious Eats, an image of an empty bag of grapes ... with a EULA.

The recipient of the produce contained in this package agrees not to propagate or reproduce any portion of the produce, including (but not limited to) seeds, stems, tissue and fruit.
To me this is particularly amusing because they're seedless grapes ...

Wednesday, September 17, 2008

djatoka

There is major buzz around the announcement that the Los Alamos National Laboratory Research Library has released djakota, a "reuse friendly" open source JPEG2000 Image Server. It's available on SourceForge under a GNU Lesser Public General License.

There's an excellent D-Lib article that fully describes the server. I love the first sentence of the article: "The Digital Library Research & Prototyping Team at the Los Alamos National Laboratory (LANL) enjoys tackling challenging problems." Now there's an understatement!

We did some explorations with Kakadu (one of the components of djakota) when I was at UVA, and we use Aware at LC. I plan to take a long, hard look at this.

smithsonian digitization initiative

There's an announcement on CNN that the Smitshonian plans to put its 137 million object collection online. The new Smithsonian Secretary G. Wayne Clough said in an interview that they do not yet know how long it will take or how much it will cost to digitize the full 137 million-object collection and will do it as money becomes available. A team will prioritize which artifacts are digitized first. They plan to focus on making the collections usable for the K-12 audience.

When I was at the Smithsonian yesterday for David Weinberger's talk, this seemed to be a buzzing topic of discussion among audience members; one Smithsonian employee even mentioned it in a question to Weinberger, expressing a certain level of surprise.

Tuesday, September 16, 2008

small pieces loosely joined by metadata



Today I had the extreme pleasure of attending a talk by David Weinberger (The Cluetrain Manifesto, Small Pieces Loosely Joined, Everything is Miscellaneous) at the Smithsonian, entitled "Knowledge, Noise and the End of Information." It was webcast, and I strongly suggest viewing it if you can.

There was lots of interesting discussion about the definition of information, the innately social nature of the human race and how social interaction is a vital aspect of information discovery, and how the loosely joined and messy nature of the internet just reflects human nature and is not a bad thing. He also stressed that one can never know what digital information might be of importance in the future, so we should, as cultural institutions, be striving to keep as much as possible. He also touched on the importance of brand and authoritativeness, but not to equate that with control.

A word I did not expect to hear today, let alone about a hundred times, was "metadata." The cellphone image above is a shot of one of his concluding statements. He talked a lot about the importance of metadata, whether it be authoritative cataloging, community tagging, or contextual relationships through linking. Since we cannot ever imagine all the uses for our digital content we cannot possibly expend the costly effort to provide all the descriptive metadata that every community might want or need, so all three are complementary and of equal value.

One of my take-aways was that this again shows the importance of just getting digital content out there. Let the content express itself through its authoritative metadata, but also provide open access and support multiple mechanisms through which it can be incorporated into new contexts and uses and gain new descriptions.

Monday, September 15, 2008

open access to museum collections

Last Friday there was a post on Open Access News that Wake Forest University's Anthropology Museum had issued a press release about the launch of its online collections, supported by an IMLS grant.

I welcomed this news on many fronts -- there aren't enough ethnographic or archaeological collections online; the museum is using Re:discovery, a great product geared toward small museums; and I have a number of friends with ties to Wake Forest and I've visited Winston-Salem many times and have a fondness for the area.

What made me sit down to think about this for a few days was the passing description of this an an Open Access project.

I worked for many years in the museum community, and every museum that I ever worked for or consulted for wanted to make its collections available in one digital form or another. The Museum Computer Network was founded in 1967 to enable museums to automate their processes and convert collections records to digital form. Museums were among the earliest institutions to share their collections online in the mid 1990s. The University of California Museum of Paleontology had a web site in 1994. The Fine Arts Museums of San Francisco brought their Thinker "imagebase" online in 1996 -- and they had volunteers assist with an early form of experimental user supplied subject metadata. e.g., proto-tagging. By 1997 the National Gallery of Art provided access to over 100,000 objects in its collection, and the Los Angeles County Museum of Art experimented with converting print museum catalogs into freely available online publications.

Sure, there have been lengthy discourses about levels of access to the digital media surrogates and questions of rights and control of those new media assets, and there is some information about the acquisition of objects that's subject to privacy restrictions, but no museum wants to limit discovery of their collections -- they want to facilitate their collections' use in research and teaching.

I've just not heard it described as "open access" before.

I'm not saying that it isn't a sort of open access initiative -- it most obviously is -- but I just think of it as such a normal museum activity I don't categorize it in my mind as anything other than business as usual. Then it hit me -- for the past 15 years museums have been major players in the open access movement without necessarily always knowing it.

Labeling this an open access initiative re-contextualizes this core museum activity into a different realm -- one that I hope will make museum collections information more visible and reinforce the importance of all categories of open access content.

Friday, September 12, 2008

nsdl metadata registry

This afternoon a group of us had the opportunity to sit down with Jon Phipps, implementer of the NSDL Metadata Registry.

I knew that such a thing existed. I understand RDF. I know about SKOS. I hadn't really given a lot of thought as to how to best take advantage of it.

Today, I had one of those skies are opening and angels are singing from on high moments. RDF can be used to model relationships between concepts and potentially enforce them through schemas. This can obviously be applied to improve discoverability when a hierarchical taxonomy is employed. Then my LC colleague Clay Redding said that he was experimenting with multiple schemas and managing additional local alternative labels in addition to authoritative preferred labels. And then Jon and Ed Summers mentioned the potential for this tool to map across schemas. My a-ha moment was understanding the potential for formalized mappings across metadata schemas to improve discoverability within and across collections described with hetreogenous taxonomies and vocabularies.

I remember using Chenhall's Nomenclature in records for ethnographic objects where we recorded every level of the hierarchy in its own field -- It was madness. I remember when we were in the early days of the AAT, busily submitting new terms and building the hierarchies, our dream was searching for "case furniture" and getting results with bookcases, chests, desks, wardrobes, and every semantic child where "case furniture" never appeared in the record.

I remember some research at USC in the late 1990s about thesaurus-enabled searching. OCLC's Metadata Switch project has done some work in cross-schema mapping. I know this is very difficult to accomplish. Today was the first time I saw a tool that might make the conceptual mapping simpler. But not simple. This is a potentially massively overwhelming task if it can't be done programmatically to a large extent.

I'm coming late to the party, but now I'm really intrigued by what might be accomplished in this arena.

Tuesday, September 09, 2008

Cory Doctorow book of essays

I am a big fan of Cory Doctorow's writing -- his fiction and his essays on technology, rights, and privacy. Via BoingBoing, comes word of his new book of essays -- Content: Selected Essays on Technology, Creativity, Copyright, and the Future of the Future.-- which he is making available as a free Creative Commons licensed PDF download.

I you haven't read Cory Doctorow yet, you should. I don't always agree with everything he says, but he is thoughtful and technologically savvy and writes thorough essays on very relevant topics in an entertaining style.

I've read some of these essays before, but having them together in one beautifully-designed volume that I can always refer to is the proverbial good thing.

LoC Repository Development Group hiring

Our group has a position open. Visit the LoC jobs page and search for posting "080214". The posting does not mention our unit specifically, so this is a head's up that the job is with us. We're still a relatively new group, working on a variety of projects with many units across the Library and developing our group's role in the institution.

The application period closes on October 3, and that is an absolute deadline. You must apply using an online federal job application system -- it's a lengthy form that requires some time to fill out. Be prepared with electronic copies of your documents to cut-and paste.

EDIT (9/24/2008): This position reports to the Director of the Repository Development Group. Everyone in the team -- including me -- reports to the Director. There is no additional management structure.

Monday, September 08, 2008

google newspaper digitization

Google is digitizing newspapers.

Not only will you be able to search these newspapers, you'll also be able to browse through them exactly as they were printed -- photographs, headlines, articles, advertisements and all.

This effort expands on the contributions of others who've already begun digitizing historical newspapers. In 2006, we started working with publications like the New York Times and the Washington Post to index existing digital archives and make them searchable via the Google News Archive. Now, this effort will enable us to help you find an even greater range of material from newspapers large and small, in conjunction with partners such as ProQuest and Heritage, who've joined in this initiative. One of our partners, the Quebec Chronicle-Telegraph, is actually the oldest newspaper in North America—history buffs, take note: it has been publishing continuously for more than 244 years.

You’ll be able to explore this historical treasure trove by searching the Google News Archive or by using the timeline feature after searching Google News. Not every search will trigger this new content, but you can start by trying queries like [Nixon space shuttle] or [Titanic located]. Stories we've scanned under this initiative will appear alongside already-digitized material from publications like the New York Times as well as from archive aggregators, and are marked "Google News Archive." Over time, as we scan more articles and our index grows, we'll also start blending these archives into our main search results so that when you search Google.com, you'll be searching the full text of these newspapers as well.
It's interesting that they're working directly with publishers and with aggregators such as ProQuest to digitize and improve discoverability of back files. That's good news, but do they also plan to work with major newspaper open access projects such as the National Digital Newspaper Program? Are they digitizing any collections in addition to publisher collections?

When I last looked at the Google news archive in September 2006 I found that way too much of the content was pay-per-view, made you pay even if your institution had licensed subscription access, and didn't work with OpenURL resolvers. I don't see that any of that has changed. I hope it will.

vintage museum photos

Via BoingBoing, check out these fabulous vintage photos from the American Museum of Natural History. I love dioramas, and the exhibit installation images are just great. Taxidermy mounting, diorama background painting, articulating dinosaur bones, casting animal models ... And the vintage exhibitions! I love the images of earnest children being led around ... and doing so-called Indian dances in their construction paper bonnets. State-of-the-art, 1900s-1970s.

Saturday, September 06, 2008

ambient awareness

This week's New York Times Magazine has a piece by Clive Thompson that explores issues around ambient awareness and privacy. Facebook, twitter, flickr, dopplr, and texting and blogging more generally. Is it narcissistic to broadcast your status using awareness tools? Are these tools to improve connectedness in a more mobile and global human ecology -- the ultimate tools for building and maintaining relationships?

This is the paradox of ambient awareness. Each little update — each individual bit of social information — is insignificant on its own, even supremely mundane. But taken together, over time, the little snippets coalesce into a surprisingly sophisticated portrait of your friends’ and family members’ lives, like thousands of dots making a pointillist painting. This was never before possible, because in the real world, no friend would bother to call you up and detail the sandwiches she was eating. The ambient information becomes like “a type of E.S.P.,” as Haley described it to me, an invisible dimension floating over everyday life.
...
And when they do socialize face to face, it feels oddly as if they’ve never actually been apart. They don’t need to ask, “So, what have you been up to?” because they already know. Instead, they’ll begin discussing something that one of the friends Twittered that afternoon, as if picking up a conversation in the middle.
An interesting section focuses on the so-called "Dunbar Number" -- just how many people can you be "friends" with, anyway? According to anthropologist Robin Dunbar, about 150. Can you max out on social connectedness? Not really, since many of one's ambient connections are weak ties, not close, intimate friends. But weak ties are just an important part of social and professional networks.

I find it useful to check in on my Facebook account and see the status newsfeeds of my friends and colleagues. I have also personally met all but a handful, and I believe that they are controlling their feeds and filtering what they write in their status that maintains their chosen levels of privacy. I keep my status updated. I blog, and I know and expect that people who have never met me read it. But is the ability to follow personal newsfeeds and tweets of people you will never know a creepy invasion of privacy, making it too easy to develop parasocial relationships? Or is it all just part of ubiquitous ambient awareness where participation is increasingly not optional?

I originally refused to blog or join Facebook because I thought it was vain to assume that anyone wanted to know what I was thinking or doing, and that I'd be giving up my privacy. OK, I have given up some of my privacy, but I've also made new connections I might never have otherwise, re-established relationships that had gone dormant, and built stronger ties with geographically disparate friends. While I'm not willing to give up my privacy for a free cup of coffee, I am willing to give up some privacy to to that.

Wednesday, September 03, 2008

HathiTrust

The University of Michigan has announced that their MBooks initiative has grown into a shared repository effort called the HathiTrust (pronounced hah-TEE).

HathiTrust was originally a collaboration of the thirteen universities of the Committee on Institutional Cooperation (CIC) to establish a repository for those universities to archive and share their digitized collections. All content to date has been supplied by the University of Michigan and the University of Wisconsin, and Indiana University and Purdue University will soon be contributing their digital materials. 20% of its current content is open access and 80% is restricted. Don't look for a single search interface yet -- it's planned. As they say: "Good, useful, technology takes time.... and the strength and insight born of collaborative work."

The new HathiTrust initiative has been funded for an initial five-year period beginning January 2008, and is now open to other institutions. Partners will be charged a one-time start-up fee based on the number of volumes added to the repository, in addition to an annual fee for the curation of the data. They already support both open access and dark archive materials, and will also do so for new partners.

Their July 2008 monthly report gives a good sense of their activities. It is interesting to note that the only initial ingest workflow supported is the Google partner workflow. That's not too surprising since this work is based on the MBooks project developed in support of Google content workflows. That in and of itself ensures that there many institutions who'll be considering partnership.

This announcement is exceptionally exciting. I look forward to its development as a service.

Monday, September 01, 2008

kete 1.1

I've blogged about Kete before - version 1.1 has been released.

From the announcement:

Kete 1.1 is now available with a giant helping of new features and improvements. This is also the first release where you can grab Kete from our code repository's new home at Github.com. See http://kete.net.nz/site/topics/show/25-downloads for details or browse the code online at http://github.com/kete/kete/.

For those who haven't seen Kete in action, Kete is open source software that enables communities, whether the community is a town or a company, to collaboratively build their own digital libraries, archives and repositories. Kete combines features from Knowledge and Content Management Systems as well as collaboration tools such as wikis, blogs, tags, and online forums to make it easy to add and relate content on a Kete site. You could create a service like Google's Knol for your community using Kete.

An in-depth list of features and issues resolved can be found at http://kete.net.nz/documentation/topics/show/182-kete-11-features-and-bug-fixes , but here are some highlights:

Saturday, August 30, 2008

web archiving

The Library of Congress has a phenomenal Web Capture team, staffed with very dedicated people who take a lot of effort to identify web sites that best document an event, crawl and capture sites through partner Internet Archive, work with cataloging to get the sites described to enhance discoverability, do quality control to make sure the archived sites will run correctly, and then make the archived sites live for public access. This process can take a very long time to ensure that a site is fully captured, preserved, and accessible.

The Web Capture team is, as they have with previous years, documenting the 2008 elections. They don't crawl sites without permission, and they always send requests. A colleague at another library sent me a link to a post and series of comments on Wonkette that were a reaction to a LoC request to capture the site. The post itself is fine. It is more than a bit surreal to get such a request from LoC -- they're going to collect what I write? -- and making fun of it is OK.

Some of the comments, however, are another story.

The reaction to the notice of the request includes strings of profanity, vulgarity, and various exhortations to "archive this, LoC!" Some comment that it's possibly a fake request similar to a Nigerian scam, some liken it to FBI wiretapping, and one comment says that it's a waste of taxpayer dollars to have federal employees reading websites in order to identify what should be archived. One comment conjectures that by "capture," we mean print out the site and store it in a box next to the Ark of the Covenant. Some of the comments are obviously humorous and some are serious, and it's hard to tell with others.

I have a sense of humor, especially about political topics. Of course it's funny to the Wonkette participants that whatever is said, whether profound or mundane or profane, the Library of Congress will crawl it. I remember my own reaction when I was approached about submitting my email to the MCN archives covering the period when I was on its board, which contained such highlights as "The membership brochure is at the printer" and "Don't faint when you see how much the conference hotel wants to charge us for internet access." But for some reason this really struck a nerve because some of the commenters were so "f--- you, Library of Congress." That saddened and angered me.

It's a huge effort to collect ever-changing interactive born-digital resources compared to print materials, but we and many others libraries do it because it's an equally important form of publishing. Libraries collect whatever is relevant regardless of their form of publication. Sites like these are important because they reflect what's really being said and what people really think about the political process. What about that isn't worth collecting?

I'll cop to being a bit overly sensitive on this, but only because I place very high value on such collecting activities.

Friday, August 29, 2008

the omnivore's 100

The blog Very Good Taste has come up with a list of 100 items that every omnivore should try in his or her life. Not surprisingly, it has turned it into a meme that I found through Serious Eats. Basically, you copy the list from Very Good Taste's The Omnivore's 100 and post it to your blog, bolding the items you've tried and striking through any you would never try.

1. Venison
2. Nettle tea
3. Huevos rancheros
4. Steak tartare
5. Crocodile [I've had alligator on a number of occasions -- would that count?]
6. Black pudding [I'll eat it but I don't seek it out]
7. Cheese fondue
8. Carp
9. Borscht
10. Baba ghanoush
11. Calamari
12. Pho
13. PB&J sandwich
14. Aloo gobi
15. Hot dog from a street cart
16. Epoisses
17. Black truffle
18. Fruit wine made from something other than grapes
19. Steamed pork buns
20. Pistachio ice cream
21. Heirloom tomatoes
22. Fresh wild berries
23. Foie gras [I've never understood foie gras worship]
24. Rice and beans
25. Brawn, or head cheese [my mother loved it but couldn't convince me to eat it growing up]

26. Raw Scotch Bonnet pepper [there was the incident with some peppers past their prime, the disposal, and the resultant evacuation of my kitchen]
27. Dulce de leche
28. Oysters
29. Baklava
30. Bagna cauda
31. Wasabi peas
32. Clam chowder in a sourdough bowl
33. Salted lassi
34. Sauerkraut
35. Root beer float

36. Cognac with a fat cigar
37. Clotted cream tea
38. Vodka jelly/Jell-O
39. Gumbo
40. Oxtail
41. Curried goat

42. Whole insects
43. Phaal
44. Goat’s milk
45. Malt whisky from a bottle worth £60/$120 or more
46. Fugu
47. Chicken tikka masala
48. Eel
49. Krispy Kreme original glazed doughnut
50. Sea urchin [I don't much care for it, and Bruce is allergic to it]
51. Prickly pear
52. Umeboshi
53. Abalone
54. Paneer
55. McDonald’s Big Mac Meal
56. Spaetzle
57. Dirty gin martini
58. Beer above 8% ABV
59. Poutine
60. Carob chips
61. S’mores
62. Sweetbreads
[not a favorite, but I will eat them if I know the preparation will be excellent]
63. Kaolin
64. Currywurst
65. Durian
66. Frogs’ legs
67. Beignets, churros, elephant ears or funnel cake
68. Haggis
69. Fried plantain
70. Chitterlings, or andouillette [are you noticing a trend that offal is a category I don't much care for?]
71. Gazpacho
72. Caviar and blini
73. Louche absinthe
74. Gjetost, or brunost

75. Roadkill
76. Baijiu
77. Hostess Fruit Pie
78. Snail [only once, when a donor at an event handed it to me and I felt I had to eat it]
79. Lapsang souchong
80. Bellini
81. Tom yum
82. Eggs Benedict
83. Pocky
84. Tasting menu at a three-Michelin-star restaurant [a 2-star, yes, not yet at a 3-star]
85. Kobe beef [actually, wagyu, but I'm counting it]
86. Hare
87. Goulash
88. Flowers
89. Horse
90. Criollo chocolate
91. Spam
92. Soft shell crab
93. Rose harissa
94. Catfish
95. Mole poblano
96. Bagel and lox
97. Lobster Thermidor
98. Polenta
99. Jamaican Blue Mountain coffee
100. Snake [I had iguana once in Mexico ...]

Wednesday, August 27, 2008

dead sea scrolls

When I was growing up, my mother had a small selection of books displayed between decorative bookends on her coffee table -- a set of 4 art history overview volumes with high quality color reproductions on glossy paper, and a book on the Dead Sea Scrolls. I was fascinated by the volume on ancient art and the book on the scrolls because of their sheer antiquity. I don't remember there being many illustrations in the book, but the story of the discovery of the scrolls was a very engaging one. I don't remember every asking my Mom why she that volume on display, or, if I did, what her answer was.

The New York Times today reports on the project to digitize the Scrolls. It's interesting to read that they plan to create new digital images, as well as digitizing the infrared images created of the scrolls in the 1950s.

Tangentially, there was an article in The Australian a couple of week ago about the conservation and multi-spectral imaging of scrolls from the Villa dei Papyri at Herculaneum.

Tuesday, August 26, 2008

Executive Director of OCA named

A press release went out tonight naming Maura Marx -- founder of the Digital Library Program at the Boston Public Library -- as the first Executive Director of the Open Content Alliance.

“Maura's background in working both inside and outside the library system will help her communicate with a broad public audience the shape of the new public library services in this digital age." said Brewster Kahle, Digital Librarian of the Internet Archive. “Her dynamic style, deep-seated commitment to open principles, and demonstrated success at implementing partnerships and initiatives in the digital space will be a powerful combination in taking the OCA to the next level.”
I met Maura at a meeting this spring, and I know that she's an excellent choice!

Monday, August 25, 2008

concordance as word cloud

Eric Lease Morgan posted about a cool little hack to present a text concordance as a word cloud. A visualization of a concordance -- what a nice idea! It would be interesting to see one at a larger scale -- for every word in a book. I'd like to see how the visual metaphor scales.

Eric said one thing, though, that gives me pause:

"It is a trivial example of how libraries can provide services against documents, not just the documents themselves."

He is absolutely right -- it is a trivial effort to create this useful service. What is still unfortunately not as trivial as it should be is getting access to accurate transcriptions of all the texts once might want to analyze. There are ascii transcriptions for many, many works, but there is always a question of accuracy, and if the desired edition(s) are available. There's OCR, but it's a fair amount of effort to check and correct the output. Google isn't releasing its OCR, but even if they did that, too needs correction. Keyboarding is expensive. And many works in copyright haven't been touched for fear of legal action.

We have the ability to build extraordinary analytical tools. Where is the critical mass of text content?

Mickey Mouse copyright

Via Techdirt and the L.A. Times, an interesting overview on the copyright status of Mickey Mouse. The Virginia Sports and Entertainment Law Journal article by Douglas Hedenkamp mentioned is available online through the "Opposing Copyright Extension" site, as is the original student work by Lauren Vanpelt.

vintage tech

I'm a sucker for vintage technology and vintage manuals. I have a small collection of the latter. I'm a big fan of The Computer History Museum in Mountain View, California. I love to read books about the history of computing.

The Alameda County Computer Resource Center (ACCRC) in Berkeley, California is a non-profit group that recycles hardware. The ACCRC has launched the blog "It Ain't Dead Yet" to showcase their more unusual finds, partly to share the wonder and partly to gauge the usefulness and value of the items. Now there's a feed I'm sure to read every day!

Wednesday, August 20, 2008

Registry of U.S. Government Publication Digitization Projects

I didn't know the Registry of U. S. Government Publication Digitization Projects existed:

"The Registry contains records for projects that include digitized copies of publications originating from the U.S. Government. It serves as a locator tool for publicly accessible collections of digitized U.S. Government publications; increases awareness of U.S. Government publication digitization projects that are planned, in progress, or completed; fosters collaboration for digitization projects; and provides models for future digitization projects."

The Registry has recently been updated, and they welcome additions. Institutions need to apply to contribute.

Ithaka's 2006 Studies of Key Stakeholders in the Digital Transformation in Higher Education

Ithaka has recently released the full findings from their 2006 surveys of the behavior and attitudes of faculty members and academic librarians. The faculty study focuses on the relationship between faculty and the library, faculty perceptions and uses of electronic resources, the transition from print to electronic journals, faculty publishing preferences, e-books, digital repositories, and the preservation of scholarly journals. The librarian survey complements the faculty study, exposing the similarities and differences between faculty and librarian views of key topics.

This is an extended quotes describe two very interesting key perceptual differences:

Over the course of these three surveys, we have tested three “roles” of the library – purchaser, archive and gateway. We have attempted to track how the importance of these three different roles has changed over time. Most highly rated among these roles is that of library as purchaser – faculty don’t want to have to pay for scholarly resources, a finding which holds across disciplines and has remained stable over time. There is slightly more variation by discipline in views on the importance of the library’s preservation function, but valuation of this role is also uniformly high and has remained static over time. The importance of the role of the library as a gateway for locating information, however, varies more widely and has fallen over time.

The declining importance assigned to the gateway role is cause for concern in general, and especially when considered by discipline. The importance to faculty of this role has decreased across all disciplines since 2003, most significantly among scientists. While almost 80% of humanists rate this role as very important, barely over 50% of scientists do so. Beyond the differences between these general disciplinary groups, there also exist substantial variations by individual discipline, as demonstrated by the perceptions of economists. Between 2003 and 2006, the percentage of economists indicating they found the library’s gateway role to be very important dropped almost fifteen percentage points. In 2006, the percentage of economists who believed this gateway role to be very important was actually below the average level of scientists, falling to 48%.

The decreasing importance of this gateway role to faculty is logical, given the increasing prominence of non-library discovery tools such as Google in the last several years. Since 2003, the number of scholars across disciplines who report starting their research at non-library discovery tools, either a general purpose search engine or a specific electronic resource, has increased, and the number who report starting in directly library-related venues, either the library building or the library OPAC, has decreased. Despite the rising popularity of tools like Google, overall, general purpose search engines still slightly trail the OPAC as a starting point for research, and are well behind specific electronic research resources. This overall picture, however, hides a number of variations by discipline; scientists typically prefer non-library resources, while humanists are more enthusiastic users of the library.

The declining importance of this role to faculty stands in stark contrast to the perceptions of librarians, as shown by our 2006 librarian survey. Although the importance of the library’s role as a gateway to faculty is decreasing, rather dramatically in certain fields, over 90% of librarians list this role as very important, and almost as many – only 5 percentage points less – expect it to remain very important in 5 years. Obviously there is a mismatch in perception here.

Librarians at all sizes of institutions see this gateway role as among their primary goals; this, along with the licensing of electronic resources and maintaining a catalog of their resources, are by far the roles most broadly considered important. They expect most of the roles of the library to rise in importance, or at least hold steady, over the next five years, with some notable exceptions to be found in roles focused on nondigital materials, such as roles relating to traditional print preservation and the maintenance of a local print journal collection, which are expected to decline in importance. There are some variations by institution size. Several roles, most notably the development and maintenance of special collections and several more technical tasks such as the management of datasets, are significantly more important at larger libraries than smaller ones. And unlike smaller libraries, larger libraries view licensing as their single most important activity, with less emphasis put on the gateway and catalog roles. This may be a sign that leading-edge libraries are beginning to change their priorities to match those of faculty and students. Still, the mismatch in views on the gateway function is a cause for further reflection: if librarians view this function as critical, but faculty in certain disciplines find it to be declining in importance, how can libraries, individually or collectively, strategically realign the services that support the gateway function?

... and

Perceptions of a decline in dependence are probably unavoidable as services are increasingly provided remotely, and in some ways these shifting faculty attitudes can be viewed as a sign of library success. One can argue that the library is serving faculty well, providing them with a less mediated research workflow and greater ability to perform their work more quickly and effectively. In the process, however, they may be making their own role less visible. This indicates a challenge facing libraries in the near future – as faculty needs are increasingly met without the direct intermediation of the library, the importance of the library decreases. Libraries must consider ways which they can offer new and innovative services to maintain, or in some cases recapture, the attention and support of faculty.

Read their full white paper, or review the raw data from the faculty survey or the librarian study.

Tuesday, August 19, 2008

is everything moving into the cloud?

There's an essay entitled "The Future of the Desktop" by Nova Spivack of Twine on ReadWriteWeb. It's a pretty thoughtful opinion piece on the trend where users are moving away from desktop applications towards Web-hosted ones that run in browsers.

He mentions something that I think is vital: everyone has a sense of the personal and "mine," so there has to be some sort of place that each of us can consider to be our "home." He rightly declares that it's not going to live in any one location or on any one device. His "Webtop" paradigm is that instead of launching the browser from the desktop, one would launch the "desktop" from the browser, and that desktop is the personal location where we do our work and interact with the world.

I'm not sure that I fully buy his metaphor that we'll give up being "librarians" ("filing" and managing resources) and fully become "daytraders" (discovering, filtering, and monitoring of trends), in part because search will replace the need to "file" things.

For one, librarians _actually_ do all of the above, but I'm not going to fault him just because he doesn't know what librarians do in their jobs.

What I'm having trouble with is the notion that just because we're working in the cloud we'll stop organizing resources. The "search will replace cataloging" argument that we've heard in libraries is one that I can't buy. Search doesn't work worth a damn if there isn't some level of organization and filing, aka metadata or cataloging. How will these daytraders efficiently filter what they discover and note trends if they aren't organizing and filing? It is true that we'll be managing fewer files _locally_, but we'll be organizing even more files in the cloud. He rightly identifies that there will be more shared, social spaces, and he says that communities will "seamlessly and collectively add, organize, track, manage, discuss, distribute, and search for information of mutual interest." Maybe it's a semantic distinction, but to me that's a resource management activity, just in a much larger and more social realm.

And ah, the dream of semantic search. And the dream of the smart webtop or desktop, where context is easily understood and parsed for data coming in and being queried. I want to believe. I'm waiting.

Where I do buy into the cloud is from a standpoint of portability. Even moving between work and home on two machines, I have found myself storing and organizing more of my resources out there rather than in here. Flickr. Delicious. Bloglines. LibraryThing. Web mail. It would waste more time than I could imagine to keep my life in sync between just two locations, let alone more.

I worry about security and preservation a lot. I lost my home desktop PC drive last year. What if that drive I lost was my only copy (it wasn't) AND flickr suffered a catastrophic failure? There goes the documentation of the past three and a half years of my life. As someone whose career is centered on digitization and management and use of digital files, I have been trained through experience to think in terms of the catastrophic. And to think about rights and ownership. The cloud must become more secure, aware of identities, distributed, and replicated in its file management to assuage my concerns before I fully buy in.

digital is not to blame

I just read a Wired essay entitled "The Critics Need a Reboot. The Internet Hasn't Led Us Into a New Dark Age." In one of those great moments in synchronicity, over the weekend I started reading a blog written by the daughter of a colleague: Generation Underrated. She was spurred to blog as a reaction to Mark Bauerlein's The Dumbest Generation: How the Digital Age Stupefies Young Americans and Jeopardizes Our Future (Or, Don't Trust Anyone Under 30), which is also mentioned in the Wired essay. I haven't read the book, but I am reasonably sure that it would make me crazy to do so. From Janna Brancolini's blog, referring to studies noted in the first chapter of Bauerlein's book:

A test was given to high school seniors in 1955. The same questions appeared on a Gallup survey given to college seniors in 2002. The college seniors in 2002 scored no better than the high school seniors had in 1955. (29)

In other words, the first chapter doesn’t given a single statistic that demonstrates that people under 30 know less than previous generations, either now or when the members of those generations themselves were under 30.

Bauerlein acknowledges this lack of empirical comparisons by saying, “Even if we grant the point that on some measure today’s teenagers and 20-year-olds perform no worse than yesterday’s, the implication critics make seems like a concession to inferiority. Just because sophomores 50 years ago couldn’t explain the Monroe Doctrine or identify a play by Sophocles any more than today’s sophomores doesn’t mean that today’s shouldn’t do better, far better.”

Janna greatly simplifies Bauerlein's argument thusly:
In a nutshell: we’re dumb because we don’t know anything, we’re dumb because we’re letting the Internet and cell phones be used for evil instead of good, and we shouldn’t be dumb since we have so much technology available to combat our overwhelming dumb-ness.
From the Wired essay:
But the latest crop of curmudgeons fail to acknowledge that there is not much new in this parade of the preposterous. The US has a long and colorful history of being taken in by the erroneous and irrational: Salem witches, the "War of the Worlds" radio broadcast, phrenology, and eugenics are just a few choice examples. The truth is that Americans often approach information — online and off — with a particular mindset. "Antirational junk thought has gained social respectability in the United States during the past half century," notes Susan Jacoby in The Age of American Unreason. "It has proved resistant to the vast expansion of scientific knowledge that has taken place during the same period." Jacoby argues that long-standing American values like rugged individualism and the need to question authority have metastasized into reflexive anti-intellectualism and disdain for "eggheads," "elites," and pretty much anyone who might be described as credentialed. This cancerous irrationalism isn't pretty, but it isn't technology's fault, either.
Readers of this blog know that I have a very negative reaction to generalizations like "digital generation" and "digital natives." In the same vein, blaming technology and saying that this generation is the dumbest seems ludicrous to me. IF we accept that this is the dumbest generation (and I don't), there are other places to identify causation/lay blame. Underfunded school systems with a focus on standardized testing rather than critical thinking. The self-esteem movement, which, when taken to extremes, does away with competition and realistic assessment. But technology? Please. There is increased ubiquitousness of technology and media access (and increased media targeting of younger consumers), but its use or lack of use hasn't made an entire generation less educated.

Friday, August 15, 2008

Patry restoring old posts

William Patry has decided to restore many of his posts which he deleted when he closed down his blog. He has been laboriously identifying the posts and plans to restore them very soon.

Red Island Repository Institute

This week the Red Island Repository Institute took place, with a week-long immersion in all things Fedora. The instructors were Sandy Payette, Richard Green, and Matt Zumwalt. It would be hard to think of people who are a better choice to teach the institute besides these three!

Powerpoints from presentations are online, and they provide a great overview of Fedora.

Thursday, August 14, 2008

free copyright licenses upheld

Great news from Larry Lessig:

I am very proud to report today that the Court of Appeals for the Federal Circuit (THE "IP" court in the US) has upheld a free (ok, they call them "open source") copyright license, explicitly pointing to the work of Creative Commons and others. (The specific license at issue was the Artistic License.) This is a very important victory, and I am very very happy that the Stanford Center for Internet and Society played a key role in securing it. Congratulations especially to Chris Ridder and Anthony Falzone at the Center.

In non-technical terms, the Court has held that free licenses such as the CC licenses set conditions (rather than covenants) on the use of copyrighted work. When you violate the condition, the license disappears, meaning you're simply a copyright infringer. This is the theory of the GPL and all CC licenses. Put precisely, whether or not they are also contracts, they are copyright licenses which expire if you fail to abide by the terms of the license.

Important clarity and certainty by a critically important US Court.

Wednesday, August 13, 2008

LibraryThing covers

Last week LibraryThing announced that they were making a million free book covers available. A LibraryThing Developer Key is required, which any LibraryThing member can get.

There are some rules:

  • Retrieve no more than 1,000 cover per day.
  • If covers are fetched through an automatic process (e.g., not by people hitting a web page), you may not fetch more than one cover per second.
  • Do not make LibraryThing cover images available to others in bulk. You may cache bulk quantities of covers.
  • Use must not involve or promote a LibraryThing competitor.

Tim Spalding admits that this service competes with Amazon web service and other commercial vendors, but LibraryThing’s Terms of Service are far more open.

After the announcement I wondered how this was legally possible for such a large number of covers since there are so many variations of rights regarding cover designs. Who holds the rights? The publishers? The designers? Third parties? It's likely it's a wide variety of all of the above. Should we start talking about orphan work book cover designs?

Yesterday Mary Minow posted about this at LibraryLaw Blog. She posits an interesting possibility that this could fall under section 113. Read the comments for more discussion from Peter Hirtle about whether this might also be transformative use of thumbnails that could possibly be covered under fair use. Peter also rightly mentions the market for cover images, since effect on the market is one of the tests for fair use.

Tuesday, August 12, 2008

Aurora

Mozilla Labs is sponsoring a Concept Series -- "... a forum for surfacing, sharing, and collaborating on new ideas and concepts. Our goal is to bring even more people to the table and provoke thought, facilitate discussion, and inspire future design directions for Firefox, the Mozilla project, and the Web as a whole."

The first featured concept is Aurora, from Adaptive Path, a vision for the future of browsers and the web. This isn't a product, it's a visualization of an interactive 3-D navigational paradigm tied to ideas about personalization, authentication, and mobility of a user's preferences, history, and context.

It's worth looking at. There's a quick guide to Aurora's interface and a descriptive concept document.

Sunday, August 10, 2008

on the mastering of new technologies

Dorothea wrote a post to which my only reply can be "Amen, Sister!" She references a great post by Steve Lawson.

I also feel that I'm at the upper end of the technological middle ground. I'm a journeyman scripter and not really a programmer. I still have digital content production chops. My XML markup skills and metadata fu are strong. I can tell you a lot about the inner working of Fedora. I can haul out my atrophying JavaScript, SQL, and Perl skills, and dredge up my minimal PHP skills. I fondly remember my ColdFusion days, and my days employing Lingo in Director to create Shockwave apps.

I'm not really up on the tools that are all the rage these days -- Python, Ruby, Django, or even Java. I've been so focussed on managing projects that I've lost some of my technological edge. Where do I go to regain/retain it? Especially since I am no longer following a path where I spend any time writing any forms of code. I am often asked why I don't go to code4lib -- it's not exactly that I feel over my head, but I'm just not doing the hands-on thing anymore and I don't think there's much I can contribute to a conversation about Python libraries or Lucene optimization.

I'm a fan of DLF Forums. I learn a lot about what tools folks at other institutions are using. But I don't always learn enough about why they use them and what those tools are especially good for. Something I can take back to my own projects and say "Hey, let's consider this solution because it's a great fit for XYZ."

There are things I need to learn about at a pretty deep technical level, but I may never personally apply them. Where do I go for that?

And circling around to another of Dorothea's topics ... I am still too often one of the few women in the room. A recent event I attended had 40 men and 4 women. But then, I am often just the sort of woman who thinks she's not technical enough to attend such events. Perhaps I need to face my own wariness about events like code4lib and just go. And/or stand up alongside others of my kind and start another type of event.

Friday, August 08, 2008

OpenCollection

Via Digital Koans, I came across an open-source collection management systems called OpenCollection. from their site:

OpenCollection is a full-featured collections management and online access application for museums, archives and digital collections. It is designed to handle large, heterogeneous collections that have complex cataloguing requirements and require support for a variety of metadata standards and media formats. Unlike most other collections management applications, OpenCollection is completely web-based. All cataloging, search and administrative functions are accessed using common web-browser software, untying users from specific operating systems and making cataloguing by distributed teams and online access to collections information simple, efficient and inexpensive.

...

OpenCollection is intended as an alternative to expensive proprietary software solutions that have traditionally been used for collections cataloguing and publishing by museums, archives, libraries and other organizations.
Having worked for many years in the museum community and had responsibility for the design, care and feeding of a number of collection management systems, this is pretty stripped down. It has very strong support for the linking of media files. At first I thought it was lacking elements to manage those fiddly details that were so ubiquitous in managing physical collections -- storage location, exhibition history, publication history, valuation, insurance, condition -- but once I created a test object through the basic entry screen the other screens became visible to me. The only thing I didn't find (but may have missed) are elements having to do with packing and shipping, which requires very detailed record keeping. I would have also expected to see more on condition, such as the ability to track a treatment history and document treatment, since there are professional record keeping requirements for conservators.

One annoyance -- the OpenCollection product site has this ribbon of images that kept crossing on top of the text and blocking it. I'm not even using Firefox 3.0, so who knows what caused this.

issues with blogger?

Has anyone else noticed any odd Blogger behavior? When I put up a new post I'm not seeing it on my blog site for a while, sometimes not until the next day. I first noticed it on July 29. The really odd bit it that I don't see what I just posted on my blog index page, but if I click on the current month in the archive I _do_ see it. I didn't worry about it too much until today a colleague in my department said that he noticed it, as well as for some other sites. The feeds are working fine, but the sites are not.

I thought I'd ask if anyone else has seen anything like this before we lay the blame on our IT environment.

i am rich

The brouhaha over the $999 "I am Rich" iPhone app is very amusing. Eight people bought "I am Rich" -- which presents a glowing, animated red jewel -- before Apple pulled it from the App Store. Is it a scam? Conceptual art? Just something funny to do if you've got the money to burn? Is it an apocryphal story that one of the purchasers didn't mean to buy it?

The best article I found was at the Los Angeles Times. Read it quick before it disappears behind a wall ... There's also this posting on Silicon Alley Insider and this article in the Times.

Wednesday, August 06, 2008

time for links and nothing more

I'm really swamped these days, and only have time to post some links to things that caught me eye during the past week:

vi.sualize.us seems like a really interesting social bookmarking tool for images. Perhaps they'll learn what delicious learned and give up the tortured . 's.

William Patry stopped blogging
. I'm not surprised if folks thought his personal blog was the word of Google. It's sad that he also decided to erase his archives, but I understand that he didn't want his past postings to live on and continue to be misunderstood.

It seems that Google is making some of its machine-translation technologies and translation management tools available to human translators, at least as a beta. I'm working with a project that requires translation into 7 languages. Managing this process is very challenging, and I've seen some very bad tools for the process.

Following the Digitization and the Humanities Symposium, Jennifer Schaffner and Merilee Profitt wrote a brief report, The Impact of Digitizing Special Collections on Teaching and Scholarship: Reflections on a Symposium about Digitization and the Humanities. The report acts as a summary of the symposium, and also gives some calls to action, especially about metrics for success.

Duke has launched its Open Library Environment Project with Mellon support. Its focus on back-end open tools is worhtwhile, but I'm not sure I know how this will be integrated with other activities in the community.

Wednesday, July 30, 2008

Hooray -- Fedora 3.0 released

Hooray -- the highly anticipated (at least by me) formal release of Fedora 3.0 is available.

Excerpted from the press release:

Fedora 3.0 features the Content Model Architecture (CMA), an integrated structure for persisting and delivering the essential characteristics of digital objects in Fedora. The software is available at
<http://www.fedora-commons.org/> and at <http://sourceforge.net/projects/fedora-commons>. The Fedora CMA plays a central role in the Fedora architecture, in many ways forms the over-arching conceptual framework for future development of Fedora Repositories. Fedora 3.0 features include:
  • Content Model Architecture - Provides a model-driven approach for persisting and delivering the essential characteristics of digital content in Fedora
  • Fedora REST API - A new API that exposes a subset of the Access and Management API using a RESTful Web interface contributed by MediaShelf
  • Mulgara Support - Fedora supports the Mulgara 2.0 Semantic Triplestore replacing Kowari -Migration Utility - Provides an update utility to convert existing collections for Content Model Architecture compatibility
  • Relational Index Simplification - The Fedora schema was simplified making changes easier without having to reload the database and significantly increasing scalability
  • Dynamic Behaviors - Objects may be added or removed dynamically from the system moving system checks into run-time errors
  • Error Reporting - Provides improved run-time error details
  • Multiple Owner as a CSV String - Enables using a CSV string as ownerID and in XACML policies
  • Java 6 Compatibility - Fedora may be optionally compiled using Java 6 while retaining support for Java Enterprise Edition 1.5 deployments
  • Relationships API - API-M has been extended to enable adding, removing, and discovering RDF relations between Fedora objects
  • Revised Fedora Object XML Schemas - The new schemas are simpler, supporting the CMA and removing Disseminators
  • Atom Support - Fedora objects can now be imported and exported in the Atom format
  • Messaging Support - Integrates JMS messaging for sending notification of important events
  • Validation Framework - Provides system operators a way to validate all or part of their repository, based on content models
  • 3.0-Compatible Service Releases - New versions of the OAI Provider and GSearch services are compatible with Fedora 3.0. The GSearch release also enables messaging support for GSearch, which allows for more robust and seamless integration with the Fedora repository.
I have been waiting for the CMS for some time -- this update to the architecture will greatly improve the flexibility of a Fedora implementation by removing the tight bindings between objects and disseminators and allowing for easier disseminator updating. The validation support is also key if one is interested in working with workflow engine to automatically process tasks and validate production outcomes. I am intrigued by the Atom support -- Dan commented on BagIt/SWORD as a possible repository SIP in one of our discussions at RepoCamp. This could become a very real experiment.

Tuesday, July 29, 2008

Cuil

I finally got through to Cuil, the plocaimed Google threat, this afternoon. After reading Siva's post, I decided to try searching my own name in quotation marks.

I know that there are other Leslie Johnstons. There's sales and marketing consultant, a renowned Scottish footballer from the 1940s-50s, a cancer researcher, etc. We all came up in the first 8 pages that I reviewed. Cuil obviously pulls images that it finds on various sites and associated them with results. In some cases, it correctly included my image with a link that I was associated with. In some cases my photo showed up associated with results for the other Leslie Johnstons. On the very first page a results for an article of mine in D-Lib was accompanied by a photo of my friend and colleague Sarah Shreeves who had an article (and hence a bio with a photo) in the same issue. On another page a result for a different D-Lib article was accompanied by a photo of Chris Awre. Other links for presentation I gave are accompanied by a portrait of Thomas Jefferson (I assume it keyed in on "University of Virginia"). For some links where there were no images there are little images of top nav bars from the page the result points to. In one case there are results from a usability site at UVA with my name on it, but the accompanying image is something that looks like a nav element in Cyrillic, which is definitely NOT on that usability site. Very weird and random.

When I tried the categories -- both of which were for Scottish footballers -- I still got lots of results of mine. I think I can say that none of my presentations or writing on digital library or museum activities ever mentioned Scottish or even American football.

Cuil needs some work. How will it learn?

Wordle

Everyone else has been playing with Wordle for weeks now. I kept thinking about what I might pipe through it, and finally landed on the UVA Repository Case Study that I submitted to Open Repositories 08. I like what it produced:


on NYRB article about Google Books

Jean-Claude Guédon and Boudewijn Walraven submitted letters to the New York Review of Books which have been published as "Who Will Digitize the World's Books?" They are commenting on Robert Darnton's "The Library in the New Age", and he responds to their letters.

Fedora and DSpace collaboration

A press release hit the streets today about a formal collaboration between the Fedora Commons and the DSpace Federation. Excerpt from the press release:

The decision to collaborate came out of meetings held this spring where members of DSpace and Fedora Commons communities discussed multiple dimensions of cooperation and collaboration between the two organizations. Ideas included leveraging the power and reach of open source knowledge communities by using the same services and standards in the future. The organizations will also explore opportunities to provide new capabilities for accessing and preserving digital content, developing common web services, and enabling interoperability across repositories.

In the spirit of advancing open source software, Fedora Commons and DSpace will look at ways to leverage and incubate ideas, community and culture to:

1. Provide the best technology and services to open source repository framework communities.

2. Evaluate and synchronize, where possible, both organizations'technology roadmaps to enable convergence and interoperability of key architectural components.

3. Demonstrate how the DSpace and Fedora open source repository frameworks offer a unique value proposition compared to proprietary solutions.

The announcement came on the heels of an event sponsored by the Joint Information Systems Committee's (JISC) Common Repository Interface Group (CRIG) held at the Library of Congress. The event, known as "RepoCamp," was a forum where developers gathered to discuss innovative approaches to improving interoperability and web-orientation for digital repositories. Sandy Payette, Executive Director of Fedora Commons, and Michele Kimpton, Executive Director of the DSpace Foundation, reiterated their commitment to collaboration and encouraged input and participation from both communities as work gets underway.
The full press release is available. Theres a great photo of Sandy and Michelle in a ceremonial handshake at LoC. Sandy and Michelle led a brief discussion about this last Friday at RepoCamp, and was exciting to watch this initiative launch.

Monday, July 28, 2008

Blow Up for Flickr

Through a post on ReadWriteWeb, I found Blow Up.

Blow Up uses the public Flickr API tp create a slide show presentation that allows your images to be seen in fullscreen mode while still showing thumbnails of the other images in the slideshow and navigation to your sets, while maintaining your image quality (or, in my case, showing me that, when blown up, a lot of my pictures are not quite focussed). It's a very clean UI.

The service is free and doesn't require anything other than your Flickr username to get started (I do wonder if they're storing those). because you are not logging in, Blow Up only shows images that you have set to public viewing. Other functionality include being able to download the Blow Up app to display your Flickr images on your other websites.

RepoCamp

Last Friday I spent all day at RepoCamp. There were at least 40 participants from I don't know how many institutions! Major kudos to David Flanders from the JISC Common Repository Interface Group (CRIG) who did a fabulous job organizing the event, and to Ed Summers who facilitated the LoC side.

There was a lot of great discussions around SWORD and OAI-ORE. I was happy to have the opportunity to talk about BagIt with a group who hadn't encountered it yet, and we had some really interesting discussions. Talking through use cases beyond our initial simple use case -- files from Institution A are transferred to Institution B and stored for preservation with no active access -- there is an obvious need for BagIt profiles that specify what is contained in a Bag and how it's organized for other uses -- like potentially as a SIP for ingest into a repository. Folks were also really interested in the idea of "Holey Bags" where the manifest is a list of URIs for retrieving files. Ideas were batted around about crawls that start out with a minimal manifest of URIs, capture those files, generate checksums, follow links from those files to capture more files and checksums, ending up with a Bag generated on-the-fly from that crawl so you have captured the files and record of the URIs where the files were found. A really interesting suggestion was the use of an OAI-ORE Resource Map to instantiate such a capture. Or for that matter, serve as the fetch file. Bags of course can simply be files on disk, but when it's a Bag of web resources you might want more structure than just a list of locations the files came from. After listening to the discussions I'm convinced that the work we're doing (I should say Ed is doing) with a web app for a Bag deposit service that uses SWORD is going in the right direction. I think were developing some real traction with BagIt.

They video recorded all the elevator pitches and the reports to the whole group. I don't know if, where, or when they'll be available. I am not a fan of seeing myself on video.

It was nice to see some of the Fedora team -- Sandy Payette, Dan Davis, Eddie Shin -- and some UVA colleagues. Its only been 3 or so months but it feels like I left so long ago. I was pleased to meet Ben O'Steen from Oxford (we were following his Fedora IR work when I was at UVA) but I didn't get the chance to really sit down and talk with him. I had so many other interesting conversations that I need to follow up on ...

Tuesday, July 22, 2008

what is a repository?

Yesterday a colleague was chatting with me about what make up a repository. Have we been overthinking what is needed? Can we simplify the tools we use? Recombine lightweight tools in a new way?

This was very timely because I'd seen a posting that JISC's Information Environment team is experimenting with IdeaScale to have a discussion about defining repositories to feed into JISC work on repository architecture.

First -- about IdeaScale:

It begins with an idea posted to your IdeaScale community by a user. Each idea can be expanded through comments by the community. The ultimate measure of an idea is determined by a voting system. Any idea can be voted to the top or buried back down to the bottom. It combines the "wisdom of the crowds" concept with Web 2.0 models like Digg.
I think it's interesting that JISC is trying this approach -- have discussants set out a series of statements about repositories, allow comments, and let members of the community sign up to vote +1 or -1 on the positions.

I have a love/hate relationship with the word "repository." It's next to impossible to define or describe, but I haven't been able to come up with anything better. I'm not sure this activity will produce any solid definitions, but it is generating a very interesting public discussion.

Friday, July 18, 2008

Names Project

I'm intrigued by the Names Project to identify requirements and develop a prototype service that will reliably and uniquely identify individuals and institutions for institutional and subject repositories in the UK. They report anecdotally that more than 75% of authors represented in IRs aren't in LCNAF. The goal is a straightforward and laudable one: a centralized name authority module that will plug into existing and future repository software and provide autocompletion of author names for depositors of materials and for searchers of the systems.

I found this paper by Amanda Hill to be the best introduction to the project. The project has just issued a software specification and I plan to watch its progress.

Thursday, July 17, 2008

fail whale art

I'm not a twitter user (please don't start on me), so I was completely unaware of what the "Fail Whale" is. There's a wonderful post on ReadWriteWeb on the artist behind the graphic and how her work took on a new life through the twitter community.

international copyright law and digitization

In one of those great synchronicities, I've encountered two publications on international copyright law and digitization, both of which are worth reading.

The first is an Information World Review article entitled "Scan and Deliver" about how issues of copyright clearance have affected the British Library's digitization program. (I keep hearing Adam Ant's "Stand and Deliver" in my head)

The second is the International Study on the Impact of Copyright Law on Digital Preservation just released by the Library of Congress. The report is a joint effort of the Library of Congress National Digital Information Infrastructure and Preservation Program, the Joint Information Systems Committee, the Open Access to Knowledge (OAK) Law Project, and the SURFfoundation.

Wednesday, July 16, 2008

ask a person next time

From BoingBoing, a very funny pointer to a Chinese restaurant that relied on an online translation service when they shouldn't have ...

Tuesday, July 15, 2008

ndiipp partners meeting

Last week I attended the three-day 2008 meeting for the partners in the Library of Congress National Digital Information Infrastructure and Preservation Program (NDIIPP). Yesterday a colleague who couldn't attend asked me what stood out for me in the program. I didn't take a lot of notes -- I kept forgetting to because I just wanted to listen -- but I see some patterns in the cryptic, poorly-keyed memo on my Centro. (Note to organizers -- get more wireless connections next time. I didn't bother with my laptop because there was very little chance of getting on the network)

Private LOCKKSS Networks were everywhere. MetaArchive, Arizona State Library and Archives PeDALS, Data-PASS, ETD preservation, and, of course, journal content. It's interesting to see the LOCKSS distributed and self-replicating architecture being used for all types of content.

Distributed and/or replicated storage overall was definitely a trend. iRODS was mentioned in several sessions, I learned more about Dataverse, and I attended a meeting with the FACIT partners.

The packaging and transfer of files between institutions was discussed quite a bit. I was pleased to see the positive reaction to the BagIt package standard that LoC has been working on, which has been put into use with some NDIIPP partners including CDL and Stanford. I was really intrigued with a presentation that Tom Habing did on the ECHO DEPository project. I've seen it presented before, but something really clicked this time when I saw their Hub and Spoke architecture and listened to him talk about packaging between systems and services.

What really stuck with me was something that Micah Altman from Harvard said. He was discussing selection for digital preservation and declared that we need to "select the selectors" in identifying what should be preserved, because we can't save everything. If we identify key researchers and tie preservation to their research, we're assured to capture at least some vital resources. But there are so many disciplines that no one institution can identify what should be preserved, so the corollary need is for many, many institutions to involves themselves in selection and preservation, so there is more preservation coverage for the future. I was glad to hear selection described as a necessary activity.

Wednesday, July 09, 2008

saw a kindle today

The man who sat down next to me on the Metro this morning had a Kindle. While he really just wanted to read, he politely answered a couple of questions and let me hold it. It feels lighter than I expected, and the screen is reasonable clear and high contrast. It's also not as ugly as I thought. Each screen shows about 3 paragraphs of text (better than my Palm), but the entire screen flashes every time to navigate to a new page.

If anyone who has a Kindle is interested, he was reading Ewan McGregor's Long Way Round.

Monday, July 07, 2008

is Google identifying more full text works for GBS?

Barbara Quint at Information Today wonders -- Is Google Book Search Targeting More Books for Public Domain?

Pretty much no U.S. library will make post-1922, probably in-copyright digitized material from their collections available on the open web without varying levels of risk assessment that includes a review of copyright renewal records. Now that Google has developed its own copyright renewal data, will it use that data to identify works that should be in the public domain? And will they make those works available as full-text in Google Book Search? And will they share their research results with the rest of the community so we can free our digitized copies, too?

Wednesday, July 02, 2008

Mbooks Collections

The BLT announced a new feature in MBooks at Michigan -- Collections Pages.

I love that I can browse collections that others have made public -- Perry has a great Gothic Literature collection going (no public domain Castle of Otranto yet?) -- and that you can view the full collection or limit your view to the full-text items. I also love that I can copy items into my own collections to help populate them. I couldn't find almost any MBooks items in Mirlyn to build my first attempt at a collection. Not too many public domain texts on voodoo are available yet...

PDF now an ISO standard

PDF is now officially an ISO standard, finally joining PDF/A (ISO 19005-1) . More details are offered in the press release from the ISO.

The Portable Document Format (PDF), undeniably one of the most commonly used formats for electronic documents, is now accessible as an ISO International Standard - ISO 32000-1. This move follows a decision by Adobe Systems Incorporated, original developer and copyright owner of the format, to relinquish control to ISO, who is now in charge of publishing the specifications for the current version (1.7) and for updating and developing future versions.
You can read the description of the standard.

blogging from excavations

As reported by The Chronicle for Higher Ed, students in Cotsen Institute Archaeology Field Program at UCLA will be blogging from seven of the school's excavation sites in Albania, Canada, Chile, Ecuador, Panama, Peru, and the U.S.

As a graduate alumna of the UCLA Archaeology department, I am thrilled to see this increased visibility for the program, as well as for the practice of archaeology. I know firsthand how challenging it is to explain what is it you actually do in the field...

Tuesday, June 24, 2008

Ithaka report on sustaining online academic resources

The Ithaka project has released a report called “Sustainability and Revenue Models for Online Academic Resources.” The Chronicle of Higher Education explains the core issues succinctly and bluntly:

"So you got a startup grant to get your digital monograph, e-journal, or wiki up and running. What kind of impact will that nifty new project have, and how will you keep it going once the grant money runs out?"
I've seen many, many digital scholarly projects in varying states of their life cycle. New projects flush with enthusiasm, cash, and vision. Projects where a limited scope is envisioned that grow beyond the scope with no plan or resources identified for scalability. Projects in that panicky stage where the money has just run out, hoping for institutional support for the future. Projects where someone has responsibility for care and feeding as an added side job with no recognition of what it might entail. Projects where no staff remained to keep the content -- even just the links -- current, degrading from a vibrant site to a side note. Projects set up using technologies that become problematic for support after time passes. Projects that take up institutional resources but get little or no use. And institutions stretched to the limit of what they can support, having to make very difficult decisions about what to do with the resources on its servers.

The report presents a lot of discussion about identifying targeted users and user needs, and revenue models. The report is aimed more at sustainability for new initiatives than scholarship, but it's worth reviewing by folks in both communities.

copyright renewal records

The Inside Google Book Search blog announced the availability of U.S. copyright renewal records as an XML file. Google created the set by taking advantage of scanned and keyboarded versions created by the Carnegie Mellon Universal Library Project and Project Gutenberg. Google cleaned up the files for improved parsing, and now they're available for download.

Google thinks that this set is "the best and most comprehensive set of renewal records available today." I am not sure what the difference in data and temporal coverage is between these records and the ones from the U.S. Copyright Office copyright registration database made available through the efforts of DLF and Public.Resource.Org. It is useful that Google has put these out as parsable XML.

Monday, June 23, 2008

Strategies for Sustaining Digital Libraries now in print

I am pleased to announce that the volume Strategies for Sustaining Digital Libraries, edited by Katherine Skinner and Martin Halbert is now in print. I'm pleased because 1), it's a Emory University Digital Library publication, and 2), I have a chapter in the book.

This collection of essays on sustaining digital libraries is a report of early findings from member of the community who have developed ongoing services and collections intended to be sustained over time in ways consistent with the long-held practices of print-based libraries. The essays address a variety of issues in advancing innovative information services to an ongoing programmatic mode of sustaining digital libraries for the long haul.

It's available in print and as a free PDF version, under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 License. For more information, please visit the Emory web site.

Friday, June 20, 2008

one can be too available

I loved this post of O'Reilly Radar: Phone in the Toilet?

Yes, I have a cell phone. I do carry it everywhere because it's a Palm Centro, and without it I would not know where I need to be and when. For the first time in years I am actually keeping the phone turned on all the time because my house in Charlottesville is on the market and one needs to be on call for Realtors. BUT, I don't give the number to too many people and I never use all my minutes. I do not feel the need to call people wherever I am just to chat. It's mind-boggling to me when I witness someone having a phone conversation in a public restroom. And I've seen this lots of times. And there's the woman who rides the same Metro shuttle I do that's planning her wedding: she and her mother have serious issues. I should not know this about a woman whose name I do not even know.

I also do not regularly check my email when at home, and I can go hours without going online (which isn't true of everyone in my household ). And my kitchen is where I cook, not go online, even to check recipes.

As my friend Liz would say: there's no kidney in a cooler, people. Take some time away from constant connectedness.

Tuesday, June 17, 2008

Fedora and DSpace meeting

There was a developer meeting last week with representatives of the DSpace and Fedora communities, and notes from that meeting are available. I am pleased that there was a focus on identifying and solving the needs of the community. I saw a lot of mention of work flows. There was a recognition of the need to get the communities TOGETHER to have discussions about these common needs. There was a brief mention of the upcoming RepoCamp as one such potential venue for discussion.

I also loved the final sentence: "Michele, Sandy, Brad, and Thorny will have a beer together to summarize ideas generated in this meeting and conceive of next steps." That's friendly collaboration at its finest.

Fedora 3.0 Beta 2 released

The second beta release of Fedora 3.0 is now available for testing. This release completes most of the features planned for the general 3.0 release this fall. I'm personally most excited by the new Content Model Architecture (CMA), an integrated structure for persisting and delivering the essential characteristics of digital objects in Fedora, which replaces the previous architecture for binding objects, behaviors and mechanisms. Formal instantiation of content models opens up interesting new avenues for work flows and validation. That this new architecture also allows for easier modification of mechanisms is a huge operational boon.

Read the announcement for more details on the release.