Sunday, June 15, 2008

changed my del.icio.us

For the three dozen people who subscribe and might care, I've changed my public del.icio.us from uva_digital_library to lljohnston. I've turned uva_digital_library over to UVA -- I don't know who might be taking it on. All the links remain at uva_digital_library, but I have also copied them to lljohnston. I may have missed some network connections when I was recreating them.

To vent a bit about one thing, while it was easy to export from the old account into the new account, the default is that all imported links are NOT shared, and there is NO batch sharing function, so I manually marked all 498 as shared. That took a while ...

Friday, June 13, 2008

save the date for RepoCamp!

RepoCamp will be a one-day event where folks who are interested in managing and creating digital repository software and their contents can gather and share ideas. RepoCamp will take place on July 25, 2008, at the Library of Congress in Washington D.C. PLEASE NOTE: Due to space constraints, there will only be room for 30-ish people.

For those of you who are familiar with the concept, this will be a "barcamp," where the sessions are proposed and scheduled by the attendees the day of the event. Having participated in beCamp the past couple of years, I'm very excited to see this opportunity for the repository community.

Visit the wiki for more information: http://barcamp.pbwiki.com/RepoCamp

Wednesday, June 11, 2008

book publisher's manifesto

I've been reading Sara Lloyd's "Book Publisher's Manifesto for the 21st Century."

It's a very interesting essay. These sections stood out to me:

We will need to think much less about products and much more about content; we will need to think of ‘the book’ as a core or base structure but perhaps one with more porous edges than it has had before. We will need to work out how to position the book at the centre of a network rather than how to distribute it to the end of a chain. We will need to recognise that readers are also writers and opinion formers and that those operate online within and across networks. We will need to understand that parts of books reference parts of other books and that now the network of meaning can be woven together digitally in a very real way, between content published and hosted by entirely separate entities. Perhaps most radically, we will have to consider whether a primary focus on text is enough in a world of multimedia mash-ups. In other words, publishers will need to think entirely differently about the very nature of the book and, in parallel, about how to market and sell those ‘books’ in the context of a wired world. Crucially, we will need to work out how we can add value as publishers within a circular, networked environment.
and
Publishers need to provide the tools of interaction and communication around book content and to be active within the digital spaces in which readers can discuss and interact with their content. It will no doubt become standard for digital texts to provide messaging and commenting functions alongside the core text, to enable readers to connect with other readers of the same text and to open up a dialogue with them. Readers are already connecting with each other – through blogs, discussion forums, social book-marking sites, book cataloguing sites and wikis. Publishers need to be at the centre of these digital conversations, driving their development and providing the tools for readers to engage with the text and with each other if they are to remain relevant.
The idea that texts exist as networked content that can be broken down into components that can be recombined with other networked content in a multitude of contexts is a huge focus in digital humanities scholarship. Remixing and recontextualization through mashups isn't just a scholarly activity by any means. Anybody who has bought a song from iTunes and added it to a playlist has taken a single component from a larger whole that was once considered the only possible unit of distribution (an album) and recontextualized it (a personal thematic playlist).

Publishers are finally beginning to understand that the "book as unit" model is no longer the only model for distribution -- in fact, that will soon no longer be the dominant model for any media distribution.

That said, I hope publishers don't throw the baby out with the bathwater in the rush to identify new paradigms for digital distribution and reading on the screen. I still buy and read books. They are still a content unit with meaning. Publishers need to think about how they will continue to distribute books, but in a way that they can be consumed and retained as a whole OR broken down into components for consumption and re-use.

Another major topic in the manifesto is one that I have never given any conscious consideration -- do book buyers give any thought to publisher brands? Her answer is no, they do not. I sometimes take note of the publisher or line -- Vintage Crime for example -- because I have come to associate their line with titles that I have enjoyed in the past, so I'm more likely to look at one of their books on the shelf now and in the future. I have a lot of books from Tuttle and Kodansha because they publish Japanese fiction in translation. Anyone who has shopped at the New England Mobile Book Fair in Newton Highlands, Massachusetts, knows they arrange their stock by _publisher_, so you'd better have that noted in your WTB lists when you go there. But why else would anyone ever give any thought to the publisher?

My recognition of Vintage Crime, Kodansha, and Tuttle is proof that one of Ms. Lloyd's suggestions for publishers -- deep genre niche brands -- is not far off the mark. And she suggests that publishers need to market directly to their consumers rather than letting the next link in the distribution chain, e.g., booksellers, become the recognized brand through their marketing efforts. Publishers need to learn the basics of digital promotion in addition to digital distribution. I can think of a lot of book promotion experts I know: Bella Stander, Kevin Smokler, and M. J. Rose -- who would agree.

Friday, June 06, 2008

BagIt

A press release went out this week on digitalpreservation.gov about the BagIt format specification. BagIt is a lightweight specification for the description of data packages meant for transfer between institutions. The need for such a standard was initially identified in working with NDIIPP partners such as CDL (John Kunze of CDL played a major role in the format development and testing and is one of the principal authors). Other partners have expressed interest, and we are moving forward with a prototype submission web app that will take advantage of BagIt.

My colleague Ed Summers has posted a fantastic overview of the specification to which I can add nothing except reinforcing the kudos deserved by everyone involved.

There is an official Internet-Draft and comments are welcome.

Wednesday, June 04, 2008

edupunk

I've been chewing on an article in the Chronicle of Higher Ed about the "edupunk" movement that just might be coalescing. The term was apparently discussed publicly for the first time only 10 days ago by Jim Groom at the University of Mary Washington (scroll to the bottom to see an extensive list of trackback links to discussions), was covered less than a week later in the above article, and now has a brief wikipedia article. From Jim Groom's blog post:

" ... in my mind the technology is often the means through which the communal acts are traced, recorded, and archived. The learning happens not as a by-product of the technology, it is, or rather should be, the Raison d’être of the technology. The teaching and thinking happen within the medium of texts, videos, film, images, art, conversation, game playing, computers, etc. Technology may provide new ways of delivering and accessing this information, and mark the basis of many a medium, but the idea of a community and its culture is what makes any technology meaningful and relevant.

This is why the idea that “it is about the technology” makes BlackBoard 8 so troubling to me. If it is about the technology, then capital can quickly recognize this fact and co-opt all the hard work by so many to move outside of the taylorized vision of educational technology grafted upon our institutions. If the technology is what is important, than what do we say if a faculty member or student notes that Bb can do what del.icio.us can, or can “mash up” YouTube, Flickr, and Google Earth maps like WPMu, or can make content at long last open, or has a slick AJAX interface, then we what what can we say about the technology?

...if we reduce the conversation to technology, and not really think hard about technology as an instantiation of capital’s will to power, than anything resembling an EdTech movement towards a vision of liberation and relevance is lost. For within those ideas is not a technology, but a group of people, who argue, disagree, and bicker, but also believe that education is fundamentally about the exchange of ideas and possibilities of thinking the world anew again and again, it is not about a corporate mandate to compete—however inanely or nefariously—for market share and/or power. I don’t believe in technology, I believe in people. And that’s why I don’t think our struggle is over the future of technology, it is over the struggle for the future of our culture that is assailed from all corners by the vultures of capital. Corporations are selling us back our ideas, innovations, and visions for an exorbitant price. I want them all back, and I want them now!
To quote Leslie Madsen Brooks blog, "In short, edupunk is student-centered, resourceful, teacher- or community-created rather than corporate-sourced, and underwritten by a progressive political stance." And check out D'Arcy Norman's edupunk heroes.

What I found most interesting -- in addition to its viral spread -- were the comments on the Chron article and Jim Groom's original post. Issues of ownership of content, the closed nature of some learning management systems versus open source, tools such as Edusim or ScholarPress, the value of common tools for supportability (I've heard that so often to stop oh-so-scary innovation), and whether Edupunk itself is technology for technology's sake.

The critiques are equally interesting. That edupunks need to grow up. That an edupunk movement isn't the right answer. And, in a lighter vein, that appropriating the word "punk" and its associations isn't, well, appropriate.

While Jim Groom might like to deny it, it's a full-fledged meme. It's official when it bleeds into another discipline -- Libpunk.

Regardless of whether the term is a good one or not, whether this is a movement or not, or how an organization is going support a thousand educational technology flowers blooming, I am glad to see a conversation about new pedagogies, innovation in educational technologies, open sourcing of tools and content, and the role of a Maker/DIY community in higher ed.

If it's anything at all, this blog post speaks to me perhaps the most: Edupunk is a mindset, not a technology movement.

Thursday, May 29, 2008

digital art and museums

WireTap Magazine has a great interview with Richard Rinehart from the Berkeley Art Museum/Pacific Film Archive that considers a number of topics, including why museums should care about and collect digital art, and why museums should make digital art available for remix. Richard has done some amazing work in promoting the curation and preservation of digital art, and makes amazing art himself. I've known Richard about 15 years.

I also want to make a plug for Berkeley Big Bang 08 (one of Richard's projects) on June 1-3, and for 01SJ on June 4-8 in San Jose (Steve Dietz, another museum colleague who is well-known for his ground-breaking commissioning and curation of digital art at the Walker Art Center, is the Artistic Director for the festival). Anyone in the Bay Area with any interest in digital art and new media should attend one of both of these events next week.

Tuesday, May 27, 2008

more on the end of Windows Live Search Books

A roundup of commentary:

ars technica (includes comments from Brewster Kahle)

shimenawa/Peter Brantley

TeleRead

New York Times

Friday, May 23, 2008

first sale doctrine and software

There's nothing that I can add to this ars technica posting:

A federal district judge in Washington State handed down an important decision this week on shrink-wrap license agreements and the First Sale Doctrine. The case concerned an eBay merchant named Timothy Vernor who has repeatedly locked horns with Autodesk over the sale of used copies of its software. Autodesk argued that it only licenses copies of its software, rather than selling them, and that therefore any resale of the software constitutes copyright infringement.

But Judge Richard A. Jones rejected that argument, holding that Vernor is entitled to sell used copies of Autodesk's software regardless of any licensing agreement that might have bound the software's previous owners. Jones relied on the First Sale Doctrine, which ensures the right to re-sell used copies of copyrighted works. It is the principle that makes libraries and used book stores possible. The First Sale Doctrine was first articulated by the Supreme Court in 1908 and has since been codified into statute.

Read the entire post. The key to the ruling is that the judge decided that the Autodesk software was sold, not licensed, and therefore the First Sale Doctrine applied.

There's also an excellent post on William Patry's blog, with some very lively and extensive discussion that is itself worth reading.

Windows Live Search Books going away

Peter Brantley forwarded a Microsoft message from the Live Search blog to the DLF community. Excerpt:

"Today we informed our partners that we are ending the Live Search Books and Live Search Academic projects and that both sites will be taken down next week. Books and scholarly publications will continue to be integrated into our Search results, but not through separate indexes.

"This also means that we are winding down our digitization initiatives, including our library scanning and our in-copyright book programs. We recognize that this decision comes as disappointing news to our partners, the publishing and academic communities, and Live Search users.
...

"Based on our experience, we foresee that the best way for a search engine to make book content available will be by crawling content repositories created by book publishers and libraries."
I never really used Live Search Books, but I have colleagues who said very good things about it, some of whom thought it was better than Google Books in terms of search success and consistency and delivery UI. I hope that the output of the digitization by the partners can be re-purposed into other services. Their message encourages partners to continue working with the Internet Archive, so I feel hopeful that the equipment and processes put into place for this project will also continue to produce output even without Microsoft's involvement.

Thursday, May 22, 2008

more on oclc and google

Here's an article in Information Today on the OCLC/Google agreement.

This focuses more on the addition of Google Book links into existing WorldCat records, and the creation of new records for volumes not currently in WorldCat. I wonder if this means adding 856 field links to digital surrogates, or creating separate digital resource records? All OCLC member institutions will be able to add records into their catalogs.

This article doesn't mention one aspect of the agreement -- that the agreement now allows Google Book partners to share records from their catalogs that have an OCLC provenance with Google (or let OCLC do it for them). There has always been a lot of discussion about what rights member institutions had vis-a-vis sharing OCLC-sourced records that represent their holdings, and not everyone agrees with OCLC's assertions of its rights. Given Google's need to know something about the volumes that it's digitizing to provide access, it seems unavoidable that sharing of some OCLC-sourced metadata between Google Book participants and Google has already happened. Now it's a recognized need and activity.

Wednesday, May 21, 2008

oclc and google cooperation

On Monday a press release was issued about cooperation between OCLC and Google. Excerpted:

OCLC and Google Inc. have signed an agreement to exchange data that will facilitate the discovery of library collections through Google search services.

Under terms of the agreement, OCLC member libraries participating in the Google Book Search™ program, which makes the full text of more than one million books searchable, may share their WorldCat-derived MARC records with Google to better facilitate discovery of library collections through Google.

Google will link from Google Book Search to WorldCat.org, which will drive traffic to library OPACs and other library services. Google will share data and links to digitized books with OCLC, which will make it possible for OCLC to represent the digitized collections of OCLC member libraries in WorldCat.

...

WorldCat metadata will be made available to Google directly from OCLC or through member libraries participating in the Google Book Search program.

Google recently released an API that provides links to books in Google Book Search using ISBNs, LCCNs and OCLC numbers. This API allows WorldCat.org users to link to some books that Google has scanned through a “Get It” link. The link works both ways. If a user finds a book in Google Book Search, a link can often be tracked back to local libraries through WorldCat.org.

The new agreement enables OCLC to create MARC records describing the Google digitized books from OCLC member libraries and to link to them. These linking arrangements should help drive more traffic to libraries, both online and in person.

There are a couple of big wins here for different communities.

For WorldCat users, there is direct access to Google Book Search volumes. For Google Book users, there is improved access to physical volumes.

For Google Book participant libraries, there are better mechanisms for getting metadata about collections to Google.

Even more importantly -- and I am being hopeful here and reading something into this that may not be there -- this is a potential way to get representation of volumes digitized as part of the Google project not only into WordCat but into the OCLC/DLF Registry of Digital Masters. For folks unfamiliar with that project, it's a registry of digitized volumes -- which meet certain digitization standards and are publicly accessible -- that can be used as a tool by libraries and users to determine if volumes have already been digitized and are available. It's a slowly growing registry where a devoted group of participants have been working to develop the standards for describing digital masters and the work flows for adding records. This is a service that is poised to become essential.

copyright and course materials

I've been reading up on a case out of the University of Florida where Michael P. Moulton, an associate professor of wildlife ecology and conservation, is part of a legal battle against Einstein’s Notes, a company that sells students study kits and lecture notes.

While his publisher, Faulkner Press, brought the suit, one of the more interesting issues to me what Moulton's claim of copyright on his lectures and that any notes taken by students is an infringement, but an allowable fair use infringement as long as the notes aren't sold. Moulton and Faulkner Press have a strong case for infringement when it comes to material from his published textbook that he uses in teaching. Mouton's claims that his printed lecture study guides are copyrighted absolutely has merit. But I have often heard that faculty cannot claim copyright on their teaching material -- lectures, syllabi -- because those are work-for-hire products that they create as part of their employment at a university, and that the university either holds copyright or at least holds an interest in the copyright. There's also the issue that, if not in a fixed form, like a written lecture, a lecture can be protected but not necessarily copyrighted. The University of Florida cleared his copyright registration, so this has obviously been vetted by their general counsel's office.

I once worked on a project where a dean forbade faculty from making syllabi on course sites publicly accessible because they were considered a work-for-hire. It seems that university policies have very much shifted in the years since I last worked with online course issues.

Site on the suit

Chronicle of Higher Ed Interview

Wired article


ars technica post

Tuesday, May 20, 2008

Larry Lessig on orphan works in the NY Times

Larry Lessig has written an op-ed piece that appears in the New York Times on Congress's consideration of a major reform of copyright law intended to solve the problem of orphan works. he characterizes the current attempt at reform as "both unfair and unwise."

But precisely what must be done by either the “infringer” or the copyright owner seeking to avoid infringement is not specified upfront. The bill instead would have us rely on a class of copyright experts who would advise or be employed by libraries. These experts would encourage copyright infringement by assuring that the costs of infringement are not too great. The bill makes no distinction between old and new works, or between foreign and domestic works. All work, whether old or new, whether created in America or Ukraine, is governed by the same slippery standard.

The proposed change is unfair because since 1978, the law has told creators that there was nothing they needed to do to protect their copyright. Many have relied on that promise. Likewise, the change is unfair to foreign copyright holders, who have little notice of arcane changes in Copyright Office procedures, and who will now find their copyrights vulnerable to willful infringement by Americans.

The change is also unwise, because for all this unfairness, it simply wouldn’t do much good. The uncertain standard of the bill doesn’t offer any efficient opportunity for libraries or archives to make older works available, because the cost of a “diligent effort” is not going to be cheap. The only beneficiaries would be the new class of “diligent effort” searchers who would be a drain on library budgets.

It seems that this reform introduces so many layers of complexity into the determination of copyright status that it would inevitably quash potential use of works because no one could figure out the process or afford the time or expert opinion needs for the determination.

Wednesday, May 14, 2008

who owns rss feed content?

There's a good article in pc world out of Australia by Larry Borsato on the ownership issues around RSS feeds. If a site aggregates RSS feeds from other sites as content without permission from the original content owner, what are the legal or moral issues?

it's all about having options

Andrew Pace wrote an interesting post about his take on Library 2.0: It's the data, stupid. My highly simplified version of his thesis is that Library 2.0 is about control and presentation of data, and how we might give the best access to it.

I think that there are a couple of corollaries to this that libraries have only recently begun to consider and implement. First, there is NO ONE WAY to best provide access, and that providing multiple paths and formats is necessary because we can never imagine what all the potential uses of our data are. Data should be exposed in as many ways as an institution finds sustainable, using appropriate community standards.

The second corollary is that while varied and easy access to data is vital, Library 2.0 is also about the personalization of discovery and use of data. Whether it's applying personal tags or personal filters to improve or focus discovery and retrieval, or applications that can take advantage of Identities and/or other APIs for personal or community-based mashups, it's all about how I might need to discover and use the data versus how Andy might need to work with data, and that those needs will likely be different next month than they are now. Last month I was researching Institutional Repository software solutions. This week I'm reading up on the Django web application framework. Last year I may have wanted to combine geographical and geocoding data with images. I don't know what I'm going to want to do next year.

Library 2.0 is about providing options to users, and removing barriers to innovative use and re-use.

Friday, May 09, 2008

happy birthday copyright

We've all heard that every time someone sings "Happy Birthday" that the copyright holders should be getting paid. Via William Patry's blog, a link to a remarkable article and web site from Professor Robert Brauneis of George Washington Law School that present all the evidence he could find about the copyright status of the song.

application profile for images

In the most recent issue of Ariadne there's an interesting article entitled "Towards an Application Profile for Images" by Mick Eadie. JISC has just completed the first phase of some work drafting an application profile for images aimed primarily at the repository community.

I was happy to see recognition of the complexity of digital image objects and the need to track relationships between images, the sources of those images, the content depicted in the images, etc. They looked at FRBR and at the VRA Core, and ended up creating a conceptual model with the digital file at the center that uses the language of FRBR.

I find myself disagreeing with some of their decisions:

In our model, we have renamed the FRBR Work entity as ‘Image’ for reasons of clarity, mainly to avoid confusion between notions of Work as described traditionally in image cataloguing in the cultural sector (i.e. the physical thing) and abstract Work as described in FRBR. As noted above, image as defined in the IAP is a digital image, in line with our notion of end-users searching repositories for digital images of something. Therefore our conceptual model - while still using the language of FRBR and using the areas of SWAP that have applicability across the text and image domains - places the digital image at its centre.
Architecturally, the JISC model makes sense. You have an image, which is a depiction of a site or a work of art, with manifestations as one or more formats of files. That's what you have to physically manage in a repository.

From a discovery point of view, thought, I'm not fully convinced. Having modeled image objects in a cataloging environment and in a repository architecture and discovery environment, I think that FRBR and the VRA Core have it exactly right as to what users are looking for -- images of something. Researchers are looking for images of Chartres Cathedral or Jeff Koons' "Rabbit" or Leonardo da Vinci's "Vetruvian Man." They are looking for images of those Works. When we designed the content models for images in the UVA Digital Collections Repository, the top level is that sense of work. The top level is a work object, which has child expressions for individual views of that work, which have child manifestations that are the actual media files. We found this relatively easy to manage and to build a discovery interface around. It was also the model already used in the cataloging, so it made for an easier translation from that system to the repository.

I need to review the Images Application Profile (IAP) work in more detail. I know this is aimed more at use for IRs and not for image collections, but I think such a profile can only become ubiquitous if it covers both repository scenarios. Their proposed metadata is well-positioned with its use of MIX elements to manage the image files, but I see less than I'd like to see in descriptive elements that support discovery. For example, I don't see a descriptive content date element. I cannot think of a research use case that wouldn't include searching for, say, images of 18th-century French sculpture. I think they should give more thought to incorporating more from VRA Core and CDWA.

Thursday, May 08, 2008

Kete

I recently become aware of an interesting content management system called Kete, which was developed for the Kete Horowhenua site in New Zealand. It's a repository and discovery service that supports uploading and metadata creation through a web interface. It supports the inclusion of:

  • Images
  • Audio recordings
  • Video recordings
  • Documents
  • URLs for web resources
Metadata can be "locked" so only the creator can edit it, or be open for any to edit. They are collecting some amazing biographical details for their Anzac (veterans) collection through the community. Every "topic" (a subject, a place, a person) can have its own discussion.

Kete Horowhenua was developed with Ruby on Rails, utilizes Zebra z39.50 full text indexing engine developed by IndexData, is fully compatible with Koha, and will be released under a GNU General Public License (GPL). The Kete software is available for download. They are in the process of building a release of the code without the Horowhenua project customizations that can be deployed using a web based wizard that supports customization. They are looking for funding to support this work, as they admit that they underestimated the development needs. They are even accepting PayPal donations to help the work along!

It's an interesting looking site and tool, but I don't know how much is specific to the Horowhenua version and what will be in the generalized version. The browse UI could use a little refinement (I couldn;t figure out how to sort, or if you can sort), but there's a lot of promise here.

Monday, May 05, 2008

beCamp 2008

As tired as I was by the time I got home Saturday night, I am very happy that I made it to beCamp this year. Close to 100 people attended over the 2 days, about half of whom hadn't attended last year. It was great to see so many new faces!

There was a spirited discussion of "Web 3.0" where there was a lot of discussion about advertising. Never having been in the commercial sector, this is something that pretty much never occurs to me. Some folks looked at me as if I were crazy when I suggested that Web 3.0 was about personalization and localization services for users, not about advertising. There was a brief exchange about privacy issues, especially around the topic of providing personal information in exchange for personalized and localized services. I'm not sure that anyone would trade a lot of their personal info in exchange for a free coffee coupon, but others in the room were certain of this.

Josh Malone from the National Radio Astronomy Observatory led a great discussion on large-scale storage needs. There are certain similarities between their needs and the Library of Congress's needs vis-a-vis transfer of files and creation of deliverables for users, but their data is observational and is never updated, while our metadata and media files may be updated. They are considering some systems that I know another institution is using and I need to get folks in touch with each other. I also need to learn a lot more about our storage infrastructure at LC.

Baron Schwartz led a great session on mySQL optimization. He's just joined Percona as a mySQL consultant. Buy his upcoming book! I also realized after the fact that the reason I recognized his name was that, as an undergrad, he worked with my UVA colleague Perry Roland on the Museum Encoding Initiative.

There was a roundtable on data visualization where I heard about a number of tools that I wasn't familiar with, especially two PHP graphing libraries -- jpGraph and sparkline. For anyone on twine, I have a Search and Information Visualization twine started.

Erik Hatcher led a Lucene optimization discussion. Bess asked lots of questions in preparation for getting Blacklight into production!

Steve Stedman gave an informal demo of the Expression Engine content management system. Very interesting, especially its capabilities in resizing graphics on-the-fly.

Those are the sessions I made it to -- check out the wiki for the other discussions on the schedule, and links to slideshare for available presentations. Ruby and Ramaze, High Availability Linux, Adobe Air, Google App engine ...

Friday, May 02, 2008

OhioLINK EAD repository and tools

OhioLINK has launched a Finding Aid Creation Tool and Repository for use by any institution in Ohio, which does not require membership in OhioLINK to use.

The OhioLINK Finding Aid Repository takes advantage of XTF in a clean and simple way. It's not clear what repository is behind it. The EAD FACTORy tool is available for use by Ohio institutions for EAD authoring. It doesn't say when or if the code for the tool will be released.

following up on Bridgeman

Peter Hirtle has an excellent post on the LibraryLaw Blog about a panel presentation at the New York Bar Association called "Who Owns This Image? Art, Access, and the Public Domain after Bridgeman v. Corel." I also highly recommend Rebecca Tushnet's post.

I have encountered many interpretations of Bridgeman over the years. I've heard it used to defend or attack copyright, performance, and use rights policies relating to images. I have been thinking a lot lately about the appropriateness of claiming copyright of images that are captured of book pages while we often take advantage of the lack of copyright-ability of images of 2-D works of art.

I was surprised to note in a comment on William Patry's blog post on the event that the Art Institute of Chicago has an exhibit on copyright in art publishing: "Copyright Law: Publishing Art and the Public Domain." I wish there were more details online about the content. Has anyone seen the exhibit?

Thursday, May 01, 2008

beCamp this weekend

Even though I moved away from Charlottesville last week, I'll be back this weekend -- for beCamp! We'll likely miss part or all of Friday night, but nothing could keep us away on Saturday.

http://barcamp.org/beCamp2008.

When

  • Friday, May 2nd, 5:00PM-9:00PM
  • Saturday, May 3rd, 9:00AM-5:00PM

Where

  • Charlottesville Business Innovation Council (on the Downtown Mall)
  • 501 East Main Street, Charlottesville, Virginia 22902

Tuesday, April 15, 2008

OR08 repository case studies

The Repositories Support Project has released two dozen UK, European, and North American digital repository case studies that were prepared for the Open Repositories 2008 conference. I thought it was a great idea when they solicited these for the conference, and I am very happy that my UVA case history is included.

Blacklight MATC nomination

You can still comment in support of Project Blacklight's nomination for the Mellon Foundation "Mellon Award for Technology Collaboration" (MATC). Anyone who'd like to say something positive about Blacklight, please visit the site and comment:

http://matc.mellon.org/nominate/university-of-virginia/project-blacklight

Saturday, April 12, 2008

Project Blacklight MATC nomination

Project Blacklight is one of the many worthy nominees for the Mellon Foundation "Mellon Award for Technology Collaboration" (MATC). Folks can submit comments in support of nominations, and I am encouraging anyone who'd like to say something positive about Blacklight to please visit the site and comment before 5 PM Eastern Time on Monday April 14.

http://matc.mellon.org/nominate/university-of-virginia/project-blacklight

Friday, April 11, 2008

hiatus from blogging while moving

Today is my last day at the UVA Library. I will very much miss my colleagues, the projects I had the opportunity to work on, and Charlottesville.

I keep telling people that I'll only be 2 1/2 hours away. I think I'll be on speed dial for a while. We move on or about the 21st, and I start my new position at the Library of Congress on the 28th. I plan to be at beCamp the first weekend in May (probably Saturday only), so I'm not making a very clean transition. And I still need to sell the house.

Folks who don't have other methods of contacting me can contact me here if they need to while I'm transitioning. I won't have time to blog, though, because we are way behind on packing...

Tuesday, April 08, 2008

software engineer traits

There was a great post on ReadWriteWeb about the top ten traits of a rock star software engineer.

I have one little quibble. These should be the traits of ANY software engineer, not just a rock star.

I have worked with "programmers" who posses maybe one or two of the ten traits. I've worked with great programmers who have eight or nine of the ten traits (having the time in libraries to continuously refactor code is a bit of a luxury).

This could be a great jumping off point for developing interview questions for the hiring of programmers in any institution.

where did all the stuff in my office come from?

Having been happily ensconced in a series of offices during my 6 years at UVA, I seem to have accumulated way more stuff than I imagined. Even after recycling many years of personal journal subscriptions that the Library didn't need or want, the contents of pretty much every file folder in my file drawers that wasn't important enough to give to someone else, folders and binders full of conference programs and attendee lists, and returning all the books I had checked out except for the one I am reading right now, I still have 4 boxes of stuff. Some books of my own, conference proceedings that I want to keep for now, files for some committees and publications, old backup CDs that I keep out of paranoia, personal desk organizers that work well for me, and way too many gewgaws that I seem to have accumulated. And some framed posters, not all of which are hung on the wall of my current office, but have been in past offices so I've moved them around with me.

Hi, my name is Leslie, and I am a packrat. I thought I was a recovering packrat, but apparently I've been deluding myself.

Friday, April 04, 2008

Open Repositories 08

The timing of my changing jobs this month did not allow me to attend Open Repositories this year. Thankfully, Digital Koans has a great roundup of reports on OR08 sessions.

PREMIS 2.0

The PREMIS Editorial Committee has released PREMIS Data Dictionary for Preservation Metadata, version 2.0, a revision of the May 2005 report. A draft XML schema -- still undergoing a month of review before its final release -- is also available.

Audiovisual Research Collections and Their Preservation

TAPE (Training for Audiovisual Preservation in Europe) has published Audiovisual Research Collections and Their Preservation. This report looks at the requirements for access and re-use, focusing on the potential of digitization for creating distributed content-based archives.

Tuesday, April 01, 2008

1994 again

There is something so charming about the first sites on the web with their white backgrounds, lists of blue underlined links, directories of the internet, and requests to "Click Here."

In honor of the ten year anniversary of the Mozilla project, home.mcom.com, the web site for Mosaic Communications Corporation in 1994 is now back online, as is their earlier site at http://mosaic.mcom.com/. And check out the archive of vintage Netscape browsers!

Read about the trials and tribulations required to make this all live again at http://jwz.livejournal.com/856745.html.

convert paper to ipaper

When I first saw this on BoingBoing, I and others were certain it was an April Fool's joke: "Free bulk-scanning, OCR and web-publishing service launched by Scribd."

Then someone from Scribd posted in the comments that this was, indeed, a real offer with a very badly timed annoucement.

Part of me still suspects it's a prank, but this could be a real, limited-time prototype service offer.

My thoughts about iPaper are still as undecided as when I first wrote about it.

Monday, March 31, 2008

preliminary decision to reject Blackboard patent claims

The Chronicle of Higher Ed has a lengthy article summarizing last Friday's preliminary decision by the U.S. Patent and Trademark Office that rejects all 44 claims Blackboard Inc. made regarding a controversial patent it was granted in 2006.

The patent in dispute involves a course-management system in which a single user with a single log-on could have multiple roles in multiple classes. For example, someone who was a student in one course and a teaching assistant in another could log on once and get different levels of access to all the course materials.

Desire2Learn and its supporters have argued that the patent should not have been granted because similar technology existed in 1999, when Blackboard applied for the patent.

The patent office awarded the patent in 2006, and within months, Blackboard sued Desire2Learn for infringement. However, the patent office, in its re-examination, cited several examples of "prior art," or previously available technology, that was similar to what Blackboard claimed to have invented.

Sunday, March 30, 2008

two new reports on IP issues

The final report of the Section 108 Study Group is available at http://www.section108.gov/docs/Sec108StudyGroupReport.pdf. The report examines the exceptions available in copyright law and discusses changes that may be needed in light of digital technologies.

The RLG Partner Copyright Investigation Summary Report is now available at http://www.oclc.org/programs/publications/2008-01.pdf. This report summarizes interviews conducted with RLG Partner institutions, who shared information about how and why institutions investigate and collect copyright evidence, both for mass digitization projects and for items in special collections.

Friday, March 28, 2008

becoming a digital scholar

Looking for an overview of what it means to be a digital scholar? Read Lisa Spiro's essay "Becoming a 'Digital Scholar,'" the text of a presentation that she gave at the Digital Discovery conference on March 27, 2008.

the law of unintended results

In a recent decision, the U.S. District Court for the Northern District of Georgia in Atlanta rejected Wal-Mart's claim of trademark infringement again an online critic of the company. The court found that Charles Smith’s parody Web sites (www.walocaust.com and www.walqaeda.com) and related merchandise sold through CafePress were protected speech and that a reasonable person would not confuse their use with Wal-Mart’s legitimate trademarks. The court also rejected Wal-Mart’s claim that it has trademark rights in the “smiley-face” that Smith used in one of his parodies. I wonder what the fall-out will be from that portion of the decision?

Read the history of the case.

Wednesday, March 26, 2008

a pair of posts on the semantic web

A interesting post on semantic web patterns from ReadWriteWeb, and a just as interesting response to it from Nodalities. Reading both gives a good view of the state of semantic web development.

Tuesday, March 25, 2008

review of OpenID

My interest in OpenID has recently been piqued, so I am definitely looking forward to the outcome of a JISC OpenID review:

"The primary aim of the project is to produce a report which will allow busy decision-makers to understand OpenID’s security properties well enough, quickly enough, to apply it safely and avoid its potential security pitfalls, based on first establishing by means of a survey a sound understanding of how such decision-makers are likely to proceed in the absence of such guidance. The secondary aims are to develop bridging software that will allow OpenIDs from any source to be used as identities within the production UK (SAML) federation, creating opportunities for early adopters to experiment. We will also demonstrate a library-type service modified to make use of such identities."

document migration

There's a very thoughtful post on the Digital Curation blog about the use of Open Office as a migration tool.

Conversions Plus has personally saved my hide when I needed access to older file formats, but it's not meant as a preservation tool. What tools do people use for conversion? How do they scale?

Friday, March 21, 2008

life on the internet

I was chatting with someone earlier today about identity and privacy, and how I'm thinking a lot about them these days.

In the process of buying the place we are moving to, I was asked for a photocopy of my driver's license. OK, that's not an uncommon request during a financial transaction conducted over email and fax. Then I was asked for a photocopy of my social security card. My what? Why? I couldn't actually find my social security card (note to self -- get a replacement card), so, in lieu of that, I had to sign a release form allowing my lender to confirm my identity with social security. And they needed a photocopy of my passport, too, if I wouldn't mind.

Bruce opened a new business banking account last year, and he was also asked to show his social security card in order to open the account.

Where is the fine line between confirmation of identify (albeit in the face of rampant identity fraud) and privacy?

I saw this posting on ReadWriteWeb about whether hiring officials should look at candidate's social bookmark profiles during the hiring process. That post referenced a Business Week debate on the topic. One of the more interesting things mentioned in passing in the argument against was that identities can be spoofed online, so how does an employer know if they're even looking at a real profile for that person?

Identity and privacy are completely intertwined online.

Circling back to the conversation I was having earlier, we were talking about how much of our lives are archived online -- a dumb message that I posted to a listserv in the mid 1990s is still likely out there in an archive, waiting for someone to wade through hundreds of pages of search results. My personal web page from 1996 is probably in the Wayback Machine. What about now? Is my Facebook profile personal or professional when it includes my movie and music likes but also supports great communication with my close professional colleagues around the world ? My LinkedIn connections include a few close friends from decades past that I have recently reconnected with through that service.

How much do we need to worry about managing our online identities? Either a lot or not at all, and I'm not sure which it is yet.

Friday, March 14, 2008

orphan works

Read the transcription of a statement on orphan works by Marybeth Peters, the Register of Copyrights, to the Subcommittee on Courts, the Internet, and Intellectual Property, Committee on the Judiciary.

Google Books Viewability API

Google has released a new API that supports links to volumes in Google Book Search. Web developers can use the Books Viewability API to quickly find out a book's viewability on Google Book Search and, in an automated fashion, embed a link to that book in Google Book Search on their own sites.

We'd already created our own service for our Virgo catalog for digitized UVA volumes that are available as full-text in GBS. We were thinking about how we'd potentially port that to Blacklight -- now we can look at the API in addition to what we'd done ourselves.

Monday, March 10, 2008

Texas Digital Library Repository

Via DigitalKoans, the Texas Digital Library Repository has launched with content from four of its partner institutions. There's a program that's making some real headway: an IR for ETDs (multi-institutional, even!), journal hosting, and real progress with inter-institutional Shibboleth – on top of what they’re already doing with digitization and online collections.

HP BookPrep

Via ReadWriteWeb, I cam across HP BookPrep, a prototype print-on-demand service for, and this is really a quote, "every book ever published." The work from page images and process them through their own process into PDF "eMasters" for printing. It's an interesting prototype.

The pilot collection is Foodsville, a food and cooking community site. Members can read and purchase cookbooks at the site's free library, where books can be discovered by keyword, by author, or by browsing through tags. Of course, they're not actually "free" -- the print-on-demand costs ranged from $7 to $32 when I browsed through the 141 titles currently available. You don't have to buy the books -- you can read them online. I browsed through Lafcadio Hearn's La Cuisine Creole and found it readable. (I love the recipe title "Delicate Rusks for Convalescents" on page 235)

I'm a fan of Lafcadio Hearn, so I checked Amazon, and found that the same POD version is available for $10.26 versus $13.96 on Foodsville, both listed as marked down from $19.95 There is what looks to be a different POD version available from Amazon as well -- also a facsimile of the 1885 edition -- for $21.24. I'm not sure which I'd choose.

It does not say where they're getting the page images from. I'd really like to know that.

Tuesday, March 04, 2008

anti-counterfeiting course

Via Techdirt, you have to read the article in Inside Higher Ed about a sponsored anti-counterfeiting course at Hunter College where one of the activities was to create a counterfeit blog about events that didn't happen to create a guerrilla marketing message about why counterfeiting is bad.

I can't even begin to describe all the issues I have with this. Let's hope the article and community backlash serve as a cautionary tale for any other organization considering this.

Visible Body

Visible Body launches in beta today, and looks pretty amazing. I could have used this when I took anatomy 25 years ago.

I can't say that I've really been able to experience it, because I waited over 30 minutes for the data to load and gave up. I expect they're getting a LOT of traffic today, so I am cutting them some slack. What is a shame is that it only works on Windows IE because of the choice of plugin.

There's a brief posting on ReadWriteWeb with more discussion.

Monday, March 03, 2008

Warrick

I just came across Warrick, a neat research project that takes advantage of cached web crawls to restore lost web site. Warrick is a utility for reconstructing or recovering a website when a back-up is not available. It searches the Internet Archive, Google, Live Search, and Yahoo for stored pages and images and will save them to your filesystem. It's not guaranteed and it is a research project and not a production service, but it could help when there's no other option. It falls under the general category of research that they call "Lazy Preservation," a phrase that I can see some loving and some hating.

When I followed the link to the about page and I saw that it was a research project at Old Dominion University, I immediately suspected that it was one of Michael Nelson's students, and it was. I briefly blogged Joan Smith's mod-oai work last year. Michael is always working on something interesting, especially his current work on OAI-ORE.

Sunday, March 02, 2008

visualization of statistics as art

I am quite wary of art with a pointed agenda. All art has some sort of agenda, of course, but some works seem to be more overtly political than art.

That said, you really need to look at Chris Jordan's "Running the Numbers" series. I don't know where he gets his numbers, but these are astonishing visualizations of statistics, many having to do with consumerism and wastefulness.

I am also an admirer of his Katrina aftermath series.

Saturday, March 01, 2008

iPaper

I'm not sure what I make of iPaper, described as a potential "YouTube for Documents." Jeff Young gives a succinct description in his Chronicle of Higher Ed article. TeleRead has a post. TechCrunch briefly touches on the business model. ReadWriteWeb has a longer post.

Basically, documents are uploaded (a number of formats are supported) and streamed to a Flash player for page turning. The iPaper player is also available for integration into other sites.

There is a lot of pointing to an astonishingly bad essay that has been posted for its humor value, and the comments are frequently much obscenity laden. How is this useful? One can see the YouTube comparison here dumb things people do in document form rather than video.

I see uploaded offprints of articles, sheet music, and car manuals, test answers for past medical residency exams, among other things, made publicly accessible. Is copyright status being confirmed in any way?

This could be useful -- a place to upload documents with unlimited storage for public access or sale,that supports review and commenting and social bookmarking. But how do you find what's really useful among tips of reducing your golf slice or "secret White House plans"?

Omeka

I've spent a little time looking at the recently announced online exhibition building tool Omeka from The Center for History and New Media at George Mason University.

I am very interested. This is a tool meant for cultural heritage organizations that I think could be extended into a tool for personal digital scholarship. We've talked a lot at UVA about the conceptual similarity between born-digital scholarly projects and exhibition building, and have long considered the possibilities for a virtual exhibition tool. Our Collectus tool was a start for us in that direction. Omeka allows you to create an archive, organize it into collections, and create exhibition-like presentations or illustrated essays with a RSS feed. "Items" in an archive can be compound objects and include multiple files. Its API support extension with plugins, such as support for COiNS, geolocation, or tagging. Hooks to applications such as Collex or the SIMILE Timeline would add some interesting functionality.

The biggest limit that I see is that its in-browser automatic delivery is currently limited to images, although you can include any type of media. This seems like an excellent opportunity for an institution to work with GMU to extend Omeka's capabilities and create richer media experiences.

Thursday, February 28, 2008

blogging will continue

A number of folks have asked if I will keep blogging after I move to the Library of Congress -- I certainly plan to!

For now, though, blogging will be a bit sparse. I'm working on a transition plan at UVA, my house goes on the market tomorrow (I feel like I'm living through an episode of "Designed to Sell"), I close on a new place mid-March, and I feel overall deluged with paperwork (UVA, LC, Realtors, banks, etc). Yesterday I felt so tired and blah that I just stayed home and slept.

Tuesday, February 19, 2008

leaving the university of virginia

Now that all the local announcements have been made, I can publicly announce that I'm leaving the University of Virginia Library for a position at the Library of Congress. I'll be joining the Repository Development unit in the Office of Strategic Initiatives. I am sad to leave the people that I work with and the projects that I have been working on, but I'm excited by the possibilities at LC.

The most daunting part of the transition is selling our house and buying something in a much more expensive market (fear not, we have something in the works, but we will be downsizing). The logistics have me briefly quite stressed. You wouldn't think that moving 2 1/2 hours away would be that complicated ...

Monday, February 18, 2008

MODS tools

We're chatting a lot about MODS at UVA, partly about data sharing via OAI and partly about possibly replacing some local metadata standards.

In the great synchronicity of our community, there were two posts on MODS creation tools today, one from the DIL and one from Peter Binkley. Between the tools mentioned in those posts and the DLF Aquifer MODS profile work, it's getting easier to work with MODS every day.

Friday, February 15, 2008

CrossRef Citation plugin

I'm running a WordPress and CommentPress experiment (very low key right now), and now I must look into the CrossRef Citation plugin.

why bind?

I love an effective visualization, and my colleague Holly has a very effective visualization to make the case for her binding budget.

museum data exchange study

Mellon has funded at great project and OCLC/RLG Programs for a project to prototype data exchange between museums. Here's the press release. I spent many years in the museum community dealing with this, so I am beyond thrilled to see movement in this area. I would like to have seen more vendor systems involved in the pilot, but this is still a very welcome project.

copyright infringement to sell salvaged CDs?

Via techdirt -- OK, there's a case to be made that dumpster diving for CDs (even if abandoned) and selling them is illegal, but it's a puzzle why it might also be copyright infringement.

book ripping

There was an interesting article in the Washington Post on the Atiz BookSnap book scanner and "book ripping" that I recommend.

Harvard open access policy

I am thrilled about the Harvard open access mandate. The text is available in this PDF. Robert Dranton made a strong case in a Harvard Crimson article. Here's the article reporting the approval in the Crimson and a brief article in the NY Times. Peter Suber has an excellent post roundup, another roundup, and a post on responses from Library Academic newswire. I especially recommend Dorothea's comments to all.

there will be news next week

I have not had time to blog of late. A series of brief posts will follow. Watch this spot for the news why.

Thursday, February 07, 2008

OpenID gets some major press with some major names

Via ReadWriteWeb, the OpenID Foundation announced this morning that Google, IBM, Microsoft, VeriSign and Yahoo! have taken seats as the organization's first corporate board members. Support is growing.

Michigan's millionth digitized book

I love the site that the University of Michigan Library has put up in celebration of the millionth digitized book. I especially love that they credit every one of their 436 staff members with contributing to the project and post some of their pictures on flickr.

West Virginia tax maps

Via Techdirt, it's reported that a company that acquired (via a Freedom of Information Act suit) public tax maps representing the state of West Virginia is being sued to stop them from putting the maps online. The basis of the suit?

"While government documents cannot be covered by copyright, apparently some gov't officials feel that preventing their ability to profit off of that public data is illegal."

the semantics of digital curation

I came across a reference to a blog posting entitled "The Digital Curator in Your Future." Not being familiar with the blog or its context, the first thing I thought was that this would likely be a post on digital curation, the evolving discipline around data curation and preservation. The phrase "digital curation" is all the rage these days.

Nope. This was a posting about online brand success, discussing the need for the intellectual curation, where subject experts "separate junk from art" in the online arena, curating site content to present the most relevant and essential content for their community.

I think this was the last place I expected to encounter the word curator -- describing a vital role in niche market success. That said, I think the posting was absolutely right. Whether you call them editors or curators or guides or whatever, a site aimed at a certain community or representing a brand requires selection/filtering/editing to keep focus, as long as the site also provides some mechanism for community involvement and isn't just a top-down kind of information feed.

The comments were just as enlightening. Some commented that many sites have this sort of human curation. Some commented that they use digg or delicious as an automated curation tool of sorts to locate what others have already identified as being of value (even though the posting explicitly posited that digg is not curation, it's aggregation). Some argued that it has to be expert and human to be effective curation. An interesting discussion.

Tuesday, February 05, 2008

text-mining project for historical scholarship

Many congratulations to the Center for History and New Media for receiving NEH funding for a two-year study of the potential of text-mining tools for historical scholarship, entitled “Scholarship in the Age of Abundance: Enhancing Historical Research With Text-Mining and Analysis Tools.”

I have been in a couple of conversation recently about related topics: What are the differences in practice between literary and historical etext analysis projects? I understand that there's a diverse group having some interesting discussions about that topic to identify needs with an aim to seek funding to support a prototype service project.

What do libraries need to provide in terms of collections and services for scholars working with etext resources for their research? I'm hearing that our service model should be more about helping them learn to create their own resources and tools instead of providing all content and tools for them. This doesn't mean that we shouldn't digitize our rare collections and make them available for use and analysis, but that we need to finally do away with our old service model where we did all the work for the scholars. The "teach them to fish ..." metaphor. Of course it's sensible, but it's a real switch in terms of staff activities, requires some different skills, and requires faculty and folks on their project teams to take on work that they used to rely on us to do. For some in the Library and on the projects that's like shifting course for an oil tanker -- it takes some time and it's not precision steering. Luckily, the overwhelming majority are embracing this.

Monday, February 04, 2008

privacy and Facebook

A UVA undergraduate research team in computer science has shared their findings about privacy issues and Facebook applications. It's so worth reading that the Chronicle of Higher Ed picked up the story.

Sunday, February 03, 2008

twine making the news

Back in October I wrote about Twine. I sent in a request to participate in the beta but didn't get selected. The New York Times has an article based on the experience of one of the beta testers -- Sarah Miller, a librarian at Illinois Wesleyan University. Reading the article, I still want to try it.

Monday, January 28, 2008

in other kindle news

A library cannot loan a Kindle without it being a violation of the Terms of Service. It also cannot keep patrons who borrow one from adding their own content.

ereaders

There was an interesting article about the Amazon Kindle in the New York Times -- Freed From the Page, but a Book Nonetheless. The article rightly addresses two of the factors that are key to the success of portable ebook readers -- the quality of the display and usability of the interface, and the physical portability of the unit itself. In those areas, the Kindle doesn't do too badly. Stephen King wrote a column for Entertainment Weekly where he reported on his first use of a Kindle -- he found that he was able to enjoy the book he was reading and the fact that he was using an ebook reader faded into the background.

The article also reports on Steve Jobs' attack on the Kindle not as a device, but as a tool for an untenable market niche because "people don’t read anymore." The article countered the percentage that he presented with some data from a survey conducted in August 2007 by Ipsos Public Affairs for The Associated Press.

27 percent of Americans had not read a book in the previous year. Not as bad as Mr. Jobs’s figure, but dismaying to be sure. Happily, however, the same share — 27 percent — read 15 or more books.

In fact, when we exclude Americans who had not read a single book in that year, the average number of books read was 20, raised by the 8 percent who read 51 books or more. In other words, a sizable minority does not read, but the overall distribution is balanced somewhat by those who read a lot.

Both sides are trying to make a point using statistics: Jobs' figure is high, and the NYT is getting a favorable average by excluding an entire class of respondents. But it's the wrong point. Neither article really dealt with the issues surrounding the content -- availability, formats, interoperability, and DRM. Those are what's really crippling the ebook reader market, not whether people read any more. Make the content as openly available as possible, in as many formats as possible, and transferable between desktops and devices and between people, _and_ do away with the crippling DRM, and then there will be a thriving ebook and ereader market.

already a blacklight 0.2

Bess has posted that 0.2 is out the door.

Within 24 hours of releasing Blacklight 0.1 she received a patch fixing a problem in one of the config files and augmenting the installation README. 0.2 is 0.1 but with the config file fixed, some better installation instructions, and exported from svn (as opposed to checked out).

She's ported the subversion repository over to rubyforge. The source is browseable here: http://blacklight.rubyforge.org/svn/ and you can see all your svn options here: http://rubyforge.org/scm/?group_id=5235

NIH mandate questions

T. Scott Plutchak has a great post on questions that should be addressed at the January ARL meeting on the NIH public access policy. I understand that this is an invitation-only meeting, but I wish I could find some details about this meeting, whoi's going to attend, and what is going to be discussed. If it's on the ARL web site I can't find it. If someone can point me toward more info, I'd appreciate it. Since we don't yet have an IR in place, we have questions about compliance.

texts read into Congressional record

There's an interesting post from Peter Brantley, which contains an opinion (not a legal opinion, but a knowledgeable opinion nonetheless) on the copyright status of texts read into the Congressional record. It's an interesting question -- tracking of rights when copyrighted texts are excerpted in a public document.

Friday, January 25, 2008

Cal Berkeley pilot to subsidize open access fees

Via OA Librarian, the U.C. Berkeley Research Impact Initiative (BRII) is launching an 18-month pilot program, to subsidize, in various degrees, fees charged to authors who select open access or paid access publication. The pilot will also yield data that can be used to gauge faculty interest in -- as well as the budgetary impacts of -- these new modes of scholarly communication on the Berkeley campus.

TheAtlantic.com

I was initially very excited by the announcement on BoingBoing that The Atlantic had opened its archive. I read the Editor's Note describing the decision. I followed the link to start my exploration.

It's a little misleading. The _site_ is now open to all. They have "Unbound" (web only) content and full issues back to 1995 open. But their other free content seems to be selected material. Some thing that I looked for are there (Vannevar Bush's "As We May Think"). I've seen other people's examples of known article with no results. There is still a "Premium Archive" for the full content going back to 1857.

Still, there's some great content. Try the search at http://www.theatlantic.com/a/search.mhtml

Thursday, January 24, 2008

Blacklight 0.1 source is out there

Check out Bess's post for more details -- Blacklight 0.1 source is now available at http://rubyforge.org/projects/blacklight/.

Free Government Information

When I first scanned over this email message in my inbox I thought about those old late night tv ads enjoining us to write NOW! for free government publications that could help YOU get money from the US Government!

This project is actually really intriguing:

Free Government Information is investigating the usefulness of tagging government documents that do not receive traditional cataloging and needs your help! We've posted 32 documents that the Government Printing Office (GPO) harvested from the EPA web site and posted them to the Internet Archive. Over the next three months, we'd like to see as many people as possible tag and describe these documents using the del.icio.us bookmarking service. For a full project description and instructions on how to participate, please visit http://freegovinfo.info/epatagging. We'd like to thank GPO for posting a sample of their harvested EPA documents that made this project possible.
They are going to run the project for three months and collate the following data: How many people participated in the project, how many documents were tagged, how many documents were described, and the average number of tags per document. I'm looking forward to seeing the outcome, especially alongside the LC Flickr project.

Monday, January 21, 2008

Peter Brantley on Google Books

There is a very brief interview with Peter Brantley, Director of DLF, in the January 25 issue of The Chronicle of Higher Education (subscription required). In this interview he answers a series of questions that identify his concerns about the Google Book project.

He identifies issues about the resources required of the libraries that participate, and that the cost can keep the library from other activities. He hopes that "a court [will] determine once and for all that it is fair use to digitize a copyrighted work and make a snippet of it publicly available." He notes that Google is creating content for its own use that can potentially be made into some sort of commercial service, and that its processes and standards are opaque to the user and the community. He expresses concern that the fuzzy status of orphan works put them into a category where someone will end up making money off them when he thinks that shouldn't be allowed. There's nothing there I'd argue with.

He also points out that the quality of the scans is not consistent. That's followed by this:

Q: Shouldn't Google be commended for helping to preserve library books?

A. The company is not preserving books. It is creating an archive for Google's own purposes.

I'm of two minds when it comes to responding to this. I do not know why so many people assume that this project is preserving the volumes -- Google does not say that; they say that they are creating access copies. The participant institutions know this, and we knew this went we entered into the project. More people need to keep pointing this out.

The other side of this is that IF WE DO WANT "preservation quality" digital objects that represent these books (argue amongst yourselves about how to define what that means or if such objects can actually exist), that's a an additional project with additional costs and time and wear and tear on the volumes to digitize them. And that is unfortunate.

I'm going to end up looking like a major Google apologist because I keep pointing out when folks have things wrong about the project. I just want to set the record straight when I can, so our participation can be judged (pro or con) based on accurate information.

American Literatures

American Literatures is a Mellon Foundation-funded project where five university presses—NYU, Fordham, Rutgers, Temple, and Virginia—have established an initiative designed to create new opportunities for publication in humanistic scholarship. The most innovative aspect of the program will be the establishment of a shared, centralized, external editorial service dedicated solely to managing the production of books in the initiative. This service will handle all copyediting, design, layout, and typesetting costs, and manage each title through to the point where it is ready for printing. The initiative has a web site and has announced on its "about page" (scroll down) in which areas each press is soliciting submissions.

I look forward to watching how this progresses for the UVA Press.

Saturday, January 19, 2008

more openID support news

After testing OpenIDs as logins to Blogger in a prototypet program in November 2007, Google has become an OpenID provider. Effective immediately, Blogger users can use their blogs URL as an OpenID login, after toggling the option via the draft.blogger.com admin menu. Read the posting on TechCrunch.

Friday, January 18, 2008

smARThistory

There's an interesting project out of the Fashion Institute of Technology to develop an open access online art history text using interactive technologies -- smARThistory. It's being developed under a Creative Commons Attribution-Noncommercial-Share Alike 3.0 License.

Google to Host Terabytes of Open-Source Science Data

Wired Science has a post on Google's plan to host open-source scientific datasets through a project called Palimpset. The post includes links to a number of documents and other posts.

Thursday, January 17, 2008

support for OpenID

Yahoo has thrown some support behind OpenID when used on conjunction with Yahoo IDs. ReadWriteWeb has a thorough analysis.

great headline

From an article in the Times of London on a new genus of Madagascar palm:

"Picnicking family stumbles on a suicidal monster palm tree."

THATcamp

Dan Cohen reports that the Center for History and New Media at George Mason University is having an "unconference" called THATCamp: The Humanities and Technology Camp from May 31 to June 1, 2008.

An unconference is a participant-generated collaborative program. I was involved in beCamp last year, and I am a big fan or the model. It's a great fit for digital humanities collaborations.

LC and flickr

One couldn't turn around online today without seeing a post on the Library of Congress flickr experiment today: Read Write Web, Shifted Librarian, eFoundations, Catalogablog, Digital Koans, etc, and of course on the LOC blog.

In case you've been under a rock, LC has put up 3,000 images of photographs from their Prints and Photographs division for which no rights restrictions are known to exist, and are asking folks to tag them to help in enriching the metadata. The pilot project is in a new area of flickr called "The Commons." LC has a GREAT FAQ on participation.

Three things stand out for me:

They have made use of a new rights statement "no known copyright restrictions," which they may or may not make available for other images uploaded to flickr.

They are interested in hearing from other cultural heritage institutions to gauge interest in other projects becoming part of The Commons.

This is a pilot, and they do not know how long they might keep the images on flickr. So, review and tag these while you can. They have other candidate collections in mind.

Tuesday, January 15, 2008

D-Lib article on CRATE

The January/February 2008 issue of D-Lib includes an article by Joan Smith and Michael Nelson on their proposed CRATE mod-oai utility to produce preservation metadata for web resources. I saw this work presented at Open Repositories 2007 and I thought it was interesting then. It's a good article on an interesting proposal.

Archivists' Toolkit 1.1 released

Archivists' Toolkit -- a tool that has been designed to fill a huge need in the archival EAD community -- has released version 1.1. On the surface it seems to do something so simple, but often those "simple tasks" are deceptively so, and are those that can benefit most from a tool like this that can be shared across the community.

everyone needs continuing ed in copyright

Both Open Access News and the Chronicle of Higher Ed reported on a presentation at ALA Midwinter by Trisha Davis, a librarian at Ohio State University that states that faculty frequently have no idea that they've signed away copyright and/or other rights in their contracts with publishers. This was presented in the context of getting content into the IR at Ohio State.

This is no surprise to many of us. I first really gave this issue some thought after hearing a very good presentation by Karla Hahn that in part covered a study of publishing contracts at the 2004 University of Maryland Center for Intellectual Property conference (part of Panel 2).

Not having attending Midwinter to hear the talk -- so I have no idea if this was mentioned or not -- it's NOT JUST FACULTY that need refreshers in intellectual property. I have heard some librarians from many institutions espouse some really unsound beliefs about copyright and fair use. I also feel quite strongly that "refresher" is neither the right word nor the right idea -- librarians and faculty need continuing education in intellectual property issues. The landscape is constantly changing, laws are changing, and online publishing is by nature international in its distribution.

The UVA Library's Scholars' Lab is holding a series of Copyright 101 presentations that I know are making a difference, because I've attended and heard the questions that faculty have asked and had answered. Some of us now want to introduce a series for library staff as well.

LCSH makes the Washington Post

In 2006 there a decision was made at the Library of Congress to do away with the subject headings for Scottish literature, instead suggesting "English literature" headings. There was more than a small kerfuffle over this when it was belatedly reported in the Times of London and the BBC last December.

Today the Washington Post actually took note of this, publishing an article on the reversal of the decision. The Times has also taken note. In fact, it's all over the news this morning. Interesting that a standards terminology issue has become so political.

Friday, January 11, 2008

Report on LC/SDSC Data Transfer and Storage Tests

The Library of Congress has released the report Data Center for Library of Congress Digital Holdings: A Pilot Project; Final Report. I've only had the chance to glance through it, but it looks to be a very clear and straightforward report on technical, procedural, and cultural issues encountered in this digital image collection data transfer experiment, and how they resolved those issues (or not).

http://www.digitalpreservation.gov/pdf/SDSC_LC_data-storage_report_2.pdf

chandler

There was an interesting convergence of posts on Chandler today.

The first post that I read was from Cory Doctorow on BoingBoing, talking about his use of Chandler over the past few months, how he likes it, and how it's improved over time. Then I saw, via TechDirt, that Mitch Kapor has backed out of the project, and there's a press release about the Chandler's future.

I never really looked into Chandler, but I know that folks at some places (UC Berkeley for one) were at one time very engaged with its possibilities. As someone who had been recently required to switch to Outlook and Exchange, I wish this had made more progress and gotten more traction.

Thursday, January 10, 2008

print-on-demand from open access books

PublicDomainReprints.org is offering an experimental non-commercial service that allows users to convert digital public domain books in the Internet Archive, Google Book Search, or the Universal Digital Library to print-on-demand using the service Lulu.com. You pay for the service through Lulu.

For example, you paste in a URL for a public domain volume from Google Book Search, and the process takes advantage of the existing PDF for production. Apparently the Internet Archive PDFs can't be used, so the process takes advantage of the dejavu files instead.

There's a blog post mini-interview with the founder, who has his own blog.

Nature archive online

The full text of all Nature articles back to the first issue in 1869 are now online. I wish it weren't by subscription only, but we have access at UVA and it is remarkably cool. That first issue includes a book review for M. Madsen's Antiquités préhistoriques du Danemarck on Danish Iron Age sites that makes me want to look for a copy so I can see the beautifully described illustrations.

Final Version of the LC Working Group on the Future of Bibliographic Control Report Released

The final version of the highly anticipated report from the Library of Congress Working Group on the Future of Bibliographic Control was released today -- "On the Record: Report of The Library of Congress Working Group on the Future of Bibliographic Control." I don't know when I'll have a chance to look at it, but I understand that changes were incorporated from comments received during the review period.

Wednesday, January 09, 2008

copyright and images online

There was an interesting article today in the Washington Post about folks who have posted images on flickr and in blogs, only to discover their images co-opted for commercial use. (I don't know how long the article will be publicly available)

Cases included a woman who found an image of her pub in a Santa suit used on a Fox NFL broadcast, and a Dallas teenager who found an image of herself used in a commercial ad campaign for Virgin Mobile in Australia. There have apparently been a number of cases involving the online parenting magazine Babble, which they repeatedly blame on inexperienced staff. In one case, a man who asked a Microsoft blog to remove a link to his image and got no response then replaced the image with a famous pornographic image, which got immediate action.

The article briefly describes fair use, and rightly mentions that a rights holder can give away or assign rights use. But there's this quote from the article: "Clearly, the only way to really make sure your photos on the Internet don't get splashed around is not to put them up there to begin with." (their emphasis.) And there's this description of creative Commons: "... Creative Commons, a nonprofit that licenses photos for Flickr."

Given that they spoke with Larry Lessig and quoted him in the article, could they have done the five minutes of research needed to correctly characterize CC? Could they have bothered to describe how best to declare one's rights? Advising folks to be vigilant about their rights is a good thing, but let's not advise withholding content for fear of misuse.

Wednesday, January 02, 2008

public domain day

John Mark Ockerbloom blogged a list of authors whose works in theory went into the public domain yesterday., at least in some countries. With the “life plus 70 years” and "life plus 50 years" term rules, authors who died in 1937 or in 1957 should have their works pass into the public domain in many countries. John goes on to decry that nothing like that will happen in the U.S. until 2019.

CopyrightWarch.ca also has a greet entry with an extensive list. It's a list full of fabulous names.

Tuesday, January 01, 2008

virtual 1815 Jefferson library

Working at the University of Virginia, one is never far removed from Mr. Jefferson's legacy, be it architectural or intellectual or in manuscript form.

LibraryThing announced today that the project to catalog Jefferson's 1815 library (the volumes sold to the Library of Congress to replace its burned collection) is complete, the first collection processed by the volunteer "I See Dead People['s Books]" group. Not only did they catalog them using the Gilreath/Wilson list, they used his classification scheme for tagging AND they excerpted book reviews from his letters, using the online page image versions from LC. And you can take advantage of LibraryThing social tools, like checking what books you have in common, view tag clouds, stats, etc. This is a work of digital scholarship, providing access and interaction and analysis, built using a tool that many do not take yet seriously.

I notice that one of the other projects is Tupac. I wonder where that authoritative list is from?