At a fundamental level, Weinberger focuses on the decentralization enabled through technology integration. In comparing LC's original cataloging efforts with Web 2.0 systems, the author concludes that:
The Library of Congress' carefully engineered, highly evolved processes for ordering information simply won't work in the new world of digital information. Not only is there too much information moving too rapidly, there are no centralized classification experts in charge of the new digital world we're rapidly creating for ourselves....(page 16)Much of what I do in supporting library services involves finding the best strategy for intelligently embracing this sea change. One strategy that my library organization has taken has been to move away from MARC in describing locally-created digital collections, to descriptive schemes such as Dublin Core and EAD. Another strategy has been to work with a new set of products, outside of the traditional integrated library system, for supporting online digital collections: OCLC CONTENTdm (licensed in 2000, for online digital collections); IXIASOFT TEXTML (licensed in 2003, for EAD XML delivery); and the open source DSpace application (implemented in 2005 for institutional repository use). All three of these applications have open APIs supporting local interface and application development. This ability to perform local development and customization in a secure way is and will be critical to my organization in serving its users in the decentralized environment described by Weinberger.
Weinberger's comparison of the Dewey Decimal and Amazon classification systems again illustrates the larger-than-appreciated differences between the analog and digital worlds that are at the center of Everything is Miscellaneous. For example, Amazon employs collaborative filtering to support its "customers who bought this item also brought" feature; as the author notes, recommended titles typically "would be scattered across the shelves in a Dewey library." Importantly, Weinberger continues, "Amazon brings [recommended titles] together not because they are on the same topic but because of a statistical analysis of customers' buying patterns" (page 60). As the author concludes this section, he notes the real problem: the methodology employed in analog-developed classifications like the Dewey system "unnecessarily inhibits" access systems in the digital realm (page 63).
There's an interesting discussion on faceted searching, including its history and how it works in theory, in the chapter "Lumps and Splits." There are several references to Endeca, whose faceted search product is used to support NCSU's online catalog.
I found the author's discussion of the concept "everything is metadata" (page 104) extremely relevant my archival repository work. To review, Weinberger notes the difference between analog metadata systems (for example, a catalog card describing a book) and digital metadata systems, the latter of which can use "every word in a book" as metadata. One major challenge that the NWDA is investigating is integrating finding aid and digital content. Two systems that provide support for EAD and digital object integration are OCLC CONTENTdm and the Archon open code software developed at UIUC. In both cases, significant metadata is lost when importing EAD to the database:
CONTENTdm: This document describes CONTENTdm's EAD support. Nothing beyond the first 128,000 characters of a document is imported, and the import is limited to seven vendor-selected EAD elements.
Archon: The EAD XML content is shredded upon import to map to the table and columnar structure of the relational database, resulting in the loss of information including: element and attribute ordering, comments, and processing instructions.
I believe the functionality of CONTENTdm and Archon is a reflection of the fact that they are both early generation integration products. In fact, Archon's integration functionality is quite slick in many respects; it's the use of the RDBMS database that compromises the tool. The current NWDA solution, TEXTML, handles the "everything is metadata" need extremely well, with all documents maintained as is in this native XML database. The down side: any EAD- digital object integration must be developed (either locally or by a vendor for a fee), in the same way that the NWDA's search/retrieval/presentation application was locally developed.
A final point that I drew from the book is the importance of published APIs (though the term "API" or "application programming interface" is never used). Weinberger uses the examples of Google Maps and Flickr (pages 227-8) as services that users have added significant value to through the creation of mash-ups and other API work. I was in a discussion at code4lib 2006 with technologists and automation librarians, and the top concern voiced at this discussion was the lack of open APIs in library integrated library systems. It seems like a reasonable demand, but ILS vendors, from my experience, are having trouble embracing it.
I finished Everything is Miscellaneous over the weekend, and I'd highly recommend this book. It emphasizes the revolutionary nature of the information technologies that have been created in the past half decade and the need for out of the box thinking in service delivery in such an era.

No comments:
Post a Comment