Tuesday, August 22, 2006

OAI-PMH data provider analysis

Next week, I'll be participating in a call with colleagues from the Northwest Digital Archives and the California Digital Library regarding the recently-developed NWDA OAI data provider support.

I coded (in VB.NET/ASP.NET, using the IXIASOFT TEXTML API) only the early stages of the NWDA OAI-PMH data provider software. My project colleague did the bulk of the coding and creatively implemented OAI support for the NWDA site. I decided to take a closer look at the specification in order to better understand the implementation.

Sets: As noted in the specification, "a set is an optional construct for grouping items for the purpose of selective harvesting." Given that the NWDA is a multi-institutional project, the most obvious set criterion is [by] institution. But going beyond that, NWDA browsing terms defining places, subjects, and material types can be used to as a basis for harvesting as well.

The spec notes that "it is expected that individual communities may formulate well-defined set configurations with perhaps a controlled vocabulary for setNames and setSpec...." For the NWDA implementation, the following structure is used, which closely follows the specification's suggested model:

setSpec: institution:nte
setName: Washington State University

setSpec: subject:irrigation
SetName: Irrigation

Here's the ListSets verb in action: http://nwda-db.wsulibs.wsu.edu/oaiserver?verb=ListSets

Flow control: Flow control is implemented as recommended in the specification for the ListIdentifiers and ListRecords verbs. When the expanded list of controlled terms is used for set definitions, flow control will be implemented for ListSets as well.

Here's flow control in action; the resumptionToken can be viewed at the bottom of the document: http://nwda-db.wsulibs.wsu.edu/oaiserver?verb=ListRecords&metadataPrefix=oai_dc&set=institution:nte

Metadata: The NWDA OAI data provider application uses a mapping between EAD collection-level data and unqualified Dublin Core. Research and practical work has been done on mapping component-level EAD information to Dublin Core; it is an extremely complex and problematic process, but one that has the potential for exposing valuable information. Given that OAI was designed to support "coarse granularity resource discovery," the initial support for collection-level information seems justified.

The locally-coded NWDA OAI harvester enables sharing of information with OAI-compliant repositories. OAI-PMH can also be used to support harvesting by search engines, using Google's Sitemaps program. Since the submission of the OAI-based Sitemap earlier this month, the indexing of NWDA documents in Google has improved, but the cause/effect relationship isn't completely clear.