Sunday, November 12, 2006

employing IRs to support self-archiving

Continuing with the scholarly communication reading...this evening, I read a preprint authored by Charles Bailey on open access issues. With the WSU institutional respository service, Research Exchange, being transitioned to production this fall, I focused on the role of self-archiving in supporting open access.

Bailey describes the self-archiving process: "scholars need the tools and assistance to deposit their referred journal articles in open electronic archives, a practice commonly called, self-archiving. When these archives conform to standards created by the Open Archives Initiative, then search engines and other tools can treat the separate archives as one." As I've become familiar with OAI-PMH in recent years, it has become increasingly clear to me that how an OAI-PMH implementation is coded and how a given repository is organized can both impact the ability of a data provider (or, repository) site to effectively expose its information. In my experience, subject-based harvesting of OAI-PMH repositories is seldom supported in software; the expectation is that the set organization (for example, in DSpace, a collection within a community is equivalent to an OAI set) is structured so that it enables the harvesting process. That assumption may or may not be correct.

Bailey describes in some detail the role of IRs in the self-archiving process. One asset of IRs, Bailey notes, is that "since they are formal institution functions, institutional repositories are permanent and stable." This seems to me to be a leap by the author, because the commitment to core IR functions such as preservation and identifier maintenance may vary greatly across institutions. But the IR concept does compare favorably to other self-archiving models cited by the author, such as storage/access on an author's personal website, and this should be communicated during the content recruitment process.

Bailey also writes that institutions running IRs "typically utilize free open source software, such as DSpace, Eprints, or Fedora." This may be less and less true over time, for a variety of reasons. (One being the need for IR services to compete with other digital repository efforts at a given institution.) A couple of commercial software options that I'm aware of include: OCLC CONTENTdm (employed at the University of Utah, an ARL institution, and at some other libraries; ProQuest Digital Commons (employed by a number of institutions, including the University of Pennsylvania).

In summary: Charles Bailey's article provides an interesting view of how IRs fit into the larger scholarly communication effort. Some of the assertions that I've made in the post are based upon narrow technical points, which are difficult to address when dealing other complex issues such as copyright and work practices related to scholarly communication.