Friday, October 30, 2009

Week Nine Readings

IV. OAIS Environment

Rieger’s White Paper really ties the first half of the semester together by looking at large-scale digitization initiatives (LSDIs) that are public, private, and combinations. Of the ones mentioned here, I am least familiar with Microsoft’s. Is that still going on, and what is Live Search Academic? Because I can’t really do justice to this reading in a blog posting after taking a midterm, I’m reposting some of the questions she asks of LSDIs.

“Should we commit to preserve all the digital materials created through the LSDIs, implement a selection process to identify what needs to be preserved, or assign levels of archival efforts that match use level?

Will electronic access spur new demand for materials seldom used in print?

Will LSDIs’ use of high-speed, automated digitizing processes disenfranchise materials needing special handling?

Is there a means for recording gaps in collections and within publications?

How much duplication should there be in selection and digi­tization efforts?

What legal rights do participating libraries have to preserve in-copyright content digitized through LSDIs?”

The answers to these questions and insights based on her study are illuminating. I’ll move on after one more quote.

“Should we perceive these ventures primarily as access projects, rather than as reformatting initiatives that yield high-quality digital surrogates for the original?”


Like the White Paper, Hedstrom (pdf) and Jones are concerned about current technologies becoming outdated, forcing a mass migration of preserved materials to a new format, and that the digital library lacks a business model. Perhaps Google Book Search can change that in time, but for now I agree. Hedstrom also raises an interesting point when asking about how a database itself might be preserved.


Lavoie (pdf) provides a history and then a blueprint of OAIS, which

  • “establish criteria for determining which materials are appropriate for inclusion in the archival store.”
  • “determine the scope of its primary user community.”
    • ensuring that the information is preserved in a form that is independently understandable to these users”
  • establish and document clear policies and procedures for carrying out the preservation of the information in its custody”
  • “committed to making the contents of its archival store available to its intended user community”

Functions include

  1. Ingest, the set of processes responsible for accepting information submitted by Producers and preparing it for inclusion in the archival store”
  2. “Archival Storage. This is the portion of the archival system that manages the long-term storage and maintenance of digital materials entrusted to the OAIS. More specifically, the Archival Storage function is responsible for ensuring that archived content resides in appropriate forms of storage – e.g., online, near-line, off-line – and that the bit streams comprising the preserved information remain complete and renderable over the long-term.”
  3. “Data Management function maintains databases of descriptive metadata identifying and describing the archived information in support of the OAIS’s finding aids; it also manages the administrative data supporting the OAIS’s internal system operations, such as system performance data or access statistics.”
  4. “Preservation Planning. This service is responsible for mapping out the OAIS’s preservation strategy, as well as recommending appropriate revisions to this strategy”
  5. “Access function manages the processes and services by which Consumers – and especially the Designated Community – locate, request, and receive delivery of items residing in the OAIS’s archival store.”
  6. Administration function is responsible for managing the day-to-day operations of the OAIS, as well as coordinating the activities of the other five high-level OAIS services.”


Finally, Littman’s article is about the implementation and operation of a digital library and what went wrong along the way. It serves as a model for the rest of us, I suppose. Corrupted files, operator errors… it’s not pretty, although the National Digital Newspaper Program (NDNP) seems to work fine now.

Tuesday, October 27, 2009

Muddiest Point #7

You talked about personal digital libraries in the lecture. Does Flickr fit the bill as one of these? Your holdings on Google docs? What about a lack of searchable metadata, or a lack of metadata overall? How important is metadata to a library, digital or not?

And how does this tie into cloud computing, in which we do everything online – is one's cloud a digital library?

How much of this is conceptual stretching, meaning at this point what is not a digital library? Do these slides jibe with the definitions we read and discussed in the first few weeks of the course?

Sunday, October 25, 2009

Muddiest Point #6

Online I keeping seeing Metadata Encoding and Transmission Standard (METS) and Analyzed Layout and Text Object (ALTO) mentioned as an increasingly popular digital library standard. I know that the LOC has used it, too. Is this something that's worth looking into?

Thursday, October 22, 2009

Muddiest Point #5

I am wondering how metadata is expressed for multimedia objects mentioned in the Fast Track Weekend lecture. I assume it's in XML, and perhaps Dublin Core, but would like a little more information on this.

Week Eight Postings

The chapter on OAI-PHM is useful for comparing and contrasting that approach to that of Z39.50, which is then discussed by Lynch. I think we should have read the Lynch article much earlier in the semester, if not much earlier in the program. In short, Z39.50 “is a protocol which specifies data structures and interchange rules that allow a client machine (called an "origin" in the standard) to search databases on a server machine (called a "target" in the standard) and retrieve records that are identified as a result of such a search.”

From this protocol, all things flow, including the subject of the next three articles, federated searches. As someone who is about to purchase and implement a federated search engine at an academic library, these made for interesting reads. In particular, Miller’s final sentences spoke to me: “The paradox demonstrated so elegantly by Google is that the most powerful information access approach also happens to be the simplest and easiest. The most complex and least intuitive interfaces wind up securing information, not facilitating information access.” I hadn’t thought of Google as being a federated search, but it’s true. Occam’s razor is certainly in effect here and we must always take patrons into account.
Hane’s article was strong on the limits of federated searching, but for my purposes in a library setting such a tool introduces users to resources they have not yet explored, and thus may help maximize a library’s returns on its investments in terms of electronic holdings/databases.
Finally, Lossau had me wondering if I shouldn’t be including Google Scholar and Google Books in any federated search I set up. Asking how libraries determine what is academic is a simple, but powerful question, but this article is also a bit dated; increasingly librarians, myself included, see libraries as portals and gateways as opposed to collections.

Friday, October 9, 2009

Muddiest Point #4

A question about BUBL Link: has this caught on at all, or are BUBL, Scorpion and the like localized, that is, not widely used outside of their origins? In other words, who uses these classification schemes, and do their creators hope to spread and disseminate them?

Tuesday, October 6, 2009

Week Seven Readings

Posting a bit early, but here goes:

The gist of the Lesk reading, aside from the various methods and standards, is that it is not only difficult to store sounds, pictures, and moving images, but also that it can be hard to find them as well. My version of Lesk is the second edition, published in 2005. I wonder if Dublin Core and/or other metadata standards have made searching and indexing these non-text objects easier.


Here’s a link to one of my favorite digital libraries for sound and pictures at Cornell, the Macauly Library, which has thousands of animal sounds and images.


The links to the other readings don’t work, but they can be found through Pitt’s ULS.


Both articles by Hawking are about search engines; the first focuses on crawling the web while the second focuses on indexing the results of the crawl. One thing I did not know before reading part 2 is how prevalent caching is in the world of search engines.

Caching. There is a strong economic incentive for search engines to use caching to reduce the cost of answering queries. In the simplest case, the search engine precomputes and stores HTML results pages for thousands of the most popular queries. A dedicated machine can use a simple in-memory lookup to answer such queries.” (90)

I’d like to read more about that.


I teach information literacy courses, and I use some the same data as Henzinger when discussing Google – 85% of users don’t make it past the first page of results. Overall I found the discussion of search engines versus spammers pretty interesting. It’s nice to know the tricks of the trade on both sides.