Wednesday, December 2, 2009
Closing time
Next time this course is taught, please make weekly comments on others' blogs part of the grading schema, as was the case for LIS2600. It creates more of a community and forces you to think in greater detail about the postings of classmates.
Thank you.
Muddiest Point #10
Thursday, November 19, 2009
Week Eleven Readings
A lot of big ideas this week. I’m not sure where to start with the UCLA piece, but we’ve read a lot of William Arms’ work this semester and I’m pretty sure that this was the best one. In this article Arms is a powerful advocate for users of digital libraries, who, at some base level, really want to use one and only one digital library. Another way of thinking of this is federation (EBSCO calls it integration). On the technical side, this means interoperability not just in terms of search protocols, like Z39.50, but in terms of interfaces as well.
Elsewhere Roush, that link doesn’t work – get it through ULS, vacillates between cheerleading the Google Book Search and giving some librarians a chance to express their reservations. I think he’s correct that public and private digital librarieswill coexist, with occasional linkages, although it seems like he thinks Google Book Search is public when, in fact, it is not. So much for Arms’ users’ viewpoint. He also fails to account for Digital Rights Management.
Muddiest Point #9
In theory DLs should be easier to adapt/modify based on user critiques than brick and mortar libraries, or does most of this take place in testing/beta prior to the launch?
Friday, November 6, 2009
Week Ten Readings
All libraries are for and serve a specified population, even if they are digital. We were introduced to this in week one, but we’ve been ignoring it since… until now.
You can find the Arms chapter here, as Blackboard has the wrong URL. Panel 8.1 in particular is useful, a nice glossary of terms we’ll be using this week.
Kling and Elliot eloquently and succinctly make their point that one should not use a one size fits all model when designing a digital library. This article was a welcome breath of social science in this course. It’s nice to read something I fully understand.
Saracevic (pdf) asks why DLs aren’t really evaluated (although they will be in this class) and examines some of the evaluations out there. It’s a bit dry, but the points are well-made.
Finally, it’s ironic that all this week is about usability and how things look, but only the final reading (Sheiderman and Plaisant) has pictures. More of those would have been useful, especially in the Arms chapter.
Muddiest Point #8
A course-related muddiest point/rant: I have zero grades for this course. We are 2/3rds of the way through the semester. This is not acceptable.
A lecture-related muddy point: digitization overall and digital preservation in particular are very expensive, although the costs are declining. Given that, there have been very few mentions of the digital divide in this course, in that Google can spend its money and large research universities can spend their as well as grant money, but where does that leave small colleges, mid-sized universities, and public libraries?
Friday, October 30, 2009
Week Nine Readings
Rieger’s White Paper really ties the first half of the semester together by looking at large-scale digitization initiatives (LSDIs) that are public, private, and combinations. Of the ones mentioned here, I am least familiar with Microsoft’s. Is that still going on, and what is Live Search Academic? Because I can’t really do justice to this reading in a blog posting after taking a midterm, I’m reposting some of the questions she asks of LSDIs.
“Should we commit to preserve all the digital materials created through the LSDIs, implement a selection process to identify what needs to be preserved, or assign levels of archival efforts that match use level?
Will electronic access spur new demand for materials seldom used in print?
Will LSDIs’ use of high-speed, automated digitizing processes disenfranchise materials needing special handling?
Is there a means for recording gaps in collections and within publications?
How much duplication should there be in selection and digitization efforts?
What legal rights do participating libraries have to preserve in-copyright content digitized through LSDIs?”
The answers to these questions and insights based on her study are illuminating. I’ll move on after one more quote.
“Should we perceive these ventures primarily as access projects, rather than as reformatting initiatives that yield high-quality digital surrogates for the original?”
Like the White Paper, Hedstrom (pdf) and Jones are concerned about current technologies becoming outdated, forcing a mass migration of preserved materials to a new format, and that the digital library lacks a business model. Perhaps Google Book Search can change that in time, but for now I agree. Hedstrom also raises an interesting point when asking about how a database itself might be preserved.
Lavoie (pdf) provides a history and then a blueprint of OAIS, which
- “establish criteria for determining which materials are appropriate for inclusion in the archival store.”
- “determine the scope of its primary user community.”
- ensuring that the information is preserved in a form that is independently understandable to these users”
- establish and document clear policies and procedures for carrying out the preservation of the information in its custody”
- “committed to making the contents of its archival store available to its intended user community”
Functions include
- Ingest, the set of processes responsible for accepting information submitted by Producers and preparing it for inclusion in the archival store”
- “Archival Storage. This is the portion of the archival system that manages the long-term storage and maintenance of digital materials entrusted to the OAIS. More specifically, the Archival Storage function is responsible for ensuring that archived content resides in appropriate forms of storage – e.g., online, near-line, off-line – and that the bit streams comprising the preserved information remain complete and renderable over the long-term.”
- “Data Management function maintains databases of descriptive metadata identifying and describing the archived information in support of the OAIS’s finding aids; it also manages the administrative data supporting the OAIS’s internal system operations, such as system performance data or access statistics.”
- “Preservation Planning. This service is responsible for mapping out the OAIS’s preservation strategy, as well as recommending appropriate revisions to this strategy”
- “Access function manages the processes and services by which Consumers – and especially the Designated Community – locate, request, and receive delivery of items residing in the OAIS’s archival store.”
- “Administration function is responsible for managing the day-to-day operations of the OAIS, as well as coordinating the activities of the other five high-level OAIS services.”
Finally, Littman’s article is about the implementation and operation of a digital library and what went wrong along the way. It serves as a model for the rest of us, I suppose. Corrupted files, operator errors… it’s not pretty, although the National Digital Newspaper Program (NDNP) seems to work fine now.
Tuesday, October 27, 2009
Muddiest Point #7
You talked about personal digital libraries in the lecture. Does Flickr fit the bill as one of these? Your holdings on Google docs? What about a lack of searchable metadata, or a lack of metadata overall? How important is metadata to a library, digital or not?
And how does this tie into cloud computing, in which we do everything online – is one's cloud a digital library?
How much of this is conceptual stretching, meaning at this point what is not a digital library? Do these slides jibe with the definitions we read and discussed in the first few weeks of the course?
Sunday, October 25, 2009
Muddiest Point #6
Thursday, October 22, 2009
Muddiest Point #5
Week Eight Postings
From this protocol, all things flow, including the subject of the next three articles, federated searches. As someone who is about to purchase and implement a federated search engine at an academic library, these made for interesting reads. In particular, Miller’s final sentences spoke to me: “The paradox demonstrated so elegantly by Google is that the most powerful information access approach also happens to be the simplest and easiest. The most complex and least intuitive interfaces wind up securing information, not facilitating information access.” I hadn’t thought of Google as being a federated search, but it’s true. Occam’s razor is certainly in effect here and we must always take patrons into account.
Hane’s article was strong on the limits of federated searching, but for my purposes in a library setting such a tool introduces users to resources they have not yet explored, and thus may help maximize a library’s returns on its investments in terms of electronic holdings/databases.
Finally, Lossau had me wondering if I shouldn’t be including Google Scholar and Google Books in any federated search I set up. Asking how libraries determine what is academic is a simple, but powerful question, but this article is also a bit dated; increasingly librarians, myself included, see libraries as portals and gateways as opposed to collections.
Friday, October 9, 2009
Muddiest Point #4
Tuesday, October 6, 2009
Week Seven Readings
The gist of the Lesk reading, aside from the various methods and standards, is that it is not only difficult to store sounds, pictures, and moving images, but also that it can be hard to find them as well. My version of Lesk is the second edition, published in 2005. I wonder if Dublin Core and/or other metadata standards have made searching and indexing these non-text objects easier.
Here’s a link to one of my favorite digital libraries for sound and pictures at Cornell, the Macauly Library, which has thousands of animal sounds and images.
The links to the other readings don’t work, but they can be found through Pitt’s ULS.
Both articles by Hawking are about search engines; the first focuses on crawling the web while the second focuses on indexing the results of the crawl. One thing I did not know before reading part 2 is how prevalent caching is in the world of search engines.
“Caching. There is a strong economic incentive for search engines to use caching to reduce the cost of answering queries. In the simplest case, the search engine precomputes and stores HTML results pages for thousands of the most popular queries. A dedicated machine can use a simple in-memory lookup to answer such queries.” (90)
I’d like to read more about that.
I teach information literacy courses, and I use some the same data as Henzinger when discussing Google – 85% of users don’t make it past the first page of results. Overall I found the discussion of search engines versus spammers pretty interesting. It’s nice to know the tricks of the trade on both sides.
Friday, October 2, 2009
Week Six Readings
Bryan’s article mentions the three things that go into an XML file: the processing instructions, which denote which version is being used; a document type declaration; and a document instance, which should be fully tagged, meaning it has matching tags at the beginning and ending.
Ogbuji’s article isn’t really an article per se, it’s much more a collection of links, and the tutorials he offers up are excellent – learning by doing is often the only way on the internet.
The link to Bergholz is broken, but Googling the title and author takes you to a pdf of what we’re looking for. His examples of XML are most welcome and helpful. I found examples 4 and 5 particularly helpful in terms of digital libraries – just think of it as more metadata.
Finally, I recall using w3schools a lot in 2600. Pretty much everything I’ve seen on this site has made me better understand markup languages. Bookmark it.
Friday, September 25, 2009
Wednesday, September 23, 2009
Week Four Readings
The first thing that caught my attention in the Witten readings is that he mentions that retrieval in brick and mortar libraries is more powerful than in virtual ones, thanks in no small part to the physical organization of books, which is often by call number or author (47). But then again, he also mentions (55) that digital libraries have no need to “shelve” based on subject. The discussion of what goes into a bibliographic structure and a bibliography is useful.
Much of chapter five strikes me as a rehash of things that were discussed in LIS2001 – Organizing Information, but it’s always good to read them again, and since that course has now merged with 2002 – Retrieving Information, perhaps this course becomes the one with the best overview of MARC and Dublin Core. As is the case with other readings,
As a cataloger, I found Gilland a little frustrating. The point of a catalog is to help users find something, but she seems to reinterpret metadata from “data about data” to “data for the sake of data” throughout the article. Anybody else get that impression?
Nonetheless, Tables 2 and 3 concerning the types of metadata and their uses are excellent resources I’ll be referring to for a long time. She’s also got a great quote towards the end of the article: “Metadata is like interest: it accrues over time. To stretch the metaphor further, wise investments generate the best return on intellectual capital. Carefully crafted metadata results in the best information management—and the best end-user access—in both the short and the long term.”
Finally, none of this week’s readings problematized Library of Congress Subject Headings, instead taking them at face vlaue. Anyone who catalogs, tags, or adds metadata to an item/object should definitely take a look at these resources, which do an excellent job of exposing some of the biases behind LCSHs.
Tuesday, September 22, 2009
Muddiest Point #3
On a semi-related note, Wikipedia uses DOIs.
Thursday, September 17, 2009
Week Three Readings
The Lesk readings remind me of my youth, spent working for both ProQuest and the
At
A question about Lesk: on page 34 there’s a chart that shows the number of online databases diminishing in 1997 and 1998. Anybody know why this is?
Unlike last week, Arms writes in clear, plain English and it’s very easy to following along, surprising given the topic. This chapter about how text and pages are digitally represented was thorough, especially regarding the rise of Unicode and PDFs. Random question, how come Facebook won’t display “>”? It’s ASCII, right?
A question about Lynch: does he conflate identifiers and handles? It certainly seems that way to me. So much so that I went and read some of this week’s background reading just to be clear.
Paskin’s article on DOIs was not exactly a page-turner, but it got the job done and seemed to clear up the relationship between identifiers and handles (A handle is a specific type of DOI). Also important to remember: even if an object is not digital, it can be identified digitally. One reason why DOI matters.
Tuesday, September 15, 2009
Muddiest Point #2
For the group term project, are we addressing 3 groups of users, or 1 group of users interested in something interdisciplinary that requires 3 collections?
