Showing posts with label Collections. Show all posts
Showing posts with label Collections. Show all posts

2014-01-23

Welcome news from the Wellcome Collection

Many of us have had the following experience. We find an interesting illustration and use it for a scholarly presentation or for a class. Then a well meaning compliance officer will notify us that if we are going to put the presentation on the web, we will have to secure the rights to the illustration for that purpose. Often the expedient response is to delete the illustration or withdraw the entire presentation from the public domain.

In this context this British invasion is a hopeful glimpse of the future. One hundred thousand pictures were just made freely available (so long as you correctly attribute their provenance). So, if you want to illustrate Harvey’s anatomical exercises of the heart and circulation, or contrast modern medical advertising to that of a well regarded phrenologist, the materials are there for you (see below), for free and without any administrative overhead. The other flank of the British open access image invasion is led by the British Library from which the charming picture below of the anatomy of the leech was taken.

L0035170 Page 1 & 2, Harvey\'s discussion on the motion heartL0040553 Leaflet for Prof. Thomas Moores, a Practical PhrenologiLeech from the British Library

2012-08-08

Billions and billions of gene expression measurements.

Let's say you are looking for a disease biomarker. Hopefully, one better than prostate specific antigen. Next time you or your student reach for a pipette to see if a gene is expressed in a particular tissue or disease, perhaps you should first check with the public databases of gene expression. As outlined in this article, we now have hit the one million array mark. That is, one million arrays measuring gene expression across thousands of conditions (tissues, diseases, pharmacological or environmental perturbation). And each array has tens of thousands of genes so these corpora have billions of gene expression measurements. That means you'll immediately be able to see if your favorite gene is uniquely expressed in a tissue in a specific disease. Or not.

Another way to think about these corpora is that they constitute one of the largest open access biomedical libraries. A model for clinical research to emulate?

Hat tip: Atul Butte.

Growth Microarray Data

2012-04-04

Copyright + book = disappearing act?

Many of us worry that the last three decades of scholarship will be lost to posterity because we have yet to provide an institutionalized set of mechanisms for preserving the digital output of academia that are as durable and accessible as our accreted analog/paper-based records. Now here comes word of an equally worrisome trend, the empirical evidence that worries about the effect of copyright on the access to important cultural, literary and scientific works, may be well-founded. A recent article in the Atlantic Monthly describes a yawning gulf in the sale of new books from recent decades such that there are more new books sold by Amazon from 1850, than there are from 1950. In their analyses, the investigators provide intriguing circumstantial evidence suggesting that the copyright laws may be responsible. More detail in this lecture by Paul Heald.

Hat tip: David Osterbur

Amazon copyright hole

2011-10-10

You are not the boss of me!

This assertion, made by millions of children with respect to their siblings, has its echos in many contemporary conversations about scholarship and authoritativeness. Whether it is the dominance of genetic or environmental effects in child development, or the value of a prostate specific antigen test, or the meaning of a genetic mutation, one group's truth is another's discredited hypothesis or outdated approximation. Of course, there some facts whose authoritativeness are beyond doubt, such as the speed of light. Well, maybe. Certainly, until that eschatological moment when we will all know the absolute truth, those of us who are entrusted with the curation and dissemination of knowledge, in its various guises, will have to provide tools to manage the growing multiplicity of perspectives.

In that spirit, I have unearthed a piece I wrote with Russ Altman over a decade ago about authoritativeness in the peer review process and how it could be managed ecumenically. We'll see if our futurology was authoritative.

2011-07-22

The unbearable effectiveness of data

Researchers in artificial intelligence (AI) of the 1980’s, librarians and aficionados of the Semantic Web have a shared faith: The unique value of human-designed knowledge structures whether they be taxonomies, ontologies or metadata. These knowledge representations are seen as providing important leverage in information retrieval, knowledge discovery, and decision-support. In this context, I was recently reminded by Alal Eran of an article by researchers at Google about the value of BIG data. These researchers (one of whom wrote a wonderful book on Common Lisp—Paradigms of Artificial Intelligence Programming: Case Studies in Common Lisp—widely appreciated by the AI community, which includes applications for expert systems) describe how statistical methods applied to trillion-word corpora can automatically support the aforementioned information tasks without requiring human annotation/categorization. It may be that the combination of human-derived annotations (whether crowd-sourced from the web or carefully curated in the monasteries of the ivory tower) can be used synergistically with the purely statistic-learning methods, but that has yet to be convincingly demonstrated. Until then, those of us working on genomic research will see how far we can get just with data, particularly those obtained in the course of healthcare.

For those of us in libraries and those of us who are librarians, there is now an active debate that has yet to achieve resolution on what value there is in human annotations and metadata. If there is value, at what cost? And if it is cost-effective, how do we demonstrate the efficacy? Our Universities' leaders will be interested in the answers and so will our colleagues at Google.

2011-05-13

Stuffed full of information

That bird rendered by taxidermy in a museum is not only of visual interest. This report by an enterprising undergraduate points to some unexpected public health insights gleaned from the feathers of these museum specimens about the trajectory over centuries of mercury contamination in seabirds. Perhaps needless to say, this could not be done with a purely virtual collection.

2010-03-24

Who's Gonna Pay for these Journals?

Town Hall Meeting: Who’s Gonna Pay for these Journals?

2:00-3:30 PM

Thurs., April 8, 2010

TMEC, Walter Amphitheater

Harvard Medical School


Scholarly Communication is broken.


Your access to articles in Brain Research, Tetrahedron Letters, Cell, Nature publications, and other leading journals in every discipline is paid for to the tune of millions of dollars per year by Harvard libraries.


Journal costs are skyrocketing and free market processes are failing to control costs--STM publishers can charge what they want without regard to value because they have a monopoly on the content that you have given them.  More and more you will be seeing "This article is not included in your organization's subscription " because libraries can no longer afford to buy back the content that has been freely given to the publishers. Commercial publishers make up to 40% profit on work produced here at Harvard and other research institutions. How can we establish some control over these costs and at the same time make it easier for you to regain control of your rights to use your own work?


Come to a Town Meeting and discuss what we can do to fix scholarly communication!

2010-03-11

We will miss our book-filled bookshelves (at least for a while)

This article on the proliferation of incompatible electronic book formats and digital rights management (DRM) standards for book content is a stark reminder that much of what is digital is incompletely archived, if at all. This is particularly the case for our personal collections. I have some favorite novels and text books which I can still pick off my shelves in my office months or decades after I first purchased them. Until that unlikely day that all digital books are sold without DRM or that one DRM/encoding is adopted by all electronic book vendors, I will have to hope against hope that the vendor of the e-book I also use will remain in business. Otherwise, I will have to regularly repurchase all the books I purchased in electronic form every few years. Or just reread those older and less evanescent paper-based books.

2009-10-23

A physician of distinction: Oliver Wendell Holmes

An invitation to celebrate the life, the accomplishments, and the continuing relevance of the literary and scientific contributions of Dr. Oliver Wendell Holmes.

Oliver Wendell Holmes (1809–1894)

Oliver Wendell Holmes (1809–1894) spent parts of the nineteenth century as America’s best-known physician and best-selling author. Sir William Osler praised him as “the most successful combination which the world has ever seen, of the physician and man of letters.” Henry James, Sr., called him “intellectually the most alive man I ever knew.” Today, he is remembered as a physician for his investigation of the contagiousness of puerperal fever (two decades before the advent of the germ theory), his advocacy for therapeutic skepticism and rationalism, and for coining such terms as “anesthesia.” He is celebrated as a literary and cultural figure for such poems as “Old Ironsides” (considered responsible for saving the U.S.S. Constitution), for his early forays into what would be considered a new depth psychology, and for terming Boston the “Hub of the solar system” and describing its “Brahmin” caste.

Join us to help celebrate the life, the accomplishments, and the continuing relevance of the literary and scientific contributions of Dr. Oliver Wendell Holmes.

Dr. Oliver Wendell Holmes and the Spirit of Skepticism:

November 17, 2009, 1:00 PM- 5:00 PM, reception 5:00-6:30

Location: Countway Library, 10 Shattuck St., Boston

2009-09-23

Never Ending STories

In preparation for a conference on substitutable platforms in health IT, I was directed to an instance of a growing form of self-publication that we call the Never Ending STory (NEST). This instance of NEST is the knol which has become an increasingly popular venue for publications including ones that look a lot like standard peer reviewed journals. More generally, a NEST starts as an embryonic paper. With iteration and with the help of co-author and reader suggestions, it incubates a mature manuscript. Unlike a blog, it is not just a snapshot of a narrative perspective in a sequence of snapshots, but a single integrated document. Unlike a wikipedia article it does not claim encyclopedic authoritativeness (or at least sole authoritativeness so that disagreeing contributors have to battle it out) but only the moderated perspective of the authors. Unlike a standard peer review article, it's publication does not signify the end of its incubation and the hatching of a fully mature narrative. And it is timely and time efficient to make NEST's more prevalent. How often, have you read a scientific article from five years ago and wondered if more recent developments had influenced the authors' perspective on their prior results and/or conclusions? Would it not be more effective to allow the author to update their articles (while maintaining an archival history of all prior versions) so that they continue to be current? Or if there were additional data that bolstered the case of the original article, the author could add these data to that article without having to go through an entire process of a new publication just for the incremental data. That would reduce unnecessary publication noise and increase the value of the article to the reader.

Although, right now, we are using the knol as the infrastructure for our NESTs, we can hope that academic publishers will provide vehicles of similar functionality. Until, then we will just have to incubate our own.

2009-08-26

What are medical libraries expected to do?

This is not an abstract question about the future of libraries, although that is also an interesting question. It is a question about what the medical school accrediting organizations have determined. "The Liaison Committee on Medical Education (LCME) is the nationally recognized accrediting authority for medical education programs leading to the M.D. degree in U.S. and Canadian medical schools. The LCME is sponsored by the Association of American Medical Colleges and the American Medical Association." and this is what they had to say (the bold face is mine for emphasis):

D. Information Resources and Library Services
ER-11 The medical school must have access to well-maintained library and information facilities,
sufficient in size, breadth of holdings, and information technology to support its education and
other missions.

There should be physical or electronic access to leading biomedical, clinical, and
other relevant periodicals, the current numbers of which should be readily
available. The library and other learning resource centers must be equipped to
allow students to access information electronically, as well as to use self-instructional
materials.

ER-12 The library and information services staff must be responsive to the needs of the faculty, residents
and students of the medical school.

A professional staff should supervise the library and information services, and
provide training in information management skills. The library and information
services staff should be familiar with current regional and national information
resources and data systems, and with contemporary information technology.
[Revised annotation approved by the LCME in October 2007 and effective immediately.]
Both school officials and library/information services staff should facilitate access
of faculty, residents, and medical students to information resources, addressing
their needs for information during extended hours and at dispersed sites.
(This is taken from:
http://www.lcme.org/functions2008jun.pdf found at:http://www.lcme.org/standard.htm
Hat tip David Osterbur.)


These are important recommendations and ones which foreshadow trends from the very near future. We have embraced this educational mission from access of electronic resources to teaching biomedical researchers how to perform bioinformatics-enabled research (see the bioinformatics nanocourses offered to all by Reddy Galli— details here ). The central question is whether librarian training will embrace the information technology that will be required to keep libraries current and relevant to their patrons. The answer to that question will determine where the future librarians are trained and that will in turn determine how central libraries remain to the academic mission.


2009-07-16

2009-04-28

The price of knowledge is worth knowing

We all pay a lot of money for the product of our own collective academic enterprise. With 2.5 million downloads of pdf's by Harvard University patrons from our top three publisher packages, I wondered what the costs might be. Well, thanks to Betsy Eggleston, we now have a better idea:

Elsevier package 2008 article downloads: ($.76/download)
Wiley package 2008 article downloads: ($1.52/download)
Springer package 2008 article downloads: ($2.98/download)

That is a lot of money per click but several questions pose themselves:

a) Do all libraries have a similar cost per download?

b) Is the relative cost per download similarly ordered for each of these three publishers in other libraries?

c) What is the equivalent cost for a circulated book/monograph per patron-use? Is that a fair comparator?

d) What is the equivalent cost per download for open access publications (including the author cost)?

I suspect that knowing the answers to these questions is a source of leverage and power. How can we make decisions with and on behalf of our researchers, faculty and public without knowing these answers? Should we not insist on greater transparency of the relationship of academic value and cost. If you have any additional data, feel free to enter a comment regarding this post or send me an email and I will add it to this post.



2009-03-20

Medical Museum on the Web





(Originally uploaded by otisarchives1 )
An interesting example of how historical medical data can be shared more widely for scholars worldwide. If all medical libraries followed this lightweight formula for dissemination, we could avoid repeating many mistakes of the past and learn how better to deal with old challenges revisited (e.g. epidemics). Kudos to the Medical Museum.



2009-01-31

The power to tag is the power to control

This temporary proscription of large segments of the web to users of the Google search engine vividly illustrates the application of the power of ontologizing. Whole segments of our collective electronic corpus can be made obscure or brought to the top of our societal awareness merely by changing a few bits on an electronic tag. Librarians of the world unite! Yours is the power to [re-]organize.

2008-12-15

Foundations of Policy

"Twenty-first century leaders in medicine and government are confronted by questions of enormous magnitude: What are the determinants of disease and its distribution?  How should health outcomes be measured? How are we to optimize health care delivery and financing, and how are we to ensure access to the fruits of medical science to the poor of this country and the developing world? However, such twenty-first century dilemmas are not new."


The Center of the History of Medicine of the Countway Library has been growing under the leadership of Scott Podolsky and Kathryn Baker Hammond and most recently they were awarded a grant from the Andrew Mellon Foundation that will enable, for the first time, research in the manuscript collections of four influential leaders in public health: Leona Baumgartner, Alan Macy Butler, Howard Hiatt, and David Rutstein. This adds to the growing list of new initiatives by the Center, An important step towards understanding the current and future challenges in public health.



2008-11-24

'Spinning' up the tissues at Harvard

Ever since the late 1990's we have been working on a variety of methods for retrieving data across disparate institutions that are often not even part of the same corporation. In response to an RFA from the National Cancer Institute, called the Shared Pathology Informatics Network, we developed a toolkit that is now in widespread use, specifically to enable genomic and other biological studies on the millions of specimens that are archived across healthcare institutions. As is often the case, we were late in bringing our own tools to use in our own backyard, but that has now happened. The Pathology Specimen Locator (PSL) is now live at our CTSA Catalyst portal. As shown in this screenshot, with authorized credentials, I was able to see that there are over 10,000 lung cancer samples (you can query for any tissue type and disease) across a wide range of ages. It is this sort of IRB-protected, informatics-enabled data liquidity that will accelerate our translational research efforts. Hats off to the entire team but particularly Andy, Frank, John, and Mark.


spin-psl-vsl



2008-11-16

Thousands of free books

As per the Lexcyle website "Stanza is a free application for your iPhone and iPod Touch. Use it to download from a vast selection of over 40,000 books and periodicals, and read them right on your phone. It’s a wireless electronic library that stays open 24/7." A quick search for science fiction reveals classics by H. G. Wells and pulp fiction by "Doc" E. E. Smith. All without Digital Rights Management. A real treat.

2008-11-12

Matching high-throughput genomics to high-throughput mining of the literature

In this study (alas, for-fee-access) , Dennis Wall demonstrates how to mine the literature (aka the biomedical bibliome) to focus the analyses of noisy genomic modalities which by virtue of measuring thousands of genes have to be aggressively corrected for multiple hypothesis testing [ed. Disclosure: I am a co-author]. By examining gene expression analyses of individuals with autism through the lens of the prior literature on neuro-psychiatric-behavioral disorders, he is able to identify genes significantly differentially expressed in individuals with autism, both known and previously non-implicated genes. This is one of a growing list of publications that are attempting to match the high throughput qualities of genomic measurements with an equally efficient automated "reading" of all the painstakingly obtained biomedical investigational literature. It also suggests that an even more detailed annotation by librarians of the existing literature (analogously to what the National Library of Medicine has done for years for the broad addition of meta data) will be productively leveraged in future investigations. I suppose this is where my colleagues from the Semantic Web have another opportunity to feed the search engines of Google.



























2008-10-30

The Joy of Collecting

Every once in a while, I get a notice that reminds me that there are pleasant avocations that don't quite make it to "Reality TV" fare. Here is one such announcement.

2008_10_30_15_19_16