Showing posts with label Computer science. Show all posts
Showing posts with label Computer science. Show all posts

2012-12-03

Take this ontology and shove it. Or, why classification matters.

I was recently called out by one of my colleagues for saying that ontologies were boring, this despite my own doctoral work on knowledge representation. Motivating my glib comment was an image of a group of pasty-faced individuals gathered around a large boardroom table and discussing which angel fit on which pin. Events from this past weekend are a reminder why such glibness is not helpful.

The American Psychiatric Association has just approved a set of updates, revisions and changes to the reference manual (DSM5) used to diagnose mental disorders. Among the changes are those redefining the inclusion and exclusion criteria for autistic disorders. By changing which children are classified as having an autistic disorder, parents will be made to feel more or less comfortable having a child carrying the diagnosis. Just as importantly, insurance companies and school programs might shift their criteria that determine which child and family gets what kind of support and at what cost. In the near term, clinical trials for the treatment of autism may not include the same patients as they would have prior to this retaxonomization.

So, are ontologies boring? Perhaps. But they certainly belong to the class of hugely important societal constructs.

Hat tip: David Osterbur.

2012-03-21

The passing of clean taxonomies.

Among the most productive constructs of the enlightenment are the modern taxonomies. These have been helpful in bringing order to the chaos of signs and symptoms and other clinical findings and were central tools in achieving our 20th century understanding of pathophysiology. They have also have an influential role to play in reimbursement for medical services. With the dawn of high-throughput molecular diagnostics many of us recognize that we are going to be able to be far more precise in our diagnostic and therefore therapeutic approach to diseases and their prevention.

Nonetheless, as we approach the systematization of medicine, we will be reminded often that nature may not hew to the simplified models that we are developing. This recent study in the New England Journal of Medicine, just does that by demonstrating directly that within a "single" tumor there exists a large multiplicity of tumor types, each with its own genomic characteristics and therefore particular therapeutic responsiveness (or lack of it). It can be argued that this is another instance of the tension between the "neats" and the "scruffies" but perhaps it is a foreshadowing of the decreased effectiveness of taxonomies as a cognitive tool for biomedical discovery and clinical care. If indeed, the underlying substructure of physiology is best represented by a probabilistic network model that can only be best grasped and managed through the use of computational tools, we have to seriously re-evaluate both our approach to disease definition and biomedical education.

2012-03-01

Get paid to play

Earlier, I described the SHRINE distributed query system across 6 million patients with 10 billion facts. If you are a member of the Harvard Medical School faculty (with employment at one of the affiliated hospitals) you now have the opportunity to get money and glory (more the latter than the former) to spin clinical data into biomedical gold. Details on the context can be found here: http://catalyst.harvard.edu/services/pilotfunding/shrine.html

If you have questions, use this email contact.

2012-02-16

Research by the numbers

What if you could mine the 10 billion medical facts across 6 million (anonymous) patients in five Harvard affiliated hospitals to ask an important and timely question? What are the other diseases or disorders associated with autism? How has the pharmacological treatment of inflammatory diseases changed over the last five years? Are there gender differences in prevalence of the infections in autoimmune diseases? How is the prevalence of diabetes mellitus changing in young adults?

Now, for the first time, if you are an eligible faculty member (or one of their fellows) in one of the five hospitals, you can now productively seek answers to these questions. The Shared Health Research Information Network (SHRINE) helps researchers overcome one of the greatest problems in population-based research: Compiling large groups of well-characterized patients. Eligible investigators may use the SHRINE web-based query tool to determine the aggregate total number of patients at participating hospitals who meet a given set of inclusion and exclusion criteria. The criteria are currently demographics, diagnoses, medications, and selected laboratory values. Because counts are aggregate, patient privacy is protected.

So, whether you are seeking a study cohort, preliminary studies for a grant proposal, or evaluating an epidemiological hypothesis, take this new tool for a spin and start translating this large mass of hard-won data into useful biomedical knowledge.

2011-10-09

Faster, cheaper and in control

Let's say you have a problem (e.g aligning the world's literature to defining the phylogenesis of the components of the current world-wide written corpus for scholarly attribution and automatic detection of plagiarism) that requires a computational solution. But it's taking days for the software to run. Buying a faster, bigger computer might provide some speed up, but what if you could get a 1000 fold improvement through a better implementation of the algorithm at the core of your software? Here's your chance to see if it can be done through a contest hosted by the Harvard Catalyst. Will the Overmind answer your most difficult computational questions?

XKCD Wikipedian protestor

2011-07-22

The unbearable effectiveness of data

Researchers in artificial intelligence (AI) of the 1980’s, librarians and aficionados of the Semantic Web have a shared faith: The unique value of human-designed knowledge structures whether they be taxonomies, ontologies or metadata. These knowledge representations are seen as providing important leverage in information retrieval, knowledge discovery, and decision-support. In this context, I was recently reminded by Alal Eran of an article by researchers at Google about the value of BIG data. These researchers (one of whom wrote a wonderful book on Common Lisp—Paradigms of Artificial Intelligence Programming: Case Studies in Common Lisp—widely appreciated by the AI community, which includes applications for expert systems) describe how statistical methods applied to trillion-word corpora can automatically support the aforementioned information tasks without requiring human annotation/categorization. It may be that the combination of human-derived annotations (whether crowd-sourced from the web or carefully curated in the monasteries of the ivory tower) can be used synergistically with the purely statistic-learning methods, but that has yet to be convincingly demonstrated. Until then, those of us working on genomic research will see how far we can get just with data, particularly those obtained in the course of healthcare.

For those of us in libraries and those of us who are librarians, there is now an active debate that has yet to achieve resolution on what value there is in human annotations and metadata. If there is value, at what cost? And if it is cost-effective, how do we demonstrate the efficacy? Our Universities' leaders will be interested in the answers and so will our colleagues at Google.

2011-03-08

Let the games begin!

Do you think that you can create the new software app that will revolutionize healthcare? Do you agree that substitutability will allow us all to innovate healthcare practice? As detailed on the challenge.gov website, there is now a very short term opportunity to "walk the talk" for a modest prize and immodest glory.

t SMArt Challenge

2011-03-03

Neat or scruffy?

Is your desk topped by the monumental accreta of your work or does it retain it's pure sheen of Scandinavian simplicity? It turns out that the dichotomy between the "Neats" and the "Scruffies" cuts across several broad swathes of the human condition. Among these are the archane arts of taxonomization and representation so well known to librarians, botanists, and engineers working on electronic health record interoperability. On the latter topic, the President's Council of Advisors on Science and Technology, (PCAST) report has issued a report on how health information technology will or will not be effectively used to improve healthcare. Given the work we are pursuing on substitutability, our own Ben Adida shared a perspective on the report.