Showing posts with label Data Re-use. Show all posts
Showing posts with label Data Re-use. Show all posts

2013-07-30

A thoughtful and useful report on the Aaron Swartz tragedy

This (http://swartz-report.mit.edu/docs/report-to-the-president.pdf) is a report that Professor Abelson helped author on the behalf of MIT. It is chock full of lessons for education institutions, libraries and the larger academic ecosystem, including the ancillary industries. Although section V is focused on questions for MIT, many of those same questions will find wider resonance.

2013-04-08

Getting Big About Mapping Dengue

Here's a very nice application of lightly used data sources about Dengue, a scourge of underdeveloped countries. As in so many areas of public health, this huge health burden is woefully under-documented. In the absence of a vetted vaccine, understanding where it is endemic is essential for the application of scarce preventive resources. This group of investigators have cleverly used a number of public but under-used data sources, including the published literature, to create a predictive map of where 390 million infections per year are occurring.

) Dengue Map

2012-12-03

Take this ontology and shove it. Or, why classification matters.

I was recently called out by one of my colleagues for saying that ontologies were boring, this despite my own doctoral work on knowledge representation. Motivating my glib comment was an image of a group of pasty-faced individuals gathered around a large boardroom table and discussing which angel fit on which pin. Events from this past weekend are a reminder why such glibness is not helpful.

The American Psychiatric Association has just approved a set of updates, revisions and changes to the reference manual (DSM5) used to diagnose mental disorders. Among the changes are those redefining the inclusion and exclusion criteria for autistic disorders. By changing which children are classified as having an autistic disorder, parents will be made to feel more or less comfortable having a child carrying the diagnosis. Just as importantly, insurance companies and school programs might shift their criteria that determine which child and family gets what kind of support and at what cost. In the near term, clinical trials for the treatment of autism may not include the same patients as they would have prior to this retaxonomization.

So, are ontologies boring? Perhaps. But they certainly belong to the class of hugely important societal constructs.

Hat tip: David Osterbur.

2012-08-08

Billions and billions of gene expression measurements.

Let's say you are looking for a disease biomarker. Hopefully, one better than prostate specific antigen. Next time you or your student reach for a pipette to see if a gene is expressed in a particular tissue or disease, perhaps you should first check with the public databases of gene expression. As outlined in this article, we now have hit the one million array mark. That is, one million arrays measuring gene expression across thousands of conditions (tissues, diseases, pharmacological or environmental perturbation). And each array has tens of thousands of genes so these corpora have billions of gene expression measurements. That means you'll immediately be able to see if your favorite gene is uniquely expressed in a tissue in a specific disease. Or not.

Another way to think about these corpora is that they constitute one of the largest open access biomedical libraries. A model for clinical research to emulate?

Hat tip: Atul Butte.

Growth Microarray Data

2012-07-24

Unstandardized standards

An insightful naïf learning about the difficulty of sharing one electronic health record from one hospital to another might reasonably ask "Why don't they just create a standard for data sharing so that I can install or delete health apps at will and view my data on several different electronic health record systems?" An expert will then inform that impertinent naïf that it's much more complicated than she understands and that the standards already exist. When challenged, the expert will cite several august committees which have ratified standards such as the Continuity of Care Document (CCD). At this point, our naive protagonist should refer the expert to this blog entry by Josh Mandel. If by then, the expert is not holding his hands to his ears, he will explain that all standards are evolving entities and that these challenges are just the expected missteps on the path of convergent evolution to interoperable samadhi.

2012-06-13

Dearth of Death: A Fatal Wound to Medical Research?

My esteemed colleague L.J. Wei often reminds us that health outcomes which are not as hard-edged as death can be misleading. For example, the early press, decades ago, about the uncovering of early cancer by the Prostate Specific Antigen (PSA) was used to justify the surgical removal of hundreds of thousands of prostates. In hindsight, neither the PSA test nor much of the ensuing expensive and occasionally morbid surgeries made a significant dent in lifespan.

One might therefore reasonably conclude that the government, the census bureau, Social Security Administration or the Department of Health and Human Services would therefore place the highest premium on the accurate reporting of death, and its causes, for our citizens. Surely, those data are the incontrovertible evidentiary base for our public health monitoring, medical treatment evaluations (whether of drug, device or procedure), and projections of the fundamental demographics of our nation. So, it might be all too easy for most of us to overlook or dismiss the following innocuous-appearing bureaucratese-laden announcement

IMPORTANT NOTICE The National Technical Information Service (NTIS) has been notified by the Social Security Administration (SSA) of an upcoming important change in the Death Master File data. NTIS, a cost-recovery government agency, disseminates the DMF data on behalf of SSA. Please see the attachment, provided by SSA, for an explanation of the change. The implementation date of this change is November 1, 2011. Should you have any questions, please email me at wstrickland@ntis.gov and I will be happy to forward any questions not answered by the attachment to the Social Security Administration for reply.

What does this mean? It means that there is no longer a single, federal authoritative source of death records. Most of the operational details have now devolved to individual states without guarantees of consistency of reporting or a one-stop-shop for researchers looking for the national distribution of the Grim Reaper. Will we have to resort to crowd-sourcing death now in order to perform accurate population research?

Hat tip: Shawn Murphy

Death workflow

2012-03-21

The passing of clean taxonomies.

Among the most productive constructs of the enlightenment are the modern taxonomies. These have been helpful in bringing order to the chaos of signs and symptoms and other clinical findings and were central tools in achieving our 20th century understanding of pathophysiology. They have also have an influential role to play in reimbursement for medical services. With the dawn of high-throughput molecular diagnostics many of us recognize that we are going to be able to be far more precise in our diagnostic and therefore therapeutic approach to diseases and their prevention.

Nonetheless, as we approach the systematization of medicine, we will be reminded often that nature may not hew to the simplified models that we are developing. This recent study in the New England Journal of Medicine, just does that by demonstrating directly that within a "single" tumor there exists a large multiplicity of tumor types, each with its own genomic characteristics and therefore particular therapeutic responsiveness (or lack of it). It can be argued that this is another instance of the tension between the "neats" and the "scruffies" but perhaps it is a foreshadowing of the decreased effectiveness of taxonomies as a cognitive tool for biomedical discovery and clinical care. If indeed, the underlying substructure of physiology is best represented by a probabilistic network model that can only be best grasped and managed through the use of computational tools, we have to seriously re-evaluate both our approach to disease definition and biomedical education.

2012-03-01

City as organism

This video from Geneva, Switzerland is a beautiful instance of the repurposing of data. Shown are the data flows between cell phones across the city over night and day. This glimpse of the interactions over time also suggests new frontiers in real-time epidemiology. What if these (anonymized) data could be tagged with symptoms (e.g. cough, sneeze) could we track the spread of infections? Public health would then start to look a lot more like intensive care medicine, providing real-time monitoring of cities or nations (taking the "pulse" of the population, evaluating the activity and coherence of its "neural" activity). Will there be a new research and medical discipline that fuses the sciences of population ecology and population health? And should populations be empowered to forego such intensive study?

(Hat tip: Joshua Parker).

x

Ville Vivante from Interactive Things on Vimeo.

p.s. As a someone who grew up in Geneva, I was quite surprised to see a lot of activity at 2 AM. Is this the consequence of the Swiss work ethic?

Get paid to play

Earlier, I described the SHRINE distributed query system across 6 million patients with 10 billion facts. If you are a member of the Harvard Medical School faculty (with employment at one of the affiliated hospitals) you now have the opportunity to get money and glory (more the latter than the former) to spin clinical data into biomedical gold. Details on the context can be found here: http://catalyst.harvard.edu/services/pilotfunding/shrine.html

If you have questions, use this email contact.

2012-02-16

Research by the numbers

What if you could mine the 10 billion medical facts across 6 million (anonymous) patients in five Harvard affiliated hospitals to ask an important and timely question? What are the other diseases or disorders associated with autism? How has the pharmacological treatment of inflammatory diseases changed over the last five years? Are there gender differences in prevalence of the infections in autoimmune diseases? How is the prevalence of diabetes mellitus changing in young adults?

Now, for the first time, if you are an eligible faculty member (or one of their fellows) in one of the five hospitals, you can now productively seek answers to these questions. The Shared Health Research Information Network (SHRINE) helps researchers overcome one of the greatest problems in population-based research: Compiling large groups of well-characterized patients. Eligible investigators may use the SHRINE web-based query tool to determine the aggregate total number of patients at participating hospitals who meet a given set of inclusion and exclusion criteria. The criteria are currently demographics, diagnoses, medications, and selected laboratory values. Because counts are aggregate, patient privacy is protected.

So, whether you are seeking a study cohort, preliminary studies for a grant proposal, or evaluating an epidemiological hypothesis, take this new tool for a spin and start translating this large mass of hard-won data into useful biomedical knowledge.

2012-01-10

Ancestral betrayal

We share many things with our ancestors, including a fraction of their genetic code. This genetic link to the past has further invigorated an already large industry and hobby in the exploration of genealogy and historical provenance. This piece from today's news, shows how these same records can be used to leverage your ancestors to identify you. In this instance, a murder suspect is potentially fingered by his ancestors from the Mayflower.

Hat tip: Ben Reis

2011-12-21

Crystallizing Methods for Drug Effects

Yet another powerful example of the novel syntheses in which the aggregation of individual knowledge sources becomes a lot more than the sum of its parts: Using only the most well accepted and conventional of compedia of drug characteristics, Cami et al. demonstrate the power of integrative reasoning and systematic applications of "guilt-by-association" to predict accurately, across an interval of 5 years, hundreds of novel adverse events across more than 800 drugs. This performance is likely to improve further as they proceed to integrate healthcare data from electronic health records from consumers. The addition of molecular characterizations are also anticipated to provide further precision and individualization of these predictions.

drugAE

These developments will be further accelerated if we can liberate data gathered with public funding for the public good.

2011-11-16

2011-09-27

Out in the Open

This impressive compilation from GOOD (the data issue) documents the impressive growth of Application Programming Interfaces that provide third party software developers with access to, and the ability to repurpose, large and very useful data sets. This growth is driven both by altruism and self-interest and represents a dramatic refutation of the skepticism towards the open data movement of merely a decade ago.

Hat tip David Kreda

openapi

2011-05-13

Stuffed full of information

That bird rendered by taxidermy in a museum is not only of visual interest. This report by an enterprising undergraduate points to some unexpected public health insights gleaned from the feathers of these museum specimens about the trajectory over centuries of mercury contamination in seabirds. Perhaps needless to say, this could not be done with a purely virtual collection.

2011-02-21

Hall or House of Mirrors? The citation perspective.

Kudos to the analysts at SCImago. They have provided an outstanding, entertaining and educational perspective on worldwide academic publishing. I'll focus here on only one aspect: citations. Although the United States is the leader in citations at 87M citations, it is a surprising laggard in self-citation (32% citations are self-citations). The leaders are China (62%), Lithuania (38%), and Iran (37%). However, in the domain of medicine, authors in the United States are considerably less reticent and rack up a self-citation rate of 47%, earning them second place. The map below of the subject areas of the USA publications indicates where the action is.

Hat tip: Peter Park

USA-publlications

2010-05-31

What about the environment?

Hundreds of millions of dollars have been spent on comparing the frequencies of genetic variants in disease-afflicted and control populations to suss out the genetic basis of diseases, common or rare. The results for diseases with diabetes have been mixed which despite the high concordance of this disease in identical twins raises yet again the question of the role of the environment in the etiology of diabetes. Of course, we know that the inherited component of diabetes risk is contingent on environmental factors (notably diet) but these are much harder to quantify and moreover there is a universe of environmental risks that is potentially much larger than the entirety of the genome. So how to go about capturing more of the environmental risk as it pertains to real human beings (as opposed to petri dish or rodent studies)?

Atul Butte at Stanford shows us how, through the lens of informatics, systematic approaches can be taken to understand the environmentally-borne determinants of our disease burden. He leverages one of the most admirable public national studies of US citizens, the NHANES study. Butte and colleagues performed analyses analoguously to genome wide association studies but instead examined each of the chemicals (e.g. pesticides, heavy metals, vitamin metabolites) measured in the urine and blood of the members of the NHANES groups comparing cases of diabetes vs. controls. As summarized in the figure below, after comparing multiple groups they found a consistent increased risk of diabetes with PCB and pesticide derivatives and a decreased risk with some carotene derivatives (related to vitamin A). Although, just as for their genome-wide analogues, the results of this Environment Wide Association Study (EWAS) should be taken with due caution, they add to a growing array of methodologies that provide a complement to conventional randomized and controlled studies.

Environement Wide Association Study

2010-03-08

the new conditions of the Harvard Medical School promise to be as nearly ideal as the forethought of man can plan

or so it says in this 1905 article from Popular Science now available to all through the Popular Science archive viewer (courtesy of Google). Worth reading also for many wonderful quotes including "A medical student so trained [ in regularly reading selected up-to-date publications] in the use of medical literature can hardly be content to depend on antiquated text-book knowledge in his practise in after years." Amen. But do our students currently know how to (and have the culture) get up to date genetic and genomic relevant knowledge from the web? Perhaps our libraries can continue to lead in this regard.

Popsci1905

2010-02-21

What the Tell-Tale Heart Tells Us About Healthcare

At the end of his short story, Edgar Allen Poe's haunted murderous protagonist lets loose in front of the unsuspecting police officers:

"Villains!" I shrieked, "dissemble no more! I admit the deed! -- tear up the planks! -- here, here! -- it is the beating of his hideous heart!"

It is this story that inspired the title of a paper we published a few years ago on how we could detect the increase (and subsequent decrease) of myocardial infarctions coincident with the rise and fall of the use of Vioxx. This investigation relied solely on the informational byproducts of healthcare delivery. This weekend a Senate report on the risk of Avandia and what was known by GSK came to light. This resonated because we had recently published an article in Diabetes Care about the rapid identification of an increased risk of myocardial infarction with Avandia in patients with diabetes mellitus as compared to other drugs, even drugs in the same class, such as Pioglitazone. Whereas the current headlines are about what the pharmaceutical company knew or dissimulated, there is a broader question that needs addressing: Should every healthcare system not be instrumented so that the clinical leaders of these systems should know whether there are unexpected changes in the risks and health status of their patient populations?

Often, the state of clinical practice is compared unfavorably to the practice of commercial air travel, but what remains underemphasized is that as a system, we are flying blind. There is no local, regional or national air-traffic-controller-equivalent for the healthcare system. Should not the local hospital, and Department of Public Health be the first to know if there is about to be a local health collision or crash? Should not such local surveillance systems run in parallel to regional and national systems? Do we not need multi-level redundancy and open communication to avoid tens of thousands of unnecessary deaths? Further, even as the federal government becomes more aware of the need of such instrumentation, we all expect that our local healthcare systems and health authorities should know of any untoward health trends. Perhaps healthcare systems will start to compete on being able to provide timely and localized health trend data to their customers. Unfortunately right now, the major investments in such health market intelligence is to the payors (i.e. the insurers) who quite reasonably want to know what are the local risks, performance and trends for each of their contracts. Will it take regulation or market competition to make such data extraction and return to patients a matter of course? Let's hope we do not have to wait for a post-mortem Poe'esque orgy of recrimination to find out.

2009-09-18

Once we have electronic medical records implemented, what then?

We, as nation, are in the process of investing several billion dollars into the implementation of electronic health records. If all goes well, there will be a lot of individual data buried in these care systems. This begs the question of what utility, if any, this data has for research whether for genomics, comparative effectiveness research, pharmacovigilance, or public health. The NIH is hosting a conference at the end of October (entitled “Widening the Use of Electronic Health Record Data for Research”) to attempt to answer the question. All are interested parties are invited.