Showing posts with label epidemiology. Show all posts
Showing posts with label epidemiology. Show all posts

2014-08-06

Near real-time tracking of Ebola

If you want to see just how Ebola is spreading (and soon regressing?) , this article provides the view and insight of how it’s done. Ten years ago, these data would percolate slowly, compiled by public health authorities and eventually appear in an academic or lay publication Now through the foundations of computational epidemiology established by Dr. Brownstein, we are able to directly monitor this in-progress disaster. Now that we have a good afferent circuit, let’s see hope the efferent circuit is even more robust.

2013-04-08

Getting Big About Mapping Dengue

Here's a very nice application of lightly used data sources about Dengue, a scourge of underdeveloped countries. As in so many areas of public health, this huge health burden is woefully under-documented. In the absence of a vetted vaccine, understanding where it is endemic is essential for the application of scarce preventive resources. This group of investigators have cleverly used a number of public but under-used data sources, including the published literature, to create a predictive map of where 390 million infections per year are occurring.

) Dengue Map

2013-01-08

Epidemic or epiphenomenon?

A number of crowd-sourced infection monitors such as FluNearYou (by our own Dr. J. Brownstein) have reported an apparent upsurge in influenza-like illnesses over the last week. The CDC has not yet reported the same trend. If the CDC then confirms this early warning, it will represent a transition from proof of principle of the citizen as health monitor (see here and here) to general public health utility. If not, then we may be witnessing an outbreak of hypochondria.

2012-06-13

Dearth of Death: A Fatal Wound to Medical Research?

My esteemed colleague L.J. Wei often reminds us that health outcomes which are not as hard-edged as death can be misleading. For example, the early press, decades ago, about the uncovering of early cancer by the Prostate Specific Antigen (PSA) was used to justify the surgical removal of hundreds of thousands of prostates. In hindsight, neither the PSA test nor much of the ensuing expensive and occasionally morbid surgeries made a significant dent in lifespan.

One might therefore reasonably conclude that the government, the census bureau, Social Security Administration or the Department of Health and Human Services would therefore place the highest premium on the accurate reporting of death, and its causes, for our citizens. Surely, those data are the incontrovertible evidentiary base for our public health monitoring, medical treatment evaluations (whether of drug, device or procedure), and projections of the fundamental demographics of our nation. So, it might be all too easy for most of us to overlook or dismiss the following innocuous-appearing bureaucratese-laden announcement

IMPORTANT NOTICE The National Technical Information Service (NTIS) has been notified by the Social Security Administration (SSA) of an upcoming important change in the Death Master File data. NTIS, a cost-recovery government agency, disseminates the DMF data on behalf of SSA. Please see the attachment, provided by SSA, for an explanation of the change. The implementation date of this change is November 1, 2011. Should you have any questions, please email me at wstrickland@ntis.gov and I will be happy to forward any questions not answered by the attachment to the Social Security Administration for reply.

What does this mean? It means that there is no longer a single, federal authoritative source of death records. Most of the operational details have now devolved to individual states without guarantees of consistency of reporting or a one-stop-shop for researchers looking for the national distribution of the Grim Reaper. Will we have to resort to crowd-sourcing death now in order to perform accurate population research?

Hat tip: Shawn Murphy

Death workflow

2012-03-21

The passing of clean taxonomies.

Among the most productive constructs of the enlightenment are the modern taxonomies. These have been helpful in bringing order to the chaos of signs and symptoms and other clinical findings and were central tools in achieving our 20th century understanding of pathophysiology. They have also have an influential role to play in reimbursement for medical services. With the dawn of high-throughput molecular diagnostics many of us recognize that we are going to be able to be far more precise in our diagnostic and therefore therapeutic approach to diseases and their prevention.

Nonetheless, as we approach the systematization of medicine, we will be reminded often that nature may not hew to the simplified models that we are developing. This recent study in the New England Journal of Medicine, just does that by demonstrating directly that within a "single" tumor there exists a large multiplicity of tumor types, each with its own genomic characteristics and therefore particular therapeutic responsiveness (or lack of it). It can be argued that this is another instance of the tension between the "neats" and the "scruffies" but perhaps it is a foreshadowing of the decreased effectiveness of taxonomies as a cognitive tool for biomedical discovery and clinical care. If indeed, the underlying substructure of physiology is best represented by a probabilistic network model that can only be best grasped and managed through the use of computational tools, we have to seriously re-evaluate both our approach to disease definition and biomedical education.

2012-03-01

City as organism

This video from Geneva, Switzerland is a beautiful instance of the repurposing of data. Shown are the data flows between cell phones across the city over night and day. This glimpse of the interactions over time also suggests new frontiers in real-time epidemiology. What if these (anonymized) data could be tagged with symptoms (e.g. cough, sneeze) could we track the spread of infections? Public health would then start to look a lot more like intensive care medicine, providing real-time monitoring of cities or nations (taking the "pulse" of the population, evaluating the activity and coherence of its "neural" activity). Will there be a new research and medical discipline that fuses the sciences of population ecology and population health? And should populations be empowered to forego such intensive study?

(Hat tip: Joshua Parker).

x

Ville Vivante from Interactive Things on Vimeo.

p.s. As a someone who grew up in Geneva, I was quite surprised to see a lot of activity at 2 AM. Is this the consequence of the Swiss work ethic?

2012-02-16

Research by the numbers

What if you could mine the 10 billion medical facts across 6 million (anonymous) patients in five Harvard affiliated hospitals to ask an important and timely question? What are the other diseases or disorders associated with autism? How has the pharmacological treatment of inflammatory diseases changed over the last five years? Are there gender differences in prevalence of the infections in autoimmune diseases? How is the prevalence of diabetes mellitus changing in young adults?

Now, for the first time, if you are an eligible faculty member (or one of their fellows) in one of the five hospitals, you can now productively seek answers to these questions. The Shared Health Research Information Network (SHRINE) helps researchers overcome one of the greatest problems in population-based research: Compiling large groups of well-characterized patients. Eligible investigators may use the SHRINE web-based query tool to determine the aggregate total number of patients at participating hospitals who meet a given set of inclusion and exclusion criteria. The criteria are currently demographics, diagnoses, medications, and selected laboratory values. Because counts are aggregate, patient privacy is protected.

So, whether you are seeking a study cohort, preliminary studies for a grant proposal, or evaluating an epidemiological hypothesis, take this new tool for a spin and start translating this large mass of hard-won data into useful biomedical knowledge.

2012-01-10

Ancestral betrayal

We share many things with our ancestors, including a fraction of their genetic code. This genetic link to the past has further invigorated an already large industry and hobby in the exploration of genealogy and historical provenance. This piece from today's news, shows how these same records can be used to leverage your ancestors to identify you. In this instance, a murder suspect is potentially fingered by his ancestors from the Mayflower.

Hat tip: Ben Reis

2011-12-21

Crystallizing Methods for Drug Effects

Yet another powerful example of the novel syntheses in which the aggregation of individual knowledge sources becomes a lot more than the sum of its parts: Using only the most well accepted and conventional of compedia of drug characteristics, Cami et al. demonstrate the power of integrative reasoning and systematic applications of "guilt-by-association" to predict accurately, across an interval of 5 years, hundreds of novel adverse events across more than 800 drugs. This performance is likely to improve further as they proceed to integrate healthcare data from electronic health records from consumers. The addition of molecular characterizations are also anticipated to provide further precision and individualization of these predictions.

drugAE

These developments will be further accelerated if we can liberate data gathered with public funding for the public good.

2011-12-14

Alternative Senior Rounds 2011

Let us re-imagine senior rounds for the 21st century.

What I am about to describe does not require any new technologies or biomedical insights; it requires merely a different use of existing resources, different emphases in training and a national focus on the real-time use of clinical data as the evidentiary basis for clinical decision-making.

Senior rounds, in which the department chair meets with the senior residents to review the cases and processes of the prior day, are a decades-old tradition in medicine, and a valuable one. But as currently practiced in most residency programs, each senior resident reporting from their written notes or electronic health record system, the rounds fall far short of what they could be. Here, I offer an “alternative reality” for senior rounds, in hopes of catalyzing a discussion about why it is not the standard of care today. As the citations attest, the findings and techniques described here are all already available. So what are the principal obstacles then to the realization of this scenario?

Soma looked around the table, picking out the Seniors whom he would ask to give reports. He glanced at the screen to the side of the conference table listing the admissions and discharges of the prior day, wait times in the emergency department, diagnoses, and laboratory work ups.

He turned to Charles. “I see you admitted a 2-month-old for ‘rule-out’ meningitis but those laboratory results don’t look particularly worrisome.” Charles nodded but directed his tablet to throw up a local map with 5 red “X”s on the room’s screen1.

“I thought so too,” Charles said, “but the intern pointed out to me the 5 cases of N. meningitis detected in the last week by the State Department of Public Health, including one case in the same day-care center as this infant, so we thought it was prudent. We’re going to wait for culture results.” 2

Soma grunted non-committedly and moved on to Dolores: “I heard you had a little argument with the diabetes service over the discharge of Mr. Smith. Care to share what happened?”

Dolores gave him a quizzical look for this ‘softball.’ “Yes, they were quite emphatic about switching Mr. Smith to a different oral hypoglycemic agent. I argued that its safety profile was far from as well established as the generics in the same structural drug class, and shared several publications with them that made the same point. But it’s only when I showed them that the risk for myocardial infarction, over the last four years, at our hospital was 50% higher for that drug as compared to others that they relented,” 3.

Soma looked at the curve of the myocardial infarction incidence of the drug in question, portrayed in red, and the six green lines showing the myocardial infarction incidence in the same hospital for the other oral hypoglycemic agents. The red curve rose above the green curves, well beyond the reach of their error bars.

He surveyed the seniors and returned to Dolores with a conspiratorial raised eyebrow. “You might want to share with the diabetes service that the FDA just reviewed these data and 20 data sets like it from other academic health centers. They all pointed in the same direction, and when they then reviewed the post-marketing data from the pharmaceutical company manufacturing the drug, the same trend was apparent. Chalk one up for evidence-based medicine. Speaking of which, Harvey, Mrs. Jones’ s headache ended up looking like a glioma on imaging. What are you telling her and her primary care provider about prognosis?”

Harvey directed the screen to replace the map with three graphs. “We’re scheduling the biopsy but it does look like Glioblastoma Multiforme (GBM), less than 2 cm in its largest dimension. The graph on the right shows the outcomes obtained at this hospital over the last 15 years for patients presenting with headaches not attributable to mass effect, like Mrs. Jones. The graph in the middle shows the other patients with GBM at this hospital without this ‘incidental’ presentation. The graph on the right suggests a better outcome but this might be due to the location of these incidental tumors. 4Regardless, it’s a tough prognosis but I shared this perspective with Mrs. Jones and her doctor. By the way, for reference you can see the national outcomes on the leftmost graph and you can see that ours are on average about 20% better as measured by survival times.”

Soma shared the slightest of winks and gestured towards Virginia.

“What about the infant you had discharged from the newborn service a week ago? I see that she was readmitted last night.” Virginia, looked up from the muffin that she had been steadily deconstructing, “Yes, that was unfortunate but not completely unexpected. We had not found any cause for the earlier episode of ventricular tachycardia in the first day of life. Because the tachycardia resolved spontaneously within 20 minutes, we decided to observe for another 72 hours. As there was no recurrence and no structural anomalies of the heart on imaging, we discharged the infant with a follow-up appointment with cardiology for a month from now. The ventricular tachycardia event did trigger an automatic rule from our electronic health record system (EHR) 5,6 which recommended a genetic screen for mutations in the depolarizing sodium and/or calcium channels. We checked the genotyping results on readmission and they are positive for a mutation in a calcium channel—CaCNB2b7—that was found in over one hundred children with ventricular tachycardia as per the National Registry and in no control cases. And by the way, as per our EHR data warehouse this is the fifth case in the last decade in our hospital alone. Although we were able to convert the infant back to sinus rhythm within 10 minutes, the cardiology service is considering use of an implantable cardioverter-defibrillator because of the chanelopathy.”

Soma interjected “Wasn’t the QRS interval abnormal after the first episode?” Virginia flicked the ECG from the EHR view on her tablet to the conference room screen. “No, as you can see, it was not, and there are several similar reports from the literature.” She followed by displaying several PubMed abstracts describing cases of normal ECG in infants with a chanelopathy.

Soma, turned towards the Chief Resident, “Mary Lee, are we going to have enough beds to keep up with all the activity in the ED?”

This question was asked so often that Mary Lee had already displayed the current bed census, as well as the projected lengths of stay based on several morbidity indices and predictors, on the conference room’s screen. “We’re in good shape. Worst case scenario still gives us 32 free beds by noon today8,9. Even with seasonal adjustment for influenza 10 we have at least 8 free beds including 2 in the ICU by the time the evening shift ends. That’s within the 95% confidence interval.”

Charles was glancing repeatedly at the smartphone he kept mostly hidden under the conference table.

“Is there a problem?” Soma asked, girding himself to deliver his well-worn diatribe on the distractions of modern communications.

Charles, stood up, pointing at the smartphone “Actually, there is. The ventilation requirements for one of the preemies is trending higher and the attending pediatrician is suggesting a caffeine infusion but I don’t think it is warranted based on the data. I had better go and check in with the team to see if they are on top of it.” 11

Soma leaned back with a smile. “I should warn you against ‘dismissing long-established clinical opinions without understanding the basis for their existence’12. But go ahead, rounds are over.”

(Thanks to Carey Goldberg for very constructive comments)

1.         Brownstein JS, Freifeld CC, Madoff LC. Digital disease detection--harnessing the Web for public health surveillance. N Engl J Med 2009;360:2153-5, 7.

2.         Fine AM, Nizet V, Mandl KD. Improved diagnostic accuracy of group a streptococcal pharyngitis with use of real-time biosurveillance. Annals of internal medicine 2011;155:345-52.

3.         Brownstein JS, Murphy SN, Goldfine AB, et al. Rapid identification of myocardial infarction risk associated with diabetes medications using electronic medical records. Diabetes Care 2010;33:526-31.

4.         Potts MB, Smith JS, Molinaro AM, Berger MS. Natural history and surgical management of incidentally discovered low-grade gliomas. J Neurosurg 2011.

5.         Ullman-Cullere MH, Mathew JP. Emerging landscape of genomics in the Electronic Health Record for personalized medicine. Human mutation 2011;32:512-6.

6.         Overby CL, Tarczy-Hornoch P, Hoath JI, Kalet IJ, Veenstra DL. Feasibility of incorporating genomic knowledge into electronic medical records for pharmacogenomic clinical decision support. BMC bioinformatics 2010;11 Suppl 9:S10.

7.         Kanter RJ, Pfeiffer R, Hu D, Barajas-Martinez H, Carboni MP, Antzelevitch C. Brugada-Like Syndrome in Infancy Presenting with Rapid Ventricular Tachycardia and Intraventricular Conduction Delay. In: Circulation; 2011.

8.         Mackay M, Lee M. Choice of models for the analysis and forecasting of hospital beds. Health Care Manag Sci 2005;8:221-30.

9.         Littig SJ, Isken MW. Short term hospital occupancy prediction. Health Care Manag Sci 2007;10:47-66.

10.       Reis BY, Pagano M, Mandl KD. Using temporal context to improve biosurveillance. Proceedings of the National Academy of Sciences of the United States of America 2003;100:1961-5.

11.       Larkin H. mHealth. Hosp Health Netw 2011;85:22-6, 2.

12.       Weiss S, Hatcher RA. Tincture of digitalis and the infusion of therapeutics. JAMA 1921;76:508-13.

2011-12-09

Insignificant significance

An amusing look at correlation. Less amusing when you realize that this kind of analysis is often used to drive public debate. Hat tip: Carey Goldberg

etc_correlation50__01__960

2010-10-04

Holding our breath for this diabetes risk

A recent study exemplifies the leverage that can be obtained from mining existing, public data sets to further our national healthcare agenda. As described by the NY Times, our colleague John Brownstein obtained data from the Centers for Disease Control and Prevention (CDC) and U.S. Environmental Protection Agency (EPA) and found a consistent relationship between the amount of air pollution (particulate matter in the air) and population risk for diabetes (after correcting for the usual suspects such as income and ethnicity). This and other large-scale populations studies such as the one we recently reported by Atul Butte suggest that we might be insufficiently including the larger environment in our study of the diabetic plague that has afflicted us.

It also suggests that we have insufficiently taken advantage of freely available public data to pursue relevant and timely medical research.

2010-05-31

What about the environment?

Hundreds of millions of dollars have been spent on comparing the frequencies of genetic variants in disease-afflicted and control populations to suss out the genetic basis of diseases, common or rare. The results for diseases with diabetes have been mixed which despite the high concordance of this disease in identical twins raises yet again the question of the role of the environment in the etiology of diabetes. Of course, we know that the inherited component of diabetes risk is contingent on environmental factors (notably diet) but these are much harder to quantify and moreover there is a universe of environmental risks that is potentially much larger than the entirety of the genome. So how to go about capturing more of the environmental risk as it pertains to real human beings (as opposed to petri dish or rodent studies)?

Atul Butte at Stanford shows us how, through the lens of informatics, systematic approaches can be taken to understand the environmentally-borne determinants of our disease burden. He leverages one of the most admirable public national studies of US citizens, the NHANES study. Butte and colleagues performed analyses analoguously to genome wide association studies but instead examined each of the chemicals (e.g. pesticides, heavy metals, vitamin metabolites) measured in the urine and blood of the members of the NHANES groups comparing cases of diabetes vs. controls. As summarized in the figure below, after comparing multiple groups they found a consistent increased risk of diabetes with PCB and pesticide derivatives and a decreased risk with some carotene derivatives (related to vitamin A). Although, just as for their genome-wide analogues, the results of this Environment Wide Association Study (EWAS) should be taken with due caution, they add to a growing array of methodologies that provide a complement to conventional randomized and controlled studies.

Environement Wide Association Study

2010-02-21

What the Tell-Tale Heart Tells Us About Healthcare

At the end of his short story, Edgar Allen Poe's haunted murderous protagonist lets loose in front of the unsuspecting police officers:

"Villains!" I shrieked, "dissemble no more! I admit the deed! -- tear up the planks! -- here, here! -- it is the beating of his hideous heart!"

It is this story that inspired the title of a paper we published a few years ago on how we could detect the increase (and subsequent decrease) of myocardial infarctions coincident with the rise and fall of the use of Vioxx. This investigation relied solely on the informational byproducts of healthcare delivery. This weekend a Senate report on the risk of Avandia and what was known by GSK came to light. This resonated because we had recently published an article in Diabetes Care about the rapid identification of an increased risk of myocardial infarction with Avandia in patients with diabetes mellitus as compared to other drugs, even drugs in the same class, such as Pioglitazone. Whereas the current headlines are about what the pharmaceutical company knew or dissimulated, there is a broader question that needs addressing: Should every healthcare system not be instrumented so that the clinical leaders of these systems should know whether there are unexpected changes in the risks and health status of their patient populations?

Often, the state of clinical practice is compared unfavorably to the practice of commercial air travel, but what remains underemphasized is that as a system, we are flying blind. There is no local, regional or national air-traffic-controller-equivalent for the healthcare system. Should not the local hospital, and Department of Public Health be the first to know if there is about to be a local health collision or crash? Should not such local surveillance systems run in parallel to regional and national systems? Do we not need multi-level redundancy and open communication to avoid tens of thousands of unnecessary deaths? Further, even as the federal government becomes more aware of the need of such instrumentation, we all expect that our local healthcare systems and health authorities should know of any untoward health trends. Perhaps healthcare systems will start to compete on being able to provide timely and localized health trend data to their customers. Unfortunately right now, the major investments in such health market intelligence is to the payors (i.e. the insurers) who quite reasonably want to know what are the local risks, performance and trends for each of their contracts. Will it take regulation or market competition to make such data extraction and return to patients a matter of course? Let's hope we do not have to wait for a post-mortem Poe'esque orgy of recrimination to find out.

2009-10-26

Screening to distraction: Greater focus on the incidentalome

Last week, Gina Kolata of the New York Times summarized a growing controversy around the value of some of the types of medical screening tests currently employed.The salutory effect of this and related articles is a growing awareness of the tradeoff between increased sensitivity and specificity. It also is a shot across the bow as we contemplate the growth in the number of incidental findings that are going to occur as we test hundreds if not thousands of genetic variants today and in the future. Also, today, Gina Kolata reviewed how little we know about diseases as extensively studied as cancer. Some of them do disappear. Do these spontaneously regressed tumors contribute to the surprisingly high false positive rates for screening?

2009-04-14

Doctors do not bill to make insurance companies smarter

Those of us who have worked with electronic healthcare data have been long aware of the limitations of billing data (aka claims data, aka administrative data) for research. They are often too coarse grained for clinical research and are inherently biased to maximize income. It is motivated by these limitations that Natural Language Processing (NLP) has become increasingly important in mining clinical records for research. What a doctor writes in her notes is much more revealing of her patient's state than what she bills for. Notwithstanding there are some significant challenges in the de-identification of textual records and in transforming these records into standardized clinical categories (e.g. SNOMED). Yet the appeal of using the clinical narrative text rather than claims data is compelling. In our work in i2b2, we have seen significant overrepresentation of diagnostic codes where a diagnostic encounter to "rule out" a disease was codified as that disease in the claims data. For example, a radiologist asked to rule out rheumatoid arthritis based on an X-ray will often classify the X-ray with a billing code corresponding to rheumatoid arthritis when perusal of the full narrative text of the radiologist's notes that there were NO findings consistent with rheumatoid arthritis.

A recent article in the Boston Globe points out additional challenges in using claims data for personal medical records. The same limitations of claims data for research appear to impinge on their utility for clinical care. My colleague John Halamka makes several useful suggestions on how to improve the use of such data, including recruiting patients themselves as collaborators in refining the categorization of their clinical records or even removing gross errors. Notwithstanding, a small number of codes are likely to be quite limiting and it may be that codifying the patient's record by using the entirety of their clinical documentation (i.e. what their care providers wrote about them) will ensure the most nuanced and most faithful representation available of what the clinician was thinking about in each clinical encounter with that patient.

2009-01-08

R U Ready for Open Source Computation?

Perhaps it is completely analogous to the ascendancy of the blogosphere in news reporting, but it is nonetheless shockingly satisfying to see this article on the ascendancy of the R statistical programming language. As a former user of its commercial competitors, I became disenchanted with their licensing policies, not least of which was the hassle of reconfiguring the license manager every time I had to re-install the software on a new laptop. But when I jumped ship, it also quickly became apparent that through the enthusiastic, almost cult-like, devotion of legion programmers and statisticians across the globe, the R software code base provides a far richer set of functionality than any other. Particularly when it comes to rapidly evolving areas of quantitative science such as genomics, R software (especially the huge set of R modules collectively called Bioconductor) have provided the best and most up-to-date functionality. Moreover, the educational programs that have grown around R have been equally first rate. I have recommended this book on learning R for introductory statistics to many colleagues who have then confessed that it was the first time that they had truly understood basic statistics.


I am quite sure that anybody presenting the R "business model" eight years ago would have been scoffed at by most experts in software distribution and dissemination. Yet it is but one example of what can happen if the the producers and consumers of a knowledge or information product are the very same academics. Perhaps we can one day achieve the same efficiencies in disseminating our scholarly publications.






2008-11-24

'Spinning' up the tissues at Harvard

Ever since the late 1990's we have been working on a variety of methods for retrieving data across disparate institutions that are often not even part of the same corporation. In response to an RFA from the National Cancer Institute, called the Shared Pathology Informatics Network, we developed a toolkit that is now in widespread use, specifically to enable genomic and other biological studies on the millions of specimens that are archived across healthcare institutions. As is often the case, we were late in bringing our own tools to use in our own backyard, but that has now happened. The Pathology Specimen Locator (PSL) is now live at our CTSA Catalyst portal. As shown in this screenshot, with authorized credentials, I was able to see that there are over 10,000 lung cancer samples (you can query for any tissue type and disease) across a wide range of ages. It is this sort of IRB-protected, informatics-enabled data liquidity that will accelerate our translational research efforts. Hats off to the entire team but particularly Andy, Frank, John, and Mark.


spin-psl-vsl



Transcriptome of a Trio

Hugh Rienhoff will be speaking about an interesting intersection of expression and genetic transmission studies on December 4th.


Hugh Rienhoff