Saturday, March 19, 2011

Watson

Share |
In 2005, Nico Schlaefer, a grad student at Carnegie Mellon University, built a statistical query system and wrote a thesis he called Statistical Thought Expansion later named Ephyra. IBM was impressed. Nico worked three summers on Watson. He is now a PhD candidate at CM and an IBM PhD Fellow.
In the Tourette syndrome example given below, Watson was unable to answer until they included more of the symptoms and signs in the database. Q & A as done with Watson seems analogues to Clinical Data & Differential Diagnosis. To make Watsons job easier, enter clinical data in a relational database in simple consistent terms. Likewise, list the sum total of medical diagnostic information in the same simple consistent terms. The relational database can correlate and list the match ups as diagnostic possibilities. A statistical program -- and here is where Watson comes in -- can list the probility of each. Furthermore, a statistical program can conduct an ongoing adjustment to the probable diagnosis based on realtime outcome as determined by subsequent information.
This is not to say that the computer makes the diagnosis, but it does give, at a glance, all of the possibilities. In fact quite the reverse, the statistical program improves its selections and statistics based on the clinician’s evolving and final diagnosis.
In practice, this computer directed diagnostics can be done on an off the shelf database program. Watson may be too hard to move around, and I imagine that the off the shelf database on the clinician’s own computer will be a bit less expensive. The important aspect, however, is still the statistical application. I guess that the articles about the development of Watson do not divulge all of the statistical mechanism, which makes up the AI of Watsons prenominal performance.
Simplicity, however, is the thing that works best with clinicians, and I would bet that there is already a simple statistical application that will function with a relational database. Schlaefer describes source expansion, and for us that source is medical information, all of it -- in simple database terms with criteria of diagnosis.
Found in Probably Irrelevant, from an interview. “Information Retrieval in IBM’s Watson: An interview with Nico Schlaefer,”  Posted on March 17th, 2011 by Jon Elsas
“Nico Schlaefer: Here is a question for which source expansion helped:
What is the name of the rare neurological disease with symptoms such as: involuntary movements (tics), swearing, and incoherent vocalizations (grunts, shouts, etc.)?
This is a question from the TREC 8 evaluation [pdf], but if written as a statement (”This rare neurological disease has symptoms such as …”) I think it could also pass as a Jeopardy! question. The answer is “Tourette syndrome”.
We first tried to answer this question using Wikipedia as a source, and there is indeed an article about “Tourette syndrome” in our copy of Wikipedia, but unfortunately it doesn’t mention most of the keywords in the question and Watson wasn’t able to get the answer. We then expanded Wikipedia, and “Tourette syndrome” was one of the topics that was automatically selected. The expanded article contains the following text passages which, by the way, all come from different websites:
·         Rare neurological disease that causes repetitive motor and vocal tics
·         The first symptoms usually are involuntary movements (tics) of the face, arms, limbs or trunk.
·         Tourette’s syndrome (TS) is a neurological disorder characterized by repetitive, stereotyped, involuntary movements and vocalizations called tics.
·         The person afflicted may also swear or shout strange words, grunt, bark or make other loud sounds.
These passages jointly almost perfectly cover the question keywords. I think the only content word that is not in there is “incoherent”. This made it very easy for Watson to find the answer.”

Malaria

Share |

Julia Kubanek a chemical ecologist at the Georgia Institute of Technology identified a seaweed, a red alga, Callophycus Serratus, in the oceans around Fiji that prevents the Malaria parasite from living and reproducing inside of red blood cells.

Maybe that's why Malaria is not a problem in the Fiji Islands. I thought it was the Kava. :)
http://news.sciencemag.org/sciencenow/2011/02/seaweed-a-source-of-potential.html?ref=hp

Saturday, February 26, 2011

Genomic Medicine, charting a course

Eric Green, director of the National Human Genomic Research Institute writes in Nature, “Charting a course for genetic medicine from base pairs to bedside.”[1] This research perspective comes about as close to must reading as anything in the medical literature. Not much happened in the ten years since the completion of the human genome to improve health care, but significant advances in understanding the complexities and cataloging the data, sets the stage for the next ten years during which genetic information will contribute hugely to the health care of our Nation.
As a participant, I have been away from medicine. When I sold the clinic, I came north to fly in the bush. The closest I came to medicine was the evacuation by floatplane of a fisherman with a gaff hook through his hand from a cove north of Kodiak Island. Away from medicine, however, I had time to think about the problems. I don’t think they have solved them yet, but I am fascinated by the potential for the electronic health record and medical information technology to solve problems of public health and cost as well as more accurate diagnosis and better treatment. I built a differential diagnosis based electronic patient record with Borland’s Paradox database back in the 80s. I did it more or less as a hobby, but it addressed one of the weaknesses in the present diagnostic coding system, the ICDA as it is presently used. It is a problem that still exists. I do not see it addressed adequately in present informatics literature.
The reader may be aware that the US ranks 46th out of 178 countries in infant mortality, --according to the CIA’s research 2009 -- and 37th in life expectancy. These numbers keep getting worse every time I look. There are many causes, by my opinion: access to the system, poor distribution of doctors and a significant population seeking deleterious alternative care or no care because of cost. More importantly, there may be a system problem over diagnosis caused by the requirement for a too early diagnosis in order to justify tests and reimbursement – even a reluctance to consider possibilities for fear of rendering the patient uninsurable. We have incredibly sophisticated treatment algorithms directed towards best evidence, but if the diagnosis is wrong, these guidelines are of little use. Autopsies were once the final word on diagnosis. A hospital was ranked in quality by its autopsy rate, but that is a thing of the past. Even then, there was argument over diagnosis at clinical pathological and morbidity and mortality conferences. Genomics, more than anything else, promises to offer not only a more accurate diagnosis, but also a statistically validated differential diagnosis.
Eric Green’s “course for genomic medicine,” emphasizes the cataloging of DNA: indexing genes underlying rare and common disease, the genomes of pathogens and the mutations in tumors into structured files. The National Human Genome Research Institute (NHGRI) launched a public research consortium, the Encyclopedia of DNA Elements (ENCODE) in September 2003, to carry out a project identifying all functional elements in the human genome.  The relational database will need to correlate the structured files of the genome with similar structured files of all recognized medical diagnoses in order to associate, over time, all parameters of the human genome with human disease. Thus far, the Human Genome Project yields 3,000 monogenic (Mendelian) diseases and some 900 loci and complex multigenic traits. That leaves 98% or more of the remainder unknown as to its function. We are clearly at the beginning of a translational period that is at first learning the correlation between the parameters of patient symptoms, physical findings, tests and genomics to the malady in question.
Just as, the relational database requires complete genomic data, so too, the database requires an indexing of all known human illness, a large order.  ICDA, the current classification of disease falls short in this requirement. CMIT until its discontinuation came close. When we have these two structured databases -- the patient and the total indexing of disease -- we will be able to identify vast amounts of unsuspected relationship between the unknown parts of the human genome and the human condition, predictive and otherwise. There will be years of data mining before valid directed diagnosis becomes a reality. Eric Green predicts 2020 before the data substantially predicts, prevents and treats based on new knowledge.
Over the past ten years, the cost of the human genome has plummeted. Massive parallel DNA sequences shorten the time required as well as cost. We have come a long way in understanding the genetic basis of disease. We recognize bio-information in non-coding DNA as well as the complexity associated with structural change and its role in disease. We recognize the role of the genome in cancer and tumor subtypes, and we do routine pharmacogenetic tests before certain drug treatments.
The NHGRI  goals for 2020 include routine orders for complete genetic profiling, genomics incorporated into the electronic health record (EHR) and education of the clinicians in the use of the information.
Multiple institutions pursue these goals. I put my faith in a relational database correlating statistically the patient record with the index of all medical disease to produce a statistically validated differential diagnosis. Mine is a clinician’s viewpoint, thirty years worth. Other institutions The University of Maryland and others are working with IBM’s Watson. Watson’s ability to read and apply unstructured narrative data may mitigate the need for scrupulous indexing of both medical information and genomics. I hope it works, and look forward to reports of success. In either case, diversity is good. The more avenues pursued in solving our health care problems, the more scientific will be the outcome. Whatever the final strategy, it should stress a continuing educational flow of current medical information to the clinician. That’s where the rubber meets the road and where motivation and information is most critical.
Eric Green’s article goes on to include societal concerns and a next generation of researchers. As I said this perspective should be must reading.

http://www.nature.com/nature/journal/v470/n7333/full/nature09764.html





[1] Nature 470, 10 Feb. 2011, 204-213


Monday, February 21, 2011

Genomics, Watson & Computerized Medicine

Share |

Last week at the advances in Genome Biology and Technology meeting in Florida, Eric Schadt, CSO of Pacific Biosciences in Menlo Park touted a radically new procedure for sequencing. A much faster process, last week Schadt and his team traced the source of the Cholera in Haiti, sequencing five strains of cholera in less than an hour. It would have previously taken a week or more. Uniquely, the Pacific Bioscience machine sequences single molecules of DNA by adding fluorescent labelled bases that flash a defining color as they are added to the DNA strand.1  This technique eliminates averaging and amplification. The company projects a human genome in fifteen minutes by 2013. However, limitations of high cost, lower accuracy, 85%, and the number of sequences that they can read per run all require further evolution.

The genome and molecular biology in general will add vast amounts of raw data to the patient medical record. The implications of this vast database will be largely unknown. The challenge will be to correlate that data with the patient’s outcome as a means of advancing medical knowledge. The computer will correlate the data on an individual, clinical, regional, and presumably national level. Obviously, as the numbers grow with accumulated data over time and collated by region, the certainty of the observations will increase. On the clinic level, a simple statistical correlation over time will add knowledge, but on a regional or National scale, the studies will lead to data mining, unexpected surprises and statistical certainty. 

Enter Watson. “IBM and Nuance Communications announced Thursday a research agreement to explore, develop and commercialize the Watson computing system’s advanced analytics capabilities in the health care industry.” --- “Columbia University Medical Center and the University of Maryland School of Medicine will contribute their medical expertise and research to the collaborative effort.”2

If Watson can come to understand medical narrative, language and terminology, such would obviate the necessity of converting doctor speak into a database format. By comparing patient data with the totality of medical information, Watson can write the book on diagnosis and treatment by its massive correlation between cause and effect. Watson is a game changer, perhaps as significant as the genome. Like the computer, Hal, in Carl Sagan’s 2001, the computer takes on omnipotence in answering the question – any question.

The larger challenge, however, might be in applying the technology. Who can ask, and how much does it cost? Will we continue to bank information behind the walls of the Digital Millennium Copyright Act (DMCA) or sequester knowledge with high cost, professional- access-only?  Will clinicians and thus the patients pay dearly for access from the government, the Exchange, an insurance company, the hospital, a drug company or IBM? How many hands will be in this trough? The cost of current medical information drives the cost of patient care to no small measure. It could get worse. If on the other hand, we make Watson’s memory base affordable to all physicians and associated providers, the positive impact on quality health care will be immeasurable.

In a not too distant time, might Watson 2.0’s massive parallel circuitry answer all 300 million of our questions simultaneously? Could not everyone have access to appropriate medical information? One-viewpoint demands free Information, like free speech for all. Another view argues for professional interpretation. There may be more to Watson than meets the eye, and a hope that IBM will get it right.
 
  1. Nature 470 10Feb 2011 p155
  2. http://www.healthcareitnews.com/news/ibm-nuance-apply-watson-analytics-healthcare%E2%80%A8

Friday, January 28, 2011

Health Information Technology (HIT)

Share |

A preliminary review of the literature on electronic medical records, EMRs and health information technology (HIT) initiatives raise a couple of doubts.

One is the apparent lack of accommodation for new bio-medical and genomic advances. The lack of a current medical diagnostic database with criteria, frustrates the accommodation for new science. Such a comprehensive medical information database needs to be infinitely scalable, dynamic and freely accessable at least by the providers.

Although stage II envisions decision support, -- such is already the case in the best of the EMRs deployed so far -- stage II does not include differential diagnosis. There again the lack of a current medical information terminology / database frustrates any attempt to correlate patient data with outcome or the genome with medical illness.

Also, the plan underestimates the resistance and distrust that both patients and providers might have in relinquishing their proprietary right to the inherent value of the medical record. The hospitals, insurance companies, drug companies and HMOs will vie for control or at least access to this information. I doubt that the federal government instills more trust. The States with their medical schools and public health departments may be neutral ground, but the level of federal access remains unclear.

Obviously we need to do this. We rank behind most of the Western World in health despite having the best medical schools and high level institutions.

The office of the national coordinator (ONC) clearly defines the issue of trust. However the suggestion that the rewards of stage II will yield to penalties in stage III seems contra-productive. On the contrary studies in human behavior suggest that education and access to information motivate far better than punishment. One might say especially with highly educated and motivated professionals.

Sadly, current medical information is harder and more expensive to come by than classified government information. Journals, books and medical seminars are extraordinarily expensive and archaic compared to the Internet, yet firewalls copyright and exorbitant access fees block that information as well.

It seems illogical to view the medical record database as public information (with privacy safe guards) whilst current medical information remains proprietary.

http://motorcycleguy.blogspot.com/2010/12/language-of-healthit.html
http://ahier.blogspot.com/search/label/Office%20of%20the%20National%20Coordinator

Thursday, January 6, 2011

Room-temperature sub-diffraction-limited plasmon laser by total internal reflection

Ren-Min Ma, Rupert F. Oulton, Volker J. Sorger, Guy Bartal & Xiang Zhang
Nature Materials (2010) published online 19 December 2010

“Plasmon lasers are a new class of coherent optical amplifiers that generate and sustain light well below its diffraction limit. Their intense, coherent and confined optical fields can enhance significantly light–matter interactions and bring fundamentally new capabilities to bio-sensing, data storage, photolithography and optical communications.” http://www.nature.com/nmat/journal/vaop/ncurrent/abs/nmat2919.html

The desk top microscope just keeps getting better and better, contributing to the rapid advancements of mediocal knowledge. For every click down into the infintisimal, the scale of information expands exponentionally.

For those who think we just about know it all, one might buy a new microscope. Bio-medicine intersects with physics more and more. There lies an entirely new reality as medical science probes the molecular level and beyond --- below the light defraction limit --- into a world of atoms, particles and quantum mechanics.

Tuesday, January 4, 2011

Electronic Medical Record (EMR)

Quick Pitch
We build electronic medical records (EMR) s in many ways, but until we post patient data including genomic material into databases as discrete data points, it will not be possible to analyze the data in a meaningful way.
Patient records repeat the same words and phrases many times. The chart could be three inches thick, but if one were to reduce it to only the repeated words and phrases and index them in a database, the record might cover only a couple of pages. In a sense, this is compression, but the compressed elements are now accessible and correlated with other information in a relational database.
The same strategy applies to current medical information and terminology. As data points on a modern relational database, specific terms defining diagnosis, criteria or treatment become available for programmed analysis, statistical use, machine logic, artificial intelligence (AI), research and data mining. New biomedical information floods the system beyond the pace of human processing. New medical information is highly perishable difficult to access, expensive and time consuming.  A credible EMR must include a continuously updating database of current medical knowledge.
EMRs strive for many things. One of them involves computer decision support systems (CDSS). A successful decision support offers the clinician diagnostic possibilities, suggestions for further testing, statistical probabilities and treatment options derived from patient data and current medical information not otherwise accessible to the clinician – specifically differential diagnosis. In design, we place far too much emphasis on reimbursement, and treatment and pay not enough attention to patient care and diagnosis.
One clinician in a year will likely produce over a thousand records. The total grows year to year, so after thirty or so years the total will exceed say thirty thousand records. Such a database affords opportunity to correlate data both in real time and retrospectively.  Combine one clinician’s records with others in the region and you have a database exceeding the size of most major studies. The bigger the database, the greater grows the value. Uploading the data anonymously to a related institution, for instance the medical school makes it available for educational focus, CME and ongoing research, even an opportunity to correlate genetic data with real world pathology. Critically, medical information, current diagnostic terms and criteria must flow back down into the clinical computers.
With the government grants for deploying and substantially using EMRs, we have the opportunity to build not just an EMR but also a relational database of medical information (MIDB) that corrects itself based on actual outcome and statistical analysis.
We must keep all of these programs out from between the patient and the clinician, maintaining a sense of humanity and the art of medicine.