General AI vs Domain specific AI #ai #chatgpt #tech

General AI vs Domain specific AI #ai #chatgpt #tech

Testing Medical AI: A Comparative Study

Overview of the Study

  • Scientists tested a $3.5 billion medical AI tool called Open Evidence against regular ChatGPT, finding that it lost in every comparison.
  • Open Evidence is specifically designed for doctors and has recently raised $210 million, with hospitals actively purchasing and using it.
  • Another established tool in the medical field is UpToDate, which has been utilized by doctors for years.

Purpose of Specialized Medical AIs

  • The assumption driving the development of specialized medical AIs is that they would outperform general-purpose models like ChatGPT.
  • NYU conducted studies to test this assumption, focusing on whether specialized tools indeed provide better answers than general AI.

Methodology of the Research

  • NYU ran three studies where they presented 100 questions from actual doctors treating patients to various AI tools including Open Evidence, UpToDate, GPT, Claude, and Gemini.
  • Twelve doctors evaluated the responses without knowing which AI generated them to ensure unbiased grading.

Findings of the Study

  • The results showed that expensive medical-specific tools consistently performed worse than general AIs like ChatGPT, Claude, and Gemini.
  • Medical tools are designed to pull information from databases but can produce inferior answers if incorrect data is retrieved compared to Frontier AI models that have integrated knowledge.

Implications and Conclusions

  • The study highlights a critical gap in testing before deploying these expensive medical AIs in real-world settings; no independent evaluations were conducted prior to their use in hospitals.
  • The paper argues against the assumption that specialized means better and calls for more rigorous testing of AI tools used in healthcare.