Artificial intelligence (AI) is rapidly becoming part of everyday life—and now it’s making its way into healthcare. But an important question remains:
Can AI actually perform at the level of a trained medical professional?
A recent study featuring Dr. David Kirschenbaum explored this question by testing how well AI—specifically ChatGPT—performs on specialized hand surgery exams.
What is AI doing in healthcare today?
AI is already being used in many areas of medicine, including:
- Analyzing medical images
- Predicting patient outcomes
- Helping personalize treatment plans
With tools like ChatGPT becoming more widely used, researchers are now evaluating how reliable these systems are in real medical scenarios.
How was AI tested in this study?
Researchers evaluated ChatGPT by having it answer real hand surgery exam questions—the same types of questions used to test doctors.
They compared:
- Older versions of AI
- Newer, more advanced versions
- AI with improved instructions (“better prompts”)
- AI with access to medical literature (file search)
The goal was to see how close AI could come to human-level performance.
What the research found
AI is improving quickly
- Older AI models performed significantly worse
- Newer models (like ChatGPT 4o) performed much better
- Giving the AI better instructions improved accuracy even more
AI can match human performance (in some cases)
According to the results shown in the chart on page 4, the latest version of ChatGPT performed about as well as—or slightly better than—human test-takers on text-based questions.
Access to information makes a big difference
When AI was allowed to search medical literature:
- Accuracy improved even further
- Performance reached over 75% on some question sets
This shows that AI works best when it can combine its knowledge with reliable medical sources.
Where AI still has limitations
While results were impressive, AI is not perfect:
- It performed worse on image-based questions (like X-rays or clinical photos)
- It depends heavily on the quality of the information it uses
- It may struggle with complex or controversial medical decisions
As noted in the data on page 5, AI still lagged behind humans when interpreting images.
What this means for patients
AI is not replacing doctors—but it can support them.
In the future, tools like ChatGPT may help:
- Doctors quickly review medical information
- Assist with rare or complex cases
- Support medical education and training
This could lead to faster, more informed decision-making—while still relying on a physician’s expertise.
Why human expertise still matters
Even though AI performed well on exam questions, real-life medicine is more complex.
Doctors consider:
- Patient history
- Physical exams
- Personal preferences
- Nuanced clinical judgment
AI can provide helpful information—but it cannot replace the experience and decision-making of a trained physician.
Dr. Kirschenbaum contributed to this research to better understand how AI can be safely and effectively used in medicine. This work helps guide how new technologies can support doctors while maintaining high standards of patient care.
Read the Full Article
Matching Human Expertise: ChatGPT’s Performance on Hand Surgery Examinations
