Asset Details
MbrlCatalogueTitleDetail
Do you wish to reserve the book?
Development and Validation of a Large Language Model–Based System for Medical History-Taking Training: Prospective Multicase Study on Evaluation Stability, Human-AI Consistency, and Transparency
by
Wu, Liping
, Shi, Chujun
, Chen, Xiaoqin
, Liu, Yang
, Zhu, Yiying
, Tan, Haizhu
, Lin, Xiule
, Zhang, Weishan
in
Accuracy
/ Ambient AI Scribes and AI-Driven Documentation Technologies
/ Application programming interface
/ Artificial Intelligence
/ Automation
/ Clinical Competence
/ Clinical medicine
/ Decision making
/ Decomposition
/ Development and Evaluation of Research Methods, Instruments and Tools
/ e-Learning and Digital Medical Education
/ Feedback
/ Female
/ Humans
/ Language
/ Large Language Models
/ Machine Learning
/ Male
/ Medical education
/ Medical history
/ Medical History Taking - methods
/ Medical History Taking - standards
/ mHealth in Medical Education and Training
/ Multimedia
/ New Methods and Approaches in Medical Education
/ Original Paper
/ Pain
/ Patients
/ Preliminary Experiences with New Educational Technology
/ Prospective Studies
/ Quality standards
/ Software
/ Students, Medical
/ Systems design
/ Teaching
/ Theme Issue: ChatGPT and Generative Language Models in Medical Education
/ Virtual Patients
2025
Hey, we have placed the reservation for you!
By the way, why not check out events that you can attend while you pick your title.
You are currently in the queue to collect this book. You will be notified once it is your turn to collect the book.
Oops! Something went wrong.
Looks like we were not able to place the reservation. Kindly try again later.
Are you sure you want to remove the book from the shelf?
Development and Validation of a Large Language Model–Based System for Medical History-Taking Training: Prospective Multicase Study on Evaluation Stability, Human-AI Consistency, and Transparency
by
Wu, Liping
, Shi, Chujun
, Chen, Xiaoqin
, Liu, Yang
, Zhu, Yiying
, Tan, Haizhu
, Lin, Xiule
, Zhang, Weishan
in
Accuracy
/ Ambient AI Scribes and AI-Driven Documentation Technologies
/ Application programming interface
/ Artificial Intelligence
/ Automation
/ Clinical Competence
/ Clinical medicine
/ Decision making
/ Decomposition
/ Development and Evaluation of Research Methods, Instruments and Tools
/ e-Learning and Digital Medical Education
/ Feedback
/ Female
/ Humans
/ Language
/ Large Language Models
/ Machine Learning
/ Male
/ Medical education
/ Medical history
/ Medical History Taking - methods
/ Medical History Taking - standards
/ mHealth in Medical Education and Training
/ Multimedia
/ New Methods and Approaches in Medical Education
/ Original Paper
/ Pain
/ Patients
/ Preliminary Experiences with New Educational Technology
/ Prospective Studies
/ Quality standards
/ Software
/ Students, Medical
/ Systems design
/ Teaching
/ Theme Issue: ChatGPT and Generative Language Models in Medical Education
/ Virtual Patients
2025
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
Do you wish to request the book?
Development and Validation of a Large Language Model–Based System for Medical History-Taking Training: Prospective Multicase Study on Evaluation Stability, Human-AI Consistency, and Transparency
by
Wu, Liping
, Shi, Chujun
, Chen, Xiaoqin
, Liu, Yang
, Zhu, Yiying
, Tan, Haizhu
, Lin, Xiule
, Zhang, Weishan
in
Accuracy
/ Ambient AI Scribes and AI-Driven Documentation Technologies
/ Application programming interface
/ Artificial Intelligence
/ Automation
/ Clinical Competence
/ Clinical medicine
/ Decision making
/ Decomposition
/ Development and Evaluation of Research Methods, Instruments and Tools
/ e-Learning and Digital Medical Education
/ Feedback
/ Female
/ Humans
/ Language
/ Large Language Models
/ Machine Learning
/ Male
/ Medical education
/ Medical history
/ Medical History Taking - methods
/ Medical History Taking - standards
/ mHealth in Medical Education and Training
/ Multimedia
/ New Methods and Approaches in Medical Education
/ Original Paper
/ Pain
/ Patients
/ Preliminary Experiences with New Educational Technology
/ Prospective Studies
/ Quality standards
/ Software
/ Students, Medical
/ Systems design
/ Teaching
/ Theme Issue: ChatGPT and Generative Language Models in Medical Education
/ Virtual Patients
2025
Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy
We have requested the book for you!
Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.
Oops! Something went wrong.
Looks like we were not able to place your request. Kindly try again later.
Development and Validation of a Large Language Model–Based System for Medical History-Taking Training: Prospective Multicase Study on Evaluation Stability, Human-AI Consistency, and Transparency
Journal Article
Development and Validation of a Large Language Model–Based System for Medical History-Taking Training: Prospective Multicase Study on Evaluation Stability, Human-AI Consistency, and Transparency
2025
Request Book From Autostore
and Choose the Collection Method
Overview
History-taking is crucial in medical training. However, current methods often lack consistent feedback and standardized evaluation and have limited access to standardized patient (SP) resources. Artificial intelligence (AI)-powered simulated patients offer a promising solution; however, challenges such as human-AI consistency, evaluation stability, and transparency remain underexplored in multicase clinical scenarios.
This study aimed to develop and validate the AI-Powered Medical History-Taking Training and Evaluation System (AMTES), based on DeepSeek-V2.5 (DeepSeek), to assess its stability, human-AI consistency, and transparency in clinical scenarios with varying symptoms and difficulty levels.
We developed AMTES, a system using multiple strategies to ensure dialog quality and automated assessment. A prospective study with 31 medical students evaluated AMTES's performance across 3 cases of varying complexity: a simple case (cough), a moderate case (frequent urination), and a complex case (abdominal pain). To validate our design, we conducted systematic baseline comparisons to measure the incremental improvements from each level of our design approach and tested the framework's generalizability by implementing it with an alternative large language model (LLM) Qwen-Max (Qwen AI; version 20250409), under a zero-modification condition.
A total of 31 students practiced with our AMTES. During the training, students generated 8606 questions across 93 history-taking sessions. AMTES achieved high dialog accuracy: 98.6% (SD 1.5%) for cough, 99.0% (SD 1.1%) for frequent urination, and 97.9% (SD 2.2%) for abdominal pain, with contextual appropriateness exceeding 99%. The system's automated assessments demonstrated exceptional stability and high human-AI consistency, supported by transparent, evidence-based rationales. Specifically, the coefficients of variation (CV) were low across total scores (0.87%-1.12%) and item-level scoring (0.55%-0.73%). Total score consistency was robust, with the intraclass correlation coefficients (ICCs) exceeding 0.923 across all scenarios, showing strong agreement. The item-level consistency was remarkably high, consistently above 95%, even for complex cases like abdominal pain (95.75% consistency). In systematic baseline comparisons, the fully-processed system improved ICCs from 0.414/0.500 to 0.923/0.972 (moderate and complex cases), with all CVs ≤1.2% across the 3 cases. A zero-modification implementation of our evaluation framework with an alternative LLM (Qwen-Max) achieved near-identical performance, with the item-level consistency rates over 94.5% and ICCs exceeding 0.89. Overall, 87% of students found AMTES helpful, and 83% expressed a desire to use it again in the future.
Our data showed that AMTES demonstrates significant educational value through its LLM-based virtual SPs, which successfully provided authentic clinical dialogs with high response accuracy and delivered consistent, transparent educational feedback. Combined with strong user approval, these findings highlight AMTES's potential as a valuable, adaptable, and generalizable tool for medical history-taking training across various educational contexts.
Publisher
JMIR Publications
Subject
/ Ambient AI Scribes and AI-Driven Documentation Technologies
/ Application programming interface
/ Development and Evaluation of Research Methods, Instruments and Tools
/ e-Learning and Digital Medical Education
/ Feedback
/ Female
/ Humans
/ Language
/ Male
/ Medical History Taking - methods
/ Medical History Taking - standards
/ mHealth in Medical Education and Training
/ New Methods and Approaches in Medical Education
/ Pain
/ Patients
/ Preliminary Experiences with New Educational Technology
/ Software
/ Teaching
/ Theme Issue: ChatGPT and Generative Language Models in Medical Education
This website uses cookies to ensure you get the best experience on our website.