Development and Prospective Validation of an AI-Based Diagnostic Model for Hepato-Pancreato-Biliary Diseases
The rapid advancement of artificial intelligence (AI) has expanded its applications in healthcare, particularly in diagnostic assistance, intelligent triage, and patient interaction. Hepatobiliary and pancreatic diseases (such as liver cancer, pancreatic cancer, cirrhosis) are characterized by insidious onset, rapid progression, low early-diagnosis rates, and poor prognosis. However, grassroots medical institutions in China face challenges including physician shortages, variable patient health literacy, and incomplete initial information collection, leading to high misdiagnosis/missed diagnosis risks.
Recent breakthroughs in large language models (LLMs) and multi-agent systems (MAS) offer new solutions. LLMs enable advanced natural language processing, while MAS coordinates specialized agents for complex decision-making. Integrating MAS with medical LLMs could create intelligent pre-consultation systems that systematically collect patient symptoms, risk factors, family history, and lifestyle data to enhance diagnostic efficiency.
This study aims to develop a MAS-based pre-consultation system for hepatobiliary-pancreatic diseases featuring four specialized agents ("guidance agent," "medical history agent," "risk assessment agent," and "summary generation agent"). The system will simulate clinical reasoning to generate structured diagnostic reports for physicians.
Research Objectives:
Develop a specialized multi-agent framework combining LLMs to simulate clinical diagnostic logic and standardize symptom collection Enhance pre-consultation data integrity through intelligent dialogue focusing on key disease indicators Generate structured diagnostic summaries highlighting critical symptoms and risk factors Establish foundation for clinical validation and application through expert evaluation and user feedback This pre-diagnostic tool will assist physicians rather than replace clinical judgment, promoting safe, effective AI applications in early disease screening and tiered healthcare systems.
Study Overview
Status
Status
Conditions
Conditions
Intervention / Treatment
Intervention / Treatment
Detailed Description
With the rapid development of artificial intelligence (AI) technologies, their applications in the healthcare field have become increasingly widespread, particularly demonstrating significant potential in diagnostic assistance, intelligent triage, and patient interaction. Hepatobiliary and pancreatic diseases (e.g., liver cancer, pancreatic cancer, bile duct cancer, cirrhosis) are characterized by insidious onset, rapid progression, low early diagnostic rates, and poor prognosis. Clinical diagnosis and treatment demand high requirements for early recognition and precise evaluation. However, grassroots medical institutions in China commonly face challenges such as insufficient specialist physicians, heterogeneous patient health literacy, and inadequate initial information collection during first visits, leading to a high risk of missed or incorrect diagnoses and delays in optimal intervention timing.
In recent years, large language models (LLMs) have achieved breakthrough advancements in natural language understanding and generation, providing a technical foundation for constructing intelligent, personalized medical dialogue systems. Multi-agent systems (MAS) simulate collaborative workflows among multiple specialized roles, enabling more complex and structured decision-making processes with notable advantages in task decomposition, knowledge integration, and dynamic reasoning. Integrating multi-agent architectures with medical LLMs offers the potential to develop intelligent pre-consultation systems with domain-specific expertise, enabling systematic collection and preliminary analysis of critical patient information-including symptoms, risk factors, family history, and lifestyle-to enhance the efficiency and completeness of medical consultations.
This study aims to develop an intelligent pre-consultation LLM system for hepatobiliary and pancreatic diseases based on a multi-agent architecture. By establishing clearly defined "guidance agents," "medical history collection agents," "risk assessment agents," and "physician summary generation agents," the system will simulate clinical diagnostic logic, proactively guide patients through structured disease presentations, and automatically generate initial diagnostic reports compliant with clinical standards for physician reference. This system seeks to assist grassroots physicians in improving early recognition of hepatobiliary and pancreatic diseases, optimize outpatient resource allocation, and enhance patient healthcare experiences.
Research Objectives:
This study aims to design, develop, and preliminarily validate an intelligent pre-consultation LLM system for hepatobiliary and pancreatic diseases based on a multi-agent architecture, exploring its feasibility and potential value in clinical pre-consultation scenarios. Specific objectives include:
Establishing a specialized multi-agent collaborative framework: Leveraging LLM technology, design a multi-agent system capable of division of labor and collaboration, incorporating modules such as guidance, medical history collection, preliminary risk screening, and report generation. By simulating the diagnostic logic and reasoning pathways of clinical physicians, the system will achieve systematic, structured information collection for hepatobiliary and pancreatic disease-related symptoms.
Enhancing the completeness and accuracy of pre-consultation information: Through intelligent dialogue, proactively guide patients to review and articulate their medical history, prioritizing key risk factors for hepatobiliary and pancreatic diseases-including jaundice, abdominal pain, weight loss, alcohol consumption history, history of viral hepatitis, and family history of cancer-to improve patient self-reporting completeness and reduce information omissions caused by memory biases or unclear expression.
Generating structured initial diagnostic assistance reports: Automatically integrate patient-reported symptoms and medical history into pre-consultation summaries adhering to clinical documentation standards, highlighting critical physical signs and risk alerts. This will assist physicians in rapidly grasping patient profiles, shortening outpatient consultation times, and improving diagnostic efficiency.
Laying the foundation for subsequent clinical validation and application: Through small-scale simulation testing and expert evaluation, preliminarily verify the system's usability and clinical alignment, collect feedback from physicians and patients, and optimize interactive workflows. This will provide scientific evidence and ethical compliance support for future prospective clinical studies and product development.
This research is not intended for direct clinical diagnosis or treatment decisions but rather as an information-assistance tool for physicians prior to consultations. It aims to promote safe, effective, and equitable applications of AI in the early screening of hepatobiliary and pancreatic diseases and tiered healthcare delivery systems.
Study Type
Study Type
Enrollment (Estimated)
Enrollment
Contacts and Locations
Study Contact
Study Contact
- Name: ding yuan, doctor
- Phone Number: 18858101960
- Email: dingyuan@zju.edu.cn
Participation Criteria
Eligibility Criteria
Eligibility Criteria
Ages Eligible for Study
- Adult
- Older Adult
Accepts Healthy Volunteers
Sampling Method
Study Population
Description
Inclusion Criteria:
- Age and Gender: Patients aged 18 to 75 years, of either sex.
- Clinical Diagnosis Requirements: Suspected or confirmed hepatobiliary or pancreatic diseases (e.g., liver cancer, pancreatic cancer, cholangiocarcinoma, cirrhosis) based on preliminary clinical evaluation.
- Ability to provide a complete medical history and symptoms for AI system interaction.
- Cognitive and Physical Capacity:
- Sufficient cognitive function to complete interactions with the AI multi-agent system independently (verified by Mini-Mental State Examination [MMSE] score ≥24).
- Proficiency in Mandarin or English to ensure accurate communication with the system.
- Consent and Compliance: Willingness to participate and provide written informed consent.
- Ability to complete all study procedures, including physician consultations and follow-up assessments.
- Clinical Workflow Compatibility: Scheduled for outpatient consultation at participating healthcare facilities.
Exclusion Criteria:
- Patients with life-threatening conditions requiring immediate intervention (e.g., acute hepatic failure, severe hemorrhage).
- Presence of severe cardiovascular or cerebrovascular diseases that may interfere with study participation.
- Cognitive or Communication Barriers:
- Cognitive impairment (MMSE score <24) or language barriers preventing effective interaction with the AI system.
- Psychiatric disorders or altered mental status affecting decision-making capacity.
- Prior or Concurrent Participation:
- Enrollment in other interventional clinical trials that may confound the study outcomes.
- Current use of experimental diagnostic tools or AI systems outside the study protocol.
- Technical or Logistical Constraints:
- Inability to access or operate electronic devices required for AI system interaction (e.g., touchscreen terminals, mobile apps).
- Lack of stable internet connectivity for system access.
- Ethical or Legal Restrictions: Pregnancy or lactation (to avoid potential risks not directly related to the study).
Study Plan
How is the study designed?
Design Details
Number of groups / cohorts
Cohorts and Interventions
Group / CohortGroup / Cohort |
Intervention / TreatmentIntervention / Treatment |
|---|---|
|
Experimental Group 1
Patients in this group will first complete a full interaction with the multi-agent system described above until the Arbiter confirms the medical record is error-free.
The system will then generate a structured "Case Characteristics" summary and an initial diagnostic recommendation produced by the Oracle.
However, this complete AI-generated output will not be displayed to the subsequent attending physician.
The physician will then conduct an independent routine consultation following standard clinical protocols.
The medical records generated by the physician are solely for maintaining the integrity of clinical workflows and will not be used as evaluation metrics for this study.
The core assessment objective for this group is to evaluate the concordance between the AI-generated final medical records and the predefined gold standard.
|
Patients first complete a full interaction with the multi-agent system until the Arbiter confirms the medical record is error-free.
The system then generates a structured "Case Characteristics" (CC) summary and preliminary diagnostic recommendations via the Oracle Agent.
However, the complete AI output is not displayed to the subsequent attending physician.
The physician conducts an independent consultation following standard clinical protocols, and their medical records are solely used to maintain clinical workflow integrity and are not evaluated as part of this study.
The core assessment objective for this group is the concordance between the AI-generated final medical records and the gold-standard reference.
|
|
Experimental Group 2
Patients in this group will complete the interaction with the multi-agent system and confirm the final "Case Characteristics" (CC).
The system will then push this structured CC summary (excluding the Oracle's diagnostic recommendations to avoid excessive guidance) to the attending physician's electronic workstation in a standardized format.
Prior to the formal consultation, physicians may refer to this summary to adjust their interview priorities, verify information accuracy, or supplement missing details.
The medical records written by the physicians will serve as the primary evaluation metrics for this group.
|
After patients finalize the CC through interaction with the multi-agent system, the structured "Case Characteristics" (excluding Oracle-generated diagnostic advice to avoid over-guidance) are pushed in a standardized format to the corresponding physician's electronic workstation.
Physicians may reference this summary before formal consultation to adjust their questioning focus, verify information accuracy, or supplement missing details.
The physician's final written medical record serves as the primary evaluation object for this group.
|
|
Control Group
This group will exclude AI intervention entirely.
Physicians will independently complete the consultation and documentation from scratch, serving as the baseline reference for evaluating AI system performance.
|
What is the study measuring?
Primary Outcome Measures
Primary Outcome Measures
Outcome Measure |
Measure Description |
Time Frame |
|---|---|---|
|
Concordance Rate Between AI-Generated Medical Records and Gold Standard
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Proportion of AI-generated "Case Characteristics" summaries that match the gold standard (established by expert panels) in key diagnostic elements (e.g., symptom description, risk factors, physical findings).
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
|
Diagnostic Accuracy of Physician Documentation With AI Summary Reference
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Proportion of physician-completed medical records in Experimental Group 2 that meet predefined quality criteria (completeness, diagnostic relevance, alignment with gold standard).
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
|
Proportion of physician medical records meeting predefined quality criteria without AI support
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Proportion of physician-completed medical records that meet predefined quality criteria (completeness, diagnostic relevance, alignment with gold standard) in the absence of AI assistance.
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Secondary Outcome Measures
Secondary Outcome Measures
Outcome Measure |
Measure Description |
Time Frame |
|---|---|---|
|
Average physician consultation duration
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Comparison of average consultation times across study groups.
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
|
Average pre-consultation preparation time for physicians
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Time spent by physicians reviewing AI-generated case summaries before consultation.
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
|
Patient satisfaction measured by validated questionnaires
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Patient-reported experience measured via validated satisfaction questionnaires.
Scores range from 0 to 100; higher scores indicate greater satisfaction, assessing ease of AI interaction and perceived care quality.
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
|
Physician satisfaction with AI workflow
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Physician survey evaluating perceived efficiency, usability, and clinical utility of AI tools.
Survey scores range from 0 to 100; higher survey scores represent better satisfaction with the AI system.
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
|
Frequency of AI-related adverse events and workflow disruptions
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Documentation of AI-related errors, workflow disruptions, or safety concerns.
Count of participants experiencing safety or workflow issues.
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
|
System Usability Scale (SUS) scores for the AI interface
Time Frame: From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
System Usability Scale (SUS) is a standardized usability questionnaire; total score ranges from 0 to 100.
Higher SUS scores correspond to better perceived usability of the AI interface, completed by physicians and patients.
|
From enrollment through completion of the index outpatient visit and chart assessment, assessed up to 7 days
|
Collaborators and Investigators
Sponsor
Sponsor
Investigators
Investigators
- Study Chair: ding yuan, doctor, Second Affiliated Hospital, School of Medicine, Zhejiang University
Publications and helpful links
General Publications
- Zhou LQ, Wang JY, Yu SY, Wu GG, Wei Q, Deng YB, Wu XL, Cui XW, Dietrich CF. Artificial intelligence in medical imaging of the liver. World J Gastroenterol. 2019 Feb 14;25(6):672-682. doi: 10.3748/wjg.v25.i6.672.
- Cao LL, Peng M, Xie X, Chen GQ, Huang SY, Wang JY, Jiang F, Cui XW, Dietrich CF. Artificial intelligence in liver ultrasound. World J Gastroenterol. 2022 Jul 21;28(27):3398-3409. doi: 10.3748/wjg.v28.i27.3398.
- Rompianesi G, Pegoraro F, Ceresa CD, Montalti R, Troisi RI. Artificial intelligence in the diagnosis and management of colorectal cancer liver metastases. World J Gastroenterol. 2022 Jan 7;28(1):108-122. doi: 10.3748/wjg.v28.i1.108.
- Bo Z, Song J, He Q, Chen B, Chen Z, Xie X, Shu D, Chen K, Wang Y, Chen G. Application of artificial intelligence radiomics in the diagnosis, treatment, and prognosis of hepatocellular carcinoma. Comput Biol Med. 2024 May;173:108337. doi: 10.1016/j.compbiomed.2024.108337. Epub 2024 Mar 24.
Study record dates
Study Major Dates
Study Start (Estimated)
Study Start
Primary Completion (Estimated)
Primary Completion
Study Completion (Estimated)
Study Completion
Study Registration Dates
First Submitted
First Submitted
First Submitted That Met QC Criteria
First Submitted That Met QC Criteria
First Posted (Actual)
First Posted
Study Record Updates
Last Update Posted (Actual)
Last Update Posted
Last Update Submitted That Met QC Criteria
Last Update Submitted That Met QC Criteria
Last Verified
Last Verified
More Information
Terms related to this study
Additional Relevant MeSH Terms
Other Study ID Numbers
Other Study ID Numbers
- 2025-1371
Plan for Individual participant data (IPD)
Plan to Share Individual Participant Data (IPD)?
IPD Plan Description
Patient Privacy and Confidentiality:
IPD contains sensitive personal health information (e.g., medical history, genetic data, diagnostic results), which could risk re-identification even after anonymization. Sharing such data might violate ethical obligations under the Declaration of Helsinki and local regulations (e.g., GDPR, HIPAA).
Informed Consent Limitations:
Participants provided consent for data use within the scope of this specific study. Broad sharing of IPD for secondary purposes (e.g., unrelated research) was not explicitly authorized in the consent process, raising ethical and legal concerns.
Data Ownership and Governance:
Data may be governed by institutional or national policies restricting external access (e.g., Chinese data sovereignty laws). Sharing IPD could conflict with agreements between the study sponsor, healthcare providers, and regulatory bodies.
Risk of Misinterpretation:
Contextual or methodological details critical to interpreting IPD (e.g., AI system workfl
Drug and device information, study documents
Studies a U.S. FDA-regulated drug product
Studies a U.S. FDA-regulated device product
This information was retrieved directly from the website clinicaltrials.gov without any changes. If you have any requests to change, remove or update your study details, please contact register@clinicaltrials.gov. As soon as a change is implemented on clinicaltrials.gov, this will be updated automatically on our website as well.
Clinical Trials on Hepatic Disease
-
NCT05116826CompletedLiver Diseases | Moderate Hepatic Impairment | Severe Hepatic Impairment
-
NCT00143546No longer availableHepatic Veno-occlusive Disease
-
NCT03652636Completed
-
NCT02104154Terminated
-
NCT03773887RecruitingComparison of Inflammatory Profiles and Regenerative Potential in Alcoholic Liver Disease (TargetOH)Liver Diseases | Acute on Chronic Hepatic Failure
-
NCT01582087TerminatedAcute and Chronic Hepatic Failure With Developing Coma
-
NCT06449339RecruitingPortal Hypertension | Advanced Chronic Liver Disease | Hepatic Decompensation
-
NCT00358501CompletedSevere Hepatic Veno-Occlusive Disease
-
NCT02310542CompletedLiver Failure, Acute | Chronic Hepatic Failure
-
NCT04058327CompletedHepatic Encephalopathy | Minimal Hepatic Encephalopathy | Overt Hepatic Encephalopathy
Clinical Trials on AI Integration type 1
-
NCT06506825RecruitingThe Study Focuses on Gastric Cancer
-
NCT06275997Not yet recruiting
-
NCT06971471Not yet recruiting
-
NCT07318233Not yet recruitingExercise Training | Motivation for Physical Activity | Motivational Enhancement | Exercise Behavior | Exercise Adherence Challenges
-
NCT06749145Enrolling by invitation
-
NCT07563296Not yet recruitingPTSD and Trauma-related Symptoms
-
NCT07470463RecruitingDifferential Diagnosis | Diagnostic Accuracy
-
NCT07401199RecruitingLocally Advanced Gastric Cancer | Gastric Cancer (GC)
-
NCT07520448Not yet recruitingEpidemiology | Tuberculosis (TB) | Prisoners | TB Infection
-
NCT07559188Not yet recruitingPresbyopia | Keratoconus | Pinhole | Scleral Contact Lenses