CITE: Clinical Inference Tethered to Evidence - a Retrieve-and-verify Layer for AI Care Plans (CITE)
2026年7月16日 更新者:Sanjay Basu、Waymark
A Randomized Controlled Trial of CITE (Clinical Inference Tethered to Evidence), an Evidence-Grounding Retrieve-and-Verify Layer That Flags Unsupported and Inappropriate Recommendations in AI-Generated Care Plans, Versus AI With Safety Guardrails Alone and Unassisted Care, in Medicaid Primary Care
This trial evaluates CITE, a retrieve-and-verify layer that audits an AI-generated care plan against a full-text evidence corpus and flags patient-specific codifiable safety hazards to the clinician.
The co-primary outcomes are how accurately CITE flags these hazards (sensitivity and specificity versus blinded clinician adjudication) and its clinician alert burden and acceptance, compared with AI care plans using safety guardrails alone and with unassisted clinician care, in Medicaid primary care.
調査の概要
状態
まだ募集していません
詳細な説明
Patients are randomized 1:1:1 to (1) unassisted clinician care; (2) AI-generated care plan with safety guardrails; (3) AI-generated care plan with safety guardrails plus CITE.
CITE audits the finalized plan against a frozen, versioned evidence corpus and returns physician-facing flags for patient-specific codifiable safety hazards (a recommended drug contraindicated by this patient's diagnosis or laboratory value; a drug-allergy conflict; a dropped high-risk medication; a guideline-indicated therapy omitted for an active diagnosis; a stated quantity refuted by the corpus), each with a verbatim quote and citation; the clinician retains decision authority.
Randomization uses a deterministic HMAC permuted-block scheme; outcome assessors are blinded to arm.
The co-primary outcomes are (1) the diagnostic accuracy (sensitivity and specificity) of CITE against blinded clinician adjudication, and (2) clinician alert burden (flags per encounter) and acceptance, comparing the CITE arm with the guardrail arm; both are estimable at the enrolled sample size because they do not depend on a rare between-arm event.
The unresolved codifiable-hazard rate by arm is reported as a descriptive secondary: codifiable hazards are infrequent, so the trial is not powered for a between-arm efficacy contrast on hazard reduction.
A prior trial of a different mechanism (a generic deterministic rule-corpus that surfaced roughly 30 or more flags per encounter and was uninformative) was completed with null results and is registered separately; this trial evaluates a materially different, patient-specific intervention and set of outcomes.
Determined exempt by WCG IRB (low risk).
Analysis is pre-registered on OSF (https://doi.org/10.17605/OSF.IO/ENXCW).
研究の種類
介入
入学 (推定)
240
段階
- 適用できない
参加基準
研究者は、適格基準と呼ばれる特定の説明に適合する人を探します。これらの基準のいくつかの例は、人の一般的な健康状態または以前の治療です。
適格基準
就学可能な年齢
- 大人
- 高齢者
健康ボランティアの受け入れ
いいえ
説明
INCLUSION CRITERIA:
- Age 18 years or older.
- Medicaid-enrolled and attributed to a participating Waymark primary care site.
- Primary care encounter that requires clinical reasoning (not administrative-only).
- English-language clinical documentation.
EXCLUSION CRITERIA:
- Age less than 18 years.
- Hospice or palliative-care-exclusive care plan.
- Administrative-only or pharmacy-only encounter that does not surface a clinical decision to the supervising clinician.
- Encounter where the supervising clinician is the principal investigator.
- Enrollment in a competing AI-safety study within the prior 90 days.
研究計画
このセクションでは、研究がどのように設計され、研究が何を測定しているかなど、研究計画の詳細を提供します。
研究はどのように設計されていますか?
デザインの詳細
- 主な目的:ヘルスサービス研究
- 割り当て:ランダム化
- 介入モデル:並列代入
- マスキング:独身
武器と介入
参加者グループ / アーム |
介入・治療 |
|---|---|
|
介入なし:Arm 1: Unassisted care
Clinician develops the care plan without AI assistance.
|
|
|
アクティブコンパレータ:Arm 2: AI with safety guardrails
AI-generated care plan produced with a safety-guardrail system prompt; no CITE.
|
AI-generated care plan produced under a safety-guardrail system prompt.
|
|
実験的:Arm 3: AI with safety guardrails plus CITE
AI-generated care plan with safety guardrails, then passed through CITE, which flags unsupported/inappropriate recommendations with evidence citations for the clinician.
|
AI-generated care plan produced under a safety-guardrail system prompt.
Reads the AI care-plan text and verifies each recommendation/claim against a full-text evidence corpus; returns physician-facing flags (commission/confabulation/unsupported/omission) with verbatim quotes and citations.
Clinician retains decision authority.
|
この研究は何を測定していますか?
主要な結果の測定
結果測定 |
メジャーの説明 |
時間枠 |
|---|---|---|
|
Diagnostic accuracy of CITE against clinician adjudication
時間枠:Day 1 (index primary care encounter)
|
Sensitivity and specificity (with positive and negative predictive values) of the CITE checker for clinically consequential codifiable safety hazards, using blinded clinician adjudication of the plan as the reference standard.
Every plan contributes, so the estimate does not depend on a rare between-arm event.
Exact-binomial 95% confidence intervals; reported overall and by hazard family.
|
Day 1 (index primary care encounter)
|
|
Clinician alert burden (flags surfaced per encounter)
時間枠:Day 1 (index primary care encounter)
|
Number of safety flags surfaced to the clinician per encounter in the CITE arm versus the guardrail arm, with clinician acceptance rate.
Co-primary usability outcome: a verifier that surfaces an unmanageable number of flags is not deployable regardless of sensitivity (the prior-trial mechanism surfaced a median of about 30 per encounter).
Pre-registered acceptability ceiling: median CITE flags per encounter at or below three.
|
Day 1 (index primary care encounter)
|
二次結果の測定
結果測定 |
メジャーの説明 |
時間枠 |
|---|---|---|
|
Clinician action on CITE flags
時間枠:Day 1 (index primary care encounter)
|
Proportion of CITE flags accepted vs overridden by the clinician, by flag type (commission, confabulation, unsupported, omission).
|
Day 1 (index primary care encounter)
|
|
Unresolved codifiable safety-hazard rate by arm (descriptive)
時間枠:Day 1 (index primary care encounter)
|
Proportion of patient-specific codifiable-hazard checkpoints with an unresolved hazard in the finalized plan, by arm, with the Arm 3 minus Arm 2 difference and 95% confidence interval.
Pre-specified as descriptive and hypothesis-generating: codifiable hazards are infrequent, so the trial is not powered for a between-arm efficacy contrast on this measure at the enrolled sample size.
|
Day 1 (index primary care encounter)
|
|
Correction of codifiable hazards within 30 days
時間枠:Up to 30 days after the index encounter
|
Among checkpoints with a hazard in the finalized plan, the proportion acted on and corrected within 30 days (documented resolution, completed referral, or corrected order).
Proximal clinical effectiveness measure.
|
Up to 30 days after the index encounter
|
|
Completed referrals within 30 days
時間枠:Up to 30 days after the index encounter
|
Proportion of initiated referrals completed within 30 days.
|
Up to 30 days after the index encounter
|
|
Clinical safety composite (exploratory)
時間枠:Day 1 (index primary care encounter)
|
Four-component clinical safety composite carried from the prior trial.
Pre-specified as exploratory; underpowered at the planned sample size.
|
Day 1 (index primary care encounter)
|
|
30-day acute care utilization (exploratory)
時間枠:Up to 30 days after the index encounter
|
Emergency department visits and hospitalizations within 30 days.
Exploratory.
|
Up to 30 days after the index encounter
|
協力者と研究者
ここでは、この調査に関係する人々や組織を見つけることができます。
研究記録日
これらの日付は、ClinicalTrials.gov への研究記録と要約結果の提出の進捗状況を追跡します。研究記録と報告された結果は、国立医学図書館 (NLM) によって審査され、公開 Web サイトに掲載される前に、特定の品質管理基準を満たしていることが確認されます。
主要日程の研究
研究開始 (推定)
2026年9月1日
一次修了 (推定)
2027年6月1日
研究の完了 (推定)
2027年9月1日
試験登録日
最初に提出
2026年7月13日
QC基準を満たした最初の提出物
2026年7月16日
最初の投稿 (実際)
2026年7月22日
学習記録の更新
投稿された最後の更新 (実際)
2026年7月22日
QC基準を満たした最後の更新が送信されました
2026年7月16日
最終確認日
2026年7月1日
詳しくは
本研究に関する用語
その他の研究ID番号
- CITE-2026-02
医薬品およびデバイス情報、研究文書
米国FDA規制医薬品の研究
いいえ
米国FDA規制機器製品の研究
いいえ
米国で製造され、米国から輸出された製品。
いいえ
この情報は、Web サイト clinicaltrials.gov から変更なしで直接取得したものです。研究の詳細を変更、削除、または更新するリクエストがある場合は、register@clinicaltrials.gov。 までご連絡ください。 clinicaltrials.gov に変更が加えられるとすぐに、ウェブサイトでも自動的に更新されます。