Large Language Models for Epidural Stimulation Electrode Mapping in Spinal Cord Injury

September 8, 2026 updated by: Görkem Açar, Istanbul Gelisim University

AI-Assisted Electrode Contact Configuration Mapping for Epidural Electrical Stimulation in Spinal Cord Injury: A Comparative Evaluation of Large Language Models

This observational and methodological study aims to compare the performance of large language models in generating electrode contact configuration recommendations for epidural electrical stimulation in spinal cord injury.

Five standardized synthetic spinal cord injury scenarios will be presented to four large language models: ChatGPT-4o, Claude, Grok 3, and Gemini 2.5 Pro. Each model will receive the same standardized prompt. The generated responses will be anonymized and evaluated independently by experts with experience in spinal cord injury rehabilitation and epidural electrical stimulation.

The responses will be assessed in five main areas: clinical accuracy, technical feasibility, safety awareness, consistency with current clinical guidance, and completeness of the response. Agreement between expert evaluators will also be examined.

No real patients, human participants, clinical interventions, or personal health data are included in this study. The study is designed to explore the potential and current limitations of large language models as artificial intelligence-based clinical decision-support tools in neurorehabilitation.

Study Overview

Study Type

Observational

Enrollment (Actual)

20

Contacts and Locations

This section provides the contact details for those conducting the study, and information on where this study is being conducted.

Study Locations

    • Istanbul
      • Istanbul, Istanbul, Turkey (Türkiye), 34290
        • Istanbul Gelisim University

Participation Criteria

Researchers look for people who fit a certain description, called eligibility criteria. Some examples of these criteria are a person's general health condition or prior treatments.

Eligibility Criteria

Ages Eligible for Study

  • Child
  • Adult
  • Older Adult

Accepts Healthy Volunteers

No

Sampling Method

Non-Probability Sample

Study Population

No human study population is included. The analytical sample consists of five standardized synthetic spinal cord injury scenarios and the responses generated for these scenarios by ChatGPT-4o, Claude, Grok 3, and Gemini 2.5 Pro. Each model is evaluated under the same standardized prompting conditions. Model outputs are anonymized and independently rated by expert evaluators for clinical accuracy, technical feasibility, safety awareness, consistency with clinical guidance, and response completeness.

Description

Inclusion Criteria:

  • Responses generated for one of the five predefined standardized synthetic spinal cord injury scenarios.
  • Responses generated using the identical standardized prompt specified in the study protocol.
  • Responses generated by one of the four prespecified large language models.
  • Complete responses available for expert evaluation.

Exclusion Criteria:

  • Responses generated using prompts that differ from the standardized study prompt.
  • Incomplete, interrupted, or technically corrupted model outputs.
  • Duplicate responses or outputs not corresponding to a predefined synthetic scenario.
  • Any response generated using real patient-identifiable or personal health information.

Study Plan

This section provides details of the study plan, including how the study is designed and what the study is measuring.

How is the study designed?

Design Details

Cohorts and Interventions

Group / Cohort
Intervention / Treatment
ChatGPT-4o
Responses generated by ChatGPT-4o for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.
Claude
Responses generated by Claude for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.
Grok 3
Responses generated by Grok 3 for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.
Gemini 2.5 Pro
Responses generated by Grok 3 for five standardized synthetic spinal cord injury scenarios using the same standardized prompt. The responses will be evaluated for clinical accuracy, technical feasibility, safety awareness, guideline consistency, and completeness.
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.

What is the study measuring?

Primary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Clinical Accuracy Score of Large Language Model Responses
Time Frame: At the time of expert evaluation, within 1 week after study initiation
Clinical accuracy of the epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater clinical accuracy of the generated recommendations.
At the time of expert evaluation, within 1 week after study initiation

Secondary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Technical Feasibility Score of Large Language Model Responses
Time Frame: At expert evaluation, within 1 week after study initiation
The technical feasibility of epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater technical feasibility and applicability of the generated recommendations.
At expert evaluation, within 1 week after study initiation
Safety Awareness Score of Large Language Model Responses
Time Frame: At expert evaluation, within 1 week after study initiation
The safety awareness demonstrated in the epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater recognition and consideration of relevant safety issues.
At expert evaluation, within 1 week after study initiation
Clinical Guideline Consistency Score of Large Language Model Responses
Time Frame: At expert evaluation, within 1 week after study initiation
The consistency of the generated epidural electrical stimulation electrode contact configuration recommendations with current clinical guidance will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater consistency with current clinical guidance and relevant evidence-based recommendations.
At expert evaluation, within 1 week after study initiation
Response Completeness Score of Large Language Model Responses
Time Frame: At expert evaluation, within 1 week after study initiation
The completeness of the epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate more complete and comprehensive responses.
At expert evaluation, within 1 week after study initiation

Collaborators and Investigators

This is where you will find people and organizations involved with this study.

Study record dates

These dates track the progress of study record and summary results submissions to ClinicalTrials.gov. Study records and reported results are reviewed by the National Library of Medicine (NLM) to make sure they meet specific quality control standards before being posted on the public website.

Study Major Dates

Study Start (Actual)

September 2, 2026

Primary Completion (Estimated)

September 9, 2026

Study Completion (Estimated)

September 9, 2026

Study Registration Dates

First Submitted

September 8, 2026

First Submitted That Met QC Criteria

September 8, 2026

First Posted (Actual)

September 14, 2026

Study Record Updates

Last Update Posted (Actual)

September 14, 2026

Last Update Submitted That Met QC Criteria

September 8, 2026

Last Verified

September 1, 2026

More Information

Terms related to this study

Drug and device information, study documents

Studies a U.S. FDA-regulated drug product

No

Studies a U.S. FDA-regulated device product

No

This information was retrieved directly from the website clinicaltrials.gov without any changes. If you have any requests to change, remove or update your study details, please contact register@clinicaltrials.gov. As soon as a change is implemented on clinicaltrials.gov, this will be updated automatically on our website as well.

Subscribe