Tämä sivu käännettiin automaattisesti, eikä käännösten tarkkuutta voida taata. Katso englanninkielinen versio lähdetekstiä varten.

Feasibility of Vision-Language Models for Detecting Physical Assistance as an Observable Criterion Defining GMFCS Levels in Children Aged 2 to 6 Years With Cerebral Palsy

perjantai 11. syyskuuta 2026 päivittänyt: Jeong Yi Kwon, Samsung Medical Center

Feasibility of a Self-Hosted Vision-Language Model for Detecting Caregiver Contact and External Support, the Observable Criteria That Define GMFCS Levels, From Multi-View Movement Video in Children Aged 2 to 6 Years With Cerebral Palsy: A Two-Center Diagnostic Accuracy Study

The GMFCS sorts children with cerebral palsy into five levels of gross motor function. What mainly separates one level from the next is two things an observer can see: whether another person has to physically help the child move, and whether the child has to bear weight on something external, such as a walker, a support stand, or furniture. Scoring these takes an experienced clinician, and it is hardest between the ages of 2 and 6.

This study asks a narrow question. Can a vision-language model, meaning an artificial-intelligence model that looks at images and answers questions about them, detect those two things from short video clips of a child moving?

Twenty-six children with cerebral palsy aged 2 to 6 years were filmed at two hospitals. Each child performed everyday movements, including walking, crawling, rolling sideways, rising from the floor to standing, and lowering back to the floor, and up to three cameras recorded every attempt at the same time. Sixteen still frames were taken from each recorded attempt, always sixteen, and shown to the model without the child's identity, the name of the movement, or the child's GMFCS level.

The model answers two yes-or-no questions and nothing else: did an adult touch the child, and did the child bear weight on an external object. A fixed rule, not the model, turns those answers into one determination per attempt, namely whether the child performed the movement unaided. These determinations are compared with what a human annotator recorded for the same clips while blind to clinical information.

All rating is done with open-weight models running on hardware the investigators control, and every reported figure is tied to the specific model version that produced it.

The study does not assign a GMFCS level, and it is not built to. It tests whether the two observations the GMFCS itself relies on can be read from video reliably. If they can, they can be placed in front of a clinician as evidence, which is where later work on supporting GMFCS assessment would start.

Tutkimuksen yleiskatsaus

Tila

Aktiivinen, ei rekrytointi

Ehdot

Yksityiskohtainen kuvaus

BACKGROUND

The GMFCS-E&R separates its five levels mainly by two things: what a child cannot do without help from another person, and what a child cannot do without a hand-held mobility device. The wording is explicit in the age bands relevant here. At Level I a child moves in and out of floor sitting and standing "without adult assistance." At Level III a child "may require adult assistance to assume sitting" and needs "adult assistance for steering and turning" when walking with a walker. Adult contact and external support are therefore not proxies chosen for convenience. They are part of the definition.

Scoring them is another matter. It takes an experienced clinician, it is slow, and between 2 and 6 years of age the assignment is known to be difficult.

STUDY DESIGN

This is a two-center feasibility study of a determination made from video. The index test is a binary determination produced by an open-weight vision-language model. The reference standard is a human annotator's record of the same clips. No clinical grade is requested from the model at any point.

Four elements of the design are fixed in advance.

First, analyses are stratified by participant and by movement class. Summary statistics are computed within a stratum and pooled afterward, so movement class cannot vary with the outcome.

Second, each movement attempt is reduced to exactly 16 frames, 8 from each of two camera views. The count is fixed and is never scaled to clip duration.

Third, the model receives the frames with no identifier, no movement name and no GMFCS level, and answers two questions: was the child touched by an adult, and did the child bear load through an external object.

Fourth, a deterministic rule combines the two answers. An attempt counts as performed unaided when both answers are negative.

BLINDING

No masking of intervention assignment applies, because the study has a single group and no assignment step. What is masked is the reading of the data, and it is masked on both sides.

The rater model sees only the extracted image frames. It is given no participant identifier, no movement name, no GMFCS level and no other clinical information, and it is never shown what the reference standard recorded for that attempt. Each request carries one attempt and nothing else, so nothing learned from one attempt can be carried into the next.

The human annotator who produced the reference standard worked from the clips alone, blind to clinical information, and recorded the two features before any model output existed. The stimulus set was fingerprinted and sealed before rating began, and the rating files are recorded as absent in that seal, which is positive evidence that no model output existed at the time the stimuli were fixed.

The index test and the reference standard were therefore read independently of each other. Neither reader saw the other's output.

PARTICIPANTS AND RECORDINGS

Twenty-six children with cerebral palsy were enrolled, 15 at Samsung Medical Center and 11 at Asan Medical Center. Ages ranged from 25 to 72 months, median 48.5 months. Fifteen were male and 11 female. All five GMFCS levels are represented: 6 children at Level I, 5 at Level II, 4 at Level III, 3 at Level IV and 8 at Level V.

Five movement classes were filmed: walking, crawling, side-rolling, floor sit-to-stand and stand-to-floor-sit. The unit of annotation is the movement attempt, filmed by up to three cameras at once. The dataset holds 536 annotated attempts across 1,554 clip files, recorded at 1920 x 1080 and 30 frames per second.

REFERENCE STANDARD

One annotator, blind to clinical information, recorded two features for every attempt: whether a caregiver physically assisted the child, and whether the child used an acrylic support stand or a walker.

ANALYSIS AND IMAGE-INDEPENDENT CONTROLS

A clinical video dataset carries the answer in places that have nothing to do with the images, so every reported figure is set against baselines that read no image content at all: majority class, movement repertoire alone, a clip-duration threshold and scene composition alone.

Two of these baselines are strong in this dataset, and they are the reason the design takes the form it does. The set of movement classes a clinician chose to film for a child recovers the dichotomized GMFCS group in 25 of 26 children, which is why analyses are stratified by movement class. Mean clip duration alone recovers it in 24 of 26, which is why the frame count per attempt is fixed. For the same reason the primary reporting level is the two video-derived determinations rather than a participant-level grade. Participant-level aggregates are reported as descriptive only and always beside their image-independent control.

RATER MODELS

Rating uses open-weight vision-language models run on hardware controlled by the investigators. Each model is identified by name and version, and no accuracy figure is carried beyond the version that produced it. Whether one rater model can be substituted for another is treated as something to measure rather than assume, and agreement between raters on identical image inputs is reported.

SCOPE

The study is designed to show whether adult contact and external support can be read from video by an open-weight vision-language model closely enough to a human annotator to be useful. It does not assign a GMFCS level.

Those two features are part of how the GMFCS defines its levels, so reading them reliably from video contributes to GMFCS assessment directly rather than alongside it. If the feasibility holds, the same two determinations can be put in front of a clinician as video-derived evidence for criteria that are at present judged by eye, and can serve as the input layer for later work on supporting GMFCS assessment between 2 and 6 years of age, the range in which that assessment is most difficult.

Opintotyyppi

Havainnollistava

Ilmoittautuminen (Todellinen)

26

Yhteystiedot ja paikat

Tässä osiossa on tutkimuksen suorittajien yhteystiedot ja tiedot siitä, missä tämä tutkimus suoritetaan.

Opiskelupaikat

      • Seoul, Etelä -Korea, 05505
        • Asan Medical Center
      • Seoul, Etelä -Korea, 06351
        • Samsung Medical Center

Osallistumiskriteerit

Tutkijat etsivät ihmisiä, jotka sopivat tiettyyn kuvaukseen, jota kutsutaan kelpoisuuskriteereiksi. Joitakin esimerkkejä näistä kriteereistä ovat henkilön yleinen terveydentila tai aiemmat hoidot.

Kelpoisuusvaatimukset

Opintokelpoiset iät

  • Lapsi

Hyväksyy terveitä vapaaehtoisia

Ei

Näytteenottomenetelmä

Ei-todennäköisyysnäyte

Tutkimusväestö

Children with cerebral palsy aged 2 to 6 years under the care of the department of physical and rehabilitation medicine at Samsung Medical Center or Asan Medical Center, each with a GMFCS level assigned by the treating clinician. All five GMFCS levels are represented. Children were enrolled as they attended, without random selection.

Kuvaus

Inclusion Criteria:

  • Clinical diagnosis of cerebral palsy
  • Aged 2 to 6 years at the time of video recording
  • Under the care of the department of physical and rehabilitation medicine at Samsung Medical Center or Asan Medical Center
  • A GMFCS level assigned by the treating clinician
  • Written informed consent from a parent or legal guardian for video recording

Exclusion Criteria:

  • Consent for video recording is not given
  • GMFCS level cannot be clearly assessed
  • The child or the parent or legal guardian declines to take part in the study

Withdrawal Criteria:

  • The child or the parent or legal guardian requests withdrawal during the study
  • The collected video is not of adequate quality for analysis
  • The participant's health status changes substantially during the study period

Opintosuunnitelma

Tässä osiossa on tietoja tutkimussuunnitelmasta, mukaan lukien kuinka tutkimus on suunniteltu ja mitä tutkimuksella mitataan.

Miten tutkimus on suunniteltu?

Suunnittelun yksityiskohdat

Kohortit ja interventiot

Ryhmä/Kohortti
Interventio / Hoito
Standardized Video Assessment
All 26 enrolled children. Each child performed five prescribed gross motor tasks (walking, crawling, side-rolling, floor sit-to-stand, stand-to-floor-sit) while being recorded simultaneously by up to three cameras. Every recorded attempt was submitted to the index test, in which 16 extracted frames are read by an open-weight vision-language model. There is no comparison group.
Standardized multi-view video recording of five prescribed gross motor tasks, followed by automated reading of the recordings. From each recorded attempt, 16 frames are extracted, 8 from each of two camera views, and presented to an open-weight vision-language model. The model answers two binary questions: whether an adult touched the child during the attempt, and whether the child bore load through an external object. A fixed rule combines the two answers into a determination of whether the attempt was performed unaided. The model is given no participant identifier, no movement name and no GMFCS level, and is never asked for a clinical grade. All models run on hardware controlled by the investigators.

Mitä tutkimuksessa mitataan?

Ensisijaiset tulostoimenpiteet

Tulosmittaus
Toimenpiteen kuvaus
Aikaikkuna
Accuracy of the model-determined adult-contact call against the human annotator's record
Aikaikkuna: Single video assessment per participant; recordings collected May 2025 to June 2026
For each movement attempt in the fixed 178-attempt floor-transition set, the rater model's binary determination of whether an adult touched the child is compared with the human annotator's blind record for the same attempt. Accuracy is the proportion of attempts in agreement, reported as a count out of 178. Sensitivity and specificity against the same reference are reported alongside. Every rater model is evaluated on identical image inputs, so accuracies are directly comparable, and comparisons between rater models use McNemar's paired test. The human record is the sole reference.
Single video assessment per participant; recordings collected May 2025 to June 2026
Accuracy of the model-determined external-support call against the human annotator's record
Aikaikkuna: Single video assessment per participant; recordings collected May 2025 to June 2026
For each movement attempt in the same fixed 178-attempt set, the rater model's binary determination of whether the child bore load through an external object is compared with the human annotator's blind record of acrylic support stand or walker use. Accuracy is reported as a count out of 178, with sensitivity and specificity against the same reference.
Single video assessment per participant; recordings collected May 2025 to June 2026

Toissijaiset tulostoimenpiteet

Tulosmittaus
Toimenpiteen kuvaus
Aikaikkuna
Stratified concordance between model and human contact determinations within participant and movement strata
Aikaikkuna: Single video assessment per participant; recordings collected May 2025 to June 2026
Concordance is computed inside each (participant, movement class) stratum and pooled across strata, so that neither participant identity nor movement class can contribute to the statistic. The null value is 0.5. The statistic is reported separately for floor transitions and for non-floor movement classes (walking, side-rolling), using an exact stratified conditional test.
Single video assessment per participant; recordings collected May 2025 to June 2026
Agreement between independent rater models on identical image inputs
Aikaikkuna: Single video assessment per participant; recordings collected May 2025 to June 2026
Two rater models that share no developer organization, no language-model backbone and no vision encoder are run on identical image inputs. Their agreement with each other, and the difference in their accuracy against the human record, are reported. Between-model comparison uses McNemar's paired test.
Single video assessment per participant; recordings collected May 2025 to June 2026
Accuracy of image-independent baselines on the same attempts
Aikaikkuna: Single video assessment per participant; recordings collected May 2025 to June 2026
Four baselines that read no image content are computed on the same attempts: majority class, movement repertoire alone, a clip-duration threshold, and scene composition alone. Each reported model accuracy is presented beside these baselines, so that any accuracy attributable to properties of the dataset rather than to the images is visible.
Single video assessment per participant; recordings collected May 2025 to June 2026

Yhteistyökumppanit ja tutkijat

Täältä löydät tähän tutkimukseen osallistuvat ihmiset ja organisaatiot.

Yhteistyökumppanit

Tutkijat

  • Päätutkija: Jeong Yi Kwon, M.D., Ph.D., Samsung Medical Center

Julkaisuja ja hyödyllisiä linkkejä

Tutkimusta koskevien tietojen syöttämisestä vastaava henkilö toimittaa nämä julkaisut vapaaehtoisesti. Nämä voivat koskea mitä tahansa tutkimukseen liittyvää.

Opintojen ennätyspäivät

Nämä päivämäärät seuraavat ClinicalTrials.gov-sivustolle lähetettyjen tutkimustietueiden ja yhteenvetojen edistymistä. National Library of Medicine (NLM) tarkistaa tutkimustiedot ja raportoidut tulokset varmistaakseen, että ne täyttävät tietyt laadunvalvontastandardit, ennen kuin ne julkaistaan ​​julkisella verkkosivustolla.

Opi tärkeimmät päivämäärät

Opiskelun aloitus (Todellinen)

Perjantai 30. toukokuuta 2025

Ensisijainen valmistuminen (Todellinen)

Tiistai 23. kesäkuuta 2026

Opintojen valmistuminen (Arvioitu)

Tiistai 1. joulukuuta 2026

Opintoihin ilmoittautumispäivät

Ensimmäinen lähetetty

Perjantai 11. syyskuuta 2026

Ensimmäinen toimitettu, joka täytti QC-kriteerit

Perjantai 11. syyskuuta 2026

Ensimmäinen Lähetetty (Todellinen)

Torstai 17. syyskuuta 2026

Tutkimustietojen päivitykset

Viimeisin päivitys julkaistu (Todellinen)

Torstai 17. syyskuuta 2026

Viimeisin lähetetty päivitys, joka täytti QC-kriteerit

Perjantai 11. syyskuuta 2026

Viimeksi vahvistettu

Tiistai 1. syyskuuta 2026

Lisää tietoa

Tähän tutkimukseen liittyvät termit

Muut tutkimustunnusnumerot

  • 2025-05-096
  • 2025-0826 (Muu tunniste: Asan Medical Center Institutional Review Board)

Yksittäisten osallistujien tietojen suunnitelma (IPD)

Aiotko jakaa yksittäisten osallistujien tietoja (IPD)?

EI

IPD-suunnitelman kuvaus

The individual participant data in this study are video recordings of identifiable children and cannot be shared.

Lääke- ja laitetiedot, tutkimusasiakirjat

Tutkii yhdysvaltalaista FDA sääntelemää lääkevalmistetta

Ei

Tutkii yhdysvaltalaista FDA sääntelemää laitetuotetta

Ei

Nämä tiedot haettiin suoraan verkkosivustolta clinicaltrials.gov ilman muutoksia. Jos sinulla on pyyntöjä muuttaa, poistaa tai päivittää tutkimustietojasi, ota yhteyttä register@clinicaltrials.gov. Heti kun muutos on otettu käyttöön osoitteessa clinicaltrials.gov, se päivitetään automaattisesti myös verkkosivustollemme .

Tilaa