SOURCE results

Ordinal gait assessment after spinal cord injury

73 animals · 162 visits · 24 features
Ordinal XGBoost with animal-grouped nested cross-validation. One T9 experiment; injured, untreated controls.

Updated
Publication status

AUROC 0.854: discrimination across ordered phases

The original primary metric shows useful ranking of postoperative phases in held-out animals. Exact three-phase assignment is less reliable; these measures answer different questions.

Mean cumulative AUROC · primary endpoint
0.854

95% CI 0.796 to 0.903
Ranking across two ordered phase boundaries

Balanced accuracy
55.9%

Mean recall across the three assigned phases

Middle-phase recall
24.4%

10 of 41 Middle visits correctly assigned

AUROC evaluates Middle/Late versus Early and Late versus Early/Middle, averaging the two boundary AUROCs. It assesses probability ranking without selecting a classification threshold. Balanced accuracy uses the single highest-probability phase assigned to each visit; AUROC is not the percentage of visits classified correctly.

Why the Middle phase is difficult to assign

Recovery may progress along a continuum, while the labels divide postoperative time into discrete windows. Transitional gait patterns in the Middle phase may resemble either neighboring phase. Of 41 Middle visits, 22 were classified as Early, 10 as Middle and nine as Late. Useful boundary ranking can therefore coexist with weak Middle-phase recall.

This is a plausible explanation, not proof that biological continuity caused the errors or that low Middle recall is inevitable. Calibration, limited sample support, missing measurements and model limitations may also contribute. Fourteen of 162 visits were misclassified directly between Early and Late. Independent functional outcomes are needed to validate a continuous recovery measure, and within-phase resolution remains uncertain.

Primary model and exploratory probability and trajectory evaluation
EndpointEstimate95% interval
Mean cumulative AUROC0.8540.796 to 0.903
Balanced accuracy55.9%48.0% to 63.8%
Middle-phase recall24.4%11.1% to 40.0%
Ranked probability score0.1410.118 to 0.163
Training-prevalence reference RPS0.2150.195 to 0.234
Within-animal temporal concordance0.7250.628 to 0.814
Within-phase temporal concordance0.6070.467 to 0.743

Lower ranked probability score (RPS) indicates better probability quality. Concordance gives each eligible animal equal weight; 0.5 is the ordering reference. AUROC and balanced accuracy retain their original intervals; other rows use the later saved-prediction evaluation. All intervals use whole-animal bootstrap samples and condition on the fitted models.

Sensitivity to assigned weight adjustment

Two contemporaneous refits used the same animals, features and nested folds. Removing weight adjustment modestly worsened probability quality and improved temporal ordering in the point estimates. Both primary difference intervals include zero.

Primary endpoints of the matched sensitivity
EndpointAdjustedUnadjustedDifference (95% interval)
Visit-weighted RPS0.1360.1420.006 (-0.005 to 0.017)
Within-animal temporal concordance0.6970.7670.070 (-0.004 to 0.154)

Difference = unadjusted minus adjusted. The adjusted refit is distinct from the original primary model. Intervals containing zero do not establish equivalence. Assigned weights came from control records within the same experiment and were not matched measurements of each recipient animal.

Aggregate figures

Original primary performance
Cumulative ROC curves, confusion counts, calibration bins and ordinal errors from the original primary analysis.
Cumulative ROC curves, confusion counts, calibration bins and ordinal errors from the original primary analysis. Numerical estimates and support are available in the evidence tables.
Probabilities and within-animal ordering
Primary and Platt-calibrated probabilities, day-wise summaries, RPS and temporal concordance. Daily means contain changing sets of animals.
Primary and Platt-calibrated probabilities, day-wise summaries, RPS and temporal concordance. Daily means contain changing sets of animals. Numerical estimates and support are available in the evidence tables.
Matched weight sensitivity
Probability quality, temporal ordering, day-wise scores and paired outer-fold RPS for the adjusted and unadjusted refits.
Probability quality, temporal ordering, day-wise scores and paired outer-fold RPS for the adjusted and unadjusted refits. Numerical estimates and support are available in the evidence tables.

Evidence tables

Table S1. Cohort support by postoperative day and phase, with transported-weight summaries
StratumPhaseVisitsAnimalsWeight mean (SD), gWeight range, gWeight support / note
All included visitsAll16273——Post-baseline controls
Day 3Early4545228.7 (24.7)192.7–284.6Donor-supported
Day 7Early3838209.7 (22.9)174.7–271.3Donor-supported
Day 14Middle2020208.7 (33.2)167.3–291.5Donor-supported
Day 21Middle2121242.7 (26.8)202.7–305.5Donor-supported
Day 28Late2323248.9 (27.3)212.1–309.3Donor-supported
Day 35Late1515259.1 (34.4)216.7–327.3Extrapolated
Phase totalEarly8357——Animal counts overlap across phases
Phase totalMiddle4132——Animal counts overlap across phases
Phase totalLate3823——Animal counts overlap across phases

Weights are assigned analysis values derived from control records within the same experiment, not linked measurements for the recipient gait animals. Day 35 uses donor-specific log-linear extrapolation from days 21–28.

No missing visits were imputed. Phase-specific animal counts overlap; the cohort contains 73 unique animals.

Thirty-one animals have one visit; 42 animals support longitudinal analysis.

What the findings support

The relatively high primary AUROC supports ordinal phase discrimination. The lower balanced accuracy and Middle-phase recall limit exact phase assignment. The score describes expected postoperative phase; temporal concordance measures chronological ordering. Independent functional measurements are still needed to validate it as a biological recovery measure.

  • All evaluation is internal to one selected experiment. Earlier dataset screening and the additional analyses make the findings exploratory.
  • Missingness was phase-dependent: 39.5% of engineered values were missing overall, including 58.3% in Early, 26.4% in Middle and 12.6% in Late phase. Medians were learned within each training partition.
  • Intervals exclude full model-development, alternative-fold and donor-assignment uncertainty. Only 42 animals contributed repeated visits.
  • A SOURCE feature-saturation curve was discussed but has not been run. No reduced feature panel or optimum feature count is claimed.

Study provenance and submission status