Range of pooled out-of-fold mean cumulative AUROC estimates across the three 26-feature classifiers.
Post-acquisition CatWalk XT analysis · aggregate results
Interpretable machine learning for classifying time-defined post-injury recovery stage from gait data after spinal cord injury
Condensed methods and principal results for the controls-only cohort, animal-grouped nested validation, held-out feature interpretation, the post-selection three-feature analysis, and the equal-weight ensemble recovery index.
Study design, analytical population, and intended use
This was a post-acquisition analysis of CatWalk XT recordings from five previous thoracic spinal-cord-injury experiments, with no new intervention, allocation, or outcome acquisition. The models classified contemporaneous, time-defined recovery phase rather than predicting a future event or an independently adjudicated biological recovery state. Repeated observations were grouped by animal throughout validation.
Range of pooled out-of-fold mean cumulative AUROC estimates when three-feature selection was repeated within each outer-training partition.
Selected in every outer fold; 6 appeared at least once.
Mean of the Early and Late cumulative AUROCs from the equal-weight fixed-panel ensemble.
Values are displayed to two decimal places. Ranges summarize the three classifier implementations without selecting a single application framework. All performance values are internal, animal-grouped validation results rather than external replication.
01 · Internal discrimination
Pooled out-of-fold ROC curves for 26- and three-feature analyses.
The Early task distinguishes Early from Middle or Late recovery, and the Late task distinguishes Early or Middle from Late recovery. The three-feature curve represents the adaptive procedure in which feature selection was repeated within each outer-training partition.
02 · Reduced-feature analysis
Performance estimates for the 26-feature and adaptive three-feature procedures.
The fully nested adaptive procedure produced a descriptive performance range similar to the 26-feature analysis. This dashboard comparison is separate from the manuscript's fixed post-selection three-feature rerun and does not define a non-inferiority or equivalence margin.
Higher is better. This summarizes discrimination across both ordered thresholds.
Animal-clustered paired mean-AUROC contrasts: the 95% interval included zero for all three frameworks. Observed mean-AUC changes ranged from -1.6 pp to +1.2 pp.Intervals that include zero do not establish model equivalence.
03 · Predictive feature evidence
Adaptive selection frequencies and fixed-panel consensus ranks.
In the fully nested analysis, 1 feature was selected in every outer fold and 6 features were selected in at least one. These frequencies quantify selection stability and are distinct from a single consensus ranking.
- 100%Forepaw Intensity DensityFixed consensus rank 1
- 60%Forepaw Loading SpeedNot in fixed top three
- 60%Hindpaw Loading RateFixed consensus rank 3
- 40%Forepaw Intensity Loading FractionFixed consensus rank 2
- 20%Forepaw Peak Loading IndexNot in fixed top three
- 20%Forepaw Temporal ContactNot in fixed top three
1 feature appeared in every adaptive fold;1 of these also belonged to the fixed top three.
Harmonization and engineered features
From the canonicalized dataset to the fixed three-feature panel.
Corresponding right and left paw measurements were aggregated into forepaw and hindpaw compartments under the available-paw rule; quality-control and cohort criteria were then applied. Nineteen paw-aggregated inputs were transformed fold-locally into 26 engineered features; the fixed panel comprised the first three features in the equal-weight TreeSHAP–ablation consensus.
- 012,101 rows × 113 columnscanonicalized dataset: 110 analysis fields plus 3 row-provenance fields
- 02113 → 65 columns96 paw-specific columns replaced by 48 forepaw/hindpaw summaries across 24 metric families; no rows removed
- 031,524 → 396 rowsquality-control-retained rows followed by the frozen controls-only cohort restriction
- 0419 gait inputsprespecified raw inputs to feature engineering
- 0526 engineered featuresprimary full-feature classifier analysis, organized into four physical domains
- 063 fixed candidate featuressecondary same-cohort post-selection branch based on three registered features
Loading magnitude and distribution
How recorded intensity and contact-area proxies are distributed between limbs.
- Hindpaw peak-loading index(z(HP maximum intensity) + z(HP 15-pixel intensity)) / 2
- Forepaw peak-loading index(z(FP maximum intensity) + z(FP 15-pixel intensity)) / 2
- Forepaw intensity shareselectedFP intensity / (FP intensity + HP intensity + ε)
- Hindpaw intensity shareHP intensity / (FP intensity + HP intensity + ε)
- Forepaw peak-intensity shareFP maximum intensity / (FP maximum intensity + HP maximum intensity + ε)
- Hindpaw peak-intensity shareHP maximum intensity / (FP maximum intensity + HP maximum intensity + ε)
- Forepaw intensity densityselectedFP intensity / (FP maximum contact area + ε)
- Hindpaw intensity densityHP intensity / (HP maximum contact area + ε)
Dynamic loading and output
Load relative to speed, contact duration, and the area of the print.
- Hindpaw loading per speedHP intensity / (run-average speed + ε)
- Forepaw loading per speedFP intensity / (run-average speed + ε)
- Hindpaw functional outputHP intensity × HP print area × HP stance
- Forepaw functional outputFP intensity × FP print area × FP stance
- Forepaw integrated loadingFP intensity × FP stance
- Hindpaw integrated loadingHP intensity × HP stance
- Forepaw loading rateFP intensity / (FP stance + ε)
- Hindpaw loading rateselectedHP intensity / (HP stance + ε)
Propulsion and contact timing
How stride length, contact area, and contact duration combine during locomotion.
- Forepaw propulsion timeFP stride length / (FP stance + ε)
- Hindpaw propulsion timeHP stride length / (HP stance + ε)
- Forepaw temporal contactFP print area × FP stance
- Hindpaw temporal contactHP print area × HP stance
Symmetry and stride mechanics
Fore-hind timing differences and stride output relative to the gait cycle.
- Stand-time symmetry|FP stance − HP stance|
- Swing-time symmetry|FP swing − HP swing|
- Duty-cycle symmetry|FP duty cycle − HP duty cycle|
- Hindpaw propulsion loadingHP stride length / (HP intensity + ε)
- Hindpaw stride per cycleHP stride length / (HP stance + HP swing + ε)
- Forepaw stride per cycleFP stride length / (FP stance + FP swing + ε)
Formula key: FP/HP = mean of the available right/left paw measurements (both paws when present; one observed paw otherwise); stance is the measured contact duration; ε = 1×10−8 prevents division by zero. The z() means and standard deviations are learned within the active training partition.
Recovery-index results use the fixed three-feature panel named in the validated snapshot. The compact ROC comparison above uses the separate fully nested adaptive three-feature procedure represented in that snapshot.
04 · Descriptive feature profiles
Selected-feature distributions across predefined recovery phases.
Each point is the median of animal-level phase means, and each vertical bar is the interquartile range. These pooled summaries describe Early, Middle, and Late observations; they are not within-animal longitudinal trajectories, and missing values were not imputed.
Forepaw Intensity Density
Early: 93/93 animals observedMiddle: 58/58 animals observedLate: 33/33 animals observed
Forepaw Intensity Loading Fraction
Early: 25/93 animals observedMiddle: 40/58 animals observedLate: 22/33 animals observed
Hindpaw Loading Rate
Early: 25/93 animals observedMiddle: 36/58 animals observedLate: 22/33 animals observed
Biological interpretation limits
Hypothesis-generating context for the selected predictive features.
The manuscript treats the three-feature panel as a pragmatic integration of held-out attribution, functional removal, and cross-framework evidence rather than three independent biomarkers. The feature-specific explanations below remain hypotheses for follow-up, not mechanisms tested by this predictive analysis.
Forepaw intensity density
The pooled forepaw-intensity-density profile declines from 85.1 (IQR 74.2–101.2) in Early recovery to 78.0 (63.9–90.2) in Middle and 62.6 (52.7–66.9) in Late recovery. One hypothesis is that this pattern reflects a progressive reduction in how concentrated the recorded forepaw signal is within the maximum contact area, possibly as support or paw placement changes across recovery. Coverage is complete within each phase (93/93, 58/58, 33/33 animals), but the phase-specific cohorts differ in size and composition, so this is not a within-animal trajectory. This is a non-causal optical ratio: acquisition conditions, print segmentation, intensity, and contact area can all affect it, and it is not a direct measure of force or neural recovery.
Forepaw intensity share
Forepaw intensity share is broadly stable across the pooled profiles: 0.60 (IQR 0.55–0.65) in Early recovery, 0.62 (0.58–0.67) in Middle, and 0.60 (0.57–0.62) in Late, with substantial interval overlap. One hypothesis is that the small Middle-phase elevation reflects a transient redistribution of the recorded loading signal toward the forepaws, but the non-monotonic profile does not support a simple progressive shift. Coverage is incomplete and changes across phases (25/93, 40/58, 22/33 animals), so missingness and cohort composition may contribute to the apparent pattern. This is a non-causal compositional ratio: higher forepaw intensity, lower hindpaw intensity, or both can increase it, while optical acquisition and paw segmentation remain alternative explanations.
Hindpaw loading rate
The pooled median hindpaw loading rate is 3,710 (IQR 1,273–4,714) in Early recovery, compared with 614.4 (306.5–1,336) in Middle and 512.4 (350.6–1,002) in Late; the Middle and Late intervals overlap. One hypothesis is that the Early profile reflects brief, high-intensity or poorly sustained hindpaw contacts, with later values compatible with longer stance or lower contact-signal intensity rather than simply less hindlimb use. Coverage is incomplete and changes across phases (25/93, 36/58, 22/33 animals), so the contrast cannot be interpreted as a within-animal recovery trajectory. This is a non-causal ratio of optical intensity to segmented stance duration; locomotor speed, stance-event boundaries, and paw detection can affect it, and it is not a measured force-loading rate.
Biological context: CatWalk intensity is an optical contact/loading proxy, and published spinal-cord-injury studies report changes in forelimb support and hindlimb intensity during recovery. See the studies on hindlimb motor recovery after thoracic injury, forelimb and hindlimb gait after cervical injury, and intensity and contact area as loading measures. Independent behavioral, kinematic, electrophysiological, histological, or treatment-response measurements would be required to evaluate these hypotheses.
05 · Ensemble recovery index
Equal-weight ensemble recovery index by observed phase.
The equal-weight ensemble averages Platt-calibrated outputs from the three fixed-panel classifiers. The 0–100 score is derived from reconstructed Middle and Late probabilities and represents expected ordinal class position.
Index validation
- Early threshold AUC
- 0.87
- Late threshold AUC
- 0.83
- Mean AUC
- 0.85
Interpretation
The engineered-feature workflow showed useful internal discrimination among time-defined recovery phases.
Across the three classifier implementations, the 26-feature models showed useful pooled out-of-fold discrimination under animal-grouped nested validation. The manuscript's fixed post-selection three-feature rerun retained similar descriptive discrimination and supplied the three components of the equal-weight ensemble recovery index; the dashboard also reports the fully nested adaptive procedure as a separate validation analysis. These outputs support interpretable internal staging but do not constitute biological recovery measurement, treatment-efficacy assessment, or externally validated deployment.