Post-acquisition CatWalk XT analysis · aggregate results

Interpretable machine learning for classifying time-defined post-injury recovery stage from gait data after spinal cord injury

Condensed methods and principal results for the controls-only cohort, animal-grouped nested validation, held-out feature interpretation, the post-selection three-feature analysis, and the equal-weight ensemble recovery index.

Snapshot currentThe server-published scientific snapshot is ready.
Complete

Study design, analytical population, and intended use

This was a post-acquisition analysis of CatWalk XT recordings from five previous thoracic spinal-cord-injury experiments, with no new intervention, allocation, or outcome acquisition. The models classified contemporaneous, time-defined recovery phase rather than predicting a future event or an independently adjudicated biological recovery state. Repeated observations were grouped by animal throughout validation.

115animals
396observations
5experiments
Full-model discrimination0.83–0.86

Range of pooled out-of-fold mean cumulative AUROC estimates across the three 26-feature classifiers.

Adaptive-analysis discrimination0.85–0.85

Range of pooled out-of-fold mean cumulative AUROC estimates when three-feature selection was repeated within each outer-training partition.

Selection stability1 feature

Selected in every outer fold; 6 appeared at least once.

Recovery-index discriminationMean AUC 0.85

Mean of the Early and Late cumulative AUROCs from the equal-weight fixed-panel ensemble.

Figure conventions

Values are displayed to two decimal places. Ranges summarize the three classifier implementations without selecting a single application framework. All performance values are internal, animal-grouped validation results rather than external replication.

01 · Internal discrimination

Pooled out-of-fold ROC curves for 26- and three-feature analyses.

The Early task distinguishes Early from Middle or Late recovery, and the Late task distinguishes Early or Middle from Late recovery. The three-feature curve represents the adaptive procedure in which feature selection was repeated within each outer-training partition.

Early threshold full-model AUC 0.85 to 0.87Adaptive three-feature procedure AUC 0.84 to 0.85
26 featuresAdaptive three-feature procedureShading = descriptive range across three co-equal frameworks
Median pooled out-of-fold ROC curves across three co-equal frameworks for the full 26-feature analysis and adaptive three-feature procedure. Shaded regions show framework ranges and are not confidence intervals.False-positive rateTrue-positive rate
Curves are based on mutually exclusive outer-test predictions grouped by animal. Each line is the pointwise median across the three classifier implementations; shading is their range, not a confidence interval. The three-feature curve is the fully nested adaptive procedure. AUC is 0.50 at chance and 1.00 at perfect discrimination.

02 · Reduced-feature analysis

Performance estimates for the 26-feature and adaptive three-feature procedures.

The fully nested adaptive procedure produced a descriptive performance range similar to the 26-feature analysis. This dashboard comparison is separate from the manuscript's fixed post-selection three-feature rerun and does not define a non-inferiority or equivalence margin.

Higher is better. This summarizes discrimination across both ordered thresholds.

Full feature set0.830.86
0.84
Animal-grouped nested validation
Adaptive three-feature procedure0.850.85
0.85
Fully nested internal validation

Animal-clustered paired mean-AUROC contrasts: the 95% interval included zero for all three frameworks. Observed mean-AUC changes ranged from -1.6 pp to +1.2 pp.Intervals that include zero do not establish model equivalence.

Each line is the rounded range across the three classifier implementations. Similar-looking values are descriptive and do not establish equivalence.

03 · Predictive feature evidence

Adaptive selection frequencies and fixed-panel consensus ranks.

In the fully nested analysis, 1 feature was selected in every outer fold and 6 features were selected in at least one. These frequencies quantify selection stability and are distinct from a single consensus ranking.

  1. 100%Forepaw Intensity DensityFixed consensus rank 1
  2. 60%Forepaw Loading SpeedNot in fixed top three
  3. 60%Hindpaw Loading RateFixed consensus rank 3
  4. 40%Forepaw Intensity Loading FractionFixed consensus rank 2
  5. 20%Forepaw Peak Loading IndexNot in fixed top three
  6. 20%Forepaw Temporal ContactNot in fixed top three

1 feature appeared in every adaptive fold;1 of these also belonged to the fixed top three.

Adaptive frequencies describe fold-specific selection stability. The manuscript's fixed panel was derived from equal-weight consensus of held-out TreeSHAP and ablation ranks; neither view establishes causal or mechanistic effects.

Harmonization and engineered features

From the canonicalized dataset to the fixed three-feature panel.

Corresponding right and left paw measurements were aggregated into forepaw and hindpaw compartments under the available-paw rule; quality-control and cohort criteria were then applied. Nineteen paw-aggregated inputs were transformed fold-locally into 26 engineered features; the fixed panel comprised the first three features in the equal-weight TreeSHAP–ablation consensus.

  1. 012,101 rows × 113 columnscanonicalized dataset: 110 analysis fields plus 3 row-provenance fields
  2. 0211365 columns96 paw-specific columns replaced by 48 forepaw/hindpaw summaries across 24 metric families; no rows removed
  3. 031,524396 rowsquality-control-retained rows followed by the frozen controls-only cohort restriction
  4. 0419 gait inputsprespecified raw inputs to feature engineering
  5. 0526 engineered featuresprimary full-feature classifier analysis, organized into four physical domains
  6. 063 fixed candidate featuressecondary same-cohort post-selection branch based on three registered features

Loading magnitude and distribution

How recorded intensity and contact-area proxies are distributed between limbs.

  • Hindpaw peak-loading index
    (z(HP maximum intensity) + z(HP 15-pixel intensity)) / 2
  • Forepaw peak-loading index
    (z(FP maximum intensity) + z(FP 15-pixel intensity)) / 2
  • Forepaw intensity shareselected
    FP intensity / (FP intensity + HP intensity + ε)
  • Hindpaw intensity share
    HP intensity / (FP intensity + HP intensity + ε)
  • Forepaw peak-intensity share
    FP maximum intensity / (FP maximum intensity + HP maximum intensity + ε)
  • Hindpaw peak-intensity share
    HP maximum intensity / (FP maximum intensity + HP maximum intensity + ε)
  • Forepaw intensity densityselected
    FP intensity / (FP maximum contact area + ε)
  • Hindpaw intensity density
    HP intensity / (HP maximum contact area + ε)

Dynamic loading and output

Load relative to speed, contact duration, and the area of the print.

  • Hindpaw loading per speed
    HP intensity / (run-average speed + ε)
  • Forepaw loading per speed
    FP intensity / (run-average speed + ε)
  • Hindpaw functional output
    HP intensity × HP print area × HP stance
  • Forepaw functional output
    FP intensity × FP print area × FP stance
  • Forepaw integrated loading
    FP intensity × FP stance
  • Hindpaw integrated loading
    HP intensity × HP stance
  • Forepaw loading rate
    FP intensity / (FP stance + ε)
  • Hindpaw loading rateselected
    HP intensity / (HP stance + ε)

Propulsion and contact timing

How stride length, contact area, and contact duration combine during locomotion.

  • Forepaw propulsion time
    FP stride length / (FP stance + ε)
  • Hindpaw propulsion time
    HP stride length / (HP stance + ε)
  • Forepaw temporal contact
    FP print area × FP stance
  • Hindpaw temporal contact
    HP print area × HP stance

Symmetry and stride mechanics

Fore-hind timing differences and stride output relative to the gait cycle.

  • Stand-time symmetry
    |FP stance − HP stance|
  • Swing-time symmetry
    |FP swing − HP swing|
  • Duty-cycle symmetry
    |FP duty cycle − HP duty cycle|
  • Hindpaw propulsion loading
    HP stride length / (HP intensity + ε)
  • Hindpaw stride per cycle
    HP stride length / (HP stance + HP swing + ε)
  • Forepaw stride per cycle
    FP stride length / (FP stance + FP swing + ε)

Formula key: FP/HP = mean of the available right/left paw measurements (both paws when present; one observed paw otherwise); stance is the measured contact duration; ε = 1×10−8 prevents division by zero. The z() means and standard deviations are learned within the active training partition.

Recovery-index results use the fixed three-feature panel named in the validated snapshot. The compact ROC comparison above uses the separate fully nested adaptive three-feature procedure represented in that snapshot.

04 · Descriptive feature profiles

Selected-feature distributions across predefined recovery phases.

Each point is the median of animal-level phase means, and each vertical bar is the interquartile range. These pooled summaries describe Early, Middle, and Late observations; they are not within-animal longitudinal trajectories, and missing values were not imputed.

Forepaw Intensity Density

A line connects the phase medians and vertical bars show the interquartile range. These are pooled descriptive phase profiles, not within-animal longitudinal trajectories or confidence intervals. Early: median 85.1, interquartile range 74.2 to 101.2, observed in 93 of 93 cohort animals. Middle: median 78.0, interquartile range 63.9 to 90.2, observed in 58 of 58 cohort animals. Late: median 62.6, interquartile range 52.7 to 66.9, observed in 33 of 33 cohort animals.EarlyMiddleLateHarmonized recovery phaseEngineered-feature value

Early: 93/93 animals observedMiddle: 58/58 animals observedLate: 33/33 animals observed

Medians and interquartile ranges are calculated after averaging repeated observations within animal and phase. Coverage varies because hindpaw-dependent values can be unavailable and the contributing animals and experiments differ by phase.

Forepaw Intensity Loading Fraction

A line connects the phase medians and vertical bars show the interquartile range. These are pooled descriptive phase profiles, not within-animal longitudinal trajectories or confidence intervals. Early: median 0.60, interquartile range 0.55 to 0.65, observed in 25 of 93 cohort animals. Middle: median 0.62, interquartile range 0.58 to 0.67, observed in 40 of 58 cohort animals. Late: median 0.60, interquartile range 0.57 to 0.62, observed in 22 of 33 cohort animals.EarlyMiddleLateHarmonized recovery phaseEngineered-feature value

Early: 25/93 animals observedMiddle: 40/58 animals observedLate: 22/33 animals observed

Medians and interquartile ranges are calculated after averaging repeated observations within animal and phase. Coverage varies because hindpaw-dependent values can be unavailable and the contributing animals and experiments differ by phase.

Hindpaw Loading Rate

A line connects the phase medians and vertical bars show the interquartile range. These are pooled descriptive phase profiles, not within-animal longitudinal trajectories or confidence intervals. Early: median 3,710.3, interquartile range 1,273.2 to 4,714.2, observed in 25 of 93 cohort animals. Middle: median 614.4, interquartile range 306.5 to 1,335.7, observed in 36 of 58 cohort animals. Late: median 512.4, interquartile range 350.6 to 1,002.2, observed in 22 of 33 cohort animals.EarlyMiddleLateHarmonized recovery phaseEngineered-feature value

Early: 25/93 animals observedMiddle: 36/58 animals observedLate: 22/33 animals observed

Medians and interquartile ranges are calculated after averaging repeated observations within animal and phase. Coverage varies because hindpaw-dependent values can be unavailable and the contributing animals and experiments differ by phase.

Biological interpretation limits

Hypothesis-generating context for the selected predictive features.

The manuscript treats the three-feature panel as a pragmatic integration of held-out attribution, functional removal, and cross-framework evidence rather than three independent biomarkers. The feature-specific explanations below remain hypotheses for follow-up, not mechanisms tested by this predictive analysis.

fp_intensity_density

Forepaw intensity density

The pooled forepaw-intensity-density profile declines from 85.1 (IQR 74.2–101.2) in Early recovery to 78.0 (63.9–90.2) in Middle and 62.6 (52.7–66.9) in Late recovery. One hypothesis is that this pattern reflects a progressive reduction in how concentrated the recorded forepaw signal is within the maximum contact area, possibly as support or paw placement changes across recovery. Coverage is complete within each phase (93/93, 58/58, 33/33 animals), but the phase-specific cohorts differ in size and composition, so this is not a within-animal trajectory. This is a non-causal optical ratio: acquisition conditions, print segmentation, intensity, and contact area can all affect it, and it is not a direct measure of force or neural recovery.

fp_intensity_loading_fraction

Forepaw intensity share

Forepaw intensity share is broadly stable across the pooled profiles: 0.60 (IQR 0.55–0.65) in Early recovery, 0.62 (0.58–0.67) in Middle, and 0.60 (0.57–0.62) in Late, with substantial interval overlap. One hypothesis is that the small Middle-phase elevation reflects a transient redistribution of the recorded loading signal toward the forepaws, but the non-monotonic profile does not support a simple progressive shift. Coverage is incomplete and changes across phases (25/93, 40/58, 22/33 animals), so missingness and cohort composition may contribute to the apparent pattern. This is a non-causal compositional ratio: higher forepaw intensity, lower hindpaw intensity, or both can increase it, while optical acquisition and paw segmentation remain alternative explanations.

hp_loading_rate

Hindpaw loading rate

The pooled median hindpaw loading rate is 3,710 (IQR 1,273–4,714) in Early recovery, compared with 614.4 (306.5–1,336) in Middle and 512.4 (350.6–1,002) in Late; the Middle and Late intervals overlap. One hypothesis is that the Early profile reflects brief, high-intensity or poorly sustained hindpaw contacts, with later values compatible with longer stance or lower contact-signal intensity rather than simply less hindlimb use. Coverage is incomplete and changes across phases (25/93, 36/58, 22/33 animals), so the contrast cannot be interpreted as a within-animal recovery trajectory. This is a non-causal ratio of optical intensity to segmented stance duration; locomotor speed, stance-event boundaries, and paw detection can affect it, and it is not a measured force-loading rate.

Biological context: CatWalk intensity is an optical contact/loading proxy, and published spinal-cord-injury studies report changes in forelimb support and hindlimb intensity during recovery. See the studies on hindlimb motor recovery after thoracic injury, forelimb and hindlimb gait after cervical injury, and intensity and contact area as loading measures. Independent behavioral, kinematic, electrophysiological, histological, or treatment-response measurements would be required to evaluate these hypotheses.

05 · Ensemble recovery index

Equal-weight ensemble recovery index by observed phase.

The equal-weight ensemble averages Platt-calibrated outputs from the three fixed-panel classifiers. The 0–100 score is derived from reconstructed Middle and Late probabilities and represents expected ordinal class position.

Recovery-index distributions by observed recovery phaseThree aggregate violin plots show the distribution of the ensemble recovery index for early, middle, and late observations. Horizontal bars mark the interquartile range and the central dot marks the median.Recovery-resemblance index (0-100)Earlymedian 17.0 / 175 obs / 93 animalsMiddlemedian 47.2 / 116 obs / 58 animalsLatemedian 70.4 / 105 obs / 33 animals

Index validation

Early threshold AUC
0.87
Late threshold AUC
0.83
Mean AUC
0.85
Violins summarize repeated-observation distributions; their width shows smoothed density, not sample size. The index represents expected ordinal class position, not percentage biological recovery, and inherits the fixed panel's same-cohort post-selection limitation.

Interpretation

The engineered-feature workflow showed useful internal discrimination among time-defined recovery phases.

Across the three classifier implementations, the 26-feature models showed useful pooled out-of-fold discrimination under animal-grouped nested validation. The manuscript's fixed post-selection three-feature rerun retained similar descriptive discrimination and supplied the three components of the equal-weight ensemble recovery index; the dashboard also reports the fully nested adaptive procedure as a separate validation analysis. These outputs support interpretable internal staging but do not constitute biological recovery measurement, treatment-efficacy assessment, or externally validated deployment.