← Clinical Evidence

Gait assessment and prognosis after stroke

This report reviews gait assessment in stroke. It examines measurement properties, interpretation of change and prognostic evidence, with the limits of each study and testing protocol.

In this report
Audited and Updated

Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.

SG-C01

The 45-person derivation cohort included 21 fallers. Index D3 was also externally evaluated in a separate 30-person cohort with nine fallers: AUC 0.84 (95% CI 0.68–1.00), sensitivity 0.56 and specificity 0.95 at the prespecified cutoff of 3.5. Index C was assessed in 54 pooled participants, comprising the original 45 plus nine new patients, so that analysis is not independent external validation. A lower D3 cutoff of 2.5 was selected in the validation sample and needs fresh evaluation. These small samples still limit implementation.

Type: material omission. Audit disposition: supported material omission.

Remaining limit: D3 revised cutoff2.5 was chosen in validation cohort and needs fresh testing. Pooled bootstrap n75 is not the independent-validation denominator.

SG-C02

Wellmon et al. reported an MDC95 of 4.24 WGS points. However, their stated formula and SEM of 1.47 yield 1.96 × √2 × 1.47 = 4.075 points. Retain 4.24 only as a source-reported estimate and flag this discrepancy for clarification; neither number is a validated repeat-capture or patient-important threshold.

Type: source arithmetic inconsistency. Audit disposition: supported source discrepancy.

Remaining limit: Preserve source-reported4.24 and recomputation separately; repeat-video ratings do not establish repeat-capture or patient-important change.

SG-C03

DGI total-score test–retest and inter-rater ICCs were both 0.96. In the full article’s Table 2, item ICCs ranged from 0.56 to 1.00 for test–retest and from 0.55 to 1.00 for inter-rater reliability. The abstract’s 0.55–0.93 range omits the perfect-agreement items.

Type: source abstract body discrepancy. Audit disposition: supported source abstract body conflict.

Remaining limit: Use full-table specificity; this is not a report fabricated-value finding.

SG-C04

The two assessments were seven days apart for 48 of 50 participants and 10 or 13 days apart for the other two; three participants were assessed at different times of day.

Type: protocol precision. Audit disposition: supported protocol precision.

Remaining limit: Small precision update; not a wholesale invalidation of test-retest evidence.

SG-C05

Longitudinal measurements during inpatient rehabilitation also demonstrate that speed and symmetry need not evolve together.

Type: report timing correction. Audit disposition: supported report timing correction.

Remaining limit: Use during inpatient rehabilitation; no later prognostic inference.

SG-C06

After excluding six slow walkers, the study also removed intersystem measurement differences beyond ±1.96 SD before its final accuracy analysis. Such trimming can make the retained error distribution and limits of agreement look better than performance on all eligible observations. The reported agreement should therefore be labelled as agreement after participant and measurement-error exclusions.

Type: material methodological qualification. Audit disposition: supported material methodological qualification.

Remaining limit: Interpretation that outcome-error trimming can improve retained agreement is methodological reasoning; do not infer its unreported quantitative effect.

SG-C07

Ng et al. reported test–retest ICC(3,1) 0.85 and inter-rater ICC(2,2) 0.94. Label the inter-rater coefficient’s average-rating model explicitly; do not assume the same coefficient describes one rater’s score.

Type: measurement model qualification. Audit disposition: supported measurement model precision.

Remaining limit: Preserve average-rating versus single-rating distinction; source narrative also calls measurements single, so report formula/model label exactly rather than operationalize a new one-rater coefficient.

SG-C08

The original article narrative is now available in a ProQuest-supplied reprint. It describes one comfortable-speed trial at each of two sessions 1–3 days apart during the final rehabilitation week, with unchanged physical assistance, walking aid, orthosis and locomotion FIM. The reprint does not expose every original table and figure.

Type: access update. Audit disposition: screened not independently rechecked.

Remaining limit: Do not count this row as an independent primary-source confirmation.

SG-C09

An author-provided full article is now accessible. Eligibility began at three months, but the reported observed duration range began at 0.5 years; this distinction should remain explicit.

Type: access update. Audit disposition: screened not independently rechecked.

Remaining limit: Do not count this row as an independent primary-source confirmation.

SG-C10

The complete PMC article body is now accessible and confirms the eight-person custom-model proof of concept. This changes current source availability, while leaving the original report’s prior retrieval history unverified.

Type: access update. Audit disposition: screened not independently rechecked.

Remaining limit: Do not count this row as an independent primary-source confirmation.

SG-C11

Eighteen participants used a walking aid (seven walkers and eleven canes), and two used a lower-extremity orthosis; the orthosis type was not specified. Replace the report’s AFO-specific wording with lower-limb orthosis.

Type: report protocol overprecision. Audit disposition: supported report protocol overprecision.

Remaining limit: Use lower-limb orthosis; does not establish that the orthoses were not AFOs.

Editorial record

  • Audit status: supported protocol precision. Small precision update; not a wholesale invalidation of test-retest evidence.
  • Edited phrase under SG-C04 . Original wording: P0061
  • Audit status: screened not independently rechecked. Do not count this row as an independent primary-source confirmation.
  • Audit status: supported report protocol overprecision. Use lower-limb orthosis; does not establish that the orthoses were not AFOs.
  • Edited phrase under SG-C11 . Original wording: P0070
  • Audit status: screened not independently rechecked. Do not count this row as an independent primary-source confirmation.
  • Audit status: supported source discrepancy. Preserve source-reported4.24 and recomputation separately; repeat-video ratings do not establish repeat-capture or patient-important change.
  • Audit status: screened not independently rechecked. Do not count this row as an independent primary-source confirmation.
  • Audit status: supported measurement model precision. Preserve average-rating versus single-rating distinction; source narrative also calls measurements single, so report formula/model label exactly rather than operationalize a new one-rater coefficient.
  • Edited phrase under SG-C07 . Original wording: P0140
  • Audit status: supported material methodological qualification. Interpretation that outcome-error trimming can improve retained agreement is methodological reasoning; do not infer its unreported quantitative effect.
  • Edited phrase under SG-C06 . Original wording: P0152
  • Audit status: screened not independently rechecked. Do not count this row as an independent primary-source confirmation.
  • Audit status: supported material methodological qualification. Interpretation that outcome-error trimming can improve retained agreement is methodological reasoning; do not infer its unreported quantitative effect.
  • Edited phrase under SG-C06 . Original wording: P0191
  • Audit status: supported report timing correction. Use during inpatient rehabilitation; no later prognostic inference.
  • Edited phrase under SG-C05 . Original wording: P0196
  • Audit status: supported source discrepancy. Preserve source-reported4.24 and recomputation separately; repeat-video ratings do not establish repeat-capture or patient-important change.
  • Audit status: supported material omission. D3 revised cutoff2.5 was chosen in validation cohort and needs fresh testing. Pooled bootstrap n75 is not the independent-validation denominator.
  • Edited phrase under SG-C01 . Original wording: P0269
  • Edited phrase under SG-C01 . Original wording: P0269
  • Audit status: supported protocol precision. Small precision update; not a wholesale invalidation of test-retest evidence.
  • Audit status: screened not independently rechecked. Do not count this row as an independent primary-source confirmation.
  • Audit status: supported source abstract body conflict. Use full-table specificity; this is not a report fabricated-value finding.
  • Audit status: screened not independently rechecked. Do not count this row as an independent primary-source confirmation.
  • Audit status: supported measurement model precision. Preserve average-rating versus single-rating distinction; source narrative also calls measurements single, so report formula/model label exactly rather than operationalize a new one-rater coefficient.
  • Audit status: supported source discrepancy. Preserve source-reported4.24 and recomputation separately; repeat-video ratings do not establish repeat-capture or patient-important change.
  • Audit status: supported report protocol overprecision. Use lower-limb orthosis; does not establish that the orthoses were not AFOs.
  • Audit status: supported material methodological qualification. Interpretation that outcome-error trimming can improve retained agreement is methodological reasoning; do not infer its unreported quantitative effect.
  • Edited phrase under SG-C06 . Original wording: P0453
  • Audit status: supported material omission. D3 revised cutoff2.5 was chosen in validation cohort and needs fresh testing. Pooled bootstrap n75 is not the independent-validation denominator.

Executive assessment

Gait assessment after stroke is clinically useful, but three claims need different evidence: measuring present walking capacity, detecting meaningful change, and forecasting a later outcome. Standardized walking tests provide a defensible clinical foundation. Instrumented walkways, wearable sensors and video can add information about how walking is achieved. Neither technical sophistication nor a high correlation with a reference instrument establishes an accurate personal prognosis.

The strongest practical measurement strategy is a layered assessment of independence, speed, endurance and walking adaptability, followed by selective examination of movement quality and actual community performance. The Functional Ambulation Categories (FAC), 10-metre walk test (10mWT), 6-minute walk test (6MWT), and a complex-walking measure address related but different questions. Their scores should not be collapsed into one unvalidated recovery or safety score. Consensus recommendations support this separation, while original studies show that error depends on severity, speed, protocol and the particular output being measured. [1–17]

There is credible prospective evidence for forecasting recovery of independent walking. EPOS and TWIST have now undergone external evaluation, including important 2026 studies. Those studies also demonstrate why discrimination is insufficient: TWIST ranked outcomes reasonably in a Japanese multicentre cohort while overestimating the probability of independence at every tested horizon. A positive New Zealand temporal validation required refinement of several original probabilities. These findings support structured, qualified clinical discussion and repeated reassessment, rather than universal numerical promises or restricting rehabilitation because of a low predicted probability. [18–24]

Gait measures also carry information about future falls, but the best examined studies are small, heterogeneous and often exploratory. Some clinical dynamic-balance measures remain informative after adjustment for speed; some instrumented variables add association beyond simple measures. Evidence for a transportable, calibrated gait-only falls calculator is much less established. A recent mobile-video study is directly relevant to rehabilitation technology, but its accessible accepted manuscript contains conflicting fall counts and numerical inconsistencies, in addition to few events and extensive modelling. It should not be used to justify an automated clinical risk threshold. [25–29]

Speed categories are useful descriptive shorthand, not proof that somebody walks independently in the community. The familiar 0.4 and 0.8 m/s categories, revised 0.49 and 0.93 m/s categories, and endurance thresholds derive from different definitions and populations. Cross-sectional classification is repeatedly described as prediction in article titles and clinical summaries. Prospective community-walking studies are more informative for prognosis, but their outcomes range from a questionnaire category to actual daily strides or outings. A person can improve walking capacity without increasing community trips, participation or quality of life. [30–39]

For rehabtools, the most defensible initial claim is reproducible, parameter-specific measurement under a stated protocol. A product should retain affected-side labels, aids and orthoses, assistance, speed instructions, task, recording conditions and failure reasons. Published thresholds should be shown with population and protocol context. A video estimate should enter a published prognostic model only after validation of both the measurement substitution and the complete prediction pipeline in the intended stroke population.

Scope and approach

This report examines adults living with stroke, with particular attention to unilateral hemiparetic gait. It considers clinical walking tests, observational gait scales, instrumented walkways, optical motion capture, wearable inertial sensors, and camera-based or markerless analysis. Its outcomes include measurement validity and reproducibility, absolute measurement error, responsiveness and meaningful change, recovery of independent walking, later falls, community ambulation and participation. It is a critical narrative synthesis, not a registered systematic review, exhaustive evidence inventory, meta-analysis or formal certainty assessment.

The search was broader than the final set of intensively appraised papers. Three explicit PubMed queries covered measurement, prognosis and systematic-review mapping. The Ross Lit Search connector returned all requested pages, but duplicate identifiers meant its page totals did not represent complete unique retrieval. This was corrected by retrieving the exact-query identifier sets through official NCBI ESearch and fetching the complete records through EFetch. The three sets contained 656, 543 and 102 records, with 1,219 unique records after cross-query deduplication. Two records were book chapters and were retained in the retrieval reconciliation, rather than mistaken for missing journal articles. A title-focused Scopus search returned 286 unique records across all twelve pages. Targeted searches, original-reference checking and backward citation chasing supplemented those bounded sets. The search appendix records the exact queries and the difference between retrieval coverage and purposive appraisal.

Original studies were selected to cover the principal measurement decisions, influential thresholds, prospective outcomes and major technological claims. Reviews were used to map the field and identify originals, not as replacements for accessible source papers. Particular effort was made to retrieve original Methods, Results, tables and limitations through Ross Lit Search, PubMed Central, Europe PMC, publishers and lawful repositories. Where a connector reported an abstract or incomplete retrieval, that status was retained. An abstract is not relabelled as full text because metadata services repeat it. Important remaining gaps are listed explicitly. The report therefore has uneven depth across studies: its conclusions about a source are limited to the material actually examined.

Selected high-impact claims were checked against original methods, results and tables, with additional independent source checks. This critical narrative review did not use duplicate independent screening or a formal certainty-grading process. The tables distinguish original prospective cohorts, retrospective cohorts with genuinely later outcomes, current classification, repeated measurement and intervention-derived change analyses. Studies can be relevant to more than one question without proving all of them.

Stroke stage and eligibility

The operational stage framework is hyperacute, first 24 hours; acute, days 1–7; early subacute, after day 7 to 3 months; late subacute, 3–6 months; and chronic, beyond 6 months. Exact time after onset is more informative than the label alone. Some original articles call a cohort chronic from 3 months, or include a participant at 5 months in a nominally chronic sample. Those source definitions should be retained rather than retrospectively harmonized without qualification. The same baseline performance has different implications at day 3, week 6 and year 6.

Stroke type, lesion distribution, recurrence, prestroke disability, cognition and communication, neglect, vision, cardiopulmonary status, lower-limb motor impairment and sensation can all change test feasibility and interpretation. Many studies exclude cerebellar stroke, bilateral lesions, major cognitive difficulty, severe comorbidity or a need for substantial physical assistance. Their findings should not automatically be applied to those groups. Age and sex distributions also matter when samples are small or predominantly male.

The FAC distinguishes human assistance needed for walking; a score of 3 indicates supervision rather than physical assistance, whereas 4 indicates independent level-surface walking. Testability and independence are therefore different. The Stroke Recovery and Rehabilitation Roundtable recommendations use FAC at least 3 for the recommended speed, endurance and complex-walking assessments; below that level, these assessments can be recorded as not testable within that core protocol. Some original studies deliberately test physically assisted walkers. Their results remain valuable, but belong to an explicitly assisted protocol. Do not convert inability to perform a test into a zero speed by default merely because a particular research analysis did so. Record why performance is missing and use appropriate lower-level mobility measures. [1, 3, 6, 40]

Capacity and performance

A short supervised test estimates what a person can do under those conditions. A timed long walk estimates sustained walking capacity. A wearable estimates what the person does during recorded life, subject to wear time and algorithm accuracy. GPS can add where that activity occurs. Self-report adds confidence, perceived limitation and participation. None is a complete substitute for the others.

This distinction is especially important in people who walk quickly in a clinic but rarely leave home, or who use a wheelchair for longer distances while walking indoors. Outdoor walking also requires negotiating kerbs, turns, distractions, obstacles, crowds and transport. Capacity may be necessary without being sufficient. Environmental accessibility, opportunity, support, fatigue, motivation and fear of falling may modify translation into daily performance; a device should not attribute every low step count to neurological impairment. [31–39]

Table 1 Choosing the clinical walking measure

Selection is aligned with the construct distinctions in SRRR 3 [1] and the primary evidence discussed in the report. This is a practical synthesis, not a validated composite battery.

Table 1 Choosing the clinical walking measure
Clinical questionMeasureEssential contextWhat it does not establish
How much human support is requiredFACLevel surface versus complex terrain; distinguish FAC 3 supervision from FAC 4 independent level walkingSpeed, endurance or unrestricted community participation
How fast can the person walkComfortable 10mWT; fast condition when safeTimed versus total distance; start; aid/AFO; assistance; trialsSafe outdoor walking or a universal falls probability
How far can the person sustain walking6MWT; 2MWT when a separate shorter-test question is appropriateCorridor length, turns, instructions, rests and symptomsDirect maximal aerobic capacity; equivalence of 2MWT and 6MWT
Can walking adapt to demandsDGI or FGA plus selected relevant tasksTask version, supervision, item failure and ceiling effectsProspective falls merely because a balance score is abnormal
How is the gait achievedStructured observation plus targeted instrumented variablesParetic side; speed; aid/orthosis; event definitionA causal impairment mechanism or treatment-benefit prediction
What walking happens in daily lifeValidated activity monitoring, optionally GPS and diaryWear time, placement, slow-gait accuracy, walking locationParticipation, choice or motivation from step count alone

Minimum protocol documentation

Every repeated measurement should preserve or document the total and timed distances; standing versus moving start; acceleration and deceleration zones; comfortable versus fast instruction; footwear; walking aid and hand used; ankle-foot orthosis or stimulation; amount and position of therapist assistance; rest and familiarization; trial count and aggregation; surface, corridor length and turning; and adverse symptoms or task failure. Instrumented recordings additionally require sensor or camera locations, sampling rate, synchronization, calibration, gait-event definition, filter and algorithm version, number of valid strides, handling of turns and transitions, and quality-control exclusions.

Consistency is not a demand to withhold a clinically needed aid. It means that an aid change is visible when interpreting change. A faster follow-up walk with a new orthosis can be a useful functional improvement, but it does not isolate recovery of the underlying impairment. Record a standardized comparison condition when clinically safe and a usual-aid condition when the question is current everyday capacity.

Clinical walking tests and the meaning of change

Walking speed is robust only when its protocol is stable

Comfortable and fast speed are simple, interpretable measures. They are not interchangeable: a person may have little safe speed reserve, while someone with relatively good comfortable speed may still struggle with acceleration, endurance or turning. Measuring speed over 3, 5, 6 or 10 metres also changes the influence of starting and stopping. The denominator and timing method should always be explicit.

Flansbjer and colleagues provide a useful chronic-stroke benchmark: fifty people, 6–46 months after stroke, repeated comfortable and fast speed, TUG, stairs and a 6MWT at two assessments seven days apart for 48 participants and 10 or 13 days apart for the other two; three participants were assessed at different times of day. ICCs were high, but the full paper also documents small systematic retest improvements and speed-dependent variability. For comfortable speed, SEM was 0.07 m/s; the reported 95% smallest-real-difference interval was approximately −0.15 to +0.25 m/s around the observed positive bias. For the 6MWT, SEM was 18.6 m and the corresponding interval approximately −37.3 to +66.0 m. Reporting only ICC 0.94–0.99 would hide information essential to judging an individual change. These are measurement-error findings in fairly active chronic survivors, not universal MCIDs. The same fifty people also contributed to a Flansbjer strength-reliability paper discussed in the companion strength report, so the assays are not independent cohort replications. [2]

Fulk and Echternach examined 35 people at a mean 34.5 days after stroke using the middle 5 m of a 9-m comfortable walk. Overall MDC90 was 0.30 m/s, but it was 0.07 m/s among thirteen physically assisted walkers and 0.36 m/s among twenty-two unassisted walkers. This counterintuitive difference does not mean that assistance improves the biological precision of walking. It shows how sample variation, speed range and protocol can alter a reliability-derived threshold. The source is available here at abstract level, so those subgroup values should not be adopted without the full methods. [3]

Lewek and Sykes provide stronger full-text evidence that baseline speed matters. Their 76-person predominantly chronic sample used three comfortable and three fast GAITRite trials at each of two visits, with aids and orthoses held constant. Comfortable-speed MDC95 was 0.10, 0.15 and 0.18 m/s in groups below 0.4, between 0.4 and 0.8, and above 0.8 m/s. The whole-sample ICC was excellent, but subgroup ICCs were lower as the between-person range narrowed. There was a mean 0.05 m/s retest increase, visits were 3–57 days apart, and the fastest subgroup was small. These qualifications favour familiarization and a contextual error band rather than a rigid automated responder label. [4]

Hosoi and colleagues subsequently reported comfortable-speed MDC95 of 0.05, 0.11 and 0.21 m/s in analogous speed strata among 84 people. The protocol used closely repeated 10mWT measurements, so the estimates should not be equated with between-day variability. There are also internal reporting problems: the displayed low-speed group mean is incompatible with the stated grouping threshold, and the description of ICC model and agreement is not fully coherent. The qualitative finding of speed-dependent error is useful; the exact values should be presented as source-specific estimates with those caveats. [5]

A separate 25-person chronic-stroke comparison by Cleland and colleagues demonstrates the importance of apparatus and timing: absolute agreement across stopwatch, electronic mat and measurement-distance conditions was excellent, yet proportional bias remained and maximum speed differed by timing method. Excellent ICC does not prove that methods are interchangeable for serial decisions near a threshold. [41]

Measurement error and meaningful change answer different questions

SEM describes measurement error on the original scale. MDC or smallest detectable change combines error from repeated scores and a stated confidence level. MDC90 and MDC95 are different quantities. An anchor-based minimal important change links a score difference to an external judgement of benefit or functional change. A statistically responsive measure can improve during rehabilitation without having a validated patient-important threshold. An improvement larger than MDC supports change beyond estimated error under a matching protocol; it does not by itself establish that the change mattered to the patient.

Tilson and colleagues estimated 0.16 m/s as an important comfortable-speed increase between approximately 20 and 60 days after stroke in 283 LEAPS participants. The anchor was at least one modified Rankin Scale category improvement; sensitivity was 73.9% and specificity 57.0%. Baseline walking was markedly impaired. Usual aids or orthoses and up to maximal assistance by one person were permitted, and some unable to complete 10 m were assigned zero within the study. The value is therefore a contextual reference for early recovery, not a universal threshold for independent chronic walkers. The anchor was broader disability, not a patient-specific walking goal. [6]

Fulk and colleagues found estimates of 0.175 m/s using the participant’s global rating of walking change and 0.190 m/s using the therapist’s rating in outpatient rehabilitation, starting at a mean 56 days after stroke. Different anchors produced different values in the same setting. Bohannon and colleagues also examined inpatient change using functional assistance-related anchors. These studies add clinically relevant perspectives, but do not justify averaging thresholds into a single stroke MCID. [7, 8]

Hayashi and colleagues add acute-stage evidence that should not be compressed into the abstract’s 0.18–0.25 m/s range. The complete 2022 multicentre paper analysed speed change in 62 initially ambulatory patients, assessed at a mean four days after onset and about twelve days later. The protocol used a comfortable 10-m walk with 3-m acceleration/deceleration zones; eighteen used a walking aid (seven walkers and eleven canes) and two a lower-limb orthosis, with orthosis type unspecified. The therapist global-rating ROC estimate was 0.21 m/s, AUC 0.76 (95% CI 0.64–0.88), with sensitivity 0.78 and specificity 0.68. Patient-rated and motor-FIM anchors had AUCs of 0.68 and 0.66, so no overall ROC threshold was derived from those anchors. [42]

The 0.18–0.25 m/s range instead includes change-difference estimates, which are not the same estimand as the ROC threshold. Distribution-based MDC95 was 0.13 m/s using a 44-person retest subset; the 0.15 m/s half-SD estimate is another distribution statistic, not proof of patient importance. Retest timing and model were not fully clear. People unable to walk initially but able later were excluded from the speed-change analysis, so a major recovery milestone is not represented by these thresholds. This walking subset belongs to the same study as the BBS findings in the companion balance report; it does not constitute an independent cohort replication. [42]

An app can usefully display both change and the context of a reference estimate. It should avoid labelling a change just below MDC as definitively absent or a change just above MCID as definitively meaningful. Sampling uncertainty, baseline function, measurement bias and the relevance of the anchor remain. Patient goals and the actual performance achieved should accompany the number.

Table 2 Selected change estimates retain their original context

MDC90 and MDC95 are not interchangeable. Important change depends on the anchor. No row is a generic treatment-success or safety threshold.

Table 2 Selected change estimates retain their original context
StudyPeople and intervalEstimateTypeKey restriction
Lewek and Sykes [4]76; largely chronic; 3–57 days between visits0.10 / 0.15 / 0.18 m/s by comfortable-speed stratumMDC95Three GAITRite trials; systematic retest improvement; not patient importance
Fulk and Echternach [3]35; mean 34.5 days after stroke0.30 m/s overall; 0.07 assisted; 0.36 unassistedMDC90Middle 5 m of 9 m; small subgroups; abstract-level
Tilson [6]283; about 20–60 days after stroke0.16 m/smRS-anchored important changeSeverely impaired speed; sensitivity 73.9%, specificity 57.0%
Fulk [7]Outpatient subacute rehabilitation; mean 56 days at entry0.175 participant anchor; 0.190 therapist anchor, m/sImportant changeGlobal change ratings; not a chronic-stroke universal value
Hayashi [42]62 acute-stage ambulators; about 4 to 16 days after onset0.21 m/s therapist ROC estimate; AUC 0.76Important changePatient/FIM anchors did not support overall ROC cutoffs; 0.18–0.25 range uses a different method
Cheng [10]20 retested; median 134 days after stroke15-m corridor: 44.0 m; 30-m corridor: 67.5 mMDC95Different corridor protocols; small cohort
Fulk and He [11]2–6 months after stroke; initial speed <0.4 m/s subgroup44 m; mRS AUC 0.726MWT important changeWhole-sample 71 m has weaker discrimination; faster subgroup uncertain
Kubo [12]107; 30–60 days after strokePrior 71 m threshold failed; FAC new estimates 63.1/69.0 m6MWT MIC
validation
and update
Methods produce different results; full text still needed
Ng [17]63; mean 8.8 years after strokeAbout 5 FGA pointsMDCNot an MCID; stroke/healthy cutoff is not a falls cutoff
Wellmon [43]6 stroke videos rated by 14 therapists4.24 WGS pointsMDC95Repeat ratings of the same videos; not repeated patient capture

Sustained walking and endurance

The 6MWT is attractive because it samples sustained walking, but distance reflects several interacting constraints: pace, endurance, motor impairment, balance, turning, motivation, symptoms and rest. It is not a direct measure of maximal oxygen consumption. Corridor length matters because shorter courses impose more turns, and assistive-device handling may make those turns disproportionately costly after stroke.

In an early-rehabilitation cohort, Fulk and colleagues reported high 6MWT reliability and MDC90 54.1 m. In Cheng and colleagues’ stroke-specific protocol study, 20 participants returned for retesting 1–3 days later; median time after stroke was 134 days. MDC95 was 44.0 m on a 15-m walkway and 67.5 m on a 30-m walkway, with high reliability for both. These estimates arise from small samples and distinct procedures. A shorter corridor may improve feasibility, but it does not make two corridor protocols numerically interchangeable. [9, 10]

Fulk and He’s analysis of change from 2 to 6 months after stroke illustrates the limitations of a single 6MWT MCID. For the overall sample, estimates were 71 m using mRS and 65 m using the Stroke Impact Scale, with AUCs only 0.66 and 0.59. Among those initially below 0.4 m/s, the mRS-anchored estimate was 44 m with AUC 0.72. The authors could not estimate an accurate MCID in the faster subgroup. Their 71-m figure should not be detached from this uncertainty. [11]

Kubo and colleagues’ 2025 external evaluation is an important corrective. In 107 rehabilitation inpatients tested at 30 and 60 days, the previously reported 71-m threshold did not meet the study’s validation criterion for mRS change: LR+ 1.41 and LR− 0.77. New estimates differed substantially by anchor and analytical method. FAC-based estimates were 69.0 m by ROC and 63.1 m by adjusted predictive modelling. The abstract reports only that the FAC-linked MIC was supported; the complete body remains an important retrieval priority before implementation. A threshold’s failure to transport is useful evidence, not a reason to select the most convenient replacement. [12]

The 2MWT can reduce burden, but should not be treated as one-third of a 6MWT. Short and long tests may provoke different fatigue and pacing responses. A 2026 seven-centre study reports an anchor-based 2MWT MCID of 33 m in 150 subacute patients. Its full text contains contradictory narrative values for baseline distance and change, and a subgroup table mixes intervention-group mean improvements with MIC estimates. These issues limit confidence in the precision and transportability of that number. It belongs in an emerging-evidence note, not a default software decision rule. [13]

Independence and complex walking

The FAC directly describes how much human support walking requires. Mehrholz and colleagues followed 55 initially nonambulatory people 30–60 days after first stroke. Video-based FAC ratings were highly reproducible, and FAC at least 4 after four weeks of rehabilitation classified later community ambulation with reported sensitivity 100% and specificity 78%. This is a small selected cohort with a later outcome, not merely a cross-sectional correlation. The retrieved abstract does not establish external validation, calibration or a universal probability of outdoor independence. [40]

DGI and FGA add demands such as speed changes, head movements, obstacles, narrow-base or backward walking and stairs. They should be selected for the task construct and functional level, not because their scores are sometimes called falls tests. Jonsdottir and Cattaneo reported DGI total-score reliability of 0.96 in 25 ambulatory participants labelled chronic, although eligibility began at three months. Item-level reliability was lower. Thieme and colleagues found high FGA observer reliability in 28 subacute ambulators, using direct and video raters. These are measurement studies, not prospective fall-risk validations. [14, 15]

Lin and colleagues compared DGI, four-item DGI and FGA in separate longitudinal and reliability samples. FGA had the least floor and ceiling effects, making it useful in higher-functioning walkers. Baseline scores correlated with later Barthel Index, but this does not establish calibrated individual prognosis or incremental value over baseline disability. A newer 63-person chronic-stroke study reports FGA test–retest ICC(3,1) 0.85, inter-rater ICC(2,2) 0.94 (an average-rating model, not a one-rater coefficient) and MDC about five points. Its score of 20 distinguishing stroke from healthy controls is a current group classifier, not a prospective falls cutoff. [16, 17]

Instrumented gait and movement quality

What additional detail can and cannot tell us

Mean speed is an outcome of many possible movement strategies. Identical speeds can coexist with different paretic propulsion, foot clearance, swing coordination, joint excursions, step widths and support-time distributions. Instrumentation can reveal these differences and help clinicians formulate or test a mechanistic hypothesis. It does not automatically establish which abnormality causes future disability or which feature should be targeted to improve it.

Optical motion capture, pressure-sensitive walkways, inertial systems and cameras measure different physical signals. Their gait events can differ, especially with flat-foot contact, toe-first contact, shuffling, foot drag, circumduction or an orthosis. A walkway is not a criterion reference for all joint kinematics, and an optical model has its own marker-placement and modelling error. Validation must name the measured construct and the appropriate reference.

The 2026 wearable systematic review is helpful as a map: sixteen studies, about 300 stroke participants, predominantly chronic, with wide variation in devices, sensors, tasks and support. Its percentages describe parameter-level results, not the proportion of patients accurately measured. Within-visit, between-visit and condition-comparison designs are mixed. The clinically useful conclusion is that several pace and spatial outputs can perform well in selected protocols while phase timing, asymmetry and slow impaired gait remain more challenging. It is not evidence that all wearables are interchangeable or that a greater sensor count is universally superior. [44]

Wearables require parameter specific validation

Moore and colleagues tested a single lower-back accelerometer in laboratory and community settings. Of 25 participants, two using fixed plastic AFOs were excluded because the event-detection method failed; 23 were analysed. Agreement with GAITRite was moderate to good for selected temporal and pace outputs, but poor for several spatial, variability and asymmetry measures. Laboratory step-length asymmetry had poor repeatability. Good week-to-week free-living repeatability did not establish free-living accuracy because a simultaneous free-living criterion reference was absent. These findings make the device potentially useful for selected outputs while clearly defining where claims should stop. [45]

Buckley and colleagues analysed trunk-accelerometer asymmetry in the same-sized, apparently shared cohort and again excluded the two fixed-AFO cases. The report therefore treats this as methodological extension rather than independent replication. Asymmetry validity depends on the specific signal feature and formula. A general statement that a single trunk sensor measures hemiparetic asymmetry would exceed the evidence. [46]

Felius and colleagues studied a feasible two-minute walking protocol using foot and lower-back IMUs during stroke rehabilitation. Many spatiotemporal features were reproducible, but only a minority of frequency features performed well, and entropy and local-divergence measures were less reliable. Approximately 25 strides were used for some complexity calculations; increasing the requirement would exclude slow walkers. Their own analysis notes that sensitivity to meaningful longitudinal change and prospective prediction still needed study. Repeatability in a short test is a prerequisite, not proof of a useful progression biomarker. [47]

Lanotte and colleagues’ publicly available author-provided full article adds an important correction to its abstract shorthand. Sixteen chronic survivors and ten younger healthy controls wore foot and L5 Opal sensors during repeated 10mWTs. In the stroke group, Table 2 reports speed bias ± limits of agreement of −0.08 ±0.10 m/s and stride length −0.10 ±0.11 m; the abstract’s −0.11 m/s and −0.12 m values should not be treated as uniform stroke corrections. Proportional bias and poorer accuracy in slow, asymmetric gait remained important. Reliability involved three within-condition trials, not between-day testing. Hundreds of footfalls do not replace the sixteen independent stroke participants, and footfall-level analyses require attention to clustering. A product should not substitute wearable values into stopwatch-derived thresholds merely because ICCs are high. [48]

Igarashi and colleagues report same-day MDCs for trunk acceleration indices in nineteen subacute inpatients. The protocol used an L3 sensor, a 16-m path, two trials within thirty minutes and five central gait cycles. Stride-regularity MDCs were approximately 0.15–0.18; harmonic-ratio MDCs approximately 0.67–0.86, depending on axis. These values quantify this short standardized protocol. They are not an MCID, a between-day error estimate or a fall-risk threshold; they cannot be copied across filters, normalizations, axes or gait-cycle counts. [49]

The 2026 REEV SENSE validation further exposes the importance of exclusions. Twenty chronic survivors enrolled but six walking below 0.28 m/s were excluded, leaving fourteen, including four cane users. Before the final accuracy analysis, intersystem measurement differences beyond ±1.96 SD were also removed. In the retained observations after participant and measurement-error exclusions, temporal agreement was generally strong; spatial performance deteriorated in slow walking. Removing observations based on the size of their measurement error can improve the retained error distribution and limits of agreement; the quantitative effect of this trimming was not reported. Because cane use and speed were closely entangled in these small groups, the study cannot isolate a causal effect of the cane on error. The excluded slowest walkers are precisely a population in which clinical measurement assistance may be most valuable. Report their failure rate rather than advertising accuracy from completers alone. [50]

Video and markerless systems

Video may reduce equipment burden and preserve a record for review, but the evidence is pipeline-specific. A multicamera commercial system, a depth camera, a moving smartphone and a single fixed RGB camera are different measurement systems. Camera geometry, clothing, assistive-device occlusion, therapist position, image resolution and training data can alter performance. Their algorithms may estimate apparently plausible movement even when critical landmarks are obscured.

Alammari and colleagues compared an eight-camera KinaTrax system with an instrumented walkway in nineteen home- or community-dwelling stroke survivors. Three comfortable and three fast trials were planned, and one matched complete gait cycle per trial was analysed. Speed and stride length agreed very well; comfortable-speed limits of agreement for speed were about ±0.03 m/s. In contrast, paretic single-limb support agreement was poor, with ICC about 0.31 at comfortable speed, and stride-width agreement was poor. A general-purpose validation label would obscure the outputs that could most affect a hemiparetic stability interpretation. The study did not establish between-day reliability, responsiveness, home use or prognosis. [51]

A 2024 SMARTGAIT study directly examined a single moving smartphone against Vicon in eight people, including four using walking aids and foot-lift orthoses. It included subacute and chronic cases and FAC levels 2–5. Speed agreement was excellent, step-length agreement lower, and average joint-angle RMSEs approximately 3.5–4.6 degrees in the sagittal plane. Aid users showed greater error for several outputs, and individual gait phases could diverge despite acceptable averages. This is valuable feasibility and concurrent-validation evidence in actual stroke, rather than simulated pathological gait; eight participants still provide a narrow basis for generalization. It does not establish a clinically important change threshold or prospective risk model. [52]

Lonini and colleagues’ earlier eight-person proof-of-concept study is a further original anchor for single-camera gait estimation. Its custom pose-estimation work should not be represented as validation of an unrelated off-the-shelf application. The original body was incompletely retrieved here, so detailed performance interpretation is limited and the paper is retained as a retrieval priority rather than a foundation for a broad implementation claim. [53]

Table 3 Technology evidence is specific to the output

Device names identify tested systems, not recommendations. One validated parameter does not validate every output or the complete product.

Table 3 Technology evidence is specific to the output
Original studyScopeUseful findingMain boundary
Moore [45]Single L5 accelerometer; 23 analysedSelected temporal and pace measures agree better than asymmetryTwo fixed-AFO cases failed; repeatability is not free-living criterion validity
Felius [47]Foot/back IMUs; rehabilitation cohortMany spatiotemporal outputs reproducibleShort records weaken complexity estimates; no proven meaningful change
Lanotte [48]Commercial Opal systemStrong within-condition repeatability for several outputsStroke-specific proportional/spatial bias; greater errors in slower or asymmetric gait; no universal correction
Igarashi [49]L3 IMU; 19 subacute inpatientsSame-day error estimates for trunk indicesFive cycles; axis/filter/normalization-specific; no fall threshold
Alammari [51]Eight-camera KinaTrax; 19 survivorsExcellent speed/stride-length agreementPoor paretic single-support and stride-width agreement
SMARTGAIT [52]Single moving smartphone; 8 survivorsConcurrent kinematic and spatiotemporal feasibilitySmall sample; aid/orthosis-related errors; no between-day prognosis evidence
Marsan [50]Foot IMUs; 14/20 analysedStrong temporal agreement after participant and measurement-error exclusionsSix slowest excluded; intersystem differences beyond ±1.96 SD also removed before the final accuracy analysis; all walker users excluded; cane and speed confounded

Asymmetry is informative but not automatically a treatment target

Step-length, step-time, swing-time and stance-time asymmetry describe different features. Ratios, normalized differences and absolute differences have different scales and properties. A ratio also becomes unstable when the denominator is small. Reporting a magnitude without direction discards which side is longer or slower, and a symmetric pattern may still be severely impaired on both sides. The affected and less-affected limbs must remain identifiable. [54]

Speed alters many spatial and temporal features. A decrease in asymmetry after walking faster may reflect a speed change, a changed strategy or recovery; those interpretations require additional evidence. Useful comparisons include a usual-speed condition and, when safe and clinically justified, a speed-matched condition. Treadmill observations should not be assumed to represent overground community gait. Kesar and colleagues’ treadmill measurement-error work is relevant to repeatability under that controlled condition, but its estimates are not transferable to unassisted overground smartphone recordings. [55]

Longitudinal measurements during inpatient rehabilitation also demonstrate that speed and symmetry need not evolve together. A faster or more symmetric gait is not automatically a safer or less effortful gait, and a compensatory strategy may support current function. Observational scales and kinematic measures can help describe those strategies, but improving a surrogate feature does not by itself prove improved falls, community participation or long-term independence. [54, 56]

Observational gait scales

Wisconsin Gait Scale and Gait Assessment and Intervention Tool approaches make observation more structured than an unrecorded impression. They can identify features not represented by speed alone, and video can support repeatable review and education. Nevertheless, rater training, camera views and the movement sample determine what can be scored.

Wellmon and colleagues’ WGS video study involved fourteen therapists but only six stroke participants. Inter-rater ICC was 0.83 and intra-rater ICC 0.91; MDC95 was 4.24 points. The effective diversity of gait was six people, not the number of ratings or videos. Re-scoring the same recording estimates rating error and does not include day-to-day biological variation or new camera capture error. A later WGS clinically-important-difference study addresses another question and should not be used to transform a repeated-video MDC into evidence of patient benefit. [43, 57]

Recovery of independent walking

Prognosis begins before gait can necessarily be measured

For a person unable to walk, early prognosis may depend more on sitting control, lower-limb strength, balance and age than on a gait waveform that cannot yet be collected. This is not a limitation of clinical relevance. A gait-related outcome can be predicted by a non-gait baseline assessment. The report therefore includes EPOS and TWIST while keeping them distinct from instrumented gait prediction.

The prediction target must specify independence, surface and time. FAC at least 4 means independent walking on level ground. It does not necessarily mean independent stairs, uneven terrain, unlimited outdoor walking or full participation. Time measured from stroke onset is different from time after rehabilitation discharge. The event of first achieving independence also differs from still being independent at a later visit. Death, recurrent stroke, discharge transfer and withdrawal are relevant competing or missing outcomes, not interchangeable exclusions.

Systematic reviews of independent-walking prognosis identify many candidate factors and models. The 2024 model review is useful for highlighting inadequate external validation in much of the earlier literature, but a review predating the 2026 EPOS and TWIST evaluations cannot establish the present absence of external evidence. The appropriate synthesis is that a few clinical tools now have external data, with important transportability and calibration qualifications, while many technological models remain at development stage. [58, 59]

EPOS

The original EPOS gait study followed 154 first-ever ischemic-stroke patients unable to walk independently. It assessed clinical variables within 72 hours and again on days 5 and 9, targeting FAC at least 4 at six months. Sitting balance and paretic leg strength were the principal predictors. The frequently quoted 98% favourable probability for preserved early sitting and leg function and 27% probability when absent belong to that development population and timing; the latter declined with later assessment. Those values should not be used as fixed prognosis across all stroke. [18]

Veerbeek and colleagues’ 2022 validation applied the model in two independent Swiss cohorts, with 39 and 78 patients, but used a three-month endpoint. Performance depended strongly on assessment timing: AUC was 0.675 on day 1 in the smaller cohort and 0.921 on day 8; in the second cohort it was 0.801 on day 3 and 0.846 on day 9. Early negative predictions were particularly problematic. The cohorts mainly represented first strokes, mild-to-moderate impairment and limited prestroke disability. This is meaningful external evaluation, with a changed outcome horizon, not proof of validity for every patient admitted in the first day. [19]

The complete 2026 Vinzens paper broadens the population while retaining an important prestroke boundary. The two Swiss centres enrolled 293 patients unable to walk independently within 72 hours; thirteen dropped out, including nine deaths, leaving 280 analysed. First and recurrent, ischemic and hemorrhagic, supratentorial and infratentorial strokes were eligible, but patients dependent in walking before admission were excluded. Thus, broader prestroke disability does not mean that premorbid walking dependence was validated. All predictors were complete. The original day-three EPOS equation was applied to a three-month FAC endpoint; 217 regained independence and 63 remained dependent. [20]

At three months, AUC was 0.74 (95% CI 0.67–0.81) and Brier score 0.14. At probability 0.5, sensitivity was 0.98 and specificity 0.32: 212 true favourable predictions, five false unfavourable predictions, twenty true unfavourable predictions and forty-three false favourable predictions. Calibration-in-the-large was −0.53 on the logit scale and slope 0.64 (95% CI 0.42–0.85). The highest original probability group, 98%, was overoptimistic. The 85% probability group contained only four people, making its apparent agreement extremely uncertain. [20]

Outcome ascertainment mixed direct assessments, structured telephone interviews and rehabilitation reports. Blinding was incomplete, approximately half overall, and calibration differed between centres and assessment methods. The study stopped before its planned sample of 476 because of investigator employment changes; sixty-three non-events limited precision. Dropouts were older and more severely affected than analysed participants. These details strengthen the caution attached to apparently high sensitivity and support explicit outcome and missingness reporting. [20]

The authors recalibrated the intercept and slope in the validation cohort, retaining discrimination and the same classification at probability 0.5. The new probabilities require another independent test. Decision curves favoured the original model over default strategies over thresholds 0.28–0.90, but this is decision-analytic evidence, not proof that using it improved rehabilitation or patient outcomes. The source’s emphasis on avoiding false pessimism is reasonable; it does not authorize reducing rehabilitation for a person in a low predicted category. The Lucerne site, recruitment period and trial registration overlap the early-TWIST evaluation below, so these publications should not be treated as wholly independent patient evidence. [20, 24]

The practical lesson is to treat pessimistic early estimates cautiously and reassess. A score should structure the discussion of uncertainty and planning, not become a reason to withhold rehabilitation. External performance in a broader cohort may be lower without making the original finding false: it changes the target population and reveals the limits of transporting the original probabilities.

Distinguish the TWIST tool from the earlier TWIST algorithm

The 2022 TWIST prediction tool combines age, knee-extension strength and Berg Balance Scale score at approximately one week after stroke. Its 93-person development cohort was followed for independent walking at 4, 6, 9, 16 and 26 weeks. The final score ranges from 0 to 4. Its purpose is to estimate both whether and by when independence may occur. These inputs and horizons must be kept separate from the earlier TCT and hip-extensor TWIST algorithm. [21, 24]

Smith and colleagues’ 2026 temporal validation recruited a new cohort at a development-site hospital in Auckland. Ninety-eight were recruited and 89 retained in the main analysis; six people died before achieving independent walking, two withdrew after medical events, and one became independent between enrolment and the one-week assessment. Fifty-seven achieved independence by nine weeks and another fifteen between nine and twenty-six weeks. C-statistics ranged from 0.81 to 0.91 and Brier scores from 0.11 to 0.15. The original probabilities were too optimistic for some score–time combinations, particularly score 3. The authors revised several probabilities by combining development and validation data. [22]

This study is stronger evidence than apparent accuracy alone because it examines discrimination and calibration in new patients. It also demonstrates that a model can rank people well while misestimating their probabilities. The revised tool has been updated using these validation outcomes; its new probabilities therefore require further independent testing. The small low-score groups are especially uncertain. Outcome testing permitted sticks, quad sticks and AFOs, but not walking frames. Prestroke frame users, people with prestroke FAC below 4, and bilateral or cerebellar strokes were excluded, narrowing the meaning and applicability of independence. Excluding deaths means the model’s displayed recovery probabilities should not be casually read as unconditional probabilities for every newly assessed patient, including those who may die before walking independently.

The independent 2026 Japanese multicentre study is crucial to transportability and was subsequently recovered in complete publisher HTML. Three hospitals recruited 179 people with first-onset stroke; 34 were lost to follow-up or died before independent walking, leaving 145 in the principal analysis. Participants were independent before stroke, could understand instructions, had lower-limb weakness and were not independent by eight days. Predictors were measured at 5–8 days. Blinded experienced assessors used FAC in person during admission and standardized telephone interviews after discharge, corroborated where possible by therapists, family or residential-care staff. Eighty-eight achieved independence by nine weeks and seventeen more by twenty-six weeks; forty remained dependent. [23]

The full table is more informative than the abstract: C-statistics were 0.76, 0.82, 0.87, 0.87 and 0.89 at 4, 6, 9, 16 and 26 weeks, respectively. Corresponding Brier scores were 0.18, 0.17, 0.14, 0.13 and 0.13. Reported calibration-in-the-large was negative at all horizons, from −0.34 to −1.62; the four-week confidence interval included zero, whereas subsequent intervals did not. Calibration slopes were 0.58, 0.56, 0.67, 0.33 and 0.52, all substantially below the ideal of one. Thus, good discrimination coexisted with poor calibration, rather than proving that the original absolute probabilities transported reliably. [23]

Missing knee-extension values in eleven participants and BBS scores in nine were imputed using missForest. The primary analysis used predictor-only imputation; complete-case and predictor–outcome imputation sensitivity analyses retained the main overprediction concern. However, the exclusion of losses and deaths still limits an unconditional prognosis for everyone initially enrolled. The median total hospital stay was 79 days, versus 33 in the development cohort, and neurological severity, balance and case mix differed. Those observations support attention to setting and outcome ascertainment; they do not identify a single cause of miscalibration. [23]

Some predicted probabilities were manually adjusted using the Japanese validation data. These updated probabilities require another independent evaluation. The article also contains inconsistent narrative statements about some score-specific calibration and the independence-versus-dependence percentage comparison; the interpretation here rests on the explicit outcome counts and Table 2. The reported calibration-in-the-large definition should be clarified before computational reuse, rather than reverse-engineering a deployable correction from the headline values. [23]

The complete Rajkovic August 2026 paper evaluates a different TWIST algorithm, originally based on TCT total score and paretic hip-extension strength at one week. The Swiss validation moved predictor assessment to within 72 hours. Of 149 enrolled, three died and were excluded, leaving 146; the planned 270 were not reached. First and recurrent strokes at any localization were eligible, including cerebellar stroke, but prestroke walking dependence was excluded. Outcomes were FAC at six and twelve weeks, and aids or orthoses could be used, including rolling walkers. This differs from the no-frame outcome rule in the New Zealand validation of the newer tool. [24]

Ninety-nine people were independent by six weeks, another twenty-six became independent between six and twelve weeks, and twenty-one remained dependent. Multiclass AUC was 0.73 and overall accuracy 0.69 (95% CI 0.61–0.77). The early-independent category had sensitivity 0.83, but the intermediate category sensitivity was 0.23. More clinically revealing, only thirteen of thirty-eight people predicted to remain dependent were actually dependent at twelve weeks; twenty-five recovered independence earlier. The positive predictive value of that unfavourable category was 0.34. High specificity for a dependence category should therefore not be confused with a reliable pessimistic prediction for an individual. [24]

Only a quarter of outcome assessments were blinded. Ascertainment mixed direct assessment and structured interviews, with clear differences in severity and functional status between follow-up modes. Excluding cerebellar strokes left similar performance, but the intermediate group remained small. Therapy intensity was not directly measured. The study establishes informative early classification for some patients and substantial error about timing for others; it does not validate calibrated individual probabilities or justify withholding rehabilitation. It must remain separate from the age, knee-extension and Berg 0–4 TWIST tool. Its Lucerne recruitment in 2021–2023 and trial NCT05039047 overlap the EPOS validation programme; exact cross-publication participant overlap should be settled before pooling evidence. [20, 24]

Table 4 Walking prognosis models should not be combined

FAC independence is level-ground human-support independence. New Zealand recalibration used combined cohorts and does not itself provide new external validation of the revised probabilities.

Table 4 Walking prognosis models should not be combined
Model and evaluationPredictor timing and inputsLater targetKey external finding
EPOS development [18]Within 72 h; sitting and paretic leg strengthFAC ≥4 at 6 monthsDevelopment probabilities require population-specific interpretation
EPOS Swiss 2022 [19]Days 1/8 or 3/9; original clinical predictorsFAC ≥4 at 3 monthsAUC 0.675 at day 1 versus 0.801–0.921 at later assessments
EPOS heterogeneous 2026 [20]Within 72 h; 280 patientsIndependence at 3 monthsAUC 0.74; calibration slope 0.64; sensitivity 0.98, specificity 0.32 at 0.5
TWIST tool NZ 2026 [22]About one week; age, knee extension and Berg; 89 analysedIndependence by 4, 6, 9, 16 and 26 weeksC-statistic 0.81–0.91; some original probabilities optimistic
TWIST tool Japan 2026 [23]1 week; same 0–4 tool; 145 patientsSame five horizonsC-statistics 0.76–0.89; poor calibration; four-week overprediction estimate uncertain
Earlier TWIST algorithm 2026 [24]Within 72 h; TCT and hip-extension; 146 patientsIndependent 6 weeks /12 weeks /dependent 12 weeksMulticlass AUC 0.73; intermediate-category sensitivity 0.23; dependence-category PPV 0.34

Sensor based prediction of rehabilitation outcomes

O’Brien and colleagues prospectively collected wearable data near admission and discharge from inpatient rehabilitation. Of 55 recruited patients, modelling of ambulatory patients used 32 with relevant gait and balance recordings; nonambulatory evaluation involved only eight patients, with models trained on a combined sample of fifty. Nested leave-one-subject-out validation is a meaningful safeguard against straightforward subject leakage, but remains internal validation in a small single-service sample. [60]

The outcomes were discharge categories based on 10mWT, motor FIM and BBS. The BBS-defined outcome was called risk of falling, but actual subsequent falls were not measured. Consequently, this is prediction of a later clinical scale category, not a future-fall validation. Sensors did not improve every endpoint: the independence model remained similar to the clinical benchmark. Fifty patients were also used in a previous sensor-free regression study, so the related publications are not independent replications. Coding uncompleted assessments as zero was an analytical choice in that study, not a general clinical recommendation.

The complete 2026 Kim paper directly tests incremental instrumented-gait information in a retrospective single-centre cohort. Clinical assessments were obtained within seven days of admission and GAITRite used at least three valid comfortable-speed passes. Patients had admission FAC 2–3 and could walk 10 m with aids or physical assistance. The actual onset range was 4–14 weeks, crossing the early/late-subacute boundary used in this report. Of 138 otherwise eligible records, one implausible step-time record was excluded; 137 complete cases remained. At discharge, eighty-nine met FAC at least 4 with documented independent outdoor mobility and forty-eight remained indoor-only. Documentation of outdoor performance is important because FAC at least 4 alone is not synonymous with outdoor independence. No fixed discharge interval or time-to-event analysis was established in the examined main text. [37]

Univariate screening and AIC-based forward selection produced a clinical model containing Motricity Index and time since onset. Adding affected-side single- and double-support percentages changed AUC from 0.995 to 0.998. The nominal block likelihood-ratio test was significant (p=0.009), while the AUC comparison was not (DeLong p=0.262). Because candidates and models were selected in these data, this is an exploratory increment rather than a confirmatory test. Neither individual added coefficient was conventionally significant in the unpenalized model, and double support reversed direction from its univariate association. Firth analyses addressed separation, but the defensible interpretation remains a block-level fit increment, not an independent causal effect of either support parameter or a reason to train more double support. [37]

Bootstrap out-of-bag analysis reported corrected AUCs of 0.994 and 0.996. However, the main text does not establish whether the full screening and forward-selection process was repeated within each bootstrap sample; code and supplements would be needed to settle that point. Quoting events per final variable does not account for the larger candidate search. Brier scores were 0.029 and 0.016, while calibration intercepts of zero and slopes of one were reported in the analysis cohort. Those apparent fitted-data calibration results and very high Hosmer–Lemeshow p values do not establish calibration in new patients. Firth estimation also does not remove every form of selection optimism. [37]

Excluding admission FAC avoided including that walking classification in the chosen benchmark, but baseline FAC is a legitimate temporally available prognostic variable, not automatically outcome leakage. Consequently, the study demonstrates increment over its selected Motricity Index/time model, not necessarily over the best routine clinical assessment. A speed-adjusted sensitivity analysis was reported but itself encountered separation. The author appropriately labels derived cutoffs exploratory and calls for external validation. Equipment-specific implementation, individual probabilities and clinically important added benefit remain unproven. [37]

Future falls

Falls require an observed outcome

A fall-risk questionnaire, BBS threshold, historical fall label, laboratory slip response and prospectively recorded fall are different outcomes. Only the last directly establishes later events. First falls, any fall, recurrent falls, injurious falls and fall counts also require different analyses. A six-month prospective cohort cannot automatically justify a one-year injury-risk claim.

Exposure complicates the relationship between impairment and falls. A person with very limited mobility may encounter fewer walking hazards, while increasing activity creates more opportunities to fall. Assistance and aids are markers of impairment and also influence exposure and protection. An association between cane use and falls does not mean a cane causes falling or should be removed. Appropriate clinical interpretation includes prior falls, cognition, vision, medication and cardiovascular contributors, environment, activity exposure and support, alongside gait.

Subacute gait and dynamic balance

Bower and colleagues followed people after inpatient rehabilitation, with 81 of 96 recruited patients completing twelve-month falls follow-up. Twenty-three fell at least once and thirteen more than once. Baseline assessment occurred at a median 24 days after stroke. Ordinal regression models adjusted each candidate predictor for country, prior falls and gait assistance, with a further model including comfortable six-metre gait speed. This is a six-metre walk, not the six-minute endurance test. [25]

Smaller mediolateral pelvic displacement during a Kinect-recorded fast walk remained associated with the ordinal prospective-falls outcome after speed adjustment: IQR-scaled OR 6.75, 95% CI 2.20–20.74. Associations for stride length and step-length asymmetry did not retain significance after speed adjustment. TUG and step-test performance also remained informative. The study therefore gives more specific evidence than an unadjusted correlation and shows the importance of comparing against speed. However, 23 fallers supported multiple separate exploratory models, the capture field was short, missing gait observations occurred, and the paper itself questioned measurement accuracy for pelvic displacement. It did not externally validate a calibrated personal-risk equation.

Plummer and colleagues’ 45-person derivation cohort provides a clinically practical counterpoint. Twenty-one participants fell within three months after discharge. Candidate indices combined obstacle crossing, paretic-limb step testing, 5-m walk speed, perceived walking difficulty and sometimes device use. Apparent AUCs of approximately 0.82–0.85 are promising, but index composition, score cutoffs and competing versions were selected in the same small cohort. Counting 45 participants divided by five candidate variables is not equivalent to counting 21 outcome events per candidate variable. Reducing selected components to a composite score does not remove the optimism introduced while creating it. Index D3 was also evaluated in a separate 30-person external cohort with nine fallers: AUC 0.84 (95% CI 0.68–1.00), sensitivity 0.56 and specificity 0.95 at the prespecified cutoff of 3.5. Index C was assessed in 54 pooled participants, comprising the original 45 plus nine new patients, so that analysis was not independent external validation. A lower D3 cutoff of 2.5 was selected in the validation sample and needs fresh evaluation. These small samples still limit implementation. [28]

Chronic gait and daily life

Punt and colleagues compared clinical assessments, treadmill motion-capture features and seven-day accelerometry in forty chronic community-dwelling survivors, with fifteen fallers over six months. Falls were collected by calendar and monthly calls. Internal cross-validated AUCs were approximately 0.73 for laboratory gait and 0.72 for daily-life gait, versus 0.64 for the clinical set; combining gait sources did not materially improve performance. Five participants who needed the treadmill handrail and other incomplete cases were excluded, selecting more capable walkers. Feature screening and small-sample component modelling leave uncertainty even though PCA and regression were cross-validated. No external calibration or clinical-impact evaluation was established. [26]

This is evidence that gait quality can contain a prospective signal. It is not proof that any wearable outperforms a comprehensive stroke-specific clinical assessment. Related Punt publications on daily gait and perturbations may share recruitment or participants; their independence must be checked before counting them as replications. The present report relies on the prospective comparison as one cohort rather than inflating its evidence through multiple publications.

Tsang and colleagues followed 93 people with chronic stroke for twelve months; thirty-six fell. A dual task combining an auditory clock test with obstacle crossing classified about 80% of fall status, with 72% sensitivity and 84% specificity in the reported model. Other tested measures did not differ between fallers and nonfallers. Importantly, the selected signal was reaction time in a particular cognitive–motor task, not a general dual-task gait-speed cost. The full body was not obtained, so model complexity, calibration and validation cannot be confirmed. This supports targeted further testing, not copying the result to serial subtraction during ordinary walking. [27]

Mobile video needs a particularly cautious interpretation

Lam and colleagues’ 2026 study is a genuine prospective attempt to use mobile markerless gait measures for later falls. The accessible final accepted manuscript reports three iPad Pro recordings of three 3-m usual-speed walks, with follow-up interviews over eighteen months. Fifty stroke participants entered and forty-six were analysed; mean time since stroke was about six years, and participants needed to walk at least 3 m with or without aids. Recent fallers within three months were excluded. The publisher abstract reports 13% fallers, whereas the accessible manuscript’s Table 1 lists six single fallers and two recurrent fallers: eight of forty-six, or 17.4%, if those categories are disjoint as written. [29]

The same manuscript reports a combined AUC of 0.81 but uses numerous gait parameters and covariates, stepwise selection and cohort-derived cutoffs with very few events. Some tabulated confidence intervals exclude the null despite nonsignificant reported p values, requiring clarification. The cited prior device validation concerns upper-limb kinematics rather than these gait outputs. These are source-reporting and design limitations, not corrections that can be resolved by choosing favourable estimates. Until clarified and independently replicated, the defensible description is exploratory prospective association. It is unsuitable as a ready-made rehabtools falls calculator.

What a defensible falls model would need

A clinically usable system needs a clearly defined endpoint and horizon; prospective event collection with an operational fall definition; an appropriate baseline comparator including fall history and relevant stroke characteristics; enough events relative to all model-building choices; transparent handling of missingness, death and exposure; participant-level internal validation of every modelling step; external discrimination and calibration; and an evaluation of whether acting on the output improves care. Sensitivity and specificity alone do not show whether a predicted 30% risk is accurate, or whether the proposed threshold produces net benefit in the intended clinic.

For now, gait technology can help characterize suspected mechanisms, document change and support a wider clinical risk assessment. Its output should not deliver a reassuring low-risk label simply because a short straight walk appears normal. A patient may fail during turning, distraction, obstacle negotiation, fatigue, transfers or environmental interaction that the recording never sampled.

Community walking and participation

Historical speed categories are not prospective recovery rules

Perry and colleagues’ influential walking-handicap classification connected speed with expert-assigned functional categories. Subsequent clinical practice widely adopted 0.4 and 0.8 m/s boundaries. These are useful descriptions of walking capacity, but crossing a boundary does not by itself demonstrate a change in actual community participation. Definitions, walking aids and real environmental demands matter. [30]

Fulk and colleagues’ 441-person analysis used contemporaneous daily step data from the LEAPS and FASTEST trials to re-examine categories. Comfortable-speed thresholds of 0.49 and 0.93 m/s and endurance-based classification were better linked to that activity definition than traditional categories. The 6MWT was the strongest single discriminator, with AUC 0.82 for home versus community and 0.76 for limited versus full community categories. Despite the word “predicting” in the title, this was cross-sectional classification. Daily step volume also does not identify the location of walking. These trial-derived analyses may overlap other LEAPS measurement reports. [31]

Bansal and colleagues addressed location directly in 60 chronic survivors and 18 controls, using seven-day accelerometry plus GPS. Higher clinical capacity separated the lowest groups from others, but medium- and high-capacity groups did not consistently differ in home or community steps. This is contemporary evidence against assuming a one-to-one translation from better speed or endurance category to real-world performance. It does not show that improving capacity is unimportant; it shows that the desired community outcome should be measured separately. [32]

Prospective studies give a more relevant but still qualified forecast

Igarashi and colleagues followed 78 of 92 enrolled ambulatory patients for six months after discharge from a Japanese hospital. Baseline assessment averaged about 22 days after onset. Walking speed used a timed 10 m with 3-m auxiliary zones; 6MWT used a 30-m corridor. Later community walking was a telephone-assessed modified Functional Walking Category. For distinguishing least-limited from unlimited community walkers, proposed thresholds were 299 m for the 6MWT and 0.94 m/s for comfortable speed, with AUCs 0.896 and 0.844. Performance distinguishing the lower two categories was weaker. The thresholds were selected and assessed in the same cohort, without multivariable adjustment, external validation or calibration. [33]

The source also contains wording problems in its definition of Youden’s index and in interpreting AUC. These do not erase the reported prospective relationship, but argue against treating the derived numbers as a fully specified clinical decision rule. The outcome is a self-reported functional category, not GPS-confirmed daily walking or social participation.

Mulder and colleagues’ larger prospective cohort included 243 patients after inpatient rehabilitation. A discharge comfortable speed of at least 0.5 m/s correctly identified 181 of 193 people in the favourable branch as later independent community walkers, whereas only 27 of 50 slower patients were correctly classified as noncommunity walkers. Thus, the model was much less informative for unfavourable prognosis. A speed below the threshold should not be communicated as an inability to recover community walking. The original full methods remain a retrieval gap, including the treatment of validation and pruning in the CART model. [34]

Rosa and colleagues’ 35-person study is a useful example of different timing producing different thresholds. A baseline fast speed of at least 0.42 m/s was prognostic of independent community walking at six months, whereas a contemporaneous six-month fast speed above 0.84 m/s discriminated current status. The same paper therefore contains a future predictor and a current classifier. The outcome depended on self-report and only nine participants achieved independence, limiting precision. [35]

Felius and colleagues’ 2025 study examined later free-living outcomes rather than category labels. Thirty-five participants had longitudinal follow-up, but gait models often relied on only about twenty participants with complete measurements. Baseline and discharge clinical tests, wearable gait, and daily-life data were related to six-month daily strides and speed. Selected multivariable models explained moderate apparent variance, but no predictor worked consistently across every outcome and timepoint. Conventional clinical gait speed was not significantly related to later daily stride count, while prior activity and selected learned features contributed. These are small-sample prognostic associations without external validation; the variational-autoencoder features are not interchangeable with simple asymmetry or video measures. The paper also reports a discharge-time range of 8–379 days while defining follow-up at six months after stroke. Individual temporal ordering therefore needs clarification before every discharge-based model can be treated as a strictly prospective forecast. This ambiguity does not apply in the same way to the admission models and is not proof that leakage occurred. [36]

Participation deserves its own outcome

Mahendran and colleagues’ accelerometer–GPS cohort of 34 subacute survivors found little change in the total volume or intensity of community ambulation across the first six months after discharge, despite some later change in long walking bouts. This is longitudinal description of performance, not a validated prognostic score. Karageorge and colleagues followed 83 community-dwelling survivors with participation goals and found that baseline outings, walking capacity and age predicted outings six months later. An outing diary measures something different from speed or steps. [38, 39]

These studies support a practical sequence: measure walking capacity, ask what community activity matters to the individual, and directly measure that activity when it is the rehabilitation goal. A faster 10mWT should not be relabelled as improved participation, cognition, survival or quality of life without outcome-specific evidence. Within the sources examined, a transportable gait-derived forecast for later dementia, institutionalization or mortality after established stroke was not verified. This is a bounded evidence conclusion, not proof that no association exists anywhere in the literature.

Implications for rehabtools and clinical implementation

A defensible initial measurement offer

The product should describe exactly what it measures and in whom it has been evaluated. For example, measuring comfortable overground speed in independently ambulatory chronic-stroke users under a standardized camera protocol is a narrower and more supportable claim than assessing overall gait recovery. Each output needs its own error evidence. If speed performs well but paretic support time does not, the latter should be withheld, flagged as experimental, or displayed with a clear quality limitation rather than inheriting speed’s validation.

A useful report should combine the measured value, units, affected side, task context, aid or orthosis, assistance, trial count, confidence or quality indicator, and reason for any failed measurement. Repeat assessments should explicitly flag changed conditions. Store the underlying recording when authorized and technically appropriate for clinical audit, but do not assume that a plausible skeleton reconstruction is an accurate gait event.

Table 5 Claims suitable for a rehabilitation measurement product

These are evidence requirements and claim boundaries, not regulatory advice or an approval assessment.

Table 5 Claims suitable for a rehabilitation measurement product
Proposed claimMinimum supporting evidencePresent interpretation
Measures comfortable speedStroke-specific criterion agreement and repeatability of deployed pipelinePlausible and testable; preserve protocol and failure reporting
Measures paretic support timeParameter-specific event-detection agreement in impaired gaitCannot inherit validation from speed; some camera results are poor
Detects real individual changeAppropriate between-visit absolute error in target usersDo not transfer same-video rater MDC or same-day error blindly
Detects important recoveryAnchor-based meaningful change with credible relevant anchorContextual reference, not a universal responder switch
Predicts future fallsProspective observed falls; calibration; external validation and clinical comparatorExamined gait models remain exploratory or population-specific
Predicts independent walkingCorrect fixed model, timing, endpoint and external evaluationClinical tools have useful evidence but transportable probabilities remain qualified
Shows improved participationDirect participation or actual community-performance outcomeSpeed and step counts alone are insufficient

Clinical change should be presented as an interpretation

Software can calculate change correctly while interpreting it incorrectly. Its interface should distinguish the observed difference from an estimate of measurement error and from an anchor-based important-change reference. Reference values should identify stroke phase, baseline severity, test protocol and confidence level. A reference drawn from a treadmill, an instrumented walkway or repeated scoring of one video should not be silently transferred to fresh smartphone recordings.

For an individual near a threshold, uncertainty should be visible. A small change may warrant repeat measurement under stable conditions rather than an immediate responder label. Improvement after an aid change may be important functional progress even when it is not evidence of restoration of normal movement. Conversely, increasing speed without improved endurance, safety or participation may leave the patient’s main goal unmet.

Prognostic features require another validation programme

Measurement validity is only the first stage. A fixed prognostic model should then be evaluated with the actual deployed measurement pipeline. Testing should preserve its original inputs, definition and horizon or explicitly develop a new model. A camera-derived variable should not be inserted into an EPOS, TWIST or published gait equation simply because its name resembles the original variable. For some models the relevant input is a clinical score or strength grade that cannot be inferred reliably from a walking video at all.

The benchmark should be clinically credible and available at the intended decision time. Useful comparisons include age, exact time after stroke, prestroke mobility, stroke severity, current FAC, standard speed or endurance, prior falls and relevant cognition or balance, selected according to the outcome. Added value should be assessed beyond that benchmark using calibration, discrimination, prediction error and decision consequences. A statistically significant likelihood-ratio test or a tiny AUC increase does not by itself establish an improvement worth the cost and complexity.

Training, feature selection, imputation and hyperparameter choices must be repeated inside participant-level resampling. Strides, frames or repeated visits from the same person must not be split across training and testing in a way that leaks identity or future information. A large number of windows does not compensate for few patients or fall events. Geographic and temporal validation should include the slower walkers, aids, orthoses, cognitive–communication limitations and clinical environments expected in use.

Rehabilitation remains more than prediction

A model can support planning without determining access to treatment. Especially early after stroke, initial weakness or inability to complete a task may change rapidly. Reassessment should update the clinical picture, but an updated prediction must have an appropriate landmark or repeated-measures design; simply inserting later scores into an early model may be invalid.

Treatment studies establish whether an intervention changes an outcome under trial conditions. They do not automatically validate the measuring device, and prediction of spontaneous or usual-care recovery does not identify who benefits most from a treatment. Treatment-effect heterogeneity requires a different design and analysis. No examined gait model justifies denying therapy, removing a necessary aid, or guaranteeing a particular recovery outcome for an individual.

Priorities for the next evidence step

First, settle the intended clinical claim and the population. If the goal is dependable measurement, a prospective method-comparison and repeatability study should stratify slow and faster walkers, assistance, aids and orthoses, affected side and stroke phase. Report absolute error, bias, proportional bias, limits of agreement, failures and missingness for every output, not only correlations.

Second, evaluate clinically relevant change using the actual protocol across visits. Include external anchors that reflect the target outcome and the person’s goals. Do not derive a patient-important threshold solely from a distribution-based statistic or import a generic older-adult or Parkinson-disease value.

Third, if prognosis is intended, choose one outcome and horizon at a time. Independent level walking at nine weeks, any fall over six months and community outings after discharge need separate endpoints and different cohorts. Compare against a useful clinical baseline, preserve participant-level independence and plan external calibration. A limited feasibility study may justify continued research while remaining insufficient for an individual risk message.

Fourth, retrieve the small set of originals whose remaining access gaps materially constrain implementation: the 2025 6MWT MIC external evaluation and selected influential original measurement studies. Their access status is recorded in the linked bibliography and search appendix. Abstract-level findings already constrain claims, but exact equations, missing-data handling, outcome counts, calibration plots, device versions and cohort overlap should be verified before reuse.

Overall conclusion

Stroke gait assessment has a strong clinical role when independence, speed, sustained capacity, walking adaptability and actual performance are kept distinct. Measurement error and meaningful change are context dependent. Instrumented and video measures add potentially useful detail, but validity is specific to each parameter and algorithm, especially in slow asymmetric gait and with aids or orthoses.

Prospective recovery prediction is now supported by externally evaluated clinical tools, while those same validations reveal important calibration limits. Future-fall and community-walking studies provide useful signals but do not establish a universal gait-based risk engine. For rehabtools, the safest scientifically defensible route is transparent measurement first, contextual change interpretation second, and outcome-specific externally validated prognosis only when the complete deployed system has earned that claim.

Primary study characteristics

Selected original studies supporting the narrative are grouped by measurement or prognostic question. Population, protocol, endpoint and source access constrain interpretation. Related publications from one cohort are identified where established and should not be counted as independent replications.

Table 6 Clinical measurement primary study matrix

People, not ratings, strides or windows, define independent sample size.

Table 6 Clinical measurement primary study matrix
Study and designPopulation and protocolPrincipal findingLimits and source status
[2] Flansbjer UB 2005. Clinical gait-test repeatability.50 people, 6–46 months after stroke; retest after seven days for 48 participants and after 10 or 13 days for the other two; three were assessed at different times of day.ICCs 0.94–0.99; systematic retest change and absolute error also examined.Fairly active chronic sample. Same cohort as the companion Flansbjer strength-reliability paper; different assays do not create independent replication. Access: Full text publisher PDF examined
[3] Fulk GD 2008. Gait-speed MDC90.35 people, mean 34.5 days; middle 5 m of a 9-m comfortable walk.0.30 m/s overall; 0.07 assisted and 0.36 unassisted.Only 13 assisted and 22 unassisted; heterogeneity and protocol matter. Detailed methods remain abstract-level. Access: Abstract
[4] Lewek MD 2019. Between-visit comfortable and fast speed reliability.76 people; mean 52 months, range 5–324; 42 left and 34 right paretic. Three trials per speed; unchanged aids/AFOs.Comfortable MDC95 0.10/0.15/0.18 m/s in low/middle/high speed strata.Visits 3–57 days apart; mean 0.05 m/s retest increase; small fast subgroup. Labelled chronic but includes five-month case. Access: Full text
[5] Hosoi Y 2023. Speed and time absolute error.84 people in three speed strata of 19, 29 and 36; closely repeated 10mWT.Comfortable-speed MDC95 0.05/0.11/0.21 m/s.Low-speed table mean conflicts with stated stratum; ICC description is inconsistent. Short-term error is not between-day error. Access: Full text
[6] Tilson JK 2010. Speed change anchored to mRS improvement.283 LEAPS participants; approximately 20–60 days; 10 m timed on 14-m course; usual aids and up to one-person maximal assistance.0.16 m/s; sensitivity 73.9%, specificity 57.0%.Broad disability anchor and severely impaired baseline walking. LEAPS overlaps later secondary walking/activity analyses. Access: Full text PMC HTML
[7] Fulk GD 2011. Participant and therapist global-change anchors.Subacute outpatients, mean 56 days at entry; baseline comfortable speed 0.56 m/s.Important-change estimates 0.175 and 0.190 m/s, respectively.Anchor and recall dependence. Full source methods and sample count not independently verified here. Access: Abstract
[10] Cheng DK 2020. Stroke-specific 10mWT and 15-/30-m corridor 6MWT.21 baseline and 20 retest; median 134 days; retest 1–3 days.MDC95 0.40 m/s, 44.0 m and 67.5 m, respectively.Small sample; corridor protocols differ. Abstract-level methods. Access: Abstract
[11] Fulk GD 2018. 6MWT change anchored to mRS or SIS.Trial data across 2–6 months; baseline-speed subgroup analysis.mRS estimate 71 m overall, AUC 0.66; 44 m in initially <0.4 m/s group, AUC 0.72.Poorer anchor discrimination in faster walkers. Not a universal 71-m threshold; underlying trial overlap requires checking. Access: Abstract
[12] Kubo, Hiroki 2025. External evaluation of prior 71-m 6MWT MIC and new estimates.107 rehabilitation inpatients, tested at 30 and 60 days.Prior rule LR+ 1.41 and LR− 0.77; FAC estimates 69.0 m by ROC and 63.1 m by adjusted method.Prior threshold did not meet validation criterion. Updated estimates arise in current cohort; full anchor/method details needed. Access: Abstract
[13] Khan M 2026. Prospective 2MWT anchor analysis.150 people; seven centres; mean 92 days after stroke; 72% used aids.Reported MIC 33 m, 95% CI 30–36, AUC 0.89.Narrative baseline/change values conflict with table; subgroup table mixes mean improvements and MIC. Do not implement as default. Access: Full text with reporting discrepancies
[14] Jonsdottir J 2007. DGI repeatability and concurrent validity.25 people, eligibility from three months, able to walk 10 m with or without aid.Total-score retest/inter-rater ICC 0.96; item ICCs 0.55–0.93.Source calls population chronic but eligibility includes late subacute. No observed later falls. Access: Abstract
[15] Thieme H 2009. German FGA observer reliability.28 ambulatory people within six months; direct and video observers.Intrarater ICC 0.97 and inter-rater 0.94.Repeated scoring of a performance differs from new-capture between-day repeatability. No prospective risk model. Access: Abstract
[16] Lin JH 2010. DGI, short DGI and FGA psychometrics; later Barthel correlation.45 longitudinal outpatients, 35 completing all follow-ups; separate 48-person reliability sample.FGA had least floor/ceiling effects; baseline scores correlated with later disability.Later association does not establish calibrated personal prognosis or increment beyond baseline disability. Access: Abstract
[17] Ng SSM 2026. FGA retest and concurrent classification.63 stroke survivors, mean 8.8 years; 29 left/34 right affected; 30 controls.Retest ICC(3,1) 0.85; inter-rater ICC(2,2) 0.94 (average-rating model); MDC 5.02 points.Stroke/control discrimination is not a future-fall threshold. MDC is not patient-important change. Access: Full text
[43] Wellmon R 2015. WGS rating reproducibility.Six stroke participants and fourteen therapists; same videos rescored three weeks later.Inter-rater ICC 0.83; intra-rater 0.91; MDC95 4.24 points.Six independent gait patterns. Captures rater error, not patient variability or new recording error. Access: Full text
[42] Hayashi S 2022. Acute-stage walking-speed change with patient/therapist global ratings and motor-FIM anchors.62 initially ambulatory people; baseline mean four days after stroke and approximately twelve-day follow-up; eighteen aid users (seven walkers and eleven canes) and two lower-limb orthosis users, with orthosis type unspecified; comfortable 10 m with 3-m auxiliary zones.Therapist ROC threshold 0.21 m/s, AUC 0.76, sensitivity 0.78, specificity 0.68. Patient/FIM AUCs 0.68/0.66 did not support overall ROC thresholds.Change-difference range 0.18–0.25 is not one ROC MCID. MDC95 0.13 from 44-person retest subset; initially nonambulatory patients omitted. Same study as companion BBS evidence. Access: Complete licensed publisher HTML text; walking-speed methods and tables examined

Table 7 Instrumented measurement primary study matrix

People, not ratings, strides or windows, define independent sample size.

Table 7 Instrumented measurement primary study matrix
Study and designPopulation and protocolPrincipal findingLimits and source status
[45] Moore SA 2017. Criterion comparison and repeatability of several gait outputs.25 enrolled, 23 analysed; mean 66 months; L5 AX3, laboratory and seven-day home recordings.Selected mean/temporal features performed better than many spatial/asymmetry outputs.Two fixed-AFO cases excluded because event detection failed. Home repeatability is not home criterion validity. Cohort appears reused in [46]. Access: Full text
[46] Buckley C 2020. Single-accelerometer asymmetry algorithm comparison.23 analysed with matching cohort descriptors and AFO exclusions to [45].Validity and reliability depend on the specific feature and formula.Methodological extension, not independent replication of [45]. A generic asymmetry claim is unsupported. Access: Full text
[47] Felius RAW 2022. Test–retest of spatiotemporal, spectral, complexity and asymmetry features.31 recruited from two rehabilitation centres; minimum speed 0.05 m/s; two-minute foot/back IMU protocol.Many timing measures reproducible; several entropy and local-divergence measures less reliable.Short recordings and limited strides; meaningful longitudinal change not established. Related programme and ethics identifier to [36]; participant overlap unresolved. Access: Full text
[48] Lanotte F 2023. Criterion validity and reliability.Sixteen chronic stroke participants and ten younger healthy controls; foot/L5 Opal sensors at 128 Hz; six stroke trials across comfortable and fast 10mWT conditions; aids allowed.Stroke spatial outputs showed underestimation and proportional bias; temporal cycle/cadence estimates generally stronger than support-phase measures.Small participant sample despite many footfalls; within-condition repeated trials, not between-day reliability. Abstract bias values must not become uniform stroke correction factors; consider clustering and control-age mismatch. Access: Full author-provided article text on ResearchGate; supplements not separately appraised
[49] Igarashi T 2023. Same-day trunk-acceleration error.19 subacute inpatients; mean age 75.4; L3 sensor at 200 Hz; two walks within 30 minutes; five central cycles.MDC regularity 0.149–0.179; harmonic ratio 0.666–0.864 by axis.Axis, filter, normalization and cycle-count specific. Neither MIC nor between-day MDC nor a falls cutoff. Access: Full text
[51] Alammari BJ 2025. Concurrent spatiotemporal validity.19 survivors >1 month, FAC ≥3; eight-camera KinaTrax against Zeno walkway; one matched cycle per trial.Speed and stride length excellent; comfortable paretic support ICC about 0.31; stride width poor.Parameter-specific agreement. Few cycles, no between-day responsiveness, free-living validation or prognosis. Access: Full text
[52] Barzyk Philipp 2024. SMARTGAIT versus Vicon.Eight people, 1–130 months, FAC 2–5; four aid/orthosis users; ten trials with moving smartphone.Speed ICC 0.997; step length 0.781; sagittal joint-angle RMSE about 3.5–4.6 degrees.Small proof of concept; aid-associated and gait-phase errors. Not a validated clinical-change or prognosis tool. Access: Full text
[50] Marsan T 2026. REEV SENSE foot IMUs against optical motion capture.Twenty enrolled; fourteen analysed: ten unaided and four cane users.After participant and measurement-error exclusions, temporal agreement strong; spatial accuracy poorer in slow/cane group.Six slowest excluded, including every walker user; intersystem differences beyond ±1.96 SD were also removed before the final accuracy analysis. Aid effects are confounded with speed and severity. Access: Full text

Table 8 Walking recovery primary study matrix

People, not ratings, strides or windows, define independent sample size.

Table 8 Walking recovery primary study matrix
Study and designPopulation and protocolPrincipal findingLimits and source status
[18] Veerbeek, J M 2011. EPOS development for FAC ≥4 at six months.154 first-ever ischemic strokes; unable independent walking; predictors at days 2, 5 and 9.Preserved sitting/leg function: 98% favourable estimate; both absent: 27% on day 2, 10% on day 9.Development probabilities; original body restricted. EPOS publications reuse the development cohort; later validations must remain distinct. Access: Abstract with model details verified in validation
[19] Veerbeek, Janne M 2022. External EPOS evaluation at three months, changing original six-month horizon.Two Swiss cohorts, 39 and 78; mainly first strokes and prestroke mRS ≤2.AUC 0.675/0.921 on days 1/8 and 0.801/0.846 on days 3/9.Small samples and wide uncertainty; day-one negative predictions unreliable. Not broad validation in severe prestroke disability. Access: Full text
[20] Vinzens L 2026. External EPOS evaluation for three-month independence.293 enrolled; 280 analysed, including 217 independent and 63 dependent at three months; thirteen losses including nine deaths. Prestroke independent walking required.AUC 0.74; Brier 0.14; calibration slope 0.64 and logit calibration-in-the-large −0.53. At threshold 0.5: 212 TP, 20 TN, 43 FP and 5 FN.About half of outcomes unblinded; mixed outcome ascertainment; centre-specific calibration; only four in original 85% probability group. Recalibration uses this cohort. Lucerne site, period and trial overlap [24]; do not assume independent samples. Access: Complete licensed publisher HTML text; main tables examined
[21] Smith MC 2022. Development of age/knee-extension/Berg TWIST 0–4 tool.93 people unable independent walking; clinical predictors at one week.Predictions at 4, 6, 9, 16 and 26 weeks; reported accuracy at least 83%.Original full body restricted; tool details verified in [22]. Different from TCT/hip-extension algorithm [24]. Access: Abstract with tool details verified in validation
[22] Smith MC 2026. Temporal validation then pooled probability revision of TWIST tool.98 recruited, 89 analysed; new Auckland cohort; frames not permitted in outcome testing.72 achieved independence by 26 weeks; C-statistics 0.81–0.91; Brier 0.11–0.15.Six deaths before independence excluded; some probabilities optimistic. Revised estimates use development plus validation cohorts and need fresh validation. Access: Full text
[23] Miyata K 2026. Geographic external TWIST evaluation and local probability adjustment.179 recruited, 145 analysed across three Japanese hospitals; 34 lost or died; predictors at 5–8 days.88 independent by 9 weeks, 17 more by 26; C-statistics 0.76–0.89; slopes 0.33–0.67.Negative calibration-in-the-large point estimates at all horizons; four-week CI includes zero. Predictor imputation and sensitivity analyses checked. Manual updates require independent validation; narrative inconsistencies noted. Access: Complete licensed publisher HTML text; main tables examined
[24] Rajkovic, Lara Anka 2026. Older TWIST TCT/hip-extension algorithm, not the 0–4 tool.149 enrolled, three deaths excluded, 146 analysed; within-72-hour predictors; prestroke independence; rolling walkers allowed for outcome FAC.99 independent by six weeks, 26 more by twelve, 21 dependent. Multiclass AUC 0.73; overall accuracy 0.69. Only 13/38 predicted dependent were dependent; PPV 0.34.Intermediate-category sensitivity 0.23. Only 25% of outcomes blinded; mixed direct/telephone assessment. Different TWIST model and timing; no calibrated probabilities. Recruitment programme overlaps [20]. Access: Complete licensed publisher HTML text; main tables examined

Table 9 Rehabilitation outcomes primary study matrix

People, not ratings, strides or windows, define independent sample size.

Table 9 Rehabilitation outcomes primary study matrix
Study and designPopulation and protocolPrincipal findingLimits and source status
[60] O'Brien MK 2024. Admission sensor models for discharge 10mWT, FIM and BBS categories.55 recruited; 32 ambulatory models; only eight nonambulatory patients tested, using combined training sample of fifty.Selected sensor models improved some endpoints; FIM performance similar to clinical benchmark.Nested leave-one-person-out internal validation only. BBS risk category is not observed future falls. Fifty participants reused in earlier regression publication. Access: Full text

Table 10 Prospective falls primary study matrix

People, not ratings, strides or windows, define independent sample size.

Table 10 Prospective falls primary study matrix
Study and designPopulation and protocolPrincipal findingLimits and source status
[25] Bower K 2019. Kinect fast gait and dynamic balance versus ordinal prospective-falls outcome.96 recruited, 81 followed; median 24 days after stroke; twelve-month follow-up.23 any-fallers, 13 recurrent; pelvic displacement IQR-OR 6.75, CI 2.20–20.74 after speed adjustment.Adjusted for country, prior falls, assistance and comfortable six-metre speed. Short capture, exploratory multiple models and no external calibration; OR is ordinal, not automatically binary. Access: Full text
[26] Punt M 2017. Comparison of gait and clinical predictor sets.Forty chronic community survivors; treadmill and seven-day accelerometry; calendars and monthly calls over six months.Fifteen fallers; internal AUC 0.73 laboratory, 0.72 daily life, 0.64 clinical.Handrail-dependent and incomplete cases excluded. Small component models; no external calibration. Related Punt papers not counted as independent cohorts. Access: Full text publisher HTML
[27] Tsang CSL 2022. Level/obstacle walking with auditory clock or Stroop tasks.93 people, mean 5.6 years after stroke; monthly falls calls for twelve months.36 fallers; selected model 80% correct, sensitivity 72%, specificity 84%.Selected signal was clock-task reaction time during obstacles, not generic gait dual-task cost. Complete adjustment/validation methods not verified. Access: Abstract
[28] Plummer P 2022. Mobility indices developed from obstacle crossing, step test, speed, Walk-12 and sometimes aid type.Derivation: 45 discharged home; three-month follow-up. Independent D3 validation: 30 participants, nine fallers. Index C: 54 pooled participants, including the original 45 plus nine new patients.Derivation: 21 fallers; apparent AUCs approximately 0.82–0.85. Independent D3 validation at the prespecified cutoff of 3.5: AUC 0.84 (95% CI 0.68–1.00), sensitivity 0.56, specificity 0.95.Many index versions and cutoffs selected in the derivation cohort. D3 was externally evaluated in 30 separate participants; its revised cutoff of 2.5 was selected in that validation sample and needs fresh testing. The pooled Index C analysis was not independent external validation. Composite creation does not erase selection optimism; persons are not events. Related obstacle publication shares cohort. Access: Full text
[29] Lam WW 2026. Exploratory markerless-gait prediction.50 entered, 46 analysed; mean 72 months; three iPads and three 3-m walks; eighteen-month interviews.Accepted manuscript reports combined AUC 0.81, sensitivity 0.84, specificity 0.79.Abstract says 13% fallers; manuscript table lists six single plus two recurrent. Few events, stepwise models, CI/p-value conflicts and no external validation. Prior cited device validation was upper-limb. Access: Final accepted manuscript and publisher abstract

Table 11 Community classification primary study matrix

People, not ratings, strides or windows, define independent sample size.

Table 11 Community classification primary study matrix
Study and designPopulation and protocolPrincipal findingLimits and source status
[31] Fulk, George D 2017. Cross-sectional current activity classification.441 participants from LEAPS and FASTEST datasets.Speed cutoffs 0.49/0.93 m/s; 6MWT single-variable AUC 0.82/0.76.No future horizon. Daily step volume is not walking location. LEAPS overlaps [6] and other secondary analyses. Access: Publisher PDF text and abstract
[32] Bansal, Kanika 2024. Concurrent speed/endurance categories versus home/community steps.60 chronic stroke survivors and 18 controls; seven-day StepWatch plus GPS.Medium- and high-capacity groups did not consistently differ in location-specific activity.Shows capacity–performance mismatch; not a prospective prognosis model. Access: Full text publisher PDF

Table 12 Community prognosis primary study matrix

People, not ratings, strides or windows, define independent sample size.

Table 12 Community prognosis primary study matrix
Study and designPopulation and protocolPrincipal findingLimits and source status
[33] Igarashi T 2023. Baseline comfortable 10mWT and 30-m corridor 6MWT versus telephone community category.92 enrolled, 78 followed; mean 21.9 days; FAC >2; six months after discharge.299 m/AUC 0.896 and 0.94 m/s/AUC 0.844 for higher-category contrast.Twenty lowest, sixteen least-limited and forty-two unlimited walkers. Same-cohort ROC thresholds; no adjustment, external validation or calibration; lower contrast weak. Access: Full text
[34] Mulder M 2019. CART development for questionnaire-defined community independence.243 patients discharged from inpatient rehabilitation, followed six months.At 0.5 m/s split, 181/193 favourable-branch and 27/50 unfavourable-branch classifications correct.Poorer unfavourable predictions; full pruning/validation methods not verified. Below threshold does not mean no recovery potential. Access: Abstract
[35] Rosa MC 2015. Prospective and concurrent ROC classification in the same study.35 ambulatory first ischemic strokes within three months; nine independent community walkers at six months.Baseline fast speed ≥0.42 m/s prognostic; later concurrent >0.84 m/s discriminative.Small event count, self-report outcome and no verified external evaluation. Timing-specific thresholds are not interchangeable. Access: Abstract
[36] Felius RAW 2025. Clinical and wearable predictors of later daily strides and speed.35 participants; many models use only about twenty complete records; admission/discharge to six months.Selected apparent adjusted R² 0.36–0.60; predictors varied by outcome and timing.Small complete-case models, univariate selection and no external validation. Shared research programme with [47]; exact overlap unresolved. Activity is not participation. Reported discharge range extends to 379 days despite six-month poststroke follow-up; individual ordering for discharge-based models requires clarification. Access: Full text
[37] Kim J.S. 2026. Retrospective later-outcome model comparing clinical and GAITRite additions.137 complete-case inpatients, admission FAC 2–3, onset 4–14 weeks; assessment within seven admission days and at least three comfortable GAITRite passes. At discharge: 89 outdoor, 48 indoor-only.MI/time clinical AUC 0.995 to 0.998 with affected support percentages; LR p=0.009, DeLong p=0.262. Reported corrected AUCs 0.994/0.996 and apparent Brier scores 0.029/0.016.No external validation or fixed horizon. Full selection-within-bootstrap procedure unverified. Same-cohort calibration is apparent; Firth does not remove all selection optimism. Individual support coefficients not independently interpretable. Baseline FAC was excluded, narrowing the comparator. Access: Complete licensed publisher HTML text; main methods and four tables examined; supplements/code not separately appraised

Search and source access appendix

Searches were run on 2 October 2026. No claim of registered or exhaustive systematic review is made. The fully retrieved sets below supplied a broad discovery pool; selection for detailed appraisal was purposive and source led, rather than a formal duplicated eligibility process. Searches can change as databases update. PubMed connector pagination and official NCBI reconciliation are reported separately because they produced materially different unique-record coverage.

PubMed measurement query:

(stroke[Title/Abstract] OR poststroke[Title/Abstract]) AND (gait[Title/Abstract] OR walking[Title/Abstract]) AND (reliability[Title/Abstract] OR validity[Title/Abstract] OR "measurement error"[Title/Abstract] OR "minimal detectable"[Title/Abstract] OR "clinically important"[Title/Abstract])

PubMed prognosis query:

(stroke[Title/Abstract] OR poststroke[Title/Abstract]) AND (gait[Title/Abstract] OR walking[Title/Abstract] OR ambulation[Title/Abstract]) AND (prognosis[Title/Abstract] OR prognostic[Title/Abstract] OR prospective[Title/Abstract] OR prediction[Title/Abstract]) AND (independence[Title/Abstract] OR recovery[Title/Abstract] OR fall*[Title/Abstract] OR community[Title/Abstract] OR participation[Title/Abstract])

PubMed review-mapping query:

(stroke[Title/Abstract] OR poststroke[Title/Abstract]) AND (gait[Title/Abstract] OR walking[Title/Abstract]) AND ("systematic review"[Title/Abstract] OR "meta-analysis"[Title/Abstract]) AND (measurement[Title/Abstract] OR prognosis[Title/Abstract] OR prediction[Title/Abstract] OR reliability[Title/Abstract] OR wearable[Title/Abstract])

Official reconciliation appended this date restriction to each query: AND ("1800/01/01"[Date - Publication] : "2026/10/02"[Date - Publication]). The connector was limited through publication year 2026. ESearch counts were 656, 543 and 102, all retrieved by identifier through EFetch. Unique counts matched those totals within each query; the merged pool contained 1,219 unique records. The connector alone had yielded only 437, 380 and 86 unique identifiers despite complete paging. Two official records were book chapters, identified and accounted for separately rather than silently dropped. These are retrieval totals, not numbers of eligible studies or full texts reviewed.

Scopus title-focused query:

TITLE(stroke AND (gait OR walk*) AND (reliab* OR valid* OR "clinically important" OR "detectable change" OR predict* OR prognos*))

The Scopus year restriction was through 2026. All twelve pages were retrieved, with 286 records and 286 unique provider identifiers. The provider returned 25 records on its first page despite a requested limit of 50; subsequent pagination followed the actual page size. Title-focused retrieval is deliberately narrower than an all-fields search and does not establish complete Scopus coverage of the topic.

Targeted PubMed searches addressed influential clinical walking papers, EPOS, time-limited tests and asymmetry. Those were discovery searches rather than separately claimed exhaustive datasets. Backward citation chasing used the 2012 walking-tests review [61], and newer reviews and primary papers were checked for relevant originals. Review references, databases and publisher titles were not assumed to be error free; original Methods and Results controlled interpretation whenever available.

Source access and priority gaps

The linked bibliography identifies the actual access level for each source. Full text includes retrieved article-body sections or directly examined publisher/repository bodies, not abstracts repeated across databases. Where needed, paginated connector retrieval was continued. Important study-specific tables were checked separately, including the accessible accepted manuscript for Lam and colleagues. Its status is a final accepted manuscript in the Hong Kong Polytechnic University archive, not an independently verified copy of every page of the publisher’s final body. Repository record: https://ira.lib.polyu.edu.hk/handle/10397/118239?mode=full

After recovery of complete originals for the Japanese TWIST, heterogeneous EPOS, early TWIST-algorithm and incremental gait-parameter evaluations, the principal remaining high-value full-text gap is the 2025 6MWT important-change external validation. Ross Lit Search attempted its public full-text and configured publisher routes; public publisher pages, PDF links and exact-DOI repository searches did not provide a usable complete body. An additional full-text connector was unavailable because of account entitlement. Authorized institutional retrieval subsequently supplied complete Japanese TWIST, heterogeneous EPOS and distinct early TWIST-algorithm article text, and their analyses and access labels were updated. Complete licensed publisher HTML also resolved the incremental instrumented-gait study. Public author-provided full article text resolved the commercial-sensor study, revealing parameter- and group-specific qualifications hidden by its abstract shorthand. The final institutional check for the 6MWT article found only an alternative library-request record, not a usable subscription body; no request was placed, and its abstract-only status is retained. No access limitation is interpreted as absence of a study or of lawful full text elsewhere.

Other abstract-level sources are labelled in the bibliography. Their limited use is intentional: the report does not infer unobserved calibration, adjustment, confidence intervals, independent validation or missing-data procedures. Future access should update the specific affected interpretation, rather than retroactively implying that the present report examined those methods.

References

References are numbered in first citation order. Study specific source descriptions identify the material examined and do not constitute a study quality rating. Links identify the original publication or the explicitly named primary source version.

1. Van Criekinge T, Heremans C, Burridge J, Deutsch JE, Hammerbeck U, Hollands K, et al. Standardized measurement of balance and mobility post-stroke: Consensus-based core recommendations from the third Stroke Recovery and Rehabilitation Roundtable. International journal of stroke : official journal of the International Stroke Society. 2024;19(2):158-168. DOI 10.1177/17474930231205207 Source examined: Full publisher article text; graphical boxes and supplements not separately appraised.

Source note: SRC-723517f23983 Van Criekinge T 2024

2. Flansbjer UB, Holmbäck AM, Downham D, Patten C, Lexell J. Reliability of gait performance tests in men and women with hemiparesis after stroke. Journal of rehabilitation medicine. 2005;37(2):75-82. DOI 10.1080/16501970410017215 Source examined: Full text publisher PDF examined.

Source note: SRC-2339a5515b9d Flansbjer UB 2005

3. Fulk GD, Echternach JL. Test-retest reliability and minimal detectable change of gait speed in individuals undergoing rehabilitation after stroke. Journal of neurologic physical therapy : JNPT. 2008;32(1):8-13. DOI 10.1097/npt0b013e31816593c0 Source examined: Abstract.

Source note: SRC-94b70c9747d7 Fulk GD 2008

4. Lewek MD, Sykes R. Minimal Detectable Change for Gait Speed Depends on Baseline Speed in Individuals With Chronic Stroke. Journal of neurologic physical therapy : JNPT. 2019;43(2):122-127. DOI 10.1097/npt.0000000000000257 Source examined: Full text.

Source note: SRC-5ad9f450843f Lewek MD 2019

5. Hosoi Y, Kamimoto T, Sakai K, Yamada M, Kawakami M. Estimation of minimal detectable change in the 10-meter walking test for patients with stroke: a study stratified by gait speed. Frontiers in neurology. 2023;14:1219505. DOI 10.3389/fneur.2023.1219505 Source examined: Full text.

Source note: SRC-ce9af60bee49 Hosoi Y 2023

6. Tilson JK, Sullivan KJ, Cen SY, Rose DK, Koradia CH, Azen SP, et al. Meaningful gait speed improvement during the first 60 days poststroke: minimal clinically important difference. Physical therapy. 2010;90(2):196-208. DOI 10.2522/ptj.20090079 Source examined: Full text PMC HTML.

Source note: SRC-d747758e6c32 Tilson JK 2010

7. Fulk GD, Ludwig M, Dunning K, Golden S, Boyne P, West T. Estimating clinically important change in gait speed in people with stroke undergoing outpatient rehabilitation. Journal of neurologic physical therapy : JNPT. 2011;35(2):82-9. DOI 10.1097/npt.0b013e318218e2f2 Source examined: Abstract.

Source note: SRC-5ee760204ed9 Fulk GD 2011

8. Bohannon RW, Andrews AW, Glenney SS. Minimal clinically important difference for comfortable speed as a measure of gait performance in patients undergoing inpatient rehabilitation after stroke. Journal of physical therapy science. 2013;25(10):1223-5. DOI 10.1589/jpts.25.1223 Source examined: Full text.

Source note: SRC-af5f0fa47287 Bohannon RW 2013

9. Fulk GD, Echternach JL, Nof L, O'Sullivan S. Clinometric properties of the six-minute walk test in individuals undergoing rehabilitation poststroke. Physiotherapy theory and practice. 2008;24(3):195-204. DOI 10.1080/09593980701588284 Source examined: Abstract.

Source note: SRC-da025921f545 Fulk GD 2008

10. Cheng DK, Nelson M, Brooks D, Salbach NM. Validation of stroke-specific protocols for the 10-meter walk test and 6-minute walk test conducted using 15-meter and 30-meter walkways. Topics in stroke rehabilitation. 2020;27(4):251-261. DOI 10.1080/10749357.2019.1691815 Source examined: Abstract.

Source note: SRC-f9bf28b9afc8 Cheng DK 2020

11. Fulk GD, He Y. Minimal Clinically Important Difference of the 6-Minute Walk Test in People With Stroke. Journal of neurologic physical therapy : JNPT. 2018;42(4):235-240. DOI 10.1097/npt.0000000000000236 Source examined: Abstract.

Source note: SRC-4d501e562e7e Fulk GD 2018

12. Kubo, Hiroki, Miyata, Kazuhiro, Tamura, Shuntaro, Kobayashi, Sota, Nozoe, Masafumi, Inamoto, Asami, et al. External Validation and Update of Minimal Important Change in the 6-Minute Walk Test in Hospitalized Patients With Subacute Stroke. Archives of physical medicine and rehabilitation. 2025. DOI 10.1016/j.apmr.2025.01.002 Source examined: Abstract.

Source note: SRC-a0ced7272d0a Kubo 2025

13. Khan M, Muzamil HS, Osailan AM, Alhammad AA, Khan S, Mushtaq M, et al. Validating the 2-minute walk test MCID for subacute stroke patients: A Pakistani multicenter cohort analysis. PloS one. 2026;21(4):e0347056. DOI 10.1371/journal.pone.0347056 Source examined: Full text with reporting discrepancies.

Source note: SRC-a38cab2b1104 Khan M 2026

14. Jonsdottir J, Cattaneo D. Reliability and validity of the dynamic gait index in persons with chronic stroke. Archives of physical medicine and rehabilitation. 2007;88(11):1410-5. DOI 10.1016/j.apmr.2007.08.109 Source examined: Abstract.

Source note: SRC-96117e35c060 Jonsdottir J 2007

15. Thieme H, Ritschel C, Zange C. Reliability and validity of the functional gait assessment (German version) in subacute stroke patients. Archives of physical medicine and rehabilitation. 2009;90(9):1565-70. DOI 10.1016/j.apmr.2009.03.007 Source examined: Abstract.

Source note: SRC-27a430b4f76f Thieme H 2009

16. Lin JH, Hsu MJ, Hsu HW, Wu HC, Hsieh CL. Psychometric comparisons of 3 functional ambulation measures for patients with stroke. Stroke. 2010;41(9):2021-5. DOI 10.1161/strokeaha.110.589739 Source examined: Abstract.

Source note: SRC-a4af589e5ccb Lin JH 2010

17. Ng SSM, Chen P, Li S, Chan PWT, Tsui AKY, Lai CYY, et al. Psychometric properties of functional gait assessment in people with stroke. BMC neurology. 2026;26(1):83. DOI 10.1186/s12883-025-04591-w Source examined: Full text.

Source note: SRC-b1bacd2a6594 Ng SSM 2026

18. Veerbeek, J M, Van Wegen, E E H, Harmeling-Van der Wel, B C, Kwakkel, G, EPOS Investigators. Is accurate prediction of gait in nonambulatory stroke patients possible within 72 hours poststroke? The EPOS study. Neurorehabilitation and neural repair. 2011. DOI 10.1177/1545968310384271 Source examined: Abstract with model details verified in validation.

Source note: SRC-433b9f540903 Veerbeek 2011

19. Veerbeek, Janne M, Pohl, Johannes, Held, Jeremia P O, Luft, Andreas R. External Validation of the Early Prediction of Functional Outcome After Stroke Prediction Model for Independent Gait at 3 Months After Stroke. Frontiers in neurology. 2022. DOI 10.3389/fneur.2022.797791 Source examined: Full text.

Source note: SRC-8dcbe96b1f7d Veerbeek 2022

20. Vinzens L, Betschart M, Veerbeek JM. External Validation of the EPOS Prediction Model for Independent Gait After Stroke. Neurorehabilitation and neural repair. 2026;40(8):665-677. DOI 10.1177/15459683261425934 Source examined: Complete licensed publisher HTML text; main tables examined.

Source note: SRC-e7f1c1671038 Vinzens L 2026

21. Smith MC, Barber AP, Scrivener BJ, Stinear CM. The TWIST Tool Predicts When Patients Will Recover Independent Walking After Stroke: An Observational Study. Neurorehabilitation and neural repair. 2022;36(7):461-471. DOI 10.1177/15459683221085287 Source examined: Abstract with tool details verified in validation.

Source note: SRC-b027f11ecee1 Smith MC 2022

22. Smith MC, Scrivener BJ, Stinear CM. Temporal External Validation of the TWIST Prediction Tool for Time to Independent Walking after Stroke. Neurorehabilitation and neural repair. 2026;40(6):461-471. DOI 10.1177/15459683261417638 Source examined: Full text.

Source note: SRC-8cd4f220c3cb Smith MC 2026

23. Miyata K, Takeda R, Kotajima K, Takahashi Y, Akiyama H, Hayashi S, et al. External Validation of the Walking Independence Prognostic Model TWIST Score After Stroke: A Multicenter, Prospective Cohort Study. Neurorehabilitation and neural repair. 2026;40(6):472-481. DOI 10.1177/15459683261418670 Source examined: Complete licensed publisher HTML text; main tables examined.

Source note: SRC-49921e02886b Miyata K 2026

24. Rajkovic, Lara Anka, Veerbeek-Preuss, Silke, Bärtschi, Sandrine, Ottiger, Beatrice, Nyffeler, Thomas, Veerbeek, Janne Marieke. External Validation of the Time to Walking Independently After Stroke (TWIST) Prediction Algorithm within 72 Hours Poststroke. Neurorehabilitation and neural repair. 2026. DOI 10.1177/15459683261477077 Source examined: Complete licensed publisher HTML text; main tables examined.

Source note: SRC-30665b26974d Rajkovic 2026

25. Bower K, Thilarajah S, Pua YH, Williams G, Tan D, Mentiplay B, et al. Dynamic balance and instrumented gait variables are independent predictors of falls following stroke. Journal of neuroengineering and rehabilitation. 2019;16(1):3. DOI 10.1186/s12984-018-0478-4 Source examined: Full text.

Source note: SRC-94f1a4507ae0 Bower K 2019

26. Punt M, Bruijn SM, Wittink H, van de Port IG, van Dieën JH. Do clinical assessments, steady-state or daily-life gait characteristics predict falls in ambulatory chronic stroke survivors? Journal of rehabilitation medicine. 2017;49(5):402-409. DOI 10.2340/16501977-2234 Source examined: Full text publisher HTML.

Source note: SRC-d5aff3418ca5 Punt M 2017

27. Tsang CSL, Miller T, Pang MYC. Association between fall risk and assessments of single-task and dual-task walking among community-dwelling individuals with chronic stroke: A prospective cohort study. Gait & posture. 2022;93:113-118. DOI 10.1016/j.gaitpost.2022.01.019 Source examined: Abstract.

Source note: SRC-f679c346369c Tsang CSL 2022

28. Plummer P, Feld JA, Mercer VS, Ni P. Brief composite mobility index predicts post-stroke fallers after hospital discharge. Frontiers in rehabilitation sciences. 2022;3:979824. DOI 10.3389/fresc.2022.979824 Source examined: Full text.

Source note: SRC-659533b2a679 Plummer P 2022

29. Lam WW, Ang WT, Fong KN. Prediction for prospective falls via gait evaluation using mobile devices for stroke survivors: A markerless motion analysis study. Clinical rehabilitation. 2026;40(7):959-971. DOI 10.1177/02692155251414356 Source examined: Final accepted manuscript and publisher abstract.

Source note: SRC-22d32ba17dba Lam WW 2026

30. Perry J, Garrett M, Gronley JK, Mulroy SJ. Classification of walking handicap in the stroke population. Stroke. 1995;26(6):982-9. DOI 10.1161/01.str.26.6.982 Source examined: Abstract and historical interpretation in later primary sources.

Source note: SRC-a688ade5010b Perry J 1995

31. Fulk, George D, He, Ying, Boyne, Pierce, Dunning, Kari. Predicting Home and Community Walking Activity Poststroke. Stroke. 2017. DOI 10.1161/strokeaha.116.015309 Source examined: Publisher PDF text and abstract.

Source note: SRC-cce536ea500b Fulk 2017

32. Bansal, Kanika, Fox, Emily J, Clark, David, Fulk, George, Rose, Dorian K. Speed- and Endurance-Based Classifications of Community Ambulation Post-Stroke Revisited: The Importance of Location in Walking Performance Measurement. Neurorehabilitation and neural repair. 2024. DOI 10.1177/15459683241257521 Source examined: Full text publisher PDF.

Source note: SRC-dc3ee575429b Bansal 2024

33. Igarashi T, Takeda R, Tani Y, Takahashi N, Ono T, Ishii Y, et al. Predictive discriminative accuracy of walking abilities at discharge for community ambulation levels at 6 months post-discharge among inpatients with subacute stroke. Journal of physical therapy science. 2023;35(3):257-264. DOI 10.1589/jpts.35.257 Source examined: Full text.

Source note: SRC-9a8ed2fc7c41 Igarashi T 2023

34. Mulder M, Nijland RH, van de Port IG, van Wegen EE, Kwakkel G. Prospectively Classifying Community Walkers After Stroke: Who Are They? Archives of physical medicine and rehabilitation. 2019;100(11):2113-2118. DOI 10.1016/j.apmr.2019.04.017 Source examined: Abstract.

Source note: SRC-6b5f729067f1 Mulder M 2019

35. Rosa MC, Marques A, Demain S, Metcalf CD. Fast gait speed and self-perceived balance as valid predictors and discriminators of independent community walking at 6 months post-stroke--a preliminary study. Disability and rehabilitation. 2015;37(2):129-34. DOI 10.3109/09638288.2014.911969 Source examined: Abstract.

Source note: SRC-1f772501159e Rosa MC 2015

36. Felius RAW, Punt M, Wouda NC, Geerars M, Bruijn SM, van Dieën JH. Predicting community walking after stroke. Frontiers in stroke. 2025;4:1523242. DOI 10.3389/fstro.2025.1523242 Source examined: Full text.

Source note: SRC-a513c9e7ab55 Felius RAW 2025

37. Kim J.S. Incremental predictive value of spatiotemporal gait parameters beyond clinical measures for achieving independent outdoor ambulation in subacute stroke patients: A retrospective cohort study. Gait and Posture. 2026. DOI 10.1016/j.gaitpost.2026.110264 Source examined: Complete licensed publisher HTML text; main methods and four tables examined; supplements/code not separately appraised.

Source note: SRC-7b953768ff3d Kim J 2026

38. Mahendran N, Kuys SS, Brauer SG. Accelerometer and Global Positioning System Measurement of Recovery of Community Ambulation Across the First 6 Months After Stroke: An Exploratory Prospective Study. Archives of physical medicine and rehabilitation. 2016;97(9):1465-1472. DOI 10.1016/j.apmr.2016.04.013 Source examined: Abstract.

Source note: SRC-43c1d900cd4b Mahendran N 2016

39. Karageorge A, Vargas J, Ada L, Kelly PJ, McCluskey A. Previous experience and walking capacity predict community outings after stroke: An observational study. Physiotherapy theory and practice. 2020;36(1):170-175. DOI 10.1080/09593985.2018.1484829 Source examined: Abstract.

Source note: SRC-45288357b93c Karageorge A 2020

40. Mehrholz J, Wagner K, Rutte K, Meissner D, Pohl M. Predictive validity and responsiveness of the functional ambulation category in hemiparetic patients after stroke. Archives of physical medicine and rehabilitation. 2007;88(10):1314-9. DOI 10.1016/j.apmr.2007.06.764 Source examined: Abstract.

Source note: SRC-eb7cac2d7d1d Mehrholz J 2007

41. Cleland BT, Perez-Ortiz A, Madhavan S. Walking test procedures influence speed measurements in individuals with chronic stroke. Clinical biomechanics (Bristol, Avon). 2020;80:105197. DOI 10.1016/j.clinbiomech.2020.105197 Source examined: Full text.

Source note: SRC-e926e2d5f3a5 Cleland BT 2020

42. Hayashi S, Miyata K, Takeda R, Iizuka T, Igarashi T, Usuda S. Minimal clinically important difference of the Berg Balance Scale and comfortable walking speed in patients with acute stroke: A multicenter, prospective, longitudinal study. Clinical rehabilitation. 2022;36(11):1512-1523. DOI 10.1177/02692155221108552 Source examined: Complete licensed publisher HTML text; walking-speed methods and tables examined.

Source note: SRC-ebd84c78fce9 Hayashi S 2022

43. Wellmon R, Degano A, Rubertone JA, Campbell S, Russo KA. Interrater and intrarater reliability and minimal detectable change of the Wisconsin Gait Scale when used to examine videotaped gait in individuals post-stroke. Archives of physiotherapy. 2015;5:11. DOI 10.1186/s40945-015-0011-z Source examined: Full text.

Source note: SRC-cecc52e448d5 Wellmon R 2015

44. Martínez-Pozo V, Barbado D, Díaz-Marín C, García-Campos J, Blasco-Peris C, Ros-Arlanzón P, et al. Concurrent Validity and Reliability of Inertial Sensor-Based Wearables for Quantifying Spatial-Temporal Gait Parameters After Stroke: A Systematic Review. Brain sciences. 2026;16(7). DOI 10.3390/brainsci16070662 Source examined: Full text systematic review.

Source note: SRC-d862414ce057 Martinez-Pozo V 2026

45. Moore SA, Hickey A, Lord S, Del Din S, Godfrey A, Rochester L. Comprehensive measurement of stroke gait characteristics with a single accelerometer in the laboratory and community: a feasibility, validity and reliability study. Journal of neuroengineering and rehabilitation. 2017;14(1):130. DOI 10.1186/s12984-017-0341-z Source examined: Full text.

Source note: SRC-541422194923 Moore SA 2017

46. Buckley C, Micó-Amigo ME, Dunne-Willows M, Godfrey A, Hickey A, Lord S, et al. Gait Asymmetry Post-Stroke: Determining Valid and Reliable Methods Using a Single Accelerometer Located on the Trunk. Sensors (Basel, Switzerland). 2020;20(1). DOI 10.3390/s20010037 Source examined: Full text.

Source note: SRC-6d42d4ba8ddc Buckley C 2020

47. Felius RAW, Geerars M, Bruijn SM, van Dieën JH, Wouda NC, Punt M. Reliability of IMU-Based Gait Assessment in Clinical Stroke Rehabilitation. Sensors (Basel, Switzerland). 2022;22(3). DOI 10.3390/s22030908 Source examined: Full text.

Source note: SRC-f811749395b9 Felius RAW 2022

48. Lanotte F, Shin SY, O'Brien MK, Jayaraman A. Validity and reliability of a commercial wearable sensor system for measuring spatiotemporal gait parameters in a post-stroke population: the effects of walking speed and asymmetry. Physiological measurement. 2023;44(8). DOI 10.1088/1361-6579/aceecf Source examined: Full author-provided article text on ResearchGate; supplements not separately appraised.

Source note: SRC-015c0e26c973 Lanotte F 2023

49. Igarashi T, Tani Y, Takeda R, Asakura T. Minimal detectable change in inertial measurement unit-based trunk acceleration indices during gait in inpatients with subacute stroke. Scientific reports. 2023;13(1):19262. DOI 10.1038/s41598-023-46725-5 Source examined: Full text.

Source note: SRC-b6c4be3b2db4 Igarashi T 2023

50. Marsan T, Clauzade S, Zhang X, Grandin N, Urman T, Linton E, et al. REEV SENSE IMUs for Spatiotemporal Gait Analysis in Post-Stroke Patients: Validation Against Optical Motion Capture. Sensors (Basel, Switzerland). 2026;26(2). DOI 10.3390/s26020667 Source examined: Full text.

Source note: SRC-aca0fe42610a Marsan T 2026

51. Alammari BJ, Schoenwether B, Ripic Z, Kirk-Sanchez N, Eltoukhy M, Bishop L. Validity of AI-Driven Markerless Motion Capture for Spatiotemporal Gait Analysis in Stroke Survivors. Sensors (Basel, Switzerland). 2025;25(17). DOI 10.3390/s25175315 Source examined: Full text.

Source note: SRC-234b35cef847 Alammari BJ 2025

52. Barzyk Philipp, Boden Alina-Sophie, Howaldt Justin, Stürner Jana, Zimmermann Philip, Seebacher Daniel, et al. Steps to Facilitate the Use of Clinical Gait Analysis in Stroke Patients: The Validation of a Single 2D RGB Smartphone Video-Based System for Gait Analysis. Sensors. 2024;24(23):7819. DOI 10.3390/s24237819 Source examined: Full text.

Source note: SRC-78adf129a4e2 Barzyk Philipp 2024

53. Lonini L, Moon Y, Embry K, Cotton RJ, McKenzie K, Jenz S, et al. Video-Based Pose Estimation for Gait Analysis in Stroke Survivors during Clinical Assessments: A Proof-of-Concept Study. Digital Biomarkers. 2022;6:9–18. DOI 10.1159/000520732 Source examined: Abstract and original-study details in later primary source.

Source note: SRC-a5583b6f4498 Lonini L 2022

54. Patterson, Kara K, Gage, William H, Brooks, Dina, Black, Sandra E, McIlroy, William E. Evaluation of gait symmetry after stroke: a comparison of current methods and recommendations for standardization. Gait & posture. 2009. DOI 10.1016/j.gaitpost.2009.10.014 Source examined: Abstract.

Source note: SRC-0c7c71cff1eb Patterson 2009

55. Kesar TM, Binder-Macleod SA, Hicks GE, Reisman DS. Minimal detectable change for gait variables collected during treadmill walking in individuals post-stroke. Gait & posture. 2011;33(2):314-7. DOI 10.1016/j.gaitpost.2010.11.024 Source examined: Incomplete connector retrieval.

Source note: SRC-d2c822b8a36a Kesar TM 2011

56. Patterson KK, Mansfield A, Biasin L, Brunton K, Inness EL, McIlroy WE. Longitudinal changes in poststroke spatiotemporal gait asymmetry over inpatient rehabilitation. Neurorehabilitation and neural repair. 2015;29(2):153-62. DOI 10.1177/1545968314533614 Source examined: Abstract.

Source note: SRC-da16003b7d3d Patterson KK 2015

57. Guzik A, Drużbicki M, Wolan-Nieroda A, Przysada G, Kwolek A. The Wisconsin gait scale - The minimal clinically important difference. Gait & posture. 2019;68:453-457. DOI 10.1016/j.gaitpost.2018.12.036 Source examined: Abstract.

Source note: SRC-c1823d3a2c18 Guzik A 2019

58. Preston E, Ada L, Stanton R, Mahendran N, Dean CM. Prediction of Independent Walking in People Who Are Nonambulatory Early After Stroke: A Systematic Review. Stroke. 2021;52(10):3217-3224. DOI 10.1161/strokeaha.120.032345 Source examined: Systematic review abstract.

Source note: SRC-78c196bf2408 Preston E 2021

59. Wouda NC, Knijff B, Punt M, Visser-Meily JMA, Pisters MF. Predicting Recovery of Independent Walking After Stroke: A Systematic Review. American journal of physical medicine & rehabilitation. 2024;103(5):458-464. DOI 10.1097/phm.0000000000002436 Source examined: Systematic review abstract.

Source note: SRC-0dfcf35b0bc1 Wouda NC 2024

60. O'Brien MK, Lanotte F, Khazanchi R, Shin SY, Lieber RL, Ghaffari R, et al. Early Prediction of Poststroke Rehabilitation Outcomes Using Wearable Sensors. Physical therapy. 2024;104(2). DOI 10.1093/ptj/pzad183 Source examined: Full text.

Source note: SRC-0a5af2985469 OBrien MK 2024

61. van Bloemendaal M, van de Water AT, van de Port IG. Walking tests for stroke survivors: a systematic review of their measurement properties. Disability and rehabilitation. 2012;34(26):2207-21. DOI 10.3109/09638288.2012.680649 Source examined: Systematic review abstract and citation network.

Source note: SRC-afdd32f222fd van Bloemendaal M 2012