In this report
Audited and Updated
Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.
PD-G01
The stated longitudinal-model covariates included Hoehn–Yahr stage, but not baseline cognition or MDS-UPDRS-III.
Type: confirmed report wording error. Audit disposition: supported with scope refinement.
Remaining limit: Explicitly say longitudinal-model covariates; do not generalize the omission to all Qu analyses.
PD-G02
Replace “non-dominant” with “not-determined (ND)”.
Type: terminology correction. Audit disposition: supported.
Remaining limit: Terminology correction only; no effect on the reported slope, SE or P value.
PD-G03
The five-predictor model was selected after consulting the internal-validation C-index; this may make internal performance optimistic. The geographic external cohort remains independent.
Type: methodological precision. Audit disposition: supported.
Remaining limit: Internal-validation reuse can bias internal estimates; preserve the distinct geographic external cohort.
PD-G04
Clarify whether the tabulated RMSE is on the transformed or back-transformed count scale. Do not assert a definitive log scale or an unqualified number of falls.
Type: unit scale precision. Audit disposition: supported as scale caution.
Remaining limit: Do not assert definitively that RMSE is measured either in log counts or in numbers of falls. Say transformed modelling was used and scale requires clarification.
PD-G05
Two-year cohort follow-up; pooled measurements targeted freezing conversion within the following year.
Type: prediction horizon precision. Audit disposition: supported.
Remaining limit: Clarify cohort duration versus prediction horizon; not a uniform baseline two-year prediction.
PD-G06
No independent external validation was reported; an overlapping broader-sample sensitivity analysis was performed.
Type: validation precision. Audit disposition: supported.
Remaining limit: An overlapping sensitivity analysis is not independent external validation; no apparent-only performance claim is reversed into an external validation claim.
PD-G07
Report 120/205 from the archived main-text capture and retain the abstract discrepancy. Label the evidence as an archived licensed HTML text capture, not a newly retrieved live publisher body.
Type: access update and source discrepancy. Audit disposition: supported archived capture.
Remaining limit: Preserve capture provenance and abstract/body conflict; no live-publisher authentication asserted.
Editorial record
- Audit status: supported archived capture. Preserve capture provenance and abstract/body conflict; no live-publisher authentication asserted.
- Audit status: supported. Internal-validation reuse can bias internal estimates; preserve the distinct geographic external cohort.
- Edited phrase under PD-G03 . Original wording: P0042
- Audit status: supported archived capture. Preserve capture provenance and abstract/body conflict; no live-publisher authentication asserted.
- Audit status: supported. Internal-validation reuse can bias internal estimates; preserve the distinct geographic external cohort.
- Audit status: supported as scale caution. Do not assert definitively that RMSE is measured either in log counts or in numbers of falls. Say transformed modelling was used and scale requires clarification.
- Audit status: supported with scope refinement. Explicitly say longitudinal-model covariates; do not generalize the omission to all Qu analyses.
- Edited phrase under PD-G01 . Original wording: P0104
- Audit status: supported with scope refinement. Explicitly say longitudinal-model covariates; do not generalize the omission to all Qu analyses.
- Audit status: supported. Terminology correction only; no effect on the reported slope, SE or P value.
- Audit status: supported. An overlapping sensitivity analysis is not independent external validation; no apparent-only performance claim is reversed into an external validation claim.
- Edited phrase under PD-G06 . Original wording: P0166
- Audit status: supported. Clarify cohort duration versus prediction horizon; not a uniform baseline two-year prediction.
Editorial nomenclature update — 4 October 2026 at 11:48:46 am (Australia/Brisbane): authored condition labels and report wording use Parkinson’s disease. Published article titles, exact quotations, recorded searches, identifiers and routes are preserved. This is a terminology edit, not a scientific correction.
Executive assessment
Gait measurements contain clinically relevant prognostic information in Parkinson’s disease, but the strength of the claim depends on the outcome and the type of evidence. Among the studies examined, the clearest practical evidence concerns future falls. Detailed gait technology is promising; its added value beyond clinical history and simple tests is much less consistently established. A device can measure impairment accurately without being able to give a reliable personal forecast.
Future falls: the three-step combination of previous falls, freezing and gait speed retains risk gradients in three independent cohorts. Their refitted AUCs range from 0.69 to 0.83, while absolute category risks vary. This supports cautious stratification rather than universal probabilities. A newer five-factor first-fall model has encouraging geographic external validation, but uses a clinical gait/postural composite and does not isolate instrumented-gait benefit. [1, 2, 3, 4]
Cognition and dementia: instrumented pace and variability are associated with subsequent change in specific cognitive domains, with important limits from attrition and modelling. Larger validated dementia scores include many non-gait predictors. MoPaRDS illustrates the distinction: good overall discrimination coexists with low sensitivity at its suggested threshold, and its falls/freezing item did not contribute significantly in the all-item model. Gait testing should complement cognitive assessment rather than replace it. [5, 6]
Freezing and motor progression: small prospective cohorts identify candidate pre-freezing patterns, but individual forecasting remains uncertain. A substantial instrumented-gait study found poor later progression prediction. More recent repeated-gait results are encouraging, although their prediction timing is not clearly separated from the data used to estimate progression. Monitoring change and forecasting future change must remain distinct. [7, 8, 9, 10]
Wider outcomes: disability and survival associations exist, but gait-specific individual forecasts remain limited. Evidence for future overall quality of life and institutionalisation was not established in the examined sources. [11, 12]
For rehabilitation tools: the most defensible near-term role is reproducible measurement, contextual interpretation and tracking under consistent conditions. A video-derived measure cannot automatically substitute for a walkway or wearable input in a published model. Numerical risk claims require validation of the actual measurement pipeline and prediction rule in the intended population, at a stated horizon.
The report gives greatest weight to primary longitudinal evidence and genuine external evaluation, while retaining null results, reporting conflicts and access limitations. It is a critical narrative synthesis of selected studies, not an exhaustive systematic review or a formal certainty assessment. The detailed matrices distinguish association, prediction, monitoring and historical classification so that promising results are not asked to support more than they demonstrate.
Scope and approach
This report examines what gait measurements can tell us about future outcomes in people with established Parkinson’s disease. Its central question is whether a gait measure improves a clinically useful forecast, rather than merely describing disease severity or distinguishing Parkinson’s disease from healthy ageing. The outcomes are future falls, cognitive decline and dementia, the development of freezing of gait, and motor or functional progression. Instrumented walkways, wearable sensors, clinical walking tests and gait-related clinical ratings are considered together, while their different measurement and prediction claims are kept separate.
The report is intended to support research and rehabilitation measurement decisions. It critically examines selected original articles, their populations, outcomes, estimates and validation strategies. It is a narrative evidence synthesis, not an exhaustive systematic review. Its conclusions apply to the studies actually examined; the absence of a verified example is not proof that no such study exists.
Article selection and examination
Original longitudinal studies were prioritised when baseline gait or a gait-related measure preceded a clinically relevant outcome. Targeted searches in PubMed, Scopus and OpenAlex, together with reference-led retrieval, were used to locate primary articles across the outcome domains. Search terms combined Parkinson’s disease and gait with falls, cognitive decline or dementia, freezing, progression, disability, longitudinal follow-up and prediction. Searches were purposive and not all result pages were screened. The review also examined selected classification and longitudinal-monitoring studies because they are important sources of overextended prognostic claims. These studies are identified as contextual evidence rather than treated as equivalent forecasts.
The critical appraisal focuses on participant selection, outcome ascertainment, missingness and attrition, candidate features relative to the available events, analysis and validation, calibration, and clinical applicability. These are consistent with contemporary prediction-model appraisal principles. No completed formal PROBAST or PROBAST+AI assessment, GRADE assessment, pooled effect estimate or exhaustive coverage claim is made.
Why the results are not pooled
The studies differ in disease duration and stage, prior falls and freezing, medication state, gait task, recording duration, sensor placement, outcome definition and follow-up. A six-month prediction of any fall is not the same target as recurrent falling over a year; cognitive-score change is not incident dementia; and the development of freezing over several years is not recognition of a freezing episode seconds before it occurs. Combining these into a single performance estimate would conceal clinically important differences.
Reported performance is retained on its original scale and horizon. Hazard ratios, odds ratios, regression coefficients and classification measures are not treated as interchangeable. Within-study comparisons are generally more informative than comparisons of headline accuracy across unrelated cohorts. Where confidence intervals are wide, outcomes are rare, or the validation is incomplete, the uncertainty remains attached to the result rather than moved into a generic closing disclaimer.
Five questions that should not be combined
A measurement can be useful without being an individual prognostic model. Five different questions recur in this literature. First, does the measure distinguish people who already have an outcome, such as current freezers or retrospective fallers? This is classification. Second, does baseline gait have an association with a later outcome after accounting for other variables? This is prognostic association. Third, can a model estimate an individual person's future risk accurately at a stated time horizon? This is prediction. Fourth, does change in a measure capture change in disease over time? This is longitudinal monitoring. Fifth, does acting on the result improve care or outcomes? This is clinical utility. Evidence answering one question does not automatically answer the next.
A statistically significant hazard ratio or regression coefficient can establish a group-level association while leaving individual prediction unresolved. A model can also rank people well yet overestimate their risks. Discrimination describes ranking; calibration compares predicted probabilities with observed outcomes. Both matter for individual risk estimates. A useful decision threshold depends on the intended action and the consequences of false reassurance and unnecessary intervention. [13]
What validation does and does not establish
Apparent performance uses the development data and is vulnerable to optimism. Internal validation estimates optimism when all modelling steps are repeated within resampling. A random split from one source population is not equivalent to evaluation in a different service. External validation applies a fixed model to new data, preserving its prediction rule. Its relevance depends on the patients, measurement process, outcome and horizon matching the intended use. Discrimination, calibration and clinical usefulness require separate evaluation. [13]
Repeated walks, strides or sensor windows are not independent participants. If observations from one person appear in both training and test sets, performance may partly reflect recognition of that person rather than prediction in a new patient. The clinically relevant effective sample size depends on people and outcome events, not the number of extracted windows. Reuse of a cohort across several articles also does not create independent replication. This report therefore distinguishes participant counts, events, repeated observations and reused datasets wherever the source makes them clear. [14]
Future falls
Clinical interpretation
The prospective evidence supports clinically informative risk stratification, with more limited support for portable individual probabilities. Prior falls, freezing history and walking performance remain useful anchors. Sensors add plausible signals, particularly variability and turning, but impressive development results should not be mistaken for established clinical performance. First falls, any future fall, recurrent falls and fall counts are different targets and are considered separately below.
Simple clinical tools have external evidence
The three-step tool combines previous-year falls, recent freezing and comfortable gait speed below 1.1 m/s. Its threshold belongs to the original protocol and is not a universal definition of unsafe walking. The original Paul article was available only at abstract level; detailed interpretation therefore rests chiefly on its primary validation studies. [15, 1, 2]
Duncan and Lindholm provide independently recruited cohorts and a useful external test of risk categories. Importantly, both also refitted regression coefficients. Their headline AUCs concern these refitted analyses, rather than unequivocal validation of a fixed original probability equation. Applying the original weighted categories nevertheless produced clear risk gradients. The observed probabilities differed from the development estimates, so external stratification and transportable absolute risk should not be conflated. Slow gait contributed independently in both refitted analyses; freezing did not. Outcome ascertainment also differed: six-month recall in Duncan versus diaries and monthly calls in Lindholm. Neither demonstrated that using the score reduces falls. [1, 2]
Almeida adds a third independent cohort. The original risk categories again separated groups, but absolute probabilities were lower and the refitted model discriminated less well. The adjusted speed estimate was imprecise and crossed the null, although its point estimate was similar to Duncan's. This mixed evidence does not prove that gait has no value or that effects differ formally between cohorts. [3]
The Mini-BESTest offers complementary dynamic-balance information. Mak and Auyeung found a prospective association with recurrent falling, but the total score includes postural responses and sensory orientation as well as gait. Its gait subdomain did not significantly distinguish outcome groups. The test-alone AUC must be separated from the larger multivariable model's AUC; only 24 events supported a final model with ten predictors. A locally selected cut-point and bootstrap interval do not establish external validity or a universal clinical threshold. [16]
Instrumented gait signals are promising but often preliminary
Weiss provides a clinically relevant free-living signal among people reporting no fall in the preceding year: an anterior–posterior spectral-variability measure predicted an earlier first fall during follow-up after adjustment. The estimate was imprecise, with few events and a cohort-derived median threshold. This finding concerns a particular sensor and processing method, not interchangeable measures of gait variability or a calibrated risk calculator. [17]
Hoskovcová reported exceptionally high apparent discrimination from OFF-state stride-time variability and cadence. Adding variability improved a clinical model, making this more informative than a gait-only association. However, selection and evaluation used the same small cohort; three people could not complete OFF-state gait testing. The result warrants replication under a clinically feasible protocol, rather than direct implementation of the equation. [18]
Shah's one-week wearable study also produced high apparent discrimination, but selected combinations from 52 measures in only 34 people without independent testing. Comparison with falls history alone is not a test of added value beyond a strong clinical baseline. The report's main outcome description and its separate recurrent-faller analysis are not fully consistent, which matters before reusing outcome-specific performance. [19]
Sotirakis extended first-fall follow-up to five years, but the machine-learning samples were small after class balancing. The selected kinematic inputs included standing sway as well as gait. A conflict between the Methods and Discussion about whether feature selection preceded cross-validation leaves possible optimism unresolved; it is not proof of an implementation error. Performance remains internally evaluated, with very wide confidence intervals, few events, multiple pipelines and some recall-based follow-up. [20]
A newer externally evaluated first fall model
Wang's 2026 study is a genuine advance: a five-predictor early-PD clinical model was evaluated in an independent Chinese cohort and reported external discrimination, calibration plots and decision curves. It includes PIGD, body mass index, orthostatic hypotension, cognition and depression. It does not isolate instrumented gait or the incremental contribution of PIGD. Selection for complete fall follow-up, stepwise development and imputation before the internal split require caution. The five-predictor model was selected after consulting the internal-validation C-index, which may make internal performance optimistic; the geographic external cohort remains independent. The calibration figure's probability orientation needs clarification; a complete usable risk equation was not verified. Decision curves do not demonstrate reduced falls in practice. [4]
This study changes a blanket claim that first-fall models lack external evidence. It does not settle whether objective gait technology adds enough information to improve treatment decisions. The next interpretive step is to distinguish a model's overall performance from the contribution of one gait-related component, and to assess whether its probabilities and workflow transport to the intended clinic.
Other outcomes and important boundaries
Greene combines a very small prospective diary dataset with a separate dataset using previous-year falls. A headline result averaged across both is therefore not a single future-fall forecast. Some pretrained components were tested independently, but target-data fitting and sensitivity to extreme counts limit the claim. Predicted counts have not thereby been validated as a treatment-trial surrogate for observed falls. [21]
Kwon's de novo cohort found an adjusted association between backward speed and later falling. Without verified discrimination, calibration or an externally tested threshold, this remains prognostic-factor evidence. [22]
Mirando's 2026 study explicitly classifies prior falls. Its repeated visits and mobility measurements do not turn that target into future prognosis. It is relevant contextual classification research and is excluded from prospective accuracy conclusions. [23]
Implications for rehabilitation
Gait and balance findings can support a broader falls assessment without providing a personal percentage. Record actual falls and near-falls, freezing, circumstances, walking aids and medication state. Select the walking or balance test to match the clinical question and retain its protocol. OFF-state variability, ON-state comfortable speed, backward walking and free-living sensor features are not interchangeable.
A high-risk pattern does not identify which intervention will work, and improving a predictive feature does not necessarily prevent falls. The defensible role of current evidence is targeted assessment and cautious stratification, with emerging composite models evaluated on their own merits. The following matrix preserves the outcome, horizon and validation level behind each headline result.
Table 1 Primary evidence for future falls
Intervals are 95% confidence intervals unless labelled otherwise. Access describes the material examined, not study quality. Apparent performance uses development data; internal validation is not external validation.
| Study cohort and outcome | Measurement and main estimate | Validation limitations and access |
|---|---|---|
| Paul 2013 [15] 205 participants; event count inconsistent across accessed sources Any fall over 6 months | Prior falls, freezing and speed <1.1 m/s Development AUC 0.80 (95% CI 0.73–0.86) | Development detail incompletely verified Abstract reports 125 fallers; validation paper category counts total 120/205 Access: original abstract and later primary validation reports |
| Duncan 2015 [1] 171 participants; 66 fallers Any fall over 6 months | Three-step clinical composite Refitted AUC 0.83 (0.76–0.89) Adjusted slow-speed OR 2.38 (1.07–5.31) | External cohort; original category risks 9%, 28%, 66% Headline AUC uses refitted model; outcome recalled at follow-up Access: full text and tables |
| Lindholm 2016 [2] 138 participants; 45 fallers; 135 model cases Any fall over 6 months | Three-step clinical composite Refitted AUC 0.74 (0.65–0.84) Adjusted slow-speed OR 2.88 (1.24–6.72) | External cohort; original category risks 13%, 38%, 63% Diaries and monthly calls; mild selected population Access: full text and tables |
| Almeida 2021 [3] 137 participants; 42 fallers Any fall over 6 months; diaries and monthly calls | Three-step clinical tool Refitted AUC 0.69 (0.60–0.78); adjusted slow-speed OR 2.38 (0.86–6.57) | External original categories: 1/23, 14/55 and 27/59 fell Speed estimate imprecise; lower gait speeds and different absolute risks; no impact study Access: full primary HTML and tables |
| Mak and Auyeung 2013 [16] 110 participants; 24 recurrent fallers More than one fall over 6 months | Mini-BESTest: AUC 0.75; adjusted OR 0.75 per point, p=.014 Combined-model AUC 0.895 versus 0.868 without Mini-BESTest; difference p=.250 | Development only; 10 predictors for 24 events Cut-point 19/28 was sample-derived; total balance score is not pure gait Access: full text and tables |
| Weiss 2014 [17] 67 with no previous-year falls; 14 subsequent fallers First fall during 12-month follow-up | Three-day lower-back accelerometry Anterior–posterior spectral variability above the sample median: adjusted reported risk ratio 7.03 (1.82–37.13) | Adjusted association, no validated absolute-risk equation Few events; wide interval; specific processing method and long walking bouts Access: full text |
| Hoskovcová 2015 [18] 45 participants; 27 fallers; 42 completed OFF-state gait Any fall over 6 months | OFF-state instrumented TUG variability and cadence Apparent gait-model AUC 0.988; clinical model 0.921; combined model 0.995 AUC CIs not reported | Same-cohort selection and evaluation; no independent validation or calibration Medication-withdrawal protocol and missing OFF tests limit applicability Access: full text; supplement not separately read |
| Shah 2023 [19] 34 participants; main analysis 25 fallers; recurrent analysis 20 12 months after one-week recording | Foot and lumbar gait/turning features Best apparent AUC 0.94 (0.84–1.00); history alone 0.77 (0.50–1.00) | No independent test; extensive search across 52 features Main versus recurrent outcome definition is inconsistent; clinical-baseline added value untested Access: full text; supplement unverified |
| Sotirakis 2024 [20] 13/98 fallers at 2 years; 23/97 at 5 years Balanced modelling samples 26 and 46 | Two-minute walk and standing sway Five-year random-forest AUC 0.848 (0.519–1.000) with age; 0.812 (0.482–1.000) for kinematics alone | Internal cross-validation only; very wide intervals Feature-selection timing inconsistently reported; balancing changes prevalence Access: full text and supplementary Table 3 |
| Wang 2026 [4] Development 283/49 events; internal 120/19; external 150/31 First fall within 36 months | PIGD plus BMI, orthostatic hypotension, MoCA and depression External C-index 0.825 (0.768–0.882); 36-month AUC 0.884 (0.829–0.939) | Independent geographic evaluation; calibration plots and decision curves No model-minus-PIGD comparison or impact trial; selected complete follow-up; five-predictor selection consulted internal-validation C-index, potentially biasing internal performance; calibration display needs clarification Access: full text, tables and figures |
| Greene 2022 [21] Prospective PD1 n=15; 8 fallers and 181 falls 24-week diaries; separate PD2 n=26 historical counts | Instrumented TUG and pretrained components Prospective PD1 fall-risk mapping R² 0.50, RMSE 1.27; after outlier removal R² 0.73, RMSE 0.41 | Some independent component testing; target count mapping fitted Tiny sample, outlier sensitivity and mixed outcome timings; no surrogate validation Access: full text and XML Tables 1–3 |
| Kwon 2024 [22] 76 de novo participants; 16 fallers Any fall over 12 months | Instrumented forward and backward gait Backward-speed adjusted OR 0.9556 per cm/s (0.9166–0.9962), p=.0324 | Adjusted association; no AUC, calibration or external threshold validation 16 events; multiple candidate measures; 14 of 90 eligible participants not followed Access: full text and tables |
Paul development category probabilities were reported as 17%, 51% and 85%; Duncan reproduces actual derivation fractions of approximately 19%, 49% and 85%. These should not be interchanged. Shah’s supplementary outcome definitions remain unverified. Sotirakis supplementary Table 3 was inspected for the wide five-year AUC intervals. [15, 1, 19, 20]
Cognitive decline and dementia
Clinical interpretation
Quantitative gait analysis, clinical gait-related syndromes and multivariable dementia-risk scores provide different kinds of evidence. Instrumented studies link gait to later change in selected cognitive domains; freezing and axial phenotypes identify groups with greater cognitive vulnerability; and broader clinical models can forecast dementia. These findings do not combine into one gait-specific accuracy estimate.
The most important unresolved question is added value beyond age, disease severity and baseline cognition. The evidence supports gait as a complementary warning signal and candidate prognostic biomarker. It is weaker for a portable, calibrated gait-only dementia calculator, or a claim that gait testing should replace cognitive assessment.
Direct quantitative gait evidence
Morris studied newly diagnosed PD with instrumented single- and dual-task walking before cognitive follow-up. Selected pace and variability measures predicted domain-specific decline and improved fit over a model containing MoCA. This is longitudinal evidence, not same-day correlation. However, the outcomes were attention and memory trajectories rather than dementia conversion; executive decline was not predicted. Baseline MoCA itself predicted attention decline. Stepwise modelling, multiple comparisons, omission of motor severity from the stated base covariates and differential dropout limit interpretation. No independently validated individual-risk model was produced. [5]
Kim's newer study found a dual-task arm-swing-asymmetry effect in a model of two-year cognitive decline. Only the original abstract was accessible. The reported accuracy belongs to the full combined model; event count, decline definition, uncertainty, calibration, validation and improvement over the covariate-only model were not verified. This is a promising lead, not an established gait accuracy estimate. [24]
Clinical gait syndromes and imperfect replication
Anang's 2014 cohort linked clinical gait involvement, freezing and falls with later dementia, but Timed Up and Go was not predictive. Incremental benefit beyond baseline cognitive impairment was not demonstrated. [25]
The independent validation materially qualifies that result: the combined falls/freezing association was smaller and not statistically significant, with a wide interval. The article also reports different dementia counts and follow-up durations in its abstract and Results, and a different group size in Table 1. The matrix preserves that unresolved discrepancy rather than choosing one account silently. It does not alter the reported nonreplication of the combined item. [26]
Qu linked baseline freezing to later cognitive impairment in PPMI. Analyses of freezing progression and cognition evolving together answer a different question from a baseline forecast. The endpoint was not adjudicated dementia; the stated longitudinal-model covariates included Hoehn–Yahr stage, but not baseline cognition or MDS-UPDRS-III, and the smaller Chinese follow-up had considerable attrition. [27]
Other phenotype findings warrant proportionate confidence. Bugalho had only four dementia cases and could not fit a multivariable dementia model. Keener's tremor-dominant versus PIGD comparison did not survive false-discovery-rate correction. Michels found small phenotype-related cognitive differences, but substantial attrition and the use of broad clinical composites constrain gait-specific claims. Collectively, these studies support attention to axial dysfunction without establishing a single causal pathway or an individual dementia probability. [28, 29, 30]
A validated model can still leave gait added value unanswered
Lo's smartphone model genuinely predicted a future cognitive screening outcome. It combined gait and balance with voice, dexterity, reaction time and tremor. The headline AUC used recording-wise cross-validation that ignored within-person correlation. A separate clinic-only analysis, described as leave-one-subject-out, reported a lower AUC; it also changed feature count and recording source, so the difference cannot be attributed to validation alone. Grouping of repeated windows in that analysis remains unclear. Multiple windows are not independent patients. The outcome was crossing a MoCA threshold, not a confirmed dementia diagnosis, and isolated gait contribution was not established. [31]
Szwedo's 2025 MoPaRDS validation provides substantial positive evidence for a clinical dementia score across six incident-PD cohorts. Discrimination was reasonably good, but the suggested threshold had low sensitivity. Falls/freezing was associated with outcome alone but was not significant in the model containing all eight score items. This validates the broader score more than it validates gait's contribution. A low score cannot be treated as reassurance against long-term dementia. [6]
Bäckström similarly found a univariate PIGD association, but PIGD was not a distinct retained predictor in the selected clinical/biomarker model. Its AUC belongs to the whole model. Cohorts reused in later pooled analyses must not be counted as wholly independent replication. [32]
Evidence outside the direct prognosis question
Del Din's wearable study examined initially non-PD participants and later PD conversion. It is relevant to prodromal identification, rather than cognitive prognosis after diagnosis. Fereshtehnejad's reported ROC comparisons aligned eventual phenoconverters with healthy controls, rather than comparing baseline iRBD converters with nonconverters; they cannot be imported as prospective PD-specific conversion-risk accuracy. Dumurgier's gait and dementia study concerned the general population. These papers motivate hypotheses but do not directly validate prognosis in established PD. [33, 34, 35]
Implications for rehabilitation
New freezing, slowing, marked variability or emerging axial difficulty can justify closer assessment of cognition and everyday function in context. Gait can complement patient and informant history, cognitive screening and further evaluation when appropriate. An isolated adverse walking result should not diagnose impending dementia or determine an individual probability.
The central research requirement is a direct comparison of a suitable clinical model with and without prespecified gait measurements, using a defined cognitive endpoint and patient-level validation. Baseline cognition, disease stage, dropout and competing death need careful handling. Improving gait through rehabilitation and preventing cognitive decline are different outcomes: the observational associations reviewed here do not establish that treatment benefit.
Table 2 Primary evidence for cognitive outcomes
Intervals are 95% confidence intervals unless labelled otherwise. Access describes the material examined, not study quality. Apparent performance uses development data; internal validation is not external validation.
| Study cohort and outcome | Measurement and main estimate | Validation limitations and access |
|---|---|---|
| Morris 2017 [5] 119 newly diagnosed participants; 81 at 36 months Cognitive-domain trajectories | Instrumented two-minute single- and dual-task walking Speed × session coefficient for fluctuating attention −4.05 (SE 1.34); step length × session for visual memory 2.93 (SE 1.19) | Model-fit improvement beyond MoCA for selected measures; no dementia endpoint or external validation Differential dropout; multiple models; no stated motor-severity adjustment Access: full text |
| Kim 2024 [24] 48 PPMI participants Two-year cognitive decline; event count unverified | Dual-task arm-swing asymmetry selected, p=.030 Final combined model classified 78.7% correctly; CI not available in abstract | Clinical covariates entered first, then stepwise gait effects Decline definition, calibration, validation and change in performance unverified Access: original abstract only |
| Anang 2014 [25] 80 dementia-free participants; 27 dementia cases Mean follow-up 4.4 years | Clinical gait involvement, falls and freezing were associated with dementia TUG did not predict dementia status; exact TUG estimate unverified | Adjusted for age, sex, duration and follow-up; no demonstrated added value beyond baseline cognitive impairment Access: abstract and indexed primary-text passages; complete table image unavailable |
| Anang 2017 [26] Abstract 35/134 cases at 3.6 years; Results 30/134 at 4.2 years; Table 1 n=135 | Falls/freezing combined item Adjusted OR 1.8 (0.74–4.5), p=.20 | Independent-cohort association validation; initial association did not replicate significantly 32% loss; sample/event discrepancy unresolved Access: full primary HTML and Table 1 |
| Qu 2022 [27] PPMI longitudinal n=595; Chinese follow-up n=51 Cognitive impairment; mean PPMI follow-up 5 years | Baseline freezing: HR 1.53 (1.03–2.27) Freezing progression analysed separately from baseline status | Association evidence, not validated dementia risk model Event count unclear; stated longitudinal covariates included Hoehn–Yahr stage, but not baseline cognition or MDS-UPDRS-III; different freezing measures Access: full text |
| Lo 2019 [31] 110 unique participants; 126 cognitive windows; 28 future MoCA <26 outcomes 18-month horizon | Multicomponent smartphone battery AUC 0.97 with 998 features and recording-wise CV; 0.82 in reported clinic-only leave-one-subject-out analysis with 30 features | Internal evaluation; repeated-window grouping in leave-one-subject-out analysis unclear Features, recording source and validation change together; no gait ablation; MoCA threshold is not dementia Access: full text and tables |
| Bugalho 2013 [28] 61 reassessed from 75; 4 dementia cases 2 years | Clinical gait/postural score associated with dementia No greater FAB decline; numerical adjusted dementia effect unavailable | Too few events for multivariable dementia modelling Informative attrition; no external validation Access: full primary text; tables incompletely rendered |
| Keener 2018 [29] 224 without baseline MMSE impairment; 34 incident cases Mean detection time 5.9 years | Tremor-dominant versus PIGD HR 0.21 (0.05–0.91) Nominal p=.0368; false-discovery-rate p=.0859 | Subtype finding did not survive multiplicity correction MMSE cutoff, not adjudicated dementia; substantial illness/death-related attrition Access: full text and Table 4 |
| Michels 2022 [30] 711 baseline; 442 with sufficient 3-year visits; 142 at 6 years Cognitive trajectories | Broad motor phenotype PIGD versus not-determined (ND) MMSE slope −0.020/month (SE 0.009), p=.038 No attention phenotype effect in body results | No externally validated risk model Selective attrition; small scale-specific effects; abstract/body attention wording differs Access: full primary PDF |
| Szwedo 2025 [6] 1,108 incident PD; 350 dementia cases Up to 10.8 years; evaluation at 9.5 years | Eight-item MoPaRDS score Time-dependent AUC 0.79 (0.75–0.83); cutoff ≥4 sensitivity 21.7% and specificity 94.9% Falls/freezing not significant in all-item model | External score validation; no quantitative-gait addition Missing-item imputation and death censoring matter; overlaps earlier cohort publications Access: full text and tables |
| Bäckström 2022 [32] 143 incident PD; 64 dementia cases 10 years | PIGD univariate HR 3.2 (1.8–5.7) Selected clinical/CSF model AUC 0.86; PIGD not separately retained | Development model, no independent validation in this report AUC is not attributable to gait; NYPUM cohort overlaps later MoPaRDS validation Access: full text and tables |
Additional boundary studies discussed in the text concern prodromal conversion or general-population dementia and are not included as direct PD cognitive-prognosis rows. Fereshtehnejad was assessed from a recovered verified primary-source check; Dumurgier from its original abstract. [33, 34, 35]
Freezing motor progression and wider outcomes
Clinical interpretation
The evidence is outcome-specific. Incident freezing, motor-score change, ambulatory disability, treatment-associated improvement and survival are different prognostic targets. A measure that deteriorates over time is not necessarily a predictor of subsequent deterioration. Estimating a score at the same visit is not a forecast, and detecting an episode of freezing does not identify who will develop freezing next year.
Repeated gait assessment may reveal evolving vulnerability, but a repeat-measure model needs a clear prediction time and outcomes occurring strictly afterward. Across the sources examined, evidence for a general-purpose gait-based forecast of motor decline, dependence, quality of life or survival remains limited. This does not negate the value of measuring current impairment or monitoring change.
Motor progression and the value of negative evidence
Dewey directly tested whether baseline instrumented gait and balance forecast later clinical change, with poor held-out performance. The result challenges the assumption that measuring current severity implies forecasting progression. Its applicability is bounded by the tested ON-medication protocol and outcomes. The reported percentage accuracy was derived from Goodman–Kruskal gamma concordance; it is not ordinary classification accuracy or AUC and should not be compared casually with a 50% chance benchmark. [7]
Raschka found that two gait visits improved prediction of estimated clinical-score slopes compared with age and sex alone. However, those slopes used all longitudinal clinical data, without a clearly future-only outcome window after the second gait visit. Separate analyses in two cohorts were not external testing of a frozen predictive model. The findings support repeated assessment as a research direction while leaving temporal separation and individual prognostic performance unresolved. [8]
Sotirakis showed wearable-derived score estimates changing over 15–18 months. Features were selected using all visits, outcomes were contemporaneous scores, participant-disjoint splitting was unclear and missed visits used interpolation from preceding and subsequent visits. These are exploratory monitoring findings rather than a baseline forecast. [36]
Incident freezing and short term warning are different targets
D'Cruz prospectively observed conversion to freezing, but no individual gait marker survived multiple-comparison correction. The internally evaluated multivariable model combined clinical, cognitive and movement domains, with clinical rating and finger tapping particularly prominent. Predictors came from the assessment before conversion rather than uniformly from study entry. With only 12 conversions, the findings cannot establish isolated gait value or a fixed-horizon clinical test. [9]
Virmani found that worsening spatiotemporal measures before conversion could distinguish later freezers. Only nine converters contributed usable pre-conversion slopes, and estimates were univariable, internally unvalidated and imprecise. Slopes required repeated observations over variable intervals. Analyses using all observations also included post-conversion measurements and must not be presented as purely prospective prediction. [10]
Wang's five-year PPMI freezing model supplies clinical context. Its PIGD, fatigue, processing-speed and cerebrospinal-fluid predictors are not an instrumented-gait battery, and complete-case selection constrains application. No independent external validation was reported; an overlapping broader-sample missingness sensitivity analysis was performed. [37]
Algorithms classifying current freezers or anticipating an episode by seconds address other clinical questions. Such systems may help trigger cueing, but episode-level performance cannot establish months-to-years prognosis. Relevant evaluation includes false alarms, warning time, participant-level generalisation and benefit from acting on a warning.
Mobility and future disability
Venuto offers a useful clinical comparator: a model developed in Tracking Parkinson's was externally tested in PPMI and stratified later disability risk. Wearable activity was examined afterward and was not a prediction input. Its raw external accuracy needs the target prevalence and sensitivity/specificity beside it; an always-Stable rule would exceed the reported accuracy in this low-prevalence sample but detect no progressors. That does not make the model useless, but shows why accuracy alone is inadequate. It is evidence for clinical stratification, not incremental sensor benefit. [11]
A disability target should be explicit. A rating-scale change, Hoehn–Yahr stage 3, needing assistance and residential-care admission are not interchangeable. Related baseline and outcome items can partly reproduce existing severity. Group separation can still be informative, but a hazard ratio between predicted groups is not a calibrated personal probability of becoming dependent.
Quality of life institutionalisation and survival
Dewey's poor prognostic results included future PDQ-39 mobility and Schwab–England change. The mobility subscale is not global quality of life. Brzenczek’s quality-of-life analyses classified contemporaneous PDQ-39 rather than future quality of life. Separately, its motor-progression subgroup models using gait or integrated inputs did not exceed the clinical-only comparator. [7, 38]
Bäckström provides long-term survival context: slower baseline TUG was associated with shorter survival after age adjustment. Correlated PIGD was selected for the multivariable clinical model instead. The result does not demonstrate independent instrumented-gait value or justify estimating life expectancy from one TUG. Hazard-ratio scaling was insufficiently explicit for clinical reuse. [12]
An instrumented-gait model for institutionalisation, or a well-validated forecast of future overall quality of life, was not established in the examined sources. This is a bounded finding of this appraisal. Cross-sectional well-being associations and improvement after rehabilitation do not answer those future-outcome questions.
Treatment response is a distinct prognostic question
Cebi related preoperative levodopa responses to freezing after STN-DBS. Kinematic measures correlated with outcome but did not improve the exploratory model beyond clinical freezing response. A tiny selected sample, very high in-sample fit and a shared preoperative OFF baseline raise optimism and mathematical-coupling concerns. Follow-up measured combined medication and stimulation, not isolated DBS efficacy. [39]
Serrao linked baseline features to ten-week rehabilitation-associated change. Attrition, no untreated PD comparator and absent external validation limit causal and predictive interpretation. Because baseline value is mathematically part of change, regression to the mean or room to improve may contribute. These associations should not be used to deny rehabilitation to someone predicted to improve less. [40]
Average treatment efficacy and predicting differential benefit are separate questions. Improvement after therapy does not show that baseline gait selects the best treatment or identifies who benefits. That requires an appropriate comparison between treatment options and evaluation of how the predictor modifies benefit.
Implications for rehabilitation
Keep clinic capacity and everyday performance distinct. A short supervised walk, free-living walking quantity and real-world gait quality answer related but different questions. Serial worsening can prompt reassessment of falls, freezing, cognition and daily activity, rather than being communicated as an inevitable disease trajectory.
Interpret repeated measurements in light of medication, fatigue, intercurrent illness, orthopaedic problems, attention and environment. Patient-relevant mobility outcomes should accompany laboratory measures. A statistically detectable change, an in-sample fitted value or a group-level hazard ratio does not by itself establish meaningful benefit or a personal risk forecast.
Table 3 Primary evidence for freezing progression disability and treatment response
Intervals are 95% confidence intervals unless labelled otherwise. Access describes the material examined, not study quality. Apparent performance uses development data; internal validation is not external validation.
| Study cohort and outcome | Measurement and main estimate | Validation limitations and access |
|---|---|---|
| Dewey 2022 [7] Follow-up samples 230, 222, 164 and 177 6, 12, 18 and 24 months; continuous outcomes | ON-state instrumented TUG and sway 24-month iTUG–UPDRS gamma-based accuracy 58.21% (SD 8.87) | Repeated 10-fold cross-validation; no external test Poor later forecasts; statistic is not ordinary classification accuracy; missingness handling incompletely described Access: full text and Tables 1–3 |
| Raschka 2025 [8] Gait-linked samples 161 and 178 Estimated clinical slopes over irregular follow-up | One or two gait visits Two-visit Erlangen cross-validated R² approximately 0.45 for axial and 0.35 for UPDRS III slopes | Internal nested repeated validation; no frozen-model external test No clear future-only window; weak single-visit results; age/sex comparator Access: full text, tables and figure caption |
| D'Cruz 2020 [9] 60 enrolled; 56 analysed; 12 freezing conversions Two-year cohort follow-up; pooled measurements targeted freezing conversion within the following year | OFF-state gait plus clinical, cognitive and movement domains Gait-asymmetry AUC 0.79; composite bootstrap-averaged AUC 0.79; Brier score 0.14 | No individual marker survived multiplicity correction Internal evaluation; many candidates, few events; predictors pooled before conversion rather than fixed baseline Access: accepted full manuscript and Table 2 |
| Virmani 2023 [10] 30 baseline nonfreezers; 10 converted; 9 usable pre-conversion slopes Mean conversion time 2.3 years | Serial ON-state instrumented gait Pre-conversion swing-phase slope AUC 0.894 (0.749–1.000); stride-length slope 0.850 (0.676–1.000) | Univariable apparent estimates; no validation Variable observation windows; post-conversion-inclusive analyses are not forecasts Access: full text and tables |
| Venuto 2023 [11] Development 1,598; external 407 Approximately 4–4.5 years; Progressive prevalence 12% and 7% | 23 baseline clinical features; wearables not inputs External sensitivity 0.688, specificity 0.739, accuracy 0.735 Predicted Progressive group to ADL <80%: HR 2.6 (1.7–4.0) | Independent cohort testing; calibration/net benefit not established Class imbalance limits raw accuracy interpretation; no sensor added-value test Access: full text and tables; supplement uninspected |
| Bäckström 2018 [12] 143 incident PD; 77 deaths 8.5–13.5 years | Baseline TUG and PIGD Age-adjusted TUG HR 1.11 (1.07–1.15); multivariable PIGD HR 3.25 (1.75–6.05) | Exposure scaling insufficiently explicit for reuse TUG excluded from multivariable model because correlated with PIGD; no calibrated life-expectancy tool Access: full text and Table 4 image |
| Cebi 2020 [39] 18 implanted; 13 freezers; 12 kinematic datasets 6 months after STN-DBS | Preoperative levodopa response to freezing change Clinical-response model in-sample R² 0.952; kinematics did not improve prediction | No validation; tiny selected sample; shared baseline in change scores Combined medication/stimulation outcome; printed interval is not interpretable as an R² CI Access: full text and tables |
| Serrao 2019 [40] 50 eligible; 36 completed 10-week rehabilitation | Baseline gait and clinical features to gait change Speed-change model adjusted R² 0.561; baseline-speed coefficient −0.446 (SE 0.109), p=.001 | In-sample model; no untreated PD comparator or validation 28% incomplete follow-up; baseline–change coupling; change definition inconsistent Access: full text and tables |
Additional contextual originals were read in full for longitudinal monitoring, a clinical/biofluid freezing model and contemporaneous or motor-rate classification. Their supplementary files were not inspected. They are discussed in the chapter but are not treated as equivalent fixed-horizon gait forecasts. [36, 37, 38]
What gait technology adds
What counts as added value from gait
The key comparator is an appropriate clinical baseline. For falls, that may include falls history, freezing, a simple walking measure and other prespecified clinical predictors. For cognitive outcomes, it may include age, education and baseline cognition. A sensor-only model with a high area under the receiver operating characteristic curve does not show that sensors improve on information already available in ordinary care. Nor does a successful multimodal model prove that its gait component is responsible for the improvement.
A convincing incremental-value analysis compares a base model with the same model plus gait, using the same participants and horizon. It should assess discrimination, calibration and clinical usefulness. Feature selection, preprocessing and tuning belong within the training process, with model evaluation kept separate. These safeguards become particularly important when a brief walking task produces many correlated features. [14]
Match the measurement to the intended decision
A clinic can measure gait to describe current impairment, to monitor within-person change, or to stratify future risk. These uses should be named before selecting the technology. A longer or more technically complex recording is worthwhile only if it captures information that matters for that use. More features are not intrinsically better, and a model developed with one device, task or processing pipeline cannot be assumed to work with another.
For routine assessment, a defensible starting point is to document falls and near-falls, freezing, walking aids and the clinical context alongside the walking task. If an instrumented test is added, its contribution should be interpreted against this baseline. The evidence reviewed supports examining richer gait characteristics in selected settings; it does not justify replacing a broader assessment with a device-generated score.
Standardise conditions before interpreting change
Repeat measurements should use a consistent task, walking distance or recording duration, instructions, usual or fast walking condition, single- or dual-task protocol, walking aid, footwear where relevant, and environment. Medication state and time relative to dosing should be recorded. These are measurement safeguards: a change in protocol or context can masquerade as biological change.
Usual gait speed, turning, variability and dual-task performance are different constructs. Results should therefore retain the units and protocol that produced them. A threshold derived in one cohort should not be presented as universal. Similarly, an observed change should not be called clinically important merely because it exceeds zero or reaches statistical significance. A claim about meaningful change needs evidence appropriate to Parkinson’s disease, the instrument and the measurement conditions.
Keep three kinds of validation separate
Measurement validity asks whether the device or algorithm measures the intended gait quantity accurately and reliably. Prognostic validation asks whether a fixed combination of measurements predicts the specified future outcome in new people. Clinical utility asks whether using that result improves decisions or outcomes enough to justify its burden and harms. Success at the first level does not establish the second or third.
For rehabilitation software and tools, this distinction is especially important. A platform may responsibly display an accurately measured walking characteristic without being able to display an evidence-based probability of falling, dementia or progression. Published performance from another device or cohort does not validate the platform's own processing pipeline. A risk output needs a reproducible model, an intended population and horizon, suitable validation, calibrated probabilities and a clear action pathway. The articles examined here should inform these requirements rather than be treated as product certification.
What would change the conclusion
The most persuasive next evidence would show that gait adds value to a prespecified clinical baseline, with reliable outcomes, enough events for the model complexity, patient-level validation and transparent handling of missing data. A fixed model would then be evaluated in a different service or cohort using the intended hardware and workflow, with discrimination, calibration and uncertainty reported at a clinically meaningful horizon. If the output is intended to alter care, the final question is whether that change improves decisions or patient outcomes.
Existing cohorts, published model specifications and device-validation studies can help establish which claims are supported, reproducible and applicable to the intended workflow. A transparent measurement tool with modest claims is more defensible than a sophisticated risk label whose clinical meaning is unknown.
Bottom line for rehabilitation tools
A well-validated video or sensor measure can support assessment without carrying a validated prognostic claim. The original task, distance, pace, medication state, aid, device and predictor definitions matter. Substituting a video-derived estimate for a walkway or inertial-sensor input requires evidence of comparability and validation of the resulting prediction pipeline. The strongest reportable claim should be the one directly supported by that pipeline’s evidence.
Overall, gait is clinically informative and scientifically promising. The evidence is strongest for selected group-level associations and risk stratification, with some genuinely externally evaluated clinical composites. Evidence that detailed gait technology consistently adds portable, calibrated and decision-changing forecasts remains more limited. Measurement, prediction and clinical benefit should therefore be stated as separate claims.
Abbreviations
ADL, activities of daily living; AUC, area under the receiver operating characteristic curve; BMI, body mass index; CI, confidence interval; CSF, cerebrospinal fluid; CV, cross-validation; FAB, Frontal Assessment Battery; HR, hazard ratio; iRBD, isolated REM sleep behaviour disorder; MoCA, Montreal Cognitive Assessment; MMSE, Mini-Mental State Examination; MoPaRDS, Montreal Parkinson Risk of Dementia Scale; OR, odds ratio; PD, Parkinson’s disease; PDQ-39, Parkinson’s Disease Questionnaire; PIGD, postural instability and gait difficulty; PPMI, Parkinson’s Progression Markers Initiative; RMSE, root mean squared error; SD, standard deviation; SE, standard error; STN-DBS, subthalamic nucleus deep brain stimulation; TUG, Timed Up and Go; UPDRS, Unified Parkinson’s Disease Rating Scale.
References
Numbering follows first citation. DOI links identify original articles. The matrices and chapter notes distinguish complete primary-text access, partial primary-text access and abstract-only evidence; supplementary material was not uniformly available.
1. Duncan RP, Cavanaugh JT, Earhart GM, et al. External validation of a simple clinical tool used to predict falls in people with Parkinson disease. Parkinsonism & Related Disorders. 2015. DOI 10.1016/j.parkreldis.2015.05.008
Source note: SRC-06dd612a6dd5 Duncan RP 2015
2. Lindholm B, et al. External validation of a 3-step falls prediction model in mild Parkinson's disease. Journal of Neurology. 2016. DOI 10.1007/s00415-016-8287-9
Source note: SRC-faf11b104690 Lindholm B 2016
3. Almeida LRS, Piemonte MEP, Cavalcanti HM, Canning CG, Paul SS. A Self-Reported Clinical Tool Predicts Falls in People with Parkinson’s Disease. Movement Disorders Clinical Practice. 2021;8:427–434. DOI 10.1002/mdc3.13170
Source note: SRC-bfef263d9311 Almeida LRS 2021
4. Wang Y, Mei J, Tang Y, et al. Development and validation of a clinical prediction model for first fall in early Parkinson's disease: a study of two fall-naive cohorts. Frontiers in Aging Neuroscience. 2026;18:1735524. DOI 10.3389/fnagi.2026.1735524
Source note: SRC-08159f63c41f Wang Y 2026
5. Morris R, Lord S, Lawson RA, Coleman S, Galna B, Duncan GW, Khoo TK, Yarnall AJ, Burn DJ, Rochester L. Gait Rather Than Cognition Predicts Decline in Specific Cognitive Domains in Early Parkinson's Disease. The Journals of Gerontology, Series A. 2017;72(12):1656–1662. DOI 10.1093/gerona/glx071
Source note: SRC-b9ab344e30e0 Morris R 2017
6. Szwedo AA, Dalen I, Lawson RA, et al. Dementia risk prediction in early Parkinson's disease: Validation and genetic integration of the Montreal Parkinson risk of dementia scale (MoPaRDS). Journal of Parkinson's Disease. 2025. DOI 10.1177/1877718X251329857
Source note: SRC-239ca17a9991 Szwedo AA 2025
7. Dewey DC, et al. APDM gait and balance measures fail to predict symptom progression rate in Parkinson’s disease. Frontiers in Neurology. 2022;13:1041014. DOI 10.3389/fneur.2022.1041014
Source note: SRC-817facd83053 Dewey DC 2022
8. Raschka T, et al. Objective monitoring of motor symptom severity and their progression in Parkinson’s disease using a digital gait device. Scientific Reports. 2025;15(1):25541. DOI 10.1038/s41598-025-09088-7
Source note: SRC-98d900c5d991 Raschka T 2025
9. D’Cruz N, et al. Repetitive Motor Control Deficits Most Consistent Predictors of Conversion to Freezing of Gait in Parkinson’s Disease: A Prospective Cohort Study. Journal of Parkinson's Disease. 2020;10(2):559-571. DOI 10.3233/JPD-191759
Source note: SRC-4c57d2069017 DCruz N 2020
10. Virmani T, et al. Gait Declines Differentially in, and Improves Prediction of, People with Parkinson’s Disease Converting to a Freezing of Gait Phenotype. Journal of Parkinson's Disease. 2023;13(6):961-973. DOI 10.3233/JPD-230020
Source note: SRC-15e76b6bd8a5 Virmani T 2023
11. Venuto CS, et al. Predicting Ambulatory Capacity in Parkinson’s Disease to Analyze Progression, Biomarkers, and Trial Design. Movement Disorders. 2023;38(10):1774-1785. DOI 10.1002/mds.29519
Source note: SRC-33a028d2af8d Venuto CS 2023
12. Bäckström D, et al. Early predictors of mortality in parkinsonism and Parkinson disease: A population-based study. Neurology. 2018;91(22):e2045-e2056. DOI 10.1212/WNL.0000000000006576
Source note: SRC-8ae8f490fe88 Backstrom D 2018
13. Riley RD, Archer L, Snell KIE, et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ. 2024;384:e074820. DOI 10.1136/bmj-2023-074820
Source note: SRC-e131c017d8a9 Riley RD 2024
14. Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. DOI 10.1136/bmj-2024-082505
Source note: SRC-1edddd770a90 Moons KGM 2025
15. Paul SS, Canning CG, Sherrington C, et al. Three simple clinical tests to accurately predict falls in people with Parkinson's disease. Movement Disorders. 2013;28:655–662. DOI 10.1002/mds.25404
Source note: SRC-507aeb58eeff Paul SS 2013
16. Mak MKY, Auyeung MM. The mini-BESTest can predict parkinsonian recurrent fallers: a 6-month prospective study. Journal of Rehabilitation Medicine. 2013;45:565–571. DOI 10.2340/16501977-1144
Source note: SRC-09ed19718487 Mak MKY 2013
17. Weiss A, Herman T, Giladi N, Hausdorff JM. Objective assessment of fall risk in Parkinson's disease using a body-fixed sensor worn for 3 days. PLoS ONE. 2014;9:e96675. DOI 10.1371/journal.pone.0096675
Source note: SRC-ffa0d3447088 Weiss A 2014
18. Hoskovcová M, et al. Predicting falls in Parkinson disease: what is the value of instrumented testing in OFF medication state? PLoS ONE. 2015;10:e0139849. DOI 10.1371/journal.pone.0139849
Source note: SRC-90b7fd743cae Hoskovcova M 2015
19. Shah VV, et al. Gait and turning characteristics from daily life increase ability to predict future falls in people with Parkinson's disease. Frontiers in Neurology. 2023;14:1096401. DOI 10.3389/fneur.2023.1096401
Source note: SRC-aa0eb40379a7 Shah VV 2023
20. Sotirakis C, et al. Predicting future fallers in Parkinson's disease using kinematic data over a period of 5 years. npj Digital Medicine. 2024. DOI 10.1038/s41746-024-01311-5
Source note: SRC-683a44e25c70 Sotirakis C 2024
21. Greene BR, Premoli I, McManus K, et al. Predicting fall counts using wearable sensors: a novel digital biomarker for Parkinson's disease. Sensors. 2022;22:54 (published online December 2021). DOI 10.3390/s22010054
Source note: SRC-5e868996cddf Greene BR 2022
22. Kwon KY, et al. Association between baseline gait parameters and future fall risk in patients with de novo Parkinson's disease: forward versus backward gait. Journal of Clinical Neurology. 2024. DOI 10.3988/jcn.2022.0299
Source note: SRC-3e31a515f4a2 Kwon KY 2024
23. Mirando M, et al. Optimising fall risk classification models in Parkinson's disease using clinical and mobility outcomes. Parkinsonism & Related Disorders. 2026. DOI 10.1016/j.parkreldis.2026.108934
Source note: SRC-836345488f1f Mirando M 2026
24. Kim J, Rider JV, Zinselmeier A, Chiu YF, Peterson D, Longhurst JK. Dual-task gait has prognostic value for cognitive decline in Parkinson's disease. Journal of Clinical Neuroscience. 2024;126:101–107. DOI 10.1016/j.jocn.2024.06.006
Source note: SRC-d4f77db17022 Kim J 2024
25. Anang JBM, Gagnon JF, Bertrand JA, Romenets SR, Latreille V, Panisset M, Montplaisir J, Postuma RB. Predictors of dementia in Parkinson disease: a prospective cohort study. Neurology. 2014;83(14):1253–1260. DOI 10.1212/WNL.0000000000000842
Source note: SRC-3053d984b1a8 Anang JBM 2014
26. Anang JBM, Nomura T, Romenets SR, Nakashima K, Gagnon JF, Postuma RB. Dementia Predictors in Parkinson Disease: A Validation Study. Journal of Parkinson's Disease. 2017;7(1):159–162 (online 2016). DOI 10.3233/JPD-160925
Source note: SRC-377d1f01186b Anang JBM 2017
27. Qu Y, Li J, Chen Y, et al. Freezing of gait is a risk factor for cognitive decline in Parkinson's disease. Journal of Neurology. Published online 27 September 2022. DOI 10.1007/s00415-022-11371-w
Source note: SRC-da31fbefe85c Qu Y 2022
28. Bugalho P, Viana-Baptista M. Predictors of cognitive decline in the early stages of Parkinson's disease: a brief cognitive assessment longitudinal study. Parkinson's Disease. 2013;2013:912037. DOI 10.1155/2013/912037
Source note: SRC-b853197bbda7 Bugalho P 2013
29. Keener AM, Paul KC, Folle A, Bronstein JM, Ritz B. Cognitive Impairment and Mortality in a Population-Based Parkinson's Disease Cohort. Journal of Parkinson's Disease. 2018. DOI 10.3233/JPD-171257
Source note: SRC-4a346f16cbc5 Keener AM 2018
30. Michels J, van der Wurp H, Kalbe E, et al. Long-Term Cognitive Decline Related to the Motor Phenotype in Parkinson's Disease. Journal of Parkinson's Disease. 2022;12:905–916. DOI 10.3233/JPD-212787
Source note: SRC-a655054bd2be Michels J 2022
31. Lo C, Arora S, Baig F, et al. Predicting motor, cognitive & functional impairment in Parkinson's. Annals of Clinical and Translational Neurology. 2019. DOI 10.1002/acn3.50853
Source note: SRC-b4bd66ae9c21 Lo C 2019
32. Bäckström D, Granåsen G, Jakobson Mo S, et al. Prediction and early biomarkers of cognitive decline in Parkinson disease and atypical parkinsonism: a population-based study. Brain Communications. 2022. DOI 10.1093/braincomms/fcac040
Source note: SRC-e6c1e0c8b7e3 Backstrom D 2022
33. Del Din S, Elshehabi M, Galna B, et al. Gait analysis with wearables predicts conversion to Parkinson disease. Annals of Neurology. 2019;86:357–367. DOI 10.1002/ana.25548
Source note: SRC-2d31388c9cb7 Del Din S 2019
34. Fereshtehnejad SM, et al. Evolution of prodromal Parkinson's disease and dementia with Lewy bodies: a prospective study. Brain. 2019;142(7):2051–2067. DOI 10.1093/brain/awz111
Source note: SRC-843bc60acb63 Fereshtehnejad SM 2019
35. Dumurgier J, Artaud F, Touraine C, et al. Gait Speed and Decline in Gait Speed as Predictors of Incident Dementia. The Journals of Gerontology, Series A. 2017;72(5):655–661. DOI 10.1093/gerona/glw110
Source note: SRC-8c02124df60f Dumurgier J 2017
36. Sotirakis C, et al. Identification of motor progression in Parkinson’s disease using wearable sensors and machine learning. npj Parkinson's Disease. 2023;9(1):142. DOI 10.1038/s41531-023-00581-2
Source note: SRC-626258f8a85b Sotirakis C 2023
37. Wang F, et al. Predicting the onset of freezing of gait in Parkinson’s disease. BMC Neurology. 2022;22(1):213. DOI 10.1186/s12883-022-02713-2
Source note: SRC-2f6ce7a241dc Wang F 2022
38. Brzenczek C, et al. Integrating digital gait data with metabolomics and clinical data to predict outcomes in Parkinson’s disease. npj Digital Medicine. 2024;7(1):235. DOI 10.1038/s41746-024-01236-z
Source note: SRC-3272793948a0 Brzenczek C 2024
39. Cebi I, et al. Clinical and Kinematic Correlates of Favorable Gait Outcomes From Subthalamic Stimulation. Frontiers in Neurology. 2020;11:212. DOI 10.3389/fneur.2020.00212
Source note: SRC-97db0dd4f14a Cebi I 2020
40. Serrao M, et al. Prediction of Responsiveness of Gait Variables to Rehabilitation Training in Parkinson’s Disease. Frontiers in Neurology. 2019;10:826. DOI 10.3389/fneur.2019.00826
Source note: SRC-1aecd550e6f3 Serrao M 2019