In this report
Audited and Updated
Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.
Two supplied versions of this report are retained; neither has been selected as authoritative.
Read the other supplied version
PD-B01
The authors reported 67/72 completion for the least challenging stance. Distinguish this from the 62 usable PD recordings in Table 3 and retain the unresolved denominator inconsistency.
Type: source denominator precision. Audit disposition: supported source denominator qualification.
Remaining limit: Completion and usable recordings can differ; do not simply replace67with62. Unreconciled loss-of-balance count remains.
Editorial record
- Audit status: supported source denominator qualification. Completion and usable recordings can differ; do not simply replace67with62. Unreconciled loss-of-balance count remains.
Editorial nomenclature update — 4 October 2026 at 11:48:46 am (Australia/Brisbane): authored condition labels and report wording use Parkinson’s disease. Published article titles, exact quotations, recorded searches, identifiers and routes are preserved. This is a terminology edit, not a scientific correction.
Executive assessment
Standing-balance assessment in Parkinson’s disease helps identify impairment, direct further examination and track performance under a defined protocol. Selected findings also carry information about later falls, cognitive impairment, disability and mortality. The evidence is task- and outcome-specific: quiet sway, sensory challenge, controlled leaning, reactive recovery and multidomain scores are not interchangeable predictors.
Clinical balance and future falls: controlled leaning, challenging stance and unexpected retropulsion provide prospective signals. Jacobs found that Mini-BESTest and Brief-BESTest added to baseline recurrent-fall history, while individual pull and push-and-release tests did not. That useful incremental signal came from only 43 complete cases and 17 future recurrent fallers, with informative attrition and no external validation. Battery accuracy varies by cohort, scoring version and follow-up; a normal pull test cannot reliably exclude future falls. [1, 2, 3, 4, 5]
Instrumented standing and future falls: prospective sway and perturbation studies include useful positive and negative findings. Standing-specific added value over an appropriate clinical baseline remains uncertain. Moraca's reactive CoP measure was associated with time to falling despite weak standalone discrimination; Matinolli's unadjusted sway differences did not survive multivariable selection. Final journal methods clarify selected cohorts, internal validation, missingness and protocol dependence; those details qualify rather than erase the candidate standing signals. [6, 7, 8, 9]
Assessment selection and change: use a clinical profile and additional tasks that answer a specific question. The assessment chapter compares actual protocols, measurement properties and appropriate uses. Trial duration, foot position, medication, dyskinesia, processing and aggregation affect interpretation. Short-interval reliability does not establish between-day error, and two-week phone aggregates do not validate a single trial. Protocol-specific absolute-error estimates cannot be exported to different tasks or devices. [16, 17, 18, 19]
For rehabtools, the strongest immediate direction is transparent, purpose-specific assessment with reproducible protocols, visible failed trials and qualified change interpretation. Numerical prognosis needs separate validation of the actual measurement and prediction pipeline in the intended population, at a named horizon.
Scope and appraisal
Assessment questions and balance domains
This report examines standing-balance assessment in established Parkinson’s disease for three uses: describing present impairment, measuring change, and forecasting later outcomes. It considers clinical tasks, force-platform and pressure-mat measurements, body-worn sensors and selected low-cost systems. The practical question is which assessment and protocol can support a defensible claim in rehabilitation, including potential rehabtools outputs.
Quiet stance, challenged stance, controlled leaning and reactive recovery assess different demands on postural control. Removing vision or changing the support surface probes performance under altered sensory conditions. Narrowing the base of support increases task difficulty. Reaching and deliberate weight shifting examine voluntary control near the support boundaries; perturbations examine recovery. Multidomain scales add transfers, stepping and often gait. Results for a total scale or an axial clinical composite are therefore not results for quiet standing alone.
The report examines these domains separately before comparing assessment choices. It retains relevant null findings and distinguishes studies of future outcomes from historical faller classification. Gait-only findings do not establish standing-balance performance, although combined models help identify what has and has not been demonstrated for standing inputs.
How to interpret the evidence
Measurement validity concerns the quantity actually measured. Repeatability describes variation under stable conditions; responsiveness concerns detection of change. Minimal detectable change concerns measurement error, whereas minimal important change concerns the significance of change to the patient or clinician. These properties do not themselves establish a forecast of future falls or progression.
A prognostic factor is associated with a later outcome. A prediction model specifies an outcome, horizon, inputs and calculation; discrimination describes ranking, and calibration describes agreement between predicted and observed risks. Apparent performance is estimated in development data. Internal validation can estimate optimism, while external validation tests a fixed model in genuinely new participants. Re-fitting in a new cohort can replicate an association without validating the original equation or absolute risks. [20]
For added value, the useful comparison is the same clinical baseline with and without the standing measurements. A combined gait-and-balance model does not isolate the contribution of balance. Participant-level separation and training-only feature selection matter when recordings yield many correlated variables. The number of people and outcome events, rather than sensor samples, constrains model complexity. [21]
Study estimates are interpreted at their original horizon and for their stated population and endpoint. Any fall, recurrent falls, first falls and injurious falls are different targets. Likewise, current severity, subsequent change, treatment-associated improvement and survival cannot be collapsed into a single notion of predictive accuracy. Clinical utility ultimately requires evidence that using the result improves decisions or outcomes.
Search and source appraisal
Original studies were prioritised when assessment preceded a later outcome. Reliability, validity and responsiveness studies were examined separately to inform test and protocol selection. Full primary text, tables and available supplements were preferred. Article-level access is recorded in the evidence tables; an abstract-only estimate does not support unverified details of its protocol or analysis. Companion measurement studies and same-cohort publications are distinguished from independent prognostic replication.
Appraisal considered population and disease stage, participant and event counts, task and medication state, hardware and processing, outcome ascertainment and horizon, effect estimates and uncertainty, adjustment and comparator, validation, missingness, attrition and applicability. Prediction-model principles informed the assessment, but no completed formal PROBAST or PROBAST+AI rating or GRADE assessment is claimed. The synthesis and source checking were AI assisted rather than a duplicated independent human review. [21]
The expanded search included outcome-specific paginated searches and backward and forward citation tracing; provider duplicates and incomplete coverage still prevent an exhaustive-search claim. Searches were purposive and not every result page was screened. Failure to retrieve a study is not evidence that it does not exist. Conclusions therefore concern the evidence examined. Differences in population, testing protocol, outcome definition and analytic target preclude treating the reported estimates as a common effect or ranking tests by AUC across unrelated cohorts.
Source coverage and remaining gaps
Final journal articles were examined for Schlenstedt 2016, Almeida 2016, Jacobs 2016, Gervasoni 2015, Matinolli 2011, Sebastiá-Amat 2026 and Beretta 2018. Their Methods, Results, tables and figures support the study-specific appraisal below. Matinolli's investigator thesis supplies additional protocol detail where the final article refers to earlier methods. The final Sebastiá-Amat article governs its reported cohort and estimates; earlier thesis findings are not silently substituted for different final-journal results. The reproduced Bloem article and Schlenstedt's companion measurement manuscript remain separately identified thesis sources. [3, 22, 5, 7, 9, 23, 24, 25, 26, 27]
Fiems 2020 remains limited to its abstract and public publisher excerpts; a complete original article was not recovered. For Paul 2012, publisher Methods, Results, Tables 1–3, Discussion and the linked procedural supplement were inspected in scoped extracts; the original PDF was not recovered. Selected measurement and contextual outcome studies remain abstract-only or partly accessible, as specified beside their findings. Supplementary-material retrieval was not exhaustive. Source access is recorded separately from study quality: obtaining the full report resolves reporting questions without conferring external validation or clinical utility. [28, 29]
Clinical balance and future falls
What the clinical tasks contribute
The clinical literature supports a qualified conclusion: difficult stance, controlled leaning and reactive balance can carry information about subsequent falls, but no single bedside standing test emerges as a dependable, externally calibrated risk calculator. The strongest-looking results often belong to multidomain batteries, and their performance cannot be assigned to quiet stance. Functional reach tests anticipatory control toward the edge of the base of support; a pull or push-and-release test challenges recovery from perturbation; single-leg stance also includes the transition onto one leg. These are related but distinct tasks. Berg combines stance with transfers, reaching, turning and stepping; BESTest and its shortened forms also contain dynamic mobility components. [30, 31, 32]
Stance and controlled leaning
Paul and colleagues provide unusually relevant direct evidence. Among 205 community-dwelling participants tested when medication was working optimally, 120 fell during six months of diaries and telephone follow-up. Shorter eyes-closed near-tandem stance and poorer narrow-base standing were associated with falls in univariable analyses. However, the selected multivariable representation of standing control was coordinated stability: tracing a path with a waist-mounted swaymeter while leaning without stepping. Coordinated stability and an impaired pull test remained associated alongside freezing and orientation; apparent model AUC was 0.73 (95% CI 0.66–0.80). [1]
This supports a standing-control signal, rather than a clinically transportable tandem-stance rule. The model selected representatives from multiple domains and deliberately omitted previous falls to examine potentially remediable impairments. It therefore does not answer whether stance improves prediction beyond fall history. Functional reach had only a weak univariable association, and no validated stance cutoff or calibrated absolute-risk estimate resulted. Short duration caps also compress performance among higher-functioning participants. [1]
Schlenstedt et al. recruited 85 people with idiopathic PD, H&Y stages 1–4, tested ON medication at a German university hospital. The final analysis comprised 66 participants: 11 were excluded after receiving DBS and eight after entering rehabilitation intended to reduce falls during follow-up. Thirty-three of the 66 experienced at least one fall over six months. This was an any-fall study, not a cohort restricted to people without previous falls: 30 participants had fallen in the preceding six months, including 23 of the 33 subsequent fallers. Monthly telephone interviews supplied the outcome. Diaries were also distributed, but some were lost or incompletely filled in, so the authors relied on telephone reports. [3]
The protocol used the German FAB (/40), Mini-BESTest (/28) and Berg (/56), in that fixed order, followed by UPDRS. Two trained examiners administered the assessments, rest was offered and duplicated tasks were performed once but scored under each scale's rules. This makes the head-to-head comparison dependent on shared performances and a fixed test order. Testing excluded clinically identified cognitive impairment, other conditions affecting stance or gait, DBS at baseline and medication changes within the preceding four weeks; assistive-device use was allowed. Medication changes during follow-up were not documented. [3]
The three complete scales had similar, modest apparent discrimination for any fall: FAB AUC 0.68 (95% CI 0.55–0.81), Mini-BESTest 0.65 (0.52–0.78) and Berg 0.69 (0.56–0.82). Data-derived cutoffs were ≤27/40, ≤19/28 and ≤52/56, respectively, with sensitivity/specificity of 0.67/0.58, 0.52/0.70 and 0.64/0.67. The cutoff rule minimized squared distance from the ideal ROC corner; these are development-sample operating points. When the outcome was changed to recurrent falls, the corresponding AUCs rose to 0.72, 0.70 and 0.74. This within-study contrast shows how the endpoint changes apparent performance. [3]
The six-item AUC of 0.84 (0.75–0.94) came from a substantial selection process. Each of the 38 scale items underwent univariable screening; candidate selection required p<0.05 and a median-split odds ratio >2 or <0.5, followed by removal of highly correlated candidates. The resulting model included Mini-BESTest rise to toes, one-leg stance and backward compensatory stepping, plus Berg turning 360°, alternate foot placement and tandem stance. There were only 33 fall events for six retained predictors, before considering the initial screening. A separate forward-selection model retained only Berg tandem stance, item 13: AUC 0.71 (0.59–0.84), sensitivity 0.82 and specificity 0.61 at ≤3/4. This is an ordinal Berg-item threshold; it does not validate a new timed tandem-stance rule. Neither model was resampled or externally validated, and coefficients sufficient to reproduce individual predictions were not supplied. No incremental comparison against a fall-history model was reported. The mixed six-item result is therefore an exploratory combination, while the tandem result is a specific but unvalidated standing-test signal. [3]
Older single-leg literature is especially easy to overstate. Jacobs et al. reported 85% classification using one-leg stance and gait in 67 participants, but the outcome was falls recalled from the preceding year. Three trials were performed, with the third selected because it related most strongly to the outcomes. This is historical classification with data-dependent analysis, not future-fall prognosis. [32]
Bloem et al.'s early cohort included 59 ambulant people with PD and 55 controls, tested about one hour after usual medication, with event forms and fortnightly calls over six months. Table 3 pools 17 recurrent fallers and 97 people with zero or one fall. Tandem stance was assessed separately with eyes open and closed: sensitivity/specificity were 52.9%/83.5% and 88.2%/46.4%, respectively. The ordinary Romberg test detected only three recurrent fallers. Its apparent 100% specificity therefore accompanied 17.6% sensitivity. [33, 26]
The history–severity–Romberg combination's 65% sensitivity and 98% specificity were reported in this pooled-analysis context. The authors also state that patient-only application gave identical results, but provide no separate PD-only accuracy table. This is not independent validation. With only three Romberg-positive cases, its added value is especially fragile; the authors themselves considered it limited. The first unexpected shoulder pull detected five of 17 recurrent fallers; the second, warned pull detected one. The six-pull average had AUC 0.62 (SE 0.07). These findings show that changing vision or warning changes test behaviour without establishing a dependable individual-risk rule. [33, 26]
Functional reach and Berg
Almeida et al. analysed 225 consecutively recruited ambulatory clinic patients after four of 229 baseline participants died during follow-up. Eighty-four participants experienced at least two falls over 12 months. Falls were recorded on calendars and checked with monthly calls involving participants, families or caregivers. Mean age was 70.7 years. Testing was ON medication, about one to two hours after intake, with one physical therapist administering the self-report and performance measures in a fixed sequence during a roughly 60-minute visit. The cohort excluded dementia/cognitive impairment, severe visual or vestibular problems and other conditions affecting locomotion or balance; walking aids were permitted, but another person's assistance was not. [22]
Functional reach had AUC 0.74 (95% CI 0.67–0.79). At ≤17 cm, sensitivity was 0.56 (0.45–0.67) and specificity 0.82 (0.75–0.88). Berg had AUC 0.79 (0.73–0.84); at ≤49/56, sensitivity and specificity were both 0.74, with CIs 0.63–0.83 and 0.66–0.81, respectively. The paper defines reach as maximum forward reach beyond arm's length while maintaining a stable base of support, but does not report precise foot spacing, arm choice, trial count or aggregation. The numerical threshold should not be detached from that incomplete protocol description. All cutoffs were selected within this cohort using the Youden index. They are not externally established screening thresholds. [22]
The combination result is more specific than simply saying that two tests predict better. The authors fitted models using dichotomized tests and compared all 15 pairs and 20 triples by Akaike information criterion (AIC). The favoured pair required both Berg ≤49 and Falls Efficacy Scale–International (FES-I) >29. Its sensitivity was 0.65 (0.55–0.75) and specificity 0.83 (0.76–0.88), compared with 0.74/0.74 for Berg alone. The published AIC fell from 105 to 98, while the positive posttest probability rose from 63% to 70% at this cohort's 37.3% recurrence prevalence; a negative result left about 20% risk, versus 17% with Berg alone. Thus the pair identified a smaller, more specific high-risk group while missing more eventual recurrent fallers. Three-test versions achieved AIC 97–98 but further reduced sensitivity. The paper supplies no combined-model AUC, calibration, optimism correction, external validation or comparison against a model containing previous falls. It supports exploratory complementarity of fear-of-falling and performance information, not a validated improvement over clinical history. [22]
Two reporting cautions affect precision. The authors describe noninferiority testing but report no prespecified noninferiority margin; nonsignificant AUC differences do not by themselves establish equivalence. In addition, some table details are internally inconsistent: the BBS–TUG negative posttest probability is printed as 13%, although the accompanying likelihood ratio and cohort prevalence imply approximately 23%. The combination-table count labelled as the number testing positive appears to count true positives. These entries should not be silently corrected or used as exact operational rules. The principal BBS/FES-I sensitivity and specificity comparison above remains directly supported by the table. [22]
A related Brazilian study illustrates this distinction. It followed 130 participants who had not fallen in the preceding year, with 40 subsequent fallers and 21 recurrent fallers over 12 months. Testing occurred ON medication and falls were captured through diaries plus monthly calls. Functional reach and Berg were associated with recurrent falls univariately, but reach did not retain an independent association in the full recurrent-fall model: OR 0.95 per cm (95% CI 0.86–1.06; p=0.37). Pull-test performance was not associated with either endpoint in univariable analysis. Disability, rather than a standing measure, dominated the selected models. [34]
This is a meaningful negative result, although not proof that reach has no prognostic information. Only 21 recurrent-fall events supported a six-predictor full model, and domain-based selection can discard correlated measures. The defensible conclusion is that simple standing measures have not demonstrated a consistent, independently validated advantage for predicting an initial fall trajectory. Different cohorts and outcomes explain some of the disagreement; a single universal cutoff obscures it. [22, 34]
Reactive balance
Lindholm et al. directly compared unexpected Nutt retropulsion with the expected UPDRS item-30 shoulder pull. After each assessment, falls and near falls were recorded for six months with diaries and monthly calls. At baseline, 146 people were assessed; 58 contributed a second assessment approximately 3.5 years later. An abnormal Nutt test at a score of at least 1 had sensitivity 0.47 and specificity 0.85 for subsequent falls initially, and 0.48/0.78 later. Initial LR+ was 3.09 (95% CI 1.77–5.39), but LR− was 0.63 (0.47–0.83). A negative test therefore lowered risk only modestly. The expected pull test performed less convincingly and did not discriminate significantly at the later assessment. [2]
The second time point is repeat evaluation within a depleted cohort, not external validation. Initially, people unable to stand unsupported and those over 80 were excluded; this limits generalisation to the frailest population. Tests were scheduled at each participant's usual best time, although some reported being OFF. Protocol differences matter: an unexpected first pull, a warned pull and backward leaning followed by release are not equivalent perturbations. [2]
An earlier analysis from the same Swedish programme found retropulsion associated with future falls/near falls after adjustment (OR 2.81, 95% CI 1.07–7.37); the falls-only estimate was 3.5 (1.3–9.4). Those associations coexist with the limited stand-alone sensitivity above. Its stepwise model screened many candidates, and a nonsignificant Hosmer–Lemeshow test is not evidence of accurate calibration across risk levels. The reports should not be counted as independent replication. [35]
Jacobs et al. provide a direct test of incremental prediction beyond baseline recurrent-faller status. Of 80 people with H&Y stages I–IV, 43 had complete baseline, six-month and 12-month assessments; eight were recurrent fallers in the preceding six months and 17 over the following 12 months. Subsequent falls were collected by six-month recall at follow-up visits, without a formal diary. Attrition was informative: the 37 excluded participants had worse motor and balance scores and more baseline recurrent falling, 17/37 versus 8/43. The complete-case results therefore preferentially represent less impaired survivors of follow-up. [5]
Testing was ON medication, approximately one to 1.5 hours after intake. A trained physical-therapy student administered the full BESTest; Mini scores were recorded simultaneously and Brief scores extracted from the same performances. Here Mini-BESTest used 14 items with the worse side of bilateral tasks, giving 28 points, and Brief-BESTest used eight scored components, giving 24 points. The backward push-and-release test used the BESTest's 0–3 scale, lower scores indicating worse performance. The participant stood with feet shoulder-width apart and arms at the sides, leaned backward about the ankles with straight hips until shoulders and hips were just behind the heels, then recovered when the examiner released support at the scapulae. This differs from the original five-point push-and-release score. The MDS-UPDRS pull test used 0–4, higher scores worse, with feet comfortably apart, arms at the sides, explanation and a mild practice before the brisk shoulder pull. [5]
History alone had AUC 0.735 (95% CI 0.57–0.90). Adding Mini-BESTest raised apparent AUC to 0.842 (0.71–0.97), with a likelihood-ratio model improvement of χ²=4.733, p=0.030; adding Brief-BESTest gave 0.838 (0.70–0.97), χ²=4.209, p=0.040. These p values concern the additional model term, not a reported formal test of the AUC difference. In contrast, adding backward push-and-release or pull scores did not significantly improve the history model, p=0.10 and 0.32, respectively; neither did MDS-UPDRS motor score, p=0.089. The stand-alone push-and-release and pull AUCs were 0.682 and 0.690, with sensitivity only 41.2% and 35.3% at the reported <2 and >2 thresholds. A normal response cannot adequately rule out future recurrence. [5]
Exploratory item selection yielded history plus incline stance and pivot turning, AUC 0.898 (0.79–1.00), or history plus the Brief-BESTest functional-reach item, AUC 0.851 (0.73–0.97). These are ordinal items from an administered battery, not validation of a centimetre-based FRT threshold or a new stance protocol. Selection among items with only 17 events creates considerable optimism; no resampling, external validation, calibration or deployable equation was reported. The paper also directs readers to Duncan 2013 for dropout details and explicitly builds on those earlier analyses, supporting linkage to that research cohort. However, its 43 analysed cases differ from the 40 in Duncan's 12-month analysis, without a reconciliation in this article. It should not be counted as a clearly independent replication. The defensible advance is evidence that a multidomain balance examination may add to a simple history indicator within this small, selected sample. [5, 31]
The frequently cited OFF/ON comparison by Valkovic et al. must be kept separate. Its 82 participants were grouped using falls in the previous six months. The push-and-release test classified those historical groups better than the pull test during ON medication, but that does not demonstrate prospective superiority. Likewise, the DBS-clinic paper titled “A Prospective Evaluation” used previous-year falls as its reference standard. Prospective enrolment does not make a retrospective outcome prognostic. [36, 37]
Multidomain clinical batteries
Duncan et al.'s 2012 pilot began with 80 participants tested ON medication. Only 51 supplied six-month outcomes (14 recurrent fallers), and 40 supplied 12-month outcomes (13 recurrent fallers). Six-month AUCs were high, reaching 0.89 for BESTest and 0.87 for Mini-BESTest; at 12 months they fell to 0.68 and 0.77. Ascertainment used six-month recall at follow-up visits, rather than daily diaries. Seven of the 11 participants lost between six and 12 months had been recurrent fallers at six months. Attrition and recall complicate the apparent horizon effect. [30]
The 2013 Brief-BESTest paper reused this cohort. Its six-month AUC was 0.88 (95% CI 0.74–0.94), with sensitivity 0.71 and specificity 0.87 at ≤11/24. It is an additional analysis, not independent confirmation. The authors' proposal to test every six months is reasonable as a hypothesis, but comparison of predictive horizons does not establish that six-month screening reduces falls. Critically, these Duncan studies scored Mini-BESTest out of 32. Their ≤20/32 threshold is not interchangeable with a 28-point score. [31]
Mak and Auyeung provide an independent prospective Mini-BESTest cohort with better retention: 110 analysed participants and 24 recurrent fallers over six months. The 28-point total had AUC 0.75; a cutoff around 19 gave sensitivity 0.79 and specificity 0.67. Mini-BESTest remained associated after adjustment (OR 0.75 per point; p=0.014), but the final model included ten predictors for 24 events. Its apparent AUC of 0.895 is therefore vulnerable to overfitting; the increment from the preceding model was nonsignificant (0.868 to 0.895; p=0.250). Bootstrapping the cutoff did not externally validate the whole model. [4]
The same study found group differences in anticipatory, reactive and sensory-orientation components, but not the gait component. These comparisons were not independent, adjusted domain-specific prediction models. Nor does improvement in a total score show which component supplied its prognostic information. Later Mini-BESTest structural work also cautions against treating the four subscores as independent constructs: the 709-person study was cross-sectional, with strong inter-factor relationships and limited subscore reliability. It strengthens a measurement warning, not the prospective evidence. [4, 38]
Reliability and ceiling effects
The companion Schlenstedt measurement study is available as a primary manuscript within the author's thesis. It tested 85 people with H&Y stages 1–4 ON regular medication, using the German FAB (/40), Mini-BESTest (/28), then Berg (/56). Duplicated items were performed once and scored by each scale's rules. This makes administration efficient, but the comparison is not between independently performed examinations. Test–retest assessments in 17 participants were 3 ± 1 days apart at the same time of day with matched medication use. FAB, Mini-BESTest and Berg ICCs were 0.99, 0.98 and 0.95; these impressive relative reliabilities do not quantify an individual's measurement error or prognostic calibration. [27]
Berg reached its maximum in 15/85 participants, including four reporting previous falls; only two reached the FAB maximum and one the Mini-BESTest maximum. Thus a high Berg score can conceal clinically relevant difficulty. However, less ceiling compression does not guarantee superior future-fall discrimination: the 2016 prognosis article gives similarly modest AUCs for all three batteries. The measurement manuscript and prognosis article report closely matching samples and protocols, but neither the matching sample size nor authorship alone establishes the exact participant overlap. Their measurement and prediction findings should remain distinct. [27, 3]
Standing to dress
A 2023 study tested whether having stopped putting on trousers or a skirt while standing predicted future falling. Among 264 people with early PD, this self-reported “Pants-sign” was associated univariately (OR 2.41, 95% CI 1.31–4.41), but did not add to prior recurrent falls and motor severity. It was not associated with subsequent falls among the 189 participants without a preceding-year fall. Outcomes relied on annual recall, and this was a behavioural question rather than a measured stance task. Nevertheless, it is a useful contemporary negative: an everyday balance difficulty may reflect established disability or adaptive caution without providing independent early warning. [39]
Injurious falls
Castro et al.'s later injury analysis provides an important check on treating every fall outcome as interchangeable. The Salvador cohort included 225 participants followed for 12 months using diaries and monthly calls after ON-medication assessment. Of 1,290 reported falls, 805 had injury information, including 107 injurious events; 485 were excluded, and some circumstances were imputed. This is a reanalysis of an existing recurrent-falls cohort, not independent replication. [40]
Among recorded falls, better Berg and reach performance was associated univariately with injury rather than no injury: OR 1.11 per Berg point (95% CI 1.06–1.15) and 1.11 per reach centimetre (1.07–1.15). These are conditional event-level comparisons. They do not mean that better standing balance increases a person's overall probability of injurious falling. When the reference was zero falls, the univariable associations reversed; adjusted Berg was uninformative, OR 0.98 per point (0.89–1.08; p=0.711), and reach was not retained for adjustment. [40]
The multinomial analysis had 919 records: 805 falls plus 114 non-fallers. Frequent fallers therefore contributed repeated records. Although the paper states that its logistic model adjusted for the individual, it gives insufficient implementation detail to audit the handling of dependence fully. Moreover, fall location and perceived cause are known at the event, not baseline. The results illuminate mechanisms and exposure, but do not supply a validated bedside injury-risk calculator. Missing consequences for 38% of reported falls further limit interpretation. No incremental standing-test discrimination, calibration or external validation was established. [40]
Clinical interpretation
The practical value of a standing assessment may be greater for identifying a specific impairment than for assigning an individual fall probability. A reactive-balance deficit can justify closer examination even when a model adds little AUC. Conversely, a normal pull test or a strong brief stance performance cannot rule out falls driven by freezing, cognition, orthostasis, environmental exposure or transitions. Prognostic association does not establish that changing the measured impairment will change the outcome.
Table 1 Principal clinical balance evidence for future falls
Intervals are 95% confidence intervals unless labelled otherwise. Access describes the material examined, not study quality. Apparent performance uses development data; internal validation is not external validation.
| Study cohort and outcome | Assessment and main estimate | Validation limitations and access |
|---|---|---|
| Paul 2014 [1] 205 participants; 120 fallers Any fall over 6 months | Controlled leaning and pull test retained in a mixed impairment model Apparent AUC 0.73 (0.66–0.80) | Development only; prior falls deliberately excluded No calibrated stance-only rule Access: full text and tables |
| Schlenstedt 2016 [3] 85 enrolled; 66 analysed after 11 DBS and 8 rehabilitation exclusions 33 any-fall events; 6 months | FAB AUC 0.68 (0.55–0.81); Mini-BESTest 0.65 (0.52–0.78); Berg 0.69 (0.56–0.82) Selected six-item AUC 0.84; Berg tandem item 0.71 | 38-item screening; 6 retained predictors for 33 events; no resampling or history-adjusted increment Tandem threshold is ordinal ≤3/4, not seconds Access: final journal Methods, tables and figures |
| Almeida 2016 [22] 229 enrolled; 4 deaths; 225 analysed 84 recurrent fallers; 12 months | Reach AUC 0.74 (0.67–0.79); Berg 0.79 (0.73–0.84) Berg ≤49: sensitivity/specificity 0.74/0.74 Berg ≤49 AND FES-I >29: 0.65/0.83 | Same-sample cutoffs and 35 test combinations; pair trades sensitivity for specificity No history-baseline comparison, calibration or external validation Access: final journal Methods, tables and figures |
| Almeida 2015 [34] 130 without prior-year falls 40 any-fall and 21 recurrent-fall cases over 12 months | Adjusted reach OR 0.95 per cm (0.86–1.06) for recurrence Pull test not associated univariably | Few events for six-predictor model; no externally tested standing rule Access: full text and tables |
| Lindholm 2015 and 2021 [35, 2] 146 initially; 58 reassessed 6-month falls after each assessment | Unexpected Nutt test sensitivity 0.47/0.48; specificity 0.85/0.78 Initial LR− 0.63 (0.47–0.83) | Overlapping Swedish cohort; later assessment is not external validation Normal result does not exclude risk Access: full text and tables |
| Duncan 2012 and 2013 [30, 31] 80 baseline; 51/40 retained 14/13 recurrent fallers at 6/12 months | Mini-BESTest AUC 0.87 at 6 months; 0.77 at 12 months Brief-BESTest 6-month AUC 0.88 (0.74–0.94) | Same cohort; substantial attrition; recalled outcomes Mini-BESTest scored out of 32 Access: full text and tables |
| Mak and Auyeung 2013 [4] 110 participants; 24 recurrent fallers 6 months | 28-point Mini-BESTest AUC 0.75 Adjusted OR 0.75 per point Model AUC 0.868 to 0.895; p=0.250 | Ten predictors for 24 events; no external validation Cutoff derivation does not validate whole model Access: full text and tables |
| Jacobs 2016 [5] 43 complete of 80; 17 future recurrent fallers 12 months; six-month recall | History AUC 0.735; plus Mini-BESTest 0.842 (0.71–0.97), added-term p = 0.030; plus Brief 0.838 (0.70–0.97), p = 0.040 Pull and push-and-release did not add significantly | Informative attrition; linked Duncan programme; 28-point Mini versus earlier 32-point implementation No resampling/external validation; p values test terms, not AUC differences Access: final journal Methods, tables and Figure 1 |
| Bloem 2001 [33, 26] 59 PD and 55 controls; 15 and 2 recurrent fallers; 6 months | Pooled tandem sensitivity/specificity: eyes open 53%/84%; closed 88%/46% Romberg 18%/100%; only 3 positive recurrent cases History–severity–Romberg model 65%/98% | Pooled accuracy table; authors state identical patient-only results without a separate PD table No independent validation Access: full primary chapter reproduced in author thesis |
| Jansen 2023 [39] 264 early PD; 189 without prior-year falls Annual follow-up | Pants-sign unadjusted OR 2.41 (1.31–4.41) No increment over prior recurrent falls and severity | Behavioural self-report, not measured stance; null in prior nonfallers; annual recall Access: full text |
| Castro 2023 [40] 225 people; 12 months 805 of 1,290 falls had injury information; 107 injurious events | Injury versus no injury: univariable Berg OR 1.11 per point (1.06–1.15) Injury versus zero falls: adjusted Berg OR 0.98 (0.89–1.08) | Event-level reanalysis, not independent cohort validation Repeated records and missing injury data; no calibrated baseline risk model Access: full text and tables |
Berg and BESTest variants include functions beyond quiet stance. Same-cohort publications are grouped; companion measurement studies are discussed in the text. Patient-level fall occurrence, recurrence and event-level injury comparisons are distinct targets. Thresholds retain the source scoring version.
Instrumented standing balance and future falls
What the standing signal represents
Instrumented standing assessments can quantify an impairment that a clinician may not see, and several prospective studies associate particular sway measures with later falls. The stronger claim, that a brief standing test supplies a dependable individual prognosis beyond fall history and basic clinical assessment, remains insufficiently demonstrated. Positive findings are metric-, protocol- and population-specific. The evidence includes small derivation cohorts, mixed standing–gait models, retrospective fall classifications and meaningful negative results. A force-plate or wearable measurement should therefore be presented as a component of balance assessment, rather than an independently validated fall-probability calculator.
Start by identifying what was actually predicted
“Predictor” is often used statistically, even when the outcome occurred before measurement. Matinolli's 120-person study associated sway area with recent falls (adjusted OR 1.25, 95% CI 1.02–1.54), but it did not predict newly occurring falls. Johnson's 48-person PD cohort illustrates the distinction particularly clearly: static posturography distinguished PD fallers from healthy controls, but not PD fallers from PD non-fallers; voluntary leaning measures performed better. Neither finding establishes future prediction. [41] [42]
Likewise, Freeman's instrumented modified Clinical Test of Sensory Integration for Balance (i-mCTSIB) involved 26 participants, only five of whom reported at least two falls in the previous six months. Its composite differed between those groups (p=.04), whereas the Sensory Organization Test (SOT) composite did not (p=.31). The authors explicitly called for prospective validation. Rivera's 2026 SOT analysis is also retrospective: 40 people reported falls over the preceding two months. An association with eyes-closed, fixed-platform condition 2 is not a validated future fall forecast. These studies can inform measurement choice and hypotheses; they cannot substantiate a promised prospective risk estimate. [43] [44]
The 2025 wearable review is useful for finding studies, but combines prospective and retrospective designs, and includes Tsai's pressure-mat study among wearable studies. Claims of consistent sensor superiority therefore require checking the original devices, outcomes and analyses. [45] [46]
Direct prospective evidence for quiet standing
The studies do not measure one interchangeable variable. Kerr’s physiological profile uses body sway displacement; Matinolli’s device derives displacement from a sacral rod/inclinometer; Tsai and Sebastiá-Amat use pressure platforms; APDM studies derive trunk-acceleration features. Frequency, displacement, CoP speed and a proprietary composite score require separate validation and cannot share numerical cutoffs. [6] [47] [46] [23] [9] [25]
For any fall, only physical activity (LAPAQ) retained significance: AUC .718 (95% CI .564–.872), sensitivity 87.5% and specificity 35.7%. For recurrent falling, velocity factor 2 had OR 2.37 (1.01–5.58; p=.0474); activity was retained under the p<.1 selection rule but did not reach conventional significance (p=.0701). The mixed model yielded reported AUC .844 (.716–.972), accuracy 76.9%, sensitivity 77.8% and specificity 76.2%. Crucially, the complete Methods explicitly describe leave-one-out cross-validation. However, the Results do not distinguish apparent from cross-validated estimates, and nesting of PCA and predictor selection within folds is not described. [7]
Important ambiguities survive full-text access. Table 3 labels the recurrent comparison as at least two falls versus no falls, without explaining single-faller handling; the final analysed denominator and contingency table are absent. The fall definition also excludes dizziness, fainting and loss of consciousness. The study therefore supports an internally evaluated candidate signal in selected previous fallers, with uncertain model reconstruction. It does not establish externally calibrated risk or improvement over a prespecified complete clinical baseline. [7]
Tsai et al. used a Tekscan pressure mat, not a body-worn IMU: three 30-second eyes-open trials, barefoot with feet together, in both medication states. Of 95 people, 24 developed falls during mean follow-up of 12.4 months. ON-state path length and velocity were associated with falls individually, but neither remained significant in the multivariable Cox model. Reported AUCs were .73 (95% CI .62–.84) for length and .72 (.61–.84) for velocity; a combination with levodopa dose and Tinetti balance/gait reached .90 (.83–.98). [46]
That apparent improvement warrants considerable restraint. All 24 future fallers already had a fall history, and the ROC combination omitted that strong predictor. Comparisons against individual components do not establish added value beyond an adequate clinical baseline model. Path length and average velocity over a fixed duration are also intrinsically related, an additional reason to examine redundancy. No external validation or calibration was reported. Table denominators and some reported estimates are internally inconsistent, including a sensitivity percentage not matching the stated 24 fallers. This is promising development evidence for future falling among previous fallers, not evidence for predicting first falls or a ready-to-use .90-AUC standing test. [46]
Sebastiá-Amat et al. (online 2025; 2026 issue) followed an ambulant H&Y 1–3 cohort for one year with monthly telephone recall supported by relatives. The final journal report recruited 55 and analysed 48: two died, one could not be contacted, one reported another neurological diagnosis and three were excluded as outliers without a specified rule. There were 17 recurrent fallers (≥2) and 31 people with zero/one fall. Testing used a FreeMed pressure platform at 100 Hz, ON medication 45–90 minutes after the morning dose, shoulder-width feet, arms at sides and an eye-level target 1.5 m ahead. Three 30-second trials per eye condition were performed in randomized order, with one-minute rests and the three-trial mean used. [23]
Eyes-open CoP mean speed had AUC .81 (95% CI .70–.93) and a Youden-selected 18.1-mm/s cutoff. Its adjusted OR was 1.81 (1.11–2.96) per mm/s in a mixed model also containing fall history, H&Y and a nonsignificant TUG term. This suggests a signal after clinical adjustment, but does not demonstrate incremental predictive performance: no clinical-only versus plus-speed comparison, calibration or internal/external validation was reported. Screening by group significance and effect size, followed by two backward-selection stages, is prone to optimism with only 17 events. The sample-size calculation addressed a large two-group mean difference, not this prognostic modelling task. The final paper provides no sensitivity or specificity at the selected cutoff. [23]
The final article and earlier thesis differ in several consequential details. It explicitly reports 88.2% overall classification, whereas the earlier thesis reported 91.7%; no contingency table or model-specific denominator explains the change, and 88.2% is not reproducible as an integer fraction of the stated 48 participants. The final recruitment/exclusion account and sensor-only model also differ from the thesis and now take precedence. Final-table reporting inconsistencies remain, including the H&Y OR/interval/SE/p-value combination and an RMS axis-label reversal. These issues warrant caution about numerical implementation, rather than substitution of thesis operating characteristics. The study supports further validation of eyes-open speed, not a deployable 18.1-mm/s risk rule. [23] [48]
Beretta et al. prospectively followed 28 older people with PD for 12 months using weekly personal or telephone contact: 15 fell, with 23 falls in total and six recurrent fallers. Two AccuGait force plates sampled at 200 Hz during bipedal, adapted-tandem and unipedal tasks. Bipedal and tandem trials lasted 30 seconds; the first 10 seconds of each trial were discarded, and CoP signals were filtered at 5 Hz. Three trials were performed per condition or limb, although their aggregation is not explained. Crucially, all participants used a wooden support during unipedal standing. The positive finding therefore concerns supported single-leg balance, not an unaided clinical single-leg test. [24]
The predictor was interlimb asymmetry, calculated as 100 times the absolute difference divided by the sum of the two limb values. Only unipedal AP mean-velocity asymmetry had a significant logistic coefficient: OR 1.147 per index percentage point, with reported interval bounds 1.017–1.294 (confidence level unlabelled), p=.025. Reported classification was 75%, with Nagelkerke pseudo-R²=.414 and Hosmer–Lemeshow p=.334; the paper does not clearly establish whether classification refers to the full unipedal model or AP asymmetry alone. The same metric correlated with fall count (rs=.449, p=.017) and was entered into a Poisson model after significance screening. That model had likelihood-ratio χ²=8.956, p=.003, but no count-model coefficient, rate ratio, confidence interval or dispersion diagnostics was reported. Bipedal and adapted-tandem measures provided no corresponding significant signal. [24]
These are exploratory development results. Six asymmetry measures were examined in each of four configurations, with only 15 fallers; model-entry details, trial completion counts and missing-data handling remain unclear. There was no reported resampling or external validation, ROC analysis, sensitivity/specificity, or comparison with a clinical model containing prior falls. The nonsignificant Hosmer–Lemeshow test cannot establish calibration in such a small sample, and unquantified assistance from the support further limits reproducibility. The study supports further evaluation of supported-unipedal asymmetry as a candidate prognostic measure, but neither a transportable fall-risk threshold nor added clinical value has been demonstrated. [24]
Wearable standing metrics
Sturchio et al. studied 26 people with PD and orthostatic hypotension, assessed in best ON without troublesome dyskinesia. The recovered original supplement specifies firm-surface stance, feet about 30 cm apart, 30 seconds with eyes open and 30 seconds with eyes closed; falls-diary compliance was assessed monthly before six-month collection. Fourteen fell. Sway variables were available for 25 participants, while the TUG waist-sway component was available for only 17, so the mixed-model evidence is thinner than recruitment N suggests. Eyes-open sway centroidal frequency at a derived threshold of .88 Hz yielded AUC .81, sensitivity 84.6% and specificity 83.3%. A cluster comprising standing jerk, standing frequency and waist sway during TUG achieved AUC .87. The latter is a mixed standing–mobility result and cannot be credited wholly to quiet standing. [47]
The small sample, many comparisons without multiplicity adjustment, selection of only ROC results with AUC≥.8, absence of independent validation, and lack of a clinical-baseline-plus-standing comparison substantially limit confidence. Absence of significant clinical associations in 26 selected patients does not prove that instruments outperform a well-developed clinical model. The study is most useful as a hypothesis for PD complicated by orthostatic hypotension; transferring its frequency cutoff to general PD or another processing pipeline is premature. [47]
Sotirakis et al. offer a longer horizon and explicit first-fall target. Baseline assessment combined two minutes of walking with 30 seconds of eyes-closed standing. There were 13 future fallers among 98 at two years and 23 among 97 at five years; modelling used balanced subsets of only 26 and 46. Sway acceleration variability was selected alongside predominantly gait measures. Five-year random-forest AUC was .848 (95% CI .519–1.000) with age and .812 (.482–1.000) without age. These are estimates for the combined kinematic assessment, not a standing-only model. [49]
There was no gait-only versus gait-plus-standing comparison establishing the incremental contribution of sway. Feature-selection timing is inconsistently described: Methods/Results describe selection across the cohort before cross-validation, while Discussion states processing occurred within folds. This leaves optimism from information leakage unresolved; it is not proof of an implementation error. Resampling changes the fall prevalence, and some five-year outcomes relied on telephone recall. Internal discrimination with wide intervals does not establish calibrated absolute risk. This OxQUIP report also belongs to a research programme used for other digital-motor publications, which should not be counted as independent replication. [49]
Negative and discordant findings change the clinical conclusion
Fiems et al. tested the consumer Sway application in 59 people with PD against falls reported for each of the following six months. Although the abstract labels the design cross-sectional, its fall endpoint is prospective. Public publisher excerpts identify the mCTSIB Sway protocol and clarify that 68 initially agreed or were recruited, with six unable to be scheduled and three age-ineligible before 59 enrolled. These are pre-testing exclusions, not a follow-up attrition count. Results excerpts report H&Y stages 1–4, conflicting with the abstract's I–III; 26 participants had recurrent falls in the preceding six months, which must not be mistaken for the future-event count. [28]
Sway AUC was .65, versus .76 for ABC and .72 for Mini-BESTest. A mixed model contained Sway, prior falls and ABC; only history and ABC were significant. The abstract says that 85% of fallers were correctly classified, not that overall accuracy was 85%. The authors concluded that Sway did not improve prediction beyond established measures/history. This is useful negative evidence about incremental utility, but not proof of equivalence: future-event counts, coefficients/intervals, nested model comparisons, missingness and optimism correction remain unavailable. [28]
Institutional and bounded public-source checks did not recover the complete article. The protocol name is now verified from publisher excerpts, but exact mCTSIB conditions, device placement, app version, duration, repeats, medication state and aggregation remain unverified. A related Sway reliability study's stance sequence cannot be substituted for this paper's protocol. Nor should a proprietary Sway score be described as force-plate velocity or treated as interchangeable with raw smartphone acceleration. Measurement reliability and added prognostic value remain separate claims. [28]
Hoskovcová et al. assessed 45 people in both OFF and ON states and observed 27 fallers over six months, with complete standing and follow-up data. The original S1 supplement confirms six SOT conditions, three 20-second trials per condition, a safety harness and condition-level averaging. OFF testing followed 12-hour levodopa/COMT-inhibitor and 48-hour agonist withdrawal; ON testing followed 150% of the usual morning levodopa dose, always second. This acute challenge protocol differs from routine best-ON assessment. [50] SOT composite scores did not predict falls in either state: per-point OR .96 (95% CI .89–1.02) OFF and .93 (.84–1.02) ON. The striking discrimination of their final instrumented model came from gait cadence and stride-time variability, not standing sensory organization. It would be misleading to use that paper's overall model accuracy to endorse SOT. [50]
Matinolli’s final journal article confirms a regional cohort of 125 people with idiopathic PD, including 120 with sway measurements. There were 79 any-fall and 59 recurrent-fall cases over two years. Diaries were checked through calls every three months; five participants withdrew and eight were lost, with mean partial follow-up 14.5 months. Outcomes were nevertheless grouped across the original 125, without an exposure-length sensitivity analysis. Testing occurred ON medication using inclinometry, not force-plate CoP. The final article refers readers to earlier methods; the original-author thesis supplies the sacral-rod device, barefoot feet-together stance, arms at sides and two 60-second trials per eye condition. That supplementary protocol attribution remains explicit. [9] [25]
Final-journal Table 2 confirms higher baseline sway in recurrent fallers: velocity .64 versus .49 cm/s (p=.016), area 3.5 versus 2.1 cm² (p=.011), and path 38.6 versus 29.0 cm (p=.010). The table does not label the eye condition. After grouped forward-stepwise selection, only fall history, OR 3.02 (95% CI 1.23–7.44), and UPDRS-ADL, OR 1.13 (1.04–1.22), remained in the reported final recurrence model. No AUC, calibration or internal/external validation was supplied. At four years, 18 had died; baseline sway measures did not differ significantly by survival, while slow walking speed was the only retained age-adjusted mortality predictor. Greater sway thus associated with subsequent recurrent groups without demonstrating independent fall or mortality prediction. Exclusion of people unable to stand, incomplete follow-up and stepwise selection constrain generalization. The 2007 cross-sectional report shares this research programme and is not independent replication. [9] [41]
Reactive posturography is a different domain. Moraca’s AccuGait force plate sampled at 200 Hz during five unpredictable posterior translations among 15 trials; each 20-second trial contained a 5-cm translation at 15 cm/s at an unpredictable time. Participants stood barefoot, pelvic-width apart, with a safety harness. The response windows were defined relative to CoP onset and subsequent negative/positive peaks, rather than an undifferentiated whole-trial sway score. [8]
In Moraca et al., 75 participants underwent unexpected support-surface translations and weekly fall surveillance for 12 months; 29 fell. Late-response CoP range was associated with time to fall after adjustment (HR 1.45, 95% CI 1.01–2.09), yet standalone discrimination was weak: AUC .58 (.45–.72). This cleanly demonstrates why a statistically significant association is insufficient for clinical prediction. Of 86 initially recruited, ten lacked completed follow-up and one was excluded for low cognition. Three Cox models each included five clinical/demographic covariates plus different CoP measures; only significant markers proceeded to ROC testing. Prior fall history was not in those adjustment sets. These details explain why the result supports mechanistic investigation more readily than useful triage. [8]
Comparing the prospective findings
For near-term any-fall prediction, Kerr supplies internally validated mixed-model evidence; Sturchio is a selected orthostatic-hypotension pilot; Tsai largely predicts further falling in previous fallers; Fiems supplies an important negative commercial-app result. For recurrent falling, Matinolli shows the loss of an unadjusted sway signal after multivariable selection, Gervasoni suggests a velocity-factor contribution with incompletely described internal validation, and Sebastiá-Amat provides a small derivation study with unresolved classification reporting. Beretta supplies a supported-unipedal asymmetry signal without independent validation. For first falls at two/five years, Sotirakis evaluates a combined gait–sway pipeline. For unexpected perturbations, Moraca demonstrates limited discrimination despite an adjusted association. This is a comparison of evidence roles, not a defensible AUC league table: endpoints, case mix, follow-up and validation differ.
Table 2 Principal instrumented standing evidence for future falls
Intervals are 95% confidence intervals unless labelled otherwise. Access describes the material examined, not study quality. Apparent performance uses development data; internal validation is not external validation.
| Study cohort and outcome | Assessment and main estimate | Validation limitations and access |
|---|---|---|
| Kerr 2010 [6] 101 participants; 48 fallers 17/59 without recent falls subsequently fell 6 months | AP firm-surface eyes-open sway sensitivity 66%; specificity 68% Mixed clinical–sway model leave-one-out sensitivity 72%; specificity 76% | Internal validation only; no external calibration Physiological Profile Assessment sway measurement Original report access: abstract and author manuscript. Current independent comparison: abstract only; detailed body claims unverified. |
| Gervasoni 2015 [7] 53 entered; 46 followed 32 any-fall; 22 recurrent 6 months | Four 30 s firm/foam EO/EC conditions Velocity factor OR 2.37 (1.01–5.58) Mixed recurrence model AUC 0.844 (0.716–0.972) | LOOCV stated; PCA/selection nesting and estimate labelling unclear Recurrent comparator/model N unresolved; no calibrated clinical increment Access: final journal full text |
| Sturchio 2021 [47] 26 with orthostatic hypotension; 14 fallers 6-month diaries | Firm surface, ~30-cm stance, 30 s EO/EC Eyes-open frequency AUC 0.81; mixed standing–TUG cluster 0.87 | Sway N=25; TUG waist-sway N=17 Selected ROC results; no independent validation or clinical baseline comparison Access: full manuscript, tables and original supplement |
| Tsai 2023 [46] 95 analysed; 24 future fallers All 24 had prior falls Mean 12.4 months | Pressure-mat ON-state path AUC 0.73 (0.62–0.84); velocity 0.72 (0.61–0.84) Combined model 0.90 (0.83–0.98) | Sway not independent in adjusted Cox model; clinical comparator incomplete; reporting inconsistencies Access: full text and primary tables |
| Sotirakis 2024 [49] 98/97 followed at 2/5 years 13/23 first-fall cases Balanced model samples 26/46 | Mixed gait and sway random forest at 5 years AUC 0.848 (0.519–1.000) with age; 0.812 (0.482–1.000) without | Internal validation; feature-selection timing conflict; no standing increment or calibrated risks Access: full text and supplementary Table 3 |
| Sebastiá-Amat 2026 [23] 55 recruited → 48 analysed 17 recurrent / 31 non-recurrent 12 months; monthly calls | FreeMed 100 Hz; 3 × 30 s per eye condition; randomized and averaged EO speed AUC 0.81 (0.70–0.93); selected 18.1 mm/s cutoff | 17 events; extensive screening/backward selection; no validation or clinical-only comparison Reported 88.2% accuracy unreconstructable; no cutoff sensitivity/specificity Access: final journal full text |
| Beretta 2018 [24] 28 participants; 15 fallers 23 falls; 12 months Weekly personal/telephone contact | Supported-unipedal AP-velocity asymmetry OR 1.147 per index point Reported bounds 1.017–1.294 (level unlabelled) Apparent classification 75% | All used wooden support; trial aggregation unclear No AUC, resampling, external validation or clinical comparator Hosmer–Lemeshow p=.334 is insufficient calibration evidence Access: final journal full text |
| Fiems 2020 [28] 59 enrolled after pre-testing exclusions Six-month subsequent falls Future-event count unverified | Sway mCTSIB named in publisher excerpt Sway AUC 0.65; ABC 0.76; Mini-BESTest 0.72 No significant Sway contribution to mixed model | H&Y I–III in abstract versus 1–4 in Results excerpt 85% refers to fallers correctly classified, not verified overall accuracy Access: abstract + publisher excerpts; complete article unavailable |
| Hoskovcová 2015 [50] 45 participants; 27 fallers 6 months; complete standing/follow-up data | SOT: six conditions, 3 × 20 s each OR per point 0.96 (0.89–1.02) OFF; 0.93 (0.84–1.02) ON | Neither standing association significant Positive final model was gait-based Fixed OFF then 150%-dose ON challenge Access: full article, tables and original S1 supplement |
| Moraca 2021 [8] 75 participants; 29 fallers 12 months; weekly surveillance | Reactive late CoP range HR 1.45 (1.01–2.09) adjusted Standalone AUC 0.58 (0.45–0.72) | Association with weak discrimination; several models relative to 29 events; no external validation Access: accepted full manuscript |
| Matinolli 2011 [9, 25] 125 participants; 59 recurrent 120 completed sway 2-year falls; 4-year mortality | ON-state inclinometry; detailed 2 × 60 s protocol from thesis Higher unadjusted velocity (0.64 vs 0.49 cm/s) Final recurrence model: history and UPDRS-ADL | 13 partial follow-ups; stand-capable cohort; grouped stepwise selection No retained sway predictor; no significant sway–mortality group differences Access: final journal plus supplementary thesis protocol |
Any, recurrent and first falls are separate targets. Fiems remains limited to its abstract and public publisher excerpts. Matinolli stance and trial details are separately thesis-sourced. Pressure-platform, inclinometer and IMU metrics cannot share thresholds. Mixed gait–standing models do not isolate standing benefit.
Progression and wider outcomes
The outcome determines the prognostic claim
An isolated reactive sign, a clinical balance battery, a self-reported dual-task problem and quiet-standing acceleration carry different kinds of information. Direct clinical evidence links postural instability with later cognitive impairment and a balance subscale with mortality. Other studies provide meaningful negative results for future disability, quality-of-life change and nursing-home admission. Broader axial composites add prognostic context, but do not identify the independent contribution of standing balance.
The following appraisal distinguishes baseline assessments preceding future outcomes from clinical composites, serial monitoring, and changes evolving together. The distinction is especially important when interpreting cognitive decline, dependence and treatment response.
Cognitive decline and dementia
Reactive balance and future cognitive impairment
Urso and colleagues provide a more direct standing-balance result than a PIGD subtype comparison. They used the early, initially untreated PPMI cohort and separately examined the items composing PIGD. During a median five-year follow-up (IQR three to six), 79 cognitive-impairment events were reported. Baseline postural instability, defined by MDS-UPDRS III postural-stability item 3.12 ≥1, was associated with subsequent impairment: adjusted HR 2.045 (95% CI 1.068–3.918). Overall PIGD classification, gait, self-reported walking/balance and freezing items were not convincing predictors. The distinction matters: averaging symptoms into a phenotype can obscure an informative reactive sign. [10]
The adjusted model included age, sex, education, motor severity with gait/freezing/postural items removed, RBD symptoms, olfaction, CSF amyloid and caudate dopamine-transporter uptake. It did not adjust continuously for baseline cognitive performance. Formal cognitive categorisation began late for many participants, so some apparently incident cases may already have been impaired. Excluding suspected baseline MCI strengthened the postural-instability association, HR 2.573 (1.110–6.010), but did not eliminate all uncertainty about baseline status. The source starts with 422 participants, excludes 47 indeterminate motor phenotypes, and requires complete follow-up cognition/covariates; it does not clearly give each adjusted model's analytic denominator or event count. Its Methods also transpose gait/postural item numbers, while the Results and Discussion consistently identify postural stability as 3.12. These qualifications should accompany the positive finding. No external replication, calibration or discrimination statistic was reported for this isolated sign, and a hazard ratio is not a person's dementia probability. The endpoint was cognitive impairment based on neuropsychological thresholds, not exclusively adjudicated dementia. [10]
Quiet sway and cognitive change
Apthorp's force-plate study found that eyes-closed sway path correlated with MoCA (Spearman r = −0.64), executive/verbal fluency (−0.57) and quality of life (0.48) in 25 patients. These were exploratory, unadjusted, same-occasion associations; no cognitive follow-up or dementia conversion was studied. [51]
Pantall's longitudinal ICICLE-PD analysis is more informative than a cross-sectional study but still does not test baseline sway against subsequent cognitive decline. Thirty-five participants had ON-medication, two-minute eyes-open stance recordings at 18, 36 and 54 months; diagnosis-baseline sway was not included. Feet were self-positioned within a marked area. The strongest relevant finding was correlation between jerk change and MoCA change over the same 36-to-54-month interval (r = −0.422). At 36 months, jerk and MoCA were concurrently related (rho = −0.392), whereas the 54-month correlation was weaker and not significant. These are contemporaneous and change–change relationships. They do not establish that an earlier adverse sway test forecasts later cognitive decline. [14]
The companion 2018 report of 50 PD participants and 59 older controls investigated sample entropy. Its original abstract reports less regular standing dynamics in PD but no further entropy increase with progression and only weak cognitive/motor correlations. Full text was not obtained. Both papers arise from ICICLE-GAIT with matching follow-up occasions; treat them as overlapping investigations, not independent replication. [52]
Dewey supplies a genuinely prospective quiet-stance counterweight. Its baseline iSway models attempted to forecast later MoCA change and performed poorly at 12–24 months. At 24 months, the version-2 iSway model's reported gamma-derived percentage was 48.12%; this is neither AUC nor conventional classification accuracy. Taken together, the instrumented evidence supports research on shared cognitive-postural processes more strongly than an implementable sway-based cognitive forecast. [15]
Broader clinical signs and transportability
Keener's population-based cohort linked a postural-reflex-impairment composite to later MMSE-defined impairment: 34 events among 224 cognitively eligible participants, HR 1.38 (1.16–1.64). That score combines posture, gait and postural stability, and the main adjustment did not include baseline MMSE. Bäckström found univariable dementia associations for PIGD and first-year postural instability in the NYPUM incident cohort, but neither became a separate retained variable in the selected clinical/CSF model. First-year instability is also not strictly a baseline measure. [53, 54]
The 2026 population-based analysis by Malfer offers useful independent but differently timed evidence. All clinical risk factors were time-dependent, rather than fixed baseline tests. Among 572 PD/PDD cases there were 249 dementia outcomes; impaired postural reflexes had HR 1.26 (0.96–1.65) in the PD-only penalised model. The corresponding association reached significance in the combined parkinsonism cohort, which included DLB and other syndromes. The PD-specific result is therefore imprecise and not a simple replication of Urso's baseline pull-test finding. Retrospective symptom dating, heterogeneous cognitive ascertainment and censoring death rather than modelling competing risk further limit individual prediction. [55]
Practical meaning: an abnormal reactive response can be a reason to assess cognition and everyday function more carefully, particularly alongside other risks. It should not be labelled impending dementia. Newer evidence strengthens the need to distinguish the specific sign, disease stage, endpoint and prediction time; it does not justify substituting balance testing for cognitive assessment.
Motor progression and balance monitoring
Dewey evaluated baseline instrumented quiet stance and iTUG against later motor severity, cognition, mobility-related quality of life, activities of daily living and a global composite. Follow-up samples were 230, 222, 164 and 177 at 6, 12, 18 and 24 months. Testing was ON medication; iSway used three 30-second firm-surface trials, arms crossed and feet 26 cm apart. Ridge regression used repeated ten-fold cross-validation. At 24 months, iSway version 2's reported held-out gamma-derived percentages were 59.40% for MDS-UPDRS total and 59.79% for Schwab–England. These are not ordinary classification accuracy or AUC, and 50% is not an established chance benchmark for this reported metric. The study's overall conclusion was poor 12–24-month prediction. No externally validated increment over a well-specified clinical benchmark was demonstrated. This negative evidence is directly relevant to baseline sway prognosis, although bounded by the protocol, medication state and outcomes tested. [15]
Mancini's 2012 pilot instead followed the measurement itself. Thirteen initially untreated patients and 12 controls underwent repeated trunk-accelerometer stance testing over a year. There was no significant group-by-time interaction. Exploratory standardised response means suggested worsening mainly among five patients remaining untreated: mediolateral dispersion SRM 0.90 and jerk 0.56. Eight began medication and followed a different pattern despite subsequent overnight-OFF testing. This supports feasibility and a monitoring hypothesis, not a baseline forecast or established progression biomarker. The cohort appears to overlap the companion ISway validity sample. [56, 16]
The Pantall studies complicate any simple claim that sway must increase as PD progresses. Some acceleration amplitude and ellipse measures decreased between visits despite increasing motor burden and medication dosage; sample entropy did not show monotonic progression. Longer recordings also revealed different findings from their nested 30-second segments. Changes in foot placement, stiffening strategies, treatment and task duration can alter measured sway without mapping directly onto a single disease trajectory. This is a substantive biological and measurement limitation, not merely a need for a larger sample. [14, 52]
Sotirakis followed 74 patients over seven visits, combining two-minute walking and 30-second eyes-closed sway. Sensor-derived contemporaneous motor-score estimates detected group change earlier than the clinical rating. Feature selection across visits, uncertain participant-disjoint splitting and interpolation using surrounding visits limit forecasting interpretation. The study did not validate a future-only standing-sway risk model or isolate sway's added value. [57]
Clinical scales remain valuable for monitoring. Duncan retained 51 of 80 participants at six months and 44 at twelve. Mini-BESTest declined by 1.50 points at twelve months (95% CI −2.41 to −0.59), whereas BBS did not change. Limits of stability, reactive responses and sensory orientation were particularly susceptible; more affected participants were disproportionately lost. This is group-level responsiveness. It does not mean a 1.5-point decline can reliably distinguish deterioration from measurement error in an individual. Score version, medication state and applicable minimal detectable change still matter. [58]
Incident freezing and mobility milestones
Standing anticipatory and reactive control plausibly contribute to freezing, but identifying an existing freezer, warning seconds before an episode and predicting first freezing over years are distinct targets. Hou's 2024 APA/limits-of-stability study compared 23 patients already experiencing slight freezing with 25 non-freezers and 24 controls. Its accessible original abstract supports concurrent classification, not incident-freezing prognosis. [59]
D'Cruz did follow non-freezers prospectively: 12 converted over two years among 56 analysed participants. Mini-BESTest anticipatory and dynamic-balance measures were not convincing individual predictors; the anticipatory-domain OR was 0.70 (0.39–1.28), AUC 0.60. Predictors came from the visit preceding conversion, rather than uniformly from enrolment. A small event count and wide intervals do not rule out useful effects, but do not support implementing a standing-domain freezing forecast. [60]
The newer ONDRI report by Faerman followed 120 patients for two years, including 25 transitional freezers and 24 already freezing at baseline. Its abstract reports changing motor, affective and cognitive predictors according to the lead time before freezing. Postural-instability/gait-difficulty scores worsened in converters, but an isolated baseline standing measure's contribution was not established. The reported multivariable models had high specificity and low sensitivity; their accuracy cannot be attributed to balance. Full text was unavailable, so detailed model composition, calibration and validation remain unverified. [61]
Shao's 2024 model is adjacent clinical evidence: 606 early-PD PPMI participants without the walking-and-balance milestone at entry, a 70/30 site-based development/validation division, and a combined model including age, processing speed, PIGD, non-motor experiences and RBD. The original abstract reports C-indices of 0.75 and 0.76. It combines later freezing/falls and contains no isolated instrumented stance input. Without full text, its detailed milestone definition, event counts and calibration should remain unverified rather than borrowed into a standing-balance recommendation. [62]
Disability dependence and everyday mobility
Patient-relevant disability is not synonymous with an examination score. Becoming unable to perform basic activities, needing help at home, reporting walking difficulty and reaching Hoehn–Yahr stage 3 should be treated as separate outcomes.
Lindh-Rengifo tested baseline predictors of walking difficulties three years later in 148 participants with paired Walk-12G data; 255 had participated in the original baseline study. Postural instability on the original UPDRS pull-test item was associated with worse future walking scores univariably. In the adjusted model (n = 146), including baseline walking difficulty, its coefficient was smaller and uncertain: B 2.56 points (95% CI −0.207 to 5.33). By contrast, a self-reported question about balance problems while standing or walking during a dual task remained associated, B 4.42 (1.55–7.29). This is useful positive patient-reported evidence and a qualified null for the reactive clinical sign, not proof that a laboratory dual-task stance test predicts disability. The adjusted model explained 67.2% of variance in the development data, but this was not externally validated predictive performance. Attrition, manual variable removal, shared content between the predictor and walking questionnaire, and the absence of an independent validation sample constrain reuse. [12]
For actual functional dependence, the 2026 Parkinson's Incidence Cohorts Collaboration supplies much larger longitudinal context: 883 incident patients across six European cohorts, median follow-up 7.3 years, and 369 new sustained-dependency outcomes. In the dependency model (n = 742), baseline axial severity was associated with dependency, HR 1.18 per score point (1.12–1.25). Sustained dependency meant needing help with basic activities and persisting at subsequent visits. The axial score combines rising from a chair, posture, gait and postural instability. It was examined as an alternative to the overall motor score, not as a validated additional standing test. Multiple imputation and adjustment for clinical/genetic factors improve the analysis, but shared functions across predictor and dependency definition, cohort heterogeneity and the absence of a balance-only ablation remain important. NYPUM, PICNICS and ICICLE-PD are included, so this is not wholly independent replication of their earlier reports. [63]
COPPADIS provides another boundary check. In 463 initially independent patients, 55 became dependent over two years. Non-tremoric phenotype and freezing were associated with outcome, but models also included changes occurring during the same follow-up interval. This is not a pure baseline standing forecast; no isolated stance or pull-test contribution was established. Likewise, Shulman's older study of 618 patients described the association of disability with disease stage cross-sectionally, despite language about the evolution of disability. It should not be counted as a prospective balance model. [64, 65]
Ribeiro's independent PICNICS validation is stronger prediction-model evidence for a different target: age, animal fluency and an axial score predicted a five-year composite of postural instability, dementia or death. Among 198 baseline-eligible participants, 93 reached the composite; AUC was 0.80 (0.74–0.86). This validates a practical combined model, but does not separate the three outcomes, attribute performance to standing balance, or supply reported calibration in the inspected validation article. Lo's multimodal smartphone battery similarly forecast an 18-month help-at-home outcome, but the relevant analysis included only 39 participants and 11 events; balance was combined with gait, voice, tapping, reaction time and tremor, without a balance-ablation comparison. Its recording-wise cross-validation ignored within-person correlation, and the reported clinic-only leave-one-subject-out AUC was lower. [66, 67]
Quality of life
The distinction between current burden and future change is particularly clear here. Apthorp's quiet-sway association was concurrent; Dewey's prospective forecasts of PDQ-39 mobility change were poor. Mobility-specific quality of life is also narrower than overall well-being. [51, 15]
Bowman's rehabilitation cohort assessed 42 patients before 20 balance/gait sessions, with 37 contributing post-treatment modelling. BBS was related to baseline PDQ-39 mobility (r = −0.40), but baseline BBS did not predict its subsequent change (r = −0.06, p = 0.70); baseline balance confidence was also unconvincing for change (r = 0.09, p = 0.58). Worse initial mobility-related quality of life, sex and disease stage entered the selected change model. The result does not show that objective balance is irrelevant to lived experience. It shows that explaining current disability and predicting who will report improvement are different questions. A selected, small rehabilitation sample, regression to the mean and lack of external validation limit the model. [13]
Older DATATOP evidence provides a further temporal distinction. Marras studied 362 patients with SF-36 assessments about 1.7 years apart, beginning five to six years after trial enrolment. The original abstract identifies depression, self-rated cognition, age and functional independence as predictors of later quality-of-life change; PIGD changed concurrently with quality of life. The full text was unavailable, so this supports the temporal distinction, not a precise balance effect estimate. Across the targeted sources inspected, a calibrated, externally validated baseline standing test for future overall quality of life was not established. This is a bounded search finding, not evidence that no such study exists. [68]
Institutionalisation and survival
Clinical balance and mortality
Laurent's 2023 study specifically separated Tinetti balance from gait. Ninety-eight patients with late-onset PD, mean age 79.4 years, underwent geriatric assessment and retrospective follow-up averaging 3.4 years, up to five years. Eighteen died and 19 entered a nursing home. In the complete-case Cox model (90 participants, 17 deaths), better Tinetti balance was associated with lower mortality, HR 0.82 per point (0.66–0.96), adjusted for age, sex, comorbidity and gait speed. Gait speed did not remain significant. This is a genuine future-outcome association for a clinically administered balance battery, rather than just PIGD. [11]
It is not isolated quiet-stance evidence: Tinetti balance includes transfers, turning and perturbation-related tasks as well as standing. An exploratory eyes-closed item association lost significance after multiple-comparison correction. The post-hoc low-versus-high balance contrast was large, HR 6.65 (1.68–26.30), but imprecise, derived in the same small sample and based on internally inconsistent printed category boundaries. Five candidate covariates and 17 deaths, univariable screening, excluded early losses/alternative diagnoses, and incomplete follow-up create substantial instability. The flow diagram shows 11 lost to follow-up and 29 administratively censored; the Results prose gives a conflicting loss count. No external validation, calibration or tested clinical decision rule was established. The result supports broader frailty/functional assessment after poor performance; it does not justify a mortality threshold or life-expectancy estimate. [11]
This is also not the first physical-function/survival study. Gray's 2009 cohort of 109 patients examined seven-year mortality: Tinetti balance and several other physical tests were associated univariably, but the final reported survival model selected age, sex and Tinetti gait. Only the original abstract was available, so no unverified balance-specific hazard ratio is supplied. Together, Gray and Laurent suggest prognostic physical vulnerability while showing that which component survives adjustment can vary with age, cohort and model. [69]
Matinolli's final journal report supplies an earlier instrumented counterexample. Mortality was ascertained at four years, separately from its two-year falls follow-up: 18 of 125 participants died. Table 2 shows no statistically significant baseline sway differences between those who died and survivors, including area (p = 0.393), velocity (p = 0.166) and path length (p = 0.143). The measurement was inclinometry, not force-platform CoP. After age adjustment, slow walking speed was the only retained mortality predictor. This does not prove that standing sway has no true prognostic association, but it supplies direct negative evidence for the tested population and analysis. It also limits generalisation from Laurent's multidomain clinical balance battery to an arbitrary instrumented sway measure. [9]
Bäckström's baseline PIGD association with mortality in NYPUM remains important broader-composite context; it overlaps the cohort used for its dementia paper. Keener's postural-reflex composite also includes gait and posture and had time-varying mortality effects. These do not become independent replications of Laurent's balance subscale or of an instrumented stance protocol. [70, 53]
Nursing home admission
Laurent found no significant predictor of nursing-home admission, including Tinetti balance, across 19 admissions. The small sample limits power, but this is a measured null outcome rather than an unsearched gap. An independent mid-/late-stage cohort of 89 patients studied by Ygland Rödström found that raw PIGD severity was associated with later nursing-home entry after adjustment (HR 1.19 per point, 1.05–1.35), whereas PIGD-versus-tremor phenotype was not convincing. Crucially, the raw-score association lost significance when participants with imputed UPDRS items were excluded. This finding concerns a broader, exploratory composite, not standalone standing control. Prebaseline admissions were excluded from time-to-event analyses; reported total milestone counts must not be mistaken for incident events. [11, 71]
The older population-based Aarsland study followed 178 community-dwelling patients for four years, with 47 nursing-home admissions. Its accessible original abstract identifies age, functional impairment, dementia and hallucinations as independent predictors. It establishes the importance of broader clinical/social context, not an isolated standing test. Available family care, resources and housing can alter placement even at similar motor impairment. Accordingly, the search did not establish a validated standing-only institutionalisation model, but did identify direct null data and limited composite evidence. [72]
Rehabilitation and DBS response
Treatment-associated change is a separate prognostic target. A baseline score can be associated with improvement because it identifies room to improve, regression to the mean, measurement coupling or genuine responsiveness. Selecting which treatment benefits a person most requires a predictor-by-treatment comparison, not merely an association among treated participants.
The rehabilitation evidence is wider than one recent random forest. Löfgren's 2019 original abstract describes 47 patients receiving ten weeks of challenging balance/gait training; perceived health, TUG and cognitive performance explained 35% of balance-change variance. This was a treated-group association, not an isolated standing predictor. Joseph's 2020 HiBalance analysis included 61 intervention and 56 control participants; 53 completed intervention. Lower balance confidence and attending at least 80% of sessions were associated with a ≥2-point Mini-BESTest response. The reported model AUC was 0.84, but attendance is accumulated during treatment, so the full model cannot be used as a baseline forecast. Full texts of these two reports were unavailable, and model calibration/validation and treatment-interaction details were not independently verified. [73, 74]
Klamroth randomised 43 patients to perturbation or conventional treadmill training and found a larger reactive Mini-BESTest responder proportion with perturbation training (44% versus 10%; risk ratio 4.22, 1.03–17.28). Its original abstract says responders had lower initial balance/cognitive performance on selected measures, with some demographic patterns described only descriptively. This is encouraging domain-specific response evidence, but does not establish a validated rule for selecting patients by baseline quiet or reactive stance. The wide interval and multiple responder definitions deserve attention. [75]
Albrecht's later HiBalance analysis supplies a salutary test of individual prediction. Among 39 treated participants, the balance-response random forest had 59% out-of-bag error and weighted kappa 0.15. Baseline impairment may help generate hypotheses about response, but these results do not support reliable personalisation, and the treated-arm model did not estimate a predictor-by-treatment contrast. The 2025 responder review was used to chase these earlier originals; its heterogeneous responder definitions and study designs should not be converted into a pooled standing-balance prediction claim. [76, 77]
Yin's retrospective STN-DBS study offers a different positive signal: preoperative BBS levodopa response was associated with postoperative OFF-medication BBS and pull-test improvement at one and twelve months. The 111-person “validation” set was a complete-follow-up subset of the 261-person exploration set, not an independent cohort. ON-medication balance gains were not significant. The predictor and BBS-change outcome both contain the preoperative baseline, raising mathematical-coupling concerns. No medication-only/sham comparison or calibrated external prediction model was provided. This may guide further response research, but does not justify selecting surgery from standing-balance performance alone. [78]
The practical conclusion is to assess balance domains, offer appropriately targeted rehabilitation and measure meaningful functional response. Neither a weak predicted response nor a poor baseline score is sufficient reason to withhold treatment. Improved test performance also does not demonstrate prevention of dementia, institutionalisation or mortality; those longer-term benefits require their own evidence.
Table 3 Principal evidence for progression and wider outcomes
Intervals are 95% confidence intervals unless labelled otherwise. Access describes the material examined, not study quality. Apparent performance uses development data; internal validation is not external validation.
| Study cohort and outcome | Assessment and main estimate | Validation limitations and access |
|---|---|---|
| Dewey 2022 [15] 230/222/164/177 at 6/12/18/24 months | Baseline iSway to clinical change Poor 12–24-month forecasts with repeated ten-fold cross-validation | Reported percentages use Goodman–Kruskal gamma, not classification accuracy No external validation Access: full text and tables |
| Mancini 2012 [56] 13 initially untreated PD; 12 controls 1 year | Repeated trunk sway; no significant group-by-time interaction Exploratory worsening among 5 untreated participants | Monitoring pilot, not baseline prognosis; 8 began medication Likely overlap with ISway validity cohort Access: full manuscript |
| Duncan 2015 [58] 80 baseline; 51/44 at 6/12 months | 28-point Mini-BESTest 12-month change −1.50 (−2.41 to −0.59) Berg did not change | Group responsiveness differs from reliable individual change; informative attrition Same parent cohort as falls reports Access: full manuscript |
| Sotirakis 2023 [57] 74 analysed; 7 visits over 18 months | Walking and standing features estimate contemporaneous clinical motor scores Earlier detection of group change | No future-only stance forecast; feature selection across visits; participant separation unclear Access: full text and tables |
| Keener 2018 [53] 224 cognitively eligible; 34 impairment events Mean 5.9 years to detection | Posture–gait–stability composite HR 1.38 (1.16–1.64) for MMSE-defined impairment | Not isolated stance or pull test; baseline cognition absent from main model; informative attrition Access: full text; sensitivity supplement not inspected |
| Bäckström 2018 and 2022 [70, 54] Same 143-person NYPUM cohort 77 deaths; 64 dementia events | PIGD associated with mortality and univariably with dementia Combined clinical/CSF dementia AUC belongs to other retained predictors | Clinical composite, not sway; overlapping cohort and distinct endpoints First-year instability is not strictly baseline Access: full text and tables |
| Ribeiro 2024 [66] Independent PICNICS cohort 198 at diagnosis; 93 composite events 5 years | Fixed age, fluency and axial model AUC 0.80 (0.74–0.86) Outcome: instability, dementia or death | Positive external validation of composite model; no isolated balance contribution or reported calibration Access: full text |
| D’Cruz 2020 [60] 56 analysed; 12 incident freezers 2 years | Anticipatory balance OR 0.70 (0.39–1.28); AUC 0.60 | Visit before conversion pooled rather than uniform enrollment horizon; few events Access: full manuscript and Table 2 |
| Lo 2019 [67] 110 people in cognition analysis 39 in help-at-home analysis 18 months | Multitask phone battery includes balance Reported clinic leave-one-subject-out AUC 0.82 for cognition; 0.83 for help at home | No isolated stance contribution; repeated-window grouping unclear; no external validation Access: full text and tables |
| Albrecht 2024 [76] 39 treated participants 10-week balance response | Balance-response random forest out-of-bag error 59%; weighted kappa 0.15 | Treated-arm exploratory analysis; no treatment interaction or independent validation Access: full text and tables |
| Yin 2021 [78] 261 DBS participants; nested 111 with 12-month data | Preoperative levodopa balance response associated with postoperative OFF-state improvement | Retrospective; overlapping validation sample; shared baseline; no untreated comparator Access: full text |
| Urso 2021 [10] PPMI; 422 initially; 79 cognitive events reported; median 5 years | Baseline postural-stability item ≥1 to cognitive impairment Adjusted HR 2.05 (1.07–3.92); sensitivity analysis 2.57 (1.11–6.01) | Model-specific denominators unclear; delayed cognitive categorisation; no continuous baseline-cognition adjustment No external validation or calibrated risks Access: full primary text and tables |
| Pantall 2018 [14, 52] 35 in cognitive analysis; overlapping ICICLE-PD cohort; assessments at 18, 36 and 54 months | Two-minute ON eyes-open trunk sway Jerk change and MoCA change over the same interval correlated r = −0.422 | Concurrent change, not baseline sway predicting later cognition Medication, foot placement and duration matter Access: full primary text for cognitive analysis; companion entropy abstract |
| Lindh-Rengifo 2021 [12] 148 with paired walking scores; adjusted model 146; 3 years | Pull-test coefficient after adjustment 2.56 (−0.21 to 5.33) Self-reported dual-task balance difficulty 4.42 (1.55–7.29) | Baseline walking adjustment attenuated reactive-sign association; no laboratory dual-task prediction established Attrition; in-sample model only Access: full text and tables |
| Macleod 2026 [63] 883 across 6 incident cohorts; 369 sustained-dependency events; model 742; median 7.3 years | Baseline axial score to sustained basic-ADL dependence Adjusted HR 1.18 per point (1.12–1.25) | Chair rise, posture, gait and stability composite; no isolated stance increment Overlaps earlier cohort publications Access: full text and tables |
| Bowman 2018 [13] 42 at baseline; 37 post-treatment; 20 rehabilitation sessions | Berg correlated with current mobility-related quality of life, r = −0.40 Baseline Berg versus subsequent change r = −0.06; p = 0.70 | Direct negative finding for response prediction; small selected sample No external validation or treatment interaction Access: full text and tables |
| Laurent 2023 [11] 98 late-onset PD; mean age 79.4; mean 3.4 years; 18 deaths and 19 nursing-home admissions | Tinetti balance to mortality: adjusted HR 0.82 per point (0.66–0.96) Cox model 90 participants and 17 deaths No predictor of nursing-home admission | Multidomain balance battery; few events; selected cutoffs and reporting discrepancies No external validation/calibration Access: full text, tables and flow diagram |
| Ygland Rödström 2021 [71] 89 mid- or late-stage PD; mean follow-up 8.1 years | Raw PIGD to nursing-home entry: adjusted HR 1.19 per point (1.05–1.35) Association lost significance after excluding imputed items | Broad composite and sensitivity-analysis instability Prebaseline admissions excluded from time-to-event models Access: full text |
| Matinolli 2011 [9] 125 participants; 18 deaths and 107 survivors; 4-year mortality ascertainment | Baseline inclinometer sway did not differ significantly Area p = 0.393; velocity p = 0.166; path p = 0.143 | Direct instrumented null; no standing mortality model or external validation Four-year mortality is distinct from two-year falls follow-up Access: final journal text and Table 2 |
Clinical composites, isolated reactive signs, instrumented sway and patient-reported difficulty have different evidence roles. Same-occasion correlations, time-varying signs and contextual abstract-only studies are discussed in the text rather than counted as equivalent baseline forecasts. Several cohorts recur across publications.
Choosing standing assessments and protocols
A usable selection guide
A useful Parkinson assessment should answer a question about balance, rather than collect the greatest possible number of sway variables. For an ambulant person with mild-to-moderate PD, a defensible clinical starting point is a multidomain assessment such as the current 28-point Mini-BESTest, supplemented by a task-specific quantitative measure when that measure answers an additional question. This is a reasoned selection from the measurement evidence, not a validated universal battery or proof that using this combination improves outcomes. The scale supplies information about anticipatory control, reactive recovery, sensory orientation and gait; an instrumented stance supplies finer resolution within a much narrower task. [18, 43, 16]
- For an overall clinical balance profile: use Mini-BESTest, or the full BESTest when a more extensive system-level examination is worth the time. Berg remains useful when transfers and basic standing function are limiting, but a near-perfect Berg result can leave important deficits unresolved. In Schlenstedt's 85-person comparison, 15 participants achieved the Berg maximum, including four previous fallers; the corresponding maxima were two for FAB and one for Mini-BESTest. This supports choosing a more challenging scale when a ceiling is likely, rather than abandoning Berg in every patient. [27]
- For subtle quiet-standing control or repeated objective measurement: select one validated CoP or lumbar-sensor pipeline and a feasible, reproducible stance. Jerk and acceleration RMS have better short-interval support than median frequency in the original ISway study. A pressure-platform protocol may instead prioritize CoP velocity and a directional amplitude measure, but no universal best feature or device emerged. An acceleration value must remain an acceleration value; similar variable names do not establish equivalence with CoP. [16, 79]
- For visual/surface dependence: add matched firm/foam and eyes-open/closed conditions, usually through a defined mCTSIB protocol. Instrumentation is particularly useful when the person reaches the timed ceiling. Specialized SOT provides a controlled sensory challenge and useful absolute-error evidence, but foam and sway-referenced support are not mechanically identical. These tests characterize the balance response to altered information; they do not by themselves diagnose a specific vestibular or sensory lesion. [43, 19]
- For unilateral support or controlled leaning: use bilateral single-leg stance or a specified reach/limits-of-stability task. Record side and movement strategy. These tasks are useful for relevant rehabilitation goals and offer information that a quiet stance cannot provide; they are not substitutes for testing recovery after a perturbation. [80, 81, 82]
- For reactive instability: include an appropriately administered reactive-stepping assessment. The test version, perturbation, trial number and required assistance matter. A normal backward pull result should not end the assessment when symptoms or history suggest instability. Reactive tests require an examiner able to prevent a fall and are unsuitable for unsupported self-testing. [83, 2]
- For remote measurement: the strongest directly relevant example here is repeated, quality-controlled waist-phone testing in selected early PD. Its evidence concerns two-week aggregates, not a one-off home score. Camera-derived balance is promising, but a validated RGB-depth pipeline is not evidence for an ordinary webcam with a different pose model. Begin with a clinician-selected feasible task and establish remote safety and data quality before adding challenge. [17, 84, 85]
The accompanying table specifies what investigators actually tested. A published protocol is a reproducible starting point, not automatically the recommended protocol for all PD stages or all clinical purposes.
Trial duration and repetitions
The frequently used three-by-30-second protocol has practical advantages and original PD precedents. Sebastiá-Amat's 2023 study used three trials under each visual condition and averaged them, with a one-minute rest between trials. However, its design examined cross-sectional relationships, not whether 30 seconds or three repetitions optimized reliability. The authors explicitly selected a commonly used combination from earlier literature. ISway's measurement studies also differed: concurrent force-platform validation used three two-minute trials, whereas the clinic reliability experiment used three 30-second trials. These results should not be collapsed into one supposedly fully validated protocol. [79, 16]
Primary methodological experiments explain why duration must be matched. Carpenter studied 49 healthy young adults: lengthening samples from 15 toward 120 seconds altered RMS and frequency estimates and improved their repeatability. Van der Kooij studied ten healthy adults for 600 seconds under eyes-open and eyes-closed conditions. Stability depended on the outcome and vision: some eyes-open time-domain measures stabilized after 60 seconds, whereas eyes-closed displacement measures required much longer samples. Averaging ten 60-second segments did not reproduce the metric calculated from the continuous 600-second record. Neither study establishes that people with PD should routinely stand for several minutes; both show why short-trial values cannot simply be judged against long-trial norms. [86, 87]
For a practical service, select a tolerable fixed duration and aggregation rule, then establish repeatability for the intended population and chosen features. Longer recording is useful only if it captures information that matters without making completion, fatigue or safety unacceptable. Frequency, entropy and stabilogram-diffusion features particularly require their own duration, filtering and analysis-window justification. A person who stops early should have a recorded completion time and reason; the software should not silently present a ten-second calculation as interchangeable with a thirty-second result.
Foot position arms and surface
The retrieved PD protocols range from shoulder-width feet with arms at the sides, to a ten-centimetre heel separation with crossed arms, to feet together or tandem with hands on the hips. These are different tasks. Chiari's study of 50 healthy young adults found that most of 55 sway parameters depended on some combination of anthropometry and foot geometry; only eleven were independent of the selected biomechanical variables. This supports recording and reproducing foot placement, not prescribing one optimal width for all people with PD. [79, 16, 84, 88]
A repeated comfortable stance can answer how a person performs using their usual base of support. A fixed narrow stance answers how they cope with a prescribed constraint. If a patient later needs a wider stance to complete testing, that is clinically relevant, but the changed-width sway value is not a clean within-protocol comparison. Record heel separation and foot angle or trace foot outlines; retain footwear, arm rule, gaze target and support contact. For foam, preserve product, dimensions and firmness rather than treating all pads as equivalent. Freeman's paper prints an Airex size of 18 × 18 × 5 inches; that specification should be checked before replication rather than silently substituted with a familiar pad. [43]
Medication and involuntary movement
Use a consistent medication state and record elapsed time since the relevant dose, dyskinesia, fatigue and unusual symptoms. This is particularly important when the aim is longitudinal measurement rather than a deliberate ON/OFF comparison. Curtze's 104-person experiment showed increased sway with levodopa in the dyskinetic subgroup, while several gait features improved; the nondyskinetic subgroup did not show the same sway increase. [89]
Paul's between-day study tested 31 independently ambulant people with PD, using the same assessor one week apart at a matched point in the medication cycle when treatment was working optimally. Fifteen reported dyskinesia and six reported disabling dyskinesia. The instrument was a waist-mounted mechanical swaymeter, with barefoot hip-width standing on the floor and on reported 7 mm medium-density foam, with eyes open and closed for 30 seconds. These body-displacement path measurements are not force-platform CoP or lumbar acceleration. Across all participants, sway ICCs ranged from 0.04 to 0.51 and SEMs from 108 to 172 mm. [29]
The complete tables materially qualify this negative finding. After inspecting Bland–Altman outliers, the authors repeated the analysis excluding the six participants with disabling dyskinesia. In the remaining 25, floor eyes-closed ICC rose from 0.29 (95% CI −0.07 to 0.58) to 0.80 (0.59–0.91), while SEM fell from 136 to 24 mm. Foam-condition ICCs were 0.75 and SEMs 34–40 mm; floor eyes-open reliability remained more modest at 0.50 (0.14–0.74). This is an exploratory subgroup analysis, and 'without disabling dyskinesia' does not mean without any dyskinesia. The study supports recording involuntary movement and matching conditions, rather than dismissing all standing measures or treating the subgroup results as universally reliable. It supplies absolute error for this protocol, not a minimal important change or future-risk threshold. [29]
Lower sway also cannot always be labeled better. Álvarez's earliest-stage subgroup sometimes swayed less than controls. A constrained or stiffened movement strategy can reduce excursion while leaving limits of stability or recovery stepping impaired. Show amplitude, direction, task completion and the relevant clinical findings together rather than converting every reduction into a generic improvement score. [82, 84]
Interpreting measurement change
Paul's additional mechanical-swaymeter tasks illustrate why reliability should be reported by task. Maximal anteroposterior balance range had ICC 0.81 (95% CI 0.63–0.91) and SEM 17 mm in 28 participants. Coordinated stability had ICC 0.50 (0.18–0.72) and SEM 12 points in 31 participants, improving to 0.67 (0.38–0.84) and SEM 5 points after exclusion of disabling dyskinesia. These are task- and subgroup-specific estimates; they do not validate a new digital tracing implementation or a manufacturer limits-of-stability score. [29]
Useful absolute-error evidence comes from controlled platform tests and clinical scales, rather than a universal raw-sway threshold. In Harro's between-day PD study, SOT composite, LOS endpoint excursion and MCT latency had MDC95 values of 11.6 points, 13.8 percentage points and 7.4 ms, respectively. These estimates are tied to the system and tasks in the table; they cannot be assigned to home foam tests or manual perturbations. Harro's printed LOS reaction-time SEM and MDC do not reconcile with its stated formula, so that particular threshold is not endorsed for implementation. [19]
For the Mini-BESTest, Löfgren's independent clinical administrations give a realistic indication of uncertainty: total-score SEM was 1.5 points between raters and 1.2 on same-rater one-week retest, with corresponding smallest real differences of 4.1 and 3.4 points. Table 3 gives ICC confidence intervals of 0.37–0.87 and 0.60–0.90; these differ slightly from the abstract, and the table values are used here. Reactive-subscore error was proportionally greater than sensory-subscore error. A one-point change or a small subscore shift therefore warrants caution even when the overall ICC appears good. [18]
Godi's 148-person, four-week balance-therapy study supports a roughly four-point Mini-BESTest improvement as clinically important in that treatment setting. The source was abstract-verified, and its patient and therapist anchors did not perform identically. Mehdizadeh's later study supplies additional responsiveness estimates but contains internal reporting discrepancies; it should not be used to overwrite an application's change thresholds without resolving those details. The appropriate lesson is to retain anchor, population, direction of change and intervention context alongside each threshold. [90, 91]
Clinical simplicity does not guarantee small error. Steffen's original results give MDC95 values of nine centimetres for forward reach and seven for backward reach; these estimates used two-trial-average reliability and came from a parkinsonism cohort. They do not automatically apply to the three-trial average in the 2023 pressure-platform study. Timed Romberg can also look highly repeatable while offering little headroom or substantial absolute error. [92, 79]
Responsiveness should guide the outcome toward the rehabilitation target. Hasegawa's exercise-versus-education crossover improved several gait and anticipatory-adjustment outcomes, while quiet-sway and automatic-response measures did not improve at the prespecified significance level. Reactive responses were not specifically trained. A stable quiet-sway value therefore should not obscure real improvements in another balance domain; conversely, improvement in a practiced task does not establish a benefit in unmeasured reactive control. [93]
Reference values and task completion
The available control samples supply useful context, but most are too selected and too protocol-specific to justify a universal normal range. A reference value should identify age range, disease stage or healthy status, device, task duration, foot position, sensory condition and processing. Published faller cutoffs also need the actual outcome definition and follow-up period. A cut point that separates historical fallers, a manufacturer sensory score, an MDC and an important-change estimate answer four different questions.
The ten-second single-leg landmark illustrates this distinction. Chomiak's study found low- and high-duration clusters and related shorter stance to motor impairment; it did not follow participants to validate a future-fall boundary. Similarly, Freeman found that a conventional SOT abnormal-sway cutoff missed four of five historical recurrent fallers. A single green result must not imply absence of clinically relevant instability. [80, 43]
Completion is itself informative. In the HoloLens study, the least challenging feet-together firm eyes-closed stance was completed by 67 of 72 PD participants, while harder conditions selected out poorer balancers. Technical recording errors affected 9% of trials. Separate 'could not maintain stance', 'stopped for safety', 'tracking failed' and 'completed with a changed strategy'; each has a different interpretation. Do not assign zero sway to a missing recording, combine technical failure with clinical failure, or hide failed trials behind an apparently reassuring mean of the successful ones. [84]
Specific implications for rehabtools
A useful first implementation would support a defined clinical scale plus a small number of explicit stance protocols, each with its own identity and outputs. It should expose raw task completion and clinically meaningful observations, not merely a combined balance number. For sensors, preserve hardware, anatomical placement, calibration, sample rate, filtering, units, algorithm version, trial duration and aggregation. For cameras, add view, frame rate, coordinate reference, pose-model version and tracking exclusions.
Existing evidence supports selectable measurement modules: a clinical multidomain score; quiet CoP or trunk-acceleration sway; a sensory-condition comparison; and, where supervised, unilateral support, controlled leaning or reactive recovery. The choice is driven by the assessment question and patient capability. Home phone aggregation has a clearer precedent than unsupervised difficult stance or manual perturbations. Ferraris supports selected alignment-angle measurements against Kinect; Kaya supports particular RGB-depth COM measures against Vicon. Neither validates camera-based CoP, an arbitrary replacement camera pipeline, or an individual community-fall probability. [17, 85, 84]
Use the current official test forms and administration guidance, including the Mini-BESTest's 28-point scoring and side rules. If numerical reference or change labels are displayed, attach the supporting population and protocol; leave the label unavailable when those conditions do not match. This creates a clinically interpretable measurement tool while leaving genuine uncertainty visible. [94]
Table 4 Standing assessments and the protocols supporting their use
Protocols and properties refer to the named source and population. ICC is relative reliability; SEM and MDC concern absolute error. A tested protocol is not a universal prescription. EO means eyes open; EC means eyes closed.
| Assessment and tested protocol | Measurement properties and limits | Appropriate use |
|---|---|---|
| Quiet CoP, Sebastiá-Amat 2023: 52 ambulant PD, ON 45–90 min; shoulder-width feet, arms at sides, target 1.5 m; firm EO/EC, 3 × 30 s each, 1-min rests, mean of three; pressure platform 100 Hz [79] | Mostly moderate clinical convergence; no retest SEM/MDC or MIC in this cross-sectional study. Footwear/filter unspecified. CoP velocity, AP/ML RMS and ellipse area retain distinct meanings | Objective quiet-standing complement. Reproduce device/processing and stance; three-by-30 s is a research precedent, not proven optimum |
| Lumbar ISway: L5 sensor; reliability 17 PD ON, 3 × 30 s, same-rater 30-min retest with replacement; separate validation used 3 × 120 s, EO, 10-cm heels, crossed arms [16] | Jerk ICC .86 (.66–.95); RMS .83 (.59–.93); median frequency .35 (−.13–.70). RMS–CoP r .74, velocity–CoP r .12. No measurement SEM/MDC; short interval | Portable quiet-sway measurement. Jerk/RMS are stronger candidates than median frequency for this protocol; validate between-day error before individual change labels |
| Mechanical waist swaymeter: 31 PD, matched optimal ON, same-assessor one-week retest; barefoot hip-width stance, floor and reported 7-mm foam, EO/EC, 30 s each; path in mm [29] | All-participant ICC 0.04–0.51, SEM 108–172 mm. Excluding six with disabling dyskinesia: floor EC ICC 0.80 (0.59–0.91), SEM 24 mm; floor EO ICC 0.50; foam ICC 0.75. Post hoc subgroup; no MIC | Protocol-specific between-day measurement. Document dyskinesia and the tested population; do not transfer mechanical-swaymeter error to CoP, IMU or camera outputs. Access: scoped publisher text, tables and supplement |
| Instrumented mCTSIB: 26 PD ON; feet together, crossed arms, L5 Opal; 3 × 30 s for firm/foam × EO/EC; specific Airex foam [43] | AP-range composite–SOT r −.64; firm EC correlation nonsignificant. No new composite retest MDC/MIC. One participant completed only 1/3 firm-EC trials; only five historical recurrent fallers | Sensory-condition comparison; retain condition values and failures. Foam is not equivalent to movable support; difficult conditions need guarding |
| SOT: NeuroCom six sensory conditions, normally 3 × 20 s each, overhead harness; Harro PD retest within 10 days, ON [43, 19, 81] | Composite ICC .90 (.81–.95), SEM 4.2, MDC95 11.6 points. No MIC established. Manufacturer/reference limits are protocol- and age-specific, not community-fall cutoffs | Specialist sensory-reweighting assessment when available; interpret alongside functional balance and completion |
| Romberg / sharpened Romberg: EO/EC timed stance; Steffen 37 community parkinsonism participants, same-rater one-week retest [92] | MDC95: Romberg EO 10 s, EC 19 s; sharpened EO 39 s, EC 19 s. Original full procedural details not verified; do not transfer estimates to an improvised version. Timed ceilings limit headroom | Low-cost task-capacity screen. Use an explicitly documented form, cap and arm/foot rules; limited fine-change measurement |
| Single-leg stance: 27 PD, ON/in-between; EO, both sides, maximum 60 s, up to three trials, best retained; trial-to-trial ICC based on first two [80] | Right ICC .82 (.64–.91), left .83 (.66–.92); no MDC/MIC. Low/high clusters were cross-sectional. Around 10 s is not a validated prospective boundary | Unilateral-support assessment where relevant to goals. Record both legs and termination event; guard and provide stable support |
| Functional reach: forward reach without changing base; 2023 protocol used arms initially at 90°, wall tape, 3 trials averaged, ON [79, 92] | Steffen's different two-trial-average protocol: forward ICC .73/MDC95 9 cm, backward .67/7 cm. No general MIC. Reach and quiet CoP had weak convergence | Controlled leaning, not an overall balance surrogate. Standardize feet, arm, landmark and allowed strategy; do not borrow error across protocols |
| Instrumented LOS: eight directional visual targets; NeuroCom EPE (% limit), or Wii-board maximum CoP displacement (cm) and latency; prescribed foot-contact rules [19, 81, 82] | NeuroCom EPE ICC .87 (.76–.93), SEM 5.0, MDC95 13.8 percentage points. No common MIC. Age influenced Wii-board displacement; WBB study was not a simultaneous device-agreement experiment | Direction-specific voluntary stability. Retain direction and normalization; do not equate maximum CoP excursion with manufacturer EPE |
| Coordinated stability / near-tandem: Paul, optimal medication state; waist-swaymeter track 14 mm wide, 29 cm lateral/18 cm AP excursion, better of two; near-tandem EC twice, 10-s cap [1] | Path departures score errors; missed corner adds five. Paul 2012: ICC 0.50, SEM 12 points; without disabling dyskinesia 0.67/5 points. Separate EO narrow-base score sums three 10-s positions. No MIC [29] | Controlled multidirectional leaning and graded narrow support. Reproduce geometry/scoring; a digital remake needs validation, not inherited model performance |
| Pull / push-and-release: Jacobs compared first/third of three trials, EO; medication cycle uncontrolled; original active-push version later modified toward passive lean [83] | Independent same-day examiners, 8 PD + 3 controls: push-release ICC .84/.83 versus pull .45/.74. No CI/MDC/MIC; historical falls only | Trained reactive assessment. Preserve scoring version: Jacobs 2016 uses BESTest P&R 0–3 and warned/practised MDS-UPDRS pull 0–4, with opposite score direction. No unsupported remote testing [5] |
| Motor Control Test: NeuroCom anterior/posterior translations, three perturbation magnitudes; ON, repeat within 10 days; harness [19, 81] | Latency ICC .92 (.84–.96), SEM 2.7 ms, MDC95 7.4 ms. No MIC; high repeatability does not establish fall prediction | Specialist automatic-response quantification. Its error values do not apply to manual pull/release tests |
| Mini-BESTest: current 14-item, 28-point form; Löfgren 27 PD H&Y 2–3, ON, independent administrations and same-rater one-week retest [18, 94] | Retest ICC .80 (.60–.90), SEM 1.2, SRD95 3.4 points; inter-rater SRD 4.1. Approximately 4-point important improvement supported in one four-week programme; abstract-verified [90] | Defensible multidomain clinical comparator in ambulant PD. Use official side rules and practice reactive administration; one-point changes need caution |
| Berg / FAB / full BESTest: use complete specified forms; BBS /56, FAB /40; full BESTest for broader system assessment [92, 27, 94] | BBS one-week ICC .94, MDC95 5 points in parkinsonism. In 85 PD, ceiling BBS 15/85, FAB 2/85, Mini 1/85. Shared-item administration and population affect comparisons | Berg for basic functional limitation; Mini/FAB provide more upper-end challenge; full BESTest when a more extensive profile is useful |
| Home waist phone: Roche early drug-naïve PD; front-waist belt, arms at sides, 30 s; median of ≥3 observations in each two-week period [17] | Aggregate jerk ICC .82 (.76–.87); postural item r .14; no H&Y I/II separation, control norms or MIC. Not a single-trial reliability estimate | Selected early-PD remote monitoring feasibility. Match schedule, placement and aggregation; validate broader-stage safety and measurement |
| Mobile accelerometry / Sway: Ozinga sacral-waist iPad during SOT, 100 Hz/3.5-Hz filter; Fiems commercial app, 30 PD, one-week retest [95, 96] | Ozinga firm-EO ML acceleration normalized path length ICC .83 (.68–.91), pooled PD/controls; nonzero method bias. Fiems ICC .72/.92, abstract-only; neither establishes a class-wide phone MDC/MIC | Device- and output-specific evidence only. Chest, hand and waist placements or proprietary scores need separate validation |
| Camera balance: Kaya RGB-depth HoloLens, single 30-s challenged stances, hands on hips, OFF ≥12 h, spotter; Ferraris RGB-only alignment versus Kinect, 60-s EO/EC, ON [84, 85] | Kaya ML range/area met ±5% equivalence versus Vicon; AP did not. 9% technical errors; full-trial selection. Ferraris examines alignment, not CoP. No between-day MDC/MIC | Promising supervised measurement development. Ordinary webcam/pose-model substitution requires new validation; retain tracking failure separately from loss of balance |
See the narrative for source-access limits, protocol discrepancies, medication effects, non-PD methodological evidence and task-completion rules. Match hardware, task and processing before transferring a threshold.
Implications for rehabilitation measurement
A transparent assessment record
For rehabtools, the immediate opportunity is to make the assessment reproducible and interpretable. Each output should identify its task and quantity, retain units and direction, and record stance width and foot position, footwear, surface, visual condition, instructions, duration, repetitions, sensor placement or camera geometry, processing version, medication state and time relative to dosing. Those fields support comparison with the protocol-specific evidence above.
Retain protective steps, assistance, interrupted or technically invalid trials, and inability to complete a condition. A person who can no longer perform a task may have an important change even when no valid continuous sway value exists. Do not convert missing or failed tests into reassuring normal scores. Challenging stance and perturbations require appropriate trained supervision and fall protection; laboratory testing does not establish safety for unsupervised home use.
Interpreting a result and a change
Show domain-specific findings before constructing a summary score. Quiet-standing sway, voluntary limits of stability, sensory challenge and reactive stepping can disagree because they examine different demands. Interpret repeated results against applicable absolute error and meaningful-change estimates, while documenting changes in medication, dyskinesia, fatigue, attention or instructions. Where an appropriate threshold has not been established, report the observed change without labelling it clinically important.
Device substitution requires validation of the reported quantity under the intended protocol. Centre-of-pressure displacement, trunk acceleration, estimated centre-of-mass movement and camera-derived alignment are not interchangeable. A reference distribution, a within-person change threshold and a future-risk cutoff should remain separately labelled.
Requirements for a future risk output
A numerical prognosis should use a fixed measurement and prediction pipeline evaluated in the intended patients and setting, at a named horizon. It needs comparison with an appropriate clinical baseline, participant-level validation, calibration, uncertainty and transparent missing-data handling. If the score is intended to change care, its clinical impact must also be evaluated. Until those requirements are met, the stronger claim is a well-characterised assessment of performance and change, with the precise benefits and limitations of each task made visible.
Abbreviations
AIC, Akaike information criterion; FES-I, Falls Efficacy Scale–International; ABC, Activities-specific Balance Confidence Scale; EC, eyes closed; EO, eyes open; EPE, endpoint excursion; FAB, Fullerton Advanced Balance scale; MCT, Motor Control Test; MIC, minimal important change; ADL, activities of daily living; AP, anteroposterior; APA, anticipatory postural adjustment; AUC, area under the receiver operating characteristic curve; BBS, Berg Balance Scale; BESTest, Balance Evaluation Systems Test; CI, confidence interval; CoM, centre of mass; CoP, centre of pressure; CSF, cerebrospinal fluid; DBS, deep brain stimulation; FOG, freezing of gait; H&Y, Hoehn and Yahr; HR, hazard ratio; ICC, intraclass correlation coefficient; IMU, inertial measurement unit; LOS, limits of stability; LAPAQ, Longitudinal Aging Study Amsterdam Physical Activity Questionnaire; LOOCV, leave-one-out cross-validation; LR, likelihood ratio; mCTSIB, modified Clinical Test of Sensory Interaction on Balance; MDC, minimal detectable change; MDS, Movement Disorder Society; ML, mediolateral; MMSE, Mini-Mental State Examination; MoCA, Montreal Cognitive Assessment; OR, odds ratio; PD, Parkinson’s disease; PIGD, postural instability and gait difficulty; RMS, root mean square; ROC, receiver operating characteristic; SD, standard deviation; SEM, standard error of measurement; SOT, Sensory Organization Test; SRM, standardised response mean; STN, subthalamic nucleus; TUG, Timed Up and Go; UPDRS, Unified Parkinson’s Disease Rating Scale.
References
Numbering follows first citation. DOI links identify original articles. The matrices distinguish complete primary-text access, partial primary-text access and abstract-only evidence; supplementary material was not uniformly available.
1. Paul SS, Sherrington C, Canning CG, et al. The relative contribution of physical and cognitive fall risk factors in people with Parkinson's disease: a large prospective cohort study. Neurorehabilitation and Neural Repair. 2014. DOI 10.1177/1545968313508470
Source note: SRC-420a8a49c874 Paul SS 2014
2. Lindholm B, Franzén E, Duzynski W, et al. Clinical Usefulness of Retropulsion Tests in Persons with Mild to Moderate Parkinson's Disease. International Journal of Environmental Research and Public Health. 2021. DOI 10.3390/ijerph182312325
Source note: SRC-d81b1a105d0f Lindholm B 2021
3. Schlenstedt C, Brombacher S, Hartwigsen G, Weisser B, Möller B, Deuschl G. Comparison of the Fullerton Advanced Balance Scale, Mini-BESTest, and Berg Balance Scale to Predict Falls in Parkinson Disease. Physical Therapy. 2016;96(4):494–501. DOI 10.2522/ptj.20150249
Source note: SRC-c65e098a32f1 Schlenstedt C 2016
4. Mak MKY, Auyeung MM. The mini-BESTest can predict parkinsonian recurrent fallers: a 6-month prospective study. Journal of Rehabilitation Medicine. 2013. DOI 10.2340/16501977-1144
Source note: SRC-09ed19718487 Mak MKY 2013
5. Jacobs JV, Earhart GM, McNeely ME. Can postural instability tests improve the prediction of future falls in people with Parkinson’s disease beyond knowing existing fall history? Journal of Neurology. 2016;263(1):133–139. DOI 10.1007/s00415-015-7950-x
Source note: SRC-8776007319e0 Jacobs JV 2016
6. Kerr GK, Worringham CJ, Cole MH, Lacherez PF, Wood JM, Silburn PA. Predictors of future falls in Parkinson disease. Neurology. 2010;75:116–124. DOI 10.1212/WNL.0b013e3181e7b688
Source note: SRC-1e5e4341a16b Kerr GK 2010
7. Gervasoni E, Cattaneo D, Messina P, et al. Clinical and stabilometric measures predicting falls in Parkinson disease/parkinsonisms. Acta Neurologica Scandinavica. 2015;132:235–241. DOI 10.1111/ane.12388
Source note: SRC-6b8028cff8f7 Gervasoni E 2015
8. Moraca GAG, Beretta VS, dos Santos PCR, Nóbrega-Sousa P, Orcioli-Silva D, Vitório R, Gobbi LTB. Center of pressure responses to unpredictable external perturbations indicate low accuracy in predicting fall risk in people with Parkinson’s disease. European Journal of Neuroscience. 2021;53:2901–2911. DOI 10.1111/ejn.15143
Source note: SRC-6a18bfc3359c Moraca GAG 2021
9. Matinolli M, Korpelainen JT, Sotaniemi KA, Myllylä VV, Korpelainen R. Recurrent falls and mortality in Parkinson’s disease: a prospective two-year follow-up study. Acta Neurologica Scandinavica. 2011;123:193–200. DOI 10.1111/j.1600-0404.2010.01386.x
Source note: SRC-9d3202d37ff7 Matinolli M 2011
10. Urso, Daniele; Leta, Valentina; Batzu, Lucia et al. Disentangling the PIGD classification for the prediction of cognitive impairment in de novo Parkinson's disease. Journal of neurology. 2021. DOI 10.1007/s00415-021-10730-3
Source note: SRC-502021f2c9fb Urso 2021
11. Laurent, Louise; Koskas, Pierre; Estrada, Janina et al. Tinetti balance performance is associated with mortality in older adults with late-onset Parkinson's disease: a longitudinal study. BMC geriatrics. 2023. DOI 10.1186/s12877-023-03776-7
Source note: SRC-17da1ae149c6 Laurent 2023
12. Lindh-Rengifo, Magnus; Jonasson, Stina B; Ullén, Susann et al. Perceived walking difficulties in Parkinson's disease - predictors and changes over time. BMC geriatrics. 2021. DOI 10.1186/s12877-021-02113-0
Source note: SRC-4c4c834d9ac5 Lindh-Rengifo 2021
13. Bowman, Thomas; Gervasoni, Elisa; Parelli, Riccardo et al. Predictors of mobility domain of health-related quality of life after rehabilitation in Parkinson's disease: a pilot study. Archives of physiotherapy. 2018. DOI 10.1186/s40945-018-0051-2
Source note: SRC-88decc5d5802 Bowman 2018
14. Pantall, Annette; Suresparan, Piriya; Kapa, Leanne et al. Postural Dynamics Are Associated With Cognitive Decline in Parkinson's Disease. Frontiers in neurology. 2018. DOI 10.3389/fneur.2018.01044
Source note: SRC-bc2378102161 Pantall 2018
15. Dewey DC, Chitnis S, McCreary MC, et al. APDM gait and balance measures fail to predict symptom progression rate in Parkinson's Disease. Frontiers in Neurology. 2022. DOI 10.3389/fneur.2022.1041014
Source note: SRC-817facd83053 Dewey DC 2022
16. Mancini M, Salarian A, Carlson‐Kuhta P, et al. ISway: a sensitive, valid and reliable measure of postural control. Journal of NeuroEngineering and Rehabilitation. 2012. DOI 10.1186/1743-0003-9-59
Source note: SRC-1d77ab0e8733 Mancini M 2012
17. Lipsmeier F, Taylor KI, Postuma RB, et al. Reliability and validity of the Roche PD Mobile Application for remote monitoring of early Parkinson's Disease. Scientific Reports. 2022. DOI 10.1038/s41598-022-15874-4
Source note: SRC-3cc9d68842c3 Lipsmeier F 2022
18. Löfgren N, Lenholm E, Conradsson D, et al. The Mini-BESTest: a clinically reproducible tool for balance evaluations in mild to moderate Parkinson's disease? BMC Neurology. 2014. DOI 10.1186/s12883-014-0235-7
Source note: SRC-0c035b731273 Lofgren N 2014
19. Harro CC, Marquis A, Piper N, Burdis C. Reliability and Validity of Force Platform Measures of Balance Impairment in Individuals With Parkinson Disease. Physical Therapy. 2016;96:1955–1964. DOI 10.2522/ptj.20160099
Source note: SRC-d40263ddaf26 Harro CC 2016
20. Riley RD, Archer L, Snell KIE, et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ. 2024;384:e074820. DOI 10.1136/bmj-2023-074820
Source note: SRC-e131c017d8a9 Riley RD 2024
21. Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. DOI 10.1136/bmj-2024-082505
Source note: SRC-1edddd770a90 Moons KGM 2025
22. Almeida LRS, Valença GT, Negreiros NN, Pinto EB, Oliveira-Filho J. Comparison of Self-report and Performance-Based Balance Measures for Predicting Recurrent Falls in People With Parkinson Disease: Cohort Study. Physical Therapy. 2016;96(7):1074–1084. DOI 10.2522/ptj.20150168
Source note: SRC-625a8074f6af Almeida LRS 2016
23. Sebastiá-Amat S, Tortosa-Martínez J, Manders RJF, Pueo B. Discriminative power of static posturography for identifying recurrent fallers in Parkinson’s disease patients. Disability and Rehabilitation. 2026;48:860–871 (online August 2025). DOI 10.1080/09638288.2025.2545598
Source note: SRC-fd12ba8b1758 Sebastia-Amat S 2026
24. Beretta VS, Barbieri FA, Orcioli-Silva D, dos Santos PCR, Simieli L, Vitório R, Gobbi LTB. Can Postural Control Asymmetry Predict Falls in People With Parkinson’s Disease? Motor Control. 2018. DOI 10.1123/mc.2017-0033
Source note: SRC-da150b097101 Beretta VS 2018
25. Matinolli M. Balance, mobility and falls in Parkinson’s disease. Doctoral dissertation. Acta Universitatis Ouluensis D Medica 1027. University of Oulu; 2009. Study IV; pp 51–59, 69–74. ISBN 978-951-42-9233-0. Doctoral thesis
Source note: SRC-86f0cf135991 Matinolli M 2009
26. Grimbergen YAM. Falls in Parkinson’s disease and Huntington’s disease. Doctoral thesis, Leiden University. 2012. Chapter 3.1 reproduces Bloem et al. J Neurol. 2001;248:950–958. Doctoral thesis
Source note: SRC-bef3e15f921c Grimbergen YAM 2012
27. Schlenstedt C, Brombacher S, Hartwigsen G, Weisser B, Möller B, Deuschl G. Comparing the Fullerton Advanced Balance Scale With the Mini-BESTest and Berg Balance Scale to Assess Postural Control in Patients With Parkinson Disease. Arch Phys Med Rehabil. 2015;96(2):218–225. Primary manuscript inspected in Schlenstedt doctoral thesis (2015), Study 1. DOI 10.1016/j.apmr.2014.09.002
Source note: SRC-d7e6e75d5759 Schlenstedt C 2015
28. Fiems CL, Miller SA, Buchanan N, Knowles E, Larson E, Snow R, Moore ES. Does a Sway-Based Mobile Application Predict Future Falls in People With Parkinson Disease? Archives of Physical Medicine and Rehabilitation. 2020;101:472–478 (online October 2019). DOI 10.1016/j.apmr.2019.09.013
Source note: SRC-a19dae1b0d0a Fiems CL 2020
29. Paul SS, Canning CG, Sherrington C, et al. Reproducibility of measures of leg muscle power, leg muscle strength, postural sway and mobility in people with Parkinson's Disease. Gait & Posture. 2012. DOI 10.1016/j.gaitpost.2012.04.013
Source note: SRC-b3a13e915936 Paul SS 2012
30. Duncan RP, Leddy AL, Cavanaugh JT, et al. Accuracy of fall prediction in Parkinson disease: six-month and 12-month prospective analyses. Parkinson's Disease. 2012. DOI 10.1155/2012/237673
Source note: SRC-da91572d7a6e Duncan RP 2012
31. Duncan RP, Leddy AL, Cavanaugh JT, et al. Comparative utility of the BESTest, mini-BESTest, and brief-BESTest for predicting falls in individuals with Parkinson disease: a cohort study. Physical Therapy. 2013. DOI 10.2522/ptj.20120302
Source note: SRC-35d8cfbea35a Duncan RP 2013
32. Jacobs JV, Horak FB, Tran VK, et al. Multiple balance tests improve the assessment of postural stability in subjects with Parkinson's Disease. Journal of Neurology, Neurosurgery & Psychiatry. 2006. DOI 10.1136/jnnp.2005.068742
Source note: SRC-6b9e2276ae5a Jacobs JV 2006
33. Bloem BR, Grimbergen YA, Cramer M, et al. Prospective assessment of falls in Parkinson's Disease. Journal of Neurology. 2001. DOI 10.1007/s004150170047
Source note: SRC-d3e1d2f83322 Bloem BR 2001
34. Almeida LRS, Sherrington C, Allen NE, et al. Disability is an Independent Predictor of Falls and Recurrent Falls in People with Parkinson's Disease Without a History of Falls: A One-Year Prospective Study. Journal of Parkinson's Disease. 2015. DOI 10.3233/JPD-150651
Source note: SRC-1d84550669cd Almeida LRS 2015
35. Lindholm B, Hagell P, Hansson O, et al. Prediction of falls and/or near falls in people with mild Parkinson's Disease. PLOS ONE. 2015. DOI 10.1371/journal.pone.0117018
Source note: SRC-7830fa969260 Lindholm B 2015
36. Valkovic P, Brozová H, Bötzel K, et al. Push-and-release test predicts Parkinson fallers and nonfallers better than the pull test: comparison in OFF and ON medication states. Movement Disorders. 2008. DOI 10.1002/mds.22131
Source note: SRC-05640ef1c6ac Valkovic P 2008
37. Brandmeir NJ, Brandmeir CL, Kuzma K, et al. A Prospective Evaluation of an Outpatient Assessment of Postural Instability to Predict Risk of Falls in Patients with Parkinson's Disease Presenting for Deep Brain Stimulation. Movement Disorders Clinical Practice. 2016. DOI 10.1002/mdc3.12257
Source note: SRC-f64e6ba1f369 Brandmeir NJ 2016
38. Godi M, Arcolin I, Leavy B, et al. Insights Into the Mini-BESTest Scoring System: Comparison of 6 Different Structural Models. Physical Therapy. 2021. DOI 10.1093/ptj/pzab180
Source note: SRC-30d00cab1028 Godi M 2021
39. Jansen JAF, Tosserams A, Weerdesteyn VGM, et al. The 'Pants-Sign': A Predictor for Falling in People with Parkinson's Disease? Journal of Parkinson's Disease. 2023. DOI 10.3233/jpd-230353
Source note: SRC-97b1c9f8a782 Jansen JAF 2023
40. Castro IPR, Valença GT, Pinto EB, Cavalcanti HM, Oliveira-Filho J, Almeida LRS. Predictors of Falls with Injuries in People with Parkinson's Disease. Mov Disord Clin Pract. 2023;10(2):258–268. Published online 25 December 2022. DOI 10.1002/mdc3.13636
Source note: SRC-c28763e4d81c Castro IPR 2023
41. Matinolli M, Korpelainen JT, Korpelainen R, Sotaniemi KA, Virranniemi M, Myllylä VV. Postural sway and falls in Parkinson’s disease: a regression approach. Movement Disorders. 2007;22:1927–1935. DOI 10.1002/mds.21633
Source note: SRC-f3355a3cb54c Matinolli M 2007
42. Johnson L, James I, Rodrigues J, Stell R, Thickbroom G, Mastaglia F. Clinical and posturographic correlates of falling in Parkinson’s disease. Movement Disorders. 2013. DOI 10.1002/mds.25449
Source note: SRC-be04038c5e80 Johnson L 2013
43. Freeman L, Gera G, Horak FB, Blackinton MT, Besch M, King L. Instrumented Test of Sensory Integration for Balance: A Validation Study. Journal of Geriatric Physical Therapy. 2018. DOI 10.1519/JPT.0000000000000110
Source note: SRC-9ce346c220e9 Freeman L 2018
44. Rivera M. Examining Sensory Systems That Contribute to Falls in Parkinson Disease Using Computerized Dynamic Posturography: Secondary Analysis. JMIR Formative Research. 2026;e84659. DOI 10.2196/84659
Source note: SRC-65a20fc0405c Rivera M 2026
45. Bradley M, O’Loughlin S, Donlon E, et al. Determining Falls Risk in People with Parkinson’s Disease Using Wearable Sensors: A Systematic Review. Sensors. 2025;25:4071. DOI 10.3390/s25134071
Source note: SRC-a17d8c624ef1 Bradley M 2025
46. Tsai CL, Lai YR, Lien CY, et al. Feasibility of Combining Disease-Specific and Balance-Related Measures as Risk Predictors of Future Falls in Patients with Parkinson’s Disease. Journal of Clinical Medicine. 2023;12:127 (online December 2022). DOI 10.3390/jcm12010127
Source note: SRC-af171bae902d Tsai CL 2023
47. Sturchio A, Dwivedi AK, Marsili L, et al. Kinematic but not clinical measures predict falls in Parkinson-related orthostatic hypotension. Journal of Neurology. 2021 (online 2020). DOI 10.1007/s00415-020-10240-8
Source note: SRC-813eb9328c34 Sturchio A 2021
48. Sebastiá-Amat S. Propuesta de evaluación motriz en personas con enfermedad de Parkinson. Doctoral dissertation. Universidad de Alicante; 2021. Study 3; pp 244–261. Deposited 2024. Doctoral thesis
Source note: SRC-56503c36b92b Sebastia-Amat S 2021
49. Sotirakis C, Brzezicki MA, Patel S, Conway N, FitzGerald JJ, Antoniades CA. Predicting future fallers in Parkinson’s disease using kinematic data over a period of 5 years. npj Digital Medicine. 2024. DOI 10.1038/s41746-024-01311-5
Source note: SRC-683a44e25c70 Sotirakis C 2024
50. Hoskovcová M, Dušek P, Sieger T, et al. Predicting Falls in Parkinson Disease: What Is the Value of Instrumented Testing in OFF Medication State? PLoS ONE. 2015;10:e0139849. DOI 10.1371/journal.pone.0139849
Source note: SRC-90b7fd743cae Hoskovcova M 2015
51. Apthorp D, Smith A, Ilschner S, et al. Postural sway correlates with cognition and quality of life in Parkinson's Disease. BMJ Neurology Open. 2020. DOI 10.1136/bmjno-2020-000086
Source note: SRC-f942ec150a64 Apthorp D 2020
52. Pantall, Annette; Del Din, Silvia; Rochester, Lynn. Longitudinal changes over thirty-six months in postural control dynamics and cognitive function in people with Parkinson's disease. Gait & posture. 2018. DOI 10.1016/j.gaitpost.2018.04.016
Source note: SRC-defc42472a9b Pantall 2018
53. Keener AM, Paul KC, Folle A, et al. Cognitive Impairment and Mortality in a Population-Based Parkinson's Disease Cohort. Journal of Parkinson's Disease. 2018. DOI 10.3233/jpd-171257
Source note: SRC-4a346f16cbc5 Keener AM 2018
54. Bäckström DC, Granåsen G, Mo SJ, et al. Prediction and early biomarkers of cognitive decline in Parkinson disease and atypical parkinsonism: a population-based study. Brain Communications. 2022. DOI 10.1093/braincomms/fcac040
Source note: SRC-e6c1e0c8b7e3 Backstrom D 2022
55. Malfer, Lorenzo; Mullan, Aidan F; Lynott, Elena et al. Cumulative incidence and risk factors of dementia in Parkinson's disease and parkinsonism: a population-based study. Brain communications. 2026. DOI 10.1093/braincomms/fcag330
Source note: SRC-6b3eaee44500 Malfer 2026
56. Mancini M, Carlson‐Kuhta P, Zampieri C, et al. Postural sway as a marker of progression in Parkinson's disease: A pilot longitudinal study. Gait & Posture. 2012. DOI 10.1016/j.gaitpost.2012.04.010
Source note: SRC-76bb8367dc00 Mancini M 2012
57. Sotirakis C, Su Z, Brzezicki MA, et al. Identification of motor progression in Parkinson's disease using wearable sensors and machine learning. npj Parkinson's Disease. 2023. DOI 10.1038/s41531-023-00581-2
Source note: SRC-626258f8a85b Sotirakis C 2023
58. Duncan RP, Leddy AL, Cavanaugh JT, et al. Detecting and predicting balance decline in Parkinson disease: a prospective cohort study. Journal of Parkinson's Disease. 2015. DOI 10.3233/jpd-140478
Source note: SRC-2bbca6d4ba77 Duncan RP 2015
59. Hou W, Wu F, Wang Y, et al. Predicting slight freezing of gait in Parkinson's disease with anticipatory postural adjustments and limits of stability. Parkinsonism & Related Disorders. 2024. DOI 10.1016/j.parkreldis.2024.106949
Source note: SRC-c0360b86ad3a Hou W 2024
60. D'Cruz N, Vervoort G, Fieuws S, et al. Repetitive Motor Control Deficits Most Consistent Predictors of Conversion to Freezing of Gait in Parkinson's Disease: A Prospective Cohort Study. Journal of Parkinson's Disease. 2020. DOI 10.3233/jpd-191759
Source note: SRC-4c57d2069017 DCruz N 2020
61. Faerman, Michelle V; Cole, Cayli; Van Ooteghem, Karen et al. Motor, affective, cognitive, and perceptual symptom changes over time in individuals with Parkinson's disease who develop freezing of gait. Journal of neurology. 2025. DOI 10.1007/s00415-025-13034-y
Source note: SRC-b62f6609aee6 Faerman 2025
62. Shao, Jing-Yu; Wang, Meng-Yun; Li, Rong et al. A prediction model for the walking and balance milestone in Parkinson's disease. Parkinsonism & related disorders. 2024. DOI 10.1016/j.parkreldis.2024.107175
Source note: SRC-430d1fa79327 Shao 2024
63. Macleod, Angus D; McLernon, David J; Camacho, Marta et al. Prognosis in Parkinson's Disease: An Individual Patient Data Meta-Analysis of Six European Incidence Cohorts. Movement disorders : official journal of the Movement Disorder Society. 2026. DOI 10.1002/mds.70303
Source note: SRC-029bdc60220c Macleod 2026
64. Santos García, Diego; de Deus Fonticoba, Teresa; Cores Bartolomé, Carlos et al. Predictors of Loss of Functional Independence in Parkinson's Disease: Results from the COPPADIS Cohort at 2-Year Follow-Up and Comparison with a Control Group. Diagnostics (Basel, Switzerland). 2021. DOI 10.3390/diagnostics11101801
Source note: SRC-c78f64f822b7 Santos Garcia 2021
65. Shulman, Lisa M; Gruber-Baldini, Ann L; Anderson, Karen E et al. The evolution of disability in Parkinson disease. Movement disorders : official journal of the Movement Disorder Society. 2008. DOI 10.1002/mds.21879
Source note: SRC-c113acc13030 Shulman 2008
66. Ribeiro JA, Camacho M, Scott KM, et al. Validation of a 5‐Year Prognostic Model for Parkinson's Disease. Movement Disorders Clinical Practice. 2024. DOI 10.1002/mdc3.14215
Source note: SRC-190d0dfa1271 Ribeiro JA 2024
67. Lo CSY, Arora SS, Baig F, et al. Predicting motor, cognitive & functional impairment in Parkinson's. Annals of Clinical and Translational Neurology. 2019. DOI 10.1002/acn3.50853
Source note: SRC-b4bd66ae9c21 Lo C 2019
68. Marras, Connie; McDermott, Michael P; Rochon, Paula A et al. Predictors of deterioration in health-related quality of life in Parkinson's disease: results from the DATATOP trial. Movement disorders : official journal of the Movement Disorder Society. 2008. DOI 10.1002/mds.21853
Source note: SRC-ace7e1010ab2 Marras 2008
69. Gray, William K; Hildreth, Anthony; Bilclough, Julie A et al. Physical assessment as a predictor of mortality in people with Parkinson's disease: a study over 7 years. Movement disorders : official journal of the Movement Disorder Society. 2009. DOI 10.1002/mds.22610
Source note: SRC-613418d23812 Gray 2009
70. Bäckström DC, Granåsen G, Domellöf M, et al. Early predictors of mortality in parkinsonism and Parkinson disease. Neurology. 2018. DOI 10.1212/wnl.0000000000006576
Source note: SRC-8ae8f490fe88 Backstrom D 2018
71. Ygland Rödström, Emil; Puschmann, Andreas. Clinical classification systems and long-term outcome in mid- and late-stage Parkinson's disease. NPJ Parkinson's disease. 2021. DOI 10.1038/s41531-021-00208-4
Source note: SRC-606511e99fb7 Ygland Rodstrom 2021
72. Aarsland, D; Larsen, J P; Tandberg, E et al. Predictors of nursing home placement in Parkinson's disease: a population-based, prospective study. Journal of the American Geriatrics Society. 2000. DOI 10.1111/j.1532-5415.2000.tb06891.x
Source note: SRC-b23bc9544295 Aarsland 2000
73. Löfgren, Niklas; Conradsson, David; Joseph, Conran et al. Factors Associated With Responsiveness to Gait and Balance Training in People With Parkinson Disease. Journal of neurologic physical therapy : JNPT. 2019. DOI 10.1097/npt.0000000000000246
Source note: SRC-373dc80717a3 Lofgren 2019
74. Joseph, Conran; Leavy, Breiffni; Franzén, Erika. Predictors of improved balance performance in persons with Parkinson's disease following a training intervention: analysis of data from an effectiveness-implementation trial. Clinical rehabilitation. 2020. DOI 10.1177/0269215520917199
Source note: SRC-2549631808d6 Joseph 2020
75. Klamroth, Sarah; Gaßner, Heiko; Winkler, Jürgen et al. Interindividual Balance Adaptations in Response to Perturbation Treadmill Training in Persons With Parkinson Disease. Journal of neurologic physical therapy : JNPT. 2019. DOI 10.1097/npt.0000000000000291
Source note: SRC-96600ba6788c Klamroth 2019
76. Albrecht F, Johansson H, Poulakis K, et al. Exploring Responsiveness to Highly Challenging Balance and Gait Training in Parkinson's Disease. Movement Disorders Clinical Practice. 2024. DOI 10.1002/mdc3.14194
Source note: SRC-dfc58b44d317 Albrecht F 2024
77. Baudendistel, Sidney T; Earhart, Gammon M. Characteristics of responders to interventions for Parkinson disease: a scoping systematic review. Neurodegenerative disease management. 2025. DOI 10.1080/17582024.2025.2493465
Source note: SRC-296e2d54daff Baudendistel 2025
78. Yin Z, Bai Y, Zou L, et al. Balance response to levodopa predicts balance improvement after bilateral subthalamic nucleus deep brain stimulation in Parkinson’s disease. npj Parkinson's Disease. 2021. DOI 10.1038/s41531-021-00192-9
Source note: SRC-c3c8e1586543 Yin Z 2021
79. Sebastiá-Amat S, Tortosa-Martínez J, Pueo B. The Use of the Static Posturography to Assess Balance Performance in a Parkinson’s Disease Population. International Journal of Environmental Research and Public Health. 2023;20:981. DOI 10.3390/ijerph20020981
Source note: SRC-364eed3c8ab0 Sebastia-Amat S 2023
80. Chomiak T, Pereira FV, Hu B. The single-leg-stance test in Parkinson’s disease. Journal of Clinical Medicine Research. Published online 2014; print 2015. DOI 10.14740/jocmr1878w
Source note: SRC-e8742e71063a Chomiak T 2014
81. Harro CC, Kelch A, Hargis C, DeWitt A. Comparing Balance Performance on Force Platform Measures in Individuals with Parkinson’s Disease and Healthy Adults. Parkinson’s Disease. 2018. DOI 10.1155/2018/6142579
Source note: SRC-784dbed619fa Harro CC 2018
82. Álvarez I, Latorre J, Aguilar M, et al. Validity and sensitivity of instrumented postural and gait assessment using low-cost devices in Parkinson’s disease. Journal of NeuroEngineering and Rehabilitation. 2020. DOI 10.1186/s12984-020-00770-7
Source note: SRC-307322df99bc Alvarez I 2020
83. Jacobs JV, Horak FB, Van Tran K, Nutt JG. An alternative clinical postural stability test for patients with Parkinson’s disease. Journal of Neurology. 2006;253:1404–1413. DOI 10.1007/s00415-006-0224-x
Source note: SRC-1821baecd44f Jacobs JV 2006
84. Kaya RD, Bazyk AS, Waltz C, et al. Using Loss of Balance to Understand Postural Instability Associated With Parkinson’s Disease. Parkinson's Disease. 2026. DOI 10.1155/padi/6998888
Source note: SRC-fc7228d70bd5 Kaya RD 2026
85. Ferraris C, Amprimo G, Olmo G, et al. From RGB-D to RGB-Only: Reliability and Clinical Relevance of Markerless Skeletal Tracking for Postural Assessment in Parkinson’s Disease. Sensors. 2026. DOI 10.3390/s26041146
Source note: SRC-65df0ebe83c3 Ferraris C 2026
86. Carpenter MG, Frank JS, Winter DA, Peysar GW. Sampling duration effects on centre of pressure summary measures. Gait & Posture. 2001;13:35–40. DOI 10.1016/S0966-6362(00)00093-X
Source note: SRC-253821887c5d Carpenter MG 2001
87. van der Kooij H, Campbell AD, Carpenter MG. Sampling duration effects on centre of pressure descriptive measures. Gait & Posture. 2011;34:19–24. DOI 10.1016/j.gaitpost.2011.02.025
Source note: SRC-172e7b63d43e van der Kooij H 2011
88. Chiari L, Rocchi L, Cappello A. Stabilometric parameters are affected by anthropometry and foot placement. Clinical Biomechanics. 2002;17:666–677. DOI 10.1016/S0268-0033(02)00107-9
Source note: SRC-13dac613d294 Chiari L 2002
89. Curtze C, Nutt JG, Carlson-Kuhta P, Mancini M, Horak FB. Levodopa Is a Double-Edged Sword for Balance and Gait in People With Parkinson’s Disease. Movement Disorders. 2015;30:1361–1370. DOI 10.1002/mds.26269
Source note: SRC-a35f6affde86 Curtze C 2015
90. Godi M, Arcolin I, Giardini M, et al. Responsiveness and minimal clinically important difference of the Mini-BESTest in patients with Parkinson’s disease. Gait & Posture. 2020. DOI 10.1016/j.gaitpost.2020.05.004
Source note: SRC-22ffaec88dd3 Godi M 2020
91. Mehdizadeh M, Fereshtehnejad SM, Landers MR, et al. Responsiveness of the mini-balance evaluation systems test, dynamic gait index, Berg balance scale, and performance-oriented mobility assessment in parkinson’s disease. Scientific Reports. 2025. DOI 10.1038/s41598-025-08463-8
Source note: SRC-206e142cb904 Mehdizadeh M 2025
92. Steffen T, Seney M. Test-retest reliability and minimal detectable change on balance and ambulation tests, the 36-item short-form health survey, and the Unified Parkinson Disease Rating Scale in people with parkinsonism. Physical Therapy. 2008;88:733–746. DOI 10.2522/ptj.20070214
Source note: SRC-03d81c10ff0e Steffen T 2008
93. Hasegawa N, Shah VV, Harker GR, et al. Responsiveness of Objective vs. Clinical Balance Domain Outcomes for Exercise Intervention in Parkinson's Disease. Frontiers in Neurology. 2020. DOI 10.3389/fneur.2020.00940
Source note: SRC-02ca14c3cf7e Hasegawa N 2020
94. BESTest. Official test forms and administration resources. Current Mini-BESTest form revised 8 March 2013; accessed 1 October 2026. Official test forms
Source note: SRC-a72e155215a5 BESTest 2013
95. Ozinga SJ, Linder SM, Alberts JL. Use of Mobile Device Accelerometry to Enhance Evaluation of Postural Instability in Parkinson Disease. Archives of Physical Medicine and Rehabilitation. Published online 2016; print 2017. DOI 10.1016/j.apmr.2016.08.479
Source note: SRC-aefc3e5d1552 Ozinga SJ 2016
96. Fiems CL, Dugan EL, Moore ES, Combs-Miller SA. Reliability and validity of the Sway Balance mobile application for measurement of postural sway in people with Parkinson disease. NeuroRehabilitation. 2018;43:147–154. DOI 10.3233/NRE-182424
Source note: SRC-169fc9a4fb9a Fiems CL 2018