← Clinical Evidence

Strength muscle power and rapid force production assessment in Parkinson’s disease

This report reviews strength assessment in Parkinson’s disease. It examines measurement properties, interpretation of change and prognostic evidence, with the limits of each study and testing protocol.

In this report
Audited and Updated

Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.

STR-C01

The abstract reports 127 participants; the main Results and descriptive tables report 129, with 127 in the final model. Mean grip was 26 versus 30 kgf (p=.02). The sample denominator discrepancy should remain explicit.

Type: source denominator conflict. Audit disposition: supported.

STR-C02

Use “sample-defined UPDRS motor cutoff of 31.7” and note the original mean/median labeling conflict if discussing how it was chosen.

Type: source mean median conflict. Audit disposition: supported.

STR-C03

Use39/51 (76.5%; published as77%) in both models. Keep the separately documented clinical-model numerator conflict explicit.

Type: source percentage arithmetic discrepancy. Current audit disposition: matched to archived main-text capture plus separate arithmetic; no live-publisher authentication.

Remaining limit: Archived main-text comparison supports the enumerated model classification fractions. Published percentage and clinical-model numerator conflicts remain; capture fidelity was not independently authenticated against the live publisher original.

STR-C04

Add “A Cross-Sectional Study” after “The Association of Grip Strength With Severity and Duration of Parkinson’s”.

Type: bibliographic completion. Audit disposition: supported.

STR-C05

Update current access notes for Mazza and Almeida, and Gallagher’s quantitative-table completeness. Preserve the historical wording accurately; none of these retrospective access statements is classified as false.

Type: current access update. Audit disposition: supported as access update.

STR-C06

Retain the precise Kerr verification gaps. The eight newly recovered sources support the exact checks recorded in audit evidence record; archived captures/transcriptions have not been independently validated against live publisher originals. Missing Bonde-Jensen/Gamborg supplements and Hammond figure pixels remain explicit, and unreported or unverified extension estimates must not be inferred.

Type: primary evidence gap. Audit disposition: supported as snapshot limit.

Remaining limit: Kerr body detail, stated supplements/figures and archived capture/transcription fidelity remain limited.

Editorial record

  • Audit status: supported as snapshot limit. Kerr body detail, stated supplements/figures and archived capture/transcription fidelity remain limited.
  • Audit status: supported as access update.
  • Audit status: supported.
  • Edited phrase under STR-C02 . Original wording: P0096
  • Audit status: supported as access update.
  • Audit status: supported.
  • Audit status: supported as access update.
  • Current audit status: archived licensed main-text comparison supports both explanatory and combined classification fractions. The 39/51 percentage discrepancy and separate clinical-model numerator conflict remain. Capture fidelity was not independently authenticated against the live publisher original.
  • Audit status: supported as snapshot limit. Kerr body detail, stated supplements/figures and archived capture/transcription fidelity remain limited.
  • Audit status: supported as snapshot limit. Kerr body detail, stated supplements/figures and archived capture/transcription fidelity remain limited.
  • Current audit status: archived licensed main-text comparison supports both explanatory and combined classification fractions. The 39/51 percentage discrepancy and separate clinical-model numerator conflict remain. Capture fidelity was not independently authenticated against the live publisher original.
  • Audit status: supported as snapshot limit. Kerr body detail, stated supplements/figures and archived capture/transcription fidelity remain limited.
  • Audit status: supported as access update.
  • Audit status: supported as snapshot limit. Kerr body detail, stated supplements/figures and archived capture/transcription fidelity remain limited.
  • Audit status: supported.
  • Edited phrase under STR-C04 . Original wording: P0615

Editorial nomenclature update — 4 October 2026 at 11:48:46 am (Australia/Brisbane): authored condition labels and report wording use Parkinson’s disease. Published article titles, exact quotations, recorded searches, identifiers and routes are preserved. This is a terminology edit, not a scientific correction.

Executive conclusions

Strength and power assessment can identify clinically relevant impairments in people with established Parkinson’s disease (PD), but interpretation must remain specific to the muscle group, task, instrument and medication state. Selected grip, dynamometry and machine-power protocols have good relative reliability. That evidence is substantially stronger than the evidence for small individual changes, universal PD-specific cutoffs or prediction of future disability. For rehabtools, the most defensible role is to explain and support a standardized, purpose-specific assessment alongside clinical and functional evaluation.

Five conclusions guide this report.

Report navigation

The report is a critical narrative evidence synthesis prepared for Associate Professor Ross Clark to inform rehabtools. It addresses measurement, interpretation and prognosis in established PD. It is not a formal systematic review, a treatment-effectiveness review or an individualized exercise prescription. Sit-to-stand test procedures and functional chair-rise power are addressed in the companion sit-to-stand report; they are included here only where needed to explain construct boundaries or relevant measurement evidence.

The report moves from definitions and impairment evidence to reproducible protocols, interpretation of change, current function and later outcomes. Tables summarize the comparisons that require direct checking; narrative sections explain their clinical meaning.

Scope and measurement framework

Evidence sources and review coverage

Strength and rapid force impairments

Published assessment protocols

Reliability and clinical change

Medication state and test feasibility

Normalization and asymmetry

Functional associations and future outcomes

Implications for rehabtools and clinical assessment

References

Abbreviations

1RM, one-repetition maximum; AUC, area under the receiver operating characteristic curve; BMI, body mass index; CI, confidence interval; DASH, Disabilities of the Arm Shoulder and Hand questionnaire; DAT, dopamine transporter; EMG, electromyography; HHD, handheld dynamometry; HY, Hoehn and Yahr; ICC, intraclass correlation coefficient; kgf, kilogram-force; LEDD, levodopa-equivalent daily dose; LOA, limits of agreement; MDC, minimal detectable change; MDS-UPDRS, Movement Disorder Society Unified Parkinson’s Disease Rating Scale; MIC, minimal important change; MVC, maximal voluntary contraction; N, newton; Nm, newton metre; OR, odds ratio; PD, Parkinson’s disease; PIGD, postural instability and gait difficulty; RFD, rate of force development; CAR, central activation ratio; ROM, range of motion; RR, risk ratio; RTD, rate of torque development; SD, standard deviation; SDC, smallest detectable change; SEM, standard error of measurement; STS, sit to stand; TD, tremor dominant; TUG, Timed Up and Go; W, watt.

Scope and measurement framework

Start with the decision the measurement must support

A force measurement can be used to describe an impairment, prescribe exercise loading, follow within-person change, classify a present condition or forecast a later event. Evidence for one use does not automatically establish another. A device can rank participants reproducibly but have insufficient precision for a small treatment-related change. A measure can differ between PD and healthy participants without identifying an individual patient accurately. An impairment can matter for rehabilitation even when it adds little to an already strong multivariable prediction model.

This report therefore considers four questions in sequence. What physical or neuromotor construct does the assay measure? How reproducible and interpretable is the result? What current clinical or functional characteristics are associated with it? Does the baseline result predict a genuinely later outcome in established PD? Keeping that sequence explicit prevents a cross-sectional regression coefficient from being presented as a clinical forecasting tool.

The term strength is used for maximal voluntary force or torque under specified task constraints. A hand dynamometer gives external grip force, whereas joint torque also depends on the perpendicular moment arm. A machine 1RM identifies the largest load successfully moved through a prescribed range and is specific to the machine geometry and movement. Mechanical power includes velocity; an isometric contraction may generate force rapidly but produces no external mechanical work when the relevant external displacement is zero. RFD and rate of torque development (RTD) quantify force-time and torque-time slopes. They should retain their analysis window, such as 0–50 ms or 100–200 ms, rather than being reduced to an unexplained label.

Force steadiness during a submaximal hold concerns fluctuations around a target. Relaxation rate concerns how force falls. Motor segmentation concerns interruptions or multiple accelerations within a rapid pulse. These measures may reveal impaired motor organization even when maximum force is relatively preserved, but none is a synonym for strength, power or a generic RFD score. Chung’s visually guided task at 15% MVC and Daniels’ repeated index-finger abduction pulses illustrate how different experimental aims produce different measures. [13, 14]

What each form of validity would establish

Known-groups differences test whether scores behave as expected across groups. Correlation with walking, chair rise or disability tests a hypothesized relation with another construct. Criterion agreement requires an appropriate reference measurement and examination of bias and individual differences. A moderate correlation between two devices cannot show that their values are interchangeable, particularly when one reports pressure and another force. Silva’s modified sphygmomanometer study provides a practical example. [15]

Responsiveness is the validity of a change score for its intended use. A significant group improvement in a training study, a mean ON/OFF medication difference or a statistically detectable mean decline does not by itself establish responsiveness to a patient-important change. The minimal detectable change addresses error under a specified measurement model. The minimal important change addresses importance. These are separate requirements, and an informative assessment may need both.

Prognostic usefulness requires more than an association with a future outcome. A clinically useful prediction model needs appropriate development, internal and preferably external validation, discrimination, calibration and evidence that adding the measure improves the intended decision. A significant adjusted strength coefficient is evidence of an independent statistical association in that model; it is not proof of a causal pathway, a universal threshold or added clinical utility. The later prognosis chapter distinguishes these levels explicitly.

Table 1 Constructs and measurement outputs

Table 1 Constructs and measurement outputs
ConstructTypical outputRequired distinction
Maximum isometric strengthN or device kgf; torque in NmPeak force at a fixed position. Torque needs the perpendicular moment arm; a distal force value alone is not joint torque.
Dynamic maximum strengthMachine 1RM in kg or device forceLargest successful load through specified range. Depends on machine leverage and movement; not interchangeable with single-joint torque.
Mechanical powerW; W/kg when normalizedForce × velocity or torque × angular velocity. State peak versus mean, load and velocity measurement.
RFD and RTDN/s or Nm/s; sometimes MVC-normalizedSlope over a stated force-time interval. Early and late windows, maximal effort and submaximal target control are different outcomes.
Force steadiness and relaxationVariability or %MVC/sFluctuation around a target or decline in force. Distinct from maximum force and force-rise capacity.
Motor segmentationSegments per force pulse and related timingAlgorithm-defined interruptions in a rapid pulse. Depends on filtering, derivatives, pulse amplitude and number of trials.
Functional performance and estimated powerSTS seconds or count; model W or W/kgComposite task performance. Anthropometry-based, encoder and force-platform estimates require separate validation.

Evidence sources and review coverage

Search and source appraisal

The three PubMed queries returned 53 measurement-property, 197 prognosis-related and 65 power/RFD records. Because connector pages overlapped, each query was independently reconciled against official NCBI ESearch identifiers and its records retrieved with EFetch. All 315 query-specific records were recovered, representing 283 unique PubMed records (281 journal articles and two book/report records).

The initial native Scopus query returned 357 unique records across 15 pages. Its separate year arguments had not applied to the raw query, so these results were treated as an all-years search. An explicit native PUBYEAR 2022–2026 query was then run and fully paged: 215 unique records across nine pages, all contained in the broader result set. Deduplication by DOI, or title when a DOI was unavailable, produced 524 records across the reconciled PubMed and Scopus searches. All returned titles were triaged, with focused abstract review and original retrieval for substantive assessment, mechanistic and longitudinal candidates.

Appraisal distinguished reliability, absolute error, construct and criterion validity, responsiveness, concurrent association and genuine future-outcome prognosis. Training effects, incident-PD population risk and sarcopenia composites were considered separately. Full original methods and tables were prioritized; abstract-only information is identified and does not support detailed protocol prescriptions or unreported adjusted estimates. Bibliographic dates distinguish online publication from journal issue dates where these differ.

This targeted search and selection process lacked the independent duplicate screening and exhaustive coverage of a formal systematic review. The queries did not cover every possible synonym, citation searches were provider-limited, and some original full texts remained unavailable. Complete retrieval of these bounded queries should not be confused with complete retrieval of the relevant literature.

What the review establishes

Gamborg and colleagues’ 2023 review provides a broad map of mechanical muscle function in PD, but its evidence ends much earlier than its publication date. Six databases were searched on 14 October 2020. Two authors independently screened and extracted data; 40 studies were included after 188 full texts were assessed. The case-control aim specifically excluded manual muscle testing, handheld dynamometry and one-repetition maximum testing. Later papers discussed in the manuscript did not update that systematic search. Its findings therefore need to be supplemented by direct assessment evidence outside those criteria and by subsequent originals. [16]

The complete main article was examined. Its supplements containing detailed searches, study characteristics, psychometric extraction and item-level quality ratings remain unavailable. This distinction matters because some main-text details require reconciliation against their cited originals.

A broad impairment signal with uneven support

The review combines different muscles, contraction types, velocities, devices, sides and medication states. Multiple actions within a study were averaged; values were divided by body mass when it was reported and otherwise retained in absolute units. Summary results used sample-size weighting, including weighted Hedges g values. The main article does not report conventional heterogeneity statistics, prediction intervals or sensitivity analyses for these aggregates. Its NIH observational-checklist assessments were broadly fair, with little blinding, limited confounder adjustment and predominantly absent or short follow-up. [16]

Lower-extremity strength averaged approximately 75% of healthy-control performance ON and 59% OFF, but these categories came from different study sets. Their contrast is not a pooled within-person medication effect. Concentric deficits were generally larger than isometric or eccentric deficits, although some comparison cells had few or no studies. Upper-extremity strength averaged about 85% ON and 87% OFF; the visually verified ratio confidence intervals cross 100% in both categories. The review therefore should not be paraphrased as showing statistically secure impairment in every muscle group or every grip assay. [16]

There were 12 RFD studies but only one power study: Lima and colleagues’ 2016 isokinetic study of ten people with PD and ten controls. The original abstract describes levodopa-naïve Hoehn and Yahr I–II participants, although the review’s figure places the power summary under ON. Levodopa-naïve does not establish the absence of other PD medication, so the precise state label still needs the full original methods. Its approximate power values of 60% of control performance for the lower limbs and 20% for the trunk arise from this one study and cannot establish a medication effect or be transferred to pneumatic-machine or chair-rise power. The range of RFD tasks also prevents a single summary deficit from specifying a clinical RFD protocol. [16, 17]

Reliability and association are narrower conclusions

Only two included studies contributed psychometric evidence: Purser 1999 and a Pang/Mak source. The review reports high isometric ICCs, but no validity or responsiveness evidence and no RFD or power measurement-property studies. Its Pang/Mak citation mismatch is confirmed: reference 23 identifies the hip-bone-density paper in 34 women, whose original describes handheld hip/knee testing with ICCs .87–.90. The review instead describes a 43-person trunk/grip study and ICCs .97–.98. A separate Pang/Mak trunk-bone-density paper has 43 PD participants and 29 controls, matching the sample description, but its full methods remain unavailable. The review's trunk/grip ICCs therefore remain unverified against the likely intended original. Dynamic 1RM familiarization, Keiser power error and later instrument-specific studies remain essential additions. [16, 1, 2, 18, 19, 20]

The 11 association studies predominantly linked lower-limb strength to concurrent function or severity. Summary squared associations with function were .24 (95% CI .17–.31) ON and .18 (−.02 to .37) OFF. This is a reported summary bound, not negative explained variance. The ON estimates for total and motor UPDRS were approximately .39 and .40. These are weighted summaries of mainly simple associations, not independently validated prognostic models or unique strength contributions after adjustment. Some review table entries preserve a negative sign after squaring correlations; their inverse direction should be stated in words, without describing a negative coefficient of determination. RFD–function evidence came from only two reports from the same research group and was weak. [16]

Original source checks remain decisive

The recovered review strengthens the need for careful source matching. Its table gives extreme PD/control RFD ratios of 3–8% for Pereira 2018. The original figure does depict very small PD RFD bars, while the same paper’s narrative and statistical summary report no PD deficit relative to the older groups. This inconsistency is not resolved by citing the review and should not be turned into a definitive numerical impairment estimate. Similarly, the review’s Pääsuke 2002 reference identifies a different title from the separately mapped PD original; those records should not be silently merged. [16, 21, 22]

Mazza and colleagues’ 2024 observational handgrip review provides additional context: its abstract identifies 27 articles, including 16 PD–control comparisons, seven of which reported significant differences. This count is not a pooled effect or evidence of equivalence in the remaining studies. Its emphasis on heterogeneous protocols and sparse longitudinal work is consistent with treating grip as a potentially useful general-capacity measure whose disease-specific interpretation requires further evidence. The full review body remains unverified. [23]

The combined evidence supports heterogeneous impairment and task-specific functional associations. It leaves major gaps in absolute error, meaningful change, advanced PD and externally validated prognosis. The reviews identify those gaps; the original studies determine which measurement and clinical claims can safely be made.

Strength and rapid force impairments in established Parkinson’s disease

Handgrip gives useful but incomplete information

Grip is feasible in many settings and can document a meaningful upper-limb impairment. Its apparent simplicity should not obscure what affects the value: hand size, age, sex, body size, pain, arthritis, position, dominance, clinical lateralization and voluntary activation. The relevant question is whether the score describes hand-force capacity, general reserve, a sarcopenia component or a PD-specific motor deficit. Those interpretations require different comparison data.

Ingram and colleagues studied 34 people with PD and 68 age- and sex-matched controls using a multidomain upper-limb physiological assessment. Mean grip was 29.4 kg OFF and 29.2 kg ON, compared with 37.8 kg in controls. The reported control-minus-PD differences were 8.4 kg (95% CI 3.7–12.8) OFF and 8.6 kg (3.7–13.1) ON. The mean ON/OFF grip difference was −0.2 kg (−1.8 to 1.1), p=.813. Thus the study supports a group deficit under its dominant-hand Jamar+ protocol, while showing no average medication effect on that particular outcome. Elbow-flexion force and several dexterity or stability tasks behaved differently. [7]

The same study related grip to self-reported upper-limb disability, but grip correlations with Hoehn and Yahr stage were not significant. This is clinically coherent: an impairment can be relevant to a person’s arm and hand function without tracking a broad ordinal disease stage closely. Its proposed assessment battery also allocated a reliability score using evidence across the adult lifespan; that score must not be mistaken for a new PD-specific test–retest experiment for every included measure. Numerous uncorrected correlations and a mostly OFF-first testing order limit stronger inference about independent associations or medication response. [7]

Roberts and colleagues provide complementary evidence from 57 people with PD recruited from a town-based population. With sex adjustment, each additional UPDRS point was associated with 0.30 kg lower grip (95% CI 0.09–0.51 lower), and each Hoehn and Yahr stage with 3.87 kg lower grip (1.21–6.54 lower). After fuller adjustment, the reported association remained significant for Hoehn and Yahr stage (p=.04), but not for UPDRS (p=.09). Duration of disease was not associated. The available primary abstract supports a relationship with present clinical severity and the importance of age and anthropometric adjustment; it cannot establish within-person progression from a cross-sectional comparison. [24]

A 2026 custom strain-gauge study makes the comparator problem especially clear. Among 45 people with PD and 51 controls, dominant-to-dominant and nondominant-to-nondominant grip comparisons were not significant. Within PD, the clinically more-affected side produced 213.68 N versus 232.33 N on the less-affected side, a mean difference of −18.65 N (95% CI −29.82 to −7.49), with Cohen d=.50. Comparing PD’s more-affected hand with controls’ dominant hand produced a significant group difference; comparing that same PD hand with controls’ nondominant hand did not. The more-affected hand was dominant in 22 participants and nondominant in 23. [8]

This does not invalidate grip assessment. It shows that dominance and disease lateralization answer different questions and must be retained separately. The standing, elbow-extended custom-device protocol also differs from a seated Jamar protocol, so its force values and comparisons cannot be transferred directly. The paper reports a positive more-affected peak-force correlation with UPDRS-III (rho=.34, 95% CI .02–.61), which is not in the anticipated direction for a simple weakness-severity argument. That reported sign should be preserved rather than silently reversed or used to assert that lower grip necessarily marks worse motor severity. [8]

Taken together, these studies support grip as a useful component of assessment, with incomplete disease specificity. A normal or stable grip score does not exclude hip, knee, rapid-force, balance or gait impairment. Conversely, a low grip score should prompt interpretation in the context of musculoskeletal disease, nutrition, general health and functional goals rather than automatic attribution to nigrostriatal degeneration.

Lower limb weakness depends on the sampled muscle and task

Salmon and colleagues compared 30 ambulatory people with mild PD with 24 controls across 12 lower-limb muscle groups using handheld dynamometry. Participants with PD were assessed ON medication, reported normal walking and exercised regularly. Average force was 78% of control values, with group values ranging from 67% to 87%; hip adduction and plantarflexion were among the largest deficits. The primary abstract therefore supports examining more than knee extension alone when lower-limb weakness is clinically suspected. It does not establish criterion validity of a device, the precision of change in one person or a validated screening threshold. [25]

The pattern is not uniform across studies. In Skinner and colleagues’ small stabilized-force study, 13 PD participants and 13 controls performed maximal tests and submaximal holds at 5%, 10% and 20% MVC. Selected body-mass-relative forces were lower in PD, including hip flexion (2.0 versus 2.6 N/kg), plantarflexion (1.74 versus 2.64 N/kg) and dorsiflexion (1.9 versus 2.3 N/kg). Greater submaximal force variability addressed control around a target rather than just maximum capacity. The original abstract establishes neither a PD-specific absolute-error threshold nor future-outcome prediction. [26]

Stevens-Lapsley and colleagues supply a more detailed example of the importance of severity and activation. They compared 17 people with PD with 17 matched controls, testing the less-involved, stronger leg optimally ON medication. Six PD participants above the sample-defined UPDRS motor cutoff of 31.7 (the source labels it median in Methods and mean in Results/Table 1) had 89.2 Nm less quadriceps torque than controls (95% CI 38.5–139.9 less). The 11 participants below that threshold differed by only −12.4 Nm (−53.7 to 29.0). For all PD participants combined, the mean difference was −39.5 Nm (−80.1 to 1.1), p=.056. Reporting only the approximately 50% high-severity deficit would give an exaggerated impression of a uniform PD-control difference. [6]

Quadriceps torque and voluntary activation correlated negatively with contemporaneous UPDRS motor score (r=−.67 and −.65). The findings support a contribution of impaired voluntary activation to the measured deficit; they do not identify the amount of structural muscle loss or show that the baseline value predicts subsequent deterioration. The small, sample-defined severity subgroups also make the apparent gradient imprecise. Because the stronger leg was selected for analysis, the study does not describe the complete bilateral burden. [6]

Fatigability illustrates a further interpretive trap. During 30 repeated isokinetic contractions, the higher-severity group increased torque by 24.6% instead of declining. Initial underactivation or progressive recruitment can produce this pattern. It would be misleading to label it superior endurance merely because the conventional fatigue index is low or negative. Subjective fatigue, voluntary activation and mechanical fatigability must remain separate outcomes, and the raw trajectory should be inspected when the result is unexpected. [6]

Isokinetic findings also depend on speed and action. Inkster’s range-averaged concentric hip and knee torque at 45 degrees per second suggested a greater hip than knee deficit in 10 men with mild PD. Frazzitta’s speed-specific peak knee torques at 90, 120 and 180 degrees per second showed a different and largely nonsignificant whole-group PD-control pattern, with more conspicuous differences in selected right-affected subgroups. Small samples, side classification and multiple comparisons make a pooled statement such as “PD causes a fixed percentage loss of knee strength” unjustified. [27, 28]

Axial testing requires its own interpretation. Bridgewater and Sharpe examined trunk range, isometric torque and isoinertial performance using the Isostation B-200. The available abstract describes altered trunk performance in early PD, but detailed sample and measurement-property extraction was unavailable. Trunk rotation or extension is not interchangeable with a lower-limb force test, and reduced trunk range cannot be assumed to be weakness without accounting for rigidity, posture and the testing task. This older evidence supports a distinct axial domain rather than a borrowed reliability threshold. [29]

Dynamic power reflects both force and movement velocity

A power deficit can arise from lower force, lower velocity or both. Allen and colleagues compared 40 people with PD with 40 neurologically normal controls during leg-extension loading and explosive lifting. The primary abstract reports 172 N lower maximum strength (95% CI 28–315 lower) and 124 W lower peak power (32–216 lower) in PD. Reduced velocity was apparent at light-to-medium loads rather than heavy loads. This gives a concrete reason to preserve the load–power relationship instead of reporting a single unexplained watt value. [30]

The implication is not that one relative load has been established as universally optimal for assessment. A light-load test emphasizes movement speed differently from a heavier-load test; the maximum power over several loads asks a different question again. If a participant has a lower 1RM, testing both participants at 40% 1RM also means testing them against different absolute resistance. Force, velocity and load are therefore needed to understand the physiological meaning of the power difference.

The original full body for this case-control study was not available for detailed protocol appraisal. Allen’s later walking/falls article explicitly identifies these same 40 participants; the two papers therefore do not represent independent cohorts. Its group contrasts support the plausibility and load dependence of power impairment, while the separate Paul repeatability study and functional studies address different validation questions. [30, 31, 1]

Rapid force impairment can coexist with preserved maximum force

The rationale for RFD assessment is strongest when the clinical task requires force before maximum force can be reached. That rationale does not establish that every RFD variable is more clinically informative than maximum strength. Early force rise is sensitive to rapid neural drive and the definition of force onset; later slopes increasingly reflect the force capacity available within that task. Relaxation and force-pulse organization introduce additional mechanisms. Evidence should therefore be evaluated by the particular assay and interval.

Earlier experiments already separated accurate force amplitude from the organization of force over time. Stelmach and colleagues used rapidly produced targets at 15%, 30%, 45% and 60% MVC and reported similar across-trial amplitude dispersion but more irregular within-trial force trajectories in PD. Park and Stelmach later used 15%, 35% and 55% MVC targets at preferred and maximal speeds, finding reduced RFD and longer time to peak as amplitude increased. These are submaximal amplitude-scaling experiments, not identical implementations of maximal knee-extension RFD. The original abstracts support this historical distinction, while exact acquisition and error details remain unavailable for transfer to a clinical protocol. [32, 33]

Rose and colleagues compared 13 men with PD with 15 matched controls using maximal knee extension/flexion, steady submaximal torque and powerful extensions. Lower RFD was associated with lower agonist activation; PD also showed increased coactivation and poorer torque steadiness. Corcos and colleagues’ ON/OFF elbow study similarly distinguished maximal strength, force generation, active force return and passive relaxation, with a greater OFF-state extensor than flexor weakness and prominent relaxation prolongation. Together, the available original abstracts support examining generation, stabilization and release separately, rather than attributing every slow force trace to low maximum strength. Numerical clinical thresholds cannot be derived from those summaries. [34, 35]

Hammond and colleagues’ full original reports final analyses of seven people with mild motor-stage PD and six controls during a voluntary and electrically stimulated quadriceps protocol. Voluntary RFD was lower in PD (p=.008, reported d=1.97), as was the voluntary-to-stimulated RFD ratio (p=.004, reported d=2.18), while maximum-force and several activation differences were not detected. This supports a dissociation between rapid and maximal output, but nonsignificant differences do not establish normal maximal strength or preserved peripheral function. The effect sizes cannot be recomputed from the recovered table because the RFD means and dispersion are graphical. The recruitment/exclusion narrative also fails to reconcile with final group and sex counts, although 17 recruited and 13 analyzed is consistent overall. The study supplies no PD-specific repeatability, meaningful-change threshold or longitudinal progression validation. [36]

The 2021 PD-subtype study similarly found lower knee peak force and RFD over 0–50, 0–100 and 0–200 ms in PD than in older controls. However, the tremor-dominant and postural-instability/gait-difficulty groups did not differ significantly from one another. Numerically worse PIGD values must not be presented as validated subtype discrimination. The reported participant counts and analysis degrees of freedom are not fully consistent, and several essential signal-processing details were absent from the retrieved Methods. These limitations matter for both replication and the confidence placed in large standardized subgroup-control effects. [37]

Chung and colleagues show why a rate-control result cannot be generalized across limbs or tasks. Twenty early-stage PD participants, tested on the more-affected side after overnight medication withdrawal, were compared with 21 controls during visually guided pinch and ankle dorsiflexion to 15% MVC. Foot force-rise rate was lower in PD, 44.87 versus 60.93% MVC/s (p=.030), but hand force-rise rate was similar, 36.32 versus 38.32% MVC/s (p=.754). Relaxation was slower for both hand and foot, while MVC did not differ. Acquisition at 50 Hz and a submaximal visual target distinguish this experiment from a high-bandwidth maximal explosive RFD test. The study supports selective rate-control deficits rather than a universal loss of hand force-rise ability. [13]

The grip study published in 2026 provides another counterweight. A selected late RFD window differed between PD’s more-affected hand and controls’ dominant hand, but most rapid-force measures were not different and did not correlate with gait. Peak grip force showed the clearer functional relationships. Exploratory window-specific findings, multiple comparisons and the sensitivity of the result to the comparator hand argue against elevating RFD above peak force by assumption. [8]

Motor segmentation is a promising specialized construct

Howard and colleagues introduced repeatability evidence for segmentation of rapid isometric index-finger abduction pulses in 10 people with PD, comparing their pattern with previously collected older-adult data. The primary abstract describes high reliability and large group differences, but exact coefficients, confidence intervals and absolute errors were not available in the retrieved material. The historical controls were older on average, and the procedure involved approximately 87 pulses. Its feasibility and precision cannot be assumed for a brief clinical substitute using a few contractions. [38]

Daniels and Knight subsequently examined 57 people with PD and 22 age-matched older adults. In the PD group, 39 of 57 participants met the study’s segmentation definition, 68% (95% CI 55–80%). The procedure used rapid index-finger abduction at submaximal amplitudes; a separate Jamar test measured grip. Segmentation was defined from the same force traces whose slowing was then compared, so the evidence is strongest for describing a force-control phenotype. It is weaker evidence for an independently validated classifier and provides no baseline-to-future-outcome prognostic validation. The observed prevalence also belongs to a largely male, mild-to-moderate, ON-medication sample. [14]

The distinction matters for implementation. A segmentation algorithm needs its filtering, derivative window, amplitude eligibility, number of pulses, onset rules and classifier fixed before evaluation. A software revision can alter the phenotype even when the raw force trace is unchanged. For rehabtools, the present value is to explain a potentially informative research construct and the required measurement architecture. It is premature to label a person as having clinically important progression from an unvalidated change in segment count.

Muscle–tendon studies, including Monte and colleagues’ combined torque, electromyography and ultrasound work, further suggest that rapid force production may reflect both neural drive and muscle mechanical behavior. The available abstract describes reduced torque and RTD in both PD limbs, with some activation differences concentrated on the more-affected side. This supports biological plausibility and cautions against treating the less-affected side as a healthy internal control. It does not supply a deployable clinical assay, a causal partition of mechanisms or a prognostic threshold. [39]

Published assessment protocols

Grip dynamometry

The clearest PD-specific repeatability protocol is Villafañe and colleagues’ Jamar hydraulic test: seated, shoulder adducted, elbow at 90 degrees, forearm and wrist neutral, handle position two. Participants completed two or three preliminary contractions, rested ten minutes, then performed three pain-free maximal three-second contractions per hand with one-minute rests. The mean of the three trials was the session score. Testing was repeated one week later, in the morning, with usual medication; no participant exhibited dyskinesia. This protocol should be described as pain-free maximal grip, because a participant who develops pain is explicitly instructed to stop before pain onset. It does not necessarily estimate unrestricted maximal force. [3]

That scoring rule differs from several other PD studies. Saarinen and colleagues used one practice attempt, two three-to-five-second efforts per hand and the better value. Ingram and colleagues used the best of three Jamar+ trials on the dominant side. A 2026 strain-gauge study tested participants standing with the elbow extended, adjusted handle separation to hand dimensions, alternated sides, and selected the stronger of two five-second efforts; RFD was then calculated from that same maximum-force trial. These are not interchangeable implementations. A maximum is sensitive to the number of opportunities, while a mean is sensitive to poorly executed trials. Selecting the strongest MVC does not necessarily select the fastest force rise. [12, 7, 8]

In practice, choose the protocol appropriate to the comparison source before testing. Record make/model, calibration, handle setting, posture, pain, trial duration, rest, instructions, every valid trial, and whether the reported value is a mean or maximum. If a posture must be modified for safety, retain the measurement but flag the modification and do not apply an incompatible reference threshold.

Fixed isometric knee strength

Stevens-Lapsley and colleagues provide a reproducible PD laboratory example: participants were seated and stabilized on a HUMAC NORM dynamometer at 60 degrees of knee flexion, completed two warm-up contractions and one practice MVC, then three maximal trials. The best trial was analyzed; verbal encouragement and visual targets slightly above the practice torque were supplied. Acquisition was at 2,000 Hz with gravity correction. Both legs were tested, although the published analysis deliberately used the less-involved/stronger leg. This is a protocol example and construct study, not a source of PD-specific SEM or MDC for routine knee testing. [6]

For any fixed knee assay, reproduce seat and backrest settings, hip and knee angles, alignment of the knee and dynamometer axes, the location and width of the distal-leg pad, and trunk/pelvic/thigh restraints. State whether zero degrees means full extension. Record the perpendicular external moment arm when converting load-cell force to joint torque. A pad moved proximally produces a different force reading for the same joint torque. Gravity correction, passive tension, the baseline subtraction method and any residual joint movement must also be specified. These mechanical requirements apply regardless of the brand name.

Handheld and belt stabilized hip and knee strength

Handheld dynamometry has substantial clinical appeal, but “portable” is not itself a reproducible protocol. The result depends on the examiner’s force, limb leverage, placement of the transducer, participant stabilization and whether the test is a make test or a break test. In a make test the participant builds force against a stationary resistance; in a break test the examiner overcomes the held position. Mixing these procedures changes the construct and can change both magnitude and safety.

An earlier study of 34 women with PD used a Nicholas MMT to measure non-dominant hip flexion and knee extension, seated with back support and hip flexion at 90 degrees. The lower trunk was strapped for hip testing; the thigh was strapped and the knee at 90 degrees for knee testing. Three trials with brief rests were averaged for each action, then summed as a composite leg-force score. The reported test–retest ICC3,1 range was 0.87–0.90, but the article did not give CIs, SEM/MDC, the repeat interval or a separate reliability sample size. Pad placement and examiner stabilization details were also incomplete. This supplies a PD protocol example, not a calibrated individual-change rule. [19]

The PD-specific Boom 2023 study included only 14 participants for intrarater and 10 for interrater analyses. Available primary abstract values range from an intrarater ICC of 0.98 (0.94–0.99) for wrist extension to 0.87 (0.43–0.97) for ankle dorsiflexion; interrater wrist flexion was −0.15 (−1.14–0.60), while grip was 0.97 (0.88–0.99). The original muscle-by-muscle table and positioning protocol remain necessary before endorsing a particular hip or knee implementation. Moreover, the interrater coefficient is reported as ICC2,2, an average-measures statistic; it cannot automatically establish reliability for a single clinician’s measurement. [40]

A sensible clinical implementation is to select a published position and preserve it, use a non-yielding belt or fixed support when the examiner cannot hold the limb stationary, mark the transducer site, measure the moment arm, and record each side separately. This is a standardization recommendation, not a claim that belt fixation and handheld resistance yield equal values. In non-PD validation studies, good reliability has coexisted with insufficient agreement between stabilization methods and between HHD and isokinetic dynamometry. Non-PD evidence also shows that tester strength can compromise agreement even in a frail sample. These findings support controlling the examiner’s contribution; their numerical errors must not be imported as PD thresholds. [41, 42, 43]

Hip tests additionally require control of pelvic rotation and trunk substitution, hip flexion/extension angle and knee position. A standing hip-abduction machine, side-lying HHD abduction and supine fixed abduction are different tasks. Apparent improvement after changing fixation or shortening the lever may be a protocol effect.

Isokinetic knee and hip torque

Inkster and colleagues used a calibrated Kin-Com for bilateral concentric hip extension in supine and knee extension seated with the hip at 90 degrees. Three submaximal and one maximal practice cycles preceded three maximal repetitions at 45 degrees per second. The outcome was average torque over the movement range, divided by body mass. Only five PD participants contributed to embedded retest assessment; the accessible original text calls the coefficients high but does not provide numerical ICCs or CIs. It is not defensible to turn this into a precise or general isokinetic reliability estimate. [27]

Frazzitta and colleagues offer a different knee protocol: Cybex Norm, hip at 90 degrees, trunk and thigh strapped, axis alignment, and gravity correction from the limb at 45 degrees of flexion. After submaximal warm-up, five maximal contractions were performed at each of 90, 120 and 180 degrees per second in randomized speed order, with one-minute recovery between maximal contractions. The best peak torque was retained, with impact artefact excluded. These speed-specific peak torques cannot be pooled numerically with Inkster’s range-averaged torque at 45 degrees per second. [28]

The operator should verify that the limb actually reaches the target speed over a usable range. Acceleration/deceleration phases and end-range impact can contaminate a nominally constant-speed measure. Specify concentric or eccentric action, range, peak versus average torque, the analyzed constant-velocity region and gravity correction. Isokinetic dynamometry offers control, but its restraints and padding can still permit movement during an intended isometric task.

Dynamic 1RM

Buckley and Hass used cable-loaded knee-extension, knee-flexion, chest-press and biceps-curl machines. Two familiarization sessions, 48–72 hours apart, each included two low-to-moderate-resistance sets. In the following week, a ten-repetition low-load warm-up preceded incremental loading. A successful repetition required complete controlled range without compensation, and 1RM was found within five attempts. No participant completed more than two 1RM tests in one day; an upper- and lower-body assessment were paired, and at least 72 hours separated tests. Medication was clinically ON, approximately 1–1.5 hours after the first dose. [2]

The paper does not supply exact joint-angle settings, load increments or inter-attempt rest. Those details should not be invented and attributed to its protocol. A local procedure must prospectively specify them using the equipment and supervising clinician’s standards. Document the largest successful load and the reason the next attempt failed: inability to generate force, loss of range, compensation, pain, fatigue or misunderstanding. A machine load in kilograms is specific to that machine’s leverage and pulley geometry, not a directly interchangeable joint torque or free-weight 1RM.

All participants completed testing without incident, but cardiovascular, musculoskeletal, vestibular and other neurological disorders were excluded, and no dyskinesia or freezing occurred during assessment. This supports feasibility for the selected cohort; it does not establish blanket safety for all PD.

Direct machine power

Paul and colleagues assessed 31 community-dwelling ambulant people with PD using Keiser A420 pneumatic equipment. Unilateral seated leg-press 1RM preceded standing hip-abduction 1RM. Left/right order was randomized once and reproduced at retest; power was then measured at six relative loads spanning 30–80% of 1RM, with peak power retained for each leg and muscle group. One assessor repeated testing one week later at the same optimal ON-medication time. “First” and “second” leg in the reliability table mean order tested, not more- and less-affected legs. The standing hip task also contains a balance and stabilization demand. [1]

The recovered Methods and supplement do not fully specify range, stabilization, load order, contraction counts at each load, rest, velocity sensing or the machine’s internal power algorithm. Consequently, they establish evidence for that study’s machine protocol but do not constitute a complete transferable operating procedure. In a new service, obtain and lock down those details. If 1RM changes at follow-up, a given percentage corresponds to a different absolute load. Report both the tested absolute loads and the relative-load definition; make clear whether a longitudinal comparison addresses peak power across a load range or power at a fixed load.

Allen 2010 illustrates why nominally identical equipment does not ensure an identical outcome. In its Keiser A420 protocol, unilateral leg-press 1RM was approached in about ten progressively heavier lifts; thirty minutes later, power was measured at eight ascending loads from 20% to 90% 1RM. The highest power across loads and the highest power at 30% 1RM were separate outcomes, then averaged across legs for analysis. Testing occurred approximately one hour after usual medication. This differs from Paul’s six-load 30–80% protocol and limb-specific reporting. Allen’s original refers fuller procedures to its 2009 report; it does not independently specify all joint angles, fixation or repetitions at each load. [31]

Encoder chair rise power

Bonde-Jensen and colleagues used a CHRONOJUMP Bosco linear encoder, software version 1.8.1, sampling at 1,000 Hz during the concentric chair-rise phase. Participants rose as fast and powerfully as possible to full standing while keeping their feet on the floor. They received two attempts separated by 30–60 seconds; the higher peak-power attempt was retained. Chair height, cable attachment, starting limb/arm position, familiarization, phase detection, filtering and force/power equations are deferred to a supplement that was unavailable for this appraisal. The recovered complete publisher text specifies these core trial rules but is not by itself a complete operating procedure. [4]

Construct validity was assessed in 69 analyzable participants with PD and 35 controls, all PD testing ON medication. Weight-normalized peak power averaged 14.1 versus 15.2 W/kg; the between-group difference was −1.1 W/kg (95% CI −2.8 to 0.5). A subgroup of ten with MDS-UPDRS III ≥33 had lower power than the mildly affected group and controls. Associations with walking tests explained 44–51% of variance and persisted after adjustment for age and sex. This is evidence about known-group differences and convergent validity. The authors explicitly identify criterion validity as untested: no concurrent force-platform or power-rig agreement comparison was performed. An encoder-derived force trace is not independently measured ground-reaction force, and the chair-rise result remains a functional power estimate. [4]

Repeatability involved twenty participants across four visits, paired in randomized medication-state order, with two ON and two approximately 12-hour OFF sessions separated by 4–16 days. Nineteen paired observations were available per state. The selected sample had mean Hoehn and Yahr stage 2, a maximum eligible stage of 3, and mean disease duration 4.3 years. One participant discontinued after OFF pain and fatigue; one ON data set was lost technically. Keeping the feet grounded may constrain maximum output in higher-functioning people, a ceiling concern acknowledged by the authors. These selection and execution constraints limit generalization to advanced disease or unrestricted explosive rising. [4]

RFD and rapid force control assays

RFD requires an explicit force-time definition. The 2021 PD-subtype knee study used a leg-extension machine with a perpendicular load cell, reported knee and hip angles of 110 and 90 degrees respectively with zero stated as full extension, one familiarization MVC, and three five-second maximal efforts separated by one minute. Participants were instructed to act rapidly and maximally; average slopes over the first 50, 100 and 200 ms were averaged across three trials. Sampling frequency, filtering, onset threshold, lever arm and a clear unilateral/bilateral specification were not present in the retrieved Methods. The reported 110-degree convention should be clarified rather than silently translated into a presumed 70-degree flexion position. These omissions limit reproduction and cross-study comparison. [37]

Hammond 2017 used a different rapid-force definition: maximal derivative of right-knee extensor force, measured through a fixed ankle-cuff/load-cell setup with the knee at 90 degrees, trunk and lap restraints, and arms crossed. Voluntary RFD was compared with electrically evoked octet RFD; superimposed triplets at the MVC plateau supplied a central activation ratio. A 30 Hz fourth-order Butterworth filter preceded differentiation, but sampling frequency and stimulus repetition frequency were not reported. The printed hip position, “100° extension” in upright sitting, also needs clarification. Its maximal derivative is not equivalent to an early fixed-window slope. [36]

In the declared final sample of seven PD participants and six controls, voluntary RFD and its voluntary/involuntary ratio were lower in PD (reported d=1.97 and 2.18; p=.008 and .004), while maximal-force and CAR differences were not detected. The cohort had mild H&Y stages but mean disease duration 7.9 years, so “mild” should not be translated into newly diagnosed. The enrollment/exclusion narrative does not reconcile with the final group and sex counts; two people discontinued for stimulation discomfort and two force records had impact artifact. This is evidence of a possible rapid-activation deficit and specialized-test feasibility constraints, not proof of preserved peripheral capacity or a validated progression biomarker. PD-specific ICC, SEM, MDC and MIC were not measured. [36]

The 2026 custom grip study used 2 kHz acquisition and distinct 0–50, 50–100 and 100–200 ms windows. The paper states that force onset was a preset threshold above baseline but does not report its numerical value in the retrieved Methods; its filtering details are also unclear. “RFD50” can mean the first 50 ms, a value at 50 ms, a moving 50-ms slope or another definition. Such labels require expansion before comparison. [8]

Chung 2023 studied visually guided force control at 15% MVC, not maximal early RFD. Hand pinch and ankle dorsiflexion were measured at 50 Hz, with reported acquisition filtering at 20 Hz and analysis filtering at 15 Hz. Force-rise and relaxation rates were calculated between manually/algorithmically defined phases and normalized to MVC. The study’s more-affected-side, OFF-state, target-matching task should not be represented as a high-bandwidth maximal explosive test. [13]

Daniels 2025 is also a distinct assay: rapid index-finger abduction, not rapid whole-hand grip. Participants used a rigid finger bumper and force transducer, with shoulder about 20 degrees and elbow about 110 degrees; acquisition was 200 Hz with a 50 Hz low-pass filter. Five 60-second trials produced pulses in different submaximal ranges. Analysis selected 41 pulses between 20 and 60% MVC, applied 50-ms moving-slope derivative windows, and used a 20% MVC/s derivative threshold for pulse onset/offset. Segmentation was classified from the median number of force-rise segments to 90% peak force. Jamar grip was measured separately. This protocol measures the organization of rapid submaximal force pulses and does not validate a knee-RFD or maximal-grip cutoff. [14]

A 2026 mechanistic study adds convergent evidence for finger-abduction segmentation: in 16 ON-medication participants with PD and 12 older controls, the number of force segments correlated with EMG bursts (rho 0.84), and first-segment peak RFD correlated with initial EMG amplitude (rho 0.83). This supports a relationship to neural activation; it does not establish individual agreement, interchangeable measurement with EMG, longitudinal responsiveness or a clinical cutoff. Only the primary abstract and publisher excerpts were available for this audit, and cohort overlap with earlier Daniels studies remains unresolved. [44]

For a new early-RFD protocol, general methodological evidence supports a rigid low-noise setup, acquisition at at least 1 kHz, a separate familiarization session, consistent fast-contraction instructions, stable-baseline checks and rejection of counter-movement or unintended pre-tension. Maffiuletti and colleagues recommend collecting at least five acceptable efforts and averaging the best three, with multiple predefined windows and minimal signal filtering. Their recommendation is methodological extrapolation, not a PD-validated optimum. A service should establish repeatability of its own complete acquisition and analysis pipeline before interpreting individual small changes. [45]

Table 2 Selected protocol differences that prevent interchangeability

Table 2 Selected protocol differences that prevent interchangeability
Assay and sourceExecution and scoreComparison boundary
Jamar grip [3]Seated; elbow 90°; handle 2. Three pain-free 3-s efforts per hand, 1-min rests; mean of three after practice.Usual medication; dominant/nondominant sides. Different from a best-of-two or standing protocol.
Fixed quadriceps [6]HUMAC NORM at knee flexion 60°; gravity corrected; two warm-ups, practice and three maxima; best.Optimal ON. Both legs tested, but stronger/less-involved leg analyzed. Construct study, not error threshold.
Isokinetic hip and knee [27, 28]Inkster: average concentric torque over range at 45°/s, Nm/kg. Frazzitta: best peak knee torque at 90, 120 and 180°/s, Nm.Range average versus peak; speed, action and normalization differ. Inkster retest subset only five PD.
Machine 1RM [2]Two familiarization visits; 10-repetition warm-up; incremental load, ≤5 attempts; controlled full range; same assessor, ≥72-h retest.Clinically ON 1–1.5 h after first dose. Machine-specific kg load; retest learning persists.
Pneumatic power [1]Unilateral seated leg press then standing hip abduction; peak across 30–80% 1RM; one-week same-assessor retest.Optimal ON. First/second legs refer to test order. Load range differs from power at a fixed 40% 1RM.
Handheld strength [40]Upper and lower muscle groups; intrarater n=14 and interrater n=10. Full positioning and aggregation not verified.Assessor design matters. Do not apply an upper-limb ICC to hip/knee testing or average-measures ICC to a single score.
Allen leg-press power [31]Keiser A420; about ten lifts to 1RM; 30-min rest; eight ascending power loads from 20–90% 1RM; bilateral mean.Peak over all loads and maximal 30%-load power were separate outcomes. Neither equals Paul’s six-load, limb-specific repeatability score.
Encoder chair rise [4]CHRONOJUMP Bosco v1.8.1, 1 kHz; fastest concentric rise with feet grounded; best of two, 30–60-s rest.69 PD/35 controls for construct analysis; 19 paired per ON/OFF state. Chair, attachment and algorithm remain in unavailable supplement.

Table 3 Rapid force assays measure different tasks and time scales

Table 3 Rapid force assays measure different tasks and time scales
AssayAcquisition and definitionImplication
Hammond knee RFD [36]Right knee, fixed ankle cuff/load cell, knee 90°; maximal derivative after 30-Hz fourth-order Butterworth filter; voluntary versus evoked octet.Sampling frequency absent. Not a fixed early-window slope; hip-angle convention and stimulation details incomplete.
Knee RFD subtype study [37]Three 5-s fast maximal efforts; average slopes 0–50, 0–100 and 0–200 ms. ON state. Sampling, filtering and numerical onset rule unavailable in inspected Methods.Maximal explosive force. Protocol omissions limit replication and window comparisons.
Custom grip [8]2 kHz; standing, elbow extended. RFD 0–50, 50–100 and 100–200 ms from strongest MVC trial. Preset onset threshold without numerical value.Sequential windows differ from cumulative 0–100/200 ms. Strongest trial may not have fastest rise.
Pinch and ankle rate control [13]50 Hz; 20-Hz acquisition and 15-Hz analysis filters; visually guided 15% MVC contractions; phase slopes normalized to MVC. OFF state.Submaximal tracking, not a high-bandwidth maximal early-RFD protocol.
Finger-abduction segmentation [14]200 Hz; 50-Hz filter; 50-ms moving derivative; onset at 20% MVC/s. Forty-one pulses at 20–60% MVC; median segments to 90% peak.Algorithm-defined rapid submaximal control. Jamar grip is a separate outcome.

Reliability and clinical change

Relative reliability is not individual precision

An ICC compares between-person variation with error variation. A heterogeneous sample can therefore have an impressive ICC even when repeated scores for one person differ substantially. Specify the statistical model, absolute-agreement versus consistency definition, single versus averaged score, rater design and confidence interval. A point estimate alone conceals uncertainty: an ICC of 0.87 with a lower confidence limit of 0.43 does not establish consistently strong repeatability. [46]

Buckley and Hass provide an especially clear PD example. Knee-extension ICC was 0.96 (0.93–0.97), but the retest increased by a mean 4.0 kg (1.9–6.2). Knee-flexion and biceps-curl gains were also significant after two familiarization sessions; individual knee-extension differences reached plus or minus 27 kg, and 51% of participants had changes of at least 5 kg across testing. Retaining participants’ rank is compatible with learning and large individual fluctuations. A single low baseline should not automatically be interpreted as deficit, and a subsequent higher score should not automatically be credited to treatment. [2]

Absolute error and meaningful change

SEM describes measurement error in the original unit. Under an appropriate stable-person, approximately homoscedastic two-measurement model, MDC95 can be calculated as 1.96 × square root of 2 × SEM. It is an estimated random-error boundary for an individual difference, not a minimal important change. MIC requires a defensible external anchor or other evidence of importance to the patient. Limits of agreement additionally make systematic bias and the spread of pairwise differences visible. If errors grow with measurement magnitude, a constant absolute threshold may be inappropriate; proportional or log-scale analysis may be required. [46, 47]

Calculated from Paul’s reported SEMs, leg-power MDC95 is approximately 97 and 100 W, about 21% of the baseline group means. Hip-abductor power MDC95 is approximately 51 and 35 W, about 49% and 33%. For leg-press strength the corresponding values are 139 and 119 N, and for hip-abduction strength 16.6 and 13.9 Nm. These are transparent calculations, not reported MICs or validated treatment-response thresholds. Paul used consistency-based ICC3,1; mean hip-strength scores increased from 58.3 to 62.0 Nm and from 59.7 to 63.5 Nm. No direct paired significance test for those shifts was recovered. The calculated MDCs are random-error benchmarks and do not establish absence of systematic retest bias. The percentages use the study’s group means for context and are not personalized percent cutoffs. A 25 W hip-power gain cannot be called definitely real merely because its ICC exceeds 0.9. [1]

Applying the same arithmetic to Buckley’s reported SEMs gives 15.8 kg for knee extension, 10.5 kg for knee flexion, 7.8 kg for biceps curl and 11.9 kg for chest press. Because the study demonstrated systematic retest gains, these are only random-error benchmarks; they do not remove learning bias. Choosing the highest of multiple baseline visits also changes the score definition and introduces its own selection effect. Decide prospectively whether to use a stabilized repeat session or a prespecified multi-session summary, then use the same rule at follow-up. [2]

Encoder reliability and its conflicting change thresholds

For body-mass-normalized chair-rise peak power, Bonde-Jensen reports absolute-agreement ICC2,1 values of 0.95 (0.87–0.98) ON and 0.93 (0.82–0.97) OFF. Table 3 labels 1.0 and 1.3 W/kg as smallest detectable change (SDC), rather than SEM or MIC. However, the reported Bland–Altman 95% limits are −3.3 to 2.7 W/kg ON and −3.3 to 3.6 W/kg OFF, with mean differences of −0.3 and +0.2 W/kg. The pairwise dispersion is substantially larger than the table SDCs imply. The authors acknowledge this discrepancy but do not resolve it in the main article; the formula details are referred to an earlier paper. [4]

Consequently, their proposed 5.7% ON and 7.9% OFF “true-change” rules should not be adopted as established individual thresholds. Report the SDC and agreement limits side by side, flag the inconsistency and retain the wider observed pairwise variability for interpretation. As an audit calculation, the approximate half-widths of those limits are 3.0 and 3.45 W/kg; under the usual stable-person two-score model these would be the corresponding random-error difference magnitudes, rather than 1.0 and 1.3. They are not validated replacement cutoffs or MICs. The high ICCs support rank repeatability within this selected cohort but do not settle individual change precision or criterion accuracy. [4]

The grip error estimate that should not be used

Villafañe’s PD grip ICCs are high: dominant 0.97 (0.92–0.99), non-dominant 0.98 (0.95–0.99). However, its displayed SEM of 0.05 kg per hand does not reproduce from its stated SD × square root of (1−ICC) formula and displayed SDs of 10.2 and 10.3 kg. Those inputs give approximately 1.77 and 1.46 kg, not 0.05 kg. These audit calculations are not replacement clinical thresholds, because raw data and the intended calculation are unavailable. [3]

The problem extends beyond that discrepancy. The text proposes ±0.092 kg as true change using two SEMs, which is not the conventional two-score MDC95 calculation. Table 2 uses inconsistent signs for the displayed differences. The Bland–Altman figure labels right and left hands, while the table uses dominance, and the plotted supposed 95% limits are visibly narrow relative to many of the individual differences. No correction was located in a targeted DOI/erratum search on 2 October 2026. The defensible conclusion is reasonable evidence of rank repeatability in a small selected cohort, with unresolved absolute-error reporting. Withhold its 0.092 kg rule and do not quote its agreement limits uncritically. [3]

Silva’s modified sphygmomanometer offers a separate caution. Its pressure output correlated with Jamar values, but correlations of 0.68 and 0.45 in PD do not establish device interchangeability. Table 3 reports pressure SEMs of 2.55 and 2.67 mmHg and MDCs of 7.06 and 7.40 mmHg, but does not clearly separate the rater/retest contrasts. Figure 2 shows PD interrater biases of −8.1 and −10.3 mmHg, with limits of −26.6 to 10.4 and −33.9 to 13.4 mmHg. Those dispersions do not reconcile with applying the small table MDCs to between-rater differences. The table values should therefore not be used as a general interrater change rule, and none is a kilogram-force threshold for Jamar. [15]

A later tip-pinch study included 50 people with PD and reported moderate correlations between modified sphygmomanometer and pinch-dynamometer values, r=0.44 and 0.48 for dominant and non-dominant sides. This is a different pinch construct and an abstract-level correlation-based validation result. It does not rescue device interchangeability, provide a grip-strength error threshold, or establish individual change precision. [48]

Table 4 Paul machine strength and power repeatability

N=31; same assessor, one-week interval and matched optimal ON timing. ICC3,1 and 95% CI. First and second mean test order. MDC95 is calculated here as 1.96 × √2 × reported SEM; it is not a published MIC. Percentages use baseline group means, not individual percent thresholds. [1]

Table 4 Paul machine strength and power repeatability
Measure and test orderICC and 95% CIReported SEMCalculated MDC95MDC95 as % mean
Leg power
first tested
0.96 (0.91–0.98)35 W97.0 W20.8%
Leg power
second tested
0.95 (0.90–0.98)36 W99.8 W20.6%
Hip abductor power
first tested
0.85 (0.71–0.92)18.5 W51.3 W48.6%
Hip abductor power
second tested
0.93 (0.87–0.97)12.5 W34.6 W33.3%
Leg press strength
first tested
0.96 (0.92–0.98)50 N138.6 N16.0%
Leg press strength
second tested
0.97 (0.94–0.99)43 N119.2 N13.5%
Hip abduction strength
first tested
0.93 (0.86–0.97)6 Nm16.6 Nm28.5%
Hip abduction strength
second tested
0.96 (0.92–0.98)5 Nm13.9 Nm23.2%

Table 5 Machine 1RM reliability and systematic retest gains

All loads in kg on the study machines. Same assessor and ON-state protocol. Calculated MDC95 uses reported SEM and does not correct the observed retest bias. Confidence intervals are 95%. Different exercises had different samples. [2]

Table 5 Machine 1RM reliability and systematic retest gains
Exercise and nICC and CIMean retest gain and CISEMCalculated MDC95
Knee extension
n=46
.96 (.93–.97)4.0 (1.9–6.2)5.715.8
Knee flexion
n=21
.91 (.79–.96)2.4 (0.2–4.7)3.810.5
Biceps curl
n=25
.97 (.92–.98)2.7 (1.2–4.1)2.87.8
Chest press
n=24
.95 (.90–.98)2.3 (−0.2–4.7)4.311.9

Table 6 Encoder chair rise repeatability and conflicting individual change estimates

Bonde-Jensen 2024: 19 paired observations per state. Absolute-agreement ICC2,1 and 95% CI. SDC values are reproduced from Table 3, not endorsed; their inconsistency with paired agreement limits prevents a validated small-change rule. No MIC or criterion agreement was established. [4]

Table 6 Encoder chair rise repeatability and conflicting individual change estimates
State and metricTest 1 to test 2 mean (SD)ICC and 95% CIReported SDCBias and 95% limits
ON W/kg17.4 (4.8) to 17.1 (4.4).95 (.87–.98)1.0 W/kg−0.3; −3.3 to +2.7 W/kg
OFF W/kg16.5 (4.8) to 16.7 (4.1).93 (.82–.97)1.3 W/kg+0.2; −3.3 to +3.6 W/kg
ON W1382 (399) to 1360 (375).96 (.89–.98)66 W−23 W; absolute-W limits not reported
OFF W1302 (386) to 1321 (340).94 (.85–.98)90 W+19 W; absolute-W limits not reported

Reported SEM percentages were 2.1% ON and 2.9% OFF for W/kg, and 1.7% ON and 2.5% OFF for W. The normalized agreement-limit half-widths are approximately 3.0 and 3.45 W/kg, calculated for audit only; these are not validated replacement thresholds. Do not endorse the authors’ 5.7%/7.9% true-change rules.

Table 7 Other measurement studies and limits on individual change

Table 7 Other measurement studies and limits on individual change
Study and designAudited findingImplication for change
Villafañe 2016
15 PD; one week [3]
Grip ICC .97 (.92–.99) dominant; .98 (.95–.99) nondominant. SEM .05 kg does not reproduce from stated formula and displayed data.Withhold proposed ±.092 kg rule. Figure and difference-label inconsistencies also prevent uncritical use of reported limits.
Silva 2015
24 PD [15]
Pressure SEM 2.55/2.67 mmHg; tabulated MDC 7.06/7.40 mmHg. Interrater bias and wide limits require separate interpretation.Pressure is not Jamar kg. Tabulated MDC must not be transferred to between-rater change or another cuff adaptation.
Boom 2023
Intra n=14; inter n=10 [40]
Intrarater wrist-extension ICC .98 (.94–.99); ankle dorsiflexion .87 (.43–.97). Interrater wrist flexion −.15 (−1.14–.60), grip .97 (.88–.99).Primary abstract only. Muscle-specific tables, scoring and absolute error unavailable. Interrater ICC2,2 is an average-measures statistic.

Medication state and test feasibility

For a longitudinal clinical series, test at the same clinically relevant medication state, document drug timing and observed state, and record dyskinesia, tremor, freezing, sleep, pain and exertion that could influence execution. “Usual medication” is a weaker standard than a stated time window plus confirmation of ON status. The purpose matters: a usual-function ON assessment and a controlled ON/OFF comparison answer different questions. Patients should not independently withhold medication simply to satisfy an assessment protocol; OFF testing requires an agreed clinical/research plan.

Medication effects are assay-specific. Ingram found no mean ON/OFF grip difference but an elbow-flexion difference, and the 2025 elbow study found higher MVC ON medication. Neither means every muscle measure must improve ON nor that grip is medication-independent. Paul included participants with dyskinesia and fluctuations; removing disabling dyskinesia did not materially rescue muscle measures in the way it changed sway results. This supports retaining those actual subgroup data rather than importing a universal dyskinesia exclusion. It does not show that severe involuntary movement never contaminates a force trace. [7, 1, 49]

Bonde-Jensen found no average ON/OFF encoder-power difference (0.4 W/kg; 95% CI −0.6 to 1.3) and similar within-state ICCs in its relatively mildly affected cohort. This does not establish equivalence of medication states, interchangeable ON/OFF measurements in an individual, or absence of state effects in more advanced PD. Continue standardizing state and timing in longitudinal testing. The main text does not supply an exact post-dose ON interval or a dyskinesia-specific reliability analysis. [4]

At baseline, measure both sides when feasible and record right/left, dominance, clinically more-affected side and testing order as separate fields. They are not synonyms. Maintain the original side labels at follow-up, even if the measured weaker side changes. Report both scores and their difference before reducing them to an average or asymmetry index. A small asymmetry can be within error; an apparently symmetric patient can have bilateral weakness. No universal 10–15% PD asymmetry threshold is justified by this evidence. The 2026 grip findings change according to whether PD’s more-affected hand is compared with control dominance or matched hand dominance, demonstrating why side matching matters. [8, 27, 28]

Record within-session trajectories. Rising values may indicate familiarization or potentiation; falling values may indicate fatigue, pain, loss of attention or medication wearing off. Subjective fatigue and performance fatigability are different. In Stevens-Lapsley’s repeated-contraction test, the high-motor-severity subgroup increased torque rather than showing the expected decline, plausibly because initial activation was inadequate. Failure to show a torque decline should not automatically be described as exceptional endurance. [6]

Feasibility and safety

No single assessment spans all levels of PD. Grip has low equipment and transfer demands, although severe rigidity, painful hand disease, poor comprehension or inability to hold the device may invalidate it. HHD is portable but needs a skilled examiner and reliable stabilization. Fixed/isokinetic testing improves control but adds setup, transfers, alignment and access requirements. Standing hip-abduction power adds balance demands; machine 1RM and repeated power trials add exertion and familiarization requirements. RFD adds acquisition and analytical expertise, rejection decisions and greater sensitivity to baseline noise.

Formal floor/ceiling frequencies were generally not reported in the core PD measurement studies. It is therefore more accurate to identify plausible measurement constraints than claim that a test is free of floor or ceiling effects. Examples include a minimum machine resistance that an individual cannot move, a sensor capacity limit, a coarse load increment, examiner force limiting a strong patient, and exclusion of people who cannot safely transfer or follow the instructions. “Unable to complete” should be recorded with its reason, not converted automatically to zero strength or power.

Normalization and asymmetry

What the score is intended to represent

Use force in N, torque in Nm, power in W, force RFD in N/s and torque RFD in Nm/s. Some dynamometers display “kg” although calibrated in kilogram-force; confirm the manufacturer’s convention. One kgf equals 9.80665 N. Mass in kg is not force, and converting force to joint torque additionally requires the perpendicular moment arm in metres. An external machine load does not become isolated joint torque by assuming a lever.

There are at least three different questions behind “normalization.” Absolute force or torque describes the capacity available to act on an external object. Force or power divided by body mass describes capacity relative to the mass that must be transported during standing or walking. Force or torque divided by a muscle-size measure asks a more mechanistic question about output relative to tissue size. A standardized z score locates a person relative to a reference distribution. None automatically answers the others. The most defensible reporting strategy for an individual is to retain the raw value and, when the functional question warrants it, add a prespecified relative value with the denominator recorded.

The importance of this distinction is visible in PD originals. Pääsuke and colleagues studied 14 sedentary women with PD and 12 age-matched controls. Absolute maximum knee-extensor force was lower in PD, whereas force relative to body weight did not differ significantly. The rate and timing of force production and chair-rise performance were still impaired. The result does not show that one normalization is correct and the other wrong. It shows that body size affects the question being tested, and that a nonsignificant small-sample ratio comparison cannot establish equivalent muscle function. The available source for these details is the original abstract; the full methods and numerical comparisons still warrant recovery. [22]

Relative strength is not a complete adjustment for confounding

Dividing by kilograms is convenient, but assumes a particular proportional relation between strength and body mass. It does not demonstrate that the ratio is independent of height, sex, age, limb length, adiposity or disease-related weight loss. The retrieved PD evidence does not establish one universally optimal allometric exponent or a validated body-size correction transferable between grip, knee torque and machine power. Allometric or regression adjustment can be useful research strategies, but should be justified for the sample and intended interpretation rather than imported as a universal clinical rule.

A ratio can also change when its denominator changes. If absolute grip remains unchanged while body mass falls, grip per kilogram rises by arithmetic alone. That may matter differently for the ability to move the body and for the interpretation of underlying neuromuscular health. For longitudinal monitoring, recording weight alongside absolute and relative strength prevents this change from being mistaken for increased force capacity. Likewise, an association between two variables sharing a body-mass denominator may partly reflect that common denominator. This is a methodological concern, not an observed correction factor estimated in these PD studies.

Direct evidence also shows that W/kg does not remove the need to consider other covariates. In Corfitsen and colleagues’ 2026 study, concentric chair-rise peak power was measured with a CHRONOJUMP linear encoder at 1000 Hz, using the best of two attempts. There were 105 participants with PD, with power data available for 101; mean relative peak power was 15.2±4.3 W/kg. Several simple power–brain-volume associations weakened after adjustment for age and sex. For whole-brain volume, the estimated association fell from 2.69 mL per W/kg (95% CI 1.28–4.10) to 0.95 (−0.34 to 2.24). The adjusted association with Symbol Digit Modalities Test performance was 0.64 points per W/kg (0.25–1.03). These findings concern concurrent association; they do not establish neuroprotection, prognostic validity or responsiveness. Some other extracted table entries contain inconsistent confidence-interval formatting and should not be quoted without checking the original layout. [50]

The contrast between ordinary and physically active PD samples is also important. Martignon and colleagues compared 10 active participants with PD with 10 controls matched for age, sex and daily energy expenditure. Mean quadriceps MVC and physiological cross-sectional area were similar, while voluntary activation and early torque development were impaired in PD. Preserved maximum output can therefore coexist with altered neural performance. This selected, small, cross-sectional study cannot establish that activity caused protection, nor that muscle-size adjustment removes all PD-related dysfunction. The abstract’s specific-torque numerical scale requires full-table checking and is deliberately not reproduced here. [51]

Sex and the choice of reference group

A mixed-sex group can show a grip or torque difference because of the balance of men and women, body size or both. Dividing by weight may leave substantial differences. Analyses should therefore make clear whether they compare sex-matched groups, adjust statistically for sex, use sex-specific reference values, or pool raw data. A study-specific global strength z score is not automatically a sex-adjusted normative score. The 2024 study of specific leg muscle strength used standardized strength values in several muscle-specific analyses; its adjustment description and table captions should be checked rather than assuming that standardization itself controlled sex. [52]

Grip also needs careful framing as a general functional measure. The 2026 study using a custom grip transducer found lower maximum force on the more affected than the less affected side within PD, but group conclusions depended on whether the comparison hand in controls was dominant or nondominant. Maximum grip was related to maximum gait speed, whereas most RFD measures were not. This illustrates the difference between a useful general-capacity association and a disease-specific signature. [8]

The newer handgrip review by Mazza and colleagues identified 27 observational articles, including 16 with a PD–control comparison; seven reported a significant between-group difference. These counts, presently verified from the review’s abstract, reinforce the need to examine protocols and reference groups. They should not be interpreted as a vote-counting meta-analysis showing equivalence in the remaining studies. Sample size, variance and confounding influence whether a comparison is significant. [23]

Dominance disease laterality and asymmetry

“Right,” “dominant,” “more affected,” “weaker,” and “first tested” name different attributes. The dominant hand can be the more or the less affected hand. A clinically more affected limb need not produce the smaller maximum force on a particular day. Selecting the weaker limb after testing also differs from prospectively selecting a limb from the clinical examination. For longitudinal interpretation, the same clinically defined side should be followed, and both raw sides should be retained where feasible.

This distinction is particularly important for the Paul reliability study. Its first- and second-leg results refer to the tested order; they cannot be relabelled more- and less-affected limbs. In other studies, bilateral averages were deliberately used to summarize whole-person capacity, while some mechanistic work tested only the more affected limb. Bilateral averaging can reduce noise but conceal a meaningful unilateral limitation. These choices are not necessarily errors, but they answer different questions. [1, 5, 11, 53]

Frazzitta and colleagues’ 2015 study is an informative warning against a simple universal asymmetry rule. Twenty-five participants with PD, all right-handed, were grouped by the clinically more affected side and compared with 15 controls during isokinetic knee testing. The pattern included subgroup-specific differences despite several nonsignificant overall PD–control comparisons. This does not validate a universal 10% or 15% asymmetry cutoff, prove that the dominant limb is always protected, or provide an individual-level diagnostic threshold. Small subgroups and multiple comparisons also make the pattern vulnerable to sampling variation. [28]

An asymmetry index must state its denominator and sign convention. A difference divided by the stronger limb is numerically different from a difference divided by the bilateral mean. Absolute-value indices discard which side is weaker; signed indices preserve direction but require a stable side definition. The retrieved PD sources do not establish the reliability, minimal detectable change or prognostic cutoff of a universal strength-asymmetry index. Recording both limb values and the clinically defined side is currently more transparent than treating a single percentage as validated disease severity.

Normalizing rapid force development changes its meaning

Absolute RFD describes how quickly force rises, in N/s or torque in Nm/s. RFD divided by MVC describes force-rise speed relative to the person’s available maximum. If MVC is reduced, a normalized RFD value may appear preserved despite a lower absolute force-generating rate. Conversely, normalization can help distinguish a slow activation pattern from a proportionate consequence of low maximal capacity. Both are useful questions, but a result in one unit does not validate a conclusion in the other.

The retrieved PD studies also use different tasks. Powerful maximal knee extensions, rapid submaximal pulses to a percentage of MVC, first-segment peak RFD, and a 50-ms moving derivative are not the same endpoint. Park and Stelmach’s target-pulse study addressed amplitude scaling at 15%, 35% and 55% MVC; the recent Daniels work addressed segmentation in rapid finger abduction, with Jamar grip measured separately in the 2025 study. The newly retrieved 2026 EMG study supports a link between force segmentation and bursts of muscle activation, but not a universal RFD measure or clinical cutoff. [33, 38, 14, 44]

In practice, an RFD report should specify the muscle action, absolute or MVC-normalized unit, contraction target, force onset rule, sampling frequency, filtering, time window, repetition selection and medication state. Dividing a noisy derivative by an uncertain MVC does not remove measurement error; uncertainty enters through both terms. The present PD literature supports mechanistic interest in rapid force generation, while device-specific repeatability and patient-level meaningful-change thresholds remain much less secure.

Functional associations and future outcomes

Current function relevant but task specific associations

Lower limb strength does not explain all mobility domains equally

Schilling and colleagues studied 17 people with PD and 10 controls using maximal isometric multi-joint leg-press force divided by body mass. Within PD, relative force correlated with Timed Up and Go performance (r=−.68, p=.003). The PD group had lower relative force (p=.044). This small, unadjusted, cross-sectional finding supports concurrent relevance to a composite mobility task. It does not establish that the test predicts a future mobility loss, nor that relative force outperforms absolute force after appropriate body-size adjustment. Full procedural detail was not recovered; these are original-abstract findings. [54]

Nocera and colleagues provide a useful contrasting pattern. In 44 optimally medicated participants, bilateral concentric knee-extension torque was measured with a stabilized, gravity-corrected KinCom at 60°/s. Strength correlated with the centre-of-pressure–centre-of-mass separation during gait initiation (r=.500, p=.001) and with Hoehn and Yahr (HY) stage (rho=−.484, p=.002). It did not significantly correlate with six-minute walk distance (r=.248, p=.127) or functional reach (r=.221, p=.177). The original tables therefore support a particular dynamic-postural association, rather than universal strength–function coupling. Although the discussion uses progression language, this was a cross-sectional comparison of current severity. Neither the gait-initiation index nor its correlation with torque was validated as a prospective falls endpoint. [55]

The practical inference is to match an impairment test to the functional question. TUG, six-minute walking, gait initiation, reach, chair rise and reactive stepping differ in endurance, balance, coordination, cueing and speed demands. A null correlation with one should not invalidate the strength assay; a positive correlation with another should not justify replacing the function test. The functional task remains the more direct observation of what the person can do, while dynamometry helps characterize one potential constraint.

Machine power has its strongest concurrent evidence in selected anticipatory tasks

Paul et al.’s functional study included 82 community-dwelling people assessed ON medication. Keiser A420 power at 40% of individual one-repetition maximum (1RM), averaged across legs, was related to current performance. Models included freezing, dyskinesia, executive function, motor severity, age, sex and height; leg-extensor and hip-abductor power entered separate models because their correlation was .81. Leg-extensor power related to chair-rise rate (standardized beta=.5, p=.003); hip-abductor power related to fast walking (beta=.3, p=.02) and choice stepping (beta=.2, p=.048). Adjusted reactive-balance models were not significant. [5]

These findings favour task specificity, not a general claim that power is the dominant determinant of all balance. The reported explanatory variance belongs to complete models, not to the independent contribution of power. There was no prospective outcome or external prediction validation. Power measured at a percentage of 1RM also embeds the strength test in the exposure definition; it cannot cleanly demonstrate that velocity is prognostically more important than force. Assigning values for people below the equipment’s minimum load or unable to complete a task further makes the lower-performing tail partly model-dependent.

Caetano and colleagues studied 54 people ON medication during projected-obstacle and stepping-target tasks. Hip-abductor power used a fixed 35-N resistance, rather than each participant’s 40% 1RM. Higher power correlated with baseline step length (r=.689) and remained in step-length models at baseline (beta=.598), walk-through (.558), and short-target conditions (.403/.338). Conversely, stepping errors were chiefly associated with inhibitory control, and short-target placement accuracy with reactive balance and executive function. These were concurrent laboratory regressions, even where the paper calls variables “predictors.” They do not establish future-falls discrimination or validate the fixed-load protocol as interchangeable with the relative-load Keiser protocol. [56]

A longer step is not automatically a safer step in every environment. The dissociation between step length and placement accuracy is especially important: more mechanical capacity may facilitate scaling movement, while rapid selection, inhibition, balance recovery and hazard perception determine whether that movement is appropriately deployed. An assessment battery should preserve this distinction rather than using machine power as a proxy for adaptive walking safety.

Handgrip has functional relevance without being a global PD severity measure

Ingram et al.’s 34-person PD study measured dominant-hand grip, elbow flexion and other upper-limb functions during ON and OFF conditions. Grip correlated with the Disabilities of the Arm, Shoulder and Hand questionnaire (DASH; r=−.562 OFF and −.549 ON); elbow-flexion associations were similar (−.576/−.531). Neither strength measure had a significant relationship with disease duration or HY stage in that sample. These findings support concurrent upper-limb construct relevance, not agreement with an objective gold standard and not future disability prediction. The study itself did not establish repeatability of this battery in its PD participants. [7]

The studies are not inherently contradictory. HY disproportionately reflects bilateral involvement, postural instability and global mobility, whereas DASH addresses the upper limb. A grip test can be relevant to the latter while weakly representing the former. Further, current severity is partly treatment-sensitive, whereas duration is an imperfect proxy for accumulated impairment. No single correlation establishes grip as an omnibus measure of PD burden.

Jones et al.’s title refers to “long-term electromyography,” but the observation was approximately 6.5 hours of daily-muscle activity in 23 participants with PD and 14 controls. Grip explained selected EMG burst features in stepwise analyses (R²=.17–.33 in PD). This is evidence about current activity, not follow-up of functional decline. Its title should not be used to populate a prognostic evidence table. [57]

Sarcopenia disability and functional composites require construct separation

Low strength or dynapenia is not synonymous with confirmed sarcopenia. In the revised European framework used by these studies, low strength raises probable sarcopenia; reduced muscle quantity or quality is needed for confirmation, and poor physical performance indicates greater severity. A questionnaire screen, measured grip, muscle mass and a composite diagnosis must therefore retain separate labels. An algorithm containing the same grip or chair-rise measure cannot independently validate that component. [58, 59]

Ozer et al. studied 70 people with PD and 85 controls cross-sectionally. Dynapenia was associated with current disease severity and Katz/Lawton disability, but the title’s “predictors of disability” does not imply future disability prediction. Sarcopenia, low muscle strength, low muscle mass and disability are not interchangeable exposures. The original abstract supports the design and direction; it does not provide the full adjustment set or an externally tested prognostic model. [60]

A recent 127-person ON-state study similarly found lower mean grip in those with poorer Short Physical Performance Battery performance (26 versus 30 kgf, p=.02). Its final model emphasized SARC-F and the postural-instability/gait-difficulty score, with reported AUC=.82. This is current-performance classification. The AUC belongs to the composite model and does not validate grip’s increment, a clinical grip threshold, or future decline. The original abstract is the access level available for this extension. [61]

Composite definitions can partly build a functional relationship into the exposure. SARC-F contains questions about walking, chair rise, stairs and falls; sarcopenia algorithms may incorporate chair-rise time or gait performance. If these are then related to a function composite or fall history, shared content must be considered before attributing the association to measured muscle force. This is particularly important in PD, where slowing, freezing, postural instability and difficulty initiating movement can impair functional performance without a one-to-one loss of maximal force.

Recalled falls clinically informative history not future outcome evidence

Allen et al. studied 40 independently walking people with mild-to-moderate PD, approximately one hour after their usual medication. A Keiser A420 seated leg press measured each leg’s 1RM force in N and power in W. After 30 minutes’ rest, participants lifted ascending loads from 20% to 90% 1RM as fast as possible. Peak power was the highest value across loads; power at 30% 1RM was the highest at that specific load. Analyses used the average of both legs. Comfortable and maximal walking speeds were timed over the middle 10 m of a 14-m walkway. Full joint-position, fixation and repetition details are referred to the earlier protocol rather than completely specified in this article. [31]

The frequently cited R²=.54 applies specifically to concurrent maximal walking speed and power at 30% 1RM. For comfortable speed, the corresponding R² was .33; strength alone explained .43 and .26 for maximal and comfortable speed, respectively. After adding UPDRS motor score, the 30%-load power coefficients were .0016 m/s per W for maximal speed (95% CI .0011–.0021, p<.001) and .0007 for comfortable speed (.0003–.0011, p=.001). Whole-model R² values were .66 and .56. Strength also remained associated with both walking speeds in separate UPDRS-adjusted models. Age and height were considered in separate two-predictor models, rather than in one fully adjusted model. The descriptive R² pattern favours low-load power for fast walking in this sample, but no formal test established an independent power increment over strength. Selection of the 30% measure for further walking analysis followed its stronger within-sample association.

Falls were recalled over the previous 12 months: seven participants reported one fall and ten reported two or more. Logistic models compared those ten recurrent fallers with participants reporting zero or one fall. For the reported univariate analyses, power and strength were dichotomized at their sample medians: peak power 392.1 W, 30%-load power 289.1 W and strength 840.5 N. Both low-power contrasts yielded OR 6.0 (95% CI 1.1–33.3, p=.04); low strength yielded OR 3.1 (.7–14.1, p=.15). These are odds of recalled recurrent-faller status, not sixfold future risk or validated clinical cutoffs.

Importantly, the adjusted models used continuous power, not the median split. The source reports the 30%-load power OR as 1.0 (95% CI .99–1.0), p=.09 with UPDRS motor score, and the same rounded OR/CI with p=.047 in the separate age-adjusted model. This coarse rounding prevents a useful precise effect-size interpretation or rescaling to 100 W; p=.09 does not establish equivalence or absence of association. With only ten recurrent fallers, a wide unadjusted interval, recalled outcomes, possible reverse causation and no prediction validation, the study supports concurrent gait relevance more securely than an independent falls claim. Its Results explicitly identify these participants as those in Allen 2009, so the two articles do not provide independent PD cohorts. The complete original article and all three tables were examined.

The 2024 specific-leg-strength study assessed 95 people with PD using MicroFET2 testing. Fall status meant falls during the preceding year; 42 of 92 participants with fall records were fallers. A forward-selected model reported a hip-extensor association with fall history (OR .562, 95% CI .352–.900), alongside associations with slow walking (.380, .203–.713) and freezing (.364, .155–.852). Hip-abductor strength related to impaired balance (.360, .204–.636). Strength was cohort-z-scored. Table 5 specifies adjustment for age and levodopa-equivalent dose; the Methods also mention sex, leaving a reporting ambiguity. Bonferroni-corrected bivariate falls associations did not remain significant. [52]

These odds ratios are per study-specific standardized exposure, not per kg or a validated clinical cutoff. Sparse freezing events, numerous candidate muscle groups and outcomes, and forward selection increase model instability. A z-score is a change of units; it does not by itself remove sex or size confounding. Most importantly, the data cannot determine whether weakness preceded falls, falls led to reduced activity and weakness, or both reflected greater PD burden.

The Brazilian 2020 study illustrates the composite problem directly. Of 218 PD participants, 210 had falls information and 92 reported a fall during the previous six months. Grip did not differ significantly between fallers and nonfallers (median 21.3 versus 22.0 kg, p=.291). Probable sarcopenia was associated in bivariate analysis, but only SARC-F positivity and disease duration survived the final falls model. Because SARC-F itself includes a falls item, its association with recalled falls cannot be treated as independent force-based prognostic validation. Grip was assessed ON, using three trials per hand and the better hand’s trial mean. [58]

Skinner and Needle’s 15-PD/15-control case-control study reported correlations of weaker ankle plantarflexor and dorsiflexor strength with fall counts (rho=−.69/−.67), and associations involving submaximal steadiness. The accessible abstract does not specify a prospective falls collection window. It is therefore retained as a concurrent/fall-history lead, not forecast evidence. The article was published online in December 2024 and in the 2025 volume; the issue year is used for its citation. [62]

Table 8 Concurrent function and recalled falls

All rows are concurrent or retrospective associations. A regression predictor in these studies does not establish a future-outcome forecast. Complete model performance is not attributed to strength or power alone.

Table 8 Concurrent function and recalled falls
StudyTemporal design and assayPrincipal findingInterpretation and access
Schilling 2009 [54]Current function; 17 PD. Isometric leg-press force/body mass.TUG r=−.68, p=.003.Unadjusted task association. Primary abstract.
Nocera 2010 [55]Current function; 44 PD. Isokinetic knee torque, 60°/s.Gait-initiation COP–COM r=.500; 6MWD r=.248, p=.127.Task-specific; no adjusted forecast. Original body and tables.
Paul 2013 [5]Current function; 82 PD. Keiser power at 40% 1RM.Leg power–STS beta=.5; hip power–fast gait beta=.3.Adjusted for motor/cognitive and demographic factors; reactive-balance models nonsignificant. Original tables.
Caetano 2019 [56]Current laboratory gait; 54 PD. Hip power at fixed 35 N.Power associated with selected step lengths; errors and accuracy mainly with cognition/balance.Stepwise in-sample models; no later falls. Full original.
Ingram 2021 [7]Current upper-limb disability; 34 PD. Dominant grip.Grip–DASH r=−.562 OFF and −.549 ON.Unadjusted; no future endpoint. Full original.
Allen 2010 [31]Prior-year recalled recurrent falls and current gait; 40 PD, ten recurrent fallers.Median-split power OR 6.0 (1.1–33.3); maximal-gait R²=.54 for 30%-1RM power.Continuous-power falls p=.09 with motor score and .047 with age in separate models. Full original; sample medians are not cutoffs.
Specific leg strength 2024 [52]Past-year falls and current motor features; 95 PD. HHD.Hip-extension z-score OR .562 (.352–.900) for past falls.Forward selection; adjustment-reporting ambiguity and multiplicity. Full original.

Future falls heterogeneous strength findings

Latt 2009 positive adjusted weaker leg force association

Latt et al. provide the clearest positive adjusted strength finding among the original prospective falls cohorts appraised here. All 113 community-dwelling participants completed 12 months of monthly falls calendars with monthly telephone verification; 51 fell at least once, reporting 2,160 falls altogether. The outcome used for logistic modelling was participant-level occurrence of at least one fall, not the number of falls. Participants could walk without an aid and had MMSE≥24. One researcher tested them, usually mid-morning, in their self-reported medication ON state. Maximal isometric knee-extension, knee-flexion and ankle-dorsiflexion force was measured using spring gauges, reported as kg force. The paper does not specify joint angles, fixation, repetitions or trial aggregation sufficiently to reconstruct the entire dynamometry protocol. The selected strength variable was the weaker leg’s knee-extension force, not a designated more-affected side. [9]

Greater weaker-leg force was associated with lower future-fall odds in both models: OR .34 (95% CI .14–.82; reported p=.009) in the explanatory model, conditional on freezing, frontal impairment, axial posture and coordinated stability; and OR .28 (.11–.68; p=.002) in the combined predictive model, conditional on freezing, coordinated stability and previous falls. These effects are per one sample-standard-deviation increase in force. Although the table labels the measurement in kg, its footnote explicitly says strength and coordinated stability were converted to sample SD scores; these are not per-kg effects, risk ratios, or a usable clinical cutoff. The univariate weaker-leg OR was .19 (.09–.41) per SD, while the stronger-leg association was not significant (.77, .52–1.13).

The models answer different questions. The explanatory candidate set deliberately excluded previous falls, UPDRS and HY because the authors regarded them as markers rather than explanations. The combined model allowed all candidate variables, including previous falls, and strength remained associated. The published analysis selected the best set of variables and reports that collinearity limited entry to one strength measure, but does not clearly specify the numerical entry/removal criteria, selection algorithm or whether confounders such as age and sex were forced into each model. It should therefore not be described as a fully prespecified age/sex-adjusted model.

Explanatory-model classification was 39/51 fallers and 51/62 nonfallers, or 76.5% sensitivity calculated from 39/51 (published as 77%) and 82% specificity. The combined model classified 39/51 and 50/62, or 76.5% calculated from 39/51 (published as 77%) and 81%. These are apparent performances in the development sample. No bootstrap correction, cross-validation, external validation, AUC, calibration or formal incremental comparison isolating strength was reported. Many candidate measurements were examined against 51 participant-level events, so selection optimism remains a concern. The paper also contains a clinical-only model discrepancy: the abstract gives 38/51 and 45/62 (75%/73%), whereas Results gives 40/51 and 47/62 (78%/76%). This inconsistency does not change the Table 5 strength estimates.

These results provide substantive prospective adjusted evidence, including persistence after accounting for past falls. They do not establish that weakness causes falls or that strengthening prevents them. It also does not directly contradict every later result: Latt used weaker-leg standardized force and odds ratios, whereas Paul’s adjusted analysis used mean bilateral force and risk ratios; Kerr excluded strength under its stricter univariate entry gate. These differences, candidate adjustment sets and participant spectra should accompany the heterogeneous overall prognosis conclusion.

Kerr 2010 prospective group difference screened out before adjustment

Kerr et al. obtained complete six-month calendars from 101 independently walking, optimally medicated participants: 48 fell, including 24 recurrent fallers. Knee-extension force was lower in future fallers (27.4±10.3 versus 33.0±16.1 kg, p=.045); knee-flexion and ankle-dorsiflexion differences were nonsignificant. Crucially, multivariable entry required univariate p<.01. Knee strength therefore did not qualify for the final model. This is not an adjusted-null strength result. [10]

The selected model combined UPDRS total, freezing, symptomatic orthostasis, Tinetti total and anteroposterior sway. Sensitivity/specificity declined from 78%/84% apparent performance to 72%/76% with leave-one-out validation. These values cannot be attributed to strength. Five participants lacked complete falls calendars; the analysed sample also excluded baseline noncompleters and walking-aid users. The narrow ambulatory spectrum and significance-gated selection limit generalization and leave strength’s increment unresolved. A variable can fail a statistical entry criterion without having demonstrated zero prognostic information.

Paul et al. measured seated maximal knee-extension force using a spring gauge, three trials per leg, retaining each leg’s best result. All 205 participants had six-month falls information; 120 fell. Mean bilateral force was lower among fallers (30.7±12.0 versus 35.2±11.1 kg). Unadjusted RR was .99/kg (95% CI .98–1.00; p=.006); standardized RR was .85. In the initial multidomain model, the strength RR was .99 (.98–1.01), p=.37, and strength was removed. [11]

The initial model contained freezing severity, coordinated stability, repeated chair rise, pull-test impairment, proprioception, strength, frontal function and orientation. It deliberately excluded prior falls to examine potentially explanatory impairments. Backward elimination used a .20 removal threshold. Final-model AUC=.73 (95% CI .66–.80) is apparent multidomain discrimination, with no external validation reported. The cohort was selected partly from exercise-trial control groups and excluded MMSE<24. This supports attenuation conditional on these selected covariates in this population, not universal causal irrelevance of weakness.

Lima 2025 prospective grip association without a reported adjusted grip estimate

A newer Brazilian cohort followed 103 people with mild-to-moderate PD for 12 months using fall forms and monthly calls: 48 fell, 23 fell recurrently, and 159 falls were recorded. ON-state grip used a SAEHAN dynamometer, three alternating trials per hand and the maximum of all six. Baseline grip was lower in future fallers (27±10 versus 31±11 kg, p=.021), but did not distinguish recurrent fallers (28±11 versus 29±10 kg, p=.598). The displayed adjusted models omit a grip coefficient; therefore they cannot be cited as a quantified adjusted-null grip effect. Confirmed sarcopenia did not predict either outcome. [59]

Several reporting issues materially affect use. Methods describe logistic regression but Table 2 calls the estimates hazard ratios; this inconsistency should remain explicit rather than silently converting their labels. AUC=.843 refers to SARC-F plus disease duration, not dynamometry. Table 3’s “handgrip strength” accuracy is the SARC-F self-report strength item, not the measured grip test. Selection after bivariate testing, correlated-variable exclusion, many candidate covariates relative to 23 recurrent events, and no reported external validation limit optimism control. The maximum reported classification accuracy of 78.64% is only slightly above the calculated 80/103=77.67% obtained by labelling everyone non-recurrent; without sensitivity, specificity and validation it is not sufficient evidence of clinical usefulness. This cohort is temporally distinct from the same centre’s 2020 cross-sectional report, but individual overlap is not established.

Cohort overlap prevents counting two Paul papers as replication

The 2013 three-test clinical prediction paper and 2014 explanatory paper concern the same 205-person cohort and six-month falls outcome. They ask different questions. The 2013 paper selected a practical tool based on previous falls, freezing and walking speed and performed bootstrap internal validation; the 2014 paper examined impairments while excluding past falls. They are not independent replications. The 2013 original Results and table identify 120 fallers; its abstract’s “125 (59%)” is internally inconsistent, so the body denominator is used here. [63]

The 82-person Paul functional study is a different analysis and must carry its own DOI and sample size. Shared investigators and apparatus also mean that studies from this research programme should not automatically be called fully independent cohorts. Exact participant overlap with the 31-person reliability study, earlier Allen sample or other machine-power reports was not established here. It should be flagged as uncertain rather than guessed.

Synthesis of prospective falls evidence

The prospective studies support a more qualified conclusion than either “strength predicts falls” or “strength has no value.” Latt supplies a positive per-SD weaker-leg force association, including a model accounting for previous falls. Kerr supplies a pre-outcome group difference that failed its stricter modelling-entry criterion. Paul supplies a directly estimated strength effect that attenuated in a multidomain model. Lima adds an any-fall grip difference but no recurrent-fall separation or reported adjusted grip estimate. Differences in participant spectrum, assay, follow-up duration, outcome coding and candidate predictors can all contribute. Their effect measures should not be pooled as though they estimate one identical exposure–outcome relationship.

For a clinically deployable prediction claim, a next study would need to show more than a significant coefficient. It should prespecify the dynamometry protocol and time horizon; measure strength before the outcome; count sufficient events; account for missingness; include established predictors; report absolute-risk calibration and discrimination; quantify what strength adds; and test the complete model in new patients. Current reviewed evidence does not establish a universal PD-specific force, power or RFD cutoff that independently changes an individual’s management.

Table 9 Strength and genuine future falls

Confidence intervals are 95%. The Paul 2013 clinical tool and Paul 2014 explanatory analysis share one 205-person cohort. Their findings are not independent replications.

Table 9 Strength and genuine future falls
StudyTemporal designStrength findingModel and interpretation
Latt 2009 [9]12 months; 113 PD, 51 fallers. ON-state spring-gauge weaker knee-extension force.Per sample SD: unadjusted OR .19 (.09–.41); explanatory .34 (.14–.82); combined .28 (.11–.68).Adjusted association persists with past falls. Apparent models 76.5%/82% and 76.5%/81% (each sensitivity calculated from 39/51; published as 77%); no validation or isolated increment. Complete original HTML.
Kerr 2010 [10]6 months; 101 PD, 48 fallers, 24 recurrent.Knee force 27.4±10.3 versus 33.0±16.1 kg, p=.045.Missed p<.01 model-entry gate. Not an adjusted-null result. Original report described author-manuscript/Table2 access; current independent comparison is abstract-only and those details remain unverified.
Paul 2014 [11]6 months; 205 PD, 120 fallers.Unadjusted RR .99/kg (.98–1.00), p=.006; adjusted .99 (.98–1.01), p=.37.Strength removed from multidomain model. Whole-model AUC .73. Original publisher tables.
Paul 2013 falls tool [63]Same 205-person cohort and 6-month endpoint.Univariate knee-force OR .97/kg (.94–.99), p=.008.Selected tool uses past falls, freezing and gait speed; bootstrap validation does not establish strength increment.
Lima 2025 [59]12 months; 103 PD, 48 fallers, 23 recurrent.Grip 27±10 versus 31±11 kg, p=.021; recurrent comparison p=.598.No adjusted grip coefficient displayed. Logistic-method/HR-label inconsistency. Full original.

Progression disability and mortality much less direct evidence

Saarinen 2025 longitudinal monitoring association without baseline prognostic validation

Saarinen et al. assessed 147 PD participants, including 106 initially unmedicated, plus 35 controls. Eighty-four returned for clinical follow-up (median 4.1 years), and 40 for repeat DAT imaging (median 6.2 years). Grip used a seated Jamar, elbow 90°, one practice and two 3–5-second maximal trials per hand, retaining the best. Mean grip was stable: 32.5→32.0 kg in the clinical cohort and 32.2→31.7 kg in the imaging subset. Individual worsening of motor score was associated with declining grip after age/sex adjustment (p=.029), whereas annual grip change was not independently related to annual DAT change (p>.62). [12]

Interpretation must respect two levels of observation. Tables 2–3 show improved mean MDS-UPDRS motor scores (34.1→29.7 and 35.0→29.2) while medication doses increased; quality of life worsened and DAT binding declined. A within-person change association does not require the whole group to worsen. The study does not demonstrate that baseline grip forecasts later motor decline, falls, disability or survival, nor that grip is a surrogate for dopamine-terminal loss.

Several cautions apply beyond statistical significance. Follow-up participants volunteered and represent 84/147 for clinical outcomes and 40/147 for imaging, so selective return may affect the observed trajectory. Intervals varied considerably. Treatment changed between observations; measured motor state is therefore not an untreated natural-history endpoint. Two-timepoint change scores contain measurement error, and a nonsignificant group mean can conceal individual deterioration and improvement. Conversely, an individual change–change correlation cannot by itself establish assay responsiveness to clinically important deterioration. The relevant supplementary coefficients/CIs were not recovered, so the report retains the verified p-values without inventing precision.

Earlier longitudinal imaging evidence separates maximum force from force control

The 2011 imaging study retrospectively analysed repeated observations over approximately four years in 26 people with PD and 11 controls. Nineteen PD participants had three sessions and seven had two; mean intersession interval was 1.8 years. Maximal pinch force was measured between thumb and index finger, whereas a separate visually guided 200-g hold quantified fine force control. Repeated-measures models linked force-control abnormalities to striatal FDOPA uptake. Maximal hand strength was similar to controls and was not the highlighted longitudinal imaging correlate. [64]

This does not contradict Saarinen’s maximal-grip findings: pinch force control and maximal whole-hand grip are different constructs, and FDOPA uptake and DAT binding are different imaging measures. Baseline motor testing was OFF medication whereas later testing used usual medication, with incomplete dose records. That state change compromises a natural-history interpretation. Mixed-effects relationships between repeated contemporaneous measurements are not a validated baseline forecast of later disability. Multiple separately fitted models, a small early-PD sample and outlier replacement also warrant caution; the article’s claim that separate models obviate multiplicity adjustment should not be adopted uncritically.

A rehabilitation cohort predicts final walking level not improvement

Gialanella et al.’s 55-person prospective outpatient rehabilitation cohort found that baseline right “Grip and Pinch” and Berg Balance Scale scores related to final six-minute walk distance (standardized beta=.36, p=.001 and .47, p<.001 respectively; model R²=.47, adjusted R²=.45). No measured variable predicted the gain in walking distance. The original abstract therefore supplies limited temporally predictive evidence for an endpoint under rehabilitation, but not treatment response, untreated progression or long-term disability. Exact strength operationalization, duration, baseline-walking-distance adjustment and model validation require the original Methods. A participant starting with better overall function may finish with better function without gaining more. [65]

Genetic evidence is mechanistic context not a dynamometer prediction rule

Wang et al. used Mendelian randomization and polygenic scores rather than baseline bedside grip as the exposure. Genetically higher right-hand grip was associated with lower levodopa-induced dyskinesia risk; the MR OR was .152 (95% CI .055–.423), with a corresponding polygenic-score association in PPMI. The analyses did not support a causal relationship with progression to HY3 or dementia. This is important context for interpreting Saarinen’s citation of grip and dyskinesia, but it does not validate a measured-kg risk threshold or show that increasing grip through training prevents dyskinesia. Lifelong genetic propensity is not equivalent to an intervention-induced change or a single clinic measurement. [66]

Mortality evidence remains limited

Gray et al. related baseline physical assessment in 109 people with PD to mortality over seven years. The battery included strength, but the abstract’s final Cox model retained age, sex and Tinetti gait. Strength is not among the abstract-listed significant physical predictors. This does not justify either a strength mortality claim or a definitive tested-null conclusion: the original strength rows, assay, event count and adjusted estimates remain unavailable. The defensible statement is that this mortality study does not presently provide verified evidence of an independent strength prognosis. [67]

Longitudinal exercise cohorts may repeatedly measure grip as an outcome while predicting it from exercise behaviour, age or disease severity. This reverses the direction required for the present prognostic question. For example, Combs-Miller and Moore’s 74-person, two-year study treats grip as an impairment outcome; it should not be cited as proof that initial grip predicts later activity or participation. Treatment-response studies and descriptive trajectories can inform future measurement design but do not substitute for clinical prognostic validation. [68]

Fracture adjustment is not an independent strength effect

Schini and colleagues examined fracture risk in a trial cohort of women aged 75 years or older. The analyses included 43 self-reported definite PD cases and four probable cases who reported dopaminergic treatment. The complete main article and supplement were available for appraisal. Maximum isometric quadriceps strength was measured with a load-cell device. The displayed hazard ratios quantify the association of PD status with later osteoporotic fracture after covariate adjustment; they are not independent hazard ratios per unit of quadriceps strength. The height- and treatment-adjusted PD hazard ratio was 2.25 (95% CI 1.24–4.08; n=5161), compared with 1.70 (0.81–3.59; n=4611) after adding maximum quadriceps strength. The smaller complete-case sample in the latter model prevents attributing all attenuation to strength. These results neither establish mediation nor validate a within-PD strength threshold for future fracture. [69]

A 2026 prospective biomarker study does not establish strengths contribution

Qaisar et al. followed 233 PD participants without freezing for 24 months; 55 developed questionnaire-defined freezing. Its primary exposure was plasma CAF22, with muscle strength included among covariates alongside age, sex, duration, motor severity, gait speed, cognition and medication. The abstract provides CAF22 and gait estimates but no independent strength coefficient or incremental-performance analysis. Consequently, this new original is a retrieval priority, not verified support for dynamometric prediction. It also illustrates why a prospective study that happens to measure strength is not automatically a strength-prognosis study. [70]

Exclude incident PD evidence from established PD prognosis

Large population cohorts relating lower grip or greater asymmetry to subsequent PD diagnosis address disease incidence among initially non-PD participants. They do not answer whether a person already diagnosed will fall, become disabled or progress more quickly. The UK Biobank grip/walking studies and a 71,702-person older-adult longitudinal study should be described as a separate etiological/prodromal literature, with attention to prediagnostic impairment and overlapping source cohorts. Their large sample sizes cannot compensate for the wrong target population when making a post-diagnosis prediction claim. [71]; [72]; [73]

Similarly, the German National Cohort paper’s mortality analysis concerns PD status, not handgrip as the prognostic exposure. A paper can provide a useful cross-sectional grip comparison and a separate mortality analysis without establishing that grip predicts mortality. General older-adult studies excluding PD, autonomic handgrip tests measuring blood-pressure response, and non-PD animal models should also remain outside the direct clinical evidence set. [74]

Table 10 Longitudinal monitoring and other future outcomes

A repeated-measures association, endpoint-level prediction under treatment and genetic association have different interpretations. They are not pooled as direct bedside-strength prognosis.

Table 10 Longitudinal monitoring and other future outcomes
StudyTemporal designMain resultDefensible interpretation
Gallagher 2011 [64]Repeated measures over about 4 years; 26 PD, 11 controls.Fine pinch-control measures related to FDOPA; maximum pinch not the highlighted correlate.Monitoring/physiology; baseline OFF and follow-up usual medication. Original body, quantitative tables incomplete.
Saarinen 2025 [12]Clinical follow-up 84/147; repeat DAT imaging 40/147.Mean grip stable; grip–motor change p=.029; adjusted grip–DAT change p>.62.Change–change association; no baseline forecasting model. Full original PDF and body.
Gialanella 2022 [65]55 PD undergoing rehabilitation; later 6MWD.Baseline right Grip and Pinch beta=.36, p=.001 for final 6MWD; no predictor of gain.Final level is not response magnitude. Baseline 6MWD adjustment unverified. Primary abstract.
Gray 2009 [67]Mortality over 7 years; 109 PD.Strength in baseline battery; final abstract model lists age, sex and Tinetti gait.Neither positive nor definitive null strength prognosis; coefficient, events and protocol unavailable.
Wang 2024 [66]Mendelian randomization and PPMI genetic-score survival.Genetically predicted right grip related to dyskinesia; no MR support for HY3 or dementia.Genetic propensity is not a measured-grip threshold or training effect. Full original.
Qaisar 2026 [70]Incident freezing over 24 months; 233 PD, 55 events.CAF22 is primary exposure; strength included as covariate without reported coefficient.Prospective design alone does not establish strength prognosis. Primary abstract/publisher preview.

Implications for rehabtools and clinical assessment

Present a purpose specific choice of assessment

Rehabtools should help a clinician choose and reproduce a test, then explain what its result can support. A short hierarchy based only on ICC would be misleading. The most reproducible assay in one small laboratory sample may be impractical for a particular patient or insufficiently precise for the change of interest. The appropriate choice depends on the impairment being investigated, access to equipment, transfer and balance requirements, the clinician’s ability to stabilize the setup and whether the intended use is description, exercise prescription, monitoring or prognosis.

A brief hand-force assessment can be valuable when upper-limb function, general reserve or a sarcopenia assessment is relevant. A patient’s difficulty with standing or stepping calls for task-relevant lower-limb assessment and direct observation of function. Hip, knee and ankle deficits need not be proportional. For exercise loading, a standardized machine 1RM gives information that a grip test does not. For rapid dynamic capacity, machine power provides velocity-sensitive information at the tested load. For early RFD or segmentation, the burden shifts to rigid fixation, a suitable force transducer, careful instructions and a validated analysis pipeline.

Functional performance should be retained beside impairment measures. Strength may be adequate while balance, coordination, cognition or freezing limits a task; conversely, compensatory movement may permit task completion despite weakness. A lower-limb test and a chair-rise test therefore complement each other. Their correlation does not justify replacing one with the other, and a computed chair-rise power score should retain its own model name and assumptions.

Table 11 Assessment choice by purpose

Table 11 Assessment choice by purpose
Clinical questionUseful measurement approachMain interpretive requirement
Hand-force impairmentStandardized bilateral gripMatch posture, handle, trials, side and comparison source; consider hand pathology.
Task-relevant lower-limb weaknessFixed or adequately stabilized isometric force/torquePreserve joint angle, lever arm, fixation and muscle action. Use alongside functional observation.
Exercise-specific loadingFamiliarized machine 1RMSame equipment and range; repeat-baseline planning; monitor learning, state and fatigue.
Dynamic rapid capacityMachine power at known loadsReport force/velocity method, absolute loads, relative-load definition and peak/mean rule.
Early force or pulse-control impairmentHigh-bandwidth fixed force/torque or specialized segmentation assayLock acquisition and processing; establish precision for the full local pipeline.
Clinical changeExact repeat protocol plus patient/function anchorSeparate random error, bias, importance and state effects; retain raw and normalized values.
Future falls or progressionMultidomain clinical assessment with validated model where availableStrength alone lacks a broadly validated cutoff; identify outcome, horizon, calibration and external validation.

Build each assessment entry around a reproducible protocol

Each rehabtools entry should identify the exact construct, population studied, equipment and scoring rule before describing results. A useful entry would state whether it concerns force, joint torque, exercise-specific maximum load, power, an RFD interval or submaximal force control. It should then identify the relevant article and distinguish directly examined protocol details from practical standardization advice supplied by the report.

The minimum record for an individual measurement should include instrument and calibration status; position and joint angles; fixation and contact point; external moment arm where relevant; side and dominance; clinically more-affected side; observed medication state and time since dose; pain or movement restrictions; practice, effort instruction, trial duration and rest; all valid trial values; and the prespecified summary rule. Power needs the load and velocity method. RFD needs acquisition rate, filtering, onset rule, baseline rejection criteria and time windows. Values cannot be reconstructed reliably from a single final score when those details are lost.

A missing or failed test is clinically informative. Record whether noncompletion reflects balance risk, an unsafe transfer, pain, cognitive or communication difficulty, minimum device resistance, fatigue, technical failure or inability to sustain the intended position. Entering zero would combine failure to measure with genuinely absent force and distort both individual follow-up and future reference datasets.

Medication state deserves more than an ON/OFF checkbox. Retest at a comparable clinically relevant state, record timing and observed fluctuations, and retain the original assessment purpose. An OFF protocol should only be used within an agreed clinical or research plan; the measurement instructions should not advise patients to withhold medication independently. A person’s usual best ON performance and the extent of OFF impairment are both potentially useful, but they answer different questions.

Communicate reliability and change without overclaiming

An evidence panel should place sample size, repeat interval, assessor design, ICC model and confidence interval beside absolute-error information. For example, Paul’s leg-power ICCs around .95–.96 describe strong same-assessor relative reliability in the tested ON-state cohort. They coexist with SEMs of 35–36 W and calculated two-score MDC95 values of approximately 97–100 W. A phrase such as “highly reliable” without that context encourages interpretation of changes smaller than the observed noise. [1]

Where a reported error statistic is ambiguous or internally inconsistent, the appropriate display is “individual change threshold not established from this study,” followed by the specific reason. Villafañe’s 0.092 kg true-change claim should not appear as a usable threshold. Bonde-Jensen’s values are explicitly labelled SDC in the full article, but they conflict with wider paired agreement limits; the proposed 5.7%/7.9% individual-change rules should not be adopted. A calculated MDC should be visibly distinguished from a published value and from an anchor-based important change. [3, 4]

A longitudinal result should show baseline and follow-up raw values, the absolute difference and any normalized values with their denominators. State differences, altered setup and systematic learning must be considered before attributing a gain to treatment or a loss to disease progression. A statistically unusual change can still be unimportant to the patient; an important functional improvement may occur without a large change in isolated force. Repeating an unexpected result under matched conditions is often a more defensible next step than immediately assigning a disease label.

Keep impairment assessment separate from risk classification

The evidence supports a role for strength and power within a multidomain assessment of a person with PD. It does not support a stand-alone “high fall risk” label based on an unvalidated force cutoff. Risk communication should identify the outcome, time horizon, derivation population and validation status of any model used. The sensitivity and specificity of an entire clinical-plus-physiological model must not be presented as the diagnostic performance of knee strength. [9, 10, 11]

The same restriction applies to progression. Weakness may be a relevant treatment target and a marker of reduced functional reserve even when it is not a specific biomarker of neurodegeneration. No appraised strength-, power- or RFD-only tool established externally validated prediction with calibration and demonstrated incremental value across later falls, disability or PD progression. This conclusion describes the evidence examined for this report; it is not a claim that future studies cannot establish such a role.

Training evidence and assessment validity

Resistance and power training trials address whether an intervention changes a measured outcome under their treatment and testing conditions. This report does not estimate their comparative effectiveness, optimum dose or long-term safety. A training study can be informative about the feasibility and range of an assay, but a significant mean gain alone does not establish device agreement, an important individual change or predictive validity.

This distinction is especially important when training and testing use the same machine or movement. Improvement can include familiarity with the task and measurement setup as well as a change in force-generating capacity. Buckley and Hass demonstrated retest gains even without an intervention, after familiarization. A well-designed treatment trial may control such effects through its comparison group, but the resulting group treatment estimate still does not supply a universal individual MDC or MIC. [2]

For rehabtools, evidence that an impairment is potentially trainable can motivate assessment, while the assessment page should remain explicit about its separate validation evidence. Neither a cross-sectional strength–function association nor genetically associated grip propensity establishes that strengthening prevents PD, dyskinesia, falls or neurodegenerative progression.

Priorities for a useful validation programme

The first priority is repeatability of the complete proposed protocol in the intended users and patient population, including different assessors where that reflects real practice. Report bias, SEM or typical error, limits of agreement, confidence intervals, failed-test rates and adverse events alongside ICC. Examine whether error depends on score magnitude and whether medication state, dyskinesia, stage or cognition alters feasibility and precision.

The next priority is agreement when substituting a new device, sensor or algorithm for an established measurement. Use paired measurements that permit inspection of bias across the relevant range; correlation alone is insufficient. Any software-derived power or RFD result needs versioned processing and independent reference validation appropriate to its physical model.

Finally, assess longitudinal interpretation with external patient or functional anchors and prospective outcome studies. Prespecify the increment being sought over a practical clinical baseline model, account for overlapping cohorts and missing follow-up, and report both discrimination and calibration. This sequence would turn an interesting impairment signal into a more defensible clinical tool without treating a group difference or a training response as a shortcut to validation.

References

Numbering follows first citation. Linked DOIs identify original articles. Complete original texts, partial primary material and abstract-only evidence are distinguished in the relevant narrative and tables. General methodological sources are explicitly separated from PD validation evidence.

1. Paul SS, Canning CG, Sherrington C, Fung VS. Reproducibility of measures of leg muscle power, leg muscle strength, postural sway and mobility in people with Parkinson's disease. Gait Posture. 2012;36(3):639–42. DOI 10.1016/j.gaitpost.2012.04.013

Source note: SRC-b3a13e915936 Paul SS 2012

2. Buckley TA, Hass CJ. Reliability in One-Repetition Maximum Performance in People with Parkinson's Disease. Parkinson's Disease. 2012;2012:928736. DOI 10.1155/2012/928736

Source note: SRC-1cbd2683da0a Buckley TA 2012

3. Villafañe JH, Valdes K, Buraschi R, Martinelli M, Bissolotti L, Negrini S. Reliability of the Handgrip Strength Test in Elderly Subjects With Parkinson Disease. Hand (N Y). 2016;11(1):54–8. DOI 10.1177/1558944715614852

Source note: SRC-d361f5579453 Villafane JH 2016

4. Bonde-Jensen F, Dalgas U, Hvid LG, Langeskov-Christensen M. Validity and reliability of linear encoder muscle power testing in persons with Parkinson's disease. Clin Rehabil. 2024;38(5):678–687. DOI 10.1177/02692155231224987

Source note: SRC-13c37debabd6 Bonde-Jensen F 2024

5. Paul SS, Sherrington C, Fung VS, Canning CG. Motor and cognitive impairments in Parkinson disease: relationships with specific balance and mobility tasks. Neurorehabil Neural Repair. 2013;27(1):63–71. DOI 10.1177/1545968312446754

Source note: SRC-0a278cb1e1a9 Paul SS 2013

6. Stevens-Lapsley J, Kluger BM, Schenkman M. Quadriceps Muscle Weakness, Activation Deficits, and Fatigue With Parkinson Disease. Neurorehabilitation and Neural Repair. 2012;26:533–541. DOI 10.1177/1545968311425925

Source note: SRC-19c61f7f2943 Stevens-Lapsley J 2012

7. Ingram LA, Carroll VK, Butler AA, Brodie MA, Gandevia SC, Lord SR. Quantifying upper limb motor impairment in people with Parkinson's disease: a physiological profiling approach. PeerJ. 2021;9:e10735. DOI 10.7717/peerj.10735

Source note: SRC-3745d69cd509 Ingram LA 2021

8. Villamil-Cabello E, Molinero-Martín E, Fogelson N, Luque-Casado A, Fernández-Del-Olmo MÁ. Maximal Handgrip Strength as a Measure of General Functional Capacity Rather than Parkinson's Disease-Specific Impairment. Brain Sci. 2026;16(8):871. DOI 10.3390/brainsci16080871

Source note: SRC-e7465cb52193 Villamil-Cabello E 2026

9. Latt MD, Lord SR, Morris JGL, Fung VSC. Clinical and physiological assessments for elucidating falls risk in Parkinson's disease. Movement Disorders. 2009;24(9):1280–1289. DOI 10.1002/mds.22561

Source note: SRC-16974e3ab727 Latt MD 2009

10. Kerr G, Worringham C, Cole M, Lacherez P, Wood J, Silburn P. Predictors of future falls in Parkinson disease. Neurology. 2010;75(2):116-124. DOI 10.1212/wnl.0b013e3181e7b688

Source note: SRC-1e5e4341a16b Kerr GK 2010

11. Paul SS, Sherrington C, Canning CG, Fung VS, Close JC, Lord SR. The relative contribution of physical and cognitive fall risk factors in people with Parkinson's disease: a large prospective cohort study. Neurorehabil Neural Repair. 2014;28(3):282–90. DOI 10.1177/1545968313508470

Source note: SRC-420a8a49c874 Paul SS 2014

12. Saarinen EK, Kuusimäki T, Niemi K, Noponen T, Jaakkola E, Myller E, et al. Hand muscle strength in Parkinson's disease: A Sarcopenic epiphenomenon or a meaningful biomarker? Parkinsonism Relat Disord. 2025;140:108021. DOI 10.1016/j.parkreldis.2025.108021

Source note: SRC-6c4515063bac Saarinen EK 2025

13. Chung JW, Knight CA, Bower AE, Martello JP, Jeka JJ, Burciu RG. Rate control deficits during pinch grip and ankle dorsiflexion in early-stage Parkinson’s disease. PLOS ONE. 2023;18(3):e0282203. DOI 10.1371/journal.pone.0282203

Source note: SRC-6dd1c23f223e Chung JW 2023

14. Daniels RJ, Knight CA. Motor segmentation: a key neuromuscular impairment in people with parkinson’s disease. Experimental Brain Research. 2025;243(12):241. DOI 10.1007/s00221-025-07189-3

Source note: SRC-86d0bdb8a372 Daniels RJ 2025

15. Silva SM, Corrêa FI, Silva PF, Silva DF, Lucareli PR, Corrêa JC. Validation and reliability of a modified sphygmomanometer for the assessment of handgrip strength in Parkinson's disease. Braz J Phys Ther. 2015;19(2):137–45. DOI 10.1590/bjpt-rbf.2014.0081

Source note: SRC-abe11fdd1716 Silva SM 2015

16. Gamborg M, Hvid LG, Thrue C, Johansson S, Franzén E, Dalgas U, et al. Muscle Strength and Power in People With Parkinson Disease: A Systematic Review and Meta-analysis. J Neurol Phys Ther. 2023;47(1):3–15. DOI 10.1097/npt.0000000000000421

Source note: SRC-bdef544d3b91 Gamborg M 2023

17. Lima LO, Cardoso F, Teixeira-Salmela LF, Rodrigues-de-Paula F. Work and power reduced in L-dopa naïve patients in the early-stages of Parkinson’s disease. Arquivos de Neuro-Psiquiatria. 2016;74(4):287–292. DOI 10.1590/0004-282x20160014

Source note: SRC-0acf3ac4a484 Lima LO 2016

18. Purser JL, Pieper CF, Duncan PW, Gold DT, McConnell ES, Schenkman MS et al. Reliability of physical performance tests in four different randomized clinical trials. Archives of Physical Medicine and Rehabilitation. 1999;80(5):557-561. DOI 10.1016/s0003-9993(99)90199-5

Source note: SRC-53b708810a4f Purser JL 1999

19. Pang MYC, Mak MKY. Muscle strength is significantly associated with hip bone mineral density in women with Parkinson’s disease: a cross-sectional study. Journal of Rehabilitation Medicine. 2009;41(4):223–230. DOI 10.2340/16501977-0311

Source note: SRC-1f2362f153f7 Pang MYC 2009

20. Pang MYC, Mak MKY. Trunk muscle strength, but not trunk rigidity, is independently associated with bone mineral density of the lumbar spine in patients with Parkinson’s disease. Movement Disorders. 2009;24:1176–1182. DOI 10.1002/mds.22531

Source note: SRC-ceac3b04c4e1 Pang MYC 2009

21. Alota Ignacio Pereira V, Augusto Barbieri F, Moura Zagatto A, Cezar Rocha Dos Santos P, Simieli L, Augusto Barbieri R, et al. Muscle Fatigue Does Not Change the Effects on Lower Limbs Strength Caused by Aging and Parkinson's Disease. Aging Dis. 2018;9(6):988–998. DOI 10.14336/ad.2018.0203

Source note: SRC-53b09c9fb7c3 Alota Ignacio Pereira V 2018

22. Pääsuke M, Mõttus K, Ereline J, Gapeyeva H, Taba P. Lower limb performance in older female patients with Parkinson's disease. Aging Clin Exp Res. 2002;14(3):185–91. DOI 10.1007/bf03324434

Source note: SRC-fa38c4e10684 Paasuke M 2002

23. Mazza RO, Silva AEL, Machado LT, de Britto VLS, Paz TSR, Corrêa CL. Handgrip strength in Parkinson’s disease: A systematic review of observational studies. Fisioterapia em Movimento. 2024;37:e37203. DOI 10.1590/fm.2024.37203

Source note: SRC-f5be59247b48 Mazza RO 2024

24. Roberts HC, Syddall HE, Butchart JW, Stack EL, Cooper C, Sayer AA. The Association of Grip Strength With Severity and Duration of Parkinson’s: A Cross-Sectional Study. Neurorehabilitation and Neural Repair. 2015;29(9):889-896. DOI 10.1177/1545968315570324

Source note: SRC-3b687b6e7c46 Roberts HC 2015

25. Salmon R, Preston E, Mahendran N, Flynn A, Ada L. People with mild PD have impaired force production in all lower limb muscle groups: A cross-sectional study. Physiotherapy Research International. 2021;26(2):e1897. DOI 10.1002/pri.1897

Source note: SRC-c87ed52d1035 Salmon R 2021

26. Skinner JW, Christou EA, Hass CJ. Lower Extremity Muscle Strength and Force Variability in Persons With Parkinson Disease. Journal of Neurologic Physical Therapy. 2019;43(1):56-62. DOI 10.1097/npt.0000000000000244

Source note: SRC-badee83c6385 Skinner JW 2019

27. Inkster LM, Eng JJ, MacIntyre DL, Stoessl AJ. Leg muscle strength is reduced in Parkinson's disease and relates to the ability to rise from a chair. Movement Disorders. 2003;18(2):157-162. DOI 10.1002/mds.10299

Source note: SRC-e3a0ee4e4758 Inkster LM 2003

28. Frazzitta G, Ferrazzoli D, Maestri R, Rovescala R, Guaglio G, Bera R, et al. Differences in muscle strength in parkinsonian patients affected on the right and left side. PLoS One. 2015;10(3):e0121251. DOI 10.1371/journal.pone.0121251

Source note: SRC-0af88cd48519 Frazzitta G 2015

29. Bridgewater KJ, Sharpe MH. Trunk Muscle Performance in Early Parkinson's Disease. Physical Therapy. 1998;78(6):566-576. DOI 10.1093/ptj/78.6.566

Source note: SRC-aa6ea834c124 Bridgewater KJ 1998

30. Allen NE, Canning CG, Sherrington C, Fung VS. Bradykinesia, muscle weakness and reduced muscle power in Parkinson's disease. Mov Disord. 2009;24(9):1344–51. DOI 10.1002/mds.22609

Source note: SRC-4255d7e8932b Allen NE 2009

31. Allen NE, Sherrington C, Canning CG, Fung VS. Reduced muscle power is associated with slower walking velocity and falls in people with Parkinson's disease. Parkinsonism Relat Disord. 2010;16(4):261–4. DOI 10.1016/j.parkreldis.2009.12.011

Source note: SRC-bb84f39a6f1f Allen NE 2010

32. Stelmach GE, Teasdale N, Phillips J, Worringham CJ. Force production characteristics in Parkinson's disease. Exp Brain Res. 1989;76(1):165–72. DOI 10.1007/bf00253633

Source note: SRC-e33a330e27c5 Stelmach GE 1989

33. Park JH, Stelmach GE. Force development during target-directed isometric force production in Parkinson's disease. Neurosci Lett. 2007;412(2):173–8. DOI 10.1016/j.neulet.2006.11.009

Source note: SRC-c10e9501693e Park JH 2007

34. Rose MH, Løkkegaard A, Sonne-Holm S, Jensen BR. Tremor irregularity, torque steadiness and rate of force development in Parkinson's disease. Motor Control. 2013;17(2):203–16. DOI 10.1123/mcj.17.2.203

Source note: SRC-ccb0e4c7ba28 Rose MH 2013

35. Corcos DM, Chen CM, Quinn NP, McAuley J, Rothwell JC. Strength in Parkinson's disease: relationship to rate of force generation and clinical status. Ann Neurol. 1996;39(1):79–88. DOI 10.1002/ana.410390112

Source note: SRC-4bffffc3f5e5 Corcos DM 1996

36. Hammond KG, Pfeiffer RF, LeDoux MS, Schilling BK. Neuromuscular rate of force development deficit in Parkinson disease. Clin Biomech (Bristol). 2017;45:14–18. DOI 10.1016/j.clinbiomech.2017.04.003

Source note: SRC-11020cc11a18 Hammond KG 2017

37. Pelicioni PHS, Pereira MP, Lahr J, Dos Santos PCR, Gobbi LTB. Assessment of Force Production in Parkinson's Disease Subtypes. Int J Environ Res Public Health. 2021;18(19):10044. DOI 10.3390/ijerph181910044

Source note: SRC-cb32ad8ed78f Pelicioni PHS 2021

38. Howard SL, Grenet D, Bellumori M, Knight CA. Measures of motor segmentation from rapid isometric force pulses are reliable and differentiate Parkinson's disease from age-related slowing. Exp Brain Res. 2022;240(7-8):2205–2217. DOI 10.1007/s00221-022-06398-4

Source note: SRC-c3fe700cb48c Howard SL 2022

39. Monte A, Magris R, Nardello F, Bombieri F, Zamparo P. Muscle shape changes in Parkinson's disease impair function during rapid contractions. Acta Physiologica. 2023;238(1):e13957. DOI 10.1111/apha.13957

Source note: SRC-b449d169e919 Monte A 2023

40. Boom M, Preston E, Salmon R, Ada L, Flynn A. Reliability of Hand-Held Dynamometry for Measuring Force Production in People with Parkinson’s Disease. Internet Journal of Allied Health Sciences and Practice. 2023;21(1):Article 3. DOI 10.46743/1540-580x/2023.2176

Source note: SRC-2c6af206d6c5 Boom M 2023

41. Stone CA, Nolan B, Lawlor PG, Kenny RA. Hand-held dynamometry: tester strength is paramount, even in frail populations. J Rehabil Med. 2011;43:808–811. DOI 10.2340/16501977-0860

Source note: SRC-972984c7bce3 Stone CA 2011

42. Martins J, da Silva JR, da Silva MRB, Bevilaqua-Grossi D. Reliability and Validity of the Belt-Stabilized Handheld Dynamometer in Hip- and Knee-Strength Tests. J Athl Train. 2017;52:809–819. DOI 10.4085/1062-6050-52.6.04

Source note: SRC-b4c5f66e6b0b Martins J 2017

43. Florêncio LL, Martins J, da Silva MRB, et al. Knee and hip strength measurements obtained by a hand-held dynamometer stabilized by a belt and an examiner demonstrate parallel reliability but not agreement. Phys Ther Sport. 2019;38:115–122. DOI 10.1016/j.ptsp.2019.04.011

Source note: SRC-c8baad2ec3f4 Florencio LL 2019

44. Daniels RJ, Knight CA. EMG analysis and correlates of motor segmentation in Parkinson's disease. J Electromyogr Kinesiol. 2026;89:103161. DOI 10.1016/j.jelekin.2026.103161

Source note: SRC-1229b0abfd1e Daniels RJ 2026

45. Maffiuletti NA, Aagaard P, Blazevich AJ, Folland J, Tillin N, Duchateau J. Rate of force development: physiological and methodological considerations. Eur J Appl Physiol. 2016;116:1091–1116. DOI 10.1007/s00421-016-3346-6

Source note: SRC-3fc01b123ef6 Maffiuletti NA 2016

46. de Vet HCW, Terwee CB, Knol DL, Bouter LM. When to use agreement versus reliability measures. J Clin Epidemiol. 2006;59:1033–1039. DOI 10.1016/j.jclinepi.2005.10.015

Source note: SRC-c5c63e3d772c de Vet HCW 2006

47. de Vet HCW, et al. Minimal changes in health status questionnaires: distinction between minimally detectable change and minimally important change. Health Qual Life Outcomes. 2006;4:54. DOI 10.1186/1477-7525-4-54

Source note: SRC-1cbd1976f47e de Vet HCW 2006

48. Rodrigues SMA, Coelho IN, Costa PHV, de Carvalho Lana R, Polese JC. Validity of the modified sphygmomanometer test for the assessment of tip pinch strength in Parkinson's disease. J Bodyw Mov Ther. 2021;28:87–91. DOI 10.1016/j.jbmt.2021.06.006

Source note: SRC-772c77dd5449 Rodrigues SMA 2021

49. Alaei P, Wile DJ, Jakobi JM. Effects of dopaminergic medication on upper limb motor function, strength and hand dexterity in people with Parkinson’s disease. Clinical Parkinsonism & Related Disorders. 2025;13:100400. DOI 10.1016/j.prdoa.2025.100400

Source note: SRC-a9df82a4b70e Alaei P 2025

50. Corfitsen AR, Nygaard MKE, Eskildsen SF, Dalgas U, Langeskov-Christensen M. Is physical fitness associated with brain structure and function in Parkinson's disease? Brain Imaging Behav. 2026;20(2):44. DOI 10.1007/s11682-026-01098-x

Source note: SRC-1b9cb84e0aa5 Corfitsen AR 2026

51. Martignon C, Ruzzante F, Giuriato G, Laginestra FG, Pedrinolla A, Di Vico IA, et al. The key role of physical activity against the neuromuscular deterioration in patients with Parkinson's disease. Acta Physiol (Oxf). 2021;231(4):e13630. DOI 10.1111/apha.13630

Source note: SRC-d7de1daef925 Martignon C 2021

52. Pongmala C, Stonsaovapak C, Luker A, Griggs A, van Emde Boas M, Haus JM, et al. Association of Specific Leg Muscle Strength and Motor Features in Parkinson's Disease. Parkinsons Dis. 2024;2024:5580870. DOI 10.1155/2024/5580870

Source note: SRC-11f67adb4a09 Pongmala C 2024

53. Moreno Catalá M, Woitalla D, Arampatzis A. Central factors explain muscle weakness in young fallers with Parkinson's disease. Neurorehabil Neural Repair. 2013;27(8):753–9. DOI 10.1177/1545968313491011

Source note: SRC-f633d457f4d6 Moreno Catala M 2013

54. Schilling BK, Karlage RE, LeDoux MS, Pfeiffer RF, Weiss LW, Falvo MJ. Impaired leg extensor strength in individuals with Parkinson disease and relatedness to functional mobility. Parkinsonism & Related Disorders. 2009;15(10):776-780. DOI 10.1016/j.parkreldis.2009.06.002

Source note: SRC-01beeeb50dda Schilling BK 2009

55. Nocera JR, Buckley T, Waddell D, Okun MS, Hass CJ. Knee extensor strength, dynamic stability, and functional ambulation: are they related in Parkinson’s disease? Archives of Physical Medicine and Rehabilitation. 2010;91:589–595. DOI 10.1016/j.apmr.2009.11.026

Source note: SRC-4c162faeb1d7 Nocera JR 2010

56. Caetano MJD, Lord SR, Allen NE, Song J, Paul SS, Canning CG, et al. Executive Functioning, Muscle Power and Reactive Balance Are Major Contributors to Gait Adaptability in People With Parkinson's Disease. Front Aging Neurosci. 2019;11:154. DOI 10.3389/fnagi.2019.00154

Source note: SRC-183ebcd206bd Caetano MJD 2019

57. Jones GR, Roland KP, Neubauer NA, Jakobi JM. Handgrip Strength Related to Long-Term Electromyography: Application for Assessing Functional Decline in Parkinson Disease. Arch Phys Med Rehabil. 2017;98(2):347–352. DOI 10.1016/j.apmr.2016.09.133

Source note: SRC-38e2132ed4f4 Jones GR 2017

58. Lima DP, de Almeida SB, Bonfadini JC, de Luna JRG, de Alencar MS, Pinheiro-Neto EB, et al. Clinical correlates of sarcopenia and falls in Parkinson's disease. PLoS One. 2020;15(3):e0227238. DOI 10.1371/journal.pone.0227238

Source note: SRC-693954445453 Lima DP 2020

59. Lima DP, Gomes VC, Luna JRG, Santos LTR, Almeida SB, Viana-Júnior AB, et al. Assessment of sarcopenia tools as predictors of falls in patients with mild to moderate Parkinson's Disease: A cohort study. Clinics (Sao Paulo). 2025;80:100776. DOI 10.1016/j.clinsp.2025.100776

Source note: SRC-1aabb83f632f Lima DP 2025

60. Ozer FF, Akın S, Gultekin M, Zararsız GE. Sarcopenia, dynapenia, and body composition in Parkinson's disease: are they good predictors of disability?: a case-control study. Neurol Sci. 2020;41(2):313–320. DOI 10.1007/s10072-019-04073-1

Source note: SRC-3b422be5c017 Ozer FF 2020

61. Almeida SB, Lima DP, Luna JRG, Viana Júnior AB, Roriz-Filho JS, Alencar ÁP, et al. Clinical correlates of physical performance and sarcopenia in Parkinson's disease: a cross-sectional study. Arq Neuropsiquiatr. 2026;84(2):1–11. DOI 10.1055/s-0046-1816034

Source note: SRC-631454774383 Almeida SB 2026

62. Skinner JW, Needle AR. Exploring the role of ankle muscle function in gait impairments and fall risk in Parkinson's disease. Hum Mov Sci. 2025;99:103316. DOI 10.1016/j.humov.2024.103316

Source note: SRC-5edbfaca544a Skinner JW 2025

63. Paul SS, Canning CG, Sherrington C, Lord SR, Close JC, Fung VS. Three simple clinical tests to accurately predict falls in people with Parkinson's disease. Mov Disord. 2013;28(5):655–62. DOI 10.1002/mds.25404

Source note: SRC-507aeb58eeff Paul SS 2013

64. Gallagher CL, Johnson SC, Bendlin BB, Chung MK, Holden JE, Oakes TR, et al. A longitudinal study of motor performance and striatal [18F]fluorodopa uptake in Parkinson’s disease. Brain Imaging Behav. 2011;5(3):203–211. DOI 10.1007/s11682-011-9124-5

Source note: SRC-55445a4e3ef4 Gallagher CL 2011

65. Gialanella B, Gaiani M, Comini L, Olivares A, Di Pietro D, Vanoglio F, et al. Walking function determinants in parkinson patients undergoing rehabilitation. NeuroRehabilitation. 2022;51(3):481–488. DOI 10.3233/nre-220103

Source note: SRC-cd3e12ddef9e Gialanella B 2022

66. Wang T, Geng J, Zeng X, Han R, Huh YE, Peng J. Exploring causal effects of sarcopenia on risk and progression of Parkinson disease by Mendelian randomization. npj Parkinson’s Disease. 2024;10:164. DOI 10.1038/s41531-024-00782-3

Source note: SRC-29f312ded023 Wang T 2024

67. Gray WK, Hildreth A, Bilclough JA, Wood BH, Baker K, Walker RW. Physical assessment as a predictor of mortality in people with Parkinson's disease: a study over 7 years. Mov Disord. 2009;24(13):1934–40. DOI 10.1002/mds.22610

Source note: SRC-613418d23812 Gray 2009

68. Combs-Miller SA, Moore ES. Predictors of outcomes in exercisers with Parkinson disease: A two-year longitudinal cohort study. NeuroRehabilitation. 2019;44(3):425–432. DOI 10.3233/nre-182641

Source note: SRC-2cae4979c483 Combs-Miller SA 2019

69. Schini M, Bhatia P, Shreef H, Johansson H, Harvey NC, Lorentzon M, et al. Increased fracture risk in Parkinson’s disease: An exploration of mechanisms and consequences for fracture prediction with FRAX. Bone. 2023;168:116651. DOI 10.1016/j.bone.2022.116651

Source note: SRC-069e9a832384 Schini M 2023

70. Qaisar R, Iqbal MS, Karim A, Ahmad F. Plasma CAF22 Predicts Incident Freezing of Gait in Parkinson's Disease: A Prospective Cohort Study. J Mol Neurosci. 2026;76(3):126. DOI 10.1007/s12031-026-02584-z

Source note: SRC-9935c024926a Qaisar R 2026

71. Liu M, He P, Ye Z, Zhang Y, Zhou C, Yang S, et al. Association of handgrip strength and walking pace with incident Parkinson’s disease. J Cachexia Sarcopenia Muscle. 2024;15(1):198–207. DOI 10.1002/jcsm.13366

Source note: SRC-dfde0f8f8f2b Liu M 2024

72. Wu KM, Kuo K, Deng YT, Yang L, Zhang YR, Chen SD, et al. Association of grip strength and walking pace with the risk of incident Parkinson's disease: a prospective cohort study of 422,531 participants. J Neurol. 2024;271(5):2529–2538. DOI 10.1007/s00415-024-12194-7

Source note: SRC-9a7e77eeaad8 Wu KM 2024

73. Mey R, Calatayud J, Casaña J, Núñez-Cortés R, Suso-Martí L, Andersen LL, et al. Is Handgrip Strength Associated With Parkinson’s Disease? Longitudinal Study of 71 702 Older Adults. Neurorehabil Neural Repair. 2023;37(10):727–733. DOI 10.1177/15459683231207359

Source note: SRC-75dcf996d55a Mey R 2023

74. Berger K, Karch A, Lill CM, Berger TL, Mueller U, Bartl M et al. A comprehensive assessment of lifecourse and mortality of Parkinson’s disease in the German National Cohort. npj Parkinson's Disease. 2026;12(1):156. DOI 10.1038/s41531-026-01446-0

Source note: SRC-3fcb84c73d99 Berger K 2026