In this report
Audited and Updated
Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.
TR-01
Replace “the sample contained eight participants per arm” with “eight participants were randomized per arm; after one BPT participant was reassigned to control following a lesion, the analyzed groups contained seven, eight and nine participants.” Add this analytical limitation to the appraisal.
Type: report qualification needed. Audit disposition: supported report qualification.
Remaining limit: Distinguish randomization from analyzed allocation; do not silently resolve separate sex-count inconsistencies.
TR-02
Add reference [35] and consider “Algorithms intended for everyday monitoring can perform less well with dyskinesia, slow or incomplete transfers” to avoid implying all cited results are unsupervised real-world PD validation.
Type: citation and setting qualification. Audit disposition: supported citation and setting precision.
Remaining limit: Add the directly relevant citation and retain intended everyday-use versus actual laboratory/home-like validation distinction.
TR-03
Delete “prespecified” or provide a dated protocol/statistical analysis plan substantiating advance specification. Do not label the analysis post hoc without evidence.
Type: unverified temporal prespecification. Audit disposition: supported provenance qualification.
Remaining limit: No dated analysis plan was established in this QA. Remove unsupported prespecified adjective or obtain dated plan; current registry absence does not prove post hoc analysis.
TR-04
If reusing the table as an outcome-specific dataset, add “dependency analysis n=176/162; mortality cohorts n=198/192.”
Type: optional denominator clarification. Audit disposition: supported optional denominator clarification.
Remaining limit: Optional precision, not a demonstrated report numerical error; axial composite HR does not isolate chair rise.
Editorial record
- Audit status: supported citation and setting precision. Add the directly relevant citation and retain intended everyday-use versus actual laboratory/home-like validation distinction.
- Edited phrase under TR-02 . Original wording: P0008
- Audit status: supported report qualification. Distinguish randomization from analyzed allocation; do not silently resolve separate sex-count inconsistencies.
- Audit status: supported provenance qualification. No dated analysis plan was established in this QA. Remove unsupported prespecified adjective or obtain dated plan; current registry absence does not prove post hoc analysis.
- Edited phrase under TR-03 . Original wording: P0178
- Audit status: supported report qualification. Distinguish randomization from analyzed allocation; do not silently resolve separate sex-count inconsistencies.
- Edited phrase under TR-01 . Original wording: P0246
- Audit status: supported optional denominator clarification. Optional precision, not a demonstrated report numerical error; axial composite HR does not isolate chair rise.
- Audit status: supported optional denominator clarification. Optional precision, not a demonstrated report numerical error; axial composite HR does not isolate chair rise.
Editorial nomenclature update — 4 October 2026 at 11:48:46 am (Australia/Brisbane): authored condition labels and report wording use Parkinson’s disease. Published article titles, exact quotations, recorded searches, identifiers and routes are preserved. This is a terminology edit, not a scientific correction.
Executive assessment
Sit-to-stand assessment in Parkinson’s disease is most useful when the transfer task and intended interpretation are explicit. A chair-rise result combines movement initiation, strategy, postural control and physical capacity. Completion time alone cannot identify the limiting impairment, and a successful fast rise can coexist with weakness, altered mechanics or unsafe settling.
Choose the task deliberately. Observe a single rise and return to sitting before selecting a repeated test. Five times sit-to-stand measures completion of a fixed sequence; 30-second and one-minute tests count sustained repeated performance; TUG-embedded rising includes the transition into walking. Chair height, arms, assistance, instructions and timing endpoints materially affect interpretation. Failure and assisted completion belong in the result rather than being discarded or converted into an apparently measured time. [1, 2, 3]
Change is protocol specific. Published PD reliability findings vary substantially. Paul's faster-of-two, 45-cm protocol had SEM 0.6 s, implying an approximately 1.7-s MDC95 by calculation; Petersen reported a 10-s MDC95 for 5×STS and three repetitions for 30-second STS. Newer trial-number data also show higher ICC estimates for averaged scores than for one trial. Statistical detectability, patient-important improvement and a change from assisted to independent transfer are separate interpretations. [4, 5, 6, 7]
Instrumentation can explain performance. Force platforms, motion analysis, IMUs and video can identify phase delays, loading and movement strategies that a stopwatch misses. Their outputs require different levels of inference: external force is measured, mechanical power is derived, and camera or equation-based power is model estimated. Good event detection or clinical-score correlation does not validate force, power or small individual change. Algorithms intended for everyday monitoring can perform less well with dyskinesia, slow or incomplete transfers; the cited PD settings were laboratory or home-like, rather than fully unattended community validation. [2, 8, 9, 10, 35]
Prognosis requires a separate judgement. Prospective chair-rise evidence includes positive associations, attenuation after clinical adjustment and null findings. Historical faller classification is not future-fall prediction. Functional transfer ability and timed performance can also have different relationships with later outcomes. In Paul and Mak, adjusted STS associations were not significant; Moraca found chance-level 30-second STS fall discrimination. A chair-rise ability rating had an unadjusted mortality association while timed rise did not. [11, 12, 13, 14]
For rehabtools, the immediate priority is a reproducible task record with visible assistance and failure, a small set of validated movement descriptors, and qualified change interpretation. Individual risk estimates need prospective validation of the complete measurement and prediction workflow.
Scope and appraisal
This report examines sit-to-stand and functional transfers in established Parkinson’s disease for present assessment, measurement of change and prediction of later outcomes. It covers a single chair rise and descent, repeated chair-rise tests, transfers embedded in Timed Up and Go, and instrumented or video approaches. The intended decisions are which task to use, what its output actually measures, and which rehabilitation or rehabtools claims the evidence supports.
A chair-rise result combines physical capacity, movement initiation and strategy, postural control, and the testing conditions. Completion, speed, use of the arms, assistance and stability after rising are complementary observations. Force, mechanically derived power, model-estimated power and task timing are therefore considered separately. A total mobility score is not an isolated chair-rise measurement; sit-to-walk adds the transition into locomotion.
Cross-sectional validity, reproducible change and prognosis answer different questions. Associations with current severity or past falls do not establish future risk. Evidence for a composite task or multivariable model is attributed to the whole task or model unless the isolated transfer contribution was tested. The report retains relevant null findings and separates independent replication from further analyses of the same cohort.
Search and source appraisal
Original Methods, Results, tables and available supplements were preferred. Study tables distinguish complete journal text, accepted manuscripts, partial primary material and abstracts; unavailable details are not inferred from a different study with the same test name. Access describes the material examined, not study quality. Search coverage and supplement retrieval were not exhaustive, so an unlocated study or outcome is not evidence of absence.
Appraisal focused on selection, disease stage, task and medication protocol, assistance, non-completion, trial aggregation, measurement error, outcome ascertainment, uncertainty, adjustment and validation. Estimates from different protocols or cohorts cannot establish a test ranking by themselves. The synthesis and source checks were AI assisted, without duplicated independent human screening or a formal COSMIN, prediction-model risk-of-bias or GRADE rating. Study findings are distinguished from calculations and implementation recommendations.
Clinical test and protocol selection
Begin with the transfer question
For an ambulatory person with Parkinson’s disease (PD), five-repetition sit-to-stand (5×STS) is a useful short assessment of repeated transfer performance. It is not a direct measure of muscle strength, leg power, or fall probability. A defensible assessment begins with one observed rise and return to sitting, documents whether the person can perform it safely without arm support, and then adds a timed or repetition-counted test appropriate to the clinical question. An individual who needs hands or assistance should still receive a functional transfer assessment; an arm-free timed test alone can exclude precisely the people with greatest transfer disability. This is a reasoned clinical selection strategy, not a validated composite battery. [1, 2]
The distinction between capacity and impairment matters. In Duncan's cross-sectional PD sample, 5×STS correlated more strongly with the Mini-BESTest (r=−0.71) and a pegboard measure used as a proxy for bradykinesia (r=0.55) than with body-weight-normalized quadriceps force (r=−0.33). Balance and pegboard performance accounted for approximately 53% of variance; adding quadriceps strength and other covariates increased the model's explained variance only modestly. This does not establish that strength is unimportant or that those correlations identify an individual's limiting mechanism. It shows why a slow chair test cannot itself diagnose weakness. If treatment decisions depend on force or power, measure those constructs separately. [1]
The 2021 review of clinical contributors is useful for tracing studies of bradykinesia, torque production, balance, posture and cognition, but it synthesizes 13 predominantly cross-sectional/case-control studies rather than validating a single clinical protocol. Its eligibility criteria also concentrate on H&Y stages I–III and people able to follow commands and rise with or without an armrest. It does not supply a universal chair-test protocol, individual change threshold or advanced-PD safety rule. [15]
Table 1 Selecting a transfer task
A reasoned selection guide rather than a validated battery. STS variants retain their own protocol and interpretation; assistance, motor non-completion and technical failure must remain visible.
| Task | Best question | What to standardize | Main interpretation limit |
|---|---|---|---|
| Single rise + descent | Can the person transfer safely, and how? | Chair/feet/arms; pace; start/end; assistance | Observer/video reliability is not day-to-day patient reliability |
| 5×STS | How quickly can five repeated transfers be completed? | Height; full stand/sit; fifth-standing vs fifth-sitting; trials/summary | Mixed balance, initiation, force and direction-switching; noncompletion matters |
| 30-second STS | How many valid transfers can be sustained for 30 s? | Arms; depth; valid repetition; partial final repetition; rest | PD retest evidence exists; 3-rep MDC is study-specific, not MIC |
| 1-minute STS | Sustained repeated-transfer performance | Same as 30 s, plus symptoms/rest and 60-s boundary | PD intervention evidence; no verified universal PD MDC/MIC or norms |
| TUG / sit-to-walk | Can rising flow safely into gait and return to sitting? | Distance/turning; aid; cue; segmentation; terminal sitting | Transition includes gait preparation; total TUG is not an isolated rise |
| Arm-assisted / higher-seat transfer | What support makes the real-life transfer possible? | Exact height, hand support and help required | Keep as explicit adapted outcome; do not compare with arm-free cutoffs |
Choosing among transfer tasks
Single sit-to-stand and the return to sitting. A single transfer is the most direct way to observe initiation, forward weight transfer, seat clearance, limb extension, stabilization and controlled descent. It can reveal failed attempts, hand placement, retropulsion, excessive forward progression, asymmetric loading or a rapid uncontrolled descent that total time misses. Its brevity also makes stopwatch reaction and endpoint decisions a relatively large component of the score.
Verheyden and colleagues tested 38 people with PD and 19 controls using one usual-way rise from a standard armchair, timed from the spoken start until standing vertical movement ceased. After one practice, a single recorded trial was independently timed by three raters, with pause and replay permitted; one rater retimed the same recording about a week later. The study therefore measures repeatability of observing a performance, not repeatability of the patient's movement on another day. It found no significant difference in single-rise timing between historical recurrent fallers and others (p=.34), which also cautions against transferring a repeated-test cutoff to this task. [16]
Stand-to-sit deserves a separately observed outcome even when the score is total 5×STS time. Record whether the person locates the seat, controls the descent, uses the hands or drops abruptly; a shorter total time does not necessarily mean better control of lowering. No independently validated PD-specific stopwatch stand-to-sit MDC or MIC was verified in this clinical evidence set. Instrumented descent measures need their own event definitions and reliability evidence rather than inheriting those of the ascent or total TUG. Chair-rise capacity also does not establish the ability to transfer into bed, a car or a low toilet; assess the relevant real-life task when that is the limitation.
A rise scored within MDS-UPDRS or Mini-BESTest answers a coarser question. These items grade the ability or strategy used to attain standing, including hand support or assistance; they do not measure watts, repeated-transfer endurance or eccentric control during descent. Use the actual scale's instructions and scoring rather than recoding an adapted test as a standard item. The Mini-BESTest's highest STS category requires unsupported rising and independent stabilization, while its lower categories distinguish first-attempt hand use from more severe difficulty. A high item score can coexist with intermittent failure over additional trials: in Iwai's study, the Mini-BESTest STS item did not significantly discriminate groups defined by worst performance across five separate rises. This is evidence of limited discrimination in that sample, not a formal estimate of the item's ceiling prevalence. [17, 18, 2]
5×STS. The principal advantage is a brief continuous time score with little equipment. It samples repeated rising, descending and switching between the two directions. That makes it broader than a single concentric rise and narrower than global mobility. Two PD protocols often cited together are materially different. Duncan used a 43-cm armless chair, arms crossed, initial back contact and a demonstrated full upright stand; timing ran from the spoken start to buttock contact after the fifth stand, and the main analysis used one trial. Paul used a 45-cm chair, arms folded, an instruction to perform as fast as possible, two trials and the faster score. Paul's recovered supplement describes timing five full stands but does not explicitly define its start event or whether the endpoint was fifth standing or subsequent sitting. A clinic cannot silently fill in those missing details and then claim exact replication of Paul's reliability estimate. [1, 4] Spagnuolo's responsiveness protocol differs again: a 43-cm chair with armrests and timing to fifth backrest contact rather than fifth seat contact. [7]
30-second chair stand. The count within a fixed duration asks how many repeated transfers a person can achieve, introducing a sustained-performance component. Unlike 5×STS, it does not require reaching five repetitions to provide a count; nevertheless, a person unable to rise arm-free still meets a functional floor if arms are prohibited. A fixed-duration test also requires explicit rules for full extension, seat contact and the partly completed repetition at the cutoff. Petersen's direct PD comparison favors 30-second STS for repeatability in its small H&Y I–III sample, with ICC(2,2)=0.94 and MDC95=3 repetitions, compared with ICC(2,2)=0.74 and MDC95=10 s for 5×STS. These are protocol- and sample-specific estimates, with complete administration details unavailable from the inspected abstract. They do not validate every 30-second chair protocol or make three repetitions a universal clinically important improvement. [5]
One-minute STS. One-minute STS has been used as an intervention outcome in PD. The September 2026 PARKEX trial randomized 24 participants, all H&Y II in the recruited sample, to two exercise programs or control. A one-minute STS count was a secondary functional outcome and improved in the exercise groups. However, this was not a test–retest, measurement-error or anchor-based interpretability study. The main paper describes proper-technique repetitions without sufficiently specifying chair height, hand use, familiarization, trial aggregation or cutoff counting to create a reproducible PD clinical standard. Its small, cognitively selected sample and baseline group imbalances also limit extrapolation. Use the one-minute version when sustained repeated transfer is genuinely the target and the protocol is documented; do not substitute pulmonary/healthy older-adult norms or MICs as PD-validated values. [3]
TUG-embedded rise and sit-to-walk. These are valuable when the question concerns getting up and moving away, turning and returning to a seat. They should not be treated as equivalent to rising and stabilizing in place. A person may carry forward momentum into the first step rather than arrest it in standing; sit-to-stand segmentation can also include part of gait initiation. Total TUG improvement can be driven by walking or turning without a corresponding improvement in the rise or controlled descent. Report the entire TUG protocol and each transition's event definitions when using instrumented subphases. The dominance of TUG in the 2026 review makes this distinction central to interpretation, rather than a minor nomenclature issue. [19]
Completion and assistance
Duncan enrolled 80 participants across H&Y I–IV, including seven who could not rise without upper-extremity use. Those seven were assigned 60 s, approximately one standard deviation beyond the slowest observed completed performance. Consequently, the reported overall mean of 20.25±14.12 s combines actual times with investigator-assigned failure scores. It is not a population norm for successful 5×STS, and a 60-s value in that study does not mean an observed one-minute completion. Excluding those participants would also have produced a more capable analysis sample. Both treatments of noncompletion change the meaning of the distribution. [1]
Iwai's 2025 study makes a useful alternative explicit. All 40 participants with PD could rise using armrests, but the study then asked for five separate comfortable-speed rises without arm assistance. Classification used the worst trial: 18 completed all five, 12 failed at least once after seat-off, and 10 failed at least once to clear the seat. Attempts were stopped if seat-off had not occurred within 10 s. The sample included H&Y IV and one H&Y V participant, unlike many studies restricted to independent performers. This does not validate a new clinical scale or establish longitudinal stages of deterioration, because it was cross-sectional. It does show that success, failure to clear the seat, failure after clearance, and inconsistency across trials are distinct observable outcomes. [2]
For routine reporting, retain at least: completed as specified; completed with hands; completed with physical assistance; incomplete because of motor difficulty; stopped for symptoms/safety; and unmeasurable because of technical error. Record valid repetitions before stopping, the level and location of assistance, attempts/rocking and whether the person ultimately achieved stable standing. A hand-assisted or higher-seat test may be clinically useful, but give it its own label and compare follow-ups under the same conditions. Do not merge adapted and unadapted times into one normative ranking. This is a practical reporting recommendation supported by the different failure patterns and protocols, rather than a validated categorical outcome set. [1, 2]
Setting up the transfer task
A fixed chair standardizes the object but does not standardize its mechanical demand relative to the participant. The same 43–45-cm seat is relatively higher for a shorter lower leg and relatively lower for a longer one. Conversely, adjusting the seat to knee height standardizes some starting geometry but changes the environmental challenge. Both approaches can be appropriate; they answer different questions. No PD-specific evidence retrieved establishes one universally optimal seat-height-to-leg-length ratio or a correction equation that makes their scores interchangeable.
The original PD laboratory protocols illustrate the issue. Inkster used a height-adjustable, armless/backless seat producing approximately 90° knee flexion, 20-cm foot separation and specified thigh support. Bhatt adjusted chair height to knee-to-floor distance, placed the feet to produce approximately 85° knee flexion, and kept hands on the iliac crests. Iwai matched height to lower-leg length, specified approximately two-thirds thigh support, 90° knees, shoulder-width bare feet with toes approximately 10° outward, and traced foot positions between trials. Martin's cueing experiment instead set knee flexion to 100° with vertical tibiae and then moved the feet 10 cm posteriorly. These are deliberate experimental choices, not interchangeable definitions of a standard chair rise. [20, 21, 22, 2, 23]
Document the loaded seat height and firmness, presence of backrest/armrests, footwear, initial seat depth/back contact, foot width and fore–aft position, and whether repositioning is allowed. For a longitudinal clinic measure, a photograph or simple floor marks can help reproduce the setup. If the relevant problem is transfer from the patient's low sofa or toilet, assess that environmental task separately and document its geometry; replacing it with a high laboratory chair could conceal the actual limitation. This is an implementation recommendation, not evidence that a particular floor-marking or photography method improves PD test reliability.
Movement instructions define the outcome. Comfortable-speed single rises, maximal-speed 5×STS and maximum-count endurance tests should not be pooled. Explicitly define full standing, whether the back must contact the backrest at every descent, whether forward trunk momentum or arm swing is allowed, the start signal and the final event. Fifth-standing and fifth-sitting endpoints differ by a descent; starting at movement onset rather than the command removes initiation latency. For stand-to-sit, report whether the endpoint is seat contact or stable seated posture and whether the descent begins from stationary standing or from turning toward the chair. Extra instructions added midway can change performance and should be recorded.
Cueing changes the test
A demonstration used for familiarization is different from synchronous modeling or continuous verbal cueing during the scored trial. The latter changes task initiation and attentional support. If cue responsiveness is the clinical question, compare a prespecified uncued condition with a clearly described cued condition, retaining both scores and the movement-quality observations. Do not introduce variable coaching into a longitudinal unassisted-capacity score.
Martin's crossover study involved 13 people with PD and 13 controls. PD assessments were confined to a window beginning 60 minutes after the usual dopaminergic medication. Three trials of each cue condition were performed. Modeling reduced coronal sway in the first 30 s after standing; two losses of balance occurred with reaching-to-target cues and one with internal-focus instructions, while no losses occurred in the uncued or modeling conditions. Two of the three episodes required examiner assistance. Event counts were small and did not differ statistically, so this does not establish that modeling is universally safest or that an entire cue category should be avoided. It does establish that cueing is not neutral and that speed or successful seat clearance alone is insufficient to describe a beneficial response. [23]
Bhatt's small training study provides complementary evidence: 21 of 31 randomized PD participants had usable force/motion data, with 13 trained and eight controls. Participants were assessed without cues before and after four weeks of audiovisual training. The study therefore concerns retained performance after training, not only the immediate benefit of a cue. Its stability findings used model-derived limits from healthy adults; they should not be recast as demonstrated fall prevention. [22]
Medication state and testing safety
Most core clinical reliability studies tested participants ON medication. Paul repeated testing one week later at the same point in the medication cycle, while Duncan specified ON without an equally detailed interval from dosing. Inkster's repeated ON/OFF laboratory assessments show why medication state belongs in the record: in a small selected mild-PD sample, mean single-rise duration was 1.97 s OFF, 1.86 s ON and 1.89 s in controls, with changes in strategy despite similar ON duration. A normal-looking time therefore need not imply normal mechanics. These experimental OFF protocols are not a recommendation to ask patients to withhold medication for a routine or home assessment. [4, 1, 21]
Record medication state, time since last dose and the person's usual fluctuation pattern, alongside dyskinesia, freezing, dizziness, pain, recent exertion and any rest periods. Repeat assessments at a comparable medication time when estimating rehabilitation-related change. If a within-day medication question is being studied, specify it prospectively and separate it from longer-term progression. Longer repetition tests also increase exposure to fatigue, cardiopulmonary symptoms and repeated orthostatic transitions; more repetitions do not automatically provide a better or safer measure.
Use a stable chair on a nonslip surface, sufficient space, and a clinician or appropriately trained helper able to prevent a fall when needed. Observe at least one safe rise and return before a rapid repeated test; stop for unsafe instability, presyncopal symptoms, pain or the need for unexpected physical assistance. This is a clinical safety framework, not a validated PD-specific stop-rule algorithm. The task is not an orthostatic-hypotension diagnostic test, and absence of dizziness does not validate cardiovascular safety for every patient.
Low cost and remote implementation
The minimum clinic equipment is a measured stable chair, stopwatch and clear scoring form. Ordinary video can add a reviewable record of initial position, hand use, full extension, seat contact, rocking, balance-recovery steps and controlled descent. A fixed view that includes the person, feet and chair is usually more useful than a close-up that omits the base of support. Keep the timing events and playback approach fixed. Video review can improve the ability to inspect an event but should not be presented as a validated force/power system, or assumed interchangeable with live stopwatch time.
Afshari's prospective feasibility study tested a concrete remote workflow: 15 H&Y II–III participants, a mandatory care partner, initial in-person instruction, gait belt, stable medication and ON-state televisits. Participants used the same 43–45-cm dining chair at home, folded their arms, fully straightened their legs and sat fully, without a requirement to touch the backrest at every repetition. Two timed trials were performed at each of four televisits. All 120 FTSTS assessments were ratable and no falls, near-falls or other adverse events occurred; 26/120 assessments were classified high risk by the study's borrowed time/arm-use criterion. This supports feasibility under that supervised, selected arrangement. It does not establish that unsupported self-testing is safe, that the home and clinic times agree within a clinically acceptable margin, or that the cutoff predicts future falls in a virtual setting. [24]
Silva-Batista and colleagues compared in-person and live videoconference home 5×STS in 62 people with H&Y I–III PD. In-person testing always came first, followed by a different remote rater 1–4 days later; both visits occurred at approximately the same time of day in the reported ON state. Participants used the same 43–45-cm armless chair against a wall or door, with initial back contact, arms crossed and feet flat. The study found r=0.95 and ICC=0.97 (95% CI 0.95–0.98). Nevertheless, the published Bland–Altman limits were −3.3 to 4.4 s, showing that excellent relative agreement can coexist with individual differences of several seconds. The prose gives a mean difference of +0.5 s, whereas Table 2 reports in-person minus remote −0.57 s from means of 11.30 and 11.87 s. Without the figure image, the sign convention cannot be resolved; the magnitude and reported limits are retained rather than silently corrected. Fixed order confounds modality, rater, day and possible practice effects. [25]
All participants completed remote assessment without a care partner and no adverse events occurred, but this was a prepared and selected workflow: technology screening, troubleshooting and an in-person home safety/setup visit preceded remote testing. Eighteen of 62 encountered technology problems, which were resolved with assistance. Eligibility excluded people unable to stand and walk without an aid, those with more than three falls per week, dementia or MoCA scores ≤19, and relevant medical contraindications. These findings support supervised videoconference testing in comparable patients; they do not establish unsupported self-testing, general safety in advanced PD, prospective fall prediction, or an individual change threshold. [25]
Instrumented and video assessment
What instrumentation adds
Instrumented sit-to-stand (STS) assessment is most defensible when it answers a specific question left unresolved by completion time: Was the delay before seat-off, during ascent, or during settling? Did the person require repeated attempts? Was the movement trunk-led, asymmetric, or associated with premature braking? Did a change in medication state or a cue change the movement strategy? These are different questions from whether a device diagnoses Parkinson’s disease (PD), reproduces a clinical severity score, or forecasts a future fall.
The evidence supports technology-specific measurement of timing and selected movement features, and exploratory characterization of PD-related transfer strategies. It does not support treating all camera, smartphone, IMU, or force-platform outputs as interchangeable measurements of leg strength, power, balance, or future risk. For rehabtools, a transparent set of timing and movement-quality descriptors is currently easier to defend than a composite “strength,” “power,” or “fall-risk” score inferred from ordinary video.
Trucco and colleagues’ 2026 scoping review provides a useful map, not a certification of measurement performance. It included 77 reports from 2015 to May 2025: 67 used IMUs, 13 optical systems and seven force/pressure sensors; categories overlapped. TUG was represented in 46 reports, whereas 17 used an STS test. Most work was outpatient-based, and no study exclusively recruited advanced-stage PD. The review did not establish a common reference standard, a pooled measurement error, or clinical thresholds for the many reported variables. Older laboratory biomechanics remain essential because its search began in 2015. [19]
What the instrument measures
Four levels must remain distinct.
Direct external-force measurement: a force platform measures ground-reaction force and moments at the contact surface. A seat platform measures chair-contact force. A two-foot setup permits separate left/right loading; a single platform under both feet does not provide limb-specific force merely because a camera shows both legs.
Mechanically derived quantities: centre of mass (CoM), joint moments and joint power require a biomechanical model. CoM velocity may be obtained from kinematics or integration of force-derived acceleration, with assumptions about body mass, initial conditions and external contacts. Joint moments usually require inverse dynamics with synchronized kinematics, force and segment inertial properties; joint power is joint moment multiplied by angular velocity. These are not direct recordings of individual muscle force.
Model-estimated power: an IMU or encoder can supply acceleration, velocity or displacement, but translating these into “power” requires an explicit mechanical or statistical model. Likewise, an equation using body mass, body height, chair height and repetition count estimates average STS power under specified assumptions. Agreement must be established for that model and population.
Performance and movement descriptors: duration, repetition count, head ascent speed, trunk angular velocity, range of motion, pauses, attempts and sway characterize task execution. They may relate to force-generating capacity, but are not themselves strength, power or rate of force development (RFD).
This distinction is particularly important before seat-off: body weight and acceleration are supported by both the seat and feet, and sometimes the arms. A foot-only force trace omits chair and arm reactions. A camera cannot determine force sharing merely from the trajectory of one landmark. Differentiating a noisy position signal to obtain acceleration, and then differentiating again to infer an RFD-like quantity, amplifies sensitivity to tracking error, frame rate, filtering and event definitions.
“RFD” is also not a universal quantity. Iwai et al. calculated a first-peak foot-loading slope from the difference between a local force minimum and the first local maximum divided by time to peak. Malling et al. calculated the mean rising-force slope within 30–70% of a selected force peak. The results cannot be exchanged without preserving those definitions. Neither is automatically the same as maximal voluntary muscle RFD from a standardized dynamometer test. [2, 26]
Table 2 Technology and protocol choices for STS assessment
This is a decision framework, not a claim that every listed configuration is clinically validated. Keep device, placement, software version, filtering, event definition and test instructions stable for monitoring.
| Technology or protocol | What is obtained | Potential added value | Required qualification |
|---|---|---|---|
| Stopwatch or reviewed video; standard repeated test | Repetition count, completion time, noncompletion, visible assistance | Low burden; supports conventional functional performance assessment | Whole-test time does not isolate ascent, descent, pauses or force generation |
| Ordinary RGB camera; fixed view | Image landmark trajectories and event times; calibrated segment speed only after validation | Attempts, phase timing and reviewable trunk strategy | Not direct force, strength or power; perspective, occlusion and filtering affect estimates |
| Depth camera or multi-camera markerless system | Three-dimensional segment positions and derived kinematics | More detailed coordination and segment-velocity assessment | System-specific validation; eight-camera evidence cannot validate a single webcam |
| Single phone or wearable IMU; fixed body site | Acceleration, angular velocity and algorithm-derived timing/orientation | Repeated cycle profiling or lower-burden activity monitoring | Body site, model, dyskinesia and transition definition matter; event accuracy is not duration agreement |
| Separate foot and seat force plates | External contact force; derived loading rate and contact timing | Weight transfer, braking, failed seat-off and bilateral foot loading | Need both foot signals for force asymmetry; pre-seat-off interpretation must include chair/arm contacts |
| Force plates plus motion capture | External forces plus modeled CoM, joint moments and joint power | Mechanistic assessment of task strategy and mechanical demands | Power and joint moments are derived, with assumptions and synchronization requirements |
| Count/anthropometry equation or IMU/encoder power model | Estimated average power in defined units | Potentially scalable functional-power descriptor | Requires independent mechanical criterion agreement, PD reliability and evidence of added value over time/count |
| Naturalistic sit-to-walk or walk-to-sit monitoring | Context-dependent transitions, counts and movement distributions | Repeated samples across ordinary activity | Different endpoint from isolated STS; furniture, carrying, opportunity and detection yield can drive apparent change |
Biomechanical evidence beyond overall time
Force generating capacity and transfer strategy are related but separable
Inkster et al. assessed ten men with mild PD and ten age-matched male controls, combining independently measured isokinetic hip/knee extensor torque with self-paced STS. In PD, hip strength correlated with STS duration in the ON and OFF states (r = −0.71 and −0.80), whereas knee strength did not. Controls showed the opposite pattern of a significant relationship with knee torque. Importantly, the PD group’s mean ON-state STS duration was essentially the same as controls (1.86 versus 1.89 seconds), despite lower measured extensor strength. This is direct evidence against interpreting a normal rise time as proof of normal strength, or treating a time–strength association as an interchangeable measurement. The sample was small, all male and selected for independent OFF-state rising. [20]
In a closely related ten-PD/ten-control experiment, synchronized motion capture, bilateral foot force plates and an instrumented seat were used for inverse-dynamics analysis. ON-state duration again resembled control performance, but movement organization differed: PD participants used more preparatory hip flexion and more anterior CoM displacement before lift-off, with less subsequent forward displacement. Knee extensor moment peaked earlier relative to seat-off, and was lower in PD, while hip extensor moment was similar. These papers have closely matching samples and should not be counted as independent replications. The implication is clinically useful: an apparently adequate time can conceal a compensatory solution. [21]
Mak et al.’s older work also separates movement speed from coordination. A 2002 study compared 20 PD participants with 20 controls, then asked 15 controls to imitate the patients’ slower speed. Many kinetic and kinematic differences disappeared at matched speed, but a longer transition from peak forward velocity to seat-off persisted in PD. A separate seven-PD/six-control inverse-dynamics study found reduced hip-flexion torque and slower torque development. These small studies make direction switching and time to generate force plausible contributors to slow rising; they do not justify assigning a causal impairment from a single camera feature. Original abstracts were verified for these two papers, but complete original Methods/Results were not obtained for this appraisal. [27, 28]
A larger 2023 study by Baizabal-Carvallo et al. included 90 PD participants and 52 controls; 20 patients had an MDS-UPDRS chair-rise score of at least 2. Lower-limb bradykinesia and reduced hip-related force were independently associated with this contemporaneous abnormality. The study used a digital force gauge and a calibrated weighing-scale leg-pressure task, rather than measured force multiplied by measured velocity. Its study-specific ≤9-kg threshold should not be treated as an externally validated prognostic cutoff or a measure of watts. It reinforces a multifactorial explanation of difficulty without converting STS into a pure strength test. [29]
Failure deserves its own outcome rather than deletion from analysis
Iwai et al. studied 40 people with PD and 23 controls using three force plates: one under the seat and one under each foot. Five separate comfortable-speed, arms-crossed attempts were performed, with seat height normalized to lower-leg length. Participants with PD could stand using armrests at inclusion; testing intentionally examined performance without that assistance. Worst performance across the five attempts classified 18 as consistently successful, 12 as failing to rise after seat-off, and ten as failing to achieve seat-off. These are contemporaneous task-defined groups, not observed longitudinal stages. [2]
Repetitive movements increased and first-peak foot-loading slope decreased across the groups. The slope was approximately 101% body weight/second in controls, 92 in successful PD, 72 in failure after seat-off and 29 in failure to reach seat-off. Reduced weight transfer and earlier braking were associated with unsuccessful performance. Lower-limb bradykinesia distinguished success from failure after seat-off, whereas balance measures were particularly informative for separating failure after seat-off from failure to achieve seat-off. The age-adjusted ROC analyses classify these same task-defined groups; they do not predict subsequent disability, disease progression or falls. [2]
The handling of missing events is instructive. The force-defined onset could not be detected in 30% of the seat-off-failure group, and successful-rise timing variables could not exist for participants who never left the chair. This is clinically meaningful noncompletion, not ordinary random missingness. An automated system that discards these attempts may preferentially remove its most impaired users and make retained performance look better than actual transfer ability.
Mak et al.’s 2011 laboratory study similarly distinguished successful from failed attempts. Its much-quoted 87.5% classification was of the immediate STS outcome using movement-state variables, not future falls. A backward return to the chair during the experiment must not be relabelled as a prospectively observed community fall. [30]
Asymmetry and instability need the correct sensor and definition
Ramsey et al.’s 13-PD/11-control force-platform, EMG and motion-analysis study found no significant overall between-group differences in the reported outcomes, but did find bilateral differences within PD in knee angle at seat-off, peak vertical force and peak torque. This supports recording sides where clinically relevant, while illustrating that within-PD asymmetry is not the same as a validated diagnostic or prognostic threshold. The original abstract was accessible; detailed original tables and multiplicity handling were not verified. [31]
A visible lateral lean is not a measure of left–right force asymmetry. Similarly, trunk acceleration variability, landmark sway area, force-platform centre-of-pressure displacement and CoM stability margin are different constructs. A person can deliberately slow or lean further to preserve stability. Conversely, faster ascent may terminate with overshoot or a corrective step. Report actual observations such as an extra step, hand support or return to sitting alongside numerical features; do not infer “instability” solely from a noisy pose trace.
Camera and markerless assessment
Criterion agreement has to be demonstrated for the particular output
Galna et al. concurrently measured nine people with PD and ten controls with Kinect and a Vicon system. This is genuine technical criterion-comparison evidence for a specific depth-camera system, rather than a correlation with a clinical score. Timing was generally accurate; gross spatial STS movement performed much better than small hand movements. The study used frontal recording at a controlled three-meter distance, and PD participants were assessed at peak medication effect. It does not validate every joint angle, a phone RGB model, a different camera placement, or camera-derived force/power. The accessible original body describes bias, ICC and limits-of-agreement analyses; not all tabulated numeric outcomes were preserved in the machine-readable body. [32]
The practical lesson is to name the quantity and reference standard: for example, “agreement of total STS duration with synchronized motion capture under this camera setup.” A strong correlation alone can coexist with consistent bias or clinically wide limits of agreement.
Morgan 2023 naturalistic video study
Morgan et al. analyzed 85 hours of video from 12 PD participants and 12 healthy spouses/friends/family members staying in pairs in an instrumented test house for five days. Participants largely lived freely, although the setting was a research house and cameras recorded selected periods. Human raters identified 813 STS episodes; this is not a verified denominator of successfully quantified episodes for every analysis. The reported group comparisons used 200 control, 75 PD-ON and 43 PD-OFF observations, whereas participant-level tables contain other counts, so the end-to-end usable-event yield is not fully reconciled. OpenPose was then applied to clips centered on manually identified events, using head/shoulder motion to quantify final-attempt duration and peak ascent speed. Duration was not the time from the first preparatory attempt; earlier unsuccessful attempts were deliberately omitted. The paper therefore validates automated quantification after event localization, not a fully unattended end-to-end detection pipeline in this cohort. [8]
The ascent signal was derived from image coordinates, normalized using skeleton height and participant stature to express speed in meters/second. It is a head/upper-body motion estimate, not whole-body CoM speed or joint power. The original Results report participant-averaged manual-versus-automatic duration correlation r = 0.419, duration bias 0.10 seconds and random error 0.19 seconds; automatic speed correlated inversely with manually annotated duration (r = −0.780). The latter compares two related but different variables, so it is not criterion agreement for speed. [8]
Clinical associations were encouraging but contextual. For non-carrying transitions, age/height-adjusted speed correlated with MDS-UPDRS III (ρ = −0.691); the association was not significant when participants were carrying something. The all-event duration–MDS-UPDRS III association became nonsignificant after age/height adjustment (ρ = 0.365, p = 0.094). Clinical correlations combined PD and control observations, which can accentuate known-group separation. Group-level ON/OFF differences were present, but only two PD participants had significantly slower automated ascent speed OFF medication; none had a significant individual duration difference. This is feasibility, concurrent construct validity and preliminary state sensitivity. It is not prospective prediction of disease progression, a validated individual monitoring threshold, or proof that a clinically meaningful change can be detected reliably. [8]
Recent markerless research adds strategy detail with limited transportability
Kantha et al. (published July 2026) studied 15 ON-medication PD participants and 15 matched controls using eight synchronized cameras at 60 Hz. This was a calibrated multi-view research system, not a single ordinary webcam. Participants completed two trials each of self-paced single STS and fast five-repetition STS from a 45-cm chair. The PD group showed earlier head/trunk initiation and slower head, hip and knee extension velocities during single STS. During repeated STS, both trunk adjustment durations were longer; apparent increases in adjustment amplitude and peak velocity did not meet the multiplicity-adjusted significance threshold. [33]
No force plates were used to establish extensor force or power. The authors’ mechanical explanation of reduced extensor force is an inference from kinematics. Sway over the final three seconds did not differ significantly between groups. Left and right kinematics were averaged, so this analysis does not validate an asymmetry metric. The paper adds a plausible description of coordination and provides precise event definitions, but no prospective prediction, treatment-response threshold or within-PD criterion-agreement validation of the full pipeline. [33]
New video scoring systems must also be interpreted at their actual level. Liu et al. used 1,731 videos from 112 PD patients for multi-item MDS-UPDRS scoring, with patient-level train/test separation within each item. The arising-from-chair/posture/global-bradykinesia recording group contained 171 videos, including 30 test videos. Strong aggregate postural-symptom accuracy is not an STS-specific accuracy estimate, and clinician-label agreement does not validate force, power or future outcomes. The work is promising for assisted contemporaneous clinical scoring; independent external and home-use validation remain separate requirements. [34]
IMU and phone measurement evidence
Pham et al. developed a lower-back IMU algorithm in five participants and evaluated it in a separate test set of 21 PD participants and 11 older adults during 90–180 minutes of daily-like activities. Video-defined detection performance deteriorated substantially with dyskinesia: accuracy was 82% in PD without dyskinesia and 47% with dyskinesia. The quoted detection and duration figures combine sit-to-stand and stand-to-sit transitions. Direction classification was 98% overall, but the reported transition-duration difference interval was wide (mean difference 0.20 seconds; −1.06 to 1.45 seconds around the reported 95% interval). The reference raters timed transitions in whole seconds. This is useful external-to-training participant validation for detection in a home-like protocol, with clear phenotype and timing-reference limitations. [35]
Adamowicz et al. provide a particularly important distinction between event detection and measurement agreement. In the held-out PD validation set, 314 complete transitions from 20 participants yielded sensitivity 0.853 and precision 0.988. Nonetheless, duration limits of agreement were approximately −1.71 to 1.30 seconds, with a mean bias of −0.20 seconds. Slow and impaired transitions were more problematic, and 105 partial transfers were excluded before validation. The often-cited finding that three monitoring days were sufficient for reliable aggregate features came from healthy participants’ home monitoring, not from PD home data. It should not become a PD monitoring prescription. ON/OFF classification by the examined features was poor despite good PD-versus-control separation. [9]
A two-IMU model from Wairagkar et al. estimated three-segment kinematics and classified sitting, standing and transitions in ten younger adults, 12 older adults and 12 people with PD. State-classification accuracy was 91.41% in PD. Reference optical kinematic validation was conducted in the younger healthy group; PD application included many home assessments. It is inappropriate to transfer the healthy optical-validation result wholesale to PD joint-angle accuracy, particularly when thigh motion was inferred rather than directly instrumented. [36]
Sher et al.’s recent single-smartphone study deserves attention because it reports more than correlation: among 35 participants, participant-mean cycle time showed ICC(2,1) = 0.998, bias −0.012 seconds and limits of agreement −0.134 to 0.110 seconds against video. However, only five participants had PD. The 660 cycles are nested observations, not 660 independent patients; PD cycle-time mean absolute error was 60.5 ms compared with smaller errors in other groups. It assessed a supervised 30-second chair-stand protocol, with no independent external PD cohort, and video frame rate is reported inconsistently in the Methods. It supports feasibility of cycle timing and strategy classification; diagnostic and personalized rehabilitation claims remain preliminary. [37]
Less favorable contemporary studies are equally relevant. Amici et al.’s 2026 sternum-IMU pilot reported 93% accuracy across all postural-transition signals in five people with PD, yet STS-specific correct classification was only 18/26 (69%); eight rises were classified as sitting down. Overall activity accuracy can conceal clinically important STS errors. Goubault et al.’s ankle-worn smartwatch study in 20 mildly affected ON-state PD participants obtained sitting-phase F-scores of 75.6–84.7 across simulated daily-activity trials. Its segment-level cross-validation is not clearly a participant-held-out external evaluation, and it validates sitting-phase segmentation rather than precise STS force or power. [38, 39]
Three-day monitoring by Bernad-Elazari et al. in 99 PD participants and 38 older controls reported strong current-group classification using everyday sit-to-walk and walk-to-sit features. These observations extend the ecological literature beyond small laboratory cohorts, but the source available for this appraisal was the original abstract: details of participant-wise validation and feature selection could not be fully audited. The classification should not be presented as future progression prediction. A 2026 continuous-video segmentation paper by Zhang et al. reports segment F1@50 of 92.59% and mean intersection-over-union of 81.78%; its full Methods were unavailable, so participant-level generalization remains unverified. This addresses the important localization problem but does not validate downstream biomechanical quantities. [40, 41]
Instrumented change and power estimates
More features can mean more variable features
Khalil et al. examined repeated instrumented mobility trials in 262 people with PD and 50 controls. TUG and cognitive TUG total times were generally reliable, but chair-transition components and other sensor features were often less consistent. Across selected quantitative features, median ICCs for TUG/cognitive TUG were only 0.36–0.61, despite good-to-excellent total-time reliability. Higher ICC in advanced PD could coexist with substantial trial changes because between-person variability was large. Their back-to-back testing was close to a best-case repeatability scenario; it does not establish day-to-day reliability. A feature can improve group classification while remaining too variable to interpret a small individual change. [42]
Marin et al.’s single-thorax-IMU scores illustrate the same issue from another angle. Twenty PD participants and 50 asymptomatic adults were studied; seven PD participants had both ON and OFF recordings. The score-based clustering distinguished PD from asymptomatic participants with sensitivity only 0.40 and specificity 0.96. Associations with MDS-UPDRS explained 24–28% of score variance, and ON/OFF score direction improved in five of seven participants. These are exploratory concurrent scores, not established progression biomarkers. [43]
Estimated power has not automatically demonstrated additive clinical value
Baltasar-Fernandez et al. compared 23 ON-state PD participants with 23 age- and sex-matched controls using 30-second STS count and Alcazar-equation power. The equation incorporated count, mass, stature and chair height; neither a force plate nor an encoder was used as a simultaneous criterion. PD participants completed approximately 3.2 fewer rises and had lower estimated absolute and relative power. However, the PD group was also 10.6 kg lighter, which directly affects absolute power. Correlations with functional measures were generally similar for count and estimated power, and none of the STS variables correlated with Mini-BESTest in the original Results. This is known-groups and concurrent-association evidence, not PD-specific criterion validation of estimated watts or proof of incremental information beyond repetitions. [10]
Serra-Añó et al.’s phone-based FallSkip procedure computed STS power from estimated CoM trajectory and anthropometry in 29 PD participants and 31 controls. Power showed short-interval repeatability (ICC 0.82 in 20 PD participants tested 15 minutes apart by different raters) but did not discriminate PD from controls (approximately 206 versus 211 W; p = 0.8). No simultaneous independent mechanical criterion was used. The article’s reliability table also contains apparent row/unit inconsistencies, so its displayed means should not become reference values. [44]
Malling et al.’s randomized trial demonstrates how force-based RFD can offer a different signal from overall time, while also illustrating the need for restraint. Eighty-two PD participants contributed to the STS analysis. A 1-kHz force plate captured six rapid cycles, with the interval from the first to sixth force peak representing five cycles. The full-group interaction for RFD did not reach conventional significance (p = 0.064), while one performance subgroup showed an interaction (p = 0.049); completion time improved across both treatment groups. This is not proof of a broadly responsive, clinically anchored RFD endpoint, and no MCID for the metric was established. [26]
Picardi et al. also reported that sit-to-walk angular velocity was responsive to inpatient rehabilitation in 20 PD participants and correlated moderately with Mini-BESTest. The available original abstract describes 81 candidate iTUG features and a pre/post design. This is clinical construct association and responsiveness, rather than criterion agreement with a biomechanical reference; it supplies neither a clinically anchored change threshold nor evidence of future fall reduction. [45]
State sensitivity and cueing do not equal progression sensitivity
Medication results depend on the sample, task and endpoint. The mild Inkster cohort changed STS duration between ON and OFF while isokinetic torque did not significantly differ. Morgan found inconsistent individual free-living state effects; Adamowicz’s features poorly separated ON from OFF. In contrast, Kikuchi et al.’s 2026 EMG/kinematic study of 14 analyzed participants selected for wearing-off and levodopa challenge found a longer OFF-state STS duration and seat-off time. That selected fluctuation cohort cannot establish a universal medication-response threshold. Medication withdrawal or challenge testing is a clinician-managed research/clinical procedure, not a default home-testing instruction. [20, 8, 9, 46]
Descent and everyday transfer context
A single STS to quiet standing, repeated STS, TUG sit-to-walk and naturalistic transfer answer different questions. Transition endpoints may be defined by upright posture, head-height plateau, cessation of trunk rotation, foot strike or entry into walking. Combining these under “STS time” can produce an apparently precise but clinically incoherent reference range.
Ascent and descent should be retained separately. Repeated-test duration combines preparation, ascent, stabilization, descent and any pauses. It cannot establish whether a change arose from concentric force generation, cautious lowering, fast uncontrolled sitting or a different trunk strategy. Wairagkar’s model explicitly analyzed both directions. Romijnders et al.’s real-life fatigue analysis also found direction-dependent associations, but only nine PD participants contributed to the pooled 86-person multidisease analysis; it does not establish a PD-specific fatigue detector. [36, 47]
Context should be data, not hidden noise. At minimum record chair height and armrests, assistance, footwear, initial foot position, cueing, natural versus maximum pace, medication state/time since dose, dyskinesia, pain/fatigue, carrying an object and whether the transfer continues into walking. Johansson et al.’s 94-person PD study found a small mean increase in TUG STS duration under serial subtraction (+0.07 seconds), whereas other TUG elements changed more. Such a group mean does not override the measurement error of a different device or become an individual abnormality threshold. [48]
Counts of real-world transfers are jointly influenced by capacity, opportunity and behavior. Fewer transfers could reflect reduced activity, a change in household routine, illness, monitoring coverage or classifier sensitivity. Slower events could reflect a lower chair, carrying a drink or a different endpoint. A monitoring display should show context and usable-event coverage alongside its summary rather than interpreting every trend as neurological deterioration.
Table 3 Selected instrumented comparisons
N refers to people unless events are explicitly specified. Agreement with a measurement reference, association with a clinical score, present-group classification and future prognosis are distinct claims. Full text status does not imply that every supplement or raw dataset was accessible.
| Study and sample | Task, quantity and main estimate | Interpretation and source access |
|---|---|---|
| Inkster 2004 10 PD + 10 controls [21] | Five self-paced rises; optical kinematics plus foot/seat force plates. Similar ON-state duration to controls coexisted with altered hip strategy, CoM displacement and knee moments. | Explains why time can conceal compensation. Joint moments were inverse-dynamics estimates. Small, probably overlapping cohort with Inkster 2003. Full original Methods/Results and tables read. |
| Iwai 2025 40 PD + 23 controls [2] | Five separate comfortable-speed attempts; three force plates. PD groups: 18 always successful, 12 failed after seat-off, 10 failed seat-off. Loading slope and repetitive movement tracked task difficulty. | Current task-defined groups, not longitudinal stages. Some events were unmeasurable in failed attempts. No joint kinematics or prospective prediction. Full original Methods/Results and tables read. |
| Galna 2014 9 PD + 10 controls [32] | Kinect depth camera against concurrent Vicon. Timing and gross STS movement showed useful criterion agreement under controlled frontal recording. | Genuine technical comparison, but does not validate ordinary RGB video, all joint angles or power. Original body read; selected numeric table cells unavailable in the inspected text. |
| Morgan 2023 12 PD + 12 controls [8] | 85 h test-house video; 813 manually identified events. Automated final-ascent duration: r = .419, bias .10 s, random error .19 s. Non-carrying ascent speed correlated with current motor severity. | Manually localized clips; final attempt only; usable-output denominator not fully reconciled. Speed correlated with duration rather than an independent speed criterion. No prospective prognosis. Full original Methods/Results and tables read. |
| Kantha 2026 15 PD + 15 controls [33] | Eight-camera, 60-Hz markerless capture; single and five-rise tasks. Earlier head/trunk initiation, slower extension and longer trunk-adjustment durations in PD. | Kinematics only; not webcam force/power validation. Amplitude/peak-velocity findings did not survive FDR adjustment; no final-sway group difference. Full publisher Methods/Results and tables read. |
| Pham 2018 Held-out test: 21 PD + 11 older adults [35] | Lower-back IMU during daily-like activities. Combined sit-to-stand and stand-to-sit detection accuracy 82% in PD without dyskinesia, 47% with dyskinesia; direction accuracy 98% overall. | Phenotype matters. Reference duration was in whole seconds and the reported duration-difference interval was wide. Home-like rather than unsupervised home setting. Full original Methods/Results and tables read. |
| Adamowicz 2020 20 PD in validation; 314 complete rises [9] | Lower-back accelerometer: sensitivity .853, precision .988. Duration bias −.20 s; limits of agreement −1.71 to 1.30 s. | Excellent precision did not ensure precise duration. 105 partial transfers excluded. Three-day monitoring reliability came from healthy adults. Full original Methods/Results and tables read. |
| Sher 2026 35 adults, only 5 PD [37] | Phone IMU during 30-s chair stands. Participant-mean cycle ICC .998, bias −.012 s; limits of agreement −.134 to .110 s against video. | Pooled technical timing validation; only five PD participants, no independent external PD cohort. Video frame rate inconsistently reported. No power measurement. Full original Methods/Results and tables read. |
| Khalil 2024 262 PD + 50 controls [42] | Back-to-back iTUG trials. Total time generally reliable, while selected sensor-feature median ICCs were .36–.61 and chair transitions often less consistent. | Current classification utility and individual monitoring reliability are different. High ICC can coexist with large within-person changes. Full original Methods/Results read; individual-feature supplement values not independently extracted. |
| Baltasar-Fernandez 2024 23 PD + 23 controls [10] | 30-s STS count and anthropometric equation power. PD: −3.2 rises, −97 W and −.85 W/kg; associations generally similar to repetition count. | No mechanical criterion validation. PD group was 10.6 kg lighter, directly affecting absolute power. No established incremental benefit over count. Original Methods/Results read. |
| Serra-Añó 2020 29 PD + 31 controls; 20 PD reliability [44] | Phone CoM/anthropometry power model. ICC .82 across two different-rater tests 15 min apart. Estimated power did not differ between PD and controls. | Short-interval repeatability, not between-day reliability or independent mechanical validation. Apparent row/unit inconsistencies in the reliability table. Original Methods/Results and tables read. |
| Malling 2018 82 PD in randomized STS analysis [26] | Force-platform RFD from 30–70% of rising force peaks during six cycles. Overall treatment interaction p = .064; high-performance subgroup p = .049. | A different mechanical endpoint from timing, with no established MCID. Subgroup result is not a positive overall trial. Custom peak-to-peak interval differs from conventional stopwatch timing. Full original Methods/Results read. |
Reliability and meaningful change
What is actually reproducible
Reliability belongs to a particular outcome, scoring method, participant spectrum and retest condition. It is not a permanent attribute of the label “sit-to-stand.” Four distinct questions recur in the PD literature: whether two observers time the same performance similarly; whether a person repeats a test similarly within a session; whether performance recurs on another day; and whether a sensor agrees with an external reference. These can have very different answers. An almost perfect observer ICC cannot be substituted for day-to-day reproducibility, and strong correlation between measurement methods is not proof of narrow individual agreement limits.
Verheyden's single-rise study illustrates observer precision directly. In 38 ON-state participants with PD, interrater ICC was 0.95 (95% CI 0.91–0.97) and intrarater ICC was 0.98 (0.96–0.99). Reported smallest detectable differences were 0.61 s and 0.39 s, respectively. These estimates came from timing the same video-recorded performance, with pause/replay allowed; the apparent one-week interval was between a rater's readings, not between patient performances. The paper explicitly limits these SDDs to that observation design. They are not between-day MDCs, live-versus-video agreement margins, or patient-important change thresholds. [16]
Duncan's 2011 study established a valuable early clinical evidence base but a small one for repeatability. Although the clinical/validity sample comprised 80 participants, the reliability subsample comprised only 10. Two raters simultaneously timed two trials, and the procedure was repeated seven days later. The paper reports interrater ICC(1,1)=0.99 and test–retest ICC(2,1)=0.76, without ICC confidence intervals, SEM or MDC. The main 80-person assessment used one trial, which the authors identified as a limitation; this should not be confused with the two-trial reliability procedure. The reported whole-cohort SD also contains seven assigned 60-s failure scores. Combining that SD with the 10-person retest ICC to manufacture an MDC would mix incompatible datasets. [1]
Paul's 2012 primary report and recovered supplement provide a stronger absolute-error anchor for one specific protocol. Thirty of 31 independently ambulant participants contributed 5×STS measurements, tested a week apart during optimal ON medication at a comparable point in the medication cycle. Using a 45-cm chair, arms folded, maximum speed and the faster of two trials, baseline time was 9.67±1.79 s and retest time 9.48±2.04 s; ICC(3,1)=0.91 (95% CI 0.82–0.96) and SEM=0.6 s. The one STS non-observation is not explained sufficiently to label it a motor failure. The sample's baseline range of 5.9–13.5 s is substantially faster and less dispersed than Duncan's mixed completer/noncompleter distribution. Study population and scoring are plausible contributors to the differing reliability estimates, but the studies did not directly randomize chair height or trial aggregation to establish the cause. [4]
A conventional individual MDC95 calculated from Paul's SEM is 1.96×√2×0.6=1.66 s. This is a derived calculation, not an MDC printed in the article, and it assumes independent measurement error and applicability of that SEM to the same fastest-of-two protocol. Because SEM was rounded to one decimal place, the derived number should be reported approximately (about 1.7 s), not with spurious hundredth-second certainty. Likewise, stopwatch recording to a millisecond in the supplement is reporting resolution, not evidence of millisecond accuracy. [4]
Petersen's separate n=22 H&Y I–III study found considerably greater 5×STS error: ICC(2,2)=0.74 and MDC95=10 s, versus ICC(2,2)=0.94 and MDC95=3 repetitions for 30-second STS, over a 6–8-day interval. The abstract does not provide the ICC confidence intervals or enough protocol detail to reconcile the difference fully. The “2” average-measures ICC notation is also important: it should not casually be quoted as a one-trial reliability coefficient. This is direct PD evidence in favor of a fixed-duration repetition count for follow-up in that study, but it does not justify pooling the two tests or averaging the discrepant MDCs into a single PD threshold. [5]
Soares and colleagues compared one complete 5×STS trial with averages of two and three after familiarisation. The paper reports 52 participants for interrater testing and 50 for 7–14-day retesting. Raters administered separate performances independently, in randomised rater order; they did not simply time the same trial simultaneously. For 5×STS, interrater ICC rose from 0.68 (95% CI 0.506–0.806) for one trial to 0.87 (0.776–0.926) for the mean of two and 0.85 (0.749–0.917) for the mean of three. Retest ICC was 0.73 (0.569–0.838), 0.86 (0.762–0.923) and 0.86 (0.757–0.922), respectively. Similar group means therefore do not establish interchangeable individual scores or equivalent agreement across the three summaries. The authors' blanket one-trial conclusion deserves qualification for STS. [6]
The original specifies rapid rises without upper-limb use, familiarisation followed by a one-minute rest, the same stopwatch and no cueing beyond the start command. Retesting used the same shift and environment, but exact chair height, timer events and medication-cycle standardisation were not reported. No SEM, MDC or limits of agreement were supplied. Although the reported retest total is 50, Table 1 sex and stage counts sum to 48; this unresolved reporting discrepancy does not justify silently changing the stated sample. The evidence supports a trade-off between assessment burden and relative reliability, rather than a universal single-trial protocol or a new individual change threshold. [6]
Practice and trial aggregation
A demonstration, one familiarization rise, a full practice 5×STS and repeated scored trials impose different learning and fatigue exposures. The published evidence does not establish one universally optimal combination across PD severity. A pragmatic clinic can use a demonstration and a safe familiarization assessment, then a prespecified number of scored trials with rest, but it should state when that is a local reproducible protocol rather than an exact published one. Retaining all scored trials preserves information about inconsistency and failed attempts even when an average or fastest value is the primary endpoint.
Fastest-of-two estimates best observed capacity and may suppress a slow or failed attempt. The mean estimates average observed performance and can reduce random error, but repeated practice can introduce systematic change. Worst-of-five in Iwai was chosen specifically to reveal intermittent failure, not to optimize a timed score's ICC. These summaries are not competing answers to the same question. Choose the question first, prespecify the summary, and do not switch from fastest to average between visits. A fixed repeat protocol does more for interpretability than selecting whichever published ICC is largest. [4, 2, 6]
Responsiveness and MIC are separate from MDC
An MDC is a statistical boundary for change beyond an estimate of measurement error; an MIC or MCID asks whether a change is important. Neither a significant group mean improvement nor a correlation with another mobility test establishes an individual clinically important change. Rehabilitation can also improve a practiced chair task through strategy, confidence or coordination without demonstrating a corresponding gain in isolated strength or power.
Spagnuolo and colleagues provide direct, but preliminary, anchor-based PD responsiveness evidence. Thirty of 35 enrolled participants completed eight weeks of group physiotherapy, including 16 one-hour sessions; four were excluded for non-attendance and one after a fall at home. The completers spanned H&Y I–IV, including three at stage IV. Assessments occurred in the afternoon, ON medication within two hours of dosing, using the same trained assessor who did not deliver the intervention. Their 5×STS used a 43-cm chair with armrests, crossed arms and initial back contact, with timing from the assessor's signal to backrest contact after the fifth rise. That endpoint differs from Duncan's fifth buttock-to-seat contact. The number of scored trials and aggregation rule were not specified. [7]
Table 3 reports 5×STS improvement from 21.3±7.8 to 17.4±5.6 s, a change of 4.0±4.5 s and SRM 0.9. Nineteen participants reported improvement on a task-specific five-point perceived-change scale administered after testing. Figure 2 reports AUC 0.92 (95% CI 0.77–0.99); Table 4 gives a >2.5-s improvement threshold with sensitivity 73.3% and specificity 100%. However, the Results text says all AUC confidence intervals included 0.50, contradicting the figure, and its narrative change values do not clearly reconcile with Table 3. The exact rule used to dichotomise the five-point anchor is insufficiently explicit. These inconsistencies, the small uncontrolled completer sample, task practice during therapy, same-sample threshold selection and absence of independent validation argue for treating >2.5 s as a preliminary intervention- and anchor-specific response threshold. It is neither a universal PD MIC nor an MDC, and cannot resolve Petersen's larger protocol-specific measurement error. [7, 5]
PARKEX demonstrates that one-minute counts can change following training, but it supplies neither reliability-based MDC nor anchor-based MIC. Its abstract calls the 31- and 19-repetition exercise-versus-control contrasts “mean differences,” whereas the methods and Table 4 define them as median differences. These group contrasts are not individual change thresholds; the outcome was secondary, eight participants were randomized per arm; after one BPT participant was reassigned to control following a lesion, the analyzed groups contained seven, eight and nine participants, and baseline clinical imbalances were present. Consequently, the trial supports considering one-minute STS as an intervention outcome for further validation rather than adopting those large contrasts as normative expectations. [3]
For an individual follow-up, report the absolute change and direction with identical protocol/medication conditions, then state which directly applicable error estimate and, separately, importance anchor is being used. A change can exceed measurement error yet be trivial to the person; an important functional gain can also be missed by a rigid arm-free timed scale. Moving from assisted to unassisted rising, or from intermittent failure to consistent success, should be reported alongside time/count change. If the protocol changed because function improved, preserve the categorical milestone and restart like-for-like numerical comparisons rather than calculating a misleading percentage improvement across different tasks.
Measurement floors and reference values
The most immediate floor is inability to perform the stipulated movement. Fixed-time counts handle fewer than five successful repetitions more naturally than a five-repetition completion time, but still have a zero-count floor when unassisted rising is impossible. A timed measure lacks a simple upper score ceiling for good performers, yet small biological differences may become indistinguishable from stopwatch and event-definition error; a three-level clinical item can lose discrimination much earlier. Formal floor/ceiling percentages and PD-specific age/sex-stratified norms were not established by the core studies reviewed here.
Duncan's 16-s threshold classified retrospective repeat-faller status, defined as at least two falls in the previous six months. AUC was 0.77, sensitivity 0.75 and specificity 0.68 in the derivation sample; confidence intervals and external validation were not supplied. It is not a validated probability of falling in the next six months, a diagnostic boundary for weakness, or a universally abnormal age-adjusted time. Afshari later reused that threshold in a remote feasibility study without independently validating its discrimination. Generic older-adult thresholds, sarcopenia algorithms and chair-stand reference charts may provide clearly labeled contextual comparisons, but they should not be relabeled as PD norms or substituted for a PD-specific prognosis model. [1, 24]
Table 4 Clinical measurement and change evidence
Intervals are 95% confidence intervals. MDC is not MIC. Access identifies the original material examined, not study quality. The narrative describes details unavailable from partial sources.
| Study and protocol | Principal result | Interpretation and source access |
|---|---|---|
| Duncan 2011 [1] 5×STS, 43-cm chair; ON; clinical n=80, reliability n=10; retest 7 days | Interrater ICC(1,1)=0.99; retest ICC(2,1)=0.76. No SEM or MDC reported. Seven non-completers assigned 60 s. | Main cohort used one trial; reliability procedure used two. Assigned values are not measured times. Original body and tables inspected. |
| Paul 2012 [4] 5×STS, 45-cm chair; faster of two; matched ON; n=30; retest 7 days | ICC(3,1)=0.91 (0.82–0.96); SEM 0.6 s. Calculated MDC95 approximately 1.7 s. Baseline mean 9.67 s. | MDC calculated from rounded SEM, not reported by authors. Start/final event insufficiently specified. Original publisher methods/results/tables and supplement inspected. |
| Petersen 2017 [5] 5×STS and 30-second STS; n=22; H&Y I–III; 6–8 days | 5×STS ICC(2,2)=0.74 and MDC95=10 s; 30-second STS ICC(2,2)=0.94 and MDC95=3 repetitions. | Average-measures ICC; not single-trial reliability or a universal change threshold. Original abstract inspected; full protocol and confidence intervals unavailable. |
| Soares 2026 [6] One versus mean of two/three complete 5×STS trials after familiarisation; reported n=52 interrater, n=50 retest; 7–14 days | Interrater ICC: .68, .87 and .85; retest ICC: .73, .86 and .86 for one, mean of two and mean of three. Single-trial CIs: .506–.806 and .569–.838, respectively. | Averaged scores had higher ICC estimates; similar means do not establish equivalent individual agreement. No SEM, MDC or LoA reported; unresolved Table 1 counts. Complete online-first original PDF and tables inspected. |
| Spagnuolo 2018 [7] 30/35 completers; H&Y I–IV; 8-week group physiotherapy; 43-cm chair with armrests; timer to fifth backrest contact; ON | 5×STS SRM .9; mean improvement 4.0 s. Figure AUC .92 (.77–.99); Table 4 threshold >2.5 s, sensitivity 73.3%, specificity 100%; 19 reported improvement. | Preliminary anchor-specific threshold, not universal MIC or MDC. Small uncontrolled sample, internal reporting contradictions and same-sample selection. Complete original PDF, tables and ROC figure inspected. |
| Verheyden 2014 [16] n=38 PD and 19 controls; usual-way single rise from armchair; one recorded performance timed by three raters, with one repeated reading | Observer interrater ICC .95 (.91–.97), intrarater ICC .98 (.96–.99); SDD .61 s and .39 s, respectively. | Precision of timing the same recording, not patient day-to-day reproducibility or live/video equivalence. Complete original PDF and tables inspected. |
| Afshari 2022 [24] Remote 5×STS; n=15; four ON televisits; two trials each; care partner and prior instruction | All 120 assessments ratable; no falls, near-falls or adverse events. Same 43–45-cm home chair used. | Selected supervised feasibility; no validated remote fall-risk cutoff or clinic–home agreement margin. Original full text inspected. |
| Silva-Batista 2026 [25] n=62; in-person first, live video second 1–4 days later; different raters, same home chair, matched ON/time of day | 5×STS r=.95; ICC=.97 (.95–.98). Published LoA −3.3 to 4.4 s; mean difference magnitude about 0.5 s. | Fixed order and several-second individual differences limit interchangeability. Prepared home/technology workflow and safety exclusions; prose and table use differing bias directions; convention unresolved. Complete publisher body, tables and captions inspected; figure images unavailable. |
Prognosis and future outcomes
Main interpretation
Sit-to-stand assessment has a defensible place in the functional examination of a person with established Parkinson’s disease (PD), but its prognostic evidence is substantially narrower than its widespread use might suggest. Slower five-times sit-to-stand (5STS) performance has preceded recurrent falls in small prospective cohorts. However, the association weakened after accounting for falls history, disease characteristics and multidomain balance in directly inspected adjusted studies. A separate prospective study of the 30-second test found essentially chance discrimination for future falls. Chair-rise ability has also appeared in mortality screening, progression composites and a recent machine-learning falls model. Those findings do not validate one interchangeable “STS predictor”: timed repetitions, a single instrumented transition, need for assistance and an ordinal clinical item measure different things. [49, 12, 13, 14, 50, 51]
The appropriate clinical inference is therefore modest: poor performance identifies a transfer problem worth investigating and may contribute to a wider risk assessment. It does not currently support a universal time threshold for forecasting falls, dependency, dementia or survival. Conversely, the evidence is not simply absent. Important positive, negative and qualified longitudinal findings are set out below, with concurrent associations and treatment monitoring kept separate.
Prospective falls evidence
Five times sit to stand and recurrent falls
Mak and Pang followed community-dwelling, independently ambulant PD participants for 12 months, contacting them monthly about falls. Seventy-four entered follow-up and 72 completed it; two died. The completed PD sample included 47 non-fallers, 12 single fallers and 13 recurrent fallers, with 133 falls overall. Baseline assessment was ON medication, within two hours of medication intake. Participants crossed their arms, stood fully and sat down five times as quickly as possible; timing ended with buttock contact on the fifth sitting. [49]
Baseline 5STS times were 13.4 ± 5.7 seconds in non-fallers, 12.9 ± 3.3 in single fallers and 17.8 ± 5.7 in recurrent fallers. Recurrent fallers were slower than non-fallers (p=.016) and single fallers (p=.048), while non-fallers and single fallers were similar. This is genuinely temporally ordered evidence, but it is a comparison between subsequent outcome groups, not an adjusted individual-risk model. Falls history, disease severity, gait, endurance and confidence also differed, and no independent 5STS effect, calibration, validated threshold or external validation was established. The reported injury characteristics were descriptive; they did not validate STS as a predictor of injury. [49]
Mak and Auyeung subsequently studied 112 eligible patients, retaining 110 after two losses to contact over six months. Twenty-four had more than one fall. Baseline arms-crossed 5STS was again tested ON medication, with timing through the fifth sitting. Future recurrent fallers averaged 21.0 ± 9.7 seconds, compared with 15.7 ± 7.3 in others; the reported between-group difference was 5.4 seconds in magnitude (95% CI 1.8–8.9). The unadjusted odds ratio per additional second was 1.073 (p=.014). [12]
The adjusted analysis changes the interpretation. The 5STS odds ratio fell to 1.062 (p=.082) after demographics, prior falls, Hoehn and Yahr stage, depression, motor severity and freezing were included, and to 1.034 (p=.368) after Mini-BESTest was added. It was therefore not an independently significant predictor in the final model. This does not prove that transfer function is irrelevant; correlated tests and only 24 events limit the precision of a ten-predictor model. It does show that an unadjusted 5STS association should not be presented as added prognostic value beyond a wider examination. The study reported no STS-specific calibration or validated 5STS cut-off. Its whole-model AUC of .895 is not a 5STS AUC. Some accuracy and likelihood-ratio entries also disagree internally, further discouraging overprecise model-performance claims. [12]
The eight variable model and the simpler clinical rule
The full Paul study makes the incremental-value question unusually clear. Two hundred and five community-dwelling participants were assessed at home ON medication; 133 were controls from two exercise trials and the remainder were other eligible volunteers. Methods specified fast five-repetition STS with arms folded. Results and Tables 1 and 4 consistently report 120 future fallers over six months, with 1,854 falls and 413 injurious falls. The abstract instead says 125; the internally consistent Results/Table denominator is used here. Monthly diaries and telephone follow-up ascertained falls. [11]
Future fallers averaged 14.5 ± 6.3 seconds versus 11.7 ± 4.8 in non-fallers. The univariate STS odds ratio was 1.11 per second (95% CI 1.04–1.18), but the adjusted odds ratio was 1.04 (.96–1.13; p=.38). The full model actually contained eight predictors: falls history, freezing, knee-extensor strength, STS, gait speed, narrow-base standing, foam sway and coordinated stability. The abstract lists only seven. STS survived backward selection in 33% of 1,000 bootstrap samples, compared with 100% for falls history, 82% for freezing and 61% for gait speed. [11]
The eight-variable AUC was .83 (.77–.88), with a bootstrap optimism-adjusted AUC of .81 (termed “zero-corrected” by the authors). The final three-variable rule, which excludes STS, had AUC .80 (.73–.86), optimism-adjusted .78; the difference from the full model was not significant (p=.14). Its Hosmer–Lemeshow test gave p=.61, but lack of a significant lack-of-fit test does not establish transportable calibration. No independent added STS contribution was demonstrated in this model once the other predictors were considered; the confidence interval still permits a modest association, and the model comparison is not a formal equivalence test. [11]
Non-completion handling also matters. Inability on a low-is-better mobility test was assigned the sample mean plus three standard deviations; other missing continuous data were estimated from similar tests. Fewer than 6% were unable to perform any given mobility measure, but the specific STS count was not stated. This is a pragmatic analysis of a mixed performance/inability score rather than wholly observed continuous STS times. Removing imputed cases reportedly did not materially change results. Initial candidate screening, pooled trial-control recruitment and restricted cognition remain limitations despite the bootstrapping. Injury counts were reported, but the fitted endpoint was any fall, not injury. [11]
External validation of the simple rule is useful context, but the rule excludes STS. For example, Duncan and colleagues studied 171 people, of whom 66 fell over six months, and reported AUC .83 (.76–.89). Those results validate the three-variable falls rule in that cohort, not the omitted STS component. [52]
Thirty second repetitions provide a useful negative result
Moraca and colleagues followed 96 community-dwelling participants for 12 months with weekly personal or telephone fall ascertainment. Thirty-six reported at least one fall, with 56 falls in total. The test was a 30-second arms-crossed chair stand, with full hip and knee extension required and verbal encouragement, tested approximately one hour after medication. It was not a five-repetition timed test. [13]
Among the 90 participants with STS data, the AUC was .50 (95% CI .38–.63; p=.96). Mean performance was approximately 15.7 repetitions in both future fallers and non-fallers. The exploratory 15.5-repetition cut-point had sensitivity .56 and specificity .41 and has no defensible role as a validated risk threshold. Combining STS, TUG, Berg balance and six-minute walking did not rescue prediction: the four-measure AUC was .52 (.38–.66), based on 73 complete cases. [13]
This is meaningful negative evidence, but its scope is specific. Participants were predominantly mildly or moderately affected, able to perform the tests and involved in an exercise programme. Eight of 104 eligible participants failed to complete follow-up, including two deaths; STS and combination models had additional missing measurements. These restrictions may narrow the range of function and limit generalisation to frailer patients. The endpoint was any fall, rather than recurrent or injurious falls, and the analysis did not test incremental value over falls history. [13]
Newer models and remote measurements
Kitaichi and colleagues analysed 543 PPMI participants, including 93 with falls at the one-year follow-up. The supplement identifies the chair measure as MDS-UPDRS item 3.9, NP3RISNG, an ordinal clinical assessment rather than timed or instrumented STS. The reported 434-person learning set was used for categorical data analysis program (CATDAP) feature selection, random-forest tuning and model development; 109 participants were held out for testing. Tuning used repeated five-fold cross-validation and SMOTE-ENC oversampling of learning data. The authors state that the whole process was repeated across 101 random splits, although they do not present the resulting distribution of performance estimates. [51]
The selected five-variable model had test sensitivity 63.2%, specificity 86.7%, positive predictive value 50%, negative predictive value 91.8% and Matthews correlation coefficient 0.456. Its 82.6% accuracy should be considered against the cohort's marked class imbalance: 82.9% were non-fallers overall. Importantly, chair rise ranked third by CATDAP's association criterion (AIC −3.18), after fall history and glaucoma; this is not a random-forest feature-importance estimate or an isolated adjusted chair effect. No chair-item ablation, incremental comparison with a simple clinical baseline, probability calibration or external cohort validation was reported. Within-fold re-selection and oversampling details, exact chair-item encoding, medication state, missing-data handling and the precise falls recall window were insufficiently described. The supplement's correlation-star footnote states p>0.05; this is retained as a reporting inconsistency rather than silently reversed. The study provides a prospective clinical-item lead within a multivariable model, not a validated standalone chair-rise risk score. [51]
Marano and colleagues provide an instructive temporal limitation. In 33 participants completing smartphone tests over four weeks, four scored at least one on original UPDRS-II item 13, “falling unrelated to freezing,” for that same period. Stand-up time averaged over the monitoring period was longer among fallers: median 2.28 versus 1.77 seconds. The univariate odds ratio was 1.7 per 0.1 second (p=.016); the authors also reported an association alongside baseline falls history. However, the sensor predictor includes measurements acquired during the outcome window, rather than a strictly prior baseline. With four events, no fall counts, no reported coefficient confidence interval or external validation, this is exploratory concurrent-period monitoring evidence, not a deployable four-week baseline risk rule. [53]
A further home-sensor study by Nouriani and colleagues included 11 people with PD, eight with normal-pressure hydrocephalus and ten controls. Although its activity-recognition algorithm identified STS transitions, its subsequent one-year falls analysis was pooled across movement disorders, with diaries available for 17 patients. The strongest novel associations concerned near-fall frequency; “sitting frequency” denoted time spent sitting, not chair-rise performance. Recognition accuracy for STS events and prospective performance of the near-fall model therefore cannot be presented as a PD-specific STS prognostic model. [54]
Independence and future functional decline
Direct chair-rise ability is relevant to current independence. Bryant and colleagues examined 88 participants cross-sectionally using original UPDRS item 27: 54 rose normally, 24 rose slowly or needed repeated attempts and 10 pushed up using their arms. No participant occupied the two most impaired categories. Arm-assisted rising remained associated with lower concurrent Schwab–England ADL performance and lower physical activity in selected regression models. The authors explicitly called for longitudinal research. These are not estimates of how soon independence will be lost, and the apparent arm-use threshold was not prospectively validated. [55]
There is real prospective evidence for broader axial impairment, including chair rise, but attribution matters. Macleod and colleagues developed mortality and dependency models in 198 newly diagnosed PINE participants and externally evaluated them in 192 ParkWest participants. Their axial score included speech, facial expression, facial tremor, neck rigidity, arising from chair, posture, gait and postural instability. A five-point higher axial score was associated with dependency (HR 1.74, 95% CI 1.28–2.35) and mortality (HR 1.33, 1.08–1.66) in the corresponding multivariable models. The models required recalibration because risk was overestimated in the external population. None of these estimates isolates chair rise, and the authors identified determining which axial components are required as future work. [56]
A stronger test of richer sensor batteries gave disappointing results. Dewey and colleagues analysed baseline APDM iTUG measurements, including STS and sitting-transition variables, in treated PD. Samples with clinical follow-up numbered 230 at six months, 222 at twelve, 164 at eighteen and 177 at twenty-four months. Repeated cross-validation of ridge-regression models showed poor out-of-sample prediction of later motor, cognitive, ADL and mobility-related quality-of-life change. This is not an isolated negative STS estimate, because the full sensor battery was modelled; it is evidence against assuming that measuring many transfer and gait features automatically produces useful prognosis. [57]
Motor and cognitive progression
Skidmore and colleagues' emerging postural instability rating scale provides a qualified item-level lead. Its seven baseline items include reported difficulty getting out of bed, a car or a deep chair, and examined arising from chair. In a PPMI derivation sample of 301 newly diagnosed, untreated patients, 85 developed Hoehn and Yahr stage 3 or worse; an additional 79 idiopathic and 141 genetic PD participants were used for validation analyses. The examined chair item received the highest integer weight, six points, in the composite. Higher composite scores were associated with later postural instability and faster decline in selected cognitive measures. [50]
This is neither a timed STS model nor a chair-specific dementia forecast. The validation cohorts came from PPMI rather than a wholly independent clinical setting, and the reported validation hazard ratios used the derivation low-quartile group as reference. The article contains inconsistent sample totals and descriptions of the outcome window and quartile construction. Requiring at least five years of data also selects people who remained observable. These issues, together with absent chair-specific calibration or ablation, make the findings hypothesis-generating for transfer assessment rather than a ready clinical risk calculator. [50]
Current cognitive correlations with single- or dual-task transfer measures should be interpreted as possible shared motor–cognitive burden. The inspected evidence did not establish a validated standalone timed or instrumented STS predictor of incident PD dementia. Findings in initially non-PD older adults, prodromal cohorts or people with mild cognitive impairment who later develop Lewy body dementia cannot answer that established-PD question.
Mortality and institutionalisation
Chair-rise ability has been examined as a predictor of mortality in PD. Gray and colleagues linked baseline physical assessment in 109 people with PD to seven-year mortality; 46 died and national records supplied complete follow-up. The baseline chair-rise ability rating was associated with mortality in unadjusted screening. Crucially, the separate timed sitting-to-standing measure was not marked significant. The ability scale distinguished unassessable performance, inability, rising with an aid and rising without an aid. The final Cox model retained age, sex and Tinetti gait, not chair rise. [14]
The defensible conclusion is that loss of transfer ability can accompany a poorer overall prognosis, while a standalone timed-STS survival tool has not been validated by this study. Macleod's axial models provide complementary composite evidence, not confirmation of an isolated STS effect. No directly validated chair-rise model for later institutionalisation was established in the studies examined. This bounded conclusion should not be confused with proof of no association. [14, 56]
Injurious falls and fractures
Fall-frequency studies cannot automatically predict injury, which also depends on bone fragility, the direction and energy of a fall, protective responses and exposure. Schini and colleagues examined 5,212 community-dwelling women aged at least 75 years, including 47 with PD, over a reported average 3.8 years; 11 women with PD sustained 12 osteoporotic fractures. Fractures were collected at six-monthly home visits and verified from clinical or radiographic records. The original Methods specify a nurse-observed single rise from an armless chair, categorised as without difficulty, with difficulty or unable. It does not establish a timed 5×STS protocol or explicitly describe hand placement. [58]
In Table 2, adding difficulty with STS to height-and-treatment adjustment left the PD-versus-no-PD fracture hazard ratio at 2.16 (95% CI 1.19–3.93; n=5,158). That is the coefficient for PD after adjustment, not the prognostic effect of STS within PD. Adding maximum quadriceps force attenuated the PD coefficient to 1.70 (0.81–3.59), but the included sample fell to 4,611 from 5,161 in the baseline adjusted model; this is not evidence of causal mediation. The complete publisher article text, all four tables and the original supplement were inspected; the figure image was unavailable, and the supplement reports sway measures rather than a separate chair-rise prognostic analysis. Recruitment to a trial and likely exclusion of more severely affected women limit generalisation. The study therefore does not establish a validated standalone STS fracture-prediction rule. [58]
Similarly, a trial that improves chair-stand performance and reduces injurious falls does not show that baseline STS predicted benefit, or that its change mediated injury reduction. In the examined evidence, no externally validated standalone STS model for injurious falls or fractures in established PD was identified.
Treatment response prediction versus measuring change
Di Lazzaro and colleagues followed 40 newly diagnosed, drug-naïve patients, with 36 retained after two withdrawals and two diagnostic revisions. A fourteen-sensor protocol included a six-metre TUG. Baseline STS duration correlated with percentage improvement in treated MDS-UPDRS III at 30–36 months (r=−.432, p=.01), but was absent from the final multivariable model. That model included other motor features and levodopa-equivalent dose measured at the final visit. It therefore did not constitute a fully baseline, validated STS forecast. [59]
The outcome mixed treatment response with disease evolution: average motor scores improved after therapy began. This makes it inappropriate to interpret the STS correlation as an untreated neurodegeneration rate. Multiple feature screening, 36 completers, no external validation and use of percentage change further limit individual prediction. It is nevertheless a genuine temporally ordered association and should be acknowledged as such. [59]
A more tentative treatment-selection lead comes from Rahimi and colleagues' uncontrolled rasagiline pilot. Eighteen optimally medicated patients with difficult freezing were enrolled for 90 days. The primary abstract reports rise-from-chair rating among six variables associated with improvement, in an outcome-defined responder comparison with adjusted R²=.9898. Subgroup counts total only 14, and full-text details were unavailable. Such near-perfect fitting in a very small selected sample is a warning for overfitting, not strong clinical validation. Without a comparator treatment and a treatment-by-baseline-STS interaction, it cannot establish that a given patient should receive rasagiline because of their chair-rise score. [60]
Most rehabilitation and ON/OFF studies use STS as an outcome: they ask whether performance changes with exercise, cueing, medication or stimulation. This is clinically useful responsiveness evidence, but is distinct from predicting who will respond. A reduction in STS time is not, by itself, a proven surrogate for fewer falls, preserved independence, slower cognitive decline or longer survival.
Table 5 Prospective and longitudinal outcome evidence
Estimates apply to the stated endpoint and protocol. Intervals are 95% confidence intervals. A composite-model result is not an isolated chair-rise effect.
| Study and later outcome | Transfer result | Interpretation and source access |
|---|---|---|
| Mak and Pang 2010 [49] 72 completers; 13 recurrent fallers; 12 months | 5×STS mean 17.8 s in recurrent fallers versus 13.4 s in non-fallers and 12.9 s in single fallers. | Prospective group differences, without an adjusted individual-risk model or validated threshold. Full accepted manuscript inspected. |
| Mak and Auyeung 2013 [12] 110 participants; 24 recurrent fallers; 6 months | Unadjusted OR 1.073 per second (p=.014); final adjusted OR 1.034 (p=.368), including Mini-BESTest. | No independent final-model association; 24 events for ten predictors. Whole-model AUC .895 is not an STS AUC. Full original PDF inspected. |
| Paul 2013 [11] 205 participants; 120 any-fall cases in Results/tables; 6 months | Unadjusted STS OR 1.11 (1.04–1.18); adjusted OR 1.04 (.96–1.13). STS retained in 33% of bootstrap samples. | Eight-variable development model; simpler three-predictor rule excludes STS. Internal bootstrap assessment; inability/missing values imputed. Full publisher HTML inspected. |
| Moraca 2022 [13] 96 followed; 90 STS records; 36 any-fall cases; 12 months | 30-second STS AUC .50 (.38–.63). Four-test combination AUC .52 (.38–.66), n=73 complete cases. | Meaningful null result in a selected mostly mild/moderate exercise-programme sample. Different task and endpoint from 5×STS recurrent-fall studies. Full original PDF inspected. |
| Kitaichi 2026 [51] PPMI n=543, 93 fallers; learning n=434, test n=109; fall status at one-year follow-up | MDS-UPDRS item 3.9 ranked third by CATDAP AIC (−3.18). Selected random forest: MCC .456, sensitivity 63.2%, specificity 86.7%. | Internal holdout, not external validation; AIC association rank is not random-forest importance or isolated incremental STS value. No calibration or chair ablation. Complete publisher HTML and original Word supplement inspected. |
| Macleod 2018 [56] PINE n=198; ParkWest n=192 external cohort; dependency and mortality | Per five-point axial score: dependency HR 1.74 (1.28–2.35); mortality HR 1.33 (1.08–1.66). Chair rise is one component. | Externally examined composite models required recalibration. Neither HR isolates transfer ability. Original body and tables inspected. |
| Dewey 2022 [57] iTUG battery; follow-up n=222 at 12 months and 177 at 24 months | Repeated cross-validation showed poor prediction of motor, cognitive, ADL and mobility-related quality-of-life change. | A negative sensor-battery result, not an isolated negative STS coefficient. Original body and tables inspected. |
| Skidmore 2022 [50] PPMI derivation n=301; validation analyses n=79 idiopathic and 141 genetic PD | Seven-item composite included chair rise with six-point weight; associated with later postural instability and selected cognitive decline. | No isolated timed-STS or dementia rule. Same-source cohorts and reporting inconsistencies limit transportability. Original body/tables inspected; supplement not separately inspected. |
| Gray 2009 [14] 109 participants; 46 deaths; 7 years | Chair-rise ability rating associated with death in unadjusted screening; separate timed-rise measure not significant. | Final Cox model retained age, sex and Tinetti gait, not chair rise. Ability and timed performance differ. Full original PDF and table inspected. |
| Di Lazzaro 2021 [59] 36 completers from 40 drug-naïve patients; 30–36 months | Baseline instrumented STS duration correlated with treated motor-score improvement, r=−.432 (p=.01); absent from final model. | Mixed treatment response and disease evolution; final model also used follow-up medication dose. No validated baseline STS forecast. Original body and tables inspected. |
Implications for rehabtools
Preserve a clinically interpretable result
Use the conventional clinical task as the anchor and add only descriptors with a clear purpose. The report should show the exact test variant, timing or counting rule, trial-level results and chosen summary. Chair height, hand support, foot position, medication timing, cues and the intended endpoint should travel with the result. Sensor placement, camera geometry, processing version and manual annotation belong in the acquisition record when technology is used.
Keep completion status beside the number. Distinguish independent completion, specified adaptation, physical assistance, motor non-completion, a safety stop and technical failure. Retain valid repetitions, unsuccessful attempts and the reason for stopping. A change from hand-assisted to independent rising can be clinically informative even when no comparable arm-free baseline time exists. Restart like-for-like numerical comparisons when the protocol changes rather than calculating a percentage improvement across different tasks.
Show what an added feature measures. A useful label might be final-ascent head velocity, foot-loading slope under a defined algorithm, or ascent duration from a stated event pair. Do not shorten these to power, strength or stability when the measurement does not support that interpretation. Keep the underlying video or trace reviewable and provide an explicit low-confidence or unscorable result when tracking or event detection fails.
Make change interpretation source specific
First check that the testing conditions match. Then show the observed change, the applicable measurement-error estimate and, separately, any patient-important change anchor. Identify calculated thresholds as calculated. If no adequately matching threshold exists, show magnitude and direction without an automatic clinically important label. A reference distribution, a change threshold and a future-risk cutoff should have distinct labels and purposes.
Naturalistic monitoring needs event coverage and context as well as a trend line. Changes in chair use, carrying, activity opportunities, dyskinesia and detection success can alter a daily summary. Fewer or slower detected transfers alone do not establish neurological deterioration.
Validate each additional claim
The evidence supports the following development sequence rather than a single technology-wide validation label.
Specify the task, event boundaries and quantity, including whether duration covers all attempts or only the final ascent
Validate detection and segmentation against synchronized reference annotation, including assisted, slow and unsuccessful transfers
Compare the output with the appropriate criterion, reporting bias and individual agreement limits in clinical units and accounting for repeated events within each person
Establish between-day error in Parkinson’s disease under the intended protocol, including relevant severity, medication and dyskinesia conditions
Evaluate treatment responsiveness and patient-important change with an external anchor
For prognostic claims, test added value beyond basic history, gait and balance, then evaluate calibration and performance prospectively in an independent cohort
Keep participants separate between development and evaluation; splitting one person's trials or video frames across both can exaggerate performance. The complete workflow, including failed acquisitions and assisted transfers, is the unit that eventually needs clinical validation. These are implementation recommendations derived from the evidence reviewed, not an already validated rehabtools battery.
Abbreviations
ADL, activities of daily living; AIC, Akaike information criterion; AUC, area under the receiver operating characteristic curve; BOS, base of support; CI, confidence interval; CoM, centre of mass; CoP, centre of pressure; DBS, deep brain stimulation; EMG, electromyography; FDR, false discovery rate; FTSTS or 5×STS, five times sit-to-stand; H&Y, Hoehn and Yahr; HR, hazard ratio; ICC, intraclass correlation coefficient; IMU, inertial measurement unit; iTUG, instrumented Timed Up and Go; LoA, limits of agreement; MCC, Matthews correlation coefficient; MCID, minimal clinically important difference; MDC, minimal detectable change; MIC, minimal important change; MDS, Movement Disorder Society; MoCA, Montreal Cognitive Assessment; OR, odds ratio; PD, Parkinson’s disease; PDQ, Parkinson’s Disease Questionnaire; PPMI, Parkinson’s Progression Markers Initiative; RFD, rate of force development; RGB, red green blue camera imagery; ROC, receiver operating characteristic; SD, standard deviation; SDD, smallest detectable difference; SEM, standard error of measurement; STS, sit-to-stand; TUG, Timed Up and Go; UPDRS, Unified Parkinson’s Disease Rating Scale.
References
References are numbered in first-citation order. Study-specific access descriptions identify the material examined and do not constitute a study-quality rating. Links identify original articles or the explicitly named primary-source version.
1. Duncan RP, Leddy AL, Earhart GM. Five times sit to stand test performance in Parkinson disease. Archives of Physical Medicine and Rehabilitation. 2011;92(9):1431–1436. Original article
Source note: SRC-dceba3bda0bb Duncan RP 2011
2. Iwai M, Tanabe S, Koyama S, et al. Clinical and biomechanical factors in the sit-to-stand decline in Parkinson’s disease. Movement Disorders Clinical Practice. 2025;12(10):1539–1550. Original article
Source note: SRC-d95d83471f02 Iwai M 2025
3. Magaña JC, Enríquez-Calzada S, Prat R, et al. Effects of basic and dual-task training programs on physical function in Parkinson’s disease: the PARKEX study. PLOS ONE. 2026. Original article
Source note: SRC-5706170a4fa5 Magana JC 2026
4. Paul SS, Canning CG, Sherrington C, Fung VSC. Reproducibility of measures of leg muscle power, leg muscle strength, postural sway and mobility in people with Parkinson’s disease. Gait & Posture. 2012;36(3):639–642. Original article
Source note: SRC-b3a13e915936 Paul SS 2012
5. Petersen C, Steffen T, Paly E, Dvorak L, Nelson R. Reliability and minimal detectable change for sit-to-stand tests and the Functional Gait Assessment for individuals with Parkinson disease. Journal of Geriatric Physical Therapy. 2017;40(4):223–226. Original article
Source note: SRC-5af3ae3d8721 Petersen C 2017
6. Soares CLA, Ribeiro IL, Benfica PAY, et al. Performance-based tests in individuals with Parkinson’s disease: outcome scores and reliability. Clinical Rehabilitation. 2026;40(2):238–245. Published online November 11, 2025. Original article
Source note: SRC-1ed0117ad15d Soares CLA 2026
7. Spagnuolo G, Faria CDCM, da Silva BA, Ovando AC, Gomes-Osman J, Swarowsky A. Are functional mobility tests responsive to group physical therapy intervention in individuals with Parkinson’s disease? NeuroRehabilitation. 2018;42(4):465–472. Original article
Source note: SRC-518f29365016 Spagnuolo G 2018
8. Morgan C, Masullo A, Mirmehdi M, et al. Automated Real-World Video Analysis of Sit-to-Stand Transitions Predicts Parkinson's Disease Severity. Digital Biomarkers. 2023. Original article
Source note: SRC-3ccdf823c4d8 Morgan C 2023
9. Adamowicz L, Karahanoglu FI, Cicalo C, et al. Assessment of Sit-to-Stand Transfers during Daily Life Using an Accelerometer on the Lower Back. Sensors. 2020. Original article
Source note: SRC-ab2b7e2d3a80 Adamowicz L 2020
10. Baltasar-Fernandez I, Parrino R, Strand K, et al. Differences in power and performance during sit-to-stand test and its relationships to functional measures in older adults with and without Parkinson's disease. Experimental Gerontology. 2024. Original article
Source note: SRC-f41c95cc400d Baltasar-Fernandez I 2024
11. Paul SS, Canning CG, Sherrington C, Lord SR, Close JCT, Fung VSC. Three simple clinical tests to accurately predict falls in people with Parkinson’s disease. Movement Disorders. 2013;28:655–662. Original article
Source note: SRC-507aeb58eeff Paul SS 2013
12. Mak MKY, Auyeung MM. The Mini-BESTest can predict parkinsonian recurrent fallers: a 6-month prospective study. Journal of Rehabilitation Medicine. 2013;45:565–571. Original article
Source note: SRC-09ed19718487 Mak MKY 2013
13. Moraca GAG, Orcioli-Silva D, Beretta VS, Zampier VC, Santos PCR, Gobbi LTB. Functional capacity components do not predict fall risk in people with Parkinson’s disease. Brazilian Journal of Motor Behavior. 2022;16:291–303. Original article
Source note: SRC-f7b2fac34ecf Moraca GAG 2022
14. Gray WK, Hildreth A, Bilclough JA, Wood BH, Baker K, Walker RW. Physical assessment as a predictor of mortality in people with Parkinson’s disease: a study over 7 years. Movement Disorders. 2009;24:1934–1940. Original article
Source note: SRC-613418d23812 Gray 2009
15. Da Cunha CP, Rao PT, Karthikbabu S. Clinical features contributing to the sit-to-stand transfer in people with Parkinson’s disease: a systematic review. Egyptian Journal of Neurology, Psychiatry and Neurosurgery. 2021;57:143. Original article
Source note: SRC-47605ea4edab Da Cunha CP 2021
16. Verheyden G, Kampshoff CS, Burnett ME, et al. Psychometric properties of 3 functional mobility tests for people with Parkinson disease. Physical Therapy. 2014;94(2):230–239. Original article
Source note: SRC-577143cc6034 Verheyden G 2014
17. Oregon Health & Science University. Mini-BESTest: item 1, Sit to Stand. Official assessment instructions and scoring portal. Official instructions
Source note: SRC-04164c66d0d7 Oregon Health Science Universit
18. International Parkinson and Movement Disorder Society. MDS-Unified Parkinson’s Disease Rating Scale (MDS-UPDRS): official scale, item 3.9, Arising from Chair. Updated 2019. Official scale
Source note: SRC-9f636e6ef1e1 International Parkinson and Move 2019
19. Trucco M, Ditella R, Dhimitriadhi A, et al. Instrumental assessment of sit-to-stand in Parkinson’s disease: a scoping review. Journal of NeuroEngineering and Rehabilitation. 2026. Original article
Source note: SRC-f6ef789b3469 Trucco M 2026
20. Inkster LM, Eng JJ, MacIntyre DL, Stoessl AJ. Leg muscle strength is reduced in Parkinson’s disease and relates to the ability to rise from a chair. Movement Disorders. 2003;18(2):157–162. Original article
Source note: SRC-e3a0ee4e4758 Inkster LM 2003
21. Inkster LM, Eng JJ. Postural control during a sit-to-stand task in individuals with mild Parkinson’s disease. Experimental Brain Research. 2004;154(1):33–38. Published online September 5, 2003. Original article
Source note: SRC-62320b175433 Inkster LM 2004
22. Bhatt T, Yang F, Mak MKY, Hui-Chan CWY, Pai YC. Effect of externally cued training on dynamic stability control during the sit-to-stand task in people with Parkinson disease. Physical Therapy. 2013;93(4):492–503. Original article
Source note: SRC-a4d9bc21197f Bhatt T 2013
23. Martin RA, Fulk G, Dibble L, Boolani A, Vieira ER, Canbek J. Modeling cues may reduce sway following sit-to-stand transfer for people with Parkinson’s disease. Sensors. 2023;23(10):4701. Original article
Source note: SRC-cfd9cc388be6 Martin RA 2023
24. Afshari M, Hernandez AV, Nonnekes J, Bloem BR, Goetz CG. Are virtual objective assessments of fall-risk feasible and safe for people with Parkinson’s disease? Movement Disorders Clinical Practice. 2022;9(6):799–804. Original article
Source note: SRC-d82e82e412c4 Afshari M 2022
25. Silva-Batista C, Scanlan KT, Stojak ME, et al. Remote assessment of mobility in people with Parkinson’s disease: a feasibility, safety, validity, agreement, and reliability study. Movement Disorders Clinical Practice. Published online July 22, 2026. Original article
Source note: SRC-de7f25bffeee Silva-Batista C 2026
26. Malling ASB, Morberg BM, Wermuth L, et al. Effect of transcranial pulsed electromagnetic fields (T-PEMF) on functional rate of force development and movement speed in persons with Parkinson's disease: A randomized clinical trial. PLOS ONE. 2018. Original article
Source note: SRC-8cf110c3a6f4 Malling ASB 2018
27. Mak MKY, Hui-Chan CWY. Switching of movement direction is central to parkinsonian bradykinesia in sit-to-stand. Movement Disorders. 2002. Original article
Source note: SRC-1699878af806 Mak MKY 2002
28. Mak MKY, Levin O, Mizrahi J, et al. Joint torques during sit-to-stand in healthy subjects and people with Parkinson's disease. Clinical Biomechanics. 2003. Original article
Source note: SRC-4e76cf862510 Mak MKY 2003
29. Baizabal-Carvallo JF, Alonso-Juarez M, Fekete R. The Role of Muscle Strength in the Sit-to-Stand Task in Parkinson’s Disease. Parkinson’s Disease. 2023. Original article
Source note: SRC-c74015c8a9a2 Baizabal-Carvallo JF 2023
30. Mak MKY, Yang F, Pai YC. Limb collapse, rather than instability, causes failure in sit-to-stand performance among patients with parkinson disease. Physical Therapy. 2011. Original article
Source note: SRC-86480332de05 Mak MKY 2011
31. Ramsey VK, Miszko TA, Horvat M. Muscle activation and force production in Parkinson's patients during sit to stand transfers. Clinical Biomechanics. 2004. Original article
Source note: SRC-e3b8e1fd6766 Ramsey VK 2004
32. Galna B, Barry G, Jackson D, et al. Accuracy of the Microsoft Kinect sensor for measuring movement in people with Parkinson's disease. Gait & Posture. 2014. Original article
Source note: SRC-4dbaa4485dae Galna B 2014
33. Kantha P, Charususin N, Bovonsunthonchai S, et al. Sit-to-stand strategies and anticipatory momentum transfer adjustments in individuals with Parkinson's disease using markerless motion capture: a cross-sectional study. Scientific Reports. 2026. Original article
Source note: SRC-5b218b026b8b Kantha P 2026
34. Liu M, Ye Y, Li H, et al. Automatic and explainable assessment for Parkinson’s disease by video-based human motion understanding. Journal of NeuroEngineering and Rehabilitation. 2026. Original article
Source note: SRC-dcf751959ac8 Liu M 2026
35. Pham MH, Warmerdam E, Elshehabi M, et al. Validation of a Lower Back "Wearable"-Based Sit-to-Stand and Stand-to-Sit Algorithm for Patients With Parkinson's Disease and Older Adults in a Home-Like Environment. Frontiers in Neurology. 2018. Original article
Source note: SRC-c68f86dcf41f Pham MH 2018
36. Wairagkar M, Villeneuve E, King R, et al. A novel approach for modelling and classifying sit-to-stand kinematics using inertial sensors. PLOS ONE. 2022. Original article
Source note: SRC-6965c9c931ef Wairagkar M 2022
37. Sher A, Rashid M, Lotfi A, et al. Cycle Metrics and Strategy Detection for Automated Chair Sit-to-Stand Test Analysis Employing a Single Smartphone. Annals of Biomedical Engineering. 2026. Original article
Source note: SRC-9f9af58d6d2f Sher A 2026
38. Amici C, Bussola R, Pollet J, et al. Automatic identification of postural transitions using a single inertial measurement unit and dynamic time warping: a pilot study in healthy individuals and people with Parkinson's disease. Frontiers in Bioengineering and Biotechnology. 2026. Original article
Source note: SRC-905e47a9fd3e Amici C 2026
39. Goubault É, Martin C, Duval C, et al. Enhanced Detection and Segmentation of Sit Phases in Patients with Parkinson’s Disease Using a Single SmartWatch and Random Forest Algorithms. Sensors. 2025. Original article
Source note: SRC-88857440692e Goubault E 2025
40. Bernad-Elazari H, Herman T, Mirelman A, et al. Objective characterization of daily living transitions in patients with Parkinson's disease using a single body-fixed sensor. Journal of Neurology. 2016. Original article
Source note: SRC-7a78c40f86bb Bernad-Elazari H 2016
41. Zhang J, Chung TM, Park H. Uncertainty-Enhanced Spatiotemporal Framework for Sit-to-Stand Temporal Segmentation in Parkinson's Disease. IEEE Transactions on Neural Systems and Rehabilitation Engineering. 2026. Original article
Source note: SRC-a08c947ea9c9 Zhang J 2026
42. Khalil RM, Shulman LM, Gruber-Baldini AL, et al. Machine Learning and Statistical Analyses of Sensor Data Reveal Variability Between Repeated Trials in Parkinson’s Disease Mobility Assessments. Sensors. 2024. Original article
Source note: SRC-906009eeb0c2 Khalil RM 2024
43. Marin F, Warmerdam E, Marin Z, et al. Scoring the Sit-to-Stand Performance of Parkinson's Patients with a Single Wearable Sensor. Sensors. 2022. Original article
Source note: SRC-476030f544d4 Marin F 2022
44. Serra-Añó P, Pedrero-Sánchez JF, Inglés M, et al. Assessment of Functional Activities in Individuals with Parkinson's Disease Using a Simple and Reliable Smartphone-Based Procedure. International Journal of Environmental Research and Public Health. 2020. Original article
Source note: SRC-d1338f1165f1 Serra-Ano P 2020
45. Picardi M, Redaelli V, Antoniotti P, et al. Turning and sit-to-walk measures from the instrumented Timed Up and Go test return valid and responsive measures of dynamic balance in Parkinson's disease. Clinical Biomechanics. 2020. Original article
Source note: SRC-8a3adacb2c23 Picardi M 2020
46. Kikuchi K, Oyama G, Shimoda S, et al. Dopaminergic medication alters muscle synergy during sit-to-stand motion in Parkinson's disease. Frontiers in Neurology. 2026. Original article
Source note: SRC-4e3c60b53275 Kikuchi K 2026
47. Romijnders R, Atrsaei A, Rehman RZU, et al. Association of real life postural transitions kinematics with fatigue in neurodegenerative and immune diseases. npj Digital Medicine. 2025. Original article
Source note: SRC-983d6286a803 Romijnders R 2025
48. Johansson H, Löfgren N, Porciuncula F, et al. Impact of dual-tasking and balance confidence on turns and transitions: a cross-sectional study in Parkinson's disease. Scientific Reports. 2026. Original article
Source note: SRC-ee3571047967 Johansson H 2026
49. Mak MKY, Pang MYC. Parkinsonian single fallers versus recurrent fallers: different fall characteristics and clinical features. Journal of Neurology. 2010;257:1543–1551. Original article
Source note: SRC-2f828a2a2efa Mak MKY 2010
50. Skidmore FM, Monroe WS, Hurt CP, et al. The emerging postural instability phenotype in idiopathic Parkinson disease. npj Parkinson’s Disease. 2022. Original article
Source note: SRC-2f15e595054a Skidmore FM 2022
51. Kitaichi N, Delizo K, Dei R, et al. A machine learning-based fall risk prediction model for Parkinson’s disease considering ophthalmic disorders. Parkinsonism and Related Disorders. 2026;147:108306. Original article
Source note: SRC-a12bf7e6b688 Kitaichi N 2026
52. Duncan RP, Cavanaugh JT, Earhart GM, et al. External validation of a simple clinical tool used to predict falls in people with Parkinson disease. Parkinsonism and Related Disorders. 2015. Original article
Source note: SRC-06dd612a6dd5 Duncan RP 2015
53. Marano M, Motolese F, et al. Remote smartphone gait monitoring and fall prediction in Parkinson’s disease during the COVID-19 lockdown. Neurological Sciences. 2021. Original article
Source note: SRC-ceff62415f01 Marano M 2021
54. Nouriani A, Jonason A, Sabal LT, et al. Real world validation of activity recognition algorithm and development of novel behavioral biomarkers of falls in aged control and movement disorder patients. Frontiers in Aging Neuroscience. 2023;15:1117802. Original article
Source note: SRC-c76e849aad8a Nouriani A 2023
55. Bryant MS, Kang GE, Protas EJ. Relation of chair rising ability to activities of daily living and physical activity in Parkinson’s disease. Archives of Physiotherapy. 2020. Original article
Source note: SRC-986c1a3683fb Bryant MS 2020
56. Macleod AD, Dalen I, Tysnes OB, Larsen JP, Counsell CE. Development and validation of prognostic survival models in newly diagnosed Parkinson’s disease. Movement Disorders. 2018;33:108–116. Published online 2017. Original article
Source note: SRC-8003da04a2bd Macleod AD 2018
57. Dewey DC, Chitnis S, McCreary MC, et al. APDM gait and balance measures fail to predict symptom progression rate in Parkinson’s disease. Frontiers in Neurology. 2022;13:1041014. Original article
Source note: SRC-817facd83053 Dewey DC 2022
58. Schini M, Bhatia P, Shreef H, et al. Increased fracture risk in Parkinson’s disease: an exploration of mechanisms and consequences for fracture prediction with FRAX. Bone. 2023;168:116651. Original article
Source note: SRC-069e9a832384 Schini M 2023
59. Di Lazzaro G, Ricci M, Saggio G, et al. Technology-based therapy-response and prognostic biomarkers in a prospective study of a de novo Parkinson’s disease cohort. npj Parkinson’s Disease. 2021. Original article
Source note: SRC-5c8c8a9349c4 Di Lazzaro G 2021
60. Rahimi F, Roberts AC, Jog M. Patterns and predictors of freezing of gait improvement following rasagiline therapy: a pilot study. Clinical Neurology and Neurosurgery. 2016;150:117–124. Original article
Source note: SRC-b5fd6ac3749d Rahimi F 2016