In this report
Audited and Updated
Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.
Editorial record
- Edited phrase under AUD-KOA-STS-C02 . Original wording
- Edited phrase under AUD-KOA-STS-P09 . Original wording
- Edited phrase under AUD-KOA-STS-C01 . Original wording
- Edited phrase under AUD-KOA-STS-C04 . Original wording
- Edited phrase under AUD-KOA-STS-D01 . Original wording
- Edited phrase under AUD-KOA-STS-D03 . Original wording
- Edited phrase under AUD-KOA-STS-D02 . Original wording
- Edited phrase under AUD-KOA-STS-P03 . Original wording
- Edited phrase under AUD-KOA-STS-P04 . Original wording
- Edited phrase under AUD-KOA-STS-C03 . Original wording
Executive assessment
Sit-to-stand (STS) testing is a useful, inexpensive way to observe knee-OA functional capacity. It can establish whether someone can rise under stated conditions, how quickly or repeatedly they do so, and, with appropriate instrumentation, how the task is distributed between limbs and joints. These are different constructs. A repetition count is not a direct measurement of quadriceps strength, a timing-derived power estimate is not measured joint power, and an attractive current-function classifier is not a validated forecast of future disability.
The 30-second chair stand is part of the OARSI minimal performance-test set. The five-rise test and single-rise biomechanics answer related questions but are not interchangeable versions of one assay. Chair height, seat depth, arm use, foot position, start/stop definitions, familiarization, pain and trial aggregation change performance. The best practical strategy is to retain a standardized count or time together with completion, assistance, symptoms and movement observations. [1–5]
Direct knee-OA reliability studies show useful relative reproducibility, but clinically relevant error and practice effects remain. Gill and colleagues found approximately 0.5–0.8 extra stands on the second attempt and within-session MDC90 values of 2.27–2.79 stands in early radiographic OA. Those values came from repeated attempts within a session at two separate study occasions; they are not six-month test–retest error. In a community knee-OA sample, one-week MDC90 was 3.2 stands and the retest mean improved by about one stand. In Tolk's preoperative cohort, same-day SDC95 was 2.4 stands. Confidence level, interval and severity matter. [3–5]
Clinical interpretability is less secure than popularity implies. Several studies found weak associations between performance and patient-reported function, and the 30-second chair stand did not meet prespecified construct/responsiveness criteria in some cohorts. Another physiotherapy study reports a positive MIC signal, but its original full text was unavailable here. A mean treatment improvement, a high ICC, an MDC and an anchor-based MIC must not be used as synonyms. [3, 4, 6–8]
Remote and video methods are promising when their workflow is made explicit. Excellent rescoring of the same recording is a different achievement from repeatability of new patient performances. A real home-sensor study found weaker home/laboratory agreement than within-home repeatability. A digital self-count study included hip as well as knee OA and showed systematic upward differences despite reassuring language about agreement. Current computer-vision work often establishes technical feasibility or concurrent classification, not a validated individual score, power measurement or prognosis. [9–14]
Prospective STS evidence must be carefully delimited. At-risk cohorts without baseline radiographic OA cannot be relabelled as prognosis in established disease. A newer OAI mediation analysis did not find a significant STS-to-later-physical-activity pathway and contains important outcome-scale reporting concerns. In Hu's longitudinal RFD study, chair-stand deterioration was an outcome, not the predictor, and its deterioration threshold was an MDC-based approximation rather than a validated chair-stand MIC. A large frequent-knee-pain cohort also found no independent seven-year chair-time association with later knee replacement. There is no defensible basis in the intensively appraised evidence for a universal knee-OA STS-only falls, progression or surgery calculator. [15–19]
Current source qualification · AUD-KOA-STS-P05: Keep source concern; obtain definitions/data clarification before using numeric prognosis. No guessed corrections.
For rehabtools, prioritize transparent measurement and repeatable task definitions. Video can add useful phase and strategy information, but automated force, joint-power, RFD, diagnostic and prognostic claims need their own validation. A particularly important implementation warning is that the original displayed equation in a knee-OA STS-power paper omits gravitational acceleration despite defining mass in kilograms and output in watts; this should not be copied without clarification. It does not prove the authors' actual calculations were wrong. [20]
Current source qualification · AUD-KOA-STS-P02: Quarantine equation, not every numerical result. Exact rendered omission remains unverified.
Current verification boundary: Text-rendered equation is mass×.9×(.5 height−chairheight)/[(30/repetitions)×.5], with kg/metres/seconds and labelled watts; no g appears. Dimensions of literal expression are kg·m/s rather than watts, assuming stated dimensionless fractions. Rendered equation unavailable, so cannot exclude text-extraction loss or certify historical PDF visual inspection. Actual calculation code not recovered. sts-B011-S03
Scope and approach
The primary population is adults with knee OA before arthroplasty. The report covers 30-second chair stands, five-rise tests, single-rise task analysis, force-platform and motion-capture measures, estimated STS power, remote testing and computer vision. It separates measurement of present capacity, movement strategy, error, meaningful change and later outcomes. Treatment studies are used when they contribute measurement or change evidence, not as the main treatment-effectiveness review.
Searches completed on 2 October 2026 combined knee OA with sit-to-stand, chair-stand and chair-rise terms. The dated PubMed query yielded 605 records, verified against the official NCBI records; a focused Scopus query yielded 426 unique records. A broader knee-OA review landscape, targeted original-source searches and bounded forward and backward citation chasing supplemented these sets. Counts describe retrieved records, not eligible or fully appraised studies.
What the chair task measures
Transfer ability repeated capacity and strategy
A single successful rise answers whether someone can transfer under those conditions. A five-rise time adds speed and repetition. A 30-second count measures repeated capacity over a fixed interval. Phase times, trunk motion, limb loading and joint kinetics describe the strategy used. A person can become faster by changing strategy without increasing the strength or power of the more painful knee.
The lower limbs, trunk and arms can compensate for one another. Pain, balance, confidence, joint range, body size and cardiopulmonary demand all influence repeated rising. The knee-OA literature does not support interpreting a low count as a uniquely identified quadriceps deficit. Relationships with dynamometry are useful construct evidence, not proof that the two measurements are interchangeable. In the community study, 30-second count correlated only moderately with body-mass-normalized aggregate knee-extensor torque. [3, 21, 22]
The output unit should therefore remain visible. Seconds, repetitions, estimated watts, measured vertical force, joint moment, joint power and rate of force development are different quantities. Even a force-platform STS test must specify whether power refers to vertical whole-body centre-of-mass motion or a particular joint. A knee moment requires an appropriate biomechanical model; joint power additionally requires angular velocity. Neither can be directly inferred from repetitions alone.
Table 1 Match the question to the STS output
These are related constructs, not interchangeable assay labels.
| Question | Useful output | Does not directly establish |
|---|---|---|
| Can the person transfer? | Completion, supports and assistance | Unsupported repetition capacity |
| How quickly or repeatedly? | Five-rise time or 30-second count | Isolated quadriceps strength |
| How is the task achieved? | Phases, trunk motion and limb loading | A universally desirable movement strategy |
| How much mechanical power? | Defined reference-measured or equation-estimated output | Interchangeability of joint and whole-body power |
| What happens later? | Validated outcome-specific prognostic model | Prognosis from a concurrent association |
Knee OA is heterogeneous
Record the diagnosis definition, radiographic grade, symptoms and affected side. Symptomatic KL 0–2 cohorts include people without definite radiographic OA. Preoperative KL 3–4 samples selected for knee replacement represent a different spectrum. A history of ACL injury, patellofemoral involvement, varus/valgus alignment and bilateral disease may change the preferred transfer strategy.
OA should not be divided into acute and chronic solely by borrowing another disease's staging framework. Instead describe longstanding disease, symptom duration, a recent flare, current pain/effusion, medication timing and recent treatment. A pain-limited session can differ from a quieter day without evidence of structural change. Conversely, stable timing does not exclude persistent pain or an altered loading strategy.
The less painful leg is not necessarily normal. Bilateral OA is common in several appraised cohorts, and person-level task performance may remain acceptable because the trunk or hips compensate. A symmetry ratio near one can arise when both limbs are impaired. Age, obesity, height, other painful joints and use of supports modify both the task demand and the relevance of published reference values. [3, 9–11, 21, 22]
Protocol determines the assay
Chair geometry and feet
The original OARSI recommendation uses an approximately 43-cm chair for the 30-second test and requires chair height and adaptations to be reported. The same absolute height is relatively lower for a taller person, changing knee and hip angles and vertical displacement. A chair adjusted to produce approximately 90° knee flexion answers a different standardized question from a universal 43-cm chair. Neither should silently replace the other during follow-up. [1]
Published protocols span 43 cm, 44 cm, 45 cm, 46 cm and height-adjusted seats. Gill's appendix specifies a 17-inch chair, foot placement chosen for comfort, arms crossed and a defined seating region. Jørgensen used a 44-cm chair; Dahlberg's digital protocol used 45 cm; Segal's motion study used a 46-cm chair; Sagawa adjusted the chair for 90° knees. This variation helps explain why a single universal time, count or power threshold is implausible. [5, 11, 20–22]
Seat firmness, backrest, depth of sitting, footwear and foot position should also be fixed or recorded. Pulling the feet farther back changes the starting mechanical demand. In Segal's study, apparent knee-range differences between faster and slower women could reflect their self-selected foot placement. Those differences should not automatically be labelled a deficit to correct. [21]
Arms and support
Arms crossed over the chest, arms at the side, hands on thighs and pushing on armrests are different conditions. Record support precisely. If a person requires arms or assistance, that is clinically important transfer information rather than an inconvenient failure to erase. A modified supported test may be useful, but its value should not be compared unqualified with an unsupported normative or change threshold.
Do not force unsupported testing when it is unsafe. Record whether the task was completed, the support used, the reason for stopping, symptoms and the achieved partial performance. A zero count, inability to begin, pain-limited stopping and a technical video failure should remain distinct records. Their prognostic meaning cannot be assumed to be identical.
Timing repetitions and familiarization
Specify whether timing begins on the command, first trunk movement or seat-off, and whether it ends at the fifth full stand or after the fifth return to sitting. Instrumented phase onset may use trunk angular velocity, marker displacement or force thresholds. These definitions are not interchangeable with a stopwatch total.
For the 30-second test, define a complete stand, return-to-seat requirement, handling of a partly completed final stand, and invalid repetitions. Gill's appendix counts a rise that is more than halfway complete at 30 seconds and includes one practice stand. Its study then evaluated three full test attempts. Aily's video protocol requires full hip/knee extension and seat contact. An app must reproduce its selected protocol rather than combine convenient elements from different papers. [5, 9]
Practice is not cost-free in painful OA. More trials may reduce unfamiliarity but also provoke pain or fatigue and selectively exclude less tolerant participants. Retain rest duration and whether symptoms returned to baseline. The best, mean and first trial have different interpretations; a reliability threshold for one aggregation should not be attached to another.
Table 2 Protocol differences that matter
Record actual geometry, assistance, start/stop, final-repetition and aggregation rules.
| Source/task | Chair/support | Trials or event rules | Interpretation limit |
|---|---|---|---|
| OARSI [1] | Approximately 43 cm; record adaptations | 30-second repetition count | Consensus protocol, not threshold validation |
| Gill [5] | 17-inch chair; arms crossed; comfortable feet | Practice rise; three attempts; >halfway final rise counts | Within-session error; repeated full tests may provoke pain |
| Suwit [3] | 43 cm; trained assessor | One-week repeat, same assessor | Retest mean about one stand higher |
| Segal [21] | 46 cm; self-selected feet; no arms | Separate self-paced instrumented rises | Different from timed five-rise stratification |
| Sagawa [22] | Adjusted for 90° knees; arms by side | Four self-paced trials; trunk-based event definitions | Not a fixed-height timed clinical test |
| Jørgensen [20] | 44 cm; arms crossed | 3–5 practice rises; one 30-second trial | Anthropometric power estimate, not force-platform power |
| Torres [14] | 18 inches; hands on chest | Five self-paced rises; last excluded | Reference biomechanics, not a stopwatch 5STS protocol |
Reproducibility and absolute error
Direct knee OA evidence
Gill and colleagues enrolled 93 people with KL-2 OA and moderate-to-severe unilateral knee pain, mean age 61.3 years and BMI 33.5 kg/m². Reliability was evaluated across three attempts separated by approximately 1–2 minutes within each session. This was done at baseline and again six months later, with 68 attending the latter occasion. The analysis was within each session, not between baseline and six months. [5]
At baseline, 83 completed all three attempts; at six months, 61 did. The second attempt improved by a mean 0.52 and 0.79 stands, respectively. ICCs were 0.92 and 0.94; SEMs 0.97 and 1.20 stands; MDC90 for an individual 2.27 and 2.79 stands. The much smaller group-level MDCs, 0.25 and 0.36, reflect averaging over a sample and must not be used for individual patients. By the third attempt, about one in ten participants did not complete the test, principally because of knee pain. This is a meaningful feasibility finding despite not crossing the authors' 15% floor-effect criterion. [5]
The paper's nominal "early-stage" label refers to radiographic severity, not short symptom duration: mean pain duration was 3.2 years. The retest estimates apply to participants tolerating all three repetitions of the entire test. Removing those whose pain prevents further testing can make reproducibility look more favourable than the full clinical workflow. The practical response is to standardize familiarization while minimizing unnecessary repeated provocation, and to report symptom-limited noncompletion. [5]
Suwit and colleagues provide one-week knee-specific evidence in 55 community participants, many with bilateral disease. The 43-cm 30-second test used the same trained assessor at both visits, blinded to earlier scores. Mean count increased from 14.6 to 15.6; ICC was 0.87, SEM 1.4 stands and MDC90 3.2 stands. The baseline-minus-retest limits of agreement were approximately −4.7 to +2.9 stands around a negative mean difference of about −0.9. These data retain both learning/systematic difference and random variation; an ICC alone would obscure the clinical issue. [3]
Tolk's preoperative study retested 30 patients after 30 minutes. The 30-second chair ICC was 0.90, SEM 0.85 and Table-2 SDC95 2.4 stands; the discussion gives 2.5. The mean rose from 9.0 to 9.8 stands. This same-day error estimate should not be substituted for a week-to-week estimate in conservatively managed OA. The baseline construct and postoperative responsiveness findings also need to be distinguished from reliability. [4]
Dobson's study involved hip and/or knee OA and excluded previous replacement. It carefully distinguished independent-rater repeated performances within a session from same-rater testing about a week later. The 30-second chair stand achieved sufficient rather than uniformly optimal reliability. Numeric original tables were not preserved in the retrieved body, so exact error values are not reproduced. Its mixed-joint results are supportive context, not a knee-only threshold. [2]
Five rise evidence and important remaining originals
Khuna and colleagues directly compared 30-second and five-rise tests in 60 older people with knee OA, with 30 retested after a week. The abstract reports near-perfect intra/interrater reliability but lower between-occasion ICCs, approximately 0.84–0.85, and explicitly describes substantial SEM/MDC. The original full methods and numeric tables were not obtained here. It would be inappropriate to treat the high scorer ICC as proof of negligible between-day error or invent exact MDCs from the abstract. [23]
Holm and colleagues' 40-person study is another relevant original acquisition gap. Its abstract reports a three-day retest interval, high ICCs, chair-count limits of agreement of approximately ±2.4 repetitions and improvement on retest. This supports the importance of familiarization and individual agreement, while the complete protocol, bias and subgroup context still need full-text confirmation before clinical implementation. [24]
What to do with error estimates
Preserve confidence level, interval, sample, protocol and trial rule beside every threshold. Three extra repetitions might exceed one study's MDC90 but not another study's error estimate. That does not establish that one paper is wrong; their designs may sample different error components. Equally, rounding a fractional MDC to a whole-repetition action rule is an implementation decision and should be documented, not presented as an independently validated threshold.
An individual change beyond an appropriate MDC is evidence against measurement noise under the study's assumptions. It is not automatically clinically important. A change below MDC can still coexist with meaningful pain relief, greater confidence, safer technique or a personally important activity gain. Conversely, a faster test with more trunk compensation can be a real capacity improvement without demonstrating restoration of knee loading.
Table 3 Selected reliability and change estimates
Population, interval, confidence level, score unit and protocol remain attached to each estimate.
| Source | Context | Estimate | Do not confuse with |
|---|---|---|---|
| Gill [5] | KL2; within-session repeated attempts | MDC90 2.27/2.79 stands | Six-month between-session error or group MDC |
| Suwit [3] | 55 community knee OA; one week | MDC90 3.2 stands; LoA −4.7 to +2.9 | Universal patient-important change |
| Tolk [4] | 30 pre-TKA; 30-minute retest | SDC95 2.4 stands | Discussion value of 2.5 or nonoperative responsiveness |
| Mostafaee [6] | 60; four-week physiotherapy; abstract only | Reported MIC 2.5 stands | An SDC or a fully appraised universal threshold |
| Aily [9] | Same-recording rescoring | ICC 1.00; zero displayed scorer MDC | Zero patient test–retest error |
| Rose [10] | Average sensor-detected rise duration | MDC95 approximately 0.25 seconds for the mean of two trials | Total five-rise time |
| Dahlberg [11] | Mixed hip/knee; digital self-report | MDC95 4.4/4.8 repetitions | Knee-only or observed-count thresholds |
Validity responsiveness and meaningful change
Correlation with self report is informative but not a gold standard
The community study's chair count was moderately associated with aggregate normalized knee-extensor torque, but weakly with KOOS physical function. Tolk likewise found good reproducibility but confirmed only 42% of prespecified chair construct-validity hypotheses. Its chair changes across TKA met only 50% of responsiveness hypotheses despite a statistically significant mean count improvement. This is a clear example of why mean change and valid responsiveness are not identical. [3, 4]
It would also be an overreaction to infer that chair testing has no clinical value. Self-report describes experienced difficulty over a period and across activities; a timed task samples observed execution under controlled conditions. The tests can be complementary. However, the construct must be explicit: a count should not be called a complete measure of knee-OA function or quality of movement. Add task-specific symptoms and observed strategy rather than using a low correlation to dismiss patient experience. [1, 3, 4]
Ramalho's 107-person study is a more recent warning against assuming the consensus core set has uniform measurement properties. Its abstract reports acceptable walk/stair validity and responsiveness, but not the chair stand, over six months. Because the original full text was unavailable, its exact hypotheses and missing-data handling are not re-appraised here. It remains direct evidence to retrieve before selecting a chair score as the sole intervention endpoint. [7]
MIC evidence must retain its treatment context
Mostafaee's abstract describes 60 knee-OA patients undergoing four weeks of physiotherapy, with a seven-point global-rating anchor. It reports chair-stand MIC 2.5 repetitions, AUCs above 0.70 for all three core tests and anchor correlations across tests of 0.43–0.63. These are relevant positive findings, explicitly at abstract level in this appraisal. The full anchor wording, stable/improved denominators, ROC precision, protocol and relationship between MIC and measurement error remain unverified. The 2.5 value should not be merged with Tolk's similarly sized SDC or presented as a universal knee-OA response threshold. [6]
Different anchors, treatment duration, starting capacity and direction of change can yield different MICs. A threshold for improvement does not automatically apply to worsening. Distribution statistics such as a fraction of a standard deviation or an MDC are not patient-importance anchors. When a study substitutes one because a validated MIC is unavailable, that substitution must remain explicit. [6, 18]
Classification thresholds do not establish prognosis
Lee and colleagues' 30-versus-30 case-control study included symptomatic KL 0–2 participants, predominantly women and all independently walking without aids. Chair count differed between groups, but the 30-second test had AUC 0.76 and a data-derived 9.8-count cut-point with sensitivity 0.50 and specificity 1.00. That result is about discrimination from intentionally asymptomatic controls, not prospective falls, future progression, or diagnosis among people with competing knee conditions. A numerical cut-point does not become a clinically actionable risk boundary merely because it is easy to enter into software. [8]
Movement strategy asymmetry and instrumentation
A normal time can conceal compensation
Segal and colleagues recruited 60 people with symptomatic radiographic knee OA from MOST; usable STS motion data were available for 49. They grouped participants by five-rise performance but analysed separate self-paced single rises from a 46-cm chair. Men and women used different associations between joint motion and task performance. Strength measures did not neatly explain all performance strata. Foot placement was self-selected, and kinetic data were available only from the more symptomatic limb because a single force plate was used. This cannot establish bilateral loading symmetry. [21]
The distinction between test and instrumented task is important. A timed five-rise test was used, whereas the motion-capture task was at preferred speed and averaged over separate rises. Relationships between the two cannot justify treating a single-rise phase time as a substitute for five-rise total time. The study is cross-sectional; it does not show that changing a particular strategy causes later mobility improvement. [21]
Current verification boundary: Timed five-rise no-arm test creates tertiles. Separate 46-cm-chair motion trials at preferred speed, mean of five trials. Self-selected foot position/seat depth; no arms; one force plate gives kinetics for more symptomatic limb only. Motion marker 60 Hz, force 300 Hz. Source does not spell out the exact rapid verbal instruction in the timed-test Methods paragraph; contrast with self-paced motion task is clear. Separate self-selected speed/mean five trials explicit. Timed clinical five-rise contrast supported; exact fast verbal instruction not reproduced in article Methods. sts-B076-S02
Sagawa and colleagues examined 101 severe knee-OA patients awaiting unilateral TKA and 27 controls using a height-adjusted chair, synchronized cameras and two force plates. An exploratory multivariate analysis described three strategy groups. A compensated group had similar task times to controls but greater trunk flexion and lateral lean. The slowest group showed substantially longer task and suspension times. This supports the clinical observation that speed alone can miss altered movement. It does not validate three universal phenotypes, a treatment-selection rule or a prognosis. The grouping was derived and compared within the same dataset. [22]
Early disease and bilateral loading
Pan and colleagues studied 24 people with unilateral KL 1–2 OA and 12 controls. They combined repeated chair stands, three-dimensional motion capture, two force plates and EMG. The OA group showed altered task timing, range, angular velocity and loading, including lower affected-side contribution. However, the sample was restricted to the dominant side being affected and BMI below 28; groups still differed in BMI. Numerous comparisons and some unusual displayed inferential values warrant caution. These findings help generate hypotheses about early compensation, not numerical targets for correcting every patient. [25]
Report the event and quantity behind an asymmetry measure: peak vertical force, impulse during ascent, mean support, external knee moment or joint power. State whether the formula is affected/unaffected, absolute difference, percentage difference or a normalized symmetry index. Denominators near zero can make ratios unstable. A bilateral total-force trace cannot identify each limb's contribution; a single side-view camera cannot directly measure ground-reaction force distribution.
Aiming for symmetry without considering pain, capacity and bilateral disease can be misleading. More equal loading is not necessarily higher total capacity, and a deliberate offloading strategy can be useful for task completion. If rehabilitation changes the strategy, evaluate the functional result and symptoms alongside the mechanical measure rather than treating visual symmetry as a standalone outcome.
Estimated STS power is not measured knee power or RFD
Jørgensen and colleagues' study, published in 2024 with a 2023 DOI, analysed pre-intervention data from 86 patients awaiting TKA. The 30-second test used a 44-cm chair, arms crossed, 3–5 familiarization repetitions and one scored trial. Estimated STS power was calculated from body size, chair height and repetition rate; knee-extensor MVC was measured separately with stabilized handheld dynamometry. The index knee was preoperative, but Table 1 records 16/86 participants with a contralateral TKA and five with a contralateral THA. Thus this was not a replacement-naïve whole-person cohort. Despite "predicts" in the title, the analysis was cross-sectional. [20]
Estimated STS power was related to concurrent functional performance in both sexes, with some patient-reported associations in men. These findings do not establish prospective prediction, sensitivity to rehabilitation or causal superiority over strengthening. The estimated power contains repetition rate by construction, so relationships with other rapid mobility tasks should not be mistaken for entirely independent physiological information. Sex-stratified models also had limited adjustment; pain and other factors could remain confounders. Reported age-adjusted R² describes the whole model, not incremental predictive value attributable to power. The regression tables express the 40-m walk in seconds despite conflicting unit wording elsewhere, so negative coefficients indicate shorter completion time. Cross-sectional estimates of a "required power gain" should not be presented as expected within-person treatment responses. [20]
The assay assumptions deserve equal attention. Timing-derived power estimates typically assume a fraction of body mass displaced vertically, approximate leg length from stature, infer displacement from chair height, and assume a concentric fraction of the repetition duration. Knee OA can alter ascent/descent timing, trunk motion and loading distribution. An equation calibrated in another population may therefore have person-specific error even when the chair count itself is accurate.
The original displayed equation in this paper omits gravitational acceleration while stating mass in kilograms, length in metres and power in watts. That expression is dimensionally inconsistent as printed. Independent inspection of the original publisher PDF confirmed that the omission is not merely text-extraction loss. It remains unknown whether the underlying calculations included the required conversion. Therefore, do not copy the displayed equation into rehabtools or infer that all numerical results are invalid without clarification. Keep this source issue distinct from the more general limitations of estimated power. [20]
Current verification boundary: Text-rendered equation is mass×.9×(.5 height−chairheight)/[(30/repetitions)×.5], with kg/metres/seconds and labelled watts; no g appears. Dimensions of literal expression are kg·m/s rather than watts, assuming stated dimensionless fractions. Rendered equation unavailable, so cannot exclude text-extraction loss or certify historical PDF visual inspection. Actual calculation code not recovered. sts-B086-S01
For a valid mechanical power calculation, mass, acceleration, displacement and time must have consistent units, and the particular model must be declared. A software implementation should have dimensional checks and numerical test cases. Estimated whole-body STS power is not isolated quadriceps power, measured knee-joint power or early-phase RFD. The latter requires a force/torque-time assay and an explicit onset/time-window definition, reviewed separately in the strength/power report.
Table 4 Power force and rate are different quantities
The original Jørgensen [20] displayed equation omits g despite kg/W definitions. Do not copy it or assume actual calculation error without clarification.
Current verification boundary: Text-rendered equation is mass×.9×(.5 height−chairheight)/[(30/repetitions)×.5], with kg/metres/seconds and labelled watts; no g appears. Dimensions of literal expression are kg·m/s rather than watts, assuming stated dimensionless fractions. Rendered equation unavailable, so cannot exclude text-extraction loss or certify historical PDF visual inspection. Actual calculation code not recovered. sts-B089-S01
| Output | What is required | Main limitation |
|---|---|---|
| Count/time | Defined repetitions and timing | Composite task capacity |
| Estimated whole-body STS power | Declared anthropometric/time equation and consistent units | Assumptions about displacement, mass fraction and concentric duration |
| Measured vertical power | Ground-reaction force plus a validated velocity/displacement method | Not isolated knee-joint power |
| Joint power | Joint moment and angular velocity from a biomechanical model | Sensitive to model, force inputs and event definitions |
| RFD/RTD | Force/torque-time signal, onset and time window | Not angular acceleration or repetition rate |
Remote assessment and digital measurement
Supported video assessment
Aily and colleagues tested 32 people with KL 2–3 OA using face-to-face and video-guided tasks on the same day in a prepared university building. The remote examiner was in another room, distances were premarked, equipment supplied and assistance nearby. Participants had previous test experience. Original supplementary criteria excluded BMI at least 30 and previous knee surgery. These details restrict extrapolation to an unprepared home or a person learning the test for the first time. [9]
For the 30-second chair stand, the mean face-to-face/video difference was −0.22 repetitions. The original Bland–Altman figure showed individual limits approximately −2.32 to +1.88 repetitions. Rescoring the same recording six weeks later produced ICC 1.00 and zero displayed scorer MDC. Those are not contradictory claims: a person can repeat the task differently while an observer counts the same recording identically. It would be incorrect to use the zero scorer MDC as evidence that any one-repetition patient change is real. [9]
The paper also contains inconsistent numeric error values for other tasks, reinforcing the need to check tables and figures rather than reproduce a general conclusion that all remote measures are validated. Technical issues occurred in 11 of 32 sessions. Recording can support later quality review, but privacy, informed use and storage must be planned separately. Feasibility in a supported research workflow is not proof of safe unsupervised testing in every home. [9]
Wearable chair rise measurements at home
Rose and colleagues supplied an armless chair, three Opal sensors, a tablet and researcher video guidance to 20 knee-OA participants. Participants completed five rises rapidly with arms crossed; the device outcome was average duration of detected chair-rise events, approximately one second, rather than total stopwatch time for five repetitions. Within-home redon/retest of the mean of two trials had ICC 0.89 and MDC95 about 0.25 seconds for average detected rise duration. These values must not be attached to five-rise total time. [10]
Home/laboratory chair-rise agreement was weaker, ICC 0.66. Eighteen had home chair data; sixteen had usable laboratory chair data, with failures related to inability, detection and signal quality. The protocol required manual correction of sensor-clock drift, and equipment setup was actively supported. Thus the result supports a carefully specified remote workflow and highlights practical failure modes. It does not establish passive detection of every everyday transfer or an autonomous falls-risk signal. [10]
Digital self counting
Dahlberg and colleagues' 2025 study examined two separate samples already familiar with a digital OA programme. The face-to-face comparison had 18 participants, 14 with knee and four with hip OA. Of 54 initially included home participants, 44 completed both tests (32 knee and 12 hip OA); ten did not complete retesting despite reminders. A 45-cm chair, app instructions and an on-screen countdown without a final alarm were used. These are mixed-joint and digitally experienced samples. [11]
Agreement ICCs were 0.87 for self versus therapist assessment and 0.88 for home retest, but individual MDC95 estimates were 4.4 and 4.8 repetitions. Participants reported 1.5 more repetitions on self-assessment and 1.2 more at the second home test; both confidence intervals excluded zero. The original text also describes no systematic bias. That reassuring wording should not erase the reported positive mean differences. The result argues for keeping self-count and observed-count series distinct, and for explicit attention to learning, counting and timing. [11]
Computer vision requires criterion validation
Zhao and colleagues analysed independently recorded five-rise videos from 104 advanced knee-OA participants using AlphaPose and VideoPose. Their work explored group differences in estimated angle time series according to WOMAC-defined stiffness/function categories. It did not include a simultaneous reference-system validation of the reconstructed joint angles. No validated individual diagnostic threshold, SEM/MDC or prospective outcome model was established. Allowing arbitrary recording angles improves convenience, but does not by itself prove that the estimated three-dimensional angles are view-invariant or sufficiently accurate. [12]
Selection and processing matter: incomplete or poorly framed videos were excluded; frequent reliance on walking aids and specified comorbidities were exclusion criteria; time series were standardized to a common length before wavelet comparison. Resampling facilitates pattern comparison but can remove or change timing information. Differences between average group curves do not demonstrate that an individual score is accurate or that a change in the score reflects clinical improvement. This is exploratory movement analysis, not a validated self-diagnostic tool. [12]
Newer smartphone and multicamera evidence
Chan and colleagues' 2026 STS-Dynamics work is more extensive, with a stated 309 participants and patient-level train/test allocation, using seven self-paced chair-rise recordings and side-view smartphone pose tracking. Separating participants before splitting repeated videos is an important safeguard against leakage. This was an internally held-out sample rather than external validation. The model discriminated symptomatic OA from non-OA with a reported AUC around 0.776. This is current classification, not a future-outcome prediction, despite a probability-like output. [13]
Several limitations prevent a stronger clinical claim. The article contains inconsistent counts of sex, control participants and retained training/test videos across sections. Some performance estimates differ between text and figure captions. Dynamic time warping was used to align video and reference motion for technical comparison, which means timing accuracy before warping is a separate question. Only the right side was filmed; eligibility and disease-side descriptions are not wholly consistent. The bootstrap sampling unit is not clear in the main methods; uncertainty for a repeated-video model should ideally account for participant clustering. These issues require original-data and supplement reconciliation before deploying an individual classifier or monitoring index. [13]
Torres and colleagues' 2026 study compared markerless and marker-based capture during walking and STS, with synchronized force plates. Twenty-one participants contributed STS data. Sagittal knee-ROM agreement was moderate, ICC 0.66, with a mean markerless difference of 4.2° and individual limits approximately −4.6 to +13.0°. Frontal knee-ROM agreement was poor, ICC 0.03. Peak KFM agreed better, but both kinetic pipelines used measured ground-reaction force. This cannot validate camera-only force or power. The study and Rose's home work share a trial programme and should not automatically be counted as fully independent replications. [10, 14]
For a product, the appropriate conclusion is that technical performance is parameter- and setup-specific. A model can yield a plausible animation while the frontal-plane metric is unreliable. High agreement after temporal warping does not establish accurate event timing. A classifier AUC does not establish calibration, clinical utility, repeatability or meaningful change.
Table 5 What the digital evidence supports
Agreement, repeatability, classification, meaningful change and prognosis need separate evidence.
| Study | Useful finding | Main boundary |
|---|---|---|
| Aily [9] | Supported video-guided testing and reproducible counting | Prepared university environment; BMI<30; same-video reliability |
| Rose [10] | Supported home sensor workflow | Detection failures, manual corrections, weaker home/lab agreement |
| Dahlberg [11] | Digital self-counting feasible in experienced users | Mixed joints; positive mean differences; individual MDC about 4–5 reps |
| Zhao [12] | Exploratory group movement patterns | No reference-system angle validation or individual threshold |
| Chan [13] | Patient-level split; promising current-status classifier | Source inconsistencies; internal classification, not prognosis |
| Torres [14] | Parameter-specific markerless/reference comparison | Frontal ROM poor; kinetics require force plates |
Prospective outcomes what STS does and does not forecast
Disease onset is not progression in established OA
Hiyama's 2026 OAI study is prospectively relevant but outside the main diagnosed-OA population: the abstract describes 1,003 people and 2,006 knees with KL 0 at baseline. Slower five-rise performance was associated with a symptom-defined onset outcome, but not with incident structural OA. Symptomatic onset was defined using a WOMAC pain threshold and increase, not necessarily the combination of pain and newly established radiographic disease. The original full text was unavailable, so detailed adjustment and within-person knee handling remain unverified. It should not be presented as proof that a slow chair test predicts progression in a person who already has knee OA. [15]
Thorstensson's earlier cohort involved 148 younger adults with chronic knee pain and a one-leg-rise test, not the bilateral five-rise or 30-second chair test. Fewer one-leg rises were associated with incident radiographic OA in unadjusted and separately adjusted models. The simultaneous age, sex, BMI and pain model was not statistically significant: OR 2.67 (95% CI 0.92–7.76), n=70. No significant progression predictor was identified in the prevalent-OA group. Baseline prevalent OA was defined as KL at least 1, another important difference from common KL-at-least-2 definitions. This is adjacent prevention evidence; the current audit recovered the full original article and its tables. [16]
Neither study licenses a universal repetition-count risk cut-point for established knee OA. The task, population, disease definition and outcome must all match before prognosis is transported.
Later activity and quality of life
Dean and colleagues' 2026 prospective mediation analysis used 782 complete OAI records from year-six and year-eight visits. The retained sample included progression, incidence and control-cohort members at original OAI enrollment, rather than a confirmed established-OA-only sample. Knee surgery between visits was excluded. The model examined baseline pain and STS together with later activity and health-related quality of life, adjusting for age, sex and race; it did not establish control of baseline activity or baseline quality of life in that model. [17]
The reported STS-to-later-moderate/vigorous-activity pathway was not significant, β=−0.02, P=0.46. An overall indirect effect in the model does not establish mediation specifically through STS. Nor does the design establish that improving chair performance will improve activity or quality of life. Baseline pain and baseline STS are contemporaneous, and activity and quality of life are measured at the later occasion; temporal ordering and residual confounding remain limitations. [17]
The original paper also describes a quality-of-life range of 0–100 but reports a mean of 103.5, and its pain-variable definition and summary statistics are not readily reconciled. These are source-level reporting concerns, not corrections that can be guessed from the narrative. The study should not support a numerical prognosis or a claim that the STS measure is a validated screening tool for later quality of life. Its own discussion calls for prospective screening evaluation. [17]
Later knee replacement
Skou and colleagues directly tested chair performance as a predictor of later knee replacement in MOST participants with frequent knee pain, including people with and without radiographic OA. Baseline chair data were available for 1,173 people and 1,564 knees. A 45-cm chair was used and the faster of two unsupported five-rise trials retained. The endpoint included total and partial knee replacement, confirmed through radiographs or medical records, over 84 months. Sex-stratified Cox models accounted for clustered knees and adjusted progressively for demographic factors, history, pain and radiographic grade. [19]
Chair time was not significantly associated with seven-year replacement risk in either sex. In secondary shorter intervals, 0–30 and 60–84 months combined, slower chair time was associated with surgery among women in crude and partially adjusted analyses, but the association attenuated after adding pain and KL grade. These secondary estimates are logistic odds ratios, despite hazard-ratio wording in parts of the prose. The distinction matters when reporting the finding. This substantial prospective study provides useful negative and confounding-sensitive evidence; it does not support a chair-time-only surgical-risk rule. [19]
Current source qualification · AUD-KOA-STS-P06: Report already correctly calls theseORs.
When chair performance is the outcome rather than the predictor
Hu and colleagues analysed baseline quadriceps RFD and 36-month function in OAI participants with or at risk of OA. Chair-stand outcome data were available for 2,017 people, using average time from two five-rise trials. The study did not identify a validated chair-stand MIC; instead it combined the cohort's baseline SD with a reliability coefficient from another study to define an MDC90-based 4.16-second deterioration threshold. Knee surgery/replacement also counted as deterioration. This is an outcome definition adopted for that analysis, not a universal chair-stand MIC. [18]
Higher RFD was associated with less subsequent WOMAC-function worsening, but not with worsening walk or chair-stand performance in the adjusted analyses. The positive patient-reported result should not be extended to the objective chair outcome. Furthermore, this study predicts later chair performance from RFD; it does not establish STS as a prognostic predictor of another outcome. The precise RFD assay and its limitations are considered in the companion strength report. [18]
OAI-derived findings across onset, mediation, gait trajectories and RFD analyses share a parent cohort, with different subsets and outcomes. They are not independent cohorts simply because different performance variables were analysed. Similarly, Segal’s mechanics subcohort and Skou’s replacement analysis both use MOST, with different subsets and questions. [19, 21] Prospective falls, later surgery and worsening mobility also require separate endpoints and models. The intensively appraised chair literature does not justify an externally validated knee-OA STS-only calculator for those outcomes.
Table 6 Separate prospective endpoints and cohort scope
Shared OAI parent data do not constitute independent replication. No STS-only established-knee-OA risk calculator is endorsed.
| Source | Predictor → outcome | Finding | Boundary |
|---|---|---|---|
| Hiyama [15] | Five-rise time → onset | Symptom-defined association; structural null | Baseline KL0; abstract only |
| Thorstensson [16] | One-leg rises → radiographic OA | Onset association nonsignificant with simultaneous adjustment; prevalent progression null | Younger chronic-pain cohort; different task; full text now recovered |
| Dean [17] | Baseline STS → later activity within mediation model | STS-to-MVPA pathway nonsignificant | Mixed OAI cohort; scale/temporal issues |
| Hu [18] | RFD → later chair deterioration | Objective chair association not significant | STS is outcome; MDC-based deterioration proxy |
| Skou [19] | Five-rise time → later total/partial knee replacement | Seven-year null; shorter-interval association attenuated after pain/KL adjustment | Frequent knee pain with/without radiographic OA; no individual risk rule |
Practical implementation for rehabilitation and rehabtools
A clinically useful record
Begin with a clear question: transfer independence, repeated capacity, strategy, mechanical output or prognosis. For a repeated clinical test, document the chair and supports, task instruction, start/stop and repetition rules, practice, rest, trial aggregation, footwear and foot placement. Record current pain, flare state, medication/treatment timing, unilateral/bilateral disease and relevant comorbid limitations.
Record raw completion and failure data before calculating derived scores. An observed count should be retained even if a power estimate is later corrected. A video that cannot be analysed is not equivalent to an unsuccessful transfer. Preserve the reason for stopping and whether a modified test was used. An arm-assisted rise may be a meaningful achievement even when it cannot be compared with the unsupported reference protocol.
Interpret count/time alongside symptoms and movement strategy. A plausible increase beyond a matching error band can strengthen confidence in real capacity change. It does not prove a particular muscle mechanism or a patient-important benefit. Where no suitable error or MIC estimate exists, say so rather than importing a geriatric, hip-OA or post-TKA number without qualification.
Validation requirements for automated outputs
For automatic counting and timing, first validate against independently adjudicated video under the exact target protocol. Assess missed rises, incomplete extension, seat contact, arm use, final partial repetitions, occlusion, fatigue and pauses. Report per-person errors and failure rates, including difficult cases. Use participant-level rather than repetition-level train/test separation.
For joint motion, specify camera placement, calibration, frame rate, body landmarks, model, filtering and phase definitions. A two-dimensional projection and reconstructed three-dimensional angle require separate validation. Angle smoothing and numerical differentiation can change peak angular velocity; those outputs should not inherit the validation of the underlying angle by assumption.
For force, power and asymmetry, obtain a suitable reference measurement and specify what is estimated. An independent validated limb-specific force reference, such as separate force plates under each foot, is needed to validate limb-loading estimates. Joint moments/power require model and kinetic inputs; estimated whole-body power requires explicit anthropometric and time assumptions. Do not relabel angular acceleration, repetition speed or trunk velocity as RFD.
For remote deployment, test the actual service: first-time users, their chairs, realistic internet conditions, lower digital literacy, pain, obesity, aids and different homes. Identify when supervision or a safer alternative is necessary. An accurate score in a selected analyzable subgroup is insufficient if the system fails disproportionately in those with the greatest disability.
For prognosis, define the later outcome and horizon, freeze the full measurement/model pipeline, and externally evaluate calibration and discrimination in the intended knee-OA population. Compare against simple clinical predictors and determine whether decisions improve. A cross-sectional probability of OA status or an internally derived movement cluster is not a personal prognosis.
Evidence gaps and conclusion
The most useful next evidence would be knee-specific between-day reliability with explicit pain/stability assessment; adequately anchored MIC estimates across severity and treatment contexts; direct comparison of five-rise and fixed-time tests; validated event and asymmetry outputs; and patient-level remote/video validation with transparent failure reporting. Full originals of Khuna, Holm, Mostafaee and Ramalho are particularly important for the clinical-measurement account. The displayed STS-power equation and several newer digital-source inconsistencies require clarification before implementation.
Chair testing can contribute substantially to knee-OA assessment when the protocol and claim remain modest and precise. A count or time is a useful observation of task capacity. Instrumentation can explain strategy and loading. Neither should be allowed to imply an unmeasured muscle property or an unvalidated future outcome. The strongest initial rehabtools product is reproducible, interpretable measurement that makes these boundaries clear.
Primary study characteristics
Primary studies supporting the narrative are grouped by measurement or prognostic question. Any consensus recommendation is explicitly identified. Population, protocol, endpoint and source access constrain interpretation. The linked bibliography identifies source-access limitations. Related publications from one cohort are not independent replications.
Table 7 Clinical measurement and interpretation
| Study and population | Protocol and timing | Main findings | Interpretive limits |
|---|---|---|---|
| Dobson F 2013 [1] Multiphase expert consensus Hip/knee OA and arthroplasty intended scope | 30-second chair stand with approximately 43-cm chair; report adaptations; No clinical follow-up | Selected for minimal performance set alongside walking and stairs | Consensus does not independently validate a universal count threshold |
| Dobson F 2017 [2] Repeated-measures reliability 59 enrolled; 51 stable principal-analysis participants with hip and/or knee OA | Separate within-session rater and one-week same-rater designs; 7–9 days | Chair test achieved sufficient rather than uniformly optimal reliability | Mixed-joint sample; original numeric tables unavailable in the article text examined |
| Suwit A 2020 [3] Knee-specific reliability and construct validity 55 community participants; 60% bilateral OA | 43-cm chair; familiarization; same trained blinded assessor after one week; One week | ICC 0.87; SEM 1.4; MDC90 3.2 stands; mean 14.6 to 15.6; LoA −4.7 to +2.9 | Systematic retest gain; moderate strength and weak self-report associations; not universal MIC |
| Tolk JJ 2019 [4] Reliability, construct validity and postoperative responsiveness 85 pre-TKA participants; 30 retest; 70 responsiveness | 43-cm chair; same-day 30-minute retest; before and 12 months after TKA; 12 months after TKA | ICC 0.90; SEM 0.85; SDC95 2.4; mean 9.0 to 9.8; 42% validity and 50% responsiveness hypotheses met | Same-day error and postoperative response; discussion gives SDC as 2.5 |
| Gill S 2022 [5] Within-session reliability on two study occasions 93 KL-2 OA; BMI 33.5; pain duration 3.2 years; 83/61 complete repeated tests | 17-inch chair; three attempts 1–2 minutes apart within baseline and six-month sessions; Within session, repeated again at six months | ICCs 0.92/0.94; MDC90 2.27/2.79; second-trial gain 0.52/0.79 stands | Not six-month test–retest error; pain-related inability near 10% by third attempt; group MDC not individual |
| Holm PM 2021 [24] Reliability and agreement study 40 radiographic and/or symptomatic knee OA, per abstract | 30-second chair, walking, stairs and muscle tests; three-day retest; Three days | Abstract LoA ±2.4 chair repetitions; practice improvement; high ICCs | Original methods/tables unavailable; threshold remains qualified abstract evidence |
| Khuna L 2024 [23] 30-second and five-rise reliability/validity 60 older knee-OA participants; 30 retested, per abstract | Rater comparisons and one-week retesting; handheld dynamometry and WOMAC; One week | Near-perfect rater ICCs; between-occasion ICC approximately 0.84–0.85; substantial error described | No exact original SEM/MDC table available; do not infer negligible patient retest error |
| Mostafaee N 2024 [6] Physiotherapy responsiveness and MIC 60 knee-OA participants, per abstract | Four-week physiotherapy; seven-point global-rating anchor; Four weeks | Chair MIC 2.5 repetitions; core-test AUCs >0.70; anchor correlations 0.43–0.63 | Full anchor, denominators, ROC precision and protocol unverified; not equivalent to SDC |
| Ramalho RB 2024 [7] Prospective construct and responsiveness analysis 107 knee-OA participants, per abstract | Three OARSI core tasks with six-month follow-up; Six months | Chair failed ≥75% hypothesis criterion; walk/stairs favourable | Original full text unavailable; not a complete methods appraisal |
| Lee SH 2022 [8] Case-control discrimination 30 symptomatic KL 0–2 and 30 controls; 90% women | 45-cm chair; OARSI tasks; independently walking without aids; None | Chair AUC 0.76; cut-point 9.8; sensitivity 0.50, specificity 1.00 | Small selected case-control sample; cut-point is not prognosis or diagnosis in competing conditions |
Table 8 Digital measurement biomechanics and estimated power
| Study and population | Protocol and timing | Main findings | Interpretive limits |
|---|---|---|---|
| Aily JB 2024 [9] Cross-method assessment and recorded-video rescoring 32 KL 2–3; BMI<30; previous test experience | Prepared university; two same-day assessments separated by a 30-minute break; recordings rescored after six weeks; Same day and later video rescoring | Chair bias −0.22 reps; Figure 2 LoA −2.32 to +1.88; rescoring ICC 1.00 | Zero scorer MDC is not zero patient error; supported setting; original numerical issues for other tests |
| Rose MJ 2023 [10] Within-home repeatability and cross-setting agreement 20 enrolled; 18 home chair records; 16 usable lab chair records | Supplied chair and three Opal sensors; video guidance; five rapid rises; redon after 15 minutes; 15-minute repeat; separate home/lab visits | Average detected event duration, averaged across two trials per set: ICC 0.89; MDC95 0.25 seconds; home/lab ICC 0.66 | Event duration is not total five-rise time; failures and manual clock correction; selected independent walkers |
| Dahlberg LE 2025 [11] Digital/therapist agreement and home repeatability Clinic 18:14 knee/4 hip; home 44 completers:32 knee/12 hip | 45-cm chair; app countdown without final alarm; digitally experienced users; home 10–14-day retest; Same-day comparison; home 10–14 days | ICCs 0.87/0.88; individual MDC95 4.4/4.8 reps; positive biases 1.5/1.2 reps | Mixed hip/knee; ten home noncompleters; positive differences conflict with 'no systematic bias' prose |
| Zhao Z 2024 [12] Cross-sectional video pattern analysis 104 advanced knee-OA participants; frequent reliance on walking aids and specified comorbidities excluded | Independent five-rise videos; AlphaPose/VideoPose; arbitrary view; common-length time-series processing; None | Group-average angle/wavelet differences by WOMAC categories | No concurrent reference validation, individual threshold, SEM/MDC or prospective model; excluded poor videos |
| Pan J 2024 [25] Case-control motion/force/EMG study 24 unilateral KL 1–2 and 12 controls; affected dominant side; BMI<28 | 43-cm chair; repeated 30-second tasks; Vicon, dual force plates and EMG; None | Altered timing, ROM, angular velocity and affected-side loading | Small sample, BMI difference, many comparisons and unusual displayed inferential values; no universal target |
| Segal NA 2013 [21] Cross-sectional strategy analysis 60 symptomatic OA from MOST; 49 usable motion records; 24 bilateral | Five-rise strata; separate self-paced rises from 46-cm chair; single force plate; None | Sex-specific strategy associations; foot placement may help explain the observed differences; strength does not explain all strata | Only more symptomatic limb kinetics; cannot establish bilateral symmetry; no causal/prospective inference |
| Sagawa Y 2017 [22] Exploratory movement-strategy grouping 101 severe knee-OA pre-TKA participants and 27 controls | Chair adjusted for 90° knees; self-paced rises; two force plates and 12 cameras; None | Compensated group had control-like time with greater trunk flexion/lean; three derived groups | Derived and compared in same data; grouping is not externally validated phenotype or treatment rule |
| Langgård Jørgensen S 2024 [20] Cross-sectional trial-baseline analysis 86 advanced OA; 16 contralateral TKA and five contralateral THA | 44-cm chair; 3–5 practice rises; one 30-second test; anthropometric power estimate; stabilized HHD; None | Estimated power associated with concurrent mobility; some male PROM associations | Not prospective or treatment response; displayed equation omits g; actual computation unverified; mixed replacement history |
| Chan LC 2026 [13] Cross-sectional technical validation and internal classification Stated 309 participants; 168 OA/141 controls; patient-level 80:20 split | Seven self-paced rises; right side-view smartphone; BlazePose; reference alignment with dynamic time warping; None | Reported AUC around 0.776 for symptomatic OA status | Count/metric inconsistencies; no external calibration or longitudinal responsiveness; warping is not unwarped timing validation |
| Torres RTG 2026 [14] Concurrent markerless/marker-based comparison 22 knee OA; 21 usable STS; trial NCT04243096 | 18-inch chair; five self-paced rises; last rise excluded; synchronized force plates; Single visit | Sagittal ROM ICC 0.66, bias 4.2°, LoA −4.6 to +13.0°; frontal ROM ICC 0.03 | Kinetics use measured forces; not camera-only; parameter-specific; related programme to [10] |
Table 9 Prospective outcomes and adjacent populations
| Study and population | Protocol and timing | Main findings | Interpretive limits |
|---|---|---|---|
| Hiyama Y 2026 [15] Prospective OAI onset study 1,003 people/2,006 knees; baseline KL0, per abstract | Five-rise time; structural and symptom-defined onset outcomes; Longitudinal onset | Slower rise associated with symptom-defined onset; structural association not clear | Original unavailable; no baseline radiographic OA; symptoms defined by WOMAC pain threshold/increase; not established-OA progression |
| Thorstensson CA 2004 [16] Population cohort with chronic knee pain 148 aged 35–54; normal versus prevalent radiographic groups, per abstract | Maximum one-leg rises, 300-m walk and one-leg stance; About five years | Unadjusted and separately adjusted onset association; simultaneous adjustment OR 2.67 (95% CI 0.92–7.76), n=70, nonsignificant; no prevalent progression predictor | One-leg task not bilateral 5STS/30CST; baseline prevalent OA KL≥1; full text now recovered |
| Dean K 2026 [17] OAI serial mediation analysis 782 complete records; progression, incidence and control cohort mix; surgery excluded | Baseline pain and STS; later activity and quality of life; age/sex/race adjustment; Two years | STS-to-later-MVPA β −0.02, P=0.46; overall indirect effect does not prove serial STS pathway | Scale/summary-statistic concerns; same-wave variable pairs; no baseline activity/quality-of-life adjustment established |
| Hu B 2018 [18] OAI RFD prognosis analysis With/at-risk OA; 2,017 chair-outcome records | Baseline 30–90% force-slope RFD; mean two five-rise times; surgery also worsening; 36 months | Positive WOMAC result, but no significant objective chair/walk worsening association | Chair worsening threshold 4.16 seconds is MDC90 proxy, not validated MIC; STS is outcome, not predictor |
| Skou ST 2016 [19] Prospective MOST cohort Frequent knee pain with/without radiographic OA; 1,173 people and 1,564 knees with chair data | 45-cm chair; best of two five-rise trials; sex-stratified cluster-adjusted Cox models; 84 months; secondary 0–30 and 60–84-month intervals | No seven-year chair association in either sex; shorter-interval female association attenuates with pain/KL adjustment | Total and partial replacement endpoint; mixed radiographic status; secondary logistic estimates are ORs, not HRs |
References
References are numbered in first citation order. Study specific source descriptions identify the material examined and do not constitute a study quality rating. Links identify the original publication or the explicitly named primary source version.
1. Dobson F, Hinman RS, Roos EM, Abbott JH, Stratford P, Davis AM, et al. OARSI recommended performance-based tests to assess physical function in people diagnosed with hip or knee osteoarthritis. Osteoarthritis and cartilage. 2013;21(8):1042-52. DOI 10.1016/j.joca.2013.05.002 Source examined: Article body complete; consensus.
2. Dobson F, Hinman RS, Hall M, Marshall CJ, Sayer T, Anderson C, et al. Reliability and measurement error of the Osteoarthritis Research Society International (OARSI) recommended performance-based tests of physical function in people with hip and knee osteoarthritis. Osteoarthritis and cartilage. 2017;25(11):1792-1796. DOI 10.1016/j.joca.2017.06.006 Source examined: Article body complete; numeric original tables not retrieved.
3. Suwit A, Rungtiwa K, Nipaporn T. Reliability and Validity of the Osteoarthritis Research Society International Minimal Core Set of Recommended Performance-Based Tests of Physical Function in Knee Osteoarthritis in Community-Dwelling Adults. The Malaysian journal of medical sciences : MJMS. 2020;27(2):77-89. DOI 10.21315/mjms2020.27.2.9 Source examined: Article body including tables complete.
4. Tolk JJ, Janssen RPA, Prinsen CAC, Latijnhouwers DAJM, van der Steen MC, Bierma-Zeinstra SMA, et al. The OARSI core set of performance-based measures for knee osteoarthritis is reliable but not valid and responsive. Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA. 2019;27(9):2898-2909. DOI 10.1007/s00167-017-4789-y Source examined: Original article including methods, results, flow diagram and Table 2.
Source note: The OARSI core set of performance-based measures for knee osteoarthritis is reliable but not valid and responsive.
5. Gill S, Hely R, Page RS, Hely A, Harrison B, Landers S. Thirty second chair stand test: Test-retest reliability, agreement and minimum detectable change in people with early-stage knee osteoarthritis. Physiotherapy research international : the journal for researchers and clinicians in physical therapy. 2022;27(3):e1957. DOI 10.1002/pri.1957 Source examined: Article body including tables and protocol appendix complete.
Source note: Gill S, Hely R, Page RS, Hely A, Harrison B, Landers S. Thirty second chair stand test: Test-retest reliability, agreement and minimum detectable change in people with early-stage knee osteoarthritis. Physiotherapy research international : the journal for researchers and clinicians in physical therapy. 2022;27(3):e1957. DOI 10.1002/pri.1957
6. Mostafaee N, Rashidi F, Negahban H, Ebrahimzadeh MH. Responsiveness and minimal important changes of the OARSI core set of performance-based measures in patients with knee osteoarthritis following physiotherapy intervention. Physiotherapy theory and practice. 2024;40(5):1028-1039. DOI 10.1080/09593985.2022.2143253 Source examined: Abstract only; original full text not retrieved.
7. Ramalho RB, Casonato NA, Montilha VB, Chaves TC, Mattiello SM, Selistre LFA. Construct Validity and Responsiveness of Performance-based Tests in Individuals With Knee Osteoarthritis. Archives of physical medicine and rehabilitation. 2024;105(10):1862-1869. DOI 10.1016/j.apmr.2024.05.024 Source examined: Abstract only; original full text not retrieved.
Source note: Construct Validity and Responsiveness of Performance-based Tests in Individuals With Knee Osteoarthritis.
8. Lee SH, Kao CC, Liang HW, Wu HT. Validity of the Osteoarthritis Research Society International (OARSI) recommended performance-based tests of physical function in individuals with symptomatic Kellgren and Lawrence grade 0-2 knee osteoarthritis. BMC musculoskeletal disorders. 2022;23(1):1040. DOI 10.1186/s12891-022-06012-2 Source examined: Article body including tables complete.
9. Aily JB, da Silva AC, de Noronha M, White DK, Mattiello SM. Concurrent Validity and Reliability of Video-Based Approach to Assess Physical Function in Adults With Knee Osteoarthritis. Physical therapy. 2024;104(6). DOI 10.1093/ptj/pzae039 Source examined: Complete original article, Figure 2 and official supplement.
10. Rose MJ, Neogi T, Friscia B, Torabian KA, LaValley MP, Gheller M, et al. Reliability of Wearable Sensors for Assessing Gait and Chair Stand Function at Home in People With Knee Osteoarthritis. Arthritis care & research. 2023;75(9):1939-1948. DOI 10.1002/acr.25096 Source examined: Article body including tables complete; supplement not separately retrieved.
11. Dahlberg LE, Karlsson O, Sirard P, Lohmander LS, Kiadaliri A. Reliability of digitally instructed self-reported 30-second chair stand test for lower extremity function. Osteoarthritis and cartilage open. 2025;7(2):100613. DOI 10.1016/j.ocarto.2025.100613 Source examined: Article body including tables complete; mixed hip/knee sample.
Source note: Dahlberg LE, Karlsson O, Sirard P, Lohmander LS, Kiadaliri A. Reliability of digitally instructed self-reported 30-second chair stand test for lower extremity function. Osteoarthritis and cartilage open. 2025;7(2):100613. DOI 10.1016/j.ocarto.2025.100613
12. Zhao Z, Yang T, Qin C, Zhao M, Zhao F, Li B, et al. Exploring the potential of the sit-to-stand test for self-assessment of physical condition in advanced knee osteoarthritis patients using computer vision. Frontiers in public health. 2024;12:1348236. DOI 10.3389/fpubh.2024.1348236 Source examined: Article body including tables complete.
Source note: Zhao Z, Yang T, Qin C, Zhao M, Zhao F, Li B, et al. Exploring the potential of the sit-to-stand test for self-assessment of physical condition in advanced knee osteoarthritis patients using computer vision. Frontiers in public health. 2024;12:1348236. DOI 10.3389/fpubh.2024.1348236
13. Chan LC, Yan J, Zhang YC, Jiang T, Zhang AY, Li HHT, et al. Smartphone-derived joint angular velocities in sit-to-stand motion provide a spatiotemporal marker for symptomatic knee osteoarthritis. Communications medicine. 2026;6(1). DOI 10.1038/s43856-026-01537-2 Source examined: Article body including tables complete; source count/performance inconsistencies flagged.
Source note: Chan LC, Yan J, Zhang YC, Jiang T, Zhang AY, Li HHT, et al. Smartphone-derived joint angular velocities in sit-to-stand motion provide a spatiotemporal marker for symptomatic knee osteoarthritis. Communications medicine. 2026;6(1). DOI 10.1038/s43856-026-01537-2
14. Torres RTG, Senderling B, Kim E, Rose M, Gheller M, Neogi T, et al. Agreement between markerless and marker-based motion capture for knee kinematics and kinetics during functional activities in knee osteoarthritis. Osteoarthritis and cartilage open. 2026;8(3):100853. DOI 10.1016/j.ocarto.2026.100853 Source examined: Article body including tables complete; supplementary model comparisons not separately examined.
15. Hiyama Y. Chair stand performance and symptomatic knee osteoarthritis onset: a prospective cohort study. Rheumatology international. 2026;46(6). DOI 10.1007/s00296-026-06142-z Source examined: Abstract only; at-risk baseline KL0 cohort; original full text not retrieved.
Source note: Hiyama Y. Chair stand performance and symptomatic knee osteoarthritis onset: a prospective cohort study. Rheumatology international. 2026;46(6). DOI 10.1007/s00296-026-06142-z
16. Thorstensson CA, Petersson IF, Jacobsson LT, Boegård TL, Roos EM. Reduced functional performance in the lower extremity predicted radiographic knee osteoarthritis five years later. Annals of the rheumatic diseases. 2004;63(4):402-7. DOI 10.1136/ard.2003.007583 Source examined: Abstract only; original full text not retrieved.
Source note: Thorstensson CA, Petersson IF, Jacobsson LT, Boegård TL, Roos EM. Reduced functional performance in the lower extremity predicted radiographic knee osteoarthritis five years later. Annals of the rheumatic diseases. 2004;63(4):402-7. DOI 10.1136/ard.2003.007583
17. Dean K, Nemati D, Sun R, Smuck M, Kushioka J, Best TM, et al. Sit-to-Stand Performance, Moderate-to-Vigorous Physical Activity, and Health-Related Quality of Life in Knee Osteoarthritis: A Prospective Mediation Analysis. Preventing chronic disease. 2026;23:E09. DOI 10.5888/pcd23.250405 Source examined: Article body including tables complete; major outcome/scale reporting concerns.
Source note: Dean K, Nemati D, Sun R, Smuck M, Kushioka J, Best TM, et al. Sit-to-Stand Performance, Moderate-to-Vigorous Physical Activity, and Health-Related Quality of Life in Knee Osteoarthritis: A Prospective Mediation Analysis. Preventing chronic disease. 2026;23:E09. DOI 10.5888/pcd23.250405
18. Hu B, Skou ST, Wise BL, Williams GN, Nevitt MC, Segal NA. Lower Quadriceps Rate of Force Development Is Associated With Worsening Physical Function in Adults With or at Risk for Knee Osteoarthritis: 36-Month Follow-Up Data From the Osteoarthritis Initiative. Archives of physical medicine and rehabilitation. 2018;99(7):1352-1359. DOI 10.1016/j.apmr.2017.12.027 Source examined: Article body including tables complete; with/at-risk OAI cohort.
19. Skou ST, Wise BL, Lewis CE, Felson D, Nevitt M, Segal NA, et al. Muscle strength, physical performance and physical activity as predictors of future knee replacement: a prospective cohort study. Osteoarthritis and cartilage. 2016;24(8):1350-6. DOI 10.1016/j.joca.2016.04.001 Source examined: Article body including tables complete; independently checked original methods and outcome labels.
20. Langgård Jørgensen S, Mechlenburg I, Bagger Bohn M, Aagaard P. Sit-to-stand power predicts functional performance and patient-reported outcomes in patients with advanced knee osteoarthritis. A cross-sectional study. Musculoskeletal science & practice. 2024;69:102899. DOI 10.1016/j.msksp.2023.102899 Source examined: Complete original article; displayed equation examined.
21. Segal NA, Boyer ER, Wallace R, Torner JC, Yack HJ. Association between chair stand strategy and mobility limitations in older adults with symptomatic knee osteoarthritis. Archives of physical medicine and rehabilitation. 2013;94(2):375-83. DOI 10.1016/j.apmr.2012.09.026 Source examined: Article body including tables complete.
Source note: Segal NA, Boyer ER, Wallace R, Torner JC, Yack HJ. Association between chair stand strategy and mobility limitations in older adults with symptomatic knee osteoarthritis. Archives of physical medicine and rehabilitation. 2013;94(2):375-83. DOI 10.1016/j.apmr.2012.09.026
22. Sagawa Y, Bonnefoy-Mazure A, Armand S, Lubbeke A, Hoffmeyer P, Suva D, et al. Variable compensation during the sit-to-stand task among individuals with severe knee osteoarthritis. Annals of physical and rehabilitation medicine. 2017;60(5):312-318. DOI 10.1016/j.rehab.2017.03.007 Source examined: Article body complete; supplementary coding appendix not separately retrieved.
Source note: Sagawa Y, Bonnefoy-Mazure A, Armand S, Lubbeke A, Hoffmeyer P, Suva D, et al. Variable compensation during the sit-to-stand task among individuals with severe knee osteoarthritis. Annals of physical and rehabilitation medicine. 2017;60(5):312-318. DOI 10.1016/j.rehab.2017.03.007
23. Khuna L, Soison T, Plukwongchuen T, Tangadulrat N. Reliability and concurrent validity of 30-s and 5-time sit-to-stand tests in older adults with knee osteoarthritis. Clinical rheumatology. 2024;43(6):2035-2045. DOI 10.1007/s10067-024-06969-6 Source examined: Abstract only; original full text not retrieved.
Source note: Khuna L, Soison T, Plukwongchuen T, Tangadulrat N. Reliability and concurrent validity of 30-s and 5-time sit-to-stand tests in older adults with knee osteoarthritis. Clinical rheumatology. 2024;43(6):2035-2045. DOI 10.1007/s10067-024-06969-6
24. Holm PM, Nyberg M, Wernbom M, Schrøder HM, Skou ST. Intrarater Reliability and Agreement of Recommended Performance-Based Tests and Common Muscle Function Tests in Knee Osteoarthritis. Journal of geriatric physical therapy (2001). 2021;44(3):144-152. DOI 10.1519/jpt.0000000000000266 Source examined: Abstract only; original full text not retrieved.
25. Pan J, Fu W, Lv J, Tang H, Huang Z, Zou Y, et al. Biomechanics of the lower limb in patients with mild knee osteoarthritis during the sit-to-stand task. BMC musculoskeletal disorders. 2024;25(1):268. DOI 10.1186/s12891-024-07388-z Source examined: Article body including tables complete; inferential inconsistencies noted.
Source note: Pan J, Fu W, Lv J, Tang H, Huang Z, Zou Y, et al. Biomechanics of the lower limb in patients with mild knee osteoarthritis during the sit-to-stand task. BMC musculoskeletal disorders. 2024;25(1):268. DOI 10.1186/s12891-024-07388-z