← Clinical Evidence

Balance assessment and prognosis after total knee arthroplasty

This report reviews balance assessment in total knee arthroplasty. It examines measurement properties, interpretation of change and prognostic evidence, with the limits of each study and testing protocol.

In this report
Audited and Updated

Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.

TKA-C01

0%–8.9% across the BESTest versions at 12 and 24 weeks (Table 4); the article narrative states 2.2%–8.9%, but the full BESTest had 0% at the maximum at 12 weeks.

Type: source table prose conflict. Audit disposition: supported with version caveat.

Remaining limit: Final published table not independently compared.

TKA-Q02

Retain 0.706 and identify Table3 as the source; note that the indexed abstract prints0.701.

Type: source table abstract conflict. Audit disposition: supported.

TKA-Q03

Keep these claims marked unverified until the dated raw search pages, returned identifiers, reconciliation logs and historical source-access artifacts are supplied. Current successful source retrieval cannot validate historical search completeness.

Type: historical provenance gap. Audit disposition: supported as limit.

TKA-C04

Table4 labels a 28-point Mini-BESTest maximum, but the Methods describes16 items scored0–2. Preserve the table label and disclose the source-version conflict; do not operationalize the threshold until the scoring protocol is clarified.

Type: source scoring description ambiguity. Audit disposition: supported with version caveat.

Remaining limit: Item-versus-task/bilateral scoring conventions could explain ambiguity; not proven scoring error.

TKA-U02

Retain separate-protocol caution and add source warning if detailed values are introduced: abstract SEM/SRD.80/2.24s differs from Table2 and Discussion1.11/3.10s.

Type: source internal conflict. Audit disposition: supported.

Editorial record

  • Audit status: supported with version caveat. Final published table not independently compared.
  • Audit status: supported.
  • Audit status: supported.
  • Audit status: supported with version caveat. Item-versus-task/bilateral scoring conventions could explain ambiguity; not proven scoring error.
  • Audit status: supported.
  • Audit status: supported.
  • Audit status: supported with version caveat. Item-versus-task/bilateral scoring conventions could explain ambiguity; not proven scoring error.
  • Audit status: supported.

Clinical interpretation

Balance after total knee arthroplasty (TKA) is a collection of task-dependent abilities rather than a single score. A patient may stand quietly with little sway, perform a fast Timed Up and Go (TUG), and still have difficulty recovering from an unexpected trip. Pain relief, increasing confidence, and improvement in everyday mobility are important outcomes, but they do not establish recovery of sensory integration, reactive stepping, or the ability to maintain equilibrium on the operated leg. The strongest practical approach is to combine a clearly specified performance assessment with information about actual falls, near-falls, walking aids, pain, and the circumstances in which instability occurs. [1–6]

For an ambulatory outpatient after primary elective TKA for osteoarthritis (OA), a BESTest-family assessment provides broader balance information than a brief mobility time alone. The full BESTest is useful when the clinical question concerns which balance systems appear limited; a correctly identified Mini-BESTest version is a more time-efficient alternative. The Berg Balance Scale (BBS) remains useful for describing basic tasks, but its ceiling can obscure recovery in independently mobile patients. Single-leg stance and step/reach tests are useful targeted supplements, provided the examiner records the stance side, support, vision, familiarization, and stopping rule. A force plate can quantify a selected standing strategy, but its output is not automatically a validated fall-risk score. [2–5, 7–13]

Three boundaries govern interpretation. First, postoperative day, surgical indication, unilateral versus bilateral replacement, and the state of the other knee must accompany every result. Second, the ability to reproduce a score, to detect change, and to forecast a future outcome are different properties. Third, published thresholds must retain their original scale and protocol. The TKA literature contains genuine measurement studies and prospective falls cohorts, but also examples in which a scale labeled “fall risk” was not tested against future falls, or a “prediction” model reproduced a concurrent therapist decision. These distinctions materially affect what a rehabilitation dashboard should display.

Scope and evidence approach

This critical narrative assessment-and-prognosis review focuses on adults after primary elective total knee replacement, principally for OA. Preoperative findings are retained when they establish the baseline, examine a predictor of later recovery, or clarify that a reliability estimate was actually obtained before surgery. Simultaneous bilateral TKA, revision TKA, unicompartmental replacement, inflammatory or fracture indications, and mixed hip/knee cohorts are separate applicability strata. Operative ligament balancing, implant alignment, and intraoperative gap measurement are not standing balance and were not treated as evidence about postural control.

The evidence backbone comprises dated PubMed and Scopus searches through 2 October 2026, targeted balance/falls/perturbation expansion, reference and forward-citation chasing, and original-source appraisal. The exact searches, coverage reconciliation, source access, and selected-study matrix accompany this manuscript. PubMed page overlap was checked against official NCBI records. The principal balance query contained 566 official records; the targeted supplement contained 76, overlapping with the principal set. These are retrieval counts, not eligible-study counts. This is not a formal systematic review: no claim of duplicate independent screening, exhaustive study inclusion, or pooled certainty grading is made.

What each assessment measures

Standing control and functional balance

Quiet bipedal standing measures control within a particular base of support. Center-of-pressure (CoP) path, velocity, displacement, and area describe different aspects of that control. They depend on foot placement, visual conditions, trial duration, surface, filtering, and whether a support device is used. Smaller sway is not invariably better: a patient may stiffen, unload the operated limb, or constrain movement. CoP is the point of application of the resultant ground reaction force, not a direct measure of whole-body center-of-mass movement. Normal quiet-standing output cannot rule out impaired stepping or poor obstacle negotiation.

Single-leg stance raises the load and balance demands on one limb. Time maintained combines postural control with strength, pain tolerance, confidence, and the ability to accept weight. A test on the dominant leg is not equivalent to a test on the operated leg. After simultaneous bilateral TKA, there is no untreated knee that can serve as a normal comparator. A short time may be informative, but inability to start safely should be recorded rather than converted into an ordinary measured zero or removed from the analysis. [5, 7]

Step and reach tests add an anticipatory weight-shift requirement. Reach distance also depends on limb length, joint motion, strength, and strategy. The Star Excursion Balance Test (SEBT) is therefore not a pure proprioception test. A person can increase reach through improved flexion or a different trunk strategy without a proportionate change in reactive balance. Conversely, a postoperative motion restriction can limit reach while equilibrium control is relatively preserved. [7, 8]

The BESTest samples multiple systems, including postural responses and sensory orientation. The Mini-BESTest emphasizes dynamic balance; the Brief-BESTest samples fewer tasks and sacrifices measurement breadth. The BBS concentrates on common functional tasks and does not fully sample rapid reactive stepping. TUG, walking tests, and chair-rise tests contain balance demands, but performance also depends strongly on gait, strength, endurance, transfers, and willingness to move quickly. Correlation with BBS supports overlap among constructs, not the conclusion that these tests are interchangeable measures of standing balance. [2, 5]

Reactive balance and actual falls

Reactive balance concerns restoring equilibrium after an external or self-generated disturbance, including a compensatory step. It cannot be inferred simply from quiet standing or planned reaching. Perturbation studies after TKA have used sudden support-surface movement, unidirectional disturbances, and tether-release forward falls. Their outcomes include body displacement, recovery-step length and velocity, muscle response, and success in recovering. These studies establish task-specific differences; they do not by themselves provide a clinical threshold for future community falls. [14–16]

A fall is an event, whereas a balance score is a measurement of performance under selected conditions. Falls additionally reflect exposure, environment, medication effects, attention, activity, and use of assistance. A more active patient can encounter more hazards despite improving capacity. An inpatient fall during first mobilization differs from a trip six months later. Record falls, near-falls and buckling separately, with the recall or surveillance period, assistance, injury, activity, and location. Balance confidence or concern about falling is another distinct construct. ABC and FES-I can complement performance testing, but a confident person is not necessarily a safe person and a fearful person is not necessarily a future faller. [6, 17–21]

Table 1 Choosing a balance assessment

Evidence and interpretation from [2–15, 17–21]. These are construct-based choices, not a universal mandatory battery.

Table 1 Choosing a balance assessment
AssessmentBest useMain limitation
BESTest familyMultidomain balance, sensory orientation and postural responsesVersion and scoring denominator must match evidence; selected aid-free cohorts
BBSBasic functional balance and lower-ability task descriptionSubstantial later ceiling; a maximum score does not rule out reactive deficits
Single-leg stanceLimb-specific weight acceptance and stance capacityPain, strength and confidence contribute; large error in bilateral TKA study
Step and reach testsAnticipatory control during weight shiftingHeight, direction, limb length, stance side and movement strategy matter
Quiet-standing CoPQuantify a specified standing strategy or sensory conditionTask and processing specific; not a stand-alone falls probability
Perturbation responseRecovery after an unexpected disturbanceSafety, direction, magnitude and anticipation limit portable use
Falls history and surveillanceActual event burden and circumstancesDifferent recall windows and exposure are not directly comparable

Reliability and measurement error

BESTest family and Berg Balance Scale

Chan and Pang studied 92 people after first TKA for end-stage OA in two groups: 46 for reliability and another 46 for validity across recovery. The reliability group was at least six months postoperative; 25 were scored concurrently by three raters, and 46 repeated testing with the same rater within a week, without intervening physiotherapy. No participant used a walking aid during testing. Reliability was expressed as ICC(2,1); SEM came from the square root of the ANOVA mean-square error, and MDC95 was SEM multiplied by 1.96 and the square root of two. The independent-rater design assessed scoring of the same performance, not three independent administrations. [2]

These results support score reproducibility under a specific later-stage protocol. They do not directly establish day-to-day error during the first postoperative fortnight, reliability in aid-dependent patients, or error for a different Mini-BESTest scoring system. Shared items were performed once and scored for several scales; this reduces independent task sampling and can raise between-scale association. High internal consistency also does not prove that every component is redundant or that a single physiological system is measured.

BBS ceiling effects were clinically substantial: 52.2% reached the maximum at 12 weeks (n = 46) and 57.8% at 24 weeks (n = 45), compared with 0%–8.9% across the BESTest versions in Table 4 at those later assessments (the source prose states 2.2%–8.9%). A stable BBS score near 56 may reflect limited headroom rather than no recovery. FGA score skew was also apparent, even when its proportion at the exact maximum was lower. Choose the scale for the expected ability range and the clinical question, rather than simply selecting the test with the highest ICC . [2]

Table 2 Direct balance measurement error estimates

[2, 5, 8, 9]. The last two rows are abstract-verified estimates, not implementation-ready thresholds. SEM is one measurement error; MDC95 concerns a two-measurement difference.

Table 2 Direct balance measurement error estimates
Test and populationModel and repeat designSEM and MDC95Interpretation boundary
BESTest percentage, ≥6 months after first OA-TKAICC(2,1) 0.96; same rater within one week, n = 462.24 and 6.22 percentage pointsNot raw points out of 108 or an early-postoperative estimate
Mini-BESTest 32, same cohortICC(2,1) 0.92; n = 461.34 and 3.71 pointsDo not apply to Mini 28
Brief-BESTest 24, same cohortICC(2,1) 0.94; n = 461.15 and 3.19 pointsError threshold, not MIC
BBS 56, same cohortICC(2,1) 0.97; n = 460.72 and 2.00 pointsLater ceiling limits evaluative utility
Single-leg stance, bilateral TKA seventh monthTwo-way mixed ICC 0.72 (0.47–0.85); one-day retest, n = 417.08 and 19.62 secondsDominant-leg eyes-open protocol; mean increased 13.96→19.47 s
Step Test, n = 56 TKAAbstract ICC 0.90; model/stage unverified0.76 and 2.11 score units reportedFull step protocol required before use
Dual-task FSST, n = 30 TKAICC(2,1) 0.97; two same-day trialsMDC95 3.43 seconds; SEM not verifiedCognitive task and exact stage require full original

Single leg stance and brief dynamic tests

Sarac and colleagues provide useful but distinct evidence from 41 patients in the seventh month after simultaneous bilateral TKA for severe OA. Participants could stand and walk without aids. Single-leg stance was performed eyes open, arms crossed, standing on the dominant leg, with a practice trial. Same-assessor retesting occurred one day later; the stated ICC model was two-way mixed, without complete single-versus-average and agreement notation. ICC was 0.72 (95% CI 0.47–0.85), SEM 7.08 seconds, and MDC95 19.62 seconds. Mean time increased from 13.96 to 19.47 seconds. Large absolute error and a substantial retest shift substantially limit small-change interpretation. [5]

This is not a unilateral operated-leg MDC. It also illustrates why ranking tests only by their correlation with BBS is inadequate: the weakest association may reflect a different construct as well as measurement noise. The original Table 3 gives the single-leg stance/BBS correlation as 0.704; the prose compresses several different correlations into an incorrect single summary. The original table, population, and protocol should be retained when building an evidence library.

For the Step Test, a TKA study of 56 people reports ICC 0.90, SEM 0.76 and MDC95 2.11, with associations with TUG and 10-m walking. The accessible abstract does not establish all details needed to implement this error threshold, including the full ICC model, exact postoperative stage, and step/trial procedure. It is evidence that direct TKA error research exists, but the number should remain non-deployable until the full protocol is obtained. [8]

The dual-task Four Square Step Test report provides same-day reliability in 30 people after TKA: ICC(2,1) 0.97 and MDC95 3.43 seconds. Its accessible abstract reports correlations with single-task TUG and the Hospital for Special Surgery knee score, not prospective falls. The cognitive task, scoring of cognitive accuracy, prioritization instructions and exact stage remain necessary before clinical replication. Modified four-square stepping has also been investigated, but a modification should never inherit the original test’s thresholds solely because the name is similar. [9, 10]

Reach testing before and after surgery

The SEBT study tested selected people with unilateral radiographic KL4 OA before TKA and followed them at six and twelve months. Retest reliability was assessed before surgery, seven days apart, using anterior, posteromedial and posterolateral reaches, barefoot stance on the affected limb, four practice trials, and the mean of three scored trials normalized to limb length. Thirty-five participants formed the final retest sample. The reported ICC range was 0.993–0.998, with small direction-specific error estimates. Those are end-stage preoperative OA reliability results, not postoperative TKA reliability results. [7]

The study did demonstrate postoperative change and examined correlations of change with other functional outcomes. This supports evaluative potential during later recovery, but its distribution-based response is not a patient-anchored MIC. Strict comorbidity and complication exclusions, substantial screening loss, and a challenging weight-bearing task limit applicability. The participant-flow arithmetic in the text also does not reconcile cleanly. Retain this study as a selected-cohort recovery and protocol source; do not make its preoperative MDC an automatic postoperative change flag.

Change detection and clinical importance

Responsiveness is not a fixed test property independent of time. Chan and colleagues assessed 134 participants at 2, 4, 8, 12 and 24 weeks. The full BESTest changed more consistently with recovery than BBS, and changes in BESTest and Mini-BESTest were more closely associated with FGA change than changes in BBS. Reported standardized response means varied by interval; the full BESTest range was 0.60–1.14, Mini-BESTest 0.40–0.94, Brief-BESTest 0.27–0.91 and BBS 0.19–0.70. These describe change relative to variability in that cohort, not individual minimal benefit. Associations between change scores across the same intervals are not forecasts of later function. [3]

The subsequent important-change study used 134 complete cases from 146 recruited, comparing two and four postoperative weeks. FGA improvement of at least four points defined the anchor-positive group; that four-point threshold had been established in an older-adult population rather than independently anchored to patient-perceived recovery after TKA. Seventy-two participants met it. BESTest, Mini-BESTest and Brief-BESTest change had AUCs of 0.811, 0.782 and 0.706 respectively; BBS discrimination was poor, with AUC 0.586. [4]

The published integer anchor thresholds were eight raw BESTest points, two Mini-BESTest points and three Brief-BESTest points. Distribution estimates were then combined with anchor estimates into reported ranges. These are not confidence intervals, and the distribution endpoint does not independently establish importance. The study used a 28-point Mini-BESTest and a raw 108-point BESTest, unlike the earlier 32-point Mini-BESTest and percentage BESTest reliability report. Printed relative percentages also do not map exactly to the rounded integer values, so this report preserves the raw, labeled values and does not silently recalculate them. [2, 4]

The meaningful interpretation is narrower than “two points means clinically important improvement.” It is an early-postoperative improvement threshold discriminating an FGA-defined response under this protocol. It is not a deterioration threshold, a patient-anchored balance MIC, a fall-prevention benefit threshold, or a universal difference exceeding measurement error. A small potentially worthwhile improvement may still be indistinguishable from noise for an individual. Conversely, an improvement beyond MDC can be real without mattering to the patient. The app should show these questions separately and suppress automatic MIC classification when the protocol, scale version, stage, or anchor differs.

Table 3 Reported early balance important change

[4] Table 4 visually checked. n = 134, postoperative weeks 2→4; anchor FGA improvement≥4, itself borrowed from older-adult evidence. Printed percentages and rounded raw cutoffs do not map exactly. Distribution/anchor ranges are not CIs. None is a falls-prevention threshold.

Table 3 Reported early balance important change
ScalePublished integer anchor cutoffAUC and 95% CIImportant qualification
BESTest raw 1088 points.811 (.739–.883)Different scale from percentage MDC
Mini-BESTest 282 points.782 (.704–.860)Table 4 labels maximum 28, but Methods describes 16 items scored 0–2; scoring protocol unresolved. Preserve the published 2-point cutoff as a source report, not an operational threshold. Different version from Mini 32 MDC
Brief-BESTest 243 points.706 (.618–.795)Limited discrimination; no patient-perceived anchor
BBS 565 points (published; not recommended by the source).586 (.490–.682)Authors withheld the combined MIC range and did not recommend the BBS anchor estimate because discrimination was poor; not a usable improvement threshold

Prognosis and future falls

The strongest direct postoperative falls cohort

Chan, Jehu and Pang measured clinical variables at four weeks and followed participants monthly for six months. Among 134 complete cases, 23 people sustained at least one fall and 31 fall episodes occurred. Most events involved walking, with slipping and tripping common. The model used univariate screening followed by backward logistic selection. Younger age, worse operated-knee proprioception, greater pain and lower BESTest sensory-orientation scores remained in the complete-case model; knee-extension strength did not. The full BESTest total score was not itself a significant univariate predictor. [6]

The reported AUC of 0.78 (95% CI 0.68–0.89) was calculated in the development sample. At the reported classification rule, sensitivity for fallers was only 21.7%, despite 98.2% specificity and 85.1% overall classification. Thus apparently good overall accuracy mostly reflected identification of non-fallers. Twenty-three events, numerous screened candidate variables, stepwise selection and lack of independent validation create substantial optimism risk. Including dropouts as non-fallers removed sensory orientation from the final model, further showing sensitivity to missing-outcome handling. The evidence supports attention to these domains, not a stand-alone validated probability calculator. [6]

Younger age in this cohort should not be turned into a general rule that older adults are safer. Activity exposure was not fully represented; a fitter or more active patient may face more opportunities to fall. Similarly, an association between pain and later falls does not show that reducing pain by a specified amount will reduce falls by the reported odds ratio. A multifactorial causal question requires different evidence from a risk-factor model.

Preoperative risk and longitudinal surveillance

Swinkels and colleagues followed 99 primary TKA recipients using monthly diaries before and for a year after surgery. A recent fall history and higher preoperative depressive-symptom scores were independently associated with postoperative falling. The widely repeated 45.8% figure concerns the subset who had already fallen preoperatively, not every TKA patient. Full methods were not recovered here, so detailed effect estimates and model calibration cannot be confirmed. Its prospective surveillance is nevertheless stronger evidence for taking a falls history than cross-sectional separation of high- and low-scoring balance groups. [17]

The larger 2022 cohort collected preoperative 12-month recall and postoperative fall calendars with reminder calls and assessments at six and twelve months. The denominators declined from 253 to 244 and 214. Fallers represented 40.3% during the retrospective preoperative year, 13.1% during the first six postoperative months and 23.4% during the subsequent six months. Comparing these percentages as a like-for-like annual reduction would be incorrect: the windows and ascertainment methods differ. Group comparisons of single-leg stance and strength do not produce a validated prospective cutoff. The paper also contains discrepancies in descriptive counts and classification totals; these should prevent overprecise secondary claims. [20]

Other studies provide important context but weaker direct support for a specific balance test. Tsonga and colleagues examined change in falls over the year after TKA, while an Osteoarthritis Initiative study examined preoperative predictors after hip or knee arthroplasty. Mixed-joint inference must remain labeled, and retrospective recall or interval outcome collection must not be described as continuous prospective fall recording. Across this literature, an observed reduction in falling after surgery is compatible with residual risk and does not imply that any isolated balance score certifies safe community mobility. [18, 19]

The 2017 unilateral OA-TKA cohort of 376 patients adds a longer horizon. Falls were recalled for the preoperative year and collected quarterly by telephone after surgery, with 321 and 350 attending the first- and second-year assessments. Contralateral radiographic OA severity was associated with falling in adjusted analyses. However, postoperative event counts were small, some odds ratios had very wide confidence intervals, and there was no independent validation or nonsurgical control. Balance-function comparisons at the yearly visit are not automatically preoperative prognostic tests. The study reinforces attention to the other knee, while leaving the causal effect of surgery, rehabilitation and fall-prevention advice unresolved. [22]

Another prospective cohort enrolled 267 older adults after elective primary knee replacement and reported 102 fallers, 200 falls and 40.6% falling over twelve months. The reported percentage implies a smaller analyzed denominator than enrollment; the inaccessible full original is needed to resolve the exact flow. The official abstract reports adjusted associations for prior falls, central-nervous-system medication count and baseline walking-aid use. This is substantive evidence that future falls involve factors beyond a standing score. Its full ascertainment protocol, covariate set, missingness handling and validation could not be inspected, so the estimates are contextual rather than a deployable risk rule. Medication use and aid use are markers in an observational model, not reasons to stop a medication or remove an aid. [23]

Inpatient events and discharge decisions

Immediate postoperative risk depends on exposure and care pathways. A 2024 institutional study found 39 patients with inpatient falls among 6,472 primary TKAs staying at least one night. Revisions, simultaneous bilateral procedures and same-day discharges were excluded. Its 0.6% inpatient proportion is not comparable with a six-month community incidence. Falls clustered early and included loss of balance, buckling, unsteadiness and vasovagal events. Analgesia, supervision, orthostatic symptoms and mobility assistance must therefore be documented alongside physical tests. Observational comparisons involving nerve blocks or opioid consumption cannot isolate medication causality. [21]

Pua and colleagues’ Wii Balance Board study is a proof of concept for a different endpoint. Eighty-nine patients were assessed on postoperative day four, with two unsupported 30-second standing trials. Mediolateral CoP variability added information to an internally bootstrapped model of the walking aid prescribed by experienced therapists masked to the board result. The optimism-corrected concordance index was 0.74. This is concurrent classification of a clinician’s prescription, not prospective demonstration that an aid prevents falls. The text and table also disagree about aid-category counts; the model should not be reimplemented from prose alone. [12]

A later prospective single-leg force-plate study of 40 TKA candidates found no association between preoperative sway and discharge destination; only five participants went to a skilled nursing facility. The unadjusted odds ratio was 0.82 (95% CI 0.27–2.11) per standardized sway increment. Small event numbers and wide uncertainty mean that the null finding does not prove balance is irrelevant. It does show that a proprietary normalized balance score should not be assumed to predict rehabilitation need merely because it produces a T-score. [13]

Table 4 Balance prognosis evidence by endpoint

[6, 12, 13, 17, 20, 21]. AUC/discrimination, calibration and clinical utility are separate; association is not proof that modifying a factor prevents falls.

Table 4 Balance prognosis evidence by endpoint
Study and timingEndpointEvidenceClinical use boundary
Chan 2018; predictors at 4 weeks, n = 134Falls during next 6 months; 23 fallersApparent AUC .78; sensitivity 21.7%, specificity 98.2%Development model; no independent validation; missingness changes selected factors
Swinkels 2009; n = 99Monthly diaries for 1 yearPreoperative falls history and GDS associated with later fallsAbstract-only detailed model; not a standing-score cutoff
Blasco 2022; n = 253→244→21412-month preoperative recall; two postoperative 6-month periods40.3%, 13.1%, 23.4% respectivelyDifferent windows/methods; group comparisons not validated prediction
Pua 2015; day 4, n = 89Concurrent walking-aid prescriptionBootstrap-corrected c-index.74Predicts a therapist decision, not future falls or aid benefit
Biometric 2024; preoperative, n = 40Skilled-nursing discharge, 5 eventsUnadjusted OR .82 (.27–2.11)Imprecise null; proprietary score not a proven triage tool
Inpatient 2024; n = 6472 overnight staysIn-hospital falls; 39 fallers0.6% patient proportionDifferent exposure from community falls; observational care-pathway context

Instrumented balance and technology

Instrumented standing is most useful when a specific question survives standardization: does the patient unload the operated side, depend strongly on vision, or use a different sway strategy under identical conditions? Preserve the raw units and processing pipeline. A platform reporting CoP velocity in cm/s, another reporting an arbitrary stability index, and a phone reporting trunk acceleration do not measure identical quantities. Validation requires agreement for the actual variable and task in the intended population, not merely a significant correlation with another device.

A recent Wii reliability study reports moderate-to-good reliability for several CoP variables over three to seven days, but its abstract conflicts internally about whether the weakest mediolateral displacement reliability belonged to eyes-open or eyes-closed stance. It also lists sex counts that do not match its stated sample size. Because the full original was not recovered, condition-specific SEM/MDC values should not be deployed. This is a precise access-and-reporting limitation, not a claim that Wii-based assessment is inherently unreliable. [11]

The Pua study used 40-Hz acquisition, low-pass filtering at 6.25 Hz and self-selected comfortable stance. Its chosen mediolateral variability cannot inherit measurement error from an unrelated stance width, sampling pipeline, or healthy validation experiment. Nor can a consumer game score inherit the evidence for research-grade extraction from the same board. A low-cost device may be clinically useful, but hardware identity alone does not establish software equivalence. [12]

Reactive-balance laboratory studies deserve particular caution. A tether-release or platform perturbation imposes a specific direction, magnitude and expectation. Repeated trials can change anticipatory preparation. Harnesses and trained personnel alter safety and feasibility. The available originals or abstracts suggest that age, timing and perturbation direction affect response, but this review did not verify a portable postoperative TKA reactive-balance MDC, patient-anchored MIC, or externally validated falls threshold. That bounded conclusion leaves open evidence inaccessible through the retrieved sources. [14–16]

Stage specific assessment strategy

Preoperative assessment

Establish the indication and surgery plan, both knees’ symptoms and previous procedures, falls during a clearly defined interval, current aids, confidence, relevant sensory or neurological problems, and the patient’s priority activities. Use a feasible standing and functional-balance baseline, but label it as preoperative OA. If a single-leg or reach task cannot be completed, retain the reason and support used. Preoperative reliability estimates do not become postoperative error limits automatically. The purpose is both baseline description and identification of issues requiring clinical evaluation; no single score should determine surgical candidacy. [1, 7, 17]

Inpatient and early outpatient recovery

First establish clinical readiness and follow the treating team’s restrictions. Record day after surgery, current analgesia or block context when known, symptoms on standing, wound or postoperative complications, weight-bearing status, and required assistance. Guarding and access to support take priority over reproducing an unsupported research task. Inability or an unsafe attempt is a meaningful result. A bedside ability to stand or transfer does not establish recovery of community reactive balance. Repeat the same protocol only when clinically appropriate; otherwise document the change in testing conditions rather than hiding it. [12, 21]

In the early outpatient weeks, select a multidomain tool that the patient can attempt safely and pair it with a concise mobility measure. The Chan data support measurement across this recovery period in selected aid-free patients, but its later-stage retest error is not a perfect estimate of early day-to-day variability. Ask about new near-falls, buckling, environmental hazards and activity exposure. A more demanding test may become appropriate as basic tasks reach ceiling; switching tests should begin a new series rather than create a false continuous trajectory. [2–4, 6]

Later recovery and persistent deficits

At three to six months and beyond, consider more challenging sensory, stepping or reach tasks when basic scales cease to discriminate. Assess both limbs where appropriate and note whether contralateral OA limits interpretation. A patient can improve relative to a severely impaired preoperative baseline without reaching age-relevant function. Small longitudinal cohorts and reviews demonstrate heterogeneous recovery; they do not establish a universal date at which balance is normal. A 2025 mixed THA/TKA cohort tracked Tinetti and TUG over a year, but its TKA subgroup was small and it did not measure future falls as the outcome. Its score changes must not be relabeled as observed changes in fall incidence. [24–26]

Practical reporting and rehabilitation tool design

The minimum record should contain surgery date and indication; unilateral, simultaneous bilateral, staged or revision status; operated side and tested side; contralateral disease; task version; footwear; visual and surface condition; arm position; support and guarding; practice and scored trials; timing or scoring rules; pain and relevant symptoms; and completion status. For a platform, additionally store sampling frequency, filtering, trial duration, CoP variable definition and units. Preserve individual trials as well as the selected summary.

Present four independent judgments: observed performance, whether a comparable prior test exists, whether change exceeds a genuinely matched error estimate, and whether any important-change or prognosis evidence is applicable. Show the source population and stage beside a threshold. Display “protocol mismatch” or “no directly applicable estimate verified” rather than borrowing a knee-OA, THA, neurological or healthy-adult cutoff. For the Mini-BESTest, require the scale denominator explicitly. For the full BESTest, distinguish raw points from percentage points.

Prospective risk communication should emphasize actionable context rather than a binary clearance. A reasonable summary is that balance and mobility have improved under the tested conditions, while falls surveillance and clinical judgment remain necessary. An application should not generate an individualized fall probability from an unvalidated development model, certify safety from a normal sway score, or infer that a treatment caused recovery from serial observational measurements. The practical value of assessment is clearer description, safer comparison and more focused clinical reasoning, not unsupported numerical certainty.

Table 5 Minimum balance protocol record

Recommended reporting framework derived from the protocol heterogeneity in the appraised originals; it is not a new validated clinical rule.

Table 5 Minimum balance protocol record
DomainRecord before comparing visits
Patient and surgeryIndication, procedure type, surgery dates, operated and tested sides, contralateral disease, relevant comorbidity
Current statePain, fatigue, symptoms on standing, relevant analgesia/block context, restrictions, aid and guarding
Clinical taskScale version and denominator, surface, vision, footwear, arm position, stance width, direction, task instructions
Trial handlingPractice trials, number of scored attempts, best/mean rule, time cap, termination rule, inability and reason
Instrumented taskDevice/software, calibration, sampling, filter, duration, CoP variable and units, support and foot placement
InterpretationProtocol match, stage match, raw change, matched error, anchor-specific importance and independently justified prognosis

Conclusions

Direct TKA research supports multidomain clinical balance testing, selected single-leg and stepping measures, and carefully standardized instrumented standing. The evidence is strongest for specific protocols and recovery stages. BESTest-family tests offer advantages over BBS when ceiling effects matter, but score-version differences and study design prevent indiscriminate threshold reuse. Direct falls cohorts support multifactorial assessment, yet the best-described postoperative model remains a small development analysis with low sensitivity at its reported classification rule. Reliable measurement, meaningful change and future risk should remain separate outputs. The accompanying tables and primary-study matrix provide the conditions under which particular numerical results may be used.

Primary study characteristics

Primary studies are grouped by measurement or prognostic question. Population, protocol, endpoint and source access constrain interpretation. The bibliography identifies access limitations. Related publications from one cohort are not independent replications.

Table 6 Clinical balance measurement

Table 6 Clinical balance measurement
Study and populationProtocol and timingMain findingsInterpretive limits and source
Chan AC 2015 [2]
Reliability and longitudinal construct validity
92 first OA-TKA recipients; reliability 46 ≥6 months, interrater 25; separate validity cohort 46 at 2 and 12 weeks, 45 at 24 weeks after one relocation; aid-free testing
BESTest percentage; Mini 32; Brief 24; BBS 56; FGA 30. Shared items scored once. Same-rater within 1 week, no interim physiotherapy; concurrent three-rater scoringICC(2,1) retest: BEST .96(.93–.98), Mini .92(.87–.96), Brief .94(.90–.97), BBS .97(.94–.98), FGA .97(.95–.98). SEM/MDC 95:2.24/6.22 percentage points; 1.34/3.71; 1.15/3.19; 0.72/2.00; 0.94/2.59
BBS maximum 52.2% at 12 wk (n = 46), 57.8% at 24 wk (n = 45). Shared constructs correlate; no falls prognosis
Interrater scoring not independent readministration. Later-stage cohort with no reported prior-year fallers; prior falls were not stated as an exclusion criterion. Mini 32 and BEST percentage cannot inherit Mini 28/raw 108 MIC. Sensory subsection zero error is not universally error-free
Source: Chan 2015 accepted manuscript pp 7–18 and Table 2 p 33
Chan ACM 2018 [3]
Prospective responsiveness cohort
134 completed; first OA TKA; 50–85 yr; 2/4/8/12/24 wk
BEST/Mini/Brief/BBS; FGA external comparator; repeated assessments during rehabilitationSRM ranges BEST .60–1.14; Mini .40–.94; Brief .27–.91; BBS .19–.70
Changes in BEST and Mini associated with same-interval FGA changes more consistently than BBS
Change association is not future prediction; SRM is not MIC. Related Chan 134-person cohort, not independent replication
Source: Chan 2018 accepted manuscript Abstract, Methods and Results, Tables
Chan ACM 2020 [4]
Anchor and distribution-based important-change study
146 recruited, 134 complete; 2→4 wk after first OA TKA; no aids during testing
BEST raw 108; Mini 28; Brief 24; BBS 56. Anchor FGA change≥4; 72 anchor-positiveReported integer anchor thresholds 8/2/3/5 points; distribution 6/1/2/2. AUC BEST .811(.739–.883); Mini .782(.704–.860); Brief .706(.618–.795); BBS .586(.490–.682)
Anchor-change r .551/.516/.402/.153. BBS range withheld by authors due to poor discrimination
FGA 4-point importance borrowed from older adults, not patient anchor validated in TKA. Ranges are not CIs. Rounded integer/% mismatch. Scale mismatch with 2015 MDC. Related cohort
Source: Chan 2020 accepted manuscript Methods, Tables 2–4; Table 4 p 24 visually checked
Sarac DC 2022 [5]
One-day same-rater test-retest and concurrent validity
41 simultaneous bilateral primary OA-TKA recipients, seventh month; independent standing/walking
SLST dominant leg, eyes open, arms crossed, nonstance knee 90°; practice; TUG/10 MWT/2 MWT/5 STSTwo-way mixed ICC unspecified unit/agreement detail: SLST .72(.47–.85), SEM 7.08 s, MDC 95 19.62 s; mean 13.96→19.47 s
SLST correlations: BBS .704; FES-I−.540. Not prospective falls
Bilateral dominant-leg result, not unilateral operated-leg threshold. Large error/retest shift; BBS ceiling. TUG Table 4 MDC 3.04 vs discussion 3.4 and prose correlation summary discrepancies
Source: Sarac 2022 Methods and Results, Tables 3–4
Bin Sheeha B 2024 [7]
Preoperative retest plus postoperative responsiveness
Selected unilateral KL 4 OA scheduled for TKA; 35 retest; 6/12 mo follow-up; extensive comorbidity exclusions
Barefoot affected-leg stance; three normalized reach directions; 4 practice+mean 3; preoperative 7 day retestTwo-way mixed absolute-agreement ICC .993–.998; reported SEM .37–.68 and MDC 1.02–1.89 normalized percentage points
Postoperative responsiveness and functional-change correlations; no patient-anchored MIC
Reliability was BEFORE TKA. Participant-flow arithmetic inconsistent; original equation rendering and direction-specific tables required for implementation. Selected completers
Source: SEBT 2024 Methods 2.1–2.6, Results/Table 1
Eymir M 2022 [8]
Test-retest and concurrent validity
56 TKA; exact stage not verified from accessible abstract
Step Test; exact step height/duration/support/trial summary not recoveredAbstract ICC .90, SEM .76, MDC 95 2.11; full ICC model not verified
r−.69 TUG, −.67 10 MWT; no verified prospective outcome
Abstract-only; threshold not deployable until protocol obtained; no MIC or falls cutoff inference
Source: Official PMID 35022951; full article unavailable
Özcan D 2025 [9]
Same-day reliability and concurrent validity
30 TKA; stage/protocol details not fully recovered
Dual-task FSST, two same-day trialsICC(2,1) .97; MDC 95 3.43 s from abstract
r .65 with single-task TUG, −.40 HSS knee score
Abstract-only; cognitive task/accuracy/prioritization need verification; no prospective fall-risk evidence
Source: Official PMID 38384122

Table 7 Instrumented and experimental balance

Table 7 Instrumented and experimental balance
Study and populationProtocol and timingMain findingsInterpretive limits and source
Zenooz GH 2025 [11]
Between-session device reliability
Stated 31 TKA; reported 6 men+27 women inconsistent
Wii bipedal stance and functional reach, eyes open/closed; 3–7 daysAbstract ICC .51–.86, outlying.29; SEM/MDC ranges mix outcomes/units
No significant mean retest differences; no verified force-plate criterion agreement or falls prognosis
Abstract contradicts eyes-open/closed mapping of weakest metric. Do not reproduce condition-specific thresholds. Nonsignificant bias does not prove narrow agreement
Source: Official PMID 39663090; full original not recovered
Pua YH 2015 [12]
Concurrent clinical-decision model with bootstrap internal validation
89 TKA inpatients postoperative day 4
Two 30 s barefoot unsupported comfortable-stance trials; 40 Hz, 6.25 Hz low-pass; mean ML CoP SD . Therapist aid prescription masked to boardNo direct reliability experiment in this cohort. Eight a priori predictors, penalized ordinal regression, bootstrap optimism correction
IQR OR 2.55(1.23–5.30) for MLSD .22→.41 cm; corrected c-index .74; calibration plot
Endpoint is concurrent prescribed aid, not future fall. Stance width self-selected. Text aid counts contradict Table 1; no external validation
Source: Pua 2015 Methods, Tables 1–2, Results; complete body
Lee JJ 2024 [13]
Prospective preoperative marker and later outcomes
40 primary TKA, 39 follow-up; 5 discharges to skilled nursing; independently ambulatory/balancing selected
Preoperative operative/nonoperative single-leg force-plate sway; proprietary normative T-scoreNo direct repeatability/MIC evidence
Unadjusted discharge OR .82(.27–2.11) per standardized operative sway; no meaningful evidence for LOS or later function
Only 5 events; wide CI, unadjusted models; null result not proof of absence. Proprietary normative score not automatically transportable
Source: 2024 Biometric study Methods and Results
Pethes Á 2015 [14]
Experimental balance response
Early postoperative knee arthroplasty; details require full original
Sudden unidirectional perturbationOriginal retrieval incomplete
Task-specific reactive balance investigation
No verified clinical MDC/MIC or community-falls risk rule; abstract-only
Source: DOI 10.1016/j.jelekin.2015.02.010
Street BD 2017 [15]
Cross-sectional experimental age comparison
59 participants including 29 unilateral TKR at 6 mo, younger/older controls
Tether-release forward-fall recovery; CoM displacement and step characteristicsNot clinical test-retest validation
Younger and older TKR groups differed in laboratory recovery
Age-group/case-control evidence is not prospective falls; no clinical cutoff validated; original retrieval incomplete
Source: Official PMID 28342974
Cho SD 2013 [24]
Single-limb longitudinal recovery
TKA; selection/timing require full-original verification
Single-limb balanceNo detailed thresholds extracted
Recovery-context source
Original retrieval incomplete; not used to define a universal normal-recovery date
Source: DOI 10.1007/s00167-012-2144-x
Vertesich K 2025 [26]
Small longitudinal mixed-joint recovery cohort
Separate THA/TKA subgroups; primary/secondary OA; selected ability to test
Tinetti, TUG, PROMs before/day 4–6/week 6/month 3/month 12No direct stable retest/MIC
TKA Tinetti 15.9 baseline→23.9 at 12 mo; TUG 14.8→11.3 s
No prospective falls outcome; TKA immediate decline nonsignificant despite broad concluding wording; do not equate with measured fall risk
Source: LongitudinalBalance 2025 Methods and Results 3.2, Table 1

Table 8 Falls and future outcomes

Table 8 Falls and future outcomes
Study and populationProtocol and timingMain findingsInterpretive limits and source
Chan ACM 2018 [6]
Prospective monthly falls surveillance and development model
146 recruited, 134 complete; clinical predictors 4 wk after first OA-TKA; 23 fallers/31 falls during 6 mo
BESTest percent/subsections, proprioception, pain, strength, ROM, ABC; monthly interviewNo new test reliability; backward model after univariate screening
Adjusted OR:age .91(.84–.99); proprioception 1.62(1.05–2.50); pain 1.68(1.07–2.64); sensory .92(.86–.99). Apparent AUC .78(.68–.89); sensitivity 21.7%, specificity 98.2%
Few events/numerous candidate screens; no external validation. Sensory predictor disappears when dropouts labeled nonfallers. Total BEST not significant; association not treatment effect
Source: Chan 2018 falls manuscript pp 13–17, Tables 2–3; Table 3 p 37 visually checked
Swinkels A 2009 [17]
Prospective preoperative and postoperative diaries
99 primary TKA; 1 yr follow-up
Monthly falls diaries; quarterly WOMAC, ABC-UK, GDSNot a performance-test measurement-property study
24.2% fell last preop quarter; 11.7–11.8% in each postop quarter; 45.8% of preop fallers fell again. History/depressive symptoms independent factors
Abstract only; 45.8% not all patients. Full coefficient/model/validation detail unavailable
Source: Official PMID 19029071
Tsonga T 2016 [18]
Longitudinal falls and clinical-factor cohort
Older patients with severe knee OA undergoing TKA
Before/one-year falls comparison and clinical variables; ascertainment must remain study-specificNo validated standing-score error or MIC established
Describes postoperative fall burden and correlates
Recall and changing activity exposure constrain causal surgery-effect inference; not a stand-alone test threshold
Source: Falls 2016 complete body
Riddle DL 2018 [19]
OAI longitudinal preoperative-risk cohort
Hip OR knee arthroplasty recipients; mixed joint inference
Preoperative clinical factors and subsequent reported fallsNot a balance-test reliability study
Preoperative factors examined against later falls
Mixed-joint pooling; interval ascertainment differs from monthly diaries. Do not relabel as TKA-only balance validation
Source: FallsOAI 2018 complete body
José‐María Blasco 2022 [20]
Retrospective preop recall plus prospective postoperative calendars
253 primary TKR candidates; 244 at 6 mo, 214 at 12 mo; selected severe OA
Preop 12 mo recall; postop calendar/reminder calls every 2 mo; clinical tests at baseline, 6/12 moNot a direct measurement-error study
Fallers 102/253 preop, 32/244 first 6 mo, 50/214 next 6 mo. Group differences in operated-limb single-leg stability
Different horizons/ascertainment; no validated predictive cutoff. Descriptive sex/classification totals and some narrative/table statements conflict; avoid precise secondary mechanism claims
Source: Blasco 2022 Methods, Tables 1–3
Lawrence KW 2024 [21]
Retrospective inpatient incident cohort
6472 primary overnight TKA; 39 fallers, 40 falls; revision/bilateral/same-day excluded
Institutional incident database; perioperative factorsNot a balance score validation
0.6% inpatient fall proportion; 2.7 falls/1000 patient-days
Exposure differs from community follow-up. Observational analgesia comparisons not causal; comorbidity included
Source: 2024 InpatientFalls Methods and Results, Tables 1–4
Si 2017 [22]
Longitudinal falls and balance-function cohort
376 unilateral primary OA-TKA patients; 321 at one year and 350 at two years
Preoperative 12-month fall recall; postoperative quarterly telephone collection and yearly confirmation; BBS, TUG, strength and questionnairesNo direct reliability, MDC or MIC experiment
Contralateral KL grade ≥3 associated with falls; 20 and 11 postoperative fallers in first and second years
Few events, wide odds-ratio uncertainty, no independent validation or nonsurgical control; yearly performance comparisons not inherently preoperative prediction; some descriptive statements conflict
Source: FallsBalance 2017 complete original Methods, Tables 3–4
Hill 2022 [23]
Prospective twelve-month falls cohort
267 enrolled older adults after elective primary knee replacement; 102 fallers reported as 40.6%; analyzed denominator needs full original
Baseline history, gait aid and medication count; falls followed after discharge; full collection details unavailableNo standing-test measurement-property estimate
Abstract adjusted OR 2.41(1.35–4.31) for fall history, 1.66(1.25–2.21) for CNS medication count; gait-aid adjusted incidence-rate ratio 2.38(1.57–3.60)
Full original unavailable; exact flow, adjustment, calibration and validation unverified. Associations do not justify withdrawing medicines or aids
Source: Official PMID 34292196; full article unavailable

Search appendix

This is an auditable critical narrative assessment and prognosis review, with targeted original-study appraisal, not a formal systematic review. Search date and cutoff: 2 October 2026. No independent duplicate screening or exhaustive eligibility count is claimed. Domain searches overlap, as do the targeted supplements; their totals must not be summed as unique eligible studies. Source matrices prioritize measurement properties, protocol definitions, recovery and prospective outcomes; treatment-only studies and off-population measurement studies can inform discovery without becoming primary evidence.

("Arthroplasty, Replacement, Knee"[MeSH Terms] OR "total knee arthroplast*"[Title/Abstract] OR "total knee replacement*"[Title/Abstract] OR TKA[Title/Abstract]) AND (balance[Title/Abstract] OR postur*[Title/Abstract] OR sway[Title/Abstract] OR "single leg"[Title/Abstract] OR "single-leg"[Title/Abstract] OR "functional reach"[Title/Abstract] OR falls[Title/Abstract] OR "force platform"[Title/Abstract]) AND (reliab*[Title/Abstract] OR valid*[Title/Abstract] OR responsiv*[Title/Abstract] OR "measurement error"[Title/Abstract] OR "minimal detectable"[Title/Abstract] OR "minimal important"[Title/Abstract] OR prognos*[Title/Abstract] OR predict*[Title/Abstract] OR longitudinal[Title/Abstract] OR recovery[Title/Abstract] OR "time course"[Title/Abstract]) AND ("1800/01/01"[Date - Publication] : "2026/10/02"[Date - Publication])

Provider total: 566. Unique connector PMIDs: 340. Official NCBI ESearch identified 566 IDs and EFetch recovered 566 records, with 0 missing from the official set. This reconciliation was necessary because successive connector pages repeated records. A returned final page was not treated as evidence of complete unique coverage.

TITLE-ABS-KEY(("total knee arthroplast*" OR "total knee replacement*" OR TKA) AND ("standing balance" OR "postural balance" OR "postural control" OR "postural sway" OR "single leg stance" OR "single-leg stance" OR "functional reach" OR falls OR "force platform" OR "balance test*") AND (reliab* OR valid* OR responsiv* OR "measurement error" OR "minimal detectable" OR "minimal important" OR prognos* OR predict* OR longitudinal OR recovery OR "time course")) AND PUBYEAR < 2027

Provider total 304; 13 archived pages; 304 unique Scopus source IDs. No missing records were apparent from the returned-page unique-count comparison, but this is not an independent database replay. The publication-year filter through 2026 is broader than the October 2 cutoff. Publication timing was checked for retained current sources where necessary; discovery results are not all treated as eligible. Individual Scopus detail access was not presumed from successful search access.

("total knee arthroplasty" OR "total knee replacement") AND ("reactive balance" OR perturbation OR "fall risk" OR "prospective falls" OR "single-limb balance") AND ("1800/01/01"[Date - Publication] : "2026/10/02"[Date - Publication])

Two connector pages were archived. Connector unique PMIDs 67 versus official ESearch total 76; official EFetch recovered 76 with no missing IDs. Supplemental records overlap with the principal query. For balance the supplement broadens reactive/perturbation/falls terminology; for strength it broadens power, activation, hip-abduction and rapid-force terminology without a measurement-property filter.

Primary elective OA-TKA is the main population. Preoperative OA reliability is labeled separately from postoperative reliability. Simultaneous bilateral, revision, UKA and mixed THA/TKA studies are not silently generalized to unilateral primary TKA. Operative ligament or gap balancing is not postural balance. Cross-sectional or repeated contemporaneous regressions are not future prognosis. Reviews and guidelines are contextual sources, not independent primary-study replications.

Access and numerical verification

The bibliography distinguishes complete articles, partially available tables, abstract-only evidence and unavailable full articles. Original articles were sought through bibliographic databases, official PMC and NCBI records, publishers and public accepted-manuscript repositories. Selected consequential estimates were checked against original tables. This does not constitute independent appraisal of every paper.

Remaining access gaps are material. Abstract-level estimates with unverified stage, positioning, ICC model or trial aggregation stay non-deployable. A complete article body can omit image-only tables or supplements; that limitation is preserved rather than inferred away. All delivered DOI/PMID references were reconciled to verified metadata; unsuccessful retrievals were not treated as source evidence. Publication year follows the journal issue when verified, with online dates retained in archived metadata.

Shared cohorts and interpretation

Chan 2018 responsiveness, Chan 2020 important change and Chan 2018 falls have matching recruitment, hospital and 134-person complete samples, consistent with a shared cohort. They do not provide three independent replications. QAB 2018, Hip 2017 and the gait report's Kittelson four-metre analysis arise from NCT01537328. Cross-report outcomes remain useful but are not independent participant samples. No neurological, healthy-athlete, nonsurgical OA or THA threshold is imported as a postoperative TKA clinical boundary.

References

References are numbered in first citation order. Source descriptions identify the material examined and do not constitute a study quality rating. Links identify the original publication or the named primary source version.

1. Jette, Diane U, Hunter, Stephen J, Burkett, Lynn, Langham, Bud, Logerstedt, David S, Piuzzi, Nicolas S, et al. Physical Therapist Management of Total Knee Arthroplasty. Physical therapy. 2020. DOI 10.1093/ptj/pzaa099 Source examined: Complete guideline; recommendation context.

Source note: SRC-27ba0555a48c Jette 2020

2. Chan AC, Pang MY. Assessing Balance Function in Patients With Total Knee Arthroplasty. Physical therapy. 2015;95(10):1397-407. DOI 10.2522/ptj.20140486 Source examined: Complete accepted manuscript; original Table 2 visually checked.

Source note: SRC-8ee9917f9259 Chan AC 2015

3. Chan ACM, Ouyang XH, Jehu DAM, Chung RCK, Pang MYC. Recovery of balance function among individuals with total knee arthroplasty: Comparison of responsiveness among four balance tests. Gait & posture. 2018;59:267-271. DOI 10.1016/j.gaitpost.2017.10.020 Source examined: Complete accepted manuscript.

Source note: SRC-024c8e7c8f48 Chan ACM 2018

4. Chan ACM, Pang MYC, Ouyang H, Jehu DAM. Minimal Clinically Important Difference of Four Commonly Used Balance Assessment Tools in Individuals after Total Knee Arthroplasty: A Prospective Cohort Study. PM & R : the journal of injury, function, and rehabilitation. 2020;12(3):238-245. DOI 10.1002/pmrj.12226 Source examined: Complete accepted manuscript; original Table 4 visually checked.

Source note: SRC-b3ddfcea9869 Chan ACM 2020

5. Sarac DC, Unver B, Karatosun V. Validity and reliability of performance tests as balance measures in patients with total knee arthroplasty. Knee surgery & related research. 2022;34(1):11. DOI 10.1186/s43019-022-00136-4 Source examined: Complete article and tables.

Source note: SRC-7757dfb24f28 Sarac DC 2022

6. Chan ACM, Jehu DA, Pang MYC. Falls After Total Knee Arthroplasty: Frequency, Circumstances, and Associated Factors-A Prospective Cohort Study. Physical therapy. 2018;98(9):767-778. DOI 10.1093/ptj/pzy071 Source examined: Complete accepted manuscript; original Table 3 visually checked.

Source note: SRC-7527b26a66ec Chan ACM 2018

7. Bin Sheeha B, Bin Nasser A, Williams A, Granat M, Johnson DS, Althomali OW, et al. Reliability of the Star Excursion Balance Test with End-Stage Knee Osteoarthritis Patients and Its Responsiveness Following Total Knee Arthroplasty. Journal of clinical medicine. 2024;13(21). DOI 10.3390/jcm13216479 Source examined: Complete article and tables; formula rendering needs original equation check before implementation.

Source note: SRC-00d351b30ecb Bin Sheeha B 2024

8. Eymir M, Yuksel E, Unver B, Karatosun V. Reliability, validity, and minimal detectable change of the Step Test in patients with total knee arthroplasty. Irish journal of medical science. 2022;191(6):2651-2656. DOI 10.1007/s11845-021-02888-6 Source examined: Abstract verified; full article unavailable.

Source note: SRC-611a7c5db50e Eymir M 2022

9. Özcan D, Unver B, Karatosun V. Balance assessment under dual task conditions in patients with total knee arthroplasty: a test-retest reliability and concurrent validity study. Physiotherapy theory and practice. 2025;41(1):93-98. DOI 10.1080/09593985.2024.2321222 Source examined: Abstract verified; full article unavailable.

Source note: SRC-56ede5232a4c Ozcan D 2025

10. Unver B, Sevik K, Yarar HA, Unver F, Karatosun V. Reliability of the Modified Four Square Step Test (mFSST) in patients with primary total knee arthroplasty. Physiotherapy theory and practice. 2021;37(4):535-539. DOI 10.1080/09593985.2019.1633713 Source examined: Abstract verified in official PubMed; detailed protocol unavailable.

Source note: SRC-e4b44b007704 Unver B 2021

11. Zenooz GH, Taheriazam A, Arab AM, Mokhtarinia H, Rezaeian T, Hosseinzadeh S, et al. Reliability of the Wii balance board for static and dynamic balance assessment in total knee arthroplasty patients. Journal of bodywork and movement therapies. 2025;41:21-28. DOI 10.1016/j.jbmt.2024.09.005 Source examined: Abstract verified; full article unavailable; abstract internal inconsistencies.

Source note: SRC-cdfc3484fa39 Zenooz GH 2025

12. Pua YH, Clark RA, Ong PH. Evaluation of the Wii Balance Board for walking aids prediction: proof-of-concept study in total knee arthroplasty. PloS one. 2015;10(1):e0117124. DOI 10.1371/journal.pone.0117124 Source examined: Complete article and tables.

Source note: SRC-83381c331695 Pua YH 2015

13. Lee JJ, Arora P, Finlay AK, Amanatullah DF. A balance focused biometric does not predict rehabilitation needs and outcomes following total knee arthroplasty. BMC musculoskeletal disorders. 2024;25(1):473. DOI 10.1186/s12891-024-07580-1 Source examined: Complete article and tables.

Source note: SRC-9c6c5bb401a9 Lee JJ 2024

14. Pethes Á, Bejek Z, Kiss RM. The effect of knee arthroplasty on balancing ability in response to sudden unidirectional perturbation in the early postoperative period. Journal of electromyography and kinesiology : official journal of the International Society of Electrophysiological Kinesiology. 2015;25(3):508-14. DOI 10.1016/j.jelekin.2015.02.010 Source examined: Retrieval incomplete; abstract-level interpretation only.

Source note: SRC-b286feba60b2 Pethes A 2015

15. Street BD, Gage W. After total knee replacement younger patients demonstrate superior balance control compared to older patients when recovering from a forward fall. Clinical biomechanics (Bristol, Avon). 2017;44:59-66. DOI 10.1016/j.clinbiomech.2017.03.006 Source examined: Retrieval incomplete; abstract-level interpretation only.

Source note: SRC-0287f727aa06 Street BD 2017

16. Gage WH, Frank JS, Prentice SD, Stevenson P. Postural responses following a rotational support surface perturbation, following knee joint replacement: frontal plane rotations. Gait & posture. 2008;27(2):286-93. DOI 10.1016/j.gaitpost.2007.04.006 Source examined: Abstract verified; original protocol unavailable.

Source note: SRC-e5569215f7e0 Gage WH 2008

17. Swinkels A, Newman JH, Allain TJ. A prospective observational study of falling before and after knee replacement surgery. Age and ageing. 2009;38(2):175-81. DOI 10.1093/ageing/afn229 Source examined: Abstract verified; full article unavailable.

Source note: SRC-a97318d926fb Swinkels A 2009

18. Tsonga T, Michalopoulou M, Kapetanakis S, Giovannopoulou E, Malliou P, Godolias G, et al. Reduction of Falls and Factors Affecting Falls a Year After Total Knee Arthroplasty in Elderly Patients with Severe Knee Osteoarthritis. The open orthopaedics journal. 2016;10:522-531. DOI 10.2174/1874325001610010522 Source examined: Complete article and tables.

Source note: SRC-e4889e845e7b Tsonga T 2016

19. Riddle DL, Golladay GJ. Preoperative Risk Factors for Postoperative Falls in Persons Undergoing Hip or Knee Arthroplasty: A Longitudinal Study of Data From the Osteoarthritis Initiative. Archives of physical medicine and rehabilitation. 2018;99(5):967-972. DOI 10.1016/j.apmr.2017.12.030 Source examined: Complete article and tables; mixed hip and knee.

Source note: SRC-2ec8da251a39 Riddle DL 2018

20. José‐María Blasco, Jose Pérez-Maletzki, Beatriz Díaz-Díaz, Antonio Silvestre-Muñoz, Ignacio Martínez-Garrido, Sergio Roig‐Casasús. Fall classification, incidence and circumstances in patients undergoing total knee replacement. Scientific Reports. 2022. DOI 10.1038/s41598-022-23258-x Source examined: Complete article and tables; descriptive inconsistencies retained.

Source note: SRC-0634d5b5c422 JoseMaria Blasco 2022

21. Lawrence KW, Link L, Lavin P, Schwarzkopf R, Rozell JC. Characterizing patient factors, perioperative interventions, and outcomes associated with inpatients falls after total knee arthroplasty. Knee surgery & related research. 2024;36(1):11. DOI 10.1186/s43019-024-00215-8 Source examined: Complete article and tables.

Source note: SRC-f1823cac2ed1 Lawrence KW 2024

22. Si, Hai-Bo, Zeng, Yi, Zhong, Jian, Zhou, Zong-Ke, Lu, Yan-Rong, Cheng, Jing-Qiu, et al. The effect of primary total knee arthroplasty on the incidence of falls and balance-related functions in patients with osteoarthritis. Scientific reports. 2017. DOI 10.1038/s41598-017-16867-4 Source examined: Complete original article and tables.

Source note: SRC-2f17bd164f3b Si 2017

23. Hill, Anne-Marie, Ross-Adjie, Gail, McPhail, Steven M, Jacques, Angela, Bulsara, Max, Cranfield, Alexis, et al. Incidence and Associated Risk Factors for Falls in Older Adults After Elective Total Knee Replacement Surgery: A Prospective Cohort Study. American journal of physical medicine & rehabilitation. 2022;101(5):454-459. DOI 10.1097/phm.0000000000001848 Source examined: Official abstract verified; full article unavailable.

Source note: SRC-11cc6a47b085 Hill 2022

24. Cho SD, Hwang CH. Improved single-limb balance after total knee arthroplasty. Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA. 2013;21(12):2744-50. DOI 10.1007/s00167-012-2144-x Source examined: Retrieval incomplete; abstract-level interpretation only.

Source note: SRC-a671e7328c50 Cho SD 2013

25. Moutzouri, M, Gleeson, N, Billis, E, Tsepis, E, Panoutsopoulou, I, Gliatis, J. The effect of total knee arthroplasty on patients' balance and incidence of falls: a systematic review. Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA. 2017. DOI 10.1007/s00167-016-4355-z Source examined: Complete review; context and citation discovery.

Source note: SRC-6f0e12ae988e Moutzouri 2017

26. Vertesich K, Staats K, Schneider E, Willegger M, Windhager R, Böhler C. Balance and Mobility in Comparison to Patient-Reported Outcomes-A Longitudinal Evaluation After Total Hip and Knee Arthroplasty. Journal of clinical medicine. 2025;14(12). DOI 10.3390/jcm14124135 Source examined: Complete article and tables; mixed hip and knee.

Source note: SRC-084a7beef7f0 Vertesich K 2025