In this report
Audited and Updated
Current annotated reading updated 4 October 2026 (Australia/Brisbane). Scoped AI-assisted narrative source review; not independently human-adjudicated.
LBP06
Clarify that 49 is the analysed/completer cohort and mention recruitment/attrition when describing study validity. The original reports say 49 participants/adults and do not explicitly say 49 enrolled. Preserve the correct change–change interpretation; this is not an established report recruitment or transcription error.
Type: cohort denominator precision. Audit disposition: supported.
LBP09
State 20 healthy enrolled but19 analysed and20 CNSLBP analysed. Update access note to accepted manuscript/Table3; retain final-abstract attribution for combined ranges. Do not merge manuscript subgroup estimates into a universal cutoff.
Type: accepted version cohort and range precision. Audit disposition: supported with version caveat.
Remaining limit: Appendices/finalversiontables not recovered at snapshot.
LBP10
Replace the inaccessible-processing note with the original acquisition/onset information while preserving the sex-restricted case-control interpretation.
Type: current access update protocol. Audit disposition: supported.
LBP11
The 55 included/11 dropouts/44 analysed flow can now be cited to the original author-posted article, rather than only to the later review. Keep four-week-versus-baseline interpretation qualified.
Type: current access update provenance. Audit disposition: supported with bounded scope.
Remaining limit: No reanalysis of full multivariable models;55 total is supported by 44+11 flow.
Editorial record
- Audit status: supported with version caveat. Appendices/finalversiontables not recovered at snapshot.
- Edited phrase under LBP09 . Original wording: P0079
- Edited phrase under LBP09 . Original wording: P0079
- Audit status: supported.
- Audit status: supported with bounded scope. No reanalysis of full multivariable models;55 total is supported by 44+11 flow.
- Audit status: supported with version caveat. Appendices/finalversiontables not recovered at snapshot.
- Edited phrase under LBP09 . Original wording: P0233
- Edited phrase under LBP09 . Original wording: P0235
- Edited phrase under LBP09 . Original wording: P0237
- Audit status: supported.
- Audit status: supported with bounded scope. No reanalysis of full multivariable models;55 total is supported by 44+11 flow.
- Edited phrase under LBP11 . Original wording: P0305
- Audit status: supported.
Executive assessment
Strength testing can contribute useful, repeatable information to low-back-pain rehabilitation when the task and intended construct are explicit. Externally stabilized dynamometry is the most direct practical route to quantifying voluntary force or torque capacity. Endurance holds, repeated chair rises, lifting, movement-control tests, mechanical power and rate of force development measure different capabilities. They should not be collapsed into a generic core-strength score.
The older conclusion that practicable trunk-strength measurement lacks symptomatic validation now needs qualification. A 2023 original study provided direct within-day and between-day evidence for an externally fixed handheld dynamometer in 20 adults with chronic nonspecific LBP. Supported positions had useful repeatability and lower absolute error than some standing procedures. However, device correlations did not establish interchangeability, and a 2025 longitudinal responsiveness study found that HHD change scores correlated only weakly with changes measured using an isokinetic dynamometer. Positive measurement findings and unresolved change validity coexist. [1‑5]
For individual follow-up, absolute error and learning matter as much as the ICC. The same named test can produce different values when fixation, lever arm, body position, contraction instructions or trial selection change. The Biering-Sørensen test is a meaningful sustained performance task, but its duration is influenced by upper-body loading, hip and trunk capacity, symptoms, self-efficacy and task engagement. It is neither isolated lumbar endurance nor maximal strength. [6‑13]
Direct RFD evidence is sparse and heterogeneous. A small female LBP case-control study found lower extension RFD, whereas trained young adults with mild acute nonspecific LBP did not differ from controls during an isometric deadlift. These results address different populations and tasks. A recent chair-transfer pilot used early ground-reaction-force slopes and cross-sectional mediation; it did not establish a causal pathway or predict future LBP. [14‑17]
The clinical-outcome prognosis literature does not justify an individual recovery score from trunk strength or endurance alone. Historical prevention findings, current group differences, predictions of same-session endurance, changes after training and true prospective pain or disability prediction must be separated. The 2025 prognostic review found low- or very-low-quality evidence for the physical measures it could synthesize. [18‑24]
For rehabtools, the strongest initial use is standardized measurement with explicit units, symptoms, completion, protocol and source-matched error information. A strength gain can be worth measuring without being a validated surrogate for pain relief, work ability or recurrence prevention. This report provides evidence interpretation and tool-design guidance, not individual exercise prescription.
Scope and search approach
The main population is adults with chronic nonspecific or primary LBP. Acute, subacute, recurrent and persistent presentations are retained separately. Radicular symptoms, spinal stenosis, lumbar surgery and specific spinal disease are not pooled into the principal inference. Studies in healthy adults, young athletes and adolescent prevention populations are included only where they clarify a boundary or the history of a frequently repeated claim.
Native PubMed measurement, prognosis and protocol or technology searches returned 187, 549 and 646 records; corresponding Scopus searches returned 432, 1,503 and 1,587. Every query was paged and reconciled, including official NCBI recovery of records missing from connector pagination. The streams overlap. Their totals describe retrieval, not eligible primary studies or an exhaustive appraisal set.
This is a critical narrative review. It does not claim registration, formal systematic-review completeness or duplicate independent screening of every retrieved record. Exact queries and reconciliation are documented in the accompanying appendix. Targeted retrieval and citation chasing emphasized original dynamometry measurement properties, endurance interpretation, rapid-force protocols, newer responsiveness evidence and actual later outcomes. The source-access inventory distinguishes complete bodies, publisher or repository PDF text, original abstracts and unresolved tables or supplements.
Maximal force and torque
A maximal voluntary isometric contraction quantifies the greatest measured external force or torque generated under a specified stationary setup. It is an observable voluntary performance, influenced by the person's available capacity, task familiarity, symptom tolerance, confidence and instructions. The value is not a direct assay of isolated muscle tissue quality and cannot prove effort was maximal in a physiological sense.
Force and torque are different quantities. A handheld load cell may report newtons, while a machine reports newton-metres. Converting force into torque requires a defined perpendicular moment arm and valid alignment assumptions. A load cell applied farther from the joint can register a different force for the same joint moment. Some papers print torque as N/m; that notation should not be copied into a tool as though newtons per metre were torque. Where the source describes a rotational moment, the physical unit is N·m, with the source's reporting inconsistency documented.
External fixation reduces dependence on the examiner's strength, but it does not remove positioning or instruction error. A strap-fixed instrument and an examiner-held instrument with the same brand name are different protocols. Pelvic restraint, knee position, hip position, upper-trunk contact and the freedom to recruit the lower limbs determine how much of the result represents a functional whole-trunk exertion rather than a more isolated lumbar task. [3, 5, 6]
Isokinetic testing constrains angular velocity over a specified range and can provide torque, work and power. Acceleration and deceleration regions, gravity correction and the portion of the movement used in analysis matter. Its reproducibility does not transfer automatically to a static handheld test. Conversely, a handheld test can be clinically useful without reproducing every output of a laboratory machine. [2, 5, 7]
Endurance and movement control
The Biering-Sørensen task measures how long a person maintains a specified unsupported horizontal trunk position. The external demand depends on upper-body mass and its lever arm, and the hip extensors contribute. Stopping because of back pain, hamstring fatigue, loss of position or unwillingness to continue are not identical outcomes. A modified support arrangement, arm position, time cap or feedback system changes the task.
The same principle applies to trunk-flexor holds, side bridges, active sit-ups and the Ito test. These can be useful low-cost performance assessments when their demands match the clinical question. They should be named and described individually rather than treated as alternate units of a common strength variable. Longer duration can reflect greater endurance, a lower relative loading demand, better task familiarity or greater tolerance. [7‑12]
Movement-control tests ask whether a person can perform or maintain an instructed movement pattern. Ultrasound thickness change and EMG describe aspects of muscle behavior; pressure biofeedback measures the pressure response to a task. None directly measures maximal trunk force. Poor performance may identify a task worth exploring, but it does not establish spinal instability or prove that one muscle is too weak. A clinical explanation should stay within the quantity actually observed.
Mechanical power and rate of force development
Mechanical power is work per unit time, or the product of force and velocity, or torque and angular velocity. An isometric contraction has no external joint movement, so its maximal force is not mechanical power. A dynamometer's isokinetic power output, an estimated sit-to-stand power and stair-ascent work divided by time each depend on their own measurement model.
Rate of force development, or RFD, is a slope of the force–time curve. Peak derivative, average slope from onset to a fixed time, a slope between percentage-force thresholds and peak force divided by time-to-peak are different definitions. Rate of torque development is the corresponding rotational quantity. The rate of EMG rise is an electrical activation measure and should not be renamed force development.
Whole-body force-platform slopes during chair rise or a deadlift also differ from isolated trunk-extension RFD. The measured signal reflects multiple joints, body movement, contact forces and task strategy. A faster chair rise may be consistent with improved capacity, but elapsed time alone does not measure early force development. A camera can describe motion and support a validated power estimate; it cannot directly recover RFD without a separately validated force-estimation method.
Table 1 Keep the performance constructs separate
A familiar test name does not justify substituting one construct for another.
| Construct | Direct output | Important distinction |
|---|---|---|
| Maximal isometric force | Newtons under a fixed task | Voluntary capacity; fixation and lever arm matter |
| Torque | Newtons multiplied by metres | Not interchangeable with force without geometry |
| Endurance | Time under a specified load and position | Includes load, symptoms and stopping strategy |
| Movement control | Ability to perform an instructed movement | Not maximal force or isolated tissue capacity |
| Power | Work/time, force×velocity or torque×angular velocity | Isometric force alone is not external mechanical power |
| RFD or RTD | Defined force–time or torque–time slope | Window, onset and filtering define the measure |
Maximal strength measurement properties
Practicable instruments and the newer direct evidence
The 2022 review by Althobaiti and colleagues identified 34 studies and 15 practicable trunk-strength measures, with much stronger evidence in asymptomatic participants than in people with spinal pain. None of the included studies evaluated responsiveness. Its conclusion is an evidence map through June 2021, not a permanent statement that symptomatic testing cannot be useful. The subsequent 2023 and 2025 original studies directly address parts of that gap. [1, 3, 4]
The 2023 HHD study used external fixation and separated within-day from between-day testing. Twenty participants with nonspecific chronic LBP completed reliability testing; separate asymptomatic samples contributed reliability and criterion-comparison data. Two HHD sessions were separated by 30 minutes, and another session occurred 5–10 days later. Testing included submaximal familiarization and three 3–5-second maximal attempts, with the highest force retained. One experienced examiner performed the assessments. [3]
Its practical positive finding is that a relatively simple fixed setup can measure trunk force repeatably. In the patient group, between-day supported-supine flexion had ICC 0.82, SEM 13.60 N and MDC95 37.72 N; prone extension had ICC 0.88, SEM 12.32 N and MDC95 34.16 N. The confidence intervals for those ICCs were 0.61–0.92 and 0.72–0.95. These figures apply to the study's exact positions, selected output and population, not to all trunk HHD testing. [3]
Standing procedures were less straightforward. Significant between-day systematic differences were reported for standing flexion and extension, so a high ICC did not remove a learning or systematic-shift problem. The criterion comparison tested both instruments in the same semi-standing position and correlated HHD force with machine torque; in LBP, correlations were 0.78 for flexion and 0.68 for extension. This supports related construct information under matched positioning, but correlation across different units does not establish agreement, a valid conversion or interchangeability. The supported-supine and prone reliability results concern different setups from this criterion comparison. The original authors also advised against using the instruments interchangeably. [3]
An HHD module should therefore support a specified fixed protocol, not advertise a generic property of the instrument. It should record the contact location, fixation, position and lever-arm convention, as well as the number of trials and whether the maximum or mean is used. If an operator cannot reproduce the setup, the software should identify the changed protocol rather than interpret the difference as recovery.
Table 2 Handheld dynamometry has useful protocol specific evidence
The supported-position reliability tests and semi-standing criterion comparison used different setups. None of these values is a patient-important change threshold.
| Finding in chronic nonspecific LBP | Result | Application limit |
|---|---|---|
| Supported 30-degree supine flexion [3] | Between-day ICC 0.82; SEM 13.60 N; MDC95 37.72 N | Fixed device; maximum of three MVICs; same-examiner 5–10-day retest |
| Prone extension [3] | Between-day ICC 0.88; SEM 12.32 N; MDC95 34.16 N | Exact contact, fixation and body position required |
| Standing procedures [3] | Systematic between-day differences in flexion and extension | High ICC does not remove bias or learning |
| Matched semi-standing device comparison [3] | HHD–Biodex correlations 0.78 flexion and 0.68 extension | Correlation of force with torque does not establish interchangeability |
Laboratory dynamometry and familiarization
Verbrugghe and colleagues enrolled 20 people with chronic nonspecific LBP and 20 healthy participants; the accepted manuscript reports 20 and 19 analyzed, respectively, using functional-trunk and isolated-lumbar Biodex protocols. The original abstract reports high reliability and comparatively small relative SEM, which supports carefully standardized machine-based testing. Its reported MDC ranges span different protocols, operators and groups; accepted-manuscript Table 3 is now available, but the combined ranges retain final-abstract attribution. This report does not implement a single error threshold from the combined range. [5]
The isokinetic reliability review also supports good reproducibility under defined conditions and emphasizes familiarization. This is compatible with, rather than a refutation of, the learning effects in older originals. A stable testing procedure after sufficient practice can be reliable even when the first unfamiliar session is an unsuitable baseline. [2]
Kienbacher's study assessed 210 patients across age groups at baseline, 1–2 days and six weeks without an intervening training programme. Direction- and age-dependent retest changes persisted despite standardized equipment and encouragement. The authors favored an early repeat baseline where feasible. The original report also documents noncompletion and lost measurements, including a pain-related dropout. Older adults should not be excluded from strength measurement by assumption, but their appropriateness for a particular demanding protocol remains a clinical decision. [6]
An initial session should consequently be considered both an assessment and an opportunity to establish reproducible task execution. If a second session is used as the reference after practice, the interface should state that. When a patient is tested only once before treatment, an early gain should be interpreted with the possibility of learning in mind. A machine does not remove that uncertainty.
Error and meaningful change
High ICC does not mean a small individual difference is real
Relative reliability is sensitive to the range of people in a study. An ICC can be excellent while two tests of the same person differ substantially. Absolute error should therefore be stated in the actual unit and, when justified, as a relative value. The ICC model, confidence interval, retest interval, trial aggregation and stability assumptions should accompany any threshold.
Keller's three-session study is an early illustration. It reported high patient ICCs but substantial critical differences, including a 57% figure for the Biering-Sørensen test and velocity-dependent values for isokinetic strength. The original full error derivation was not retrieved, so these are contextual findings rather than software cutoffs. Their useful message is that relative ranking and individual precision can lead to different judgments. [13]
The subacute recurrent-LBP study similarly found excellent MVIC ICCs alongside about 31% smallest detectable differences in the patient group and approximately 9–10% higher force on the second day. Controls differed in age and BMI, and the sample was selected for ability to perform the protocol. The study provides positive reproducibility evidence for that short paraspinal procedure while showing why its absolute error and practice effects must remain visible. It does not establish a chronic nonspecific-LBP MIC. [8]
SEM, MDC, SDC and smallest real difference describe measurement error under their stated definitions. An MDC95 derived as 1.96 × square root of 2 × SEM assumes a suitable repeated-measurement error model. Systematic bias may require separate handling. A threshold calculated for an averaged result should not be assigned to one trial, and an estimate based on stable within-day performance should not be relabelled a six-week change threshold.
Responsiveness is not patient importance
The 2025 HHD responsiveness study followed 22 adults recruited with nonspecific chronic LBP through six weeks of progressive trunk resistance training; 21 completed follow-up. The fixed HHD detected average improvements, with effect sizes from 0.40 to 0.85 and standardized response means from 0.60 to 0.74. This is useful evidence that the measure can register change under that programme. It is not evidence that every observed gain exceeds individual error or is important to the patient. [4]
External responsiveness was less convincing. HHD flexion and extension change correlated with machine change at 0.22 and 0.26, with wide confidence intervals spanning zero. The HHD and machine required different postures, and HHD was always tested first. Thus the study cannot determine whether the divergence arose from poor change measurement, different constructs, order effects or several factors together. It also lacked a patient-importance anchor and cannot supply a universal responder cutoff. [4]
Effect sizes and standardized response means depend on the sample's variability and the magnitude of change that occurred. They are not intrinsic device constants. A programme in which participants improve little may produce a smaller responsiveness statistic even when the instrument is accurate. Conversely, a large mean improvement does not establish how precisely an individual changed.
Patient-important change needs a credible anchor related to the intended strength or functional construct, adequate association between anchor and change, a clear time interval and uncertainty. A global pain-improvement rating may not be an adequate anchor for isolated force capacity. A change exceeding measurement error may matter little to a person's activities; a valued improvement in daily function may occur without a detectable force increase. No general patient-anchored threshold for trunk force, power or RFD was established by the direct evidence appraised here.
Table 3 Change evidence should not be collapsed into a responder label
Error, responsiveness, importance and treatment efficacy are distinct questions.
| Evidence | Positive finding | Remaining question |
|---|---|---|
| HHD responsiveness [4] | Average strength change detected after six weeks; ES 0.40–0.85 | Change correlations with machine only 0.22–0.26; no patient anchor |
| Machine retesting [6] | Standardized assessment feasible across adult age groups | Age/direction-dependent learning and lost or incomplete measurements |
| Subacute MVIC [8] | High patient ICCs under a brief protocol | About 31% detectable difference and 9–10% day-two gain |
| Sørensen [7, 13] | Useful sustained performance task and selected-group discrimination | Wide individual error; no universal importance cutoff |
Endurance as a useful but composite performance
The Biering Sorensen task
The Biering-Sørensen test is attractive because it is inexpensive and directly demonstrates sustained task performance. Its construct should remain visible: duration under a body-mass-dependent external demand, with position, support, feedback and stopping criteria specified. Record why the person stopped and whether the prescribed position was maintained. Do not interpret every short duration as weak lumbar extensors.
Gruther's case-control and retest study found strong discrimination of selected CLBP cases from controls with the endurance test, but its individual retest limits of agreement were wide, approximately −89 to +103 seconds in the retest sample. A high case-control AUC and little average learning therefore did not establish precise serial measurement for every patient. Its selected groups also do not constitute a diagnostic reference population for unexplained back pain. [7]
The recurrent-LBP construct studies provide an important complementary explanation. Applegate's conventional task analysis and the related virtual-reality work found that trunk mass, self-efficacy and other psychological or physiological measures could be associated with duration. In the VR study, the mild recurrent-LBP group did not have a shorter mean time to failure than matched controls. Altering engagement through the VR task changes the assessment context, not just its appearance. Possible participant overlap between these reports is not counted as independent replication. [9, 10]
Russ and colleagues analysed baseline data from 30 recurrent-LBP participants and found modest relationships between endurance time, isolated extensor performance, anthropometry and disability. Their use of duration-based future-risk categories does not convert the analysis into a prospective recurrence study. A risk label borrowed from earlier literature remains a classification until actual subsequent outcomes are observed. [11]
Other endurance tests and symptom tolerance
The office-worker study of prone and supine endurance tests reported strong retest reliability in a subacute-LBP subgroup. That is useful evidence for a specific occupational and symptom-stage application. It should not be used to assign chronic-LBP norms, compare an older frail population against young workers, or create an importance threshold from group discrimination. [12]
Submaximal EMG-fatigue protocols address a different question from time to failure. A task set at a percentage of measured MVIC depends on the accuracy of that MVIC. If the baseline maximal effort is symptom-limited, the nominal 60% load may not represent the same physiological intensity across participants. Median-frequency slopes and EMG amplitudes also have their own electrode, processing and repeatability requirements. A reliable force signal does not validate every EMG feature collected during it. [8]
Pain and fear are relevant context, not grounds for dismissing performance. The broader fear-related-performance review supports considering these relationships, but neither a weak score nor a fear score establishes intent, malingering or the cause of a deficit. Record the participant's experience and use task-specific discussion rather than telling them that an observed movement proves danger. [25]
Rapid force and power evidence
Direct trunk RFD
Rossi and colleagues compared 14 women with LBP with 14 controls during dynamometer-based isometric trunk tasks. The symptomatic group showed lower extension RFD and delayed attainment of peak RFD during flexion, with relationships to rapid EMG activation. This is direct case-control evidence for a potentially relevant capability. It is not a reliability study, a patient-important-change study or evidence that the measured deficit forecasts disability. The original acquisition, filtering and onset methods are now available. This case-control study does not establish a reliability or patient-important RFD threshold, so no such threshold is imported. [14]
The small Schilaty study combined trunk dynamometry, gait, balance and patient reports in 18 CLBP participants and 15 controls. Its breadth is useful for exploring how domains relate, but simultaneous measurements cannot establish whether strength or power causes a functional limitation. A statistically significant correlation in this sample is not enough to estimate the benefit of increasing an individual's force by a specified amount. [16]
Task and symptom stage can change the result
The isometric deadlift study tested 16 resistance-trained adults with mild acute nonspecific LBP and 19 controls. Force was acquired at 1,000 Hz, and peak RFD was defined as the highest first derivative between onset and peak force, normalized to body mass. No group difference was found in normalized peak force or peak RFD. The cohort had low disability, recent symptoms and substantial prior resistance-training experience. This does not contradict the chronic or recurrent trunk-extension findings; it addresses a different task and selected population. [15]
Its practical lesson is not that acute pain has no effect on force, nor that maximal deadlifting is appropriate for all patients. It shows that pain status does not impose a uniform deficit across tasks. The compound pull distributes demand across the trunk, hips, knees and upper limbs. Bar type changed some muscle-excitation patterns without changing the primary group comparison, illustrating why a similar external performance can arise through different strategies.
The 2025 caregiver pilot measured rapid chair transfers using a 100-Hz force platform with a five-point moving average and several 50–200-millisecond windows. Thirty-two of 49 recruited participants provided complete data, including only 13 with CLBP. Early stand-to-sit force slopes were associated with task-specific fear and current pain-group membership. Because all variables were contemporaneous, the mediation model cannot establish that fear changes RFD and thereby causes chronic pain. Numerical inconsistencies between the article's coefficient table and narrative provide an additional reason to withhold its mediation coefficients from implementation. [17]
At 100 Hz, a 50-millisecond window spans only about five sampling intervals; a five-point moving average operates on a similar time scale. This is a protocol observation, not a blanket declaration that such data are invalid. It means early-slope estimates require explicit processing validation and should not be equated with high-frequency isolated-muscle RFD. The chair-transfer algorithm's onset is also a task-defined event rather than an electromyographic or dynamometric contraction onset.
What an RFD module would need
A research-grade RFD workflow should retain the raw signal and document device bandwidth, sampling frequency, calibration, filtering, onset definition, preload, countermovement rejection, contraction instruction, analysis window and trial acceptance. Early windows are especially sensitive to onset shifts and smoothing. Peak derivative can be sensitive to noise; a longer window may be more stable but answer a different physiological question.
Report absolute and appropriately normalized values without assuming normalization removes all body-size effects. Force per body mass and torque per body mass are not dimensionally equivalent. Normalizing RFD to maximal force produces a relative rapid-force construct, and changes in its denominator can create a changed ratio even when absolute RFD is stable. A ratio of extensor to flexor torque likewise needs both component values and their errors.
Between-day reliability, error and tolerability should be established in the intended symptomatic population before a progress label is added. Clinical importance and future-outcome prediction then require separate studies. No RFD threshold from knee OA, healthy athletes or neurological rehabilitation should be transferred to LBP without being explicitly labelled indirect and independently validated.
Table 4 Rapid force findings depend on the population and task
Absolute and normalized values should be retained; different derivative definitions cannot share an unvalidated threshold.
| Original evidence | Finding | Transfer boundary |
|---|---|---|
| Rossi [14] | Lower extension RFD in a small female LBP group | No direct error, MIC or future-outcome validation |
| Acute deadlift [15] | No group force/RFD deficit in trained adults with mild acute symptoms | Whole-body 1000-Hz peak-derivative task; not chronic isolated trunk RFD |
| Caregiver pilot [17] | Early stand-to-sit slope related to current fear/CLBP | 100-Hz processed signal; cross-sectional mediation; numerical inconsistencies |
| Practical inference | Preserve force signal, onset, window, preload and processing | Do not estimate direct RFD from transfer time alone |
Prognosis and treatment response
Historical prevention is not chronic pain recovery
The seminal Biering-Sørensen population study followed adults for one year and distinguished first onset from recurrence or persistence. Its abstract suggests a protective association of better extensor endurance for first-time back trouble in men, while recurrence or persistence was chiefly related to the interval since previous episodes. It does not establish a universal endurance cutoff for chronic-LBP recovery. [24]
Lee's five-year study is even less directly applicable: the initially asymptomatic volunteers had a mean age of 17 years. The later-LBP group differed in an extensor/flexor ratio rather than absolute peak torque. This is prevention evidence in a young population, not adult chronic-LBP prognostic validation. A strength ratio should not become a clinical risk category on that basis. [22]
Takala and Viikari-Juntura followed separate nonsymptomatic and prior-LBP worker cohorts. Selected strength and balance measures were associated with future outcomes, but wide overlap limited screening usefulness. The accessible original material does not establish a calibrated, externally validated risk equation. Its different outcomes, including symptoms, consultation and sickness absence, should remain separate. [21]
Prospective pain and disability findings are mixed
The 2025 systematic review synthesized 42 prospective studies, half rated at high risk of bias. It found low-quality evidence of no clear predictive association between greater baseline back-extension endurance and better long-term pain, and very-low-quality inconsistent evidence for flexor endurance. This is not proof that physical capacity never matters. It reflects heterogeneous measurement, outcomes, imprecision and limited confounder control. The review excluded some instrumented measures and should not be treated as an exhaustive judgment on all strength technologies. [18]
The underlying studies illustrate the interpretation. Enthoven's small primary-care cohort found more useful associations from physical measures taken four weeks after presentation than from the initial examination. A four-week score incorporates early symptom evolution, treatment exposure and practice; it is not the same predictor as a baseline test. The original article reports increased pain after the initial physical battery in 18 of 44 patients. The original article and tables are now available, but these small-sample associations do not establish a validated, calibrated individual prediction model. [19]
In Strøyer's occupational study, the middle back-endurance category showed a significant association with increased pain at 30 months, while the lowest category and overall endurance association did not meet the conventional significance criterion. This nonmonotonic pattern should not be simplified to a universal dose–response relationship. The study's endpoint was worsening pain intensity, not falls or work clearance; original event counts and the complete adjustment specification were not available here. [20]
Predicting measured strength is a narrower claim
The 2025 exploratory resistance-training model explained a high proportion of follow-up trunk-flexion variance within a cohort of only 20 people. The seven-predictor final model, univariate screening, data-driven selection and replacement of one outlier make the apparent R² of 0.93 especially vulnerable to optimism. Its primary aim was explaining measured strength within a treatment cohort, not predicting disability or return to work. No external validation was reported. The paper also shifts between follow-up strength and strength-change wording, so a ready-to-use change calculator would require clarification. [23]
The model report explicitly describes a secondary analysis of the same HHD responsiveness cohort and should not be counted as independent replication. Similarly, the obese older-adult exercise study related improvement in lumbar strength to improvement in walking endurance. A change–change association under treatment does not prove that strength mediates the clinical benefit or identify who should receive a particular intervention. [4, 23, 26]
To use strength as a prognostic factor, a study should establish temporal order, specify the outcome and horizon, include relevant baseline severity and psychosocial or contextual predictors, address missing data and account for candidate-model complexity. To use it for treatment selection requires evidence of differential treatment benefit, not merely an association with outcome. External validation and calibration are prerequisites for an individual probability display; clinical utility is a further question.
Table 5 Prognostic claims need a temporal and population check
The Models 2025 report is an explicit secondary analysis of the HHD responsiveness cohort, not independent replication.
| Evidence | Actual question | Unsupported extension |
|---|---|---|
| Biering-Sørensen 1984 [24] | First onset versus recurrence in a population cohort | A universal chronic-LBP endurance cutoff |
| Lee 1999 [22] | Incident pain in initially asymptomatic young people, mean age 17 | Adult chronic-pain recovery from a strength ratio |
| Enthoven and Strøyer [19, 20] | Later pain/disability or worsening after baseline/repeat testing | An externally calibrated general recovery calculator |
| Models 2025 [23] | Strength after training in 20 people with seven final flexion predictors | Future disability prediction or treatment-selection benefit |
| Training association [26] | Strength change related to walking change | Proof that strengthening mediates the clinical response |
Practical feasibility and lower limb measures
Lower-limb strength and power may be clinically important when walking, stair climbing, transfers or falls are the concern. They should be assessed as their own constructs. A low stair-ascent time does not isolate quadriceps force, and a bilateral chair-rise estimate does not identify trunk extensor capacity. The test should match the activity limitation and the patient's ability to perform it safely.
Piva's large chronic-LBP feasibility study shows how this works in practice. The initial quadriceps-frame protocol produced inconsistent measurements; a subsequent machine protocol was burdensome, and the programme ultimately adopted stair climbing. These are legitimate operational adaptations, but the resulting tasks are not equivalent measures that can be concatenated into one strength trajectory. A device failure or protocol change is part of the data history. [27]
Hip HHD tests were not completed in approximately 21–26% of participants in that study, and the active sit-up was not completed in 26.8%. Safety screening accounted for much of the noncompletion. The quadriceps machine subgroup had particularly high noncompletion and time burden. These are observations from a supervised, mixed chronic-LBP cohort with comorbidities and previous surgery, not pure nonspecific-LBP norms or permission to reproduce the battery unsupervised. [27]
The 2024 resistance-band study offers a potentially accessible alternative for measuring progressive anti-rotation task performance. Its 18 symptomatic participants completed repeated row and Pallof-press holds without reported discomfort, and selected force measures had high repeatability. However, eligibility explicitly allowed herniated disc with claudication, stenosis, lumbar instability and instrumented arthrodesis. This is adjacent mixed clinical evidence, not direct validation for a homogeneous nonspecific-LBP population. Band calibration, stepwise load increments and compensation-based stopping rules also define the score. [28]
Table 6 Feasibility and data quality belong in the result
A supervised study does not by itself establish unrestricted unsupervised safety.
| Setting | Observed issue | Practical response |
|---|---|---|
| Large supervised cohort [27] | Hip HHD noncompletion about 21–26%; screening and refusal important | Retain denominators and specific noncompletion reasons |
| Quadriceps protocol changes [27] | Frame malfunction, then burdensome machine test, then stairs | Do not concatenate unlike tasks as one strength trajectory |
| Band tests [28] | Repeatable anti-rotation tasks in a mixed neurosurgical sample | State diagnosis/surgery mix and task calibration |
| Home or camera implementation | Timer or movement feature may be convenient | Validate the exact output, symptomatic users and remote workflow |
A practical rehabtools measurement record
For every test, retain the clinical population label, symptom stage, current pain, testing-related symptoms, relevant surgery or neurological features, and the reason the test was selected. Record the device and software version, calibration, contact position, stabilization, joint angles, contraction mode, instructions, practice, rest, accepted repetitions and aggregation. Keep units explicit and store the raw value before normalization.
Report noncompletion rather than substituting an invented zero. Useful categories include not clinically appropriate, participant declined, stopped for pain, stopped for fatigue, technique criterion not met and technical failure. A short endurance hold with a recorded pain endpoint provides different information from a completed maximum-capacity test. Repeated inability can itself be clinically relevant, even when no continuous numerical change can be calculated.
In the initial interface, show the measured result and the matched prior result, then the observed absolute and relative change. If a protocol-specific MDC is defensibly applicable, display it as an error reference with its source and confidence level. Do not label it the minimum worthwhile improvement. Show the person's reported functional progress separately, and flag when fixation, testing position or device changed.
More advanced development should prioritize reliable acquisition before automated interpretation. A phone-camera endurance timer or movement-quality feature needs event-detection and repeatability validation in symptomatic users. An estimated power measure requires a clear mechanical model and agreement assessment. RFD requires an appropriate force signal and processing pipeline. None should inherit a diagnostic or prognostic claim from the fact that a related laboratory measure differs between groups.
Limitations and conclusion
Several older original papers remain abstract-only or have incomplete numerical-table access. The accessible PDF text was checked where machine-readable bodies omitted tables, but visual table verification was not completed for every numerical source. This is recorded rather than hidden. Source inconsistencies in units, denominators, endpoint wording and pilot mediation tables are retained as limitations. Complete retrieval of the planned searches does not imply complete eligibility screening or eliminate publication and selection bias.
The practical evidence is sufficient to support selected, standardized strength and endurance assessments, particularly when they answer a concrete rehabilitation question. It is not sufficient to merge force, endurance, movement control, power and RFD or to convert any one result into a universal prognosis. The most useful tool is therefore precise about the task, honest about error and cautious about meaning, while still making real changes in voluntary performance visible to the clinician and patient.
Primary study characteristics
The study profiles preserve the population, design, protocol, measurement findings, change interpretation, later outcomes and limitations for each appraised original study. A source can be useful for one question while remaining insufficient for another. Contextual and mixed-population studies are explicitly identified, and related publications are not assumed to represent independent cohorts.
Althobaiti 2023
Study and population [3] Direct reliability and concurrent device comparison. 20 nonspecific CLBP; 20 control retests plus 15 control validity participants
Protocol Externally fixed Active Force 2; two HHD sessions 30 minutes apart, followed by a 5–10-day retest; maximum of three MVICs; randomized task order. Criterion comparison used matched semi-standing positions for HHD and Biodex.
Measurement and change Patient supine flexion ICC .82, SEM 13.60 N, MDC95 37.72 N; prone extension .88, 12.32 N, 34.16 N between days. MDC95, not MIC; standing systematic bias−25 / −29 N
Later outcomes and interpretation No prognosis. One examiner and young sample; force in N compared with torque in N·m; correlations of 0.68–0.78 do not establish interchangeability. Supported-position reliability and semi-standing criterion comparison are different setups.
Source examined Complete body; original repository PDF Tables 2–4, text extraction. Complete article text and original repository PDF text including Tables 1 to 4; figure and table image rendering not independently inspected
Althobaiti 2025
Study and population [4] Longitudinal responsiveness. 22 nonspecific CLBP enrolled / 21 completed; mean 33 y; ODI 41.1%
Protocol Same fixed HHD; 6-week progressive resistance; ID seated versus HHD varied posture; fixed instrument order
Measurement and change HHD ES .40–.85 / SRM .60–.74; change correlations .22 / .26 with ID. No patient anchor / MIC; flex correlation CI −0.22–.60; extension −0.18–.62
Later outcomes and interpretation Measurement change after treatment, not outcome prognosis. Small selected cohort; 1 withdrawal; unmatched posture; no controlled efficacy inference
Source examined Complete body Tables 2–4; Methods. Complete article text; figures and supplementary objects require separate checks
Verbrugghe 2019
Study and population [5] Direct operator / retest dynamometry. 20 CNSLBP analyzed / 20 healthy enrolled and 19 analyzed, according to the accepted manuscript
Protocol Biodex 3 functional trunk versus isolated lumbar; repeated operators
Measurement and change Abstract combined ranges ICC .94–.98; SEM 4.7–9.2%; MDC flex 14.3–29.8 Nm / ext 39.1–68.5 Nm. Accepted-manuscript Table 3 now available; combined ranges here retain final-abstract attribution; no pooled cutoff implemented
Later outcomes and interpretation No future outcome. Ranges across populations / protocols not a unique LBP error threshold
Source examined PMID 30995544; accepted manuscript and Table 3 now available. Combined ranges retain final-abstract attribution; subgroup estimates do not establish a universal cutoff
Kienbacher 2016
Study and population [6] Direct short / long-interval reliability. 210 CLBP recruited / 195 completed; 18–90 y; no training during 6 weeks
Protocol DAVID fixed-position flexion / extension / rotation; baseline, 1–2 d, 6 wk
Measurement and change Learning varies by age / direction; repeat early baseline recommended; ICC(2, 1), SEM, SRD95. SRD is error; source language equating LoA with clinical importance not adopted
Later outcomes and interpretation No clinical prognosis. 15 noncompleters; 38 lost measurements; one pain-related dropout; age-specific tables not generalized
Source examined Publisher PDF Methods/Results pp 2–4. Publisher PDF text inspected for methods and results; numerical age-specific table thresholds not implemented
Gruther 2009
Study and population [7] Reliability and known-groups accuracy. 32 CLBP / 19 controls / 15 headache controls; 21 in retest table
Protocol Biodex isometric / isokinetic 90 degrees per second; Biering-Sørensen; 3-week retest
Measurement and change Learning in isokinetic / isometric-flexion; Sørensen AUC .93 but LoA−88.69 to 103.17 s. No anchor-based MIC
Later outcomes and interpretation Current case classification, not prognosis. Case-control AUC not diagnostic cutoff; near-zero mean change coexists with wide individual error
Source examined Publisher PDF Table III and methods; 7-page original. Publisher PDF text inspected including Table III; not visually verified
Koumantakis 2021
Study and population [8] Direct force / EMG reliability. 66 subacute recurrent LBP / 26 controls
Protocol Isometric paraspinal MVIC plus brief 60%MVIC endurance; repeat days
Measurement and change Patient max ICC(3, 1).91, SEM 6.96 reported kg, SDD 31.61%; mean ICC 0.96, SDD 30.75%; day 2 gain≈9–10%. Error not MIC
Later outcomes and interpretation No later clinical outcome. Selected participants able to perform; control age / BMI differed; force display kg should not be treated as SI force
Source examined Complete body Tables 2/4/5; Methods. Complete article text; figures and supplementary objects require separate checks
Applegate 2019
Study and population [9] Concurrent construct analysis. 24 recurrent-LBP participants and 24 matched controls
Protocol Sørensen time to task failure, isolated force, EMG fatigue and self-efficacy
Measurement and change No group difference in duration; trunk mass and self-efficacy associated with duration in the LBP group. No MIC
Later outcomes and interpretation Statistical predictors of same-session duration, not subsequent recurrence. Related sample and methods to the VR study; participant overlap unresolved, so not independent replication
Source examined PMID 30447543 original abstract. Original abstract or metadata only; no original full-text numerical table validation
Applegate 2018
Study and population [10] Concurrent modified-task mechanism. 24 people with mild recurrent LBP and 24 matched controls; no recent pain above 3 / 10; mainly young adults
Protocol Virtual-reality Sørensen task with EMG and psychological measures
Measurement and change No group difference in duration; LBP performance related to trunk mass, fear and self-efficacy. No change threshold
Later outcomes and interpretation Explains same-session task duration, not clinical future. VR changes the task; small-sample stepwise regression; possible shared cohort with the conventional-task paper
Source examined Complete body Methods, Results and Limitations. Complete article text; figures and supplementary objects require separate checks
Russ 2021
Study and population [11] Secondary baseline association. 30 recurrent-LBP participants; 10 men and 20 women
Protocol Sørensen duration categories compared with isolated extensor performance
Measurement and change Duration correlations of 0.36–0.42 with endurance and fat-mass to strength measures. No MIC
Later outcomes and interpretation Duration-based risk labels; no newly observed recurrence follow-up. An RCT parent study does not make this baseline analysis prospective; several constructs contribute
Source examined PMID 33136088 original abstract. Original abstract or metadata only; no original full-text numerical table validation
del Pozo Cruz 2014
Study and population [12] Reliability and known-group validity. 190 office workers; 118 with subacute LBP and 72 controls; 31 LBP participants retested
Protocol Prone and supine isometric endurance; seven-day retest
Measurement and change ICCs above 0.90; case-group AUC near or above 0.70. No verified patient MIC
Later outcomes and interpretation No later outcome. Occupational and subacute population; not chronic-LBP norms; original full error table unavailable
Source examined PMID 24561788 original abstract. Original abstract or metadata only; no original full-text numerical table validation
Rossi 2017
Study and population [14] Cross-sectional RFD and EMG comparison. 14 women with LBP and 14 controls
Protocol Dynamometer-based isometric flexion and extension; force development and EMG rise
Measurement and change Lower extension RFD and delayed flexion peak RFD in the LBP group. No direct reliability, MDC or MIC established
Later outcomes and interpretation No later clinical endpoint. Exact filtering and onset tables inaccessible; sex-restricted case-control sample
Source examined PMID 28667897; full-text retrieval attempts unsuccessful. Original abstract or metadata only; no original full-text numerical table validation
Stock 2022
Study and population [15] Cross-sectional task comparison. 16 adults with acute nonspecific LBP and 19 controls; trained adults aged 18–35; restricted low-disability sample
Protocol Isometric conventional and hexagonal-bar pulls; 1000 Hz; peak first derivative; body-mass normalization
Measurement and change No group difference in peak force or RFD; bar type affected EMG patterns. No MIC or between-day error
Later outcomes and interpretation No clinical prognosis. Compound trunk and limb task; not isolated chronic-LBP RFD; perceived safety is not broad safety validation
Source examined Complete body Methods 2.2 and 2.4, Results. Complete article text; figures and supplementary objects require separate checks
Schilaty 2023
Study and population [16] Concurrent group comparison. 18 CLBP / 15 controls
Protocol IMU gait / balance and dynamometer trunk sensorimotor battery
Measurement and change Physical / psychological associations and group comparisons. No protocol-specific MIC established
Later outcomes and interpretation No future outcome. Small case-control study; sensor results and laboratory force / power not phone-derived strength
Source examined PMID 37245281; original abstract. Original abstract or metadata only; no original full-text numerical table validation
Abiko 2025
Study and population [17] Pilot concurrent mediation. 49 female caregivers recruited; 32 complete cases, including 13 CLBP and 19 controls
Protocol Rapid chair transfers; 100-Hz vertical ground-reaction force; five-point moving average; 50–200-ms windows
Measurement and change Early stand-to-sit force slopes associated with fear and current CLBP. No reliability or MIC
Later outcomes and interpretation Cross-sectional mediation without temporal or causal validation. 17 excluded for incomplete data; few cases; coefficient and uncertainty inconsistencies in Table 3 and prose; no coefficients implemented
Source examined Complete body Methods and Results, Table 3. Complete article text; figures and supplementary objects require separate checks
Enthoven 2003
Study and population [19] Prospective primary-care course. 55 included, 11 dropouts and 44 analyzed, now verified in the original author-posted article; mixed symptom duration
Protocol Physical battery at baseline and four weeks; pain and disability at 12 months
Measurement and change Not a reliability-validation study. No endurance MIC
Later outcomes and interpretation Four-week physical measures associated with later outcome; baseline less informative. Repeat assessment incorporates early course; small bivariate analyses; 18 of 44 reported pain increase after baseline battery
Source examined PMID 12892242 original abstract; Rashed Table 2 for initial denominator. Original abstract or metadata only; no original full-text numerical table validation
Strøyer 2008
Study and population [20] Prospective occupational association. 327 workers; 271 women; review reports 113 missing at 30 months
Protocol Extension / flexion endurance, balance; >2 / 10 increase in prior-year pain
Measurement and change Not measurement validation. No patient-important endurance threshold
Later outcomes and interpretation Middle endurance category OR 2.7, P=.034; low category OR 2.4, P=.076; overall P=.067; measured balance not associated. Nonmonotonic result; incomplete full-text adjustment / events; no deployable risk rule
Source examined PMID 18317201 abstract; Rashed table secondary attrition. Original abstract or metadata only; no original full-text numerical table validation
Takala 2000
Study and population [21] Prospective occupational cohort. 307 nonsymptomatic and 123 prior-LBP workers; two separate cohorts
Protocol Baseline strength / endurance / force-platform balance; 2-year pain, consultations and sick leave
Measurement and change Not a measurement-error study. No MIC
Later outcomes and interpretation Selected associations with future pain; workload / sex / age / anthropometrics considered. Large overlap; no calibrated external model; do not pool prevention and prior-LBP cohorts
Source examined PMID 10954645 original abstract; review tables secondary only. Original abstract or metadata only; no original full-text numerical table validation
Lee 1999
Study and population [22] Prospective prevention cohort. 67 initially asymptomatic volunteers; mean age 17 ± 2 years; 18 incident LBP cases
Protocol Isokinetic torque at 60 degrees per second; ratios; five-year follow-up
Measurement and change Peak torque did not differ by subsequent outcome; extension / flexion ratio differed. Not applicable
Later outcomes and interpretation Incident LBP rather than recovery in symptomatic adults. Young and adolescent prevention sample; no validated individual ratio threshold
Source examined PMID 9921591 original abstract. Original abstract or metadata only; no original full-text numerical table validation
Althobaiti 2025
Study and population [23] Exploratory treatment-cohort model. 20 CLBP participants, mean age 33; explicit secondary analysis of HHD-responsiveness cohort
Protocol Baseline and six-week training; MVIC, EMG and patient reports
Measurement and change Follow-up flexion model R² 0.93 with seven predictors in 20 people. No patient MIC
Later outcomes and interpretation Measured strength outcome within the development cohort; not disability or falls; no external validation. Univariate screening and selection; one outlier replaced; overfitting; follow-up versus change wording inconsistent
Source examined Complete body Study design, Methods, Results Tables 3–4 and Limitations. Complete article text; figures and supplementary objects require separate checks
Franco López 2024
Study and population [28] Direct test–retest study in adjacent mixed population. 18 symptomatic neurosurgery patients and 17 controls; eligibility permits stenosis, disc disease, instability and fusion
Protocol 72-hour retest; five-second holds; progressive band tension; 80-Hz strain gauge; row and Pallof press
Measurement and change High ICCs with reported SEM of 12.9–18.8 N across selected conditions; no criterion comparison. No patient MIC
Later outcomes and interpretation No clinical outcome. Mixed specific and surgical population; not isolated trunk MVC; calibration and stopping rules define the score
Source examined Complete body Methods 2.1–2.4 and Results. Complete article text; figures and supplementary objects require separate checks
Piva 2025
Study and population [27] Large descriptive feasibility. 1007 enrolled chronic-LBP cohort; 1006 feasibility denominator; prior surgery / comorbidities included
Protocol Supervised extensive battery; standing, hip HHD, dynamometry / stairs, active sit-up
Measurement and change Balance not done 3.9% bilateral / 8.9% single-leg; hip HHD 21.3–25.8%; active sit-up 26.8%. No reliability or MIC validation
Later outcomes and interpretation No prognosis. Safety-screening, refusal and protocol changes matter; quadriceps frame failed, replacement tasks not equivalent; descriptive means not norms
Source examined Complete body Tables 1/3/4 and protocol-change section. Complete article text; figures and supplementary objects require separate checks
Vincent 2014
Study and population [26] Treatment-related change association. 49 obese older adults with CLBP, aged 60–85 years
Protocol Four-month resistance training; lumbar strength and walking outcomes
Measurement and change Strength change explained 10.6% of walking-endurance change variance. No strength MIC
Later outcomes and interpretation Change–change association, not baseline prognosis. Small selected trial; no causal mediation or differential treatment-selection validation
Source examined PMID 24211698 original abstract. Original abstract or metadata only; no original full-text numerical table validation
Keller 2001
Study and population [13] Direct repeated-measurement reliability. 31 CLBP participants and matched controls
Protocol Three sessions 5–10 days apart; isokinetic extension, Sørensen and Astrand tests
Measurement and change Patient ICCs 0.93–0.98; Sørensen critical difference 57%; isokinetic critical differences 28–63% by speed. Critical difference measures error rather than importance
Later outcomes and interpretation No later clinical endpoint. Original error derivation unavailable; do not implement 57% as a universal rule
Source examined PMID 11295899 original abstract. Original abstract or metadata only; no original full-text numerical table validation
Biering Sørensen 1984
Study and population [24] Population-based prospective cohort. 449 men and 479 women, aged 30, 40, 50 or 60; 99% one-year questionnaire response
Protocol Strength, endurance and flexibility; first onset separated from recurrence or persistence
Measurement and change Historical testing protocol. No current chronic-LBP patient MIC
Later outcomes and interpretation Better extensor endurance associated with first-onset protection in men; recurrence chiefly related to prior episode interval. Prevention and recurrence strata differ; original abstract only; no calibrated modern risk model
Source examined PMID 6233709 original abstract. Original abstract or metadata only; no original full-text numerical table validation
Search methods and source access
Table 7 Search retrieval and reconciliation
| Slice/source | Retrieved pages | Provider total | Initial unique IDs | Official NCBI unique IDs |
|---|---|---|---|---|
| Measurement / PubMed | 4 | 187 | 149 | 187 |
| Measurement / Scopus | 18 | 432 | 432 | Not applicable |
| Prognosis / PubMed | 11 | 549 | 356 | 549 |
| Prognosis / Scopus | 61 | 1503 | 1503 | Not applicable |
| Protocol and technology / PubMed | 13 | 646 | 455 | 646 |
| Protocol and technology / Scopus | 64 | 1587 | 1587 | Not applicable |
Across the three overlapping slices: 1060 unique PubMed IDs and 2560 unique Scopus IDs. These are database records, not unique eligible or appraised studies.
Exact executed native queries
Measurement PubMed
("Low Back Pain"[MeSH Terms] OR "low back pain"[Title/Abstract] OR "low-back pain"[Title/Abstract] OR lumbago[Title/Abstract] OR "lumbar pain"[Title/Abstract]) AND ("Muscle Strength"[MeSH Terms] OR "Muscle Strength Dynamometer"[MeSH Terms] OR ("muscle strength"[Title/Abstract] OR "muscle power"[Title/Abstract] OR dynamometr*[Title/Abstract] OR isokinetic[Title/Abstract] OR isometric[Title/Abstract] OR "maximal voluntary"[Title/Abstract] OR "maximum voluntary"[Title/Abstract] OR "one repetition maximum"[Title/Abstract] OR 1RM[Title/Abstract] OR "rate of force"[Title/Abstract] OR "rate of torque"[Title/Abstract] OR RFD[Title/Abstract] OR RTD[Title/Abstract] OR "explosive strength"[Title/Abstract] OR "voluntary activation"[Title/Abstract] OR "leg power"[Title/Abstract] OR "knee extensor strength"[Title/Abstract] OR "hip abductor strength"[Title/Abstract] OR "trunk strength"[Title/Abstract] OR "trunk endurance"[Title/Abstract])) AND (reliab*[Title/Abstract] OR valid*[Title/Abstract] OR reproducib*[Title/Abstract] OR psychometr*[Title/Abstract] OR clinimetr*[Title/Abstract] OR agreement[Title/Abstract] OR "measurement error"[Title/Abstract] OR "standard error"[Title/Abstract] OR "minimal detectable"[Title/Abstract] OR "minimum detectable"[Title/Abstract] OR "smallest detectable"[Title/Abstract] OR "minimal important"[Title/Abstract] OR "minimally important"[Title/Abstract] OR "minimum important"[Title/Abstract] OR responsiv*[Title/Abstract] OR interpretabil*[Title/Abstract] OR "floor effect"[Title/Abstract] OR "ceiling effect"[Title/Abstract]) AND ("1800/01/01"[Date - Publication] : "2026/10/02"[Date - Publication])
Measurement Scopus
TITLE-ABS-KEY(("low back pain" OR "low-back pain" OR lumbago OR "lumbar pain") AND ("muscle strength" OR "muscle power" OR dynamometr* OR isokinetic OR isometric OR "maximal voluntary" OR "maximum voluntary" OR "one repetition maximum" OR 1RM OR "rate of force" OR "rate of torque" OR RFD OR RTD OR "explosive strength" OR "voluntary activation" OR "leg power" OR "knee extensor strength" OR "hip abductor strength" OR "trunk strength" OR "trunk endurance") AND (reliab* OR valid* OR reproducib* OR psychometr* OR clinimetr* OR agreement OR "measurement error" OR "standard error" OR "minimal detectable" OR "minimum detectable" OR "smallest detectable" OR "minimal important" OR "minimally important" OR "minimum important" OR responsiv* OR interpretabil* OR "floor effect" OR "ceiling effect")) AND PUBYEAR BEF 2027
Prognosis PubMed
("Low Back Pain"[MeSH Terms] OR "low back pain"[Title/Abstract] OR "low-back pain"[Title/Abstract] OR lumbago[Title/Abstract] OR "lumbar pain"[Title/Abstract]) AND ("Muscle Strength"[MeSH Terms] OR "Muscle Strength Dynamometer"[MeSH Terms] OR ("muscle strength"[Title/Abstract] OR "muscle power"[Title/Abstract] OR dynamometr*[Title/Abstract] OR isokinetic[Title/Abstract] OR isometric[Title/Abstract] OR "maximal voluntary"[Title/Abstract] OR "maximum voluntary"[Title/Abstract] OR "one repetition maximum"[Title/Abstract] OR 1RM[Title/Abstract] OR "rate of force"[Title/Abstract] OR "rate of torque"[Title/Abstract] OR RFD[Title/Abstract] OR RTD[Title/Abstract] OR "explosive strength"[Title/Abstract] OR "voluntary activation"[Title/Abstract] OR "leg power"[Title/Abstract] OR "knee extensor strength"[Title/Abstract] OR "hip abductor strength"[Title/Abstract] OR "trunk strength"[Title/Abstract] OR "trunk endurance"[Title/Abstract])) AND (prognos*[Title/Abstract] OR predict*[Title/Abstract] OR longitudinal[Title/Abstract] OR prospective[Title/Abstract] OR cohort[Title/Abstract] OR "follow up"[Title/Abstract] OR "follow-up"[Title/Abstract] OR recovery[Title/Abstract] OR deteriorat*[Title/Abstract] OR fall*[Title/Abstract] OR "natural history"[Title/Abstract] OR "return to work"[Title/Abstract] OR discharge[Title/Abstract] OR "risk factor"[Title/Abstract]) AND ("1800/01/01"[Date - Publication] : "2026/10/02"[Date - Publication])
Prognosis Scopus
TITLE-ABS-KEY(("low back pain" OR "low-back pain" OR lumbago OR "lumbar pain") AND ("muscle strength" OR "muscle power" OR dynamometr* OR isokinetic OR isometric OR "maximal voluntary" OR "maximum voluntary" OR "one repetition maximum" OR 1RM OR "rate of force" OR "rate of torque" OR RFD OR RTD OR "explosive strength" OR "voluntary activation" OR "leg power" OR "knee extensor strength" OR "hip abductor strength" OR "trunk strength" OR "trunk endurance") AND (prognos* OR predict* OR longitudinal OR prospective OR cohort OR "follow up" OR "follow-up" OR recovery OR deteriorat* OR fall* OR "natural history" OR "return to work" OR discharge OR "risk factor")) AND PUBYEAR BEF 2027
Protocol mechanism technology PubMed
("Low Back Pain"[MeSH Terms] OR "low back pain"[Title/Abstract] OR "low-back pain"[Title/Abstract] OR lumbago[Title/Abstract] OR "lumbar pain"[Title/Abstract]) AND ("Muscle Strength"[MeSH Terms] OR "Muscle Strength Dynamometer"[MeSH Terms] OR ("muscle strength"[Title/Abstract] OR "muscle power"[Title/Abstract] OR dynamometr*[Title/Abstract] OR isokinetic[Title/Abstract] OR isometric[Title/Abstract] OR "maximal voluntary"[Title/Abstract] OR "maximum voluntary"[Title/Abstract] OR "one repetition maximum"[Title/Abstract] OR 1RM[Title/Abstract] OR "rate of force"[Title/Abstract] OR "rate of torque"[Title/Abstract] OR RFD[Title/Abstract] OR RTD[Title/Abstract] OR "explosive strength"[Title/Abstract] OR "voluntary activation"[Title/Abstract] OR "leg power"[Title/Abstract] OR "knee extensor strength"[Title/Abstract] OR "hip abductor strength"[Title/Abstract] OR "trunk strength"[Title/Abstract] OR "trunk endurance"[Title/Abstract])) AND (protocol[Title/Abstract] OR biomechan*[Title/Abstract] OR kinematic*[Title/Abstract] OR kinetic*[Title/Abstract] OR symmetr*[Title/Abstract] OR asymmetr*[Title/Abstract] OR "weight bearing"[Title/Abstract] OR "weight-bearing"[Title/Abstract] OR "ground reaction"[Title/Abstract] OR "force plate"[Title/Abstract] OR "force platform"[Title/Abstract] OR sensor*[Title/Abstract] OR wearable*[Title/Abstract] OR inertial[Title/Abstract] OR acceleromet*[Title/Abstract] OR markerless[Title/Abstract] OR "motion capture"[Title/Abstract] OR camera[Title/Abstract] OR video[Title/Abstract] OR algorithm[Title/Abstract] OR electromyogra*[Title/Abstract] OR activation[Title/Abstract] OR "sampling frequency"[Title/Abstract] OR "filter cutoff"[Title/Abstract]) AND ("1800/01/01"[Date - Publication] : "2026/10/02"[Date - Publication])
Protocol mechanism technology Scopus
TITLE-ABS-KEY(("low back pain" OR "low-back pain" OR lumbago OR "lumbar pain") AND ("muscle strength" OR "muscle power" OR dynamometr* OR isokinetic OR isometric OR "maximal voluntary" OR "maximum voluntary" OR "one repetition maximum" OR 1RM OR "rate of force" OR "rate of torque" OR RFD OR RTD OR "explosive strength" OR "voluntary activation" OR "leg power" OR "knee extensor strength" OR "hip abductor strength" OR "trunk strength" OR "trunk endurance") AND (protocol OR biomechan* OR kinematic* OR kinetic* OR symmetr* OR asymmetr* OR "weight bearing" OR "weight-bearing" OR "ground reaction" OR "force plate" OR "force platform" OR sensor* OR wearable* OR inertial OR acceleromet* OR markerless OR "motion capture" OR camera OR video OR algorithm OR electromyogra* OR activation OR "sampling frequency" OR "filter cutoff")) AND PUBYEAR BEF 2027
Coverage reconciliation and source selection
All returned database pages were retrieved through the terminal page. Scopus unique-ID counts matched each provider total. PubMed pages repeatedly returned some records while omitting others; exact-query official NCBI ESearch/EFetch recovered all IDs and records without missing fetches. PubMed BookArticle records were handled where present. Native Scopus queries used publication year before 2027; individual included-source dates were checked against the 2 October cutoff, since issue year alone is not exact date eligibility.
Search slices are overlapping discovery sets. No claim is made that all retrieved records were independently screened. Intensive appraisal prioritized original measurement properties, influential change thresholds, useful clinical protocols, technology validation and prospective clinical outcomes. The earlier low-back-pain landscape search, review references and bounded citation chasing supplemented native queries.
Supplementary source discovery on 2 October 2026 used bounded OpenAlex citation searches: backward references from the 2025 HHD responsiveness paper (10.1186/s12891-025-08325-4), forward citations of Maribo 2012 (10.1007/s00586-011-1981-5), and backward references from Rashed 2025 (10.1371/journal.pone.0335535). Each call had a 100-record cap. Returned counts were HHD responsiveness references: 65; Maribo forward citations: 33; Rashed references: 53. These are discovery calls, not complete or independently screened citation universes. The Maribo 2009 clinical single-leg-stance original was additionally identified and verified during independent source appraisal.
Original full articles were sought through bibliographic services and lawful publisher or repository pages; original abstracts were used when full articles remained unavailable. HHD 2023 was retrieved as an article body; its original repository PDF text supplied Tables 1–4 omitted from that body. The 2025 HHD responsiveness, model, band-test, subacute force/EMG, deadlift, caregiver RFD, Y-Balance, balance review and hip-burden fall-study bodies were retrieved where relevant. Continuations for Park 2023 and Rashed 2025 were obtained through the end of each article.
The attempted machine-readable retrieval of Maribo 2012 did not provide an article. Its indexed original PMC methods and results, original abstract and author thesis were distinguished. Maribo 2009 publisher PDF and HHD 2023 numerical tables were independently checked. Some older subscription originals and exact instrument/error tables remain inaccessible. No original is labelled full text merely because its abstract or metadata was returned.
Cross-domain original supplements available for Moissenet 2023 are the numerical measurement-properties workbook and task protocols. They inform output-specific interpretation; they do not validate every video or sensor implementation. Article-level access is recorded in the primary study profiles and references; unresolved evidence needs are listed below.
Remaining source and validation needs
These gaps limit implementation claims. They do not invalidate useful protocol-specific strength measurement.
1. Accepted-manuscript Table 3 from Verbrugghe 2019 is now available. Preserve its group/protocol/operator estimates and the separately attributed final-abstract ranges; final-version differences and appendices remain unresolved. Do not assign a combined range to a specific patient protocol.
2. Rossi 2017 original acquisition, onset and filtering methods are now available, without a validated RFD change or prognosis threshold. Obtain Keller 2001’s original critical-difference derivation before implementing an endurance threshold.
3. HHD 2023 source-specific reliability values are verified, but supported-position reliability and matched semi-standing criterion comparison concern different setups. HHD 2025 change-score divergence warrants comparable-position replication and a relevant patient-important anchor.
4. Clarify the Models 2025 follow-up-strength versus change wording before recreating its equations. Its seven-predictor, twenty-person development result has no external validation and shares the HHD responsiveness cohort.
5. Resolve the caregiver RFD pilot's inconsistent coefficients and uncertainty values before numerical reuse. Its cross-sectional mediation cannot establish a causal pathway or subsequent risk, even if the arithmetic is repaired.
6. Improve prospective evidence for pain, disability, participation and recurrence using prespecified models, adequate sample/event counts, missing-data methods, internal validation and external calibration. Evidence that force changes after training is not proof of a clinical surrogate or treatment-selection effect.
7. Preserve separate development pathways for maximal force, endurance, mechanical power and rapid force. A practical video timer, modeled power estimate and direct force signal each require different validation.
References
References are numbered in first citation order. Source descriptions identify the material examined and are not study quality ratings. Links identify the original publication or the explicitly named original source version.
1. Althobaiti S, Rushton A, Aldahas A, Falla D, Heneghan NR. Practicable performance-based outcome measures of trunk muscle strength and their measurement properties: A systematic review and narrative synthesis. PloS one. 2022;17(6):e0270101. DOI 10.1371/journal.pone.0270101 Source examined: Complete review article; used for source discovery and measurement evidence context.
Source note: SRC-18653c0cbc60 Althobaiti S 2022
2. Reyes-Ferrada W, Chirosa-Rios L, Martinez-Garcia D, Rodríguez-Perea Á, Jerez-Mayorga D. Reliability of trunk strength measurements with an isokinetic dynamometer in non-specific low back pain patients: A systematic review. Journal of back and musculoskeletal rehabilitation. 2022;35(5):937-948. DOI 10.3233/bmr-210261 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-bd3eb0ec2100 Reyes-Ferrada W 2022
3. Althobaiti S, Falla D. Reliability and criterion validity of handheld dynamometry for measuring trunk muscle strength in people with and without chronic non-specific low back pain. Musculoskeletal science & practice. 2023;66:102799. DOI 10.1016/j.msksp.2023.102799 Source examined: Complete article text and original repository PDF text including Tables 1 to 4; figure and table image rendering not independently inspected.
Source note: SRC-9ebcc14385aa Althobaiti S 2023
4. Althobaiti S, Deane JA, Falla D. Responsiveness of hand-held dynamometry for measuring changes in trunk muscle strength in people with chronic low back pain. BMC musculoskeletal disorders. 2025;26(1):66. DOI 10.1186/s12891-025-08325-4 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-9f270df613f8 Althobaiti S 2025
5. Verbrugghe J, Agten A, Eijnde BO, Vandenabeele F, De Baets L, Huybrechts X, Timmermans A. Reliability and agreement of isometric functional trunk and isolated lumbar strength assessment in healthy persons and persons with chronic nonspecific low back pain. Physical therapy in sport : official journal of the Association of Chartered Physiotherapists in Sports Medicine. 2019;38:1-7. DOI 10.1016/j.ptsp.2019.03.009 Source examined: Original abstract verified; accepted manuscript identified but full PDF retrieval unsuccessful; group-specific table estimates withheld.
Source note: SRC-cbf134e46a6b Verbrugghe J 2019
6. Kienbacher T, Kollmitzer J, Anders P, Habenicht R, Starek C, Wolf M, Paul B, Mair P, Ebenbichler G. Age-related test-retest reliability of isometric trunk torque measurements in patiens with chronic low back pain. Journal of rehabilitation medicine. 2016;48(10):893-902. DOI 10.2340/16501977-2164 Source examined: Publisher PDF text inspected for methods and results; numerical age-specific table thresholds not implemented.
Source note: SRC-5ee9b998ed69 Kienbacher T 2016
7. Gruther W, Wick F, Paul B, Leitner C, Posch M, Matzner M, Crevenna R, Ebenbichler G. Diagnostic accuracy and reliability of muscle strength and endurance measurements in patients with chronic low back pain. Journal of rehabilitation medicine. 2009;41(8):613-9. DOI 10.2340/16501977-0391 Source examined: Publisher PDF text inspected including Table III; not visually verified.
Source note: SRC-9ae7cf956764 Gruther W 2009
8. Koumantakis GA, Oldham JA. Paraspinal strength and electromyographic fatigue in patients with sub-acute back pain and controls: Reliability, clinical applicability and between-group differences. World journal of orthopedics. 2021;12(11):816-832. DOI 10.5312/wjo.v12.i11.816 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-a371e339d8e2 Koumantakis GA 2021
9. Applegate ME, France CR, Russ DW, Leitkam ST, Thomas JS. Sørensen test performance is driven by different physiological and psychological variables in participants with and without recurrent low back pain. Journal of electromyography and kinesiology : official journal of the International Society of Electrophysiological Kinesiology. 2019;44:1-7. DOI 10.1016/j.jelekin.2018.11.006 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-fdc949ecc658 Applegate ME 2019
10. Applegate ME, France CR, Russ DW, Leitkam ST, Thomas JS. Determining Physiological and Psychological Predictors of Time to Task Failure on a Virtual Reality Sørensen Test in Participants With and Without Recurrent Low Back Pain: Exploratory Study. JMIR serious games. 2018;6(3):e10522. DOI 10.2196/10522 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-3277aca3d686 Applegate ME 2018
11. Russ DW, Amano S, Law TD, Thomas JS, Clark BC. Multiple measures of muscle function influence Sorensen Test performance in individuals with recurrent low back pain. Journal of back and musculoskeletal rehabilitation. 2021;34(1):139-147. DOI 10.3233/bmr-200079 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-e134a21671a4 Russ DW 2021
12. del Pozo-Cruz B, Mocholi MH, del Pozo-Cruz J, Parraca JA, Adsuar JC, Gusi N. Reliability and validity of lumbar and abdominal trunk muscle endurance tests in office workers with nonspecific subacute low back pain. Journal of back and musculoskeletal rehabilitation. 2014;27(4):399-408. DOI 10.3233/bmr-140460 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-1862df8d3ed9 del Pozo-Cruz B 2014
13. Keller A, Hellesnes J, Brox JI. Reliability of the isokinetic trunk extensor test, Biering-Sørensen test, and Astrand bicycle test: assessment of intraclass correlation coefficient and critical difference in patients with chronic low back pain and healthy individuals. Spine. 2001;26(7):771-7. DOI 10.1097/00007632-200104010-00017 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-63e3b0dd4388 Keller A 2001
14. Rossi DM, Morcelli MH, Cardozo AC, Denadai BS, Gonçalves M, Navega MT. Rate of force development and muscle activation of trunk muscles in women with and without low back pain: A case-control study. Physical therapy in sport : official journal of the Association of Chartered Physiotherapists in Sports Medicine. 2017;26:41-48. DOI 10.1016/j.ptsp.2016.12.007 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-8fe111e47225 Rossi DM 2017
15. Stock MS, Bodden ME, Bloch JM, Starnes KL, Rodriguez G, Girts RM. Acute, Non-Specific Low Back Pain Does Not Impair Isometric Deadlift Force or Electromyographic Excitation: A Cross-Sectional Study. Sports (Basel, Switzerland). 2022;10(11):168. DOI 10.3390/sports10110168 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-031461827266 Stock MS 2022
16. Schilaty N, Bates N, Holmes B, Nagai T. Group differences and associations between patient-reported outcomes and physical characteristics in chronic low back pain patients and healthy controls. Clinical biomechanics (Bristol, Avon). 2023;106:106009. DOI 10.1016/j.clinbiomech.2023.106009 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-0bba3d3fe41a Schilaty N 2023
17. Abiko T, Murata S, Shigetoh H, Matsumoto N, Hideaki Y, Ohyama M, Sakata E, Hing W. Role of Rate of Force Development in Mediating the Relationship Between Task-Specific Fear and Chronic Low Back Pain Among Japanese Caregivers: A Pilot Cross-Sectional Mediation Study. Cureus. 2025;17(12):e98512. DOI 10.7759/cureus.98512 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-82b6be7b75ce Abiko T 2025
18. Rashed R, Niazigharemakher A, Walton D, Kowalski K, Rushton A. Physical measures of physical functioning as prognostic factors to predict outcomes in low back pain: A systematic review and narrative synthesis. PloS one. 2025;20(10):e0335535. DOI 10.1371/journal.pone.0335535 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-2be3805c8003 Rashed R 2025
19. Enthoven P, Skargren E, Kjellman G, Oberg B. Course of back pain in primary care: a prospective study of physical measures. Journal of rehabilitation medicine. 2003;35(4):168-73. DOI 10.1080/16501970306124 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-fbf9b918621e Enthoven P 2003
20. Strøyer J, Jensen LD. The role of physical fitness as risk indicator of increased low back pain intensity among people working with physically and mentally disabled persons: a 30-month prospective study. Spine. 2008;33(5):546-54. DOI 10.1097/brs.0b013e3181657cde Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-bbb45bc8f2fb Stryer J 2008
21. Takala EP, Viikari-Juntura E. Do functional tests predict low back pain? Spine. 2000;25(16):2126-32. DOI 10.1097/00007632-200008150-00018 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-4835f47d5c38 Takala EP 2000
22. Lee JH, Hoshino Y, Nakamura K, Kariya Y, Saita K, Ito K. Trunk muscle weakness as a risk factor for low back pain. A 5-year prospective study. Spine. 1999;24(1):54-7. DOI 10.1097/00007632-199901010-00013 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-a36d2573652a Lee JH 1999
23. Althobaiti S, Jiménez-Grande D, Deane JA, Falla D. Explaining trunk strength variation and improvement following resistance training in people with chronic low back pain: clinical and performance-based outcomes analysis. Scientific reports. 2025;15(1):8657. DOI 10.1038/s41598-025-93280-2 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-4987393eef13 Althobaiti S 2025
24. Biering-Sørensen, F. Physical measurements as risk indicators for low-back trouble over a one-year period. Spine. 1984;9(2):106-119. DOI 10.1097/00007632-198403000-00002 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-e25638e15c43 Biering-Srensen 1984
25. Matheve T, Janssens L, Goossens N, Danneels L, Willems T, Van Oosterwijck J, De Baets L. The Relationship Between Pain-Related Psychological Factors and Maximal Physical Performance in Low Back Pain: A Systematic Review and Meta-Analysis. The journal of pain. 2022;23(12):2036-2051. DOI 10.1016/j.jpain.2022.08.001 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-4dcdcbbbd059 Matheve T 2022
26. Vincent HK, Vincent KR, Seay AN, Conrad BP, Hurley RW, George SZ. Back strength predicts walking improvement in obese, older adults with chronic low back pain. PM & R : the journal of injury, function, and rehabilitation. 2014;6(5):418-26. DOI 10.1016/j.pmrj.2013.11.002 Source examined: Original abstract or metadata only; no original full-text numerical table validation.
Source note: SRC-3899c92ce557 Vincent HK 2014
27. Piva SR, Alfikri Z, Anderst W, Bell KM, Carlesso C, Darwin J, Delitto A, Greco CM, Johnson ME, McKernan GP, McLoughlin R, Patterson CG, Roos RE, Schneider MJ, Smith C, Sowa GA, Vo NV, Zhou L. Feasibility of Physical Exam and Performance-Based Tests in Individuals With Chronic Low Back Pain: A Descriptive Study. JOR spine. 2025;8(3):e70096. DOI 10.1002/jsp2.70096 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-440cc699ea82 Piva SR 2025
28. Franco-López F, Durkalec-Michalski K, Díaz-Morón J, Higueras-Liébana E, Hernández-Belmonte A, Courel-Ibáñez J. Using Resistance-Band Tests to Evaluate Trunk Muscle Strength in Chronic Low Back Pain: A Test-Retest Reliability Study. Sensors (Basel, Switzerland). 2024;24(13):4131. DOI 10.3390/s24134131 Source examined: Complete article text; figures and supplementary objects require separate checks.
Source note: SRC-eb377d0c5aa4 Franco-Lopez F 2024