Measurement

Standardising your assessments with PROMs

Choosing the questionnaires, deciding when to administer them, and reading a change in score: what makes results comparable across clinicians and over time.

13 min readUpdated

Most clinics already collect questionnaires. Few can answer the question that matters: are our patients doing better than they were six months ago, and by how much.

The obstacle is almost never collection. It lies in each clinician using the instrument they prefer, at times they choose, and in the fact that an isolated score compares to nothing. Standardisation is a team decision before it is a software function.

This guide is about that decision: which instruments to keep, when to administer them, how to read a change in score, and what to do when the score contradicts your clinical impression.

Comparability matters more than completeness

A clinic using four instruments for the same condition, according to individual preference, holds four series of data none of which compares to the others. A clinic using a single one, imperfect but consistent, can follow a cohort over two years.

The same reasoning applies over time. Changing instrument mid-course makes the record unreadable: the difference between the April score and the June score no longer means anything. If a change is necessary, make it on new episodes only, never in the middle of an ongoing course of care.

The corollary is uncomfortable and has to be accepted: an instrument everyone applies is better than a superior instrument half the team works around.

This is where the opening question finds its answer, in the outcome snapshot tab. A cohort only becomes analysable if the same instruments were administered at the same moments: the standardisation described below is what makes this screen useful.

Choosing the instruments

Aim for one instrument per broad category of condition, and no more. A list of five to seven covers the essentials of a rehabilitation clinic.

Cross-cuttingPain and perceived function

A pain scale applies everywhere and is understood without training. A patient-defined measure of function is a useful complement, since it follows the activities that matter to them rather than an imposed list.

These two form a reasonable basis for almost every record.

By regionOne instrument per territory

Lumbar spine, cervical spine, upper limb, lower limb: one recognised instrument per region is enough.

The precise choice matters less than deciding once and staying with it.

Three criteria for choosing between two candidates.

Completion time. A thirty-item questionnaire is better on paper and worse in practice, because it ends up being skipped when the waiting room is full.

The existence of a published threshold. An instrument for which nobody has established the meaningful-change threshold will give you scores you will not know how to interpret. See below.

The language. An instrument validated in French is better than an in-house translation of a more prestigious one. An unvalidated translation breaks the very comparability you are trying to obtain.

The moment of administration

This is the most neglected decision, and it determines the value of all the data.

Set three moments and put them in the protocol rather than in individual memory.

MomentWhy
At the initial assessment, before treatmentWithout a baseline, no change can be calculated. This is the most frequently missing piece of data
At a fixed interval during careA regular interval, for example every four to six weeks, allows a plateau to be seen as it forms
At the end of the episodeThis is what makes the cohort analysable. It is also the most often forgotten, because the patient who is doing well does not come back

The final measurement is almost always the missing one, and its absence makes the whole set unusable. The patient who is well enough to stop treatment is precisely the one whose result matters most. Plan to obtain it remotely if the last session does not take place.

A fixed interval is better than one left to judgement. Judgement is triggered when something looks abnormal, which biases the series towards the cases that are going badly.

Questionnaires and their evolution are attached to the record, under the follow-up tab. It is this continuity that allows two measurements to be compared: a questionnaire completed outside the record compares to nothing.

Reading a change in score

An isolated score says almost nothing. It is the difference between two measurements that carries the information, and interpreting it requires knowing two thresholds that are commonly confused.

MDCThe measurement threshold

The minimal detectable change is the smallest difference that exceeds the instrument's measurement error.

Below it, the observed change may be nothing but noise. This threshold answers the question: has something genuinely moved?

MCIDThe patient threshold

The minimal clinically important difference is the smallest change the patient perceives as a real difference in their life.

This threshold answers a different question: does the change matter to them?

The two do not coincide, and which is higher depends on the instrument. A change above the MDC but below the MCID is real without being important. This is a common situation, and naming it avoids two errors: announcing progress the patient cannot feel, and concluding there is a plateau while measurable progress continues.

Here is what the clinician sees at the moment of reading.

Score analysis

ODI, low back disability

34/ 100ModerateImproved
0100
−14vs. previous

Clinically meaningful

Threshold 12,8

Copay 2008 Spine J, MCID of 12.8 on the 0 to 100 % scale.

Score analysis

Risk stratification tool

5/ 9Unchanged
0vs. previous

No meaningful-change threshold established

L'instrument ne publie ni MCID ni MDC. Le panneau le dit, plutôt que d'afficher un nombre approché.

An illustration of the « Score analysis » panel as the clinician sees it, with a fictional patient. On the left, an instrument whose threshold is published: the score, the interpretation band it falls in, the change since the previous measurement, the verdict, and the source of the threshold. On the right, the case described below, an instrument with no published threshold.

The change is given with its sign, the verdict accounts for the direction of the instrument, and the threshold comes with its source. It is that last line that allows you to answer a payer who asks what the claim of progress rests on.

The values come from the literature, not from your software. They vary according to the population studied, and the same instrument may have several published values depending on the study chosen. A tool that displays a threshold should be able to tell you where it comes from.

In Bio6, every threshold comes with its citation, and the label says what the value actually is. The 2-point threshold on the pain scale, for example, is identified as an MDC from Childs 2005, noting that Farrar 2001's MCID is 1.7. This precision looks excessive until the day a payer asks what your claim of progress rests on.

The thresholds supplied by default

These instruments arrive configured in Bio6, with the threshold and the source. They form a reasonable starting point for a short list, and the values remain editable if your team prefers another publication.

InstrumentScaleThresholdSource
Pain (NPRS)0 to 102Childs 2005, MDC. Farrar 2001's MCID is 1.7
PSFS, patient-defined function0 to 102Backman 2016 and Wright 2017, MCID of 2 to 2.2
ODI, low back disability0 to 100 %12.8Copay 2008. Monticone 2012 reports 9.5
NDI, neck disability0 to 10019Cleland 2008. Young 2009 reports 15
QuickDASH, upper limb0 to 10011Mintken 2009, MDC of 11.2
LEFS, lower limb0 to 809Binkley 1999, also an MDC of 9
Berg Balance Scale0 to 565Donoghue and Stokes 2009, MDC. No MCID published

Two details in this table deserve attention, because they illustrate the discipline described above. Several of these values are MDCs rather than MCIDs, and they are labelled as such. And the Berg scale carries the note that no MCID has been published, rather than an approximate figure.

Instruments with no threshold

Some instruments have no published threshold for change, and that is a property of the instrument rather than a shortcoming in your system.

A risk stratification tool, for example, places a patient in a category. It is not designed as a measure of change, and following its evolution point by point amounts to making it say something it does not. Likewise, some scales have a published MDC but no MCID: you know when the change is real, not whether it matters to the patient.

The right answer is to display "no threshold established" rather than a plausible number. This is a distinction the Bio6 team applies deliberately: invented values were removed from the system precisely because they appeared in no publication.

For your protocol, the consequence is simple. Use these instruments for what they do, guiding care or stratifying, and do not draw a progress curve from them.

When the score contradicts the clinician

This happens regularly, and the way the team handles these cases determines whether PROMs are adopted or worked around.

The patient is improving, the score does not move. Check first what the instrument measures. A regional questionnaire does not capture a gain in an activity it does not list. A patient-defined measure will often give a different answer on the same record.

The score improves, the clinician does not believe it. Look at the moment of completion. A questionnaire completed in the waiting room after a long journey and one completed at the end of a session do not produce the same thing. This is one more argument for fixing the moment rather than leaving it to chance.

The score deteriorates for no apparent reason. Often the best information of the week. It frequently precedes a dropout, and a call at that point is worth more than an observation three sessions later.

The principle that holds in all three cases: the score is data, not a verdict. It does not replace clinical reasoning, but a change you cannot explain deserves to be understood rather than dismissed.

Putting the protocol in place

One page is enough, and it must be written down.

It names the instruments chosen per category of condition, the three moments of administration, the person responsible for administering them, and the shared rule for reading results. Without a document, each clinician rebuilds their own practice within a few weeks and you are back to the original problem.

Start with a single condition, the one you see most. Three months on an applied protocol is worth more than a complete protocol applied halfway, and the first comparable cohort convinces the team better than any meeting.

Frequently asked questions

How many instruments in total?

Five to seven cover a general rehabilitation clinic. Beyond that, the list becomes a catalogue from which each person chooses, which brings back the problem standardisation was meant to solve.

Can we change instrument afterwards?

Yes, on new episodes. Never in the middle of ongoing care: the difference between two different instruments has no meaning.

Should questionnaires be completed at the clinic or remotely?

Remotely by preference, before the session, so as not to consume treatment time. What matters more is consistency: always at the same point relative to the session.

What about records with no baseline measurement?

Nothing retroactive is possible. Take the first available measurement as the reference, note that it is not a baseline, and exclude those records from your cohort analyses.

Are PROMs useful for payer files?

They support a narrative, they do not replace it. A change above the published threshold, with its citation, is a considerably stronger argument than a qualitative appraisal. See the guide Completing the CNESST account of care and treatment (5055).

Further reading

More guides