Fundamentals CalibrationMetrological traceabilityMeasurement uncertaintyCalibration certificateISO/IEC 17025

Calibration, Traceability and Measurement Uncertainty

Everything a calibration certificate is really worth sits in three places: the as-found readings, the uncertainty column, and the decision rule.

Metrological traceability pyramid: uncertainty accumulating layer by layer from the SI definitions down to the instrument under test
Traceability is not a certificate, it is a chain: every step down passes through one more calibration, and adds the uncertainty of that comparison.

In brief

Calibration is the comparison of an instrument's indication against a reference standard of known uncertainty, and the reporting of the difference between them; it does not by itself make the instrument any more accurate. Bringing the deviation back inside specification is adjustment; taking the calibration result and deciding whether it meets a stated requirement is verification. Whether a certificate is worth anything depends on metrological traceability: a documented, unbroken chain of comparisons running from your instrument up through the calibration laboratory's reference standards and the national metrology institute to the definitions of the SI units, with every link in date and every link stating its own uncertainty — which is why holding a calibration certificate is not the same as having traceability. Measurement uncertainty is what that chain accumulates: components evaluated from the statistics of repeated observations are Type A, components evaluated from other existing information such as calibration certificates, data sheets and resolution are Type B; both are first expressed as standard uncertainties, combined in quadrature into the combined standard uncertainty u_c, and multiplied by a coverage factor k (conventionally k = 2, approximately 95% coverage probability for a normal distribution) to give the expanded uncertainty U on the certificate. So there are three things to watch when reading a certificate: the as-found readings, the as-left readings, and U with its k at every measurement point. And pass or fail has to be stated together with the decision rule — when the measured value ± U straddles the specification limit, the laboratory either pulls the acceptance limit in with a guard band, or applies simple acceptance where the test uncertainty ratio (TUR) is 4:1 or better.

  • Calibration only measures and reports the deviation; adjustment is what removes it
  • A certificate without a complete chain of comparisons is not traceability
  • U = k · u_c; k = 2 is roughly 95%, and an unstated k cannot be compared
  • The as-found readings decide whether the past period's measurements still stand
  • When the reference is only twice as good as the item, borderline verdicts are unreliable

Calibration is a measurement, not a repair

The International Vocabulary of Metrology (VIM) defines calibration in two steps: first, under specified conditions, establish the relation between the quantity values provided by a reference standard (with their uncertainties) and the corresponding indications of the instrument (with their uncertainties); then use that relation to obtain a measurement result from an indication. In plain terms: show the instrument something whose value is known, write down what it reads, and write down the difference together with how much that difference can be trusted. Note the writing down — what calibration produces is data, not an instrument that has become accurate. This is where purchasing and engineering departments most often go wrong the first time they deal with calibration: if the certificate shows that nothing was adjusted, the instrument that comes back is exactly as accurate as the one that went out. What you bought was knowing how accurate it is.

Three words are routinely confused with calibration. Adjustment means changing the instrument's internal correction values or its hardware so that its indication lies closer to the true value. Verification means taking the data obtained by calibration, comparing it against a stated requirement and declaring conformity or non-conformity. Repair means clearing a fault so that a function that did not work works again. In practice a single visit to the laboratory is often calibration, then adjustment if the item is out of tolerance, then calibration again — which is exactly why a certificate carries two sets of figures, as-found and as-left. Conversely, if a service performed an adjustment but issued no measurement data afterwards, no calibration has taken place: you still do not know how accurate the instrument is.

There is also a clash of terminology in the RF and microwave world worth settling first. The SOLT/TOSM 'calibration' of a vector network analyzer, the zeroing and zero/cal routine of a power meter, and the self-alignment an oscilloscope runs at power-up all mean moving systematic errors to the reference plane or performing an internal self-adjustment. They are operations the user performs before each measurement. They create no traceability and they do not replace periodic metrological calibration. Answering an auditor's 'how often do you calibrate' with 'we calibrate before every measurement' is the single most common way to lose that point — they are two different things.

Traceability: a documented chain of comparisons that cannot be broken

Metrological traceability is defined as the property of a measurement result whereby the result can be related to a stated reference through a documented, unbroken chain of calibrations, each contributing to the measurement uncertainty. There are three key words in that sentence — unbroken, documented, and each contributing to the uncertainty. Lose any one of them and there is no traceability. The pyramid in the lead figure is that chain: the instrument in your hands is calibrated by an in-house reference standard or a calibration laboratory, that laboratory's reference standards are calibrated against higher-order standards, up to the national metrology institute, and finally to the definitions of the SI units.

In Taiwan the practical shape of this is: most companies send instruments to calibration laboratories accredited by the Taiwan Accreditation Foundation (TAF), and those laboratories' reference standards are traceable in turn to the national standards maintained by the National Measurement Laboratory (NML). Since 2019 all seven SI base units have been defined in terms of physical constants — the second by the hyperfine transition frequency of the caesium-133 ground state, 9 192 631 770 Hz, and the kilogram by the Planck constant. Because TAF is a signatory to the ILAC Mutual Recognition Arrangement (ILAC MRA), certificates issued within its accredited scope are usually also accepted by overseas customers, which is a real difference for anyone exporting.

There are four common ways a traceability chain breaks, and all four are visible on the certificate: it does not state which reference standard was used; the reference standard's own certificate has expired; the certificate carries only a 'pass' stamp, with no data and no uncertainty; or the issuing body has no corresponding capability and scope for the item at all. The minimum check a buyer can make is actually very cheap: ask the supplier for a sample certificate for the same instrument type and turn to the 'standards used' field, to see whether it gives the standard's identification number and the validity date of that standard's own certificate. If that field is blank or vague, the numbers that follow have nothing holding them up.

Uncertainty: Type A, Type B, combined and expanded

No measurement is perfect, so a measurement result is always a value plus an interval. Measurement uncertainty is the parameter that quantifies that interval: it characterises the dispersion of the values that could reasonably be attributed to the measurand. Note that it is not error — error is the difference between the measured value and the true value, a single number you can never know; uncertainty is an interval you can compute and write on the certificate. Nor is it the instrument's specification: a specification is the manufacturer's promise about a whole production run under stated conditions, whereas uncertainty is the dispersion of this measurement, of this unit, made this time, by this laboratory.

There are two methods of evaluation. A Type A evaluation comes from the statistical analysis of a series of repeated observations: take ten readings in a row, compute the standard deviation, and divide by √n to obtain the standard uncertainty of the mean. A Type B evaluation is derived from other existing information: the U on a reference standard's certificate divided by its own k, an instrument's resolution treated as a rectangular distribution and divided by √3, an accuracy specification from a data sheet, a temperature coefficient, the effect of mismatch in an RF measurement. Type B is not the second-class option — in RF power and S-parameter measurements the largest single line in the budget is usually a Type B contribution such as mismatch, and it does not get any smaller for taking more readings.

Once every contribution has been expressed as a standard uncertainty, they are combined in quadrature (the root sum of squares), u_c = √(u₁² + u₂² + … + uₙ²), and finally multiplied by a coverage factor to give the expanded uncertainty U = k · u_c. The convention is k = 2, which for an approximately normal distribution corresponds to a coverage probability of about 95% (95.45%, strictly), and that is why nearly every calibration certificate shows 'U (k = 2)'. But the k must be stated explicitly: for the same measurement, the figure at k = 1 is half the figure at k = 2 and looks better, while its coverage probability is only about 68%.

Combining in quadrature also has a very practical consequence: the largest contribution dominates the result. In the example in figure 1, the five components add up to 3.45 if you simply sum them, but only 1.68 in quadrature; halve the shortest few bars and U barely moves, halve the longest one and the whole figure drops immediately. So when someone asks you to bring the uncertainty down, the correct first step is to lay the budget out and find out what the longest bar is — usually the grade of the reference standard, mismatch, or temperature control on site, and not taking more readings.

Uncertainty budget bar chart: five components combined in quadrature into the combined standard uncertainty, then multiplied by k = 2 to give the expanded uncertainty
Fig. 1 Components combine in quadrature and the largest one dominates — the same set of figures adds up to 3.45 in a straight sum, but only 1.68 in quadrature.

How to read a calibration certificate

Start with the header and the scope: the laboratory's name, its accreditation number, and the thing most often skipped — whether the item and the measurement range you sent in fall inside that laboratory's accredited scope. Accreditation is granted item by item and range by range; it is not a single certificate of omnicompetence, and the same laboratory may be accredited for DC voltage and not for RF power at 40 GHz. Then look at the environmental conditions (23 ± 2 °C, 45 ± 15 %RH, for instance) and at the standards used and their traceability information — the field described above, the fulcrum of the whole chain.

Then the data table. A complete row should carry the nominal or applied value, the as-found reading, the as-left reading, the specification or tolerance, and the expanded uncertainty U for that point together with its k. The most important of these is as-found, because it is the only information that can answer the question 'do the measurements made during the last interval still count?'. A cheap certificate that gives only as-left readings, or only a pass stamp, looks no different day to day; the difference appears the one time an item really is out of tolerance and a customer asks you to state how far the impact reaches, and you find you have no evidence at all.

Finally, the verdict. ISO/IEC 17025:2017 requires that when a laboratory issues a statement of conformity, it states the decision rule it applied. ILAC-G8 divides the possibilities into several zones (see figure 2): measured value ± U entirely inside the specification is a pass; entirely outside is a fail; straddling the limit, strictly speaking, can only be reported as conditional pass or conditional fail — or the laboratory applies a guard band, pulling the acceptance limit in from the specification limit by one U, and then makes a binary decision. A guard band causes some results that 'look like a pass' to be failed; the price is conservatism and what you get back is a controlled risk of a wrong decision. The question for a buyer is not 'why was this failed' but 'which decision rule did you use'.

Guard band diagram: specification limit, acceptance limit and the four possible verdicts for a measured value plus or minus its expanded uncertainty
Fig. 2 When the measured value ± U straddles the specification limit, neither pass nor fail holds; what entitles the certificate to a verdict at that point is the decision rule it applies.

TUR and TAR: how much better than the item does the reference have to be

The older rule is the test accuracy ratio (TAR): the specification of the item under test divided by the specification of the reference standard, conventionally required to be 4:1. Its weakness is that it compares two data sheets and nothing else — it takes no account of the laboratory's repeatability, its environment, mismatch or the operator. What is in general use now is the test uncertainty ratio (TUR): the tolerance of the item under test (the ± span of the specification) divided by the expanded uncertainty U of the whole calibration process. Replacing 'the reference standard's specification' in the denominator with 'the uncertainty of the whole process' is what makes TUR more honest than TAR, and it frequently makes the real number a good deal smaller than expected.

The 4:1 threshold is not a law of physics, it is a risk-management convention (ANSI/NCSL Z540.3, for example, states 4:1, or alternatively a probability of false accept (PFA) of less than 2%). Its meaning is intuitive: at a TUR of 4:1, U is no more than a quarter of the tolerance, and the probability of a wrong decision from simple acceptance — no guard band, you accept what you measure — is still within an acceptable range. At a TUR of 2:1, U already takes up half the tolerance, and a verdict near the specification limit is close to a coin toss: an item passed has a fair chance of actually being out of tolerance, and an item failed may well have been fine.

To be honest about it, the higher the grade of the instrument, the harder 4:1 becomes. High-end power sensors, vector network analyzers, phase noise measurements, and time-and-frequency instruments whose frequency accuracy already approaches that of the reference source all run into the problem of there being no standard four times better available. The right response then is not to pretend the ratio was met, but to apply a guard band explicitly, or report a conditional verdict, and to state the TUR on the certificate. So it is worth a buyer asking one extra question when requesting a quotation: what is the TUR at these few critical measurement points, and if it is below 4:1, what decision rule do you use? A supplier who can answer that is usually the supplier who actually understands the subject.

What ISO/IEC 17025 accreditation means from the buyer's side

ISO/IEC 17025 sets out the competence, impartiality and consistent operation of testing and calibration laboratories. What an accreditation body — TAF, in Taiwan — grants after assessment is a recognition of limited scope: which items, which measurement ranges, and the best measurement capability (calibration and measurement capability, CMC) within that range. So the correct question is not 'are you accredited', it is 'is my item, over my range, inside your accredited scope, and what is the corresponding CMC?'. Anyone can answer yes to the first; the second can only be answered by producing the scope of accreditation and checking it.

When is a certificate within an accredited scope genuinely required? Usually in three cases: a customer contract or quality agreement requires it in writing; the product has to go through safety, EMC or type approval, where the traceability of the measuring equipment will be examined; and quality system audits in medical, automotive or aerospace work. For instruments used internally in R&D to watch trends and do design verification, a certificate that is traceable but outside an accredited scope is usually enough, and the price difference is not small. Sending every instrument for the highest grade of service is usually not rigour but a misallocated budget — what should actually be upgraded are the few instruments that directly decide what ships.

If a supplier has no accredited laboratory of its own — which is very common among instrument distributors and service companies in Taiwan — that does not mean it cannot handle your requirement; it will normally be arranged through the manufacturer or through a partner accredited laboratory. What the buyer needs to establish is four things: who issues the final certificate, whether the item is inside that body's accredited scope, whether the certificate includes as-found data and the uncertainty at each point, and what the obligation and the deadline are for notifying you if an item is found out of tolerance. Writing those four into the purchase contract or the quotation is far more use than looking at a photograph of a certificate.

Glossary

Calibration
The operation that, under specified conditions, establishes the relation between the quantity values provided by a reference standard and the indications of an instrument, and on that basis obtains a measurement result from an indication. What calibration produces is data and an uncertainty; it does not change the instrument. Bringing the indication back inside specification is adjustment.
Metrological traceability
The property whereby a measurement result can be related to a stated reference — usually the SI units — through a documented, unbroken chain of calibrations, each of which contributes to the uncertainty. A certificate missing any link in that chain does not constitute traceability.
Expanded uncertainty (U)
The combined standard uncertainty u_c multiplied by a coverage factor k: U = k · u_c. Calibration certificates conventionally use k = 2, roughly 95% coverage probability for an approximately normal distribution; a U with no stated k cannot be compared with any other certificate.
Guard band
The amount by which the acceptance limit is pulled in from the specification limit when a conformity decision is made, commonly one U, in order to hold down the risk of a wrong decision. ISO/IEC 17025 requires the decision rule applied — including whether a guard band was used — to be stated whenever a statement of conformity is issued.
Test uncertainty ratio (TUR)
The tolerance of the item under test divided by the expanded uncertainty of the calibration process. The industry convention is 4:1 or better, which is the condition under which simple acceptance with no guard band is safe; at 2:1 a pass/fail verdict near the specification limit no longer discriminates.

Related instruments

R&S®NRX View specifications R&S®NRPxxS / NRPxxSN / NRPxxSN-V View specifications R&S®SMA100B View specifications R&S®ZNA View specifications R&S®ZNB View specifications R&S®FSW View specifications

Further reading

Browse every technical article

Need help choosing the right instrument, or advice on an application?

Technical enquiry Product catalogue

Back to the Knowledge Centre