Calibration is a measurement, not a repair
The International Vocabulary of Metrology (VIM) defines calibration in two steps: first, under specified conditions, establish the relation between the quantity values provided by a reference standard (with their uncertainties) and the corresponding indications of the instrument (with their uncertainties); then use that relation to obtain a measurement result from an indication. In plain terms: show the instrument something whose value is known, write down what it reads, and write down the difference together with how much that difference can be trusted. Note the writing down — what calibration produces is data, not an instrument that has become accurate. This is where purchasing and engineering departments most often go wrong the first time they deal with calibration: if the certificate shows that nothing was adjusted, the instrument that comes back is exactly as accurate as the one that went out. What you bought was knowing how accurate it is.
Three words are routinely confused with calibration. Adjustment means changing the instrument's internal correction values or its hardware so that its indication lies closer to the true value. Verification means taking the data obtained by calibration, comparing it against a stated requirement and declaring conformity or non-conformity. Repair means clearing a fault so that a function that did not work works again. In practice a single visit to the laboratory is often calibration, then adjustment if the item is out of tolerance, then calibration again — which is exactly why a certificate carries two sets of figures, as-found and as-left. Conversely, if a service performed an adjustment but issued no measurement data afterwards, no calibration has taken place: you still do not know how accurate the instrument is.
There is also a clash of terminology in the RF and microwave world worth settling first. The SOLT/TOSM 'calibration' of a vector network analyzer, the zeroing and zero/cal routine of a power meter, and the self-alignment an oscilloscope runs at power-up all mean moving systematic errors to the reference plane or performing an internal self-adjustment. They are operations the user performs before each measurement. They create no traceability and they do not replace periodic metrological calibration. Answering an auditor's 'how often do you calibrate' with 'we calibrate before every measurement' is the single most common way to lose that point — they are two different things.
Traceability: a documented chain of comparisons that cannot be broken
Metrological traceability is defined as the property of a measurement result whereby the result can be related to a stated reference through a documented, unbroken chain of calibrations, each contributing to the measurement uncertainty. There are three key words in that sentence — unbroken, documented, and each contributing to the uncertainty. Lose any one of them and there is no traceability. The pyramid in the lead figure is that chain: the instrument in your hands is calibrated by an in-house reference standard or a calibration laboratory, that laboratory's reference standards are calibrated against higher-order standards, up to the national metrology institute, and finally to the definitions of the SI units.
In Taiwan the practical shape of this is: most companies send instruments to calibration laboratories accredited by the Taiwan Accreditation Foundation (TAF), and those laboratories' reference standards are traceable in turn to the national standards maintained by the National Measurement Laboratory (NML). Since 2019 all seven SI base units have been defined in terms of physical constants — the second by the hyperfine transition frequency of the caesium-133 ground state, 9 192 631 770 Hz, and the kilogram by the Planck constant. Because TAF is a signatory to the ILAC Mutual Recognition Arrangement (ILAC MRA), certificates issued within its accredited scope are usually also accepted by overseas customers, which is a real difference for anyone exporting.
There are four common ways a traceability chain breaks, and all four are visible on the certificate: it does not state which reference standard was used; the reference standard's own certificate has expired; the certificate carries only a 'pass' stamp, with no data and no uncertainty; or the issuing body has no corresponding capability and scope for the item at all. The minimum check a buyer can make is actually very cheap: ask the supplier for a sample certificate for the same instrument type and turn to the 'standards used' field, to see whether it gives the standard's identification number and the validity date of that standard's own certificate. If that field is blank or vague, the numbers that follow have nothing holding them up.
Uncertainty: Type A, Type B, combined and expanded
No measurement is perfect, so a measurement result is always a value plus an interval. Measurement uncertainty is the parameter that quantifies that interval: it characterises the dispersion of the values that could reasonably be attributed to the measurand. Note that it is not error — error is the difference between the measured value and the true value, a single number you can never know; uncertainty is an interval you can compute and write on the certificate. Nor is it the instrument's specification: a specification is the manufacturer's promise about a whole production run under stated conditions, whereas uncertainty is the dispersion of this measurement, of this unit, made this time, by this laboratory.
There are two methods of evaluation. A Type A evaluation comes from the statistical analysis of a series of repeated observations: take ten readings in a row, compute the standard deviation, and divide by √n to obtain the standard uncertainty of the mean. A Type B evaluation is derived from other existing information: the U on a reference standard's certificate divided by its own k, an instrument's resolution treated as a rectangular distribution and divided by √3, an accuracy specification from a data sheet, a temperature coefficient, the effect of mismatch in an RF measurement. Type B is not the second-class option — in RF power and S-parameter measurements the largest single line in the budget is usually a Type B contribution such as mismatch, and it does not get any smaller for taking more readings.
Once every contribution has been expressed as a standard uncertainty, they are combined in quadrature (the root sum of squares), u_c = √(u₁² + u₂² + … + uₙ²), and finally multiplied by a coverage factor to give the expanded uncertainty U = k · u_c. The convention is k = 2, which for an approximately normal distribution corresponds to a coverage probability of about 95% (95.45%, strictly), and that is why nearly every calibration certificate shows 'U (k = 2)'. But the k must be stated explicitly: for the same measurement, the figure at k = 1 is half the figure at k = 2 and looks better, while its coverage probability is only about 68%.
Combining in quadrature also has a very practical consequence: the largest contribution dominates the result. In the example in figure 1, the five components add up to 3.45 if you simply sum them, but only 1.68 in quadrature; halve the shortest few bars and U barely moves, halve the longest one and the whole figure drops immediately. So when someone asks you to bring the uncertainty down, the correct first step is to lay the budget out and find out what the longest bar is — usually the grade of the reference standard, mismatch, or temperature control on site, and not taking more readings.
How to read a calibration certificate
Start with the header and the scope: the laboratory's name, its accreditation number, and the thing most often skipped — whether the item and the measurement range you sent in fall inside that laboratory's accredited scope. Accreditation is granted item by item and range by range; it is not a single certificate of omnicompetence, and the same laboratory may be accredited for DC voltage and not for RF power at 40 GHz. Then look at the environmental conditions (23 ± 2 °C, 45 ± 15 %RH, for instance) and at the standards used and their traceability information — the field described above, the fulcrum of the whole chain.
Then the data table. A complete row should carry the nominal or applied value, the as-found reading, the as-left reading, the specification or tolerance, and the expanded uncertainty U for that point together with its k. The most important of these is as-found, because it is the only information that can answer the question 'do the measurements made during the last interval still count?'. A cheap certificate that gives only as-left readings, or only a pass stamp, looks no different day to day; the difference appears the one time an item really is out of tolerance and a customer asks you to state how far the impact reaches, and you find you have no evidence at all.
Finally, the verdict. ISO/IEC 17025:2017 requires that when a laboratory issues a statement of conformity, it states the decision rule it applied. ILAC-G8 divides the possibilities into several zones (see figure 2): measured value ± U entirely inside the specification is a pass; entirely outside is a fail; straddling the limit, strictly speaking, can only be reported as conditional pass or conditional fail — or the laboratory applies a guard band, pulling the acceptance limit in from the specification limit by one U, and then makes a binary decision. A guard band causes some results that 'look like a pass' to be failed; the price is conservatism and what you get back is a controlled risk of a wrong decision. The question for a buyer is not 'why was this failed' but 'which decision rule did you use'.
TUR and TAR: how much better than the item does the reference have to be
The older rule is the test accuracy ratio (TAR): the specification of the item under test divided by the specification of the reference standard, conventionally required to be 4:1. Its weakness is that it compares two data sheets and nothing else — it takes no account of the laboratory's repeatability, its environment, mismatch or the operator. What is in general use now is the test uncertainty ratio (TUR): the tolerance of the item under test (the ± span of the specification) divided by the expanded uncertainty U of the whole calibration process. Replacing 'the reference standard's specification' in the denominator with 'the uncertainty of the whole process' is what makes TUR more honest than TAR, and it frequently makes the real number a good deal smaller than expected.
The 4:1 threshold is not a law of physics, it is a risk-management convention (ANSI/NCSL Z540.3, for example, states 4:1, or alternatively a probability of false accept (PFA) of less than 2%). Its meaning is intuitive: at a TUR of 4:1, U is no more than a quarter of the tolerance, and the probability of a wrong decision from simple acceptance — no guard band, you accept what you measure — is still within an acceptable range. At a TUR of 2:1, U already takes up half the tolerance, and a verdict near the specification limit is close to a coin toss: an item passed has a fair chance of actually being out of tolerance, and an item failed may well have been fine.
To be honest about it, the higher the grade of the instrument, the harder 4:1 becomes. High-end power sensors, vector network analyzers, phase noise measurements, and time-and-frequency instruments whose frequency accuracy already approaches that of the reference source all run into the problem of there being no standard four times better available. The right response then is not to pretend the ratio was met, but to apply a guard band explicitly, or report a conditional verdict, and to state the TUR on the certificate. So it is worth a buyer asking one extra question when requesting a quotation: what is the TUR at these few critical measurement points, and if it is below 4:1, what decision rule do you use? A supplier who can answer that is usually the supplier who actually understands the subject.
What ISO/IEC 17025 accreditation means from the buyer's side
ISO/IEC 17025 sets out the competence, impartiality and consistent operation of testing and calibration laboratories. What an accreditation body — TAF, in Taiwan — grants after assessment is a recognition of limited scope: which items, which measurement ranges, and the best measurement capability (calibration and measurement capability, CMC) within that range. So the correct question is not 'are you accredited', it is 'is my item, over my range, inside your accredited scope, and what is the corresponding CMC?'. Anyone can answer yes to the first; the second can only be answered by producing the scope of accreditation and checking it.
When is a certificate within an accredited scope genuinely required? Usually in three cases: a customer contract or quality agreement requires it in writing; the product has to go through safety, EMC or type approval, where the traceability of the measuring equipment will be examined; and quality system audits in medical, automotive or aerospace work. For instruments used internally in R&D to watch trends and do design verification, a certificate that is traceable but outside an accredited scope is usually enough, and the price difference is not small. Sending every instrument for the highest grade of service is usually not rigour but a misallocated budget — what should actually be upgraded are the few instruments that directly decide what ships.
If a supplier has no accredited laboratory of its own — which is very common among instrument distributors and service companies in Taiwan — that does not mean it cannot handle your requirement; it will normally be arranged through the manufacturer or through a partner accredited laboratory. What the buyer needs to establish is four things: who issues the final certificate, whether the item is inside that body's accredited scope, whether the certificate includes as-found data and the uncertainty at each point, and what the obligation and the deadline are for notifying you if an item is found out of tolerance. Writing those four into the purchase contract or the quotation is far more use than looking at a photograph of a certificate.