Skip to main content

Nonsensia.pl

NIST certifies the mass fraction of its benzoic acid primary standard PS1 as 999.92 mg/g, with a 95 percent level-of-confidence expanded uncertainty of 0.05 mg/g. That is 0.005 percent relative. The figure is small, but it is not decoration: it is the widest part of the answer, and the only part of the certificate that tells a buyer how much of the stated value can be relied on.

Most laboratory certificates carry a number of that shape, usually written after a plus-or-minus sign, usually with a small letter k somewhere nearby. Understanding what an expanded uncertainty on a certificate is supposed to mean requires three separate ideas: the standard uncertainty underneath it, the coverage factor that multiplied it, and the coverage probability that the product is supposed to deliver. Certificates often print the first and third and leave the second implicit, which is the single most common way a stated interval becomes uninterpretable.

What Expanded Uncertainty on a Certificate Actually Means

The international vocabulary of metrology defines measurement uncertainty as a non-negative parameter characterising the dispersion of the quantity values being attributed to a measurand, based on the information used. The last clause matters. Uncertainty is a statement about the state of knowledge behind a value, not a tolerance band on a manufacturing process and not an error.

Standard uncertainty is that parameter expressed as a standard deviation. Combined standard uncertainty is what results when the standard uncertainties of several input quantities are put through the measurement model. Expanded uncertainty, written U, is defined as the product of a combined standard uncertainty and a factor larger than one. That factor is the coverage factor, symbol k.

The Guide to the Expression of Uncertainty in Measurement states the relation as U = k times the combined standard uncertainty, and says a result should then be reported as an estimate plus or minus U. The interval so defined is expected to encompass a large fraction of the distribution of values that could reasonably be attributed to the measurand. Notice what is absent from that sentence: any promise about a specific percentage.

Why k Equals 2 Is a Convention and Not a Constant

The Guide says the coverage factor will in general lie in the range 2 to 3, and that for special applications it may fall outside that range. It then spends an annex explaining why picking a value that delivers an exactly known level of confidence is difficult in practice: doing so requires detailed knowledge of the probability distribution of every input quantity, which laboratories rarely have.

The working compromise is set out explicitly. Where the output distribution can be treated as approximately normal, and where the effective degrees of freedom of the combined standard uncertainty are of significant magnitude, said to be greater than 10, one may adopt k = 2 and assume the interval has a level of confidence of approximately 95 percent, or adopt k = 3 for approximately 99 percent.

The word approximately is load-bearing. For a strictly normal distribution the factors 1, 2 and 3 encompass 68.27, 95.45 and 99.73 percent respectively. When degrees of freedom are limited, the t-distribution factor needed for 95 percent is larger than 2: the Guide tabulates 2.23 at ten degrees of freedom, 2.78 at four, and 12.71 at one. At eleven degrees of freedom, k = 2 undershoots the true 95 percent factor by roughly 10 percent. A certificate quoting k = 2 for a value derived from a handful of replicates is quoting a slightly optimistic interval.

Degrees of freedom t factor for 95 percent t factor for 99 percent
1 12.71 63.66
4 2.78 4.60
10 2.23 3.17
20 2.09 2.85
50 2.01 2.68
infinite 1.960 2.576

Values as tabulated in Annex G of the Guide. The last row is the normal-distribution limit, which is where the familiar 1.96 comes from and why 2 is a convenient rounding of it.

Why the Interval Is Not a Confidence Interval

The Guide is unusually careful here. The terms confidence interval and confidence level have specific statistical definitions that apply to the interval defined by U only when certain conditions are met, including that every component contributing to the combined standard uncertainty came from a Type A evaluation. Because that condition almost never holds in chemical measurement, the Guide avoids the word confidence as a modifier of interval and uses level of confidence instead.

The vocabulary goes further and gives the interval its own name: a coverage interval, an interval containing the set of true quantity values of a measurand with a stated probability, based on the information available. It adds a note saying a coverage interval should not be termed a confidence interval, and another saying it need not be centred on the measured value at all.

Type A and Type B are not synonyms for random and systematic. A Type A evaluation is a statistical analysis of measured quantity values obtained under defined measurement conditions. A Type B evaluation is anything else: the vocabulary lists authoritative published values, the value of a certified reference material, a calibration certificate, information about drift, an instrument accuracy class, and limits deduced from personal experience. A certificate value used as an input to somebody else’s budget is, by definition, a Type B contribution to it.

What a Certificate Should State Alongside the Number

Both the Guide and NIST Technical Note 1297 list the same minimum. Report the estimate and the units. Report U and the units. Give the value of k used to obtain U, or give k together with the combined standard uncertainty. Give the approximate level of confidence and say how it was determined. NIST states the policy in one line: report U together with the coverage factor k used to obtain it, or report the combined standard uncertainty.

The Guide also warns against the bare plus-or-minus format for a standard uncertainty, precisely because that notation has traditionally signalled a high level of confidence and will be read as an expanded uncertainty whether or not a caveat is attached. Two digits are normally enough for an uncertainty; the estimate should then be rounded to match.

  • Interval without k: uninterpretable, because the same half-width can represent one, two or three standard uncertainties.
  • k without a stated level of confidence: usable, but the 95 percent reading is an assumption about the distribution.
  • k = 2 with very few degrees of freedom: an interval narrower than the 95 percent claim implies.
  • Relative and absolute forms mixed in one document: a frequent source of factor-of-ten errors when transcribed.

What the Certified Uncertainty Does Not Cover

This is the part that most often surprises a first-time buyer. NIST Technical Note 1297 states that for standards sent to NIST for calibration, the quoted uncertainty should not normally include estimates of uncertainties introduced by returning the standard to the customer and using it there. Transport, mechanical damage, the passage of time and differences in environmental conditions between the issuing laboratory and the user’s bench are explicitly outside the number, unless the report says otherwise.

The same document adds that it is generally not possible to know all the uses a result will be put to, so it is usually inappropriate for the issuing laboratory to build assumptions about downstream use into the figure it quotes. The uncertainty on a certificate is the uncertainty obtained at the issuing laboratory, for the measurand as that laboratory defined it.

Real budgets show how the pieces assemble. A 2026 reference material for amoxicillin in lyophilised bovine milk was assigned 4.10 micrograms per kilogram, with the uncertainty combining contributions from characterisation, between-bottle homogeneity and short-term stability into a single figure of 0.13 micrograms per kilogram. Homogeneity and stability are not separate reassurances printed beside the value; they are terms inside it. Laboratories sourcing characterised material for their own budgets will find the supply conditions in a supplier’s published terms of use, which is a different document from any certificate.

Frequently asked questions

What does k equals 2 mean on a certificate of analysis?

It means the stated interval is two combined standard uncertainties wide on each side of the value. Under the assumption that the underlying distribution is approximately normal and that the effective degrees of freedom are reasonably large, that interval is taken to have a level of confidence of about 95 percent. The Guide to the Expression of Uncertainty in Measurement presents this as a practical convention, not an exact result.

Is expanded uncertainty the same as a confidence interval?

No. Confidence interval has a specific statistical meaning that applies only when all contributing components were evaluated statistically, which is rare in chemistry. The metrological vocabulary calls the interval a coverage interval and states plainly that it should not be called a confidence interval. The associated probability is called coverage probability, or level of confidence in the Guide’s wording.

What is the difference between standard and expanded uncertainty?

Standard uncertainty is expressed as one standard deviation. Expanded uncertainty is that quantity, after combination across all input quantities, multiplied by a coverage factor larger than one. The expanded form exists because commercial, regulatory, health and safety contexts need an interval expected to encompass a large fraction of the plausible values rather than a bare standard deviation.

Why is the coverage factor not always 2?

Because the factor that delivers a chosen level of confidence depends on the effective degrees of freedom of the combined standard uncertainty. With few replicates the t-distribution factor is substantially larger: 2.78 at four degrees of freedom and 2.23 at ten, against 1.960 in the limit. A laboratory that needs a genuine 95 percent interval from a small dataset must use the larger factor and say so.

Does the uncertainty on a certificate cover my own measurement?

No. It covers the value as determined by the issuing laboratory for the measurand that laboratory defined. NIST states that uncertainties arising from transport, from the passage of time and from conditions at the user’s laboratory are not normally included. Those contributions belong in the user’s own budget, alongside the certified value treated as a Type B input.

What is the difference between a Type A and a Type B evaluation?

A Type A evaluation obtains a component by statistical analysis of measured values gathered under defined conditions. A Type B evaluation obtains it by any other means, such as a calibration certificate, a certified reference material value, an accuracy class or published data. The labels describe how the component was evaluated, not whether the underlying effect is random or systematic.

How many digits should a stated uncertainty have?

Usually at most two significant digits, with the reported estimate rounded to match. The Guide notes that additional digits are sometimes retained to avoid rounding errors in later calculations, and that rounding an uncertainty upwards is occasionally appropriate. A value carrying more digits than its uncertainty supports is reporting precision that the measurement does not have.

References

  1. JCGM 100:2008, Evaluation of measurement data – guide to the expression of uncertainty in measurement (GUM), BIPM
  2. JCGM 200:2012, International vocabulary of metrology (VIM), 3rd edition, BIPM
  3. Taylor and Kuyatt, NIST Technical Note 1297, 1994 edition – guidelines for evaluating and expressing the uncertainty of NIST measurement results
  4. Beauchamp, Camara, Carney et al., NIST Special Publication 260-136, 2021 – metrological tools for the reference materials and reference instruments of the NIST Material Measurement Laboratory
  5. Wei, Zhang, Suo, Wang, Veterinary Sciences, 2026 – preparation of a lyophilized bovine milk reference material for quality control of amoxicillin detection

Research use only. Nonsensia Lab supplies analytical reference standards for laboratory and research applications. This article is published for scientific and educational purposes. It is not medical advice, it does not describe any use in humans, and nothing in it should be read as a recommendation to administer any substance to a person or animal.

Filed under: Reference Standards

This article is part of our guide to The Complete Guide to Analytical Reference Standards.