Rare Defects: Headline Accuracy Can Outrun the Evidence
A high overall accuracy score is easy to publish; evidence about rare defects, missed cases and their operational consequences is harder to obtain. Original NIST work documents this disclosure problem in semiconductor defect metrology. The resulting information imbalance concerns who can translate a test result into a production decision. Closing it requires independently labeled cases, explicit denominators, deployment conditions and a human owner of the release decision—not another reassuring percentage.
The Asymmetry in One Sentence
The party presenting an inspection score can control the test narrative while the production owner still lacks the rare-case evidence and local consequences needed to trust the decision.
Scope and Why It Matters
The signal behind this investigation is a disclosure problem identified in NIST’s 2023 semiconductor-defect research: important evaluation inputs can remain outside the public record. This brief examines binary industrial inspection, where positive means defective and an alert prompts review, hold or rework.
US federal and international research from 2015–2023, checked October 4, 2026, establishes a documented barrier and transferable statistical mechanism—not its current factory prevalence, a supplier’s behavior or a local defect rate.
Missed defects and false alarms have different consequences. Procurement, developers, process engineers and quality owners hold different parts of that context. The score becomes useful when those parts can be examined together.
Explain It Simply
Imagine inspecting boxes of screws. Almost every screw is fine. A machine that always says “fine” can be right almost all the time while finding none of the damaged screws.
Another machine catches most damaged screws but also flags good ones. To judge it, you need three answers: how many damaged screws it misses, how many good screws it sends for unnecessary review, and what each mistake causes. “Right 99% of the time” answers none of those questions by itself.
The difficult evidence is the independently checked damaged screw, especially one the machine confidently accepted.
Evidence Map: What Is Established
- Documented disclosure barrier: Barnes and Henn’s NIST paper, published April 27, 2023, says defect-metrology literature may withhold datasets, misclassification costs and class imbalance for industrial reasons. Quantified industrial costs were unavailable; the experiments used public surrogate data, not production defect imagery.
- Evaluation principle: NIST AI RMF 1.0, January 2023, calls for realistic test conditions, documented methods and evaluation during deployment. This is risk-management guidance, not certification of an inspection product.
- Statistical mechanism: Saito and Rehmsmeier’s 2015 original study explains why precision–recall analysis exposes positive-case performance on imbalanced datasets. Its applications include bioinformatics, not an industrial performance survey.
- Uncertainty: NIST Technical Note 2119, September 2020, develops confidence intervals and bounds for detection and false-alarm probabilities.
- Score interpretation: Guo and colleagues’ 2017 experiments distinguish neural-network confidence from calibrated probability. They do not show that every model is miscalibrated.
Our inference: a buyer cannot resolve a context-sensitive inspection decision from an aggregate score when denominators, rare-case labels and consequences remain unavailable. The size of that disadvantage must be established locally.
Who Knows What—and Where Proof Stops
The developer knows the model, training choices and test. The factory knows its product mix, process changes and escaped-defect consequences. Inspectors know reference methods and disputed labels. Procurement may receive only the summary.
The imbalance can run both ways: the supplier may lack local costs; the plant may lack evaluation access. Confidentiality is a plausible barrier; absent public data establishes neither concealment nor misconduct.
The path is observation → reference label → score → threshold → action → outcome. Missing methods, undisclosed thresholds and unobserved escapes weaken different links. Locate the broken link rather than requesting an undifferentiated “AI accuracy report.”
A Worked Example: 98.91% Accuracy, 47.6% Precision
Illustration only, not a benchmark: assume 10,000 inspections, 1% truly defective items, 90% sensitivity and 99% specificity at one fixed threshold. Sensitivity is the share of defects detected; specificity is the share of nondefective items correctly cleared.
| Actual condition | Flagged | Cleared |
|---|---|---|
| 100 defective | 90 true positives | 10 false negatives |
| 9,900 nondefective | 99 false positives | 9,801 true negatives |
There are 189 alerts. Precision—the fraction of alerts that are real defects—is 90 ÷ 189 = 47.6%. Overall accuracy is (90 + 9,801) ÷ 10,000 = 98.91%. Specificity of 99% does not mean an alert has a 99% chance of identifying a defect.
Always clearing every item would score 99% accuracy and miss all 100 defects. It is a revealing baseline, not an acceptable operating policy. These illustrative frequencies are model assumptions, not guaranteed counts in the next batch.
The Base Rate Changes the Meaning of an Alert
Precision depends on prevalence as well as sensitivity and specificity. In a second illustration, take 100,000 inspections with 0.1% defects and assume the same two operating rates. The expected counts are 90 detected defects, 10 missed defects, 999 false alarms and 98,901 correctly cleared items. Precision becomes 90 ÷ 1,089 = 8.3%.
The detector has not become worse under these assumptions; defects have become rarer relative to false alarms. In an actual process change, sensitivity and specificity can change too. A calculation that holds them constant isolates the base-rate effect; it is not a distribution-shift forecast.
More Images Can Still Mean Little Rare-Case Evidence
A large test dominated by easy negatives can leave sensitivity poorly determined. Report the number of independently labeled positives by important defect type, not only total image count. Ten near-identical views of one defect do not supply ten independent examples of production variation.
Deliberately enriching a test with defects can help estimate detection performance. It must be disclosed. Precision measured on that enriched mix cannot be presented directly as deployment precision; the deployment prevalence and sampling design must be accounted for.
Confidence bounds describe uncertainty under statistical assumptions. They do not cover every unseen defect, mislabeled case or future operating condition. Additional examples should answer a specific uncertainty, rather than merely enlarge a reassuring denominator.
Costs and Incentives: Who Carries the Mistake?
For a fixed scenario, a simple error-cost calculation is false positives × consequence per false positive + false negatives × consequence per false negative. In the first illustration, that is 99 × C_FP + 10 × C_FN. Cost inputs require documented local assumptions; the brief invents no prices, savings or optimal threshold.
Review time, unnecessary holds, rework and downstream escapes belong to different owners. Add normal inspection, escalation and implementation costs when comparing full workflows. If an escape already includes its downstream rework, do not charge that same loss twice. Safety or contractual requirements may impose limits that money cannot trade away.
A team rewarded for throughput can prefer fewer alerts; a developer rewarded for accuracy can prefer the majority class; a quality team bears the escapes. These are incentive hypotheses to test, not allegations about an organization. A threshold should reflect agreed consequences and reviewer capacity, not whichever metric looks strongest.
Why the Gap Can Persist
Difficult evidence is costly to create and easy to lose. Rare defects take time to encounter; reference inspection can be destructive or disputed; data can reveal proprietary processes. More ordinary images do not automatically resolve those limits.
Feedback is selective. Alert review reveals true and false positives; defects among cleared items remain invisible without another check. Apparent improvement can follow inconvenient cases disappearing from evaluation.
Confidential validation, agreed summaries and sampled review of cleared items can narrow the gap without publishing secrets. Persistence is conditional. Transfer requires checking reference quality, frequency and consequences; semiconductor evidence does not establish identical economics elsewhere.
Constraints: Labels, Independence and Change
Define what counts as a defect before scoring. Visual disagreement, tolerances and hidden internal flaws require an adjudication process. A human label is evidence with a method and limitations, not automatic ground truth.
Separate training, threshold selection and the frozen evaluation. Split correlated parts or images by an appropriate production unit, such as lot or time window. Tuning on the final test turns independent evidence into development feedback.
Check performance across relevant machines, materials, lighting and defect types. Distribution shift can alter rates and calibration. A model score of 0.9 is not necessarily a 90% defect probability; validate that interpretation on relevant labeled data. Neither calibration nor a single drift score authorizes production release. Evidence access, proprietary rights and approved data handling remain part of the evaluation design.
What Most People Miss
The decisive missing datum may be the label on a confidently cleared part. An alert dashboard is evidence about alerts, not evidence that accepted output is safe. The population absent from review can dominate the risk.
There is also a difference between model quality and workflow quality. Low precision may be tolerable when review is cheap and missing a defect is costly. High precision may be useless if achieved by detecting only obvious defects. The relevant comparison includes the existing inspection workflow, its blind spots and the downstream human decision.
Sharing more data is not the sole remedy. Sharing an auditable sampling method, per-type counts, uncertainty and a controlled independent evaluation can be more informative than an inaccessible warehouse of images.
Critical View: A Useful Detector Can Have an Unflattering Score
This argument establishes no supplier misconduct. Accuracy, ROC curves and aggregate metrics have legitimate uses when scope is explicit. Precision–recall complements them; no plot answers every decision.
The illustrative rates neither measure current systems nor rank models. The scoped 2023 disclosure observation does not establish the same missing information in a particular 2026 deployment.
A stronger reference can be costly or imperfect. Unrealistic proof can exclude a useful system or smaller provider without improving safety. Require proportionate evidence and permit controlled evaluation. An opaque evaluator score would merely move the asymmetry.
Sidy’s Synthesis: Give the Rare Case a Route to the Decision
My synthesis is a chain of responsibility: denominator → reference → consequence → decision owner. This is an analytical model, not a standard or performance formula.
The denominator says which population a claim covers. The reference says how we know the item’s condition. The consequence says what an error does in this workflow. The owner decides what evidence is sufficient and remains accountable for release. A missing link means the score has traveled farther than its proof.
This changes the investigation. Instead of requesting “better accuracy,” request the next piece of evidence that could change the action: an independent test of a costly defect, a review of cleared items, or a labeled trial under changed conditions. The useful output may be a restricted deployment, a human-review requirement or a decision to wait.
The rule: a rare defect deserves a path into the evidence, even when it barely moves the headline metric. That path can reduce an information disadvantage. It is not yet proof that a new commercial verification service has demand or sustainable economics.
AI & Future Lens
Now: AI can help organize authorized test records, surface missing denominators, draft reproducible calculations and group cases for expert inspection. Check calculations with deterministic code. Generated labels or synthetic defects may support development; they cannot stand in for independently observed production evidence. A language model can invent a plausible defect or label, so the quality owner must validate the reference and authorize any operating change.
In 5 years—2031, conditional scenario: if traceable inspection records and independent review become easier to exchange, evidence packages could travel with model updates. Verification effort may fall, but only where confidentiality, sampling and label ownership are resolved. More automated reporting could also make weak tests look more polished.
In 10 years—2036, conditional scenario: if sensing and adjudication improve, rare cases may become cheaper to characterize. Models could then optimize a richer set of consequences. Changing products and processes would still generate unfamiliar failure modes; synthetic coverage would need confirmation on real observations.
In 20 years—2046, conditional scenario: if inspection, provenance and downstream feedback are integrated, the scarce capability may be governing changing evidence rather than producing scores. A closed loop could narrow the asymmetry—or institutionalize it if one platform controls labels and access. These horizons identify dependencies, not adoption forecasts.
Build From This: A Rare-Case Evidence Register
Problem and hypothesis: evaluation summaries lose the link between a claim, its rare-case denominator and the release decision. Test whether a shared register makes those gaps visible enough to change review. This is a proposed internal pilot, not a validated product.
Inputs: authorized historical inspection records; model and threshold versions; lot or time identifiers; reference method and disagreements; positive and negative sample selection; defect type; proposed action; and documented consequence assumptions. Keep sensitive imagery under approved access controls.
Output and owner: a versioned evidence sheet showing the four outcome counts, per-type sample size, uncertainty method, sampling limits, unresolved labels and disposition. A quality engineer owns adjudication; the designated production authority owns release. No automatic model or supplier ranking.
Minimal pilot: replay one defined product family against its existing inspection workflow. Freeze the model and threshold before evaluation; include an independently selected sample of cleared items. Enrich rare cases if needed, while disclosing selection and keeping deployment prevalence separate.
Acceptance evidence: an independent reviewer reproduces counts from approved records; each critical defect claim identifies its reference and uncertainty; predefined escape and review-capacity criteria are evaluated. Insufficient positives yield “insufficient evidence,” not a passing score. Pilot acceptance does not itself authorize production.
Feedback: record reviewer disagreements, discovered escapes and context changes. Adjudicate them, revise the sampling plan, and test changes on new independent data rather than silently retuning the original holdout.
Actions: Replace a Score Request with a Proof Request
- Today: take one accuracy claim and request threshold, four outcome counts, positive-case count, reference method and test population. Record unavailable fields explicitly.
- This week: review one costly defect type with process and quality owners. Identify how misses among cleared items would be discovered and who may release the product.
- Over time: repeat independent evaluation when model, product, sensor or process conditions change. Preserve earlier versions so improvement claims remain comparable.
Where evidence cannot be shared publicly, agree a controlled review. Do not interpret confidentiality as proof of failure or accept a polished summary as proof of adequacy.
What Would Close the Gap—or Reopen the Decision?
The local asymmetry thesis weakens when evaluator and production owner can inspect independent, representative evidence, rare-case uncertainty and agreed consequences. Transparent access can remove the information disadvantage even though irreducible uncertainty remains.
Reopen acceptance after a new critical defect, an escape cluster, disputed reference labels, changed prevalence, product or imaging conditions, a model update or review queues exceeding agreed capacity. Investigate cause rather than automatically blaming the model.
Evidence of reliable current performance would support a bounded use. Evidence that an independent reference cannot be obtained would narrow the claim further. Neither result warrants pretending that every possible defect has been covered.
Remember This
- Accuracy counts correct decisions; precision counts real defects among alerts.
- Rare positives require their own denominators and uncertainty.
- Reviewing alerts alone hides false negatives.
- Expected performance is conditional, not a batch guarantee.
- A score earns decision authority through its evidence and accountable owner.
Primary sources
Facts, figures and quotations should be traceable to the sources below. Sidy's synthesis is labeled as synthesis and does not replace sourced facts.
- Barnes & Henn — Addressing misclassification costs in machine learning through asymmetric loss functions — NIST / SPIE (2023-04-27)
- NIST — Artificial Intelligence Risk Management Framework, AI RMF 1.0, §§3.1 and MEASURE 2 — NIST (2023-01)
- Saito & Rehmsmeier — The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets — PLOS ONE (2015-03-04)
- Lu — Estimating Instrument Performance: with Confidence Intervals and Confidence Bounds, NIST TN 2119 — NIST (2020-09)
- Guo, Pleiss, Sun & Weinberger — On Calibration of Modern Neural Networks — PMLR / ICML (2017)
