Confidence is not enough: why every recommendation needs provenance

A confidence score invites over-trust. Pairing it with evidence turns a number into something a reviewer can actually judge.

Confidence is not enough: why every recommendation needs provenance

Early in the design of our recommendation system we showed annotators a suggested label and a confidence score. Acceptance rates were high. Too high. On a recommendation type we knew to be about eighty percent accurate, annotators were accepting ninety-seven percent of suggestions. They were not evaluating the recommendation. They were evaluating the number, and the number looked fine.

What a confidence score actually communicates

A calibrated confidence score is useful information. It tells a reviewer where to spend attention. But presented alone, it does something else: it transfers authority from the person to the system. A score of 0.91 reads as "this is probably right," and a busy annotator has no way to check. So they do not.

Provenance changes the question

We redesigned the recommendation card to show evidence alongside confidence. For a cuboid suggestion: the three similar frames the agent matched against, the LiDAR extent it measured, the guideline rule it applied. Now the annotator's question changes from "do I trust 0.91?" to "do those three frames actually look like this one?" That is a question a person can answer, and it is a question that surfaces the agent's mistakes.

Acceptance rates on the same recommendation type dropped to eighty-four percent — close to its true accuracy. Annotators were catching the errors. The agent was doing its job of narrowing attention; the person was doing theirs.

A number is not a reason. Evidence is a reason.

Provenance after the decision

Provenance also runs forward. Once a label is accepted, edited, or rejected, that decision — with the person's name, the timestamp, the guideline version, and any note they left — becomes part of the label's record. Six months later, when an evaluation slice fails and someone asks why the training data looked the way it did, the answer is there. Not reconstructed. Recorded.

Design principles we arrived at

Show confidence and evidence together, never confidence alone. Phrase recommendations as actions in the guideline's vocabulary. Make rejecting a recommendation as easy as accepting it. Record every decision with its reason. Track acceptance rates per recommendation type and per annotator, and treat an acceptance rate far above known accuracy as a warning.

These are not features. They are the difference between a tool that makes annotation faster and a tool that makes annotation faster while quietly making the dataset worse.

Published November 4, 2025 · All posts

← OlderIntroducing the agentic feedback loopNewer →LiDAR annotation for robotics: what image tools get wrong

Working on the same problems?

We'd like to see your data.