guide

Writing annotation guidelines for multimodal robotics data

Guidelines that hold across annotators, sensors, and months — and how to keep them alive as the dataset grows.

Agentuor team · Reading time: 4 minutes

Robotics guidelines fail differently from image guidelines. They must describe events in time, objects across sensors, and outcomes with physical meaning. Here is what works.

Write for the hardest sensor

If a label must be consistent between a wrist camera and a LiDAR sweep, define it in terms both can satisfy — typically 3D extent and timing — and derive the 2D appearance from that, not the reverse.

Define events by observable triggers

"Grasp begins when the fingers start closing" is observable. "Grasp begins when the robot intends to grasp" is not. Every event definition should name the signal that triggers it.

Give examples of the ambiguous cases, not the easy ones

Annotators do not need ten examples of a clean pick. They need the partial grasp, the object that slips and is recovered, the occluded placement. Collect these from your own review queue and add them to the guideline with the decision that was made.

Version the guideline and record which version applied

When a rule changes, older labels are not wrong; they followed a different rule. Recording the guideline version on every label lets you decide later whether to relabel or to slice by version.

Close the loop with reviewers

Reviewer corrections are the best source of guideline clarifications. Cluster them, resolve each ambiguity once, and push the clarification to every annotator. A guideline that is updated weekly from real disagreements stays useful; one written once becomes folklore.

Keep it short enough to be read

A forty-page guideline is a guideline nobody consults. Put definitions and triggers up front, ambiguous examples next, and everything else in an appendix.

← All guides

Put these ideas to work on your dataset.

Start with a walkthrough on your own data.