Collect, annotate, and ship
robot-ready data.
Agentuor is data infrastructure for Physical AI. Agents watch annotation as it happens and turn every reviewer decision into a better dataset.
Agents that watch the work, not just the output.
Agentuor agents analyze annotation as it happens, fold in reviewer feedback, and turn every correction into an insight the next annotator — and the next model — can use.
Analyze in real time
Every stroke, box, and event marker is checked against the guideline and the rest of the dataset while the annotator is still on the frame.
Recommend with provenance
Suggested labels and workflow actions carry a confidence score and a traceable reason. Nothing is applied without a person confirming it.
Improve continuously
Reviewer decisions retrain the recommendations, tighten guidelines, and surface the failure modes worth fixing next.
Specific actions, not vague scores.
Instead of a red flag, annotators get a concrete suggestion they can accept, edit, or reject — each one logged with who decided and why.
What makes Agentuor stand out
Humans decide, always
Agentuor is built so that every label is a human decision. Agents propose; a person confirms. No auto-apply, no silent changes — a dataset you can audit and defend.
Recommendations in real time
Suggestions arrive while the annotator is still on the frame — a boundary to add, an event to split, an outcome to mark — each with confidence and the evidence behind it.
Truly multimodal
2D, 3D, LiDAR, and video on one synchronized timeline with joint states and force-torque. Many tools bolt point clouds onto an image editor. We started from the geometry.
Enterprise-grade governance
Role-based access, complete audit trails, region choice, and hybrid deployments that keep raw data in your storage — from the first research dataset onward.
What can Agentuor do for you?
Provenance on every label
See what evidence produced a suggestion, who accepted or changed it, and which guideline version applied. Audit trails export with the data and attach to model cards.
Quality that compounds
Automated validation catches the obvious. Experts see only the cases that need judgement. Every decision recalibrates the next recommendation, so quality rises with each batch.
Coverage you can measure
Agents chart the corpus by lighting, clutter, and object class, deprioritize redundant footage, and turn blind spots into the next collection session.
Feel the simplicity
You don't need a data-engineering team to get going. Connect a robot, a rig, a simulator, or a bucket through our SDK; episodes are segmented and described automatically. We give you room to run collection, annotation, review, and evaluation in one workspace — with three delivery models available on the same dataset.
Integrate, watch, act.
Plug it in
Push episodes with a few lines of code
project: "warehouse-pick",
episode: recording,
metadata: { site: "seville-1", shift: "pm" }
});
// agents begin profiling immediately
QUALITY
FLAGS
Act when it matters
Ambiguity routed with full context
Wrist-camera projection disagrees with head-camera view by 6 cm. Matches 14 similar disputed cases in this dataset.
ep_4812v7 · rule 4.2- Routed to a perception lead
- Decision recorded with reason
- Clarification reaches 42 open tasks
One schema, lab to fleet
The dataset a research team starts with is the dataset the production fleet runs on. Growth adds rows, never rebuilds tables.
Find the patterns hiding in millions of frames
Agents cluster ambiguous examples and recurring failures by cause, so you fix the class of problem — not one instance at a time.
Guidance in the annotator's language
Not "confidence 0.6" but "split at frame 1,248". Recommendations speak the vocabulary of your guideline, with evidence attached.
Your team, ours, or both
Keep sensitive work in-house and burst volume to our managed workforce. One guideline, one quality standard, one provenance trail.
One platform from first recording to fielded robot.
Data collection
Teleop, sim, and fleet logs ingested with synchronized sensors and episode metadata.
Learn moreQuality & review
Automated validation plus expert review, scaled by agents that route the right cases.
Learn moreModel evaluation
Edge-case and real-world condition testing to harden robustness before deployment.
Learn moreEverything multimodal data needs to become training data.
Multimodal annotation
Bounding boxes, cuboids, segmentation, keypoints, event timelines, and LiDAR point labels — synchronized across every sensor on the robot.
Learn morePattern and failure-mode discovery
Agents cluster ambiguous examples and recurring mistakes across millions of frames so you fix the class of problem, not one instance.
Learn moreExpert review at scale
Automated validation catches the obvious; domain experts get only the cases that need judgement, with full context attached.
Learn moreReal-world model evaluation
Slice performance by lighting, clutter, occlusion, and embodiment. Know where a policy breaks before the robot does.
Learn moreEnterprise-grade governance
Role-based access, audit logs, data residency options, and controls built for regulated and safety-critical programs.
Learn moreResearch to production
Start with a single research dataset and grow into fleet-scale pipelines without changing tools or schemas.
Learn moreRobot data is long, multimodal, and unforgiving.
An image dataset is a pile of pictures. A robotics dataset is minutes of synchronized sensors, dozens of joints, ambiguous outcomes — and a bad label shows up as a robot doing something it shouldn't. Agentuor was built for that reality, not adapted to it.
Every stream on one clock.
Head and wrist cameras, LiDAR, joint encoders, force-torque — aligned at ingest with drift correction, so a boundary placed on the video is the same boundary in the point cloud and the telemetry.
- Dropped frames and clock skew flagged automatically
- Episodes segmented from continuous recordings
- Metadata inferred with confidence, confirmed by an operator
Stop labeling what you already have.
Agents cluster the corpus by lighting, clutter, object class, and motion, then chart it. Redundant footage is deprioritized; blind spots become the next collection session. Annotation budget goes where the robot is weakest.
About agent intelligenceFaint cells are the scenarios your robot hasn't seen enough of. Agents point collection there next.
If a label is in the dataset, you can see why.
Every recommendation, edit, flag, and decision is recorded with who made it and what evidence they had. Audit trails export with the data and attach to model cards.
- Pre-label proposedagentcuboid "tote", confidence 0.91, evidence: 3 similar frames + LiDAR extent
- Accepted with editannotatorheight adjusted +4 cm to match point cloud
- Flagged for reviewagentdisagreement with wrist-camera projection
- Adjudicatedexpertkept edited cuboid; guideline clarified: "measure to rim, not lid"
- Clarification pushedsystemreaches 42 open tasks with the same object class
Left: the life of one label. Right: failure modes grouped by cause, each with a recommended fix.
Evaluation talks back to collection.
A failing slice in evaluation links to the episodes and labels behind it — and becomes a collection or relabeling task in the same workspace. Nothing is lost between tools, because there is only one.
Before and after one platform.
beforeFive tools, zero memory
- Recorder, labeling vendor, QA spreadsheet, eval notebook, ticket queue
- Boundaries drift between annotators and across weeks
- Reviewers see labels, not the reasons behind them
- Evaluation reports an average; nobody knows which condition failed
- The same ambiguity is re-argued on every batch
afterOne workspace, compounding quality
- Collection, annotation, review, evaluation under one schema
- Agents propose boundaries from state; people confirm
- Every decision carries confidence, evidence, and a name
- Performance by slice, traced to source episodes
- Ambiguities resolved once and pushed to every open task
Built with robotics teams, for robotics data.
Mobile manipulation
Sub-task boundaries, grasp outcomes, and object tracks across long episodes.
ExploreHumanoids & legged
Whole-body keypoints, contact events, terrain segmentation, gait-aware evaluation.
ExploreWarehouse & logistics
SKU-level coverage, exception mining, managed capacity for peak season.
ExploreAutonomous mobile robots
Obstacle classes, dynamic agent tracks, evaluation by lighting and crowd density.
ExploreTeleoperation programs
Operator quality tracking and demonstration filtering before training.
ExploreSim-to-real transfer
Shared ontology, gap metrics per condition, randomization guidance.
ExploreFour principles behind every feature.
Humans decide
Agents recommend. A person confirms. No label enters the dataset any other way.
Provenance is not optional
Confidence, evidence, and a name on every decision — exportable, auditable, permanent.
One schema, lab to fleet
The research dataset and the production pipeline never need a migration.
Quiet software
The best data tool is the one annotators stop noticing.
Short, opinionated reading for data teams.
Written from the problems we see most often — episode structure, guideline design, condition-based evaluation, and oversight for agentic labeling.
All guidesStructuring robot episodes for learning
Episodes, attempts, sub-tasks — and why boundaries should anchor to state, not appearance.
Evaluating policies by condition, not by average
Why 92% success can still mean 60% under glare, and how to build frozen suites.
Human oversight in agentic labeling
Recommend, never apply. Show evidence with confidence. Watch acceptance rates.
Frequently asked
Do agents ever change my data without approval?
No. Recommendations are never applied automatically. Every change is confirmed by a person and logged with a reason.
Which data types are supported?
2D images, 3D data, LiDAR point clouds, and video, plus synchronized telemetry such as joint states and force-torque readings.
Can I start in SaaS and add managed capacity later?
Yes — it's the most common path. Ontologies, guidelines, and quality history carry over unchanged.
Where is my data stored?
In the region you choose, or in your own storage under a hybrid deployment.
Your workspace, or ours — or both.
Self-serve SaaS
For teams who want full in-workspace control
- Your annotators, your guidelines, your queue
- Agent recommendations inside every task
- Dashboards, exports, and API access
Managed services
For teams who want end-to-end execution
- We staff, train, and run the annotation workforce
- Domain expert review included
- Delivery against agreed quality SLAs
Hybrid
For teams who want to mix both
- Keep sensitive or expert tasks in-house
- Burst volume to our managed team
- One dataset, one quality standard
Bring your first dataset. We'll show you what the agents find.
A 30-minute walkthrough on your own data, or on a sample robotics dataset if yours isn't ready yet.