Why we started Agentuor

Robot learning teams were spending more time moving data between tools than improving it. We built the platform we wished existed.

Why we started Agentuor

Every robotics team we worked with had the same shape of problem. The recorder lived on the robot. Labels came back from a vendor in a format that needed converting. Quality checks lived in a spreadsheet. Evaluation lived in a notebook that only one engineer could run. Between each of those steps, context leaked: nobody could say why a label had been chosen, which frames had been disputed, or whether the failing test scenario was even in the training data.

None of these tools were bad. They were built for a different kind of data — flat images, simple classes, short clips — and robotics data is none of those things. Episodes are minutes long. Sensors are many and must stay synchronized. Outcomes are ambiguous and consequential. A bad label does not lower a benchmark score; it shows up as a robot doing something it should not.

The observation that became a company

In the spring of 2025 we noticed that the highest-leverage decisions in a robotics data program were being made by the people with the least support: annotators placing a boundary on a long episode, reviewers adjudicating an occlusion, engineers choosing which scenario to record next. Each of them was working from partial information, because the information that would have helped was in a different tool.

The idea was simple to state. Put the whole data journey — collection, annotation, review, evaluation — under one schema, and let software agents watch the work as it happens. Not to replace the people, but to hand them the context they were missing: here is how similar frames were labeled, here is where two annotators disagreed, here is the evaluation slice this episode belongs to.

Agents recommend. People decide. The dataset remembers.

What we committed to

Three commitments shaped everything that followed. First, no recommendation is ever applied without a person confirming it. Second, every decision carries provenance — the evidence, the confidence, the guideline version, and a name. Third, the schema a research lab starts with is the same schema a production fleet runs on, so growth never requires a migration.

These sound like values. They are also engineering constraints, and they turned out to be the constraints that made the rest of the design fall into place.

Where we are

Agentuor is based in Orcera, up in the Sierra de Segura in Andalucía. We are a small team working closely with a handful of early robotics partners on mobile manipulation, humanoid, and logistics data. Over the coming months we will use this blog to write about the problems we are solving, the decisions we are making, and the things we get wrong along the way.

If you are building embodied AI and recognize the shape of the problem above, we would like to talk.

Published May 20, 2025 · All posts

Newer →The anatomy of a robot episode

Working on the same problems?

We'd like to see your data.