Back to overview

PLATFORM

How Roborecs captures the data.

From a first-person human demonstration to a labelled, license-ready training episode: the pipeline, the seven capture channels, the Sofia facility, and the specialist models the corpus is built to train.

How it works

How a demonstration becomes training-ready data, the phase roadmap, and every channel a session records.

OPERATION PIPELINE

From human demonstration to robot training data.

01 / CAPTURE

Egocentric demonstrations

Trained operators record real two-handed tasks from a first-person view on a wearable head-and-wrist rig: synchronized RGB, depth, 3D hand pose and skeleton, motion, and audio. Scales linearly with trained operators, at a fraction of teleoperation’s cost per hour, no lab or robot needed to record.

HEAD + WRIST RIG · MULTI-VIEW RGB · SUB-MS SYNC
02 / SYNC & LABEL

Align and annotate

Sub-millisecond timestamps align every channel; faster streams are downsampled and slower ones interpolated. Each episode is segmented into tasks and steps, with object and contact tags, consent, and provenance.

SUB-MS SYNC · TASK + STEP LABELS · PROVENANCE
03 / DELIVER

License-ready episodes

Output: labelled training episodes in the format the ecosystem already uses, ready to drop into humanoid foundation-model pipelines.

GDPR · CONSENT · HDF5 / LeRobot v3

THE ROADMAP · CAPTURE TO CORPUS TO MODELS

One data engine, built in three turns.

STARTINGPHASE 1

Egocentric capture

First-person human demonstrations, captured on a wearable head-and-wrist rig. Far cheaper and more scalable than teleoperation for producing the base of every humanoid model.

The wide base of the corpus. Where we start.

NEXTPHASE 2

Deployment and the loop

Robots deployed on real factory tasks will capture the deployment failures a wearable rig never sees. Each cycle improves the robots and compounds the corpus with data no rig-only capture can produce.

The loop that compounds the data lead.

ROADMAPPHASE 3

Specialist models

The corpus compounds and stays ours. When it reaches scale, and with the ML research leadership we bring on for this phase, we train the vertical models it powers, compatible with NVIDIA Isaac GR00T N1.7, Physical Intelligence π0, and Hugging Face LeRobot.

The data moat becomes a model moat.

TASK EXECUTION

This is what robotics used to be.

Action
-
Cube
-
From → To
-
Cycles
-
Frames rendered
-
CAPTURE PRIMITIVESRB-CAP / FUSION MODULE

Seven synchronized channels, merged into one labelled frame. The multimodal supply that humanoid robots train on.

Roborecs is building the pipeline that captures these channels first-person and fuses them into training frames in the format the industry already uses.

SPEC
240 Hz CAPTURE · 30 Hz OUTPUT
7 CHANNELS · 1 NEXT · EU

Capture specification: the target channels, their rates, and the fused output. Illustrative, not live or recorded sensor data.

01IN RIG
3D Hand Pose & Skeleton
SK-01 · RB-CAP
POSE60 Hz
02IN RIG
3D Scene Geometry
DG-02 · RB-CAP
DEPTH30 Hz
03IN RIG
Object Tracking & Velocity
OT-03 · RB-CAP
VELOCITY100 Hz
04IN RIG
Scene Segmentation
SG-04 · RB-CAP
SEGMENTS30 Hz
05IN RIG
Sound & Audio Environment
SA-06 · RB-CAP
DB48 kHz
06IN RIG
Spatial Localization
LC-08 · RB-CAP
POSITION60 Hz
NEXTGA-07 · RB-CAP
Gaze & Attention

Needs eye-tracking the camera rig does not have yet. Ships when the phase-2 evaluation layer opens.

FUSION PIPELINE · RB-MERGE / SENSOR FUSION
7 channels in. One labelled frame out. Synchronized to LeRobot-standard 30 Hz.

Each channel captures at its native rate. Sub-millisecond timestamps align every sample. Faster channels are downsampled; slower ones interpolated. The output is a single multimodal frame, 30 times per second.

SK-01
60 HzIN RIG
DG-02
30 HzIN RIG
OT-03
100 HzIN RIG
SG-04
30 HzIN RIG
SA-06
48 kHzIN RIG
GA-07
120 HzNEXT
LC-08
60 HzIN RIG
OUTPUT · MERGED FRAME
Labelled training frame
FORMATLeRobot v3 · HDF5
RATE240 Hz capture · 30 Hz output
COMPATIBLEGR00T · π0 · LeRobot
JURISDICTIONEU · GDPR
Modern humanoid foundation models, NVIDIA Isaac GR00T N1.7, Physical Intelligence π0, Hugging Face LeRobot, all train on multimodal physical data: vision, proprioception, motion, and depth. The bottleneck is supply. The internet has trillions of text tokens; humanoid robots have a fraction of the physical data they need. Roborecs supplies it, first-person and multimodal, in LeRobot-compatible format from the EU.
What ships

A Roborecs dataset is delivered as documented, training-ready episodes, not raw dumps. What is available depends on the capture phase.

Modalities
  • 7-camera RGB video
  • Depth (from multi-view)
  • IMU at 500 Hz
  • Spatial trajectories
  • Audio
  • Hand pose and skeleton
Annotations
  • Task and step segmentation
  • Object and interaction tags (pick, place, insert, route)
  • Contact and grasp events
  • Per-episode metadata (task, environment, operator)
  • Custom taxonomies for your model or benchmark
Delivery
  • LeRobot v3 / HDF5
  • Dataset cards and manifests
  • Consent and provenance artifact per episode
  • EU-jurisdiction delivery
  • Cloud bucket or encrypted transfer

Delivery spec for the capture program. Available modalities and annotations depend on capture phase.

REQUEST CAPTURE SPECIFICATION →7 channels · Sofia, Bulgaria · LeRobot v3 compatible · EU GDPR

The fidelity layer

The Sofia facility, and the multimodal pipeline behind training-ready data.

Rendering of the planned Roborecs data capture facility at Sofia Tech Park
THE FACILITY · RB-FAC

Built for physical AI, in Europe.

A purpose-built data capture facility planned for Sofia Tech Park, 1,100 m² at launch scaling to 5,000 m². Direct access to STEM talent, energy infrastructure, and EU logistics. Targeted operational from Q3 2027.

1,100 m² launch·5,000 m² planned·Sofia Tech Park·Q3 2027
TASK LIBRARY5 CATEGORIES · CUSTOM AVAILABLE

Bimanual industrial assembly, the task catalog.

The contact-rich, two-handed tasks that need real industrial demonstration, connector work, fastening, precision handling, to be captured in the industrial settings robots deploy into. Clients can commission task libraries across the categories below, with custom commissions for proprietary needs.

Flagship
Electronics & Precision Assembly
  • Connector mating & cable routing
  • Precision fastening with torque
  • PCB & component handling
High demand
Industrial & Warehouse
  • Pick & place assembly
  • Quality inspection
  • Bin sorting & packing
High value
Dexterous Fine Motor
  • Bimanual coordination
  • Tool manipulation
  • Fine alignment
Emerging
Human Interaction
  • Collaborative task handoff
  • Assistance & guiding
  • Safe proximity work
Available
Custom Commissions
  • OEM-defined task specs
  • Bespoke capture sessions
  • Full IP assignment
Don’t see your task?

Custom Commissions are available for OEM-specific task libraries. Full IP assignment. EU GDPR-compliant provenance from day one.

DISCUSS COMMISSION →
DEPLOYMENT · THE LOOPROBOT IN THE LOOP

The robot on the task.

This is the deployment end of the loop: a robot running a manipulation task. Move your cursor to guide it; in deployment we capture what happens as it works, the wins and the failures, and feed those back to improve the next version. That is where the corpus is designed to compound.

Deployment · closed loopconcept illustration
your inputrobot · executing
REC · EP 001
hold near object · grab  |  release on pad · place  |  scroll · depth  |  drag · orbit
shoulder_pitch_R·
elbow_pitch_R·
ee_accel·
contact_evt·
gripper_aperture·
operator_cmd·
rgb_cam·
depth_cam·

illustrative telemetry · derived live from demo kinematics, not recorded sensor data

Episodes recorded
0
Synchronized channels
8
Robot joint states
28 DOF · 240 Hz
SPECIALIST MODELS · THE COMPOUNDING LAYERSTATUS · ROADMAP

First the corpus. Then the models built on it.

Capture and deployment build one thing: a proprietary, action-labeled corpus of real industrial manipulation. Robot model architecture is commoditizing, open GR00T, open LeRobot, open π0. The data underneath it is not. What compounds is the loop, real deployments feeding the corpus, and the specialist models built on it.

CAPTURE
CORPUS
SPECIALIST MODEL
DEPLOYMENT
Every deployment feeds the next capture.

Nothing in the corpus exists until we capture it ourselves, one interaction at a time, and it grows with every task we run. Each turn of the loop widens the data lead.

STATUS · ROADMAP

No model trains on Roborecs data yet, and no Roborecs model ships today. The model layer opens on two gates: the corpus crosses the scale published results show a specialist policy needs, and we add the ML research leadership to build it. We would rather state the gates than imply a model, or a team, that does not exist yet.

NEUTRAL SUPPLY LAYER

Our specialist models are task policies, not a robot, and they run on any LeRobot-compatible stack. The data is the product: any corpus a customer licenses is carved out of our own model roadmap, so we never ship a policy that competes with a data customer. The policies we do build prove what the corpus can produce.

EXAMPLE POLICIES · ROADMAP

RB-VLA-01

Electronics assembly

A specialist vision-language-action policy for connector insertion, cable routing, and screwdriving. Post-trained on the multimodal channels of the assembly corpus, where a millimetre decides success. Delivered LeRobot-compatible.

RB-VLA-02

Precision kitting

Bin-to-fixture part placement under tight tolerance. The same corpus, a different task family.

The models sit downstream of the open foundation models. Post-trained on the corpus, delivered in the same LeRobot-compatible format. Compatible with, never competing with, NVIDIA Isaac GR00T N1.7 and Physical Intelligence π0.

NEXT

Run a pilot.

Tell us your robot and target tasks. We scope a capture program, agree a sample spec, and deliver an evaluation set before any volume commitment.