PLATFORM
How Roborecs captures the data.
From a first-person human demonstration to a labelled, license-ready training episode: the pipeline, the seven capture channels, the Sofia facility, and the specialist models the corpus is built to train.
How it works
How a demonstration becomes training-ready data, the phase roadmap, and every channel a session records.
From human demonstration to robot training data.
Egocentric demonstrations
Trained operators record real two-handed tasks from a first-person view on a wearable head-and-wrist rig: synchronized RGB, depth, 3D hand pose and skeleton, motion, and audio. Scales linearly with trained operators, at a fraction of teleoperation’s cost per hour, no lab or robot needed to record.
HEAD + WRIST RIG · MULTI-VIEW RGB · SUB-MS SYNCAlign and annotate
Sub-millisecond timestamps align every channel; faster streams are downsampled and slower ones interpolated. Each episode is segmented into tasks and steps, with object and contact tags, consent, and provenance.
SUB-MS SYNC · TASK + STEP LABELS · PROVENANCELicense-ready episodes
Output: labelled training episodes in the format the ecosystem already uses, ready to drop into humanoid foundation-model pipelines.
GDPR · CONSENT · HDF5 / LeRobot v3THE ROADMAP · CAPTURE TO CORPUS TO MODELS
One data engine, built in three turns.
Egocentric capture
First-person human demonstrations, captured on a wearable head-and-wrist rig. Far cheaper and more scalable than teleoperation for producing the base of every humanoid model.
The wide base of the corpus. Where we start.
Deployment and the loop
Robots deployed on real factory tasks will capture the deployment failures a wearable rig never sees. Each cycle improves the robots and compounds the corpus with data no rig-only capture can produce.
The loop that compounds the data lead.
Specialist models
The corpus compounds and stays ours. When it reaches scale, and with the ML research leadership we bring on for this phase, we train the vertical models it powers, compatible with NVIDIA Isaac GR00T N1.7, Physical Intelligence π0, and Hugging Face LeRobot.
The data moat becomes a model moat.
This is what robotics used to be.
Seven synchronized channels, merged into one labelled frame.
The multimodal supply that humanoid robots train on.
Roborecs is building the pipeline that captures these channels first-person and fuses them into training frames in the format the industry already uses.
Capture specification: the target channels, their rates, and the fused output. Illustrative, not live or recorded sensor data.
Needs eye-tracking the camera rig does not have yet. Ships when the phase-2 evaluation layer opens.
Each channel captures at its native rate. Sub-millisecond timestamps align every sample. Faster channels are downsampled; slower ones interpolated. The output is a single multimodal frame, 30 times per second.
A Roborecs dataset is delivered as documented, training-ready episodes, not raw dumps. What is available depends on the capture phase.
- 7-camera RGB video
- Depth (from multi-view)
- IMU at 500 Hz
- Spatial trajectories
- Audio
- Hand pose and skeleton
- Task and step segmentation
- Object and interaction tags (pick, place, insert, route)
- Contact and grasp events
- Per-episode metadata (task, environment, operator)
- Custom taxonomies for your model or benchmark
- LeRobot v3 / HDF5
- Dataset cards and manifests
- Consent and provenance artifact per episode
- EU-jurisdiction delivery
- Cloud bucket or encrypted transfer
Delivery spec for the capture program. Available modalities and annotations depend on capture phase.
The fidelity layer
The Sofia facility, and the multimodal pipeline behind training-ready data.

Built for physical AI, in Europe.
A purpose-built data capture facility planned for Sofia Tech Park, 1,100 m² at launch scaling to 5,000 m². Direct access to STEM talent, energy infrastructure, and EU logistics. Targeted operational from Q3 2027.
Bimanual industrial assembly, the task catalog.
The contact-rich, two-handed tasks that need real industrial demonstration, connector work, fastening, precision handling, to be captured in the industrial settings robots deploy into. Clients can commission task libraries across the categories below, with custom commissions for proprietary needs.
- →Connector mating & cable routing
- →Precision fastening with torque
- →PCB & component handling
- →Pick & place assembly
- →Quality inspection
- →Bin sorting & packing
- →Bimanual coordination
- →Tool manipulation
- →Fine alignment
- →Collaborative task handoff
- →Assistance & guiding
- →Safe proximity work
- →OEM-defined task specs
- →Bespoke capture sessions
- →Full IP assignment
Custom Commissions are available for OEM-specific task libraries. Full IP assignment. EU GDPR-compliant provenance from day one.
DISCUSS COMMISSION →The robot on the task.
This is the deployment end of the loop: a robot running a manipulation task. Move your cursor to guide it; in deployment we capture what happens as it works, the wins and the failures, and feed those back to improve the next version. That is where the corpus is designed to compound.
illustrative telemetry · derived live from demo kinematics, not recorded sensor data
First the corpus. Then the models built on it.
Capture and deployment build one thing: a proprietary, action-labeled corpus of real industrial manipulation. Robot model architecture is commoditizing, open GR00T, open LeRobot, open π0. The data underneath it is not. What compounds is the loop, real deployments feeding the corpus, and the specialist models built on it.
Nothing in the corpus exists until we capture it ourselves, one interaction at a time, and it grows with every task we run. Each turn of the loop widens the data lead.
No model trains on Roborecs data yet, and no Roborecs model ships today. The model layer opens on two gates: the corpus crosses the scale published results show a specialist policy needs, and we add the ML research leadership to build it. We would rather state the gates than imply a model, or a team, that does not exist yet.
Our specialist models are task policies, not a robot, and they run on any LeRobot-compatible stack. The data is the product: any corpus a customer licenses is carved out of our own model roadmap, so we never ship a policy that competes with a data customer. The policies we do build prove what the corpus can produce.
EXAMPLE POLICIES · ROADMAP
Electronics assembly
A specialist vision-language-action policy for connector insertion, cable routing, and screwdriving. Post-trained on the multimodal channels of the assembly corpus, where a millimetre decides success. Delivered LeRobot-compatible.
Precision kitting
Bin-to-fixture part placement under tight tolerance. The same corpus, a different task family.
The models sit downstream of the open foundation models. Post-trained on the corpus, delivered in the same LeRobot-compatible format. Compatible with, never competing with, NVIDIA Isaac GR00T N1.7 and Physical Intelligence π0.
NEXT
Run a pilot.
Tell us your robot and target tasks. We scope a capture program, agree a sample spec, and deliver an evaluation set before any volume commitment.