Roboto to LeRobot Contract
The roboto-to-lerobot-v2_1 and roboto-to-lerobot-v3_0 actions are driven by a contract YAML stored in the invocation dataset — the dataset the action is invoked on. The contract tells the action which topics to read from each recording, how to extract values from each message, how to align streams onto a common timeline, and what transforms to apply along the way. Every event in the input collection becomes one episode in the output LeRobot dataset, and the contract is applied identically to every episode.
This page is the full schema reference. For a step-by-step walkthrough of authoring a contract and running the action, see Convert to LeRobot.
Full example
The contract below exercises the full schema: a camera stream, numeric observations that concatenate onto a shared key, a role-bound spec, per-stream alignment and transforms, an action stream, and a task channel. Each piece is unpacked in the sections that follow. (For a minimal starting point, see the user guide instead.)
name: pick_place
version: 1
fps: 20
robot_type: dual_arm
action_lead_steps: 1 # read actions one frame ahead of observations
observations:
# Camera -> a video feature; the image: block routes it to the video pipeline.
- key: observation.images.exo
topic: /camera/exo/image_raw/compressed
type: sensor_msgs/msg/CompressedImage
image:
resize: [480, 640] # [height, width]
align: {method: nearest, tolerance_ms: 100}
# Left arm -> the first three columns of observation.state.
- key: observation.state
topic: /left_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [left_1, left_2, left_3]
align: {method: hold, tolerance_ms: 100}
transforms:
- type: butterworth_lowpass_causal # causal: reproducible at inference
cutoff_hz: 5
order: 2
fs_hz: 50 # freeze the design rate into the contract
# Right arm -> concatenated onto the same feature (shared key).
- key: observation.state
topic: /right_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [right_1, right_2, right_3]
align: {method: hold, tolerance_ms: 100}
# Base velocity, bound by role instead of topic.
# No align: -> hold with the auto-bounded tolerance.
- key: observation.base_velocity
role: base_twist
type: geometry_msgs/msg/Twist
selector:
names: [linear.x, angular.z]
lerobot_names: [base_vx, base_wz]
actions:
- key: action
topic: /left_arm/joint_commands
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [left_cmd_1, left_cmd_2, left_cmd_3]
align: {method: linear, tolerance_ms: null} # null = no tolerance bound
transforms:
- type: resample_uniform
stage: pre # pre only; must be set explicitly
- type: butterworth_lowpass # zero-phase; offline-only
stage: pre
cutoff_hz: 30
order: 2
- type: finite_difference # stage defaults to "post"
tasks:
- key: prompt # optional; defaults to the topic name
topic: /task/prompt
type: std_msgs/msg/String
Top-level fields
| Field | Type | Default | Description |
|---|---|---|---|
name |
string | "contract" |
Logical contract name. Recorded with each episode for traceability. |
version |
int | 1 |
Contract version. Bump it when you change the schema in a way that breaks downstream consumers. |
fps |
float | 20.0 |
Target sampling frequency, in Hz. A reference timeline of one frame per 1/fps seconds is built at this rate, and every observation and action is aligned onto it. |
robot_type |
string | null |
Free-form label written into the LeRobot dataset metadata. |
action_lead_steps |
int | 0 |
Shift action sampling N frames into the future to compensate for control delay. 0 means action[t] is read at the same instant as observation[t]. |
observations |
list | [] |
Observation streams. Entries with an image: block become video features; entries without become numeric observation features. See Observation specs. |
actions |
list | [] |
Action streams. Become numeric action features. See Action specs. |
tasks |
list | [] |
Optional task channels used to derive each episode’s task label. See Task specs. |
Observation specs
Each entry in observations: describes one stream. There are two flavors, distinguished by whether an image: block is present:
- Numeric observations (no
image:block) become a single LeRobot feature perkey. Multiple entries that share akeyare concatenated along the feature axis — see Selectors and lerobot_names. - Video / image observations (with an
image:block) become adtype: videofeature. See Video and image specs.
| Field | Required | Description |
|---|---|---|
key |
yes | LeRobot feature key. Conventionally observation.state, observation.images.<name>, and so on. |
topic |
yes | ROS topic name to read the stream from. Not required if the spec binds by role: instead; see Role-based binding. |
type |
yes | Message type string used to select the decoder. See Supported message types. |
selector |
no | {names: [...], lerobot_names: [...]}. Selects which fields the decoder extracts and how they are labeled. See Selectors and lerobot_names. |
image |
no | Image / video options, chiefly resize: [H, W]. Its presence routes the spec to the video pipeline. See Video and image specs. |
align |
no | {method, tolerance_ms}. Defaults to hold with the auto-bounded tolerance. See Alignment. |
transforms |
no | A list of transforms applied to the stream. See Transforms. |
Video and image specs
An entry under observations: becomes a video feature when it carries an image: block. The output dtype is always video (HWC uint8 RGB at the configured resize).
observations:
- key: observation.images.exo
topic: /camera/exo/image_raw/compressed
type: sensor_msgs/msg/CompressedImage
image:
resize: [480, 640] # [height, width]
align: {method: nearest, tolerance_ms: 100}
The image: block currently accepts a single field:
| Field | Required | Description |
|---|---|---|
resize |
yes | [height, width] in pixels. Frames are resized to this exact shape. |
A camera that publishes a compressed video stream rather than per-message stills is declared the same way, with its own type: — see Compressed video.
Restrictions on image streams:
align.method: linearis rejected — useholdornearest.- Depth encodings are not yet supported. Supported raw
sensor_msgs/msg/Imageencodings arergb8,bgr8,mono8,rgba8,bgra8, and8UC1. Animage.depth:block is currently rejected at contract load; depth support is planned for a future release.
Action specs
Action specs follow the same shape as numeric observations — the same fields, with no image: block.
actions:
- key: action
topic: /robot/joint_commands
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2]
lerobot_names: [arm_1, arm_2]
align: {method: hold, tolerance_ms: 100}
As with observations, multiple entries that share a key are concatenated along the feature axis. Set topic: and type: directly on the action spec, as above.
Alignment on action streams:
- All four alignment methods apply —
hold(the default),nearest,linear, andnone. The inference caveat that limits observation streams does not apply to actions; see the note under Alignment for why every method is safe here. - To shift action sampling relative to observations — for example, to compensate for control delay — set
action_lead_stepsat the top level rather than per spec. See Top-level fields.
Task specs
tasks: is optional and lets the contract derive a per-episode LeRobot task label from a message stream — for example, a language prompt published on a topic.
tasks:
- key: prompt # optional; defaults to the topic name
topic: /task/prompt
type: std_msgs/msg/String
A task spec sets topic: and type: — use std_msgs/msg/String; other message types fail when the stream is decoded — plus an optional key:, which defaults to the topic name.
For each episode, the task label is resolved in the following order:
- The event’s
taskmetadata field, when present and non-empty. This is an explicit per-episode annotation and always wins; thetasks:block is not consulted. - Otherwise, the payload of the first
std_msgs/msg/Stringmessage that falls within the episode’s own time window, taken from the first declared task spec. If the contract declares more than one task spec, only the first is used and the others are ignored (a warning names them). - Otherwise, the literal string
"default"— when there is notaskmetadata, notasks:block, and no in-window message.
Selectors and lerobot_names
selector: controls which fields the decoder pulls out of each message:
selector:
names: [joint_1, joint_2, joint_3] # what to extract
lerobot_names: [arm_1, arm_2, arm_3] # how to label them (optional)
Two rules to keep in mind:
- The meaning of
namesdepends on the message type. ForJointStatethey are joint names (with an optionalposition.<joint>/velocity.<joint>/effort.<joint>prefix; a bare name defaults toposition). ForImu,Odometry, andTwistthey are dotted paths into the message (for example,linear.x). See Supported message types for the full list. - The
lerobot_namesfield is optional but strongly recommended. Without it, a feature is labeled<topic>/<selector_name>, which is verbose and ties feature names to your ROS topology. With it, you get clean LeRobot-facing names.lerobot_namesmust have the same length asnames, and everylerobot_namemust be unique across the entire contract.
When multiple specs share a key (for example, two observation.state entries from two arms), all of their lerobot_names are concatenated in order to name the combined feature.
observations:
# Left arm -> the first three columns of observation.state
- key: observation.state
topic: /left_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [left_1, left_2, left_3]
# Right arm -> concatenated onto the same feature
- key: observation.state
topic: /right_arm/joint_states
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [right_1, right_2, right_3]
The combined observation.state feature then has six columns, labeled left_1, left_2, left_3, right_1, right_2, right_3 in that order.
Alignment
Each stream is merged onto the reference timeline (one frame per 1/fps seconds) using its align: block:
align: {method: hold, tolerance_ms: 100}
| Method | Behavior |
|---|---|
hold |
Last-observation-carried-forward (backward as-of join). Good for slow state signals. Default. |
nearest |
Pick the closest sample in time, in either direction. Good for cameras and high-rate signals. |
linear |
Linearly interpolate between the two bracketing samples. Numeric streams only — not allowed for images. |
none |
Exact timestamp matches only; everything else becomes NaN. |
tolerance_ms is the maximum gap, in milliseconds, between the reference frame timestamp and the matched sample before the result becomes NaN. Set it to roughly one period of your slowest signal — for example, 100 for a 10 Hz signal.
| Tolerance_ms | Behavior |
|---|---|
omitted (or align: omitted entirely) |
Auto-bounded default of max(2/fps, 50 ms). This is usually a reasonable starting point. |
null |
Unlimited — always carry forward, or always pick the nearest sample, however far away it is. |
0 |
Rejected. In older contracts 0 meant “unlimited”; that spelling collided with “zero tolerance” and is no longer accepted. Use null for unbounded, or omit tolerance_ms for the auto-bounded default. |
| a positive number | That bound, in milliseconds, unchanged. |
Transforms
Transforms are applied per stream, in the order they appear, before the stream becomes part of the LeRobot dataset:
actions:
- key: action
topic: /robot/joint_commands
type: sensor_msgs/msg/JointState
selector: {names: [j1, j2], lerobot_names: [arm_1, arm_2]}
transforms:
- type: resample_uniform
stage: pre
- type: butterworth_lowpass
stage: pre
cutoff_hz: 30
order: 2
- type: finite_difference # stage defaults to "post"
Stages
A transform’s stage says whether it runs before or after alignment — the step that merges each raw stream onto the shared reference timeline of one frame per 1/fps seconds (see Alignment). A pre transform therefore sees the raw, irregularly timed samples; a post transform sees the regular, fps-rate timeline.
| Stage | When it runs | Notes |
|---|---|---|
pre |
Before alignment, on the raw stream. | Operates on the raw sample timestamps. Use it for resampling and pre-alignment filtering. |
post |
After alignment, on the reference timeline. | Operates at fps. This is the default when stage: is omitted. |
Built-in transforms
| Type | Stage | Causal | Params | Description |
|---|---|---|---|---|
resample_uniform |
pre only |
no | — | Resamples an irregularly sampled signal onto a uniform grid (the median inter-sample interval) using nearest-neighbor. Always set stage: pre explicitly — this transform cannot run after alignment, and the stage: default is post. |
butterworth_lowpass |
any | no | cutoff_hz (required), order (default 2) |
Zero-phase Butterworth low-pass filter. No phase distortion, but each output sample depends on both past and future samples. |
butterworth_lowpass_causal |
any | yes | cutoff_hz (required), order (default 2), fs_hz (optional) |
Causal (forward-only) Butterworth low-pass filter. Each output sample depends only on past samples, at the cost of some phase lag. |
finite_difference |
any | no | — | Frame-to-frame difference (out[i] = in[i+1] - in[i]). The last row is repeated to keep the shape constant. |
Causal vs non-causal transforms. A causal transform computes each output sample from the current and past samples only, so the identical operation can also run online — sample-by-sample at inference time — and reproduce exactly what training saw. A non-causal transform additionally looks at future samples, so it can only ever run offline over a complete recording. The Causal column above marks which is which. When a stream feeds a deployed policy and the same processing must run at both training and inference time, choose a causal transform — for example butterworth_lowpass_causal, the causal counterpart of butterworth_lowpass. When only offline processing matters — for example, cleaning an observation that plays no role in a deployed policy — a non-causal transform is fine, and because a zero-phase filter such as butterworth_lowpass introduces no phase lag, it is often preferable.
Cutoff bounds. cutoff_hz must be strictly between 0 and the Nyquist frequency (half the sample rate). For a stage: pre transform the sample rate is the measured raw topic rate; for stage: post it is fps.
Sample-rate override (causal filter only). fs_hz declares the sample rate the filter is designed for, overriding the rate the transform would otherwise use. For a stage: pre causal filter, declaring fs_hz is what makes the filter reproducible during deployment: it freezes the design rate into the contract so offline and online both build identical filter coefficients. If the declared fs_hz deviates from the stream’s actual measured rate by more than 10%, a warning names both numbers.
Ordering. When you chain a stage: pre transform that changes the time axis (resample_uniform) with later filters, place resample_uniform first and any low-pass filter after it. The sample rate passed to the later transforms is recomputed automatically from the resampled timestamps.
Role-based binding
Instead of pinning a spec to a specific topic:, you can bind it by role — a logical label you attach to the source file. This helps when the same logical stream lives under different topic names across recordings: you bind the spec to a stable role once, and each recording maps that role to whatever its actual topic happens to be.
To use it, set role: on the spec in place of topic:, then tag each source file with a matching role by setting a role field in the file’s Roboto metadata to the same value. For example, a spec with role: arm_joints binds to the file whose metadata has role: arm_joints. You can set this metadata manually — in the web UI, CLI, or SDK — so no companion action is required, though one may be provided to populate roles automatically.
observations:
- key: observation.state
role: arm_joints # instead of topic:
type: sensor_msgs/msg/JointState
selector:
names: [joint_1, joint_2, joint_3]
lerobot_names: [arm_1, arm_2, arm_3]
Rules and failure modes:
- A spec must set exactly one of
topic:orrole:. Setting both, or neither, is rejected at contract load. - Image and video streams cannot be role-bound — bind them by
topic:. - Within a dataset, exactly one file may carry a given role. Zero matching files, or more than one, stops the run with an error naming the role.
- The matched file must have exactly one topic whose fields cover the spec’s
selector.names; zero or multiple covering topics is an error.
Supported message types
type: selects the decoder. The selector semantics depend on the type.
Image and video
type: |
Required image: fields |
|---|---|
sensor_msgs/msg/CompressedImage |
resize |
sensor_msgs/msg/Image |
resize |
video, avi_video, mp4_video |
resize. File-backed video, one still frame per row. |
foxglove_msgs/msg/CompressedVideo |
resize. A compressed video stream (H.264, H.265, VP9, or AV1). See Compressed video. |
Supported sensor_msgs/msg/Image encodings: rgb8, bgr8, mono8, rgba8, bgra8, and 8UC1.
Compressed video
Topics that Roboto ingestion tags compressedVideo store one encoded video access unit per message rather than a standalone still image, so a frame in the middle of a group of pictures (GOP) can only be decoded together with the frames back to its keyframe. The action handles that: it decodes each episode’s whole time range in one pass, walking back up to 10 seconds to find the keyframe that anchors the range.
Declare the topic’s real schema name. Three spellings are accepted — foxglove_msgs/msg/CompressedVideo, foxglove_msgs/CompressedVideo, and foxglove.CompressedVideo. Declaring a compressed-video topic as sensor_msgs/msg/CompressedImage is rejected with an explanatory error rather than silently producing broken frames.
Apart from the type:, the spec is an ordinary video spec — an image: block carrying resize, plus an align: block:
observations:
- key: observation.images.external
topic: /camera/exo/video
type: foxglove_msgs/msg/CompressedVideo
image:
resize: [480, 640] # [height, width]
align: {method: nearest, tolerance_ms: 150}
- key: observation.images.gripper
topic: /camera/hand/video
type: foxglove_msgs/msg/CompressedVideo
image:
resize: [480, 640]
align: {method: nearest, tolerance_ms: 150}
Notes and limits:
- The codec is read from each message’s
formatfield. A stream in a codec other than H.264, H.265, VP9, or AV1 fails the conversion rather than dropping the camera. resizeis applied as the frames are decoded, which keeps an episode’s frames in memory at the target resolution rather than at the source resolution.- Frames whose keyframe is unreachable are skipped with a warning. A range where nothing decodes fails the conversion.
Numeric streams
type: |
What selector.names means |
|---|---|
sensor_msgs/msg/JointState |
Joint name. Optional position.<joint> / velocity.<joint> / effort.<joint> prefix; a bare name defaults to position. Without names: all joint positions in message order. |
trajectory_msgs/msg/JointTrajectory |
Joint name. Each trajectory point becomes its own sample at header.stamp + time_from_start. Intended for action streams. |
control_msgs/msg/MultiDOFCommand |
DOF name. Optional values.<dof> / values_dot.<dof> prefix; a bare name defaults to values. Without names: all values followed by all values_dot. |
sensor_msgs/msg/Imu |
Dotted path into the message, for example orientation.x or angular_velocity.z. Without names: returns [quat, ang_vel, lin_acc] (10 values). |
nav_msgs/msg/Odometry |
Dotted path. Without names: returns [pos.xyz, quat.xyzw] (7 values). |
geometry_msgs/msg/Twist |
Dotted path, for example linear.x. Without names: returns [linear.xyz, angular.xyz] (6 values). |
std_msgs/msg/Float32MultiArray |
not used; the full data array is emitted as float32. |
std_msgs/msg/Float64MultiArray |
not used; full data as float64. |
std_msgs/msg/Int32MultiArray |
not used; full data as int32. |
std_msgs/msg/Float32 / Float64 / Int32 / Int64 |
not used; a single-element array containing data. |
std_msgs/msg/String |
not used; emitted as a string. Intended for tasks:. |
string_typed_msg |
Series index keys, each read as a float. names is required for this type. |
anymal_msgs/AnymalState |
ANYmal joint name into joints. Optional position.<joint> / velocity.<joint> / acceleration.<joint> / effort.<joint> prefix; a bare name defaults to position. Without names: all joint positions in message order. |
series_elastic_actuator_msgs/SeActuatorReadings |
ANYmal joint name, mapped positionally — the message carries no per-actuator name — onto the fixed order LF/RF/LH/RH × HAA/HFE/KFE. Optional field-path prefix, for example commanded.velocity.<joint> or state.joint_position.<joint>; a bare name defaults to commanded.position. Without names: all 12 commanded.position. |
Validation rules and common errors
The contract loader checks a few things up front. When it rejects a contract, the error message names the offending spec.
- Binding: every observation and action spec must set exactly one of
topic:orrole:(see Role-based binding). Image and video streams must usetopic:. - Image alignment: image streams cannot use
align.method: linear. - Compressed video: a topic that stores compressed video must be declared with a compressed-video
type:. Declaring it assensor_msgs/msg/CompressedImage— or as a compressed-video spelling that has no decoder — stops the run, naming the offending spec’skey(see Compressed video). - Names:
lerobot_namesmust matchnamesin length, and everylerobot_namein the contract must be globally unique. A duplicate name names both conflicting specs; a length mismatch prints both lists. - Video keys: every video / image spec must have its own
key. Two video specs sharing akeyare rejected — only numeric specs concatenate onto a sharedkey(see Selectors and lerobot_names). - Alignment method:
align.methodmust be one ofhold,nearest,linear, ornone. - Zero tolerance:
align.tolerance_ms: 0is rejected — see Alignment. - Legacy keys: the
strategyandtol_msalignment keys are rejected — usemethodandtolerance_ms. A legacypublish:block on an action spec is likewise rejected. - Required fields: a missing or wrong-typed required field on a spec (for example, no
keyor notype) raises an error naming that spec’s position and key, such asobservations[2] (key 'observation.state'): .... - At runtime, every topic referenced by the contract must be present on at least one file in the dataset, otherwise the action stops before doing any work.