---
title: Roboto to LeRobot Contract
sidebar:
  order: 58
---
The `roboto-to-lerobot-v2_1` and `roboto-to-lerobot-v3_0` actions are driven by a **contract YAML** stored in the **invocation dataset** — the dataset the action is invoked on. The contract tells the action which topics to read from each recording, how to extract values from each message, how to align streams onto a common timeline, and what transforms to apply along the way. Every [event](/docs/learn/concepts#events-section) in the input [collection](/docs/learn/concepts#collections-section) becomes one episode in the output LeRobot dataset, and the contract is applied identically to every episode.

This page is the full schema reference. For a step-by-step walkthrough of authoring a contract and running the action, see [Convert to LeRobot](/docs/user-guides/convert-to-lerobot).

## Full example

The contract below exercises the full schema: a camera stream, numeric observations that concatenate onto a shared key, a role-bound spec, per-stream alignment and transforms, an action stream, and a task channel. Each piece is unpacked in the sections that follow. (For a minimal starting point, see the [user guide](/docs/user-guides/convert-to-lerobot) instead.)

```yaml
name: pick_place
version: 1
fps: 20
robot_type: dual_arm
action_lead_steps: 1     # read actions one frame ahead of observations

observations:
  # Camera -> a video feature; the image: block routes it to the video pipeline.
  - key: observation.images.exo
    topic: /camera/exo/image_raw/compressed
    type: sensor_msgs/msg/CompressedImage
    image:
      resize: [480, 640]   # [height, width]
    align: {method: nearest, tolerance_ms: 100}

  # Left arm -> the first three columns of observation.state.
  - key: observation.state
    topic: /left_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [left_1, left_2, left_3]
    align: {method: hold, tolerance_ms: 100}
    transforms:
      - type: butterworth_lowpass_causal   # causal: reproducible at inference
        cutoff_hz: 5
        order: 2
        fs_hz: 50        # freeze the design rate into the contract

  # Right arm -> concatenated onto the same feature (shared key).
  - key: observation.state
    topic: /right_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [right_1, right_2, right_3]
    align: {method: hold, tolerance_ms: 100}

  # Base velocity, bound by role instead of topic.
  # No align: -> hold with the auto-bounded tolerance.
  - key: observation.base_velocity
    role: base_twist
    type: geometry_msgs/msg/Twist
    selector:
      names:         [linear.x, angular.z]
      lerobot_names: [base_vx, base_wz]

actions:
  - key: action
    topic: /left_arm/joint_commands
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [left_cmd_1, left_cmd_2, left_cmd_3]
    align: {method: linear, tolerance_ms: null}   # null = no tolerance bound
    transforms:
      - type: resample_uniform
        stage: pre                  # pre only; must be set explicitly
      - type: butterworth_lowpass   # zero-phase; offline-only
        stage: pre
        cutoff_hz: 30
        order: 2
      - type: finite_difference     # stage defaults to "post"

tasks:
  - key: prompt          # optional; defaults to the topic name
    topic: /task/prompt
    type: std_msgs/msg/String
```

## Top-level fields

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `name` | string | `"contract"` | Logical contract name. Recorded with each episode for traceability. |
| `version` | int | `1` | Contract version. Bump it when you change the schema in a way that breaks downstream consumers. |
| `fps` | float | `20.0` | Target sampling frequency, in Hz. A reference timeline of one frame per `1/fps` seconds is built at this rate, and every observation and action is aligned onto it. |
| `robot_type` | string | `null` | Free-form label written into the LeRobot dataset metadata. |
| `action_lead_steps` | int | `0` | Shift action sampling `N` frames into the future to compensate for control delay. `0` means `action[t]` is read at the same instant as `observation[t]`. |
| `observations` | list | `[]` | Observation streams. Entries with an `image:` block become video features; entries without become numeric observation features. See [Observation specs](#observation-specs). |
| `actions` | list | `[]` | Action streams. Become numeric action features. See [Action specs](#action-specs). |
| `tasks` | list | `[]` | Optional task channels used to derive each episode's task label. See [Task specs](#task-specs). |

## Observation specs

Each entry in `observations:` describes one stream. There are two flavors, distinguished by whether an `image:` block is present:

- **Numeric observations** (no `image:` block) become a single LeRobot feature per `key`. Multiple entries that share a `key` are concatenated along the feature axis — see [Selectors and lerobot\_names](#selectors-and-lerobot_names).
- **Video / image observations** (with an `image:` block) become a `dtype: video` feature. See [Video and image specs](#video-and-image-specs).

| Field | Required | Description |
| --- | --- | --- |
| `key` | yes | LeRobot feature key. Conventionally `observation.state`, `observation.images.<name>`, and so on. |
| `topic` | yes | ROS topic name to read the stream from. Not required if the spec binds by `role:` instead; see [Role-based binding](#role-based-binding). |
| `type` | yes | Message type string used to select the decoder. See [Supported message types](#supported-message-types). |
| `selector` | no | `{names: [...], lerobot_names: [...]}`. Selects which fields the decoder extracts and how they are labeled. See [Selectors and lerobot\_names](#selectors-and-lerobot_names). |
| `image` | no | Image / video options, chiefly `resize: [H, W]`. Its presence routes the spec to the video pipeline. See [Video and image specs](#video-and-image-specs). |
| `align` | no | `{method, tolerance_ms}`. Defaults to `hold` with the auto-bounded tolerance. See [Alignment](#alignment). |
| `transforms` | no | A list of transforms applied to the stream. See [Transforms](#transforms). |

## Video and image specs

An entry under `observations:` becomes a video feature when it carries an `image:` block. The output dtype is always `video` (HWC `uint8` RGB at the configured resize).

```yaml
observations:
  - key: observation.images.exo
    topic: /camera/exo/image_raw/compressed
    type: sensor_msgs/msg/CompressedImage
    image:
      resize: [480, 640]   # [height, width]
    align: {method: nearest, tolerance_ms: 100}
```

The `image:` block currently accepts a single field:

| Field | Required | Description |
| --- | --- | --- |
| `resize` | yes | `[height, width]` in pixels. Frames are resized to this exact shape. |

A camera that publishes a compressed video stream rather than per-message stills is declared the same way, with its own `type:` — see [Compressed video](/docs/reference/roboto-to-lerobot-contract#lerobot-compressed-video).

Restrictions on image streams:

- `align.method: linear` is rejected — use `hold` or `nearest`.
- Depth encodings are not yet supported. Supported raw `sensor_msgs/msg/Image` encodings are `rgb8`, `bgr8`, `mono8`, `rgba8`, `bgra8`, and `8UC1`. An `image.depth:` block is currently rejected at contract load; depth support is planned for a future release.

## Action specs

Action specs follow the same shape as numeric observations — the same fields, with no `image:` block.

```yaml
actions:
  - key: action
    topic: /robot/joint_commands
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2]
      lerobot_names: [arm_1, arm_2]
    align: {method: hold, tolerance_ms: 100}
```

As with observations, multiple entries that share a `key` are concatenated along the feature axis. Set `topic:` and `type:` directly on the action spec, as above.

Alignment on action streams:

- All four alignment methods apply — `hold` (the default), `nearest`, `linear`, and `none`. The inference caveat that limits observation streams does not apply to actions; see the note under [Alignment](#alignment) for why every method is safe here.
- To shift action sampling relative to observations — for example, to compensate for control delay — set `action_lead_steps` at the top level rather than per spec. See [Top-level fields](#top-level-fields).

## Task specs [#task-specs]

`tasks:` is optional and lets the contract derive a per-episode LeRobot task label from a message stream — for example, a language prompt published on a topic.

```yaml
tasks:
  - key: prompt          # optional; defaults to the topic name
    topic: /task/prompt
    type: std_msgs/msg/String
```

A task spec sets `topic:` and `type:` — use `std_msgs/msg/String`; other message types fail when the stream is decoded — plus an optional `key:`, which defaults to the topic name.

For each episode, the task label is resolved in the following order:

1. The event's `task` metadata field, when present and non-empty. This is an explicit per-episode annotation and always wins; the `tasks:` block is not consulted.
2. Otherwise, the payload of the first `std_msgs/msg/String` message that falls within the episode's own time window, taken from the **first** declared task spec. If the contract declares more than one task spec, only the first is used and the others are ignored (a warning names them).
3. Otherwise, the literal string `"default"` — when there is no `task` metadata, no `tasks:` block, and no in-window message.

## Selectors and lerobot\_names

`selector:` controls which fields the decoder pulls out of each message:

```yaml
selector:
  names:         [joint_1, joint_2, joint_3]   # what to extract
  lerobot_names: [arm_1,   arm_2,   arm_3]     # how to label them (optional)
```

Two rules to keep in mind:

1. **The meaning of** `names` **depends on the message type.** For `JointState` they are joint names (with an optional `position.<joint>` / `velocity.<joint>` / `effort.<joint>` prefix; a bare name defaults to `position`). For `Imu`, `Odometry`, and `Twist` they are dotted paths into the message (for example, `linear.x`). See [Supported message types](#supported-message-types) for the full list.
2. **The** `lerobot_names` **field is optional but strongly recommended.** Without it, a feature is labeled `<topic>/<selector_name>`, which is verbose and ties feature names to your ROS topology. With it, you get clean LeRobot-facing names. `lerobot_names` must have the same length as `names`, and every `lerobot_name` must be unique across the entire contract.

When multiple specs share a `key` (for example, two `observation.state` entries from two arms), all of their `lerobot_names` are concatenated in order to name the combined feature.

```yaml
observations:
  # Left arm -> the first three columns of observation.state
  - key: observation.state
    topic: /left_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [left_1, left_2, left_3]

  # Right arm -> concatenated onto the same feature
  - key: observation.state
    topic: /right_arm/joint_states
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [right_1, right_2, right_3]
```

The combined `observation.state` feature then has six columns, labeled `left_1, left_2, left_3, right_1, right_2, right_3` in that order.

## Alignment

Each stream is merged onto the reference timeline (one frame per `1/fps` seconds) using its `align:` block:

```yaml
align: {method: hold, tolerance_ms: 100}
```

| Method | Behavior |
| --- | --- |
| `hold` | Last-observation-carried-forward (backward as-of join). Good for slow state signals. **Default.** |
| `nearest` | Pick the closest sample in time, in either direction. Good for cameras and high-rate signals. |
| `linear` | Linearly interpolate between the two bracketing samples. Numeric streams only — not allowed for images. |
| `none` | Exact timestamp matches only; everything else becomes `NaN`. |

:::note
**Reproducing alignment at inference.** If you plan to deploy the trained policy for online, streaming inference, only `hold` and `nearest` can be reproduced sample-by-sample on a live input; `linear` and `none` need the complete recording and are therefore offline-only — the same distinction as non-causal vs causal [Transforms](#transforms). Use `hold` or `nearest` for any **observation** stream whose processing must match between training and deployment. **Action** streams are exempt — a deployed policy emits actions rather than aligning them — so all four methods are always fine there.
:::

`tolerance_ms` is the maximum gap, in milliseconds, between the reference frame timestamp and the matched sample before the result becomes `NaN`. Set it to roughly one period of your slowest signal — for example, `100` for a 10 Hz signal.

| Tolerance\_ms | Behavior |
| --- | --- |
| omitted (or `align:` omitted entirely) | Auto-bounded default of `max(2/fps, 50 ms)`. This is usually a reasonable starting point. |
| `null` | Unlimited — always carry forward, or always pick the nearest sample, however far away it is. |
| `0` | **Rejected.** In older contracts `0` meant "unlimited"; that spelling collided with "zero tolerance" and is no longer accepted. Use `null` for unbounded, or omit `tolerance_ms` for the auto-bounded default. |
| a positive number | That bound, in milliseconds, unchanged. |

:::note
The legacy field names `strategy` and `tol_ms` are no longer accepted. Use `method` and `tolerance_ms`.
:::

## Transforms

Transforms are applied per stream, in the order they appear, before the stream becomes part of the LeRobot dataset:

```yaml
actions:
  - key: action
    topic: /robot/joint_commands
    type: sensor_msgs/msg/JointState
    selector: {names: [j1, j2], lerobot_names: [arm_1, arm_2]}
    transforms:
      - type: resample_uniform
        stage: pre
      - type: butterworth_lowpass
        stage: pre
        cutoff_hz: 30
        order: 2
      - type: finite_difference     # stage defaults to "post"
```

### Stages

A transform's `stage` says whether it runs before or after **alignment** — the step that merges each raw stream onto the shared reference timeline of one frame per `1/fps` seconds (see [Alignment](#alignment)). A `pre` transform therefore sees the raw, irregularly timed samples; a `post` transform sees the regular, `fps`-rate timeline.

| Stage | When it runs | Notes |
| --- | --- | --- |
| `pre` | Before alignment, on the raw stream. | Operates on the raw sample timestamps. Use it for resampling and pre-alignment filtering. |
| `post` | After alignment, on the reference timeline. | Operates at `fps`. **This is the default** when `stage:` is omitted. |

### Built-in transforms

| Type | Stage | Causal | Params | Description |
| --- | --- | --- | --- | --- |
| `resample_uniform` | `pre` only | no | — | Resamples an irregularly sampled signal onto a uniform grid (the median inter-sample interval) using nearest-neighbor. Always set `stage: pre` explicitly — this transform cannot run after alignment, and the `stage:` default is `post`. |
| `butterworth_lowpass` | any | no | `cutoff_hz` (required), `order` (default `2`) | Zero-phase Butterworth low-pass filter. No phase distortion, but each output sample depends on both past and future samples. |
| `butterworth_lowpass_causal` | any | yes | `cutoff_hz` (required), `order` (default `2`), `fs_hz` (optional) | Causal (forward-only) Butterworth low-pass filter. Each output sample depends only on past samples, at the cost of some phase lag. |
| `finite_difference` | any | no | — | Frame-to-frame difference (`out[i] = in[i+1] - in[i]`). The last row is repeated to keep the shape constant. |

**Causal vs non-causal transforms.** A **causal** transform computes each output sample from the current and past samples only, so the identical operation can also run online — sample-by-sample at inference time — and reproduce exactly what training saw. A **non-causal** transform additionally looks at future samples, so it can only ever run offline over a complete recording. The `Causal` column above marks which is which. When a stream feeds a deployed policy and the same processing must run at both training and inference time, choose a causal transform — for example `butterworth_lowpass_causal`, the causal counterpart of `butterworth_lowpass`. When only offline processing matters — for example, cleaning an observation that plays no role in a deployed policy — a non-causal transform is fine, and because a zero-phase filter such as `butterworth_lowpass` introduces no phase lag, it is often preferable.

**Cutoff bounds.** `cutoff_hz` must be strictly between `0` and the Nyquist frequency (half the sample rate). For a `stage: pre` transform the sample rate is the measured raw topic rate; for `stage: post` it is `fps`.

**Sample-rate override (causal filter only).** `fs_hz` declares the sample rate the filter is designed for, overriding the rate the transform would otherwise use. For a `stage: pre` causal filter, declaring `fs_hz` is what makes the filter reproducible during deployment: it freezes the design rate into the contract so offline and online both build identical filter coefficients. If the declared `fs_hz` deviates from the stream's actual measured rate by more than 10%, a warning names both numbers.

**Ordering.** When you chain a `stage: pre` transform that changes the time axis (`resample_uniform`) with later filters, place `resample_uniform` first and any low-pass filter after it. The sample rate passed to the later transforms is recomputed automatically from the resampled timestamps.

:::note
The built-in transform library will be expanded in future releases.
:::

## Role-based binding

Instead of pinning a spec to a specific `topic:`, you can bind it by **role** — a logical label you attach to the source file. This helps when the same logical stream lives under different topic names across recordings: you bind the spec to a stable role once, and each recording maps that role to whatever its actual topic happens to be.

To use it, set `role:` on the spec in place of `topic:`, then tag each source file with a matching role by setting a `role` field in the file's [Roboto metadata](/docs/learn/concepts#files-section) to the same value. For example, a spec with `role: arm_joints` binds to the file whose metadata has `role: arm_joints`. You can set this metadata manually — in the web UI, CLI, or SDK — so no companion action is required, though one may be provided to populate roles automatically.

```yaml
observations:
  - key: observation.state
    role: arm_joints          # instead of topic:
    type: sensor_msgs/msg/JointState
    selector:
      names:         [joint_1, joint_2, joint_3]
      lerobot_names: [arm_1, arm_2, arm_3]
```

Rules and failure modes:

- A spec must set **exactly one** of `topic:` or `role:`. Setting both, or neither, is rejected at contract load.
- **Image and video streams cannot be role-bound** — bind them by `topic:`.
- Within a dataset, **exactly one file** may carry a given role. Zero matching files, or more than one, stops the run with an error naming the role.
- The matched file must have exactly one topic whose fields cover the spec's `selector.names`; zero or multiple covering topics is an error.

## Supported message types

`type:` selects the decoder. The selector semantics depend on the type.

### Image and video

| `type:` | Required `image:` fields |
| --- | --- |
| `sensor_msgs/msg/CompressedImage` | `resize` |
| `sensor_msgs/msg/Image` | `resize` |
| `video`, `avi_video`, `mp4_video` | `resize`. File-backed video, one still frame per row. |
| `foxglove_msgs/msg/CompressedVideo` | `resize`. A compressed video stream (H.264, H.265, VP9, or AV1). See [Compressed video](#lerobot-compressed-video). |

Supported `sensor_msgs/msg/Image` encodings: `rgb8`, `bgr8`, `mono8`, `rgba8`, `bgra8`, and `8UC1`.

### Compressed video [#lerobot-compressed-video]

Topics that Roboto ingestion tags `compressedVideo` store one encoded video *access unit* per message rather than a standalone still image, so a frame in the middle of a group of pictures (GOP) can only be decoded together with the frames back to its keyframe. The action handles that: it decodes each episode's whole time range in one pass, walking back up to 10 seconds to find the keyframe that anchors the range.

Declare the topic's real schema name. Three spellings are accepted — `foxglove_msgs/msg/CompressedVideo`, `foxglove_msgs/CompressedVideo`, and `foxglove.CompressedVideo`. Declaring a compressed-video topic as `sensor_msgs/msg/CompressedImage` is rejected with an explanatory error rather than silently producing broken frames.

Apart from the `type:`, the spec is an ordinary video spec — an `image:` block carrying `resize`, plus an `align:` block:

```yaml
observations:
  - key: observation.images.external
    topic: /camera/exo/video
    type: foxglove_msgs/msg/CompressedVideo
    image:
      resize: [480, 640]   # [height, width]
    align: {method: nearest, tolerance_ms: 150}

  - key: observation.images.gripper
    topic: /camera/hand/video
    type: foxglove_msgs/msg/CompressedVideo
    image:
      resize: [480, 640]
    align: {method: nearest, tolerance_ms: 150}
```

Notes and limits:

- The codec is read from each message's `format` field. A stream in a codec other than H.264, H.265, VP9, or AV1 fails the conversion rather than dropping the camera.
- `resize` is applied as the frames are decoded, which keeps an episode's frames in memory at the target resolution rather than at the source resolution.
- Frames whose keyframe is unreachable are skipped with a warning. A range where *nothing* decodes fails the conversion.

### Numeric streams

| `type:` | What `selector.names` means |
| --- | --- |
| `sensor_msgs/msg/JointState` | Joint name. Optional `position.<joint>` / `velocity.<joint>` / `effort.<joint>` prefix; a bare name defaults to `position`. Without `names`: all joint positions in message order. |
| `trajectory_msgs/msg/JointTrajectory` | Joint name. Each trajectory point becomes its own sample at `header.stamp + time_from_start`. Intended for action streams. |
| `control_msgs/msg/MultiDOFCommand` | DOF name. Optional `values.<dof>` / `values_dot.<dof>` prefix; a bare name defaults to `values`. Without `names`: all `values` followed by all `values_dot`. |
| `sensor_msgs/msg/Imu` | Dotted path into the message, for example `orientation.x` or `angular_velocity.z`. Without `names`: returns `[quat, ang_vel, lin_acc]` (10 values). |
| `nav_msgs/msg/Odometry` | Dotted path. Without `names`: returns `[pos.xyz, quat.xyzw]` (7 values). |
| `geometry_msgs/msg/Twist` | Dotted path, for example `linear.x`. Without `names`: returns `[linear.xyz, angular.xyz]` (6 values). |
| `std_msgs/msg/Float32MultiArray` | not used; the full `data` array is emitted as `float32`. |
| `std_msgs/msg/Float64MultiArray` | not used; full `data` as `float64`. |
| `std_msgs/msg/Int32MultiArray` | not used; full `data` as `int32`. |
| `std_msgs/msg/Float32` / `Float64` / `Int32` / `Int64` | not used; a single-element array containing `data`. |
| `std_msgs/msg/String` | not used; emitted as a string. Intended for `tasks:`. |
| `string_typed_msg` | Series index keys, each read as a float. `names` is required for this type. |
| `anymal_msgs/AnymalState` | ANYmal joint name into `joints`. Optional `position.<joint>` / `velocity.<joint>` / `acceleration.<joint>` / `effort.<joint>` prefix; a bare name defaults to `position`. Without `names`: all joint positions in message order. |
| `series_elastic_actuator_msgs/SeActuatorReadings` | ANYmal joint name, mapped **positionally** — the message carries no per-actuator name — onto the fixed order `LF/RF/LH/RH × HAA/HFE/KFE`. Optional field-path prefix, for example `commanded.velocity.<joint>` or `state.joint_position.<joint>`; a bare name defaults to `commanded.position`. Without `names`: all 12 `commanded.position`. |

## Validation rules and common errors

The contract loader checks a few things up front. When it rejects a contract, the error message names the offending spec.

- **Binding:** every observation and action spec must set exactly one of `topic:` or `role:` (see [Role-based binding](#role-based-binding)). Image and video streams must use `topic:`.
- **Image alignment:** image streams cannot use `align.method: linear`.
- **Compressed video:** a topic that stores compressed video must be declared with a compressed-video `type:`. Declaring it as `sensor_msgs/msg/CompressedImage` — or as a compressed-video spelling that has no decoder — stops the run, naming the offending spec's `key` (see [Compressed video](/docs/reference/roboto-to-lerobot-contract#lerobot-compressed-video)).
- **Names:** `lerobot_names` must match `names` in length, and every `lerobot_name` in the contract must be globally unique. A duplicate name names both conflicting specs; a length mismatch prints both lists.
- **Video keys:** every video / image spec must have its own `key`. Two video specs sharing a `key` are rejected — only numeric specs concatenate onto a shared `key` (see [Selectors and lerobot\_names](#selectors-and-lerobot_names)).
- **Alignment method:** `align.method` must be one of `hold`, `nearest`, `linear`, or `none`.
- **Zero tolerance:** `align.tolerance_ms: 0` is rejected — see [Alignment](#alignment).
- **Legacy keys:** the `strategy` and `tol_ms` alignment keys are rejected — use `method` and `tolerance_ms`. A legacy `publish:` block on an action spec is likewise rejected.
- **Required fields:** a missing or wrong-typed required field on a spec (for example, no `key` or no `type`) raises an error naming that spec's position and key, such as `observations[2] (key 'observation.state'): ...`.
- **At runtime,** every topic referenced by the contract must be present on at least one file in the dataset, otherwise the action stops before doing any work.
