Skip to content
Roboto
Esc
↑↓navigate↵open⌘Jpreview
On this page

roboto.experimental.ingest

Describe what an uploaded file carries, so the platform can register it without opening the file.

A caller declares the topics a file contributes data to, what each of them carries, and which files a read of each opens; the types that compose those files into sessions live in roboto.experimental.sessions.

Submodules

Package Contents

DeclaredTimelineSource

roboto.experimental.ingest.DeclaredTimelineSource#View Source

One timeline source a file’s topic data carries, with the bounds it spans in this file.

Field

class roboto.experimental.ingest.Field(/, **data)#View Source

Bases: pydantic.BaseModel

One column of a topic’s data, identified by name, type, and unit.

path lists the names from the schema root down to this field, so a nested field’s path extends its parent’s. name and path state one fact twice (the last path element is the field’s name), so either may be omitted and derives from the other: a top-level field needs only name, and a nested field needs only path. A vector column (e.g. a LeRobot observation.state feature) is expressed as a parent field holding the array plus one child field per named element.

This is what a caller declares. SchemaFieldRecord is what the platform returns for a field it has stored, and carries the identifiers it assigns.

Parameters

data Any

Attributes

Field.canonical_data_type

Roboto’s normalized type for the field, used for cross-format reads and visualization.

Field.data_type

data_type str = None #

Native type of the field as recorded by the source format (e.g. "float32").

Field.model_config

model_config #

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

Field.name

name str = '' #

Name of the field. Defaults to the last element of path when only path is given; at least one of name and path must be declared.

Field.path

path list[str] = None #

The names from the schema root down to this field; a nested field’s path extends its parent’s path. Omitted, None, or empty defaults to [name]; when given, every element must be non-empty and the last element must equal name.

Field.unit

unit str | None = None #

Unit of the field’s values (e.g. "rad"). A field typed Timestamp must carry a TimeUnit value (one of "s", "ms", "us", "ns").

FileTopicDeclaration

class roboto.experimental.ingest.FileTopicDeclaration(/, **data)#View Source

Bases: TopicDeclaration

One topic a file contributes data to, declared on the file itself rather than inside a session.

Adds to TopicDeclaration the wall-clock instant the data was captured at, which a topic declared inside a session takes from the entry enclosing it.

Parameters

data Any

Attributes

FileTopicDeclaration.anchor_ns

anchor_ns roboto.time._EpochNanosecondsFromTime | None = None #

Optional wall-clock anchor for the data this declaration names: the real-world time, in nanoseconds since the Unix epoch, at which that data’s time 0 occurred. Must fall after the Unix epoch, and be small enough to fit in the signed 64-bit integer the platform stores it in. Also accepts any roboto.time.Time at runtime, read as roboto.time.to_epoch_nanoseconds() reads it; convert with that function first to satisfy a type checker. When omitted, the data keeps the anchor it already carries from an earlier declaration, and keeps an offset of 0 when it carries none: its timestamps read exactly as declared.

An anchor covers the whole slice named by data_range, not this topic alone, so every topic declared over that slice moves to it; two declarations sharing a slice may not name different anchors.

MAX_FILES_AND_TOPICS_PER_REQUEST

roboto.experimental.ingest.MAX_FILES_AND_TOPICS_PER_REQUEST = 500#View Source

Cap on how many files and topics one request may name, counting each file entry and each topic declaration in it.

Each file and topic named costs the platform another round of writes, and a request has a fixed amount of time to finish them all, so a request naming more than this is refused outright rather than left to run out of time; split a larger batch across several calls. The request models below and in roboto.experimental.sessions apply the cap when the request body is constructed, and the platform checks it again on every call that declares files or topics, so a body built without these models is held to the same cap.

MESSAGE_ENVELOPE_TIMELINE_SOURCES

roboto.experimental.ingest.MESSAGE_ENVELOPE_TIMELINE_SOURCES: dict[str, tuple[roboto.domain.topics.record.TimelineSourceKind, str]]#View Source

Stored kind and stored name of each timeline source a container stamps on its records, keyed by declared kind.

These sources sit in the message envelope rather than in the topic’s data columns. The stored kind says which of the envelope’s two timestamps a source is, log time or publish time, and reads select a source by its stored name.

MCAP log time and MP4 presentation time each name the timeline their container stamps on every record it holds: for MCAP, the instant the recorder wrote the record; for MP4, the instant the frame is shown. The platform stores one such timeline per topic schema, so the two share a stored kind, differ only in the name a read selects them by, and cannot both be declared on one topic.

The request models in this module check each declaration against the kind and name in this table, and the platform stores the declaration under exactly those.

McapLogTimeSource

class roboto.experimental.ingest.McapLogTimeSource(/, **data)#View Source

Bases: _DeclaredSourceBase

MCAP’s log time: when the recorder wrote each record to the file.

MCAP carries these timestamps in its message envelope rather than in a data column, so this source names no field.

Parameters

data Any

Attributes

McapLogTimeSource.kind

kind Literal['mcap_log_time'] = 'mcap_log_time' #

Discriminator identifying this entry as MCAP log time.

McapPublishTimeSource

class roboto.experimental.ingest.McapPublishTimeSource(/, **data)#View Source

Bases: _DeclaredSourceBase

MCAP’s publish time: when each message was published on the bus.

MCAP carries these timestamps in its message envelope rather than in a data column, so this source names no field.

Parameters

data Any

Attributes

McapPublishTimeSource.kind

kind Literal['mcap_publish_time'] = 'mcap_publish_time' #

Discriminator identifying this entry as MCAP publish time.

Mp4PresentationTimeSource

class roboto.experimental.ingest.Mp4PresentationTimeSource(/, **data)#View Source

Bases: _DeclaredSourceBase

Presentation time: when each frame is shown, relative to the start of the media.

These timestamps sit in the message envelope rather than in a data column, so this source names no field. The platform stores it as a message-envelope timeline source named presentation_time, which is the name a read selects it by.

Data declared with this source is readable from an MCAP of encoded frames whose messages’ log times are the presentation times, when the topic lists that MCAP in TopicDeclaration.representations. The MP4 itself cannot be listed, since no storage format a representation can state describes it. A topic with no representations is registered, and a read of it returns no rows, as TopicDeclaration.representations describes.

Parameters

data Any

Attributes

Mp4PresentationTimeSource.kind

kind Literal['mp4_presentation_time'] = 'mp4_presentation_time' #

Discriminator identifying this entry as MP4 presentation time.

RepresentationDeclaration

class roboto.experimental.ingest.RepresentationDeclaration(/, **data)#View Source

Bases: pydantic.BaseModel

One representation of a topic’s data: a file a read of the topic can open, and how that file holds the data.

Listed in TopicDeclaration.representations when a topic is declared, and in TopicRepresentations.representations when a topic’s representations are replaced. The platform does not open the file when it is listed, so what this states is taken on trust, and a read of the topic that picks this representation opens file_id and decodes it as stated.

Every file named by a representation that is untransformed, or whose every transformation is an encode, must hold the topic’s rows:

  1. At the same positions in each of those files, so a TopicDeclaration.data_range names the same rows whichever one a read opens. In an MCAP, positions count only the topic’s own messages.
  2. Decoding to the topic’s declared schema, carrying the timestamps its timeline sources describe: not rebased, not converted to other units, not rounded.

A representation whose file breaks either is accepted, and reads of it return the wrong rows or fail.

A representation covers the whole topic unless it states a field_path, in which case it covers that field and everything under it. Its file still holds every row, each with its timestamp: a read that takes fields from several files pairs those files row by row, and fails when their row numbers or timestamps differ.

Four things identify a representation among those of one topic: what it covers (the whole topic, or one field), its storage_format, its content_format and its transformations. Over each part of a file it is declared on, a topic holds at most one representation per combination of the four: two with the same combination cannot be listed together, and one declared later takes the earlier one’s place.

A read decodes a PARQUET representation’s file as Parquet; the file needs no .parquet extension. A read can decode an MCAP representation’s file when all three of the following hold:

  1. The file is chunked and carries a summary section indexing those chunks.
  2. Exactly one of the file’s channels carries the topic’s name.
  3. That channel’s schema is in an encoding the platform can decode: ros1msg, ros2msg, ros2idl, omgidl, jsonschema, or json.

A recording holding several topics is therefore read one topic at a time, each from the channel carrying its name; the file’s other channels are neither decoded nor checked.

Parameters

data Any

Attributes

RepresentationDeclaration.content_format

content_format str | None = None #

Format of the data inside the container, such as "jpeg" for re-encoded images or "compressedVideo" for passed-through video frames, or None when unspecified. A read names it through content_format.

RepresentationDeclaration.field_path

field_path list[str] | None = None #

Field of the topic’s schema this representation covers, as the names from the schema root down to it. None, the same as leaving it out, makes this a representation of the whole topic; an empty list and an empty name are refused.

The path must equal the path of a field the topic’s schema declares. That field may have fields under it, and the representation then covers those too. A level of the schema that only the paths of deeper fields pass through, with no field declared at it, cannot be named.

A read takes each field from the representation covering the deepest field that contains it, so a representation of one field supplies that field whatever its transformations, and the representations of the whole topic supply the rest. A read asking for untransformed data (RepresentationSelector.raw()) leaves out every representation that states a transformation.

RepresentationDeclaration.file_id

file_id str = None #

File a read opens to get the topic’s data: the file the topic is declared on, when its own bytes hold the data, or another file of the org, such as a per-topic MCAP converted out of a PX4 ULog. Its status must be Available, and the caller must be able to edit it. A read through the SDK or the web app downloads this file with the reader’s own download permission on it, and Roboto’s AI tools read it for anyone who can read the topic.

RepresentationDeclaration.model_config

model_config #

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

RepresentationDeclaration.storage_format

Container file_id holds the data in. Within one request, every representation naming a file must state the same format for it, whichever of the request’s topics, files or sessions lists it; a request stating two formats for one file is refused when it is built.

RepresentationDeclaration.transformations

transformations list[str] = None #

The transformations applied to the topic’s data to produce what file_id holds, in the order applied, as "<kind>:<param>" descriptors such as ["downsample:0.7", "encode:jpeg"]; TransformationKind lists the kinds. Empty for untransformed data. A read names them through transformations, and with no selector the platform prefers the representation with the fewest.

A representation can be read by row position when it is untransformed or every transformation is an encode. One with any other transformation, such as a downsample, cannot, and a read of a topic declared over a data_range never uses it: on such a topic, everything it covers must also be covered by representations that can.

Within one topic, every representation naming a file must state the same transformations for it: a read that takes two representations from one file and finds their transformations differ raises RobotoReadPlanExecutionException with kind inconsistent-scan-tasks-on-file.

Schema

class roboto.experimental.ingest.Schema(/, **data)#View Source

Bases: pydantic.BaseModel

The structure of one topic’s data: the columns it carries.

However a schema is produced, whether hand-written field by field or converted from a source format’s own metadata, the registered result is the same: schemas are content-addressed server-side. Identity covers every attribute of every field (name, path, source data type, canonical type, and unit), so identical declarations collapse to a single stored schema no matter how many times they are repeated, while declarations differing in any field attribute are stored separately.

A column a timeline source reads is declared by typing it Timestamp with a TimeUnit unit; nothing else marks it. Which of a topic’s timeline sources reads fall back to is not part of the schema: it is stated on timeline_sources and can be changed later, so the same columns are one schema no matter which source is preferred.

This is what a caller declares, so it carries no checksum: the platform computes that from the fields. TopicSchemaRecord is the stored schema the platform returns, carrying that checksum and the identifiers it assigns.

Parameters

data Any

Attributes

Schema.fields

fields list[Field] = None #

Declared columns of the topic’s data. At least one is required, and every field’s path must be unique within the schema.

Schema.model_config

model_config #

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

Schema.name

name str | None = None #

Informational label for the schema (often the topic name). Not part of schema identity.

SchemaFieldSource

class roboto.experimental.ingest.SchemaFieldSource(/, **data)#View Source

Bases: _DeclaredSourceBase

Timestamps read from a column of the topic’s own data, such as a timestamp field.

Parameters

data Any

Attributes

SchemaFieldSource.field_path

field_path list[str] = None #

Path of the field holding this source’s timestamps: the names from the schema root down to that field. The field must be declared on the topic’s schema and typed Timestamp.

SchemaFieldSource.kind

kind Literal['field'] = 'field' #

Discriminator identifying this entry as timestamps read from a data column.

TopicDeclaration

class roboto.experimental.ingest.TopicDeclaration(/, **data)#View Source

Bases: pydantic.BaseModel

A topic a file contributes data to.

The platform registers the data from the declaration alone, never opening the file, so the declaration states the topic’s name, the structure of its rows, the bounds of every timeline source those rows carry, and which part of the file they occupy.

List a file’s topic declarations on the SessionFile for that file; that entry anchors everything it declares. To register a file’s topic data without naming a session, through declare_topics(), build a FileTopicDeclaration instead; with no enclosing entry to anchor it, that form carries its own anchor.

The file a topic is declared on is the one its data belongs to: the data’s slices, bounds and anchors, and the windows sessions hold it over, are all stated against that file. Which files a read opens to get the data is a separate statement, representations, so data can belong to a file the platform cannot decode (a PX4 ULog, a ROS .bag, a CSV) and be read from files converted out of it.

Parameters

data Any

Attributes

TopicDeclaration.data_range

data_range roboto.domain.topics.record.DataRange | None = None #

The part of the file this topic’s data occupies, as (start, end): start is the first covered position, and end is one past the last. Set this when one file packs a topic’s data into slices, such as a single episode inside a LeRobot v3 data file. Omit it when the data covers whatever encloses this declaration: the slice claimed by the SessionFile carrying it, or the whole file when the topics are declared on the file itself through declare_topics(). A range set inside a SessionFile must sit within that entry’s own data_range.

Values are in the file’s own units: stored-row positions (counted from 0), or nanoseconds of the file’s media time for video. Each file uses exactly one of the two, so the pair needs no unit marker; which one applies follows from the file’s format. A read applies the range to the file it opens, and it opens only files named by representations that can be read by row position, as RepresentationDeclaration.transformations describes: each of those files must hold the same rows at the same positions. Over a range, everything a topic’s representations cover must be covered by ones that can be read by row position; the platform refuses the declaration otherwise.

TopicDeclaration.model_config

model_config #

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

TopicDeclaration.representations

representations list[RepresentationDeclaration] = None #

The files a read of this topic’s data opens, each with how it holds that data. List the file the topic is declared on here, naming its file_id, when its own bytes are readable; list other files when the data is read from them, such as the per-topic MCAPs converted out of a PX4 ULog. A topic may list several, and a read picks among them by RepresentationSelector.

A topic with no representations is still registered: its data counts toward the bounds of every session holding the file, and a read of it returns no rows, or is refused when the read’s RepresentationSelector sets a criterion. A video, or a proprietary log with no file converted out of it, is declared that way.

Redeclaring the topic over the same part of the file adds the representations listed to the ones it has. A listed representation takes the place of the stored one that covers what it covers (the whole topic, or the same field) under the same storage format, content format and transformations, and of every stored one that names the same file, whatever that one covers; every other stored representation stays. Redeclaring the topic with the new files a conversion wrote therefore replaces each stored representation with the listed one that differs from it only in its file_id; redeclaring it with a file uploaded again in another storage format replaces the stored representation that names that file; and resending the same declaration leaves the stored representations as they are. To remove a representation, or to replace a topic’s representations outright, use set_representations().

A representation of one field is stored against that field of the schema it was declared under. While one is stored, the platform refuses a declaration of the topic over the same part of the file with a changed topic_schema, unless the declaration lists a representation that takes the stored one’s place. To change the schema, list the representation of that field again in the same declaration, or remove it first with set_representations().

Refused when the model is built:

  1. Two representations covering the same thing (both the whole topic, or the same field) under the same storage format, content format and transformations.
  2. Two representations naming one file under different transformations.
  3. A representation whose field_path names no field of topic_schema.
  4. A PARQUET representation on a topic declaring any timeline source but SchemaFieldSource: the others take their timestamps from the message envelope, which a Parquet file does not have.

TopicDeclaration.timeline_sources

timeline_sources list[DeclaredTimelineSource] = None #

Timeline sources this file’s topic data carries, each with its own bounds. Data timestamped several ways (a message’s publish time and the time the recorder wrote it, say) declares one entry per source; data timestamped one way declares a list of one. Reads pick a source by name, and resolve to the one marked is_default_for_reads when they do not.

TopicDeclaration.topic_name

topic_name str = None #

Topic this file (or slice of it) contributes data to. Topic names are unique within an org; files declaring the same name contribute to the same topic.

TopicDeclaration.topic_schema

Structure of the topic’s data. Repeat the same schema on every file that uses it; the platform stores each distinct schema once, so repetition costs nothing extra.

TopicRepresentations

class roboto.experimental.ingest.TopicRepresentations(/, **data)#View Source

Bases: pydantic.BaseModel

The complete set of representations one topic’s data on a file is read from.

Handed to set_representations(), which leaves the topic, over the part of the file data_range names, with exactly these representations. RepresentationDeclaration states what the file each one names must hold.

Parameters

data Any

Attributes

TopicRepresentations.data_range

data_range roboto.domain.topics.record.DataRange | None = None #

The part of the file over which the topic’s representations are replaced, or None for data declared over the whole file. It must equal a range the topic is already declared over on the file: the topic’s TopicDeclaration.data_range, or, for a topic that stated none inside a SessionFile, that entry’s data_range.

TopicRepresentations.model_config

model_config #

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

TopicRepresentations.representations

representations list[RepresentationDeclaration] #

The representations the topic ends up with; its representations not listed are removed, those of one field included. An empty list removes every representation: the topic stays registered and still counts toward the bounds of every session holding the file, but reads of it return no rows. Two representations may not cover the same thing (both the whole topic, or the same field) under the same storage format, content format and transformations, or name one file under different transformations; both are refused when the model is built. set_representations() states what the platform refuses.

TopicRepresentations.topic_name

topic_name str = None #

Topic whose representations are replaced. It must already be declared on the file.

Was this page helpful?