roboto.domain.topics.record
Module Contents
CanonicalDataType
Bases: enum.Enum
Normalized data types used across different robotics frameworks.
Well-known and simplified data types that provide a common vocabulary for describing message path data types across different frameworks and technologies. These canonical types are primarily used for UI rendering decisions and cross-platform compatibility.
The canonical types abstract away framework-specific details while preserving the essential characteristics needed for data processing and visualization.
References
- ROS 1 field types: http://wiki.ros.org/msg
- ROS 2 field types: https://docs.ros.org/en/iron/Concepts/Basic/About-Interfaces.html#field-types
- uORB: https://docs.px4.io/main/en/middleware/uorb.html#adding-a-new-topic
Example mappings:
float32->CanonicalDataType.Numberuint8[]->CanonicalDataType.Arraysensor_msgs/Image->CanonicalDataType.Imagegeometry_msgs/Pose->CanonicalDataType.Objectstd_msgs/Header->CanonicalDataType.Objectstring->CanonicalDataType.Stringchar->CanonicalDataType.Stringbool->CanonicalDataType.Booleanbyte->CanonicalDataType.Byte
Attributes
CanonicalDataType.Boolean
CanonicalDataType.Byte
CanonicalDataType.Categorical
Data that can take a limited, fixed set of values. To be interpreted correctly by Roboto clients, a MessagePathRecord with this type must have a "categories" metadata key on the MessagePathRecord, which must be the ordered list of values that the Categorical can take.
For example, a signal that is logged as either “off” or “on” could be represented as a Categorical with the metadata "categories"=["off", "on"]. This allows Roboto to map the value “off” to 0 and “on” to 1 –each corresponding to their index position in the metadata array– and therefore visualize these state transitions as a plot.
The default visual representation of Categorical data will be the same as String data, but the Roboto visualizer will be capable of rendering Categorical data in a plot.
CanonicalDataType.Image
Special purpose type for data that can be rendered as an image.
CanonicalDataType.LatDegFloat
Geographic point in degrees. E.g. 47.6749387 (used in ULog ver_data_format >= 2)
CanonicalDataType.LatDegInt
Geographic point in degrees, expressed as an integer. E.g. 317534036 (used in ULog ver_data_format < 2)
CanonicalDataType.LonDegFloat
Geographic point in degrees. E.g. 9.1445274 (used in ULog ver_data_format >= 2)
CanonicalDataType.LonDegInt
Geographic point in degrees, expressed as an integer. E.g. 1199146398 (used in ULog ver_data_format < 2)
CanonicalDataType.Number
CanonicalDataType.NumberArray
CanonicalDataType.String
CanonicalDataType.Timestamp
Time elapsed since the Unix epoch, identifying a single instant on the time-line. Roboto clients will look for a "unit" metadata key on the MessagePath record, and will assume “ns” if none is found. If the timestamp is in a different unit, add the following metadata to the MessagePath record: { "unit": "s"|"ms"|"us"|"ns" } The unit must be a known value from TimeUnit.
DataRange
A slice of one file’s contents, as (start, end) with start included and end excluded.
A position is expressed in whatever the file’s format uses to address its contents: stored-row positions counted from 0, or nanoseconds of media time for video. The pair alone does not say which of the two applies, so a reader takes that from the file’s format. The half-open form matches LeRobot’s dataset_from_index and dataset_to_index, so row ranges read from LeRobot episode metadata carry over unchanged.
FieldPath
A schema field’s path components, in order from the schema root to the leaf.
MessagePathMetadataWellKnown
Bases: roboto.compat.StrEnum
Well-known metadata key names (with well-known semantics) that may be set in metadata.
These are most often set by Roboto’s first-party ingestion actions and used by Roboto clients.
Attributes
MessagePathMetadataWellKnown.Categories
An ordered list of values that a Categorical can take.
Usage
"categories"=["off", "on"]"categories"=["left", "up", "right", "down"]
MessagePathMetadataWellKnown.ColumnName
The original name or path to this field in the source data schema. May differ from message_path if character substitutions were applied to conform to naming requirements.
Notes
- Use of this metadata field is soft-deprecated as of SDK v0.24.0.
- Prefer use of
source_pathandpath_in_schemainstead. Those attributes are now first-class fields onMessagePathRecordand can be specified viaAddMessagePathRequest.
MessagePathRecord
Bases: pydantic.BaseModel
Record representing a message path within a topic.
Defines a specific field or signal within a topic’s data schema, including its data type, metadata, and statistical information. Message paths use dot notation to specify nested attributes within complex message structures.
Message paths are the fundamental units for accessing individual data elements within time-series robotics data, enabling fine-grained analysis and visualization of specific signals or measurements.
Parameters
data AnyAttributes
MessagePathRecord.canonical_data_type
Normalized data type, used primarily internally by the Roboto Platform.
MessagePathRecord.created
MessagePathRecord.created_by
MessagePathRecord.data_type
‘Native’/framework-specific data type of the attribute at this path. E.g. “float32”, “uint8[]”, “geometry_msgs/Pose”, “string”.
MessagePathRecord.message_path
Dot-delimited path to the attribute within the datum record.
MessagePathRecord.message_path_id
MessagePathRecord.metadata
Key-value pairs to associate with this metadata for discovery and search, e.g. { ‘min’: ‘0.71’, ‘max’: ’1.77 }
MessagePathRecord.modified
MessagePathRecord.modified_by
MessagePathRecord.org_id
This message path’s organization ID, which is the organization ID of the containing topic.
MessagePathRecord.parents()
Logical message path ancestors of this path.
Usage
Given a deeply nested field root.sub_obj_1.sub_obj_2.leaf_field:
field = "root.sub_obj_1.sub_obj_2.leaf_field"
record = MessagePathRecord(message_path=field) # other fields omitted for brevity
print(record.parents())
# ['root.sub_obj_1.sub_obj_2', 'root.sub_obj_1', 'root']Parameters
delimiter strReturn type
Attributes
MessagePathRecord.path_in_schema
List of path components representing the field’s location in the original data schema. Unlike message_path, which must conform to Roboto-specific naming requirements and assumes dots separated path parts imply nested data, this preserves the exact path from the source data for accurate attribute access. This is expected to be the split representation of source_path.
MessagePathRecord.representations
Zero to many Representations of this MessagePath.
MessagePathRecord.source_path
The original name of this field in the source data schema. May differ from message_path if character substitutions were applied to conform to naming requirements.
This is the preferred field to use when specifying message_path_include or message_path_exclude to the get_data or get_data_as_df methods of Topic and Event.
MessagePathRecord.to_field_selection()
Translate this record into the FieldSelection the format decoders accept.
Return type
Attributes
MessagePathRecord.topic_id
MessagePathRepresentationMapping
Bases: pydantic.BaseModel
Mapping between message paths and their data representation.
Associates a set of message paths with a specific representation that contains their data. This mapping is used to efficiently locate and access data for specific message paths within topic representations.
Parameters
data AnyAttributes
MessagePathRepresentationMapping.message_paths
MessagePathRepresentationMapping.representation
MessagePathStatistic
Bases: enum.Enum
Statistics computed by Roboto in our standard ingestion actions.
Which of these a given message path actually carries depends on the ingestion path that produced it, so treat every one as optional: read them with metadata.get(...) or through the corresponding MessagePath property, both of which yield None when the statistic was never written, and never assume a missing value means zero. Indexing metadata directly raises KeyError for a statistic that was never written.
Attributes
MessagePathStatistic.Count
MessagePathStatistic.Max
MessagePathStatistic.Mean
MessagePathStatistic.Median
MessagePathStatistic.Min
MessagePathStatistic.P25
MessagePathStatistic.P75
MessagePathStatistic.P95
MessagePathStatistic.P99
MessagePathStatistic.Stddev
RepresentationRecord
Bases: pydantic.BaseModel
Record representing a data representation for topic content.
A representation is a pointer to processed topic data stored in a specific format and location. Representations enable efficient access to topic data by providing multiple storage formats optimized for different use cases.
Most message paths within a topic point to the same representation (e.g., an MCAP or Parquet file containing all topic data). However, some message paths may have multiple representations for analytics or preview formats.
Representations are versioned and associated with specific files or storage locations through the association field.
Parameters
data AnyAttributes
RepresentationRecord.association
Identifier and entity type with which this Representation is associated. E.g., a file, a database.
RepresentationRecord.created
RepresentationRecord.format
Content format descriptor for this representation. For image topics: the image encoding (e.g. “jpeg”, “png”) for simplified representations, or the ROS schema name (e.g. “sensor_msgs/Image”) for original/passthrough representations. None for non-image topics or legacy representations.
RepresentationRecord.modified
RepresentationRecord.representation_id
RepresentationRecord.storage_format
RepresentationRecord.topic_id
RepresentationRecord.transformations
Ordered list of transformation descriptors applied to produce this representation. Empty for original/passthrough representations.
Each entry is a "<kind>:<param>" string where <kind> is a TransformationKind member. Construct entries via TransformationKind.with_param() and parse them via TransformationKind.parse() to keep the vocabulary centralized.
Example: ["downsample:0.5", "encode:jpeg"]
RepresentationRecord.version
RepresentationSelector
Bases: pydantic.BaseModel
Criteria for selecting among multiple representations of the same data.
When a message path has multiple representations (e.g., both raw sensor data and a processed JPEG encoding), this is a hard filter: only matching representations qualify, and message paths with no matching representation are dropped from selection results — callers must handle empty or partial output.
Legacy carve-out for ``content_format``: representations with no format set (i.e., predating the field) are treated as matching any content_format request. This keeps older data accessible. When both an explicit format match and a legacy representation are available for the same message path, the explicit match wins.
Instances are immutable (frozen=True) so they can be safely shared — including as default arguments to methods like Topic.get_data().
Parameters
data AnyAttributes
RepresentationSelector.content_format
If set, only representations whose format field matches this value qualify (e.g., "jpeg"). Representations with no format also qualify under the legacy carve-out. None means no constraint.
RepresentationSelector.matches()
Check whether a representation satisfies this selector’s criteria.
A representation matches when each non-None selector field is satisfied. For content_format, representations with no format set are treated as matching (legacy carve-out — see class docstring).
Parameters
representation RepresentationRecordReturn type
Attributes
RepresentationSelector.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
RepresentationSelector.raw()
Select representations with no transformations applied (original data).
Return type
RepresentationSelector.select_representations()
Select one representation per message path that matches this selector.
When the API returns multiple representations for the same message paths (e.g., both a raw MCAP and a processed JPEG MCAP for an image topic), this method picks a matching representation for each path and deduplicates so each message path appears in exactly one mapping.
Non-matching representations are excluded. When both an explicit format match and a legacy representation (no format set) cover the same message path, the explicit match wins. Message paths covered by no matching representation are dropped — callers must handle empty or partial results.
Parameters
mappings list[MessagePathRepresentationMapping]All representation mappings, potentially with overlapping message paths.
Returns
Deduplicated mappings of message paths to matching representations. Empty if no representation matches.
Attributes
RepresentationSelector.transformations
If set, only representations whose transformations field matches exactly qualify. [] matches representations with no transformations (i.e., raw/original data). None means no constraint.
RepresentationStorageFormat
Bases: enum.Enum
Supported storage formats for topic data representations.
Defines the available formats for storing and accessing topic data within the Roboto platform. Each format has different characteristics and use cases.
SchemaFieldRecord
Bases: pydantic.BaseModel
A single field within a topic schema.
One entry per unique field path within a schema; field paths are deduplicated across topics that share the schema.
Parameters
data AnyAttributes
SchemaFieldRecord.canonical_data_type
Normalized data type used for cross-framework compatibility and UI rendering decisions.
SchemaFieldRecord.created
SchemaFieldRecord.created_by
SchemaFieldRecord.data_type
Native, framework-specific data type of the field. E.g. “float32”, “uint8[]”, “geometry_msgs/Pose”.
SchemaFieldRecord.field_id
SchemaFieldRecord.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
SchemaFieldRecord.modified
SchemaFieldRecord.modified_by
SchemaFieldRecord.name
Human-readable display name of the field (typically the final component of path_in_schema).
SchemaFieldRecord.org_id
SchemaFieldRecord.path_in_schema
Path components locating this field in the source data schema. Each component is a schema-native attribute name, in order from the schema root to the leaf.
SchemaFieldRecord.schema_id
SchemaFieldRecord.unit
Optional unit of the field’s values (e.g., "ns", "m/s"). None if the field is unitless or unknown.
TimelineExtentRecord
Bases: pydantic.BaseModel
Min/max timestamp bounds for one topic partition measured against one timeline source.
Written by ingest when a partition’s timestamps are summarized for a given source (e.g., a schema timestamp field, or message log/publish time).
Stored timestamps come through verbatim from the data source: they may be absolute nanoseconds since the Unix epoch, or partition-relative (e.g., monotonic from zero). unix_epoch_offset_ns is the calibration that projects stored values onto Unix-epoch wall-clock: session_time_ns = stored_time_ns + unix_epoch_offset_ns. A value of 0 means the stored timestamps are already absolute Unix-epoch ns, or that no calibration has been applied yet.
Parameters
data AnyAttributes
TimelineExtentRecord.created
TimelineExtentRecord.created_by
TimelineExtentRecord.max_timestamp
Largest stored timestamp in this extent, in nanoseconds. Absolute or partition-relative per the source.
TimelineExtentRecord.min_timestamp
Smallest stored timestamp in this extent, in nanoseconds. Absolute or partition-relative per the source.
TimelineExtentRecord.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
TimelineExtentRecord.modified
TimelineExtentRecord.modified_by
TimelineExtentRecord.org_id
TimelineExtentRecord.timeline_extent_id
TimelineExtentRecord.timeline_source_id
ID of the timeline source these bounds are measured against.
TimelineExtentRecord.topic_part_id
ID of the topic partition these bounds apply to.
TimelineExtentRecord.unix_epoch_offset_ns
Nanoseconds to add to each stored timestamp to obtain Unix-epoch wall-clock time: session_time_ns = stored_time_ns + unix_epoch_offset_ns. 0 when stored timestamps are already absolute Unix-epoch ns, or when no calibration has been recorded for this partition/source pair.
TimelineSourceKind
Discriminator for how a TimelineSourceRecord derives its timestamps.
"schema_field" points at a timestamp field inside the schema (field_id is set). "message_log_time" and "message_publish_time" point at the message envelope’s log or publish timestamp respectively (field_id is None).
TimelineSourceRecord
Bases: pydantic.BaseModel
A registered timeline source for a schema.
A timeline source either points at a timestamp field inside the schema (source="schema_field", field_id set) or at the message envelope’s log or publish timestamp (source in {"message_log_time", "message_publish_time"}, field_id is None). Timeline sources are scoped to a schema, not a topic, so topics that share a schema share their timeline sources.
Parameters
data AnyAttributes
TimelineSourceRecord.created
TimelineSourceRecord.created_by
TimelineSourceRecord.field_id
ID of the schema field supplying timestamps. Set when source == "schema_field"; otherwise None.
TimelineSourceRecord.is_default
Whether this timeline source is the default for its schema when no source is specified explicitly.
TimelineSourceRecord.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
TimelineSourceRecord.modified
TimelineSourceRecord.modified_by
TimelineSourceRecord.org_id
TimelineSourceRecord.schema_id
ID of the schema this timeline source is registered against.
TimelineSourceRecord.source
Where timestamps come from: a schema field ("schema_field"), or the message envelope’s log or publish timestamp ("message_log_time" / "message_publish_time").
TimelineSourceRecord.timeline_source_id
TopicIdentityRecord
Bases: pydantic.BaseModel
A durable identity for a topic.
Within an organization, topic names are unique: data logged under the same topic name in different files shares a single identity record.
Parameters
data AnyAttributes
TopicIdentityRecord.created
TopicIdentityRecord.created_by
TopicIdentityRecord.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
TopicIdentityRecord.modified
TopicIdentityRecord.modified_by
TopicIdentityRecord.name
Human-readable topic name (e.g., "/camera/image_raw"). Unique within an organization.
TopicIdentityRecord.org_id
TopicPartitionRecord
Bases: pydantic.BaseModel
One file’s data for a topic.
Pairs a topic identity with a file and carries the facts that vary from file to file: the schema the file’s messages follow (schema_id), the device that produced them, and the data_range locating them inside the file, for formats that pack several slices of data into one shared file. A partition references a file, not a specific version; reads always resolve to the current version.
Parameters
data AnyAttributes
TopicPartitionRecord.created
TopicPartitionRecord.created_by
TopicPartitionRecord.data_range
The slice of the file this partition’s data occupies, as (start, end), or None for the whole file.
start alone identifies the partition within its (topic, file) pair, since a slice’s starting position is stable across re-ingest: re-declaring a slice that begins at the same position updates the existing partition instead of adding a second, overlapping one.
TopicPartitionRecord.device_id
ID of the device that produced this partition’s data, if known.
TopicPartitionRecord.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
TopicPartitionRecord.modified
TopicPartitionRecord.modified_by
TopicPartitionRecord.org_id
TopicPartitionRecord.topic_part_id
TopicRecord
Bases: pydantic.BaseModel
Record representing a topic in the Roboto platform.
A topic is a collection of timestamped data records that share a common name and association (typically a file). Topics represent logical data streams from robotics systems, such as sensor readings, robot state information, or other time-series data.
Data from the same file with the same topic name are considered part of the same topic. Data from different files or with different topic names belong to separate topics, even if they have similar schemas.
When source files are chunked by time or size but represent the same logical data collection, they will produce multiple topic records for the same “logical topic” (same name and schema) across those chunks.
Parameters
data AnyAttributes
TopicRecord.association
Identifier and entity type with which this Topic is associated. E.g., a file, a dataset.
TopicRecord.created
TopicRecord.created_by
TopicRecord.default_representation
Default Representation for this Topic. Assume that if a MessagePath is not more specifically associated with a Representation, it should use this one.
TopicRecord.end_time
Timestamp of oldest message in topic, in nanoseconds since epoch (assumed Unix epoch).
TopicRecord.message_count
TopicRecord.message_paths
Zero to many MessagePathRecords associated with this TopicSource.
TopicRecord.modified
TopicRecord.modified_by
TopicRecord.org_id
TopicRecord.schema_checksum
Checksum of topic schema. May be None if topic does not have a known/named schema.
TopicRecord.schema_id
ID of the schema record for this topic. May be None if the topic has no schema, or if the schema record has not yet been populated.
TopicRecord.schema_name
Type of messages in topic. E.g., “sensor_msgs/PointCloud2”. May be None if topic does not have a known/named schema.
TopicRecord.start_time
Timestamp of earliest message in topic, in nanoseconds since epoch (assumed Unix epoch).
TopicRecord.topic_id
TopicRecord.topic_name
TopicSchemaRecord
Bases: pydantic.BaseModel
A content-addressed topic schema.
Within an organization, two schemas with identical fields share a single record (identified by a deterministic checksum of the fields). name is a mutable, informational label (last-writer-wins) and is not part of the schema’s identity.
Parameters
data AnyAttributes
TopicSchemaRecord.checksum
Deterministic checksum computed over the schema’s fields; identical schemas share a checksum.
TopicSchemaRecord.created
TopicSchemaRecord.created_by
TopicSchemaRecord.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
TopicSchemaRecord.modified
TopicSchemaRecord.modified_by
TopicSchemaRecord.name
Informational label for the schema (e.g., "sensor_msgs/PointCloud2"). Not part of identity.
TopicSchemaRecord.org_id
TopicTimeBounds
Bases: pydantic.BaseModel
Earliest start and latest end, in epoch nanoseconds, across a set of topics.
The aggregate of the start_time and end_time of every topic in the set, computed server-side so a caller does not have to page the whole set to fold them.
Either field is None when no topic in the set carries that timestamp — because the set is empty, or because every topic in it left that bound unset.
Parameters
data AnyTransformationKind
Bases: roboto.compat.StrEnum
Canonical vocabulary of transformations that can be applied when producing a representation.
A transformation is serialized into RepresentationRecord.transformations as a "<kind>:<param>" string (e.g. "downsample:0.5", "encode:jpeg"). This enum is the source of truth for the set of supported kinds; the parameter tail remains free-form because different kinds carry different parameter shapes (floats, format tokens, etc.).
Producers should construct transformation strings via with_param() and consumers should destructure them via parse() to keep the vocabulary centralized.
Usage
TransformationKind.DOWNSAMPLE.with_param(0.5)
# 'downsample:0.5'
TransformationKind.parse("encode:jpeg")
# (<TransformationKind.ENCODE: 'encode'>, 'jpeg')TransformationKind.parse()
Parse a "<kind>:<param>" transformation descriptor into its kind and raw parameter.
Parameters
descriptor strRaises
ValueErrorIf the kind prefix is not a known TransformationKind member.
Return type
TransformationKind.with_param()
Construct a transformation descriptor string for this kind with the given parameter.
Parameters
param objectReturn type
validate_data_range()
Check that data_range is a pair of non-negative positions with start < end.
Both positions must fit in a signed 64-bit integer, which is how the platform stores them.
Parameters
data_range tuple[int, int]The (start, end) pair to check, with start included and end excluded.
Returns
data_range, unchanged.
Raises
ValueErrorstart is negative, end is not greater than start, or a position is too large for a signed 64-bit integer.