roboto.domain.files
Submodules
Package Contents
CreateDirectoryRequest
Bases: pydantic.BaseModel
Request payload to create a directory among the files of one association.
Parameters
data AnyAttributes
CreateDirectoryRequest.create_intermediate_dirs
If True, creates intermediate directories in the path if they don’t exist. If False, requires all parent directories to already exist.
CreateDirectoryRequest.error_if_exists
CreateDirectoryRequest.name
CreateDirectoryRequest.origination
CreateDirectoryRequest.parent_path
CreateLinkRequest
Bases: pydantic.BaseModel
Request body for PUT /v1/files/association/id/<association_id>/link.
Parameters
data AnyAttributes
CreateLinkRequest.relative_path
Where the link sits among the association’s files. Missing parent directories are created.
CreateLinkRequest.target_file_id
ID of the file the link points at. It must be a file, not a link or a directory, in the same org.
CreateLinkRequest.target_version
Version of the target to pin. Defaults to the target’s current version.
DeleteFileRequest
Bases: pydantic.BaseModel
Request payload for deleting a file from the platform.
This request is used internally by the platform to delete files and their associated data. The file is identified by its storage URI.
Parameters
data AnyAttributes
DeleteFileRequest.uri
Storage URI of the file to delete (e.g., ‘s3://bucket/path/to/file.bag’).
DirectoryContentsPage
Bases: pydantic.BaseModel
Response containing the contents of a dataset directory page.
Represents a paginated view of files and subdirectories within a dataset directory. Used when browsing dataset contents hierarchically.
Parameters
data AnyAttributes
DirectoryContentsPage.directories
Subdirectories contained in this directory page.
DirectoryContentsPage.files
Files contained in this directory page.
DirectoryContentsPage.next_token
Token for retrieving the next page of results, if any.
DirectoryRecord
Bases: pydantic.BaseModel
Wire-transmissible representation of a directory within a dataset.
DirectoryRecord represents a logical directory structure within a dataset, containing metadata about the directory’s location and contents. Directories are used to organize files hierarchically within datasets.
Directory records are typically returned when browsing dataset contents or when performing directory-based operations like bulk deletion.
Parameters
data AnyAttributes
DirectoryRecord.association_id
DirectoryRecord.created
DirectoryRecord.created_by
DirectoryRecord.description
DirectoryRecord.directory_id
DirectoryRecord.metadata
DirectoryRecord.modified
DirectoryRecord.modified_by
DirectoryRecord.org_id
DirectoryRecord.origination
DirectoryRecord.parent_id
DirectoryRecord.relative_path
DirectoryRecord.status
DirectoryRecord.storage_type
DirectoryRecord.tags
DirectoryRecord.upload_id
FSType
Bases: roboto.compat.StrEnum
File system type enum
Attributes
FSType.Directory
FSType.File
FSType.Link
A pointer to one version of another file, which may live under a different dataset, device, or org.
A link stores no object of its own. Its record’s uri is roboto://file/<target_file_id>?v=<version> and its size is 0; downloading it fetches the target at that version.
File
Represents a file within the Roboto platform.
Files are the fundamental data storage unit in Roboto. They can be uploaded to datasets, imported from external sources, or created as outputs from actions. Once in the platform, files can be tagged with metadata, post-processed by actions, added to collections, visualized in the web interface, and searched using the query system.
Files contain structured data that can be ingested into topics for analysis and visualization. Common file formats include ROS bags, MCAP files, ULOG files, CSV files, and many others. Each file has an associated ingestion status that tracks whether its data has been processed and made available for querying.
Files are versioned entities - each modification creates a new version while preserving the history. Files are associated with datasets and inherit access permissions from their parent dataset.
The File class provides methods for downloading, updating metadata, managing tags, accessing topics, and performing other file operations. It serves as the primary interface for file manipulation in the Roboto SDK.
Parameters
roboto_client Optional[roboto.file_service Optional[roboto.File.add_topic()
Create a Topic from a pandas DataFrame and associate it with this file.
If a topic with the same name already exists for this file, it will be updated with the new data and schema.
Parameters
topic_name strName for the topic. Must be unique within this file.
df pandas.pandas DataFrame containing the data to ingest. Must include a timestamp column (either explicitly specified or automatically detectable).
timestamp_column Optional[str]Name of the column to use as the timestamp. If not provided, the method will attempt to automatically detect a timestamp column by looking for the first column that is a timezone-aware timestamp type.
timestamp_unit Optional[Union[str, roboto.Unit of the timestamp column values. Required when timestamp_column contains numeric values (int, float, decimal). Valid values include “s”, “ms”, “us”, “ns”. Not needed for datetime columns or when timestamp_column is not specified.
Returns
The created or updated Topic instance.
Raises
If the timestamp column cannot be determined, is not present in the DataFrame, has an invalid type, or if the timestamp unit is required but not provided.
ImportErrorIf pandas or pyarrow are not installed. Install with pip install roboto[ingestion] to use this feature.
This file is associated with a device or with the org itself, not with a dataset; only a dataset’s files hold topics.
If the caller lacks permission to create topics or upload files to this file’s dataset.
Notes
- Requires installing this package using the
roboto[ingestion]extra - Topic names are unique within a file
- Schema and statistics are automatically inferred from the DataFrame
Usage
Create a topic with explicit timestamp column and unit:
import pandas as pd
from roboto import File
file = File.from_id("file_abc123")
df = pd.DataFrame(
{
"timestamp": [1763947309.4198897, 1763947316.7686195, 1763947335.0095527],
"temperature": [20.5, 21.0, 20.8],
"humidity": [45.2, 46.1, 45.8],
}
)
topic = file.add_topic(
topic_name="sensor_data", df=df, timestamp_column="timestamp", timestamp_unit="s"
)
print(f"Created topic: {topic.name}")
# Created topic: sensor_dataCreate a topic with automatic timestamp detection:
import pandas as pd
from roboto import File
file = File.from_id("file_abc123")
df = pd.DataFrame(
{
"ts": pd.date_range("2025-11-24", periods=3, freq="1s", tz="UTC"),
"velocity": [10.5, 11.2, 10.8],
"acceleration": [0.5, 0.3, -0.2],
}
)
topic = file.add_topic("motion_data", df)Retrieve the data back
retrieved_df = topic.get_data_as_df()
print(f"Retrieved {len(retrieved_df)} rows")
# Retrieved 3 rowsAdd derived data as a new topic to the same file, using the original topic’s timestamp index:
import pandas as pd
from roboto import File
file = File.from_id("file_abc123")
# Get existing topic data as DataFrame
original_topic = file.get_topic("sensor_data")
original_df = original_topic.get_data_as_df()
# Create derived data
derived_df = pd.DataFrame(
{
"temp_category": original_df["temperature"].apply(lambda x: "hot" if x > 25 else "not_hot"),
},
index=original_df.index,
)
derived_topic = file.add_topic(
"temperature_categories",
derived_df,
)Properties
File.association
The dataset, device, or org this file is associated with.
Every file has exactly one association, inferred from the prefix of its association ID. Read file.association.association_type to branch on it.
File.created
Timestamp when this file was created.
Returns the UTC datetime when this file was first uploaded or created in the Roboto platform. This timestamp is immutable.
File.created_by
Identifier of the user who created this file.
Returns the user ID or identifier of the person or service that originally uploaded or created this file in the Roboto platform.
File.dataset_id
Identifier of the dataset that contains this file.
Valid only for a file associated with a dataset; files associated with a device or with the org itself have no dataset. Prefer association, which works for every file.
Raises
This file is not associated with a dataset.
Return type
File.declare_topic()
Register one topic this File contributes data to, without naming a Session.
The singular form of declare_topics(), taking the fields of one FileTopicDeclaration as separate arguments. That class documents what each field means; declare_topics() documents what the platform does with it.
Parameters
topic_name strTopic this File contributes data to. Topic names are unique within an org.
topic_schema roboto.Structure of the topic’s data.
timeline_sources collections.Timeline sources this File’s topic data carries, each with the bounds it spans in this File, stated in the File’s own timestamps.
data_range Optional[tuple[int, int]]The part of the File this topic’s data occupies, or None for the whole File.
anchor Optional[roboto.Optional wall-clock instant the data this declaration names was captured at: an int of nanoseconds since the Unix epoch, or any other Time, read as to_epoch_nanoseconds() reads it (a datetime or ISO 8601 string is that instant; a float, Decimal, or numeric string is seconds since the epoch). Must fall after the Unix epoch.
representations collections.The files a read of this topic’s data opens, each with how it holds that data: this File, when its own bytes are readable, and other files when the data is read from them, such as files converted out of it. Empty lists none, and reads of a topic with no representations return no rows.
Returns
The topic this declaration registered against.
Raises
TypeErrorIf anchor is not one of the Time types.
ValueErrorIf anchor is a boolean, a negative number (an int, float, Decimal, or numeric string), or a string that is neither a number of seconds nor an ISO 8601 timestamp. Raised before anything is sent to the platform.
OverflowErrorIf anchor is an infinite float, Decimal, or string, such as "inf". Raised before anything is sent to the platform.
pydantic.ValidationErrorIf these arguments do not form a valid FileTopicDeclaration, for instance two representations of the whole topic, or of one field, sharing a storage format, content format and transformations, or any timeline source but SchemaFieldSource beside a PARQUET representation, or if representations names one file in two storage formats. Enforced before anything is sent to the platform.
Whatever the platform refused this declaration with.
Usage
from roboto.domain.files import File
from roboto.domain.topics import CanonicalDataType, RepresentationStorageFormat
from roboto.experimental.ingest import Field, RepresentationDeclaration, Schema, SchemaFieldSource
timestamp = Field(
name="timestamp",
data_type="float64",
canonical_data_type=CanonicalDataType.Timestamp,
unit="s",
)
file = File.from_id("fl_0123456789ab")
topic = file.declare_topic(
topic_name="observation.state",
topic_schema=Schema(
name="observation.state",
fields=[timestamp, Field(name="observation.state", data_type="float32")],
),
timeline_sources=[
SchemaFieldSource(
field_path=["timestamp"],
min_file_timestamp_ns=0,
max_file_timestamp_ns=4_000,
)
],
representations=[
RepresentationDeclaration(
file_id=file.file_id,
storage_format=RepresentationStorageFormat.PARQUET,
)
],
)
print(topic.topic_id)File.declare_topics()
Register the topic data this File carries, without naming a Session.
One call states everything the platform needs to serve this File’s topic data: for each topic, the structure of its rows, the timeline sources those rows carry with the bounds they span in this File, and, when the File packs its data into slices, which slice the topic occupies. The platform does not open the File when topics are declared on it, so this declaration is all it knows about the File’s contents.
The platform applies each declaration on its own: one it refuses leaves the others registered, and the response says what became of each. Nothing about the call involves a Session, so declarations on the Files of one recording can run concurrently, and a File’s topic data can be registered before the Session holding it exists. A Session takes that data on by attaching the File, through add_file() or a SessionFile carrying no topics. A Session already holding the File takes on what this call declares before the call returns, with its time bounds recomputed to cover the newly declared data.
Resending the same call is safe: the platform identifies the data a declaration registers by the topic plus the slice of the File that declaration names, so a resend converges on what the first attempt registered rather than duplicating it, and a corrected redeclaration replaces what it corrects.
A declaration states its bounds in the File’s own timestamps, read as nanoseconds since the Unix epoch; the platform never invents a wall-clock time. Data whose timestamps start at 0 therefore sits at the epoch until it is anchored. To place it at the wall-clock time it was captured, supply anchor_ns, which anchors the whole slice it names rather than the one topic declaring it, so the topics sharing a slice must agree on it.
The topic data belongs to this File: its bounds, anchors and slices are stated against it, and a Session holding this File holds the data. What makes the data readable is each topic’s representations: the files a read opens to get it, each decoded in the storage format its representation states. A file’s name and extension are not used. A topic lists this File when its own bytes are readable, and other files when the data is read from them, such as the per-topic MCAPs converted out of a PX4 ULog. A topic with no representations is still registered and still counts toward the bounds of the Sessions holding the File, but reads of it return no rows. RepresentationDeclaration states what a representation’s file must hold and when a read can decode an MCAP representation’s file.
Redeclaring a topic adds the representations listed to the ones it has, each taking the place of the stored ones it matches, as representations describes. To remove a representation, or to replace a topic’s representations outright, use set_representations().
Parameters
topics collections.One declaration per topic and slice of this File. An empty sequence returns an empty response without contacting the platform.
Returns
One element per declaration, in request order, holding either the topic it registered against or why the platform refused it.
Raises
pydantic.ValidationErrorIf more than MAX_FILES_AND_TOPICS_PER_REQUEST topics are given, one topic is declared twice over the same slice, two topics anchor one slice at different instants, or representations name one file in two storage formats. These are enforced when the request body is constructed, before anything is sent to the platform; each declaration’s own rules are enforced earlier, when the caller builds it.
If this File no longer exists, or a representation names a file that does not exist in this File’s org or whose status is not Available. Nothing is registered.
If the caller lacks edit access to this File or to a file a listed representation names, or lacks topic edit access in this File’s org while a declaration states is_default_for_reads on a timeline source.
Usage
Register the two topics a LeRobot episode file carries, each read from the file’s own bytes:
from roboto.domain.files import File
from roboto.domain.topics import CanonicalDataType, RepresentationStorageFormat
from roboto.experimental.ingest import (
Field,
FileTopicDeclaration,
RepresentationDeclaration,
Schema,
SchemaFieldSource,
)
timestamp = Field(
name="timestamp",
data_type="float64",
canonical_data_type=CanonicalDataType.Timestamp,
unit="s",
)
file = File.from_id("fl_0123456789ab")
from_file = RepresentationDeclaration(
file_id=file.file_id,
storage_format=RepresentationStorageFormat.PARQUET,
)
registered = file.declare_topics(
[
FileTopicDeclaration(
topic_name="observation.state",
topic_schema=Schema(
name="observation.state",
fields=[timestamp, Field(name="observation.state", data_type="float32")],
),
timeline_sources=[
SchemaFieldSource(
field_path=["timestamp"],
min_file_timestamp_ns=0,
max_file_timestamp_ns=4_000,
)
],
representations=[from_file],
),
FileTopicDeclaration(
topic_name="action",
topic_schema=Schema(
name="action",
fields=[timestamp, Field(name="action", data_type="float32")],
),
timeline_sources=[
SchemaFieldSource(
field_path=["timestamp"],
min_file_timestamp_ns=0,
max_file_timestamp_ns=4_000,
)
],
representations=[from_file],
),
],
)
print([topic.topic_id for topic in registered.succeeded])Register a topic over the slice of a shared file that holds one episode, anchored at the instant that episode was recorded:
registered = file.declare_topics(
[
FileTopicDeclaration(
topic_name="observation.state",
topic_schema=Schema(
name="observation.state",
fields=[timestamp, Field(name="observation.state", data_type="float32")],
),
timeline_sources=[
SchemaFieldSource(
field_path=["timestamp"],
min_file_timestamp_ns=0,
max_file_timestamp_ns=4_000,
)
],
data_range=(0, 80),
anchor_ns=1_785_974_400_000_000_000,
representations=[from_file],
),
],
)File.delete()
Delete this file from the Roboto platform.
Permanently removes the file and all its associated data, including topics and metadata. This operation cannot be undone.
For files that were imported from customer S3 buckets (read-only BYOB integrations), this method does not delete the file content from S3. It only removes the metadata and references within the Roboto platform.
Raises
File does not exist or has already been deleted.
Caller lacks permission to delete the file.
Return type
Usage
file = File.from_id("file_abc123")
file.delete()
# # File is now permanently deletedProperties
File.description
Human-readable description of this file.
Returns the optional description text that provides details about the file’s contents, purpose, or context. Can be None if no description was provided.
File.device_id
Identifier of the device that generated this data.
Returns the optional identifier of the device that generated the data contained within this file. Can be None if the file was not generated by a device.
File.download()
Download this file to a local path.
Downloads the file content from cloud storage to the specified local path. The parent directories are created automatically if they don’t exist.
For a link, downloads the version of the target file that the link pins.
Parameters
local_path pathlib.Local filesystem path where the file should be saved.
print_progress boolWhether to show a progress bar during download.
Raises
This file is a link whose target, at the pinned version, no longer exists.
Caller lacks permission to download the file, or a link’s target.
FileNotFoundErrorFile content is not available in storage.
Usage
import pathlib
file = File.from_id("file_abc123")
local_path = pathlib.Path("/tmp/downloaded_file.bag")
file.download(local_path)
print(f"Downloaded to {local_path}")Properties
File.file_id
Unique identifier for this file.
Returns the globally unique identifier assigned to this file when it was created. This ID is immutable and used to reference the file across the Roboto platform.
File.from_id()
Create a File instance from a file ID.
Retrieves file information from the Roboto platform using the provided file ID and optionally a specific version.
Parameters
file_id strUnique identifier for the file.
version_id Optional[int]Specific version of the file to retrieve. If None, gets the latest version.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
Returns
File instance representing the requested file.
Raises
File with the given ID does not exist.
Caller lacks permission to access the file.
Usage
file = File.from_id("file_abc123")
print(file.relative_path)
# 'data/sensor_logs.bag'old_version = File.from_id("file_abc123", version_id=1)
print(old_version.version)
# 1File.from_path_and_dataset_id()
Create a File instance from a file path and dataset ID.
Retrieves file information using the file’s relative path within a specific dataset. This is useful when you know the file’s location within a dataset but not its file ID.
Parameters
file_path Union[str, pathlib.Relative path of the file within the dataset.
dataset_id strID of the dataset containing the file.
version_id Optional[int]Specific version of the file to retrieve. If None, gets the latest version.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
Returns
File instance representing the requested file.
Raises
File at the given path does not exist in the dataset.
Caller lacks permission to access the file or dataset.
Usage
file = File.from_path_and_dataset_id("logs/session1.bag", "ds_abc123")
print(file.file_id)
# 'file_xyz789'file = File.from_path_and_dataset_id(pathlib.Path("data/sensors.csv"), "ds_abc123")
print(file.relative_path)
# 'data/sensors.csv'File.get_signed_url()
Generate a signed URL for direct access to this file.
Creates a time-limited URL that allows direct access to the file content without requiring Roboto authentication. Useful for sharing files or integrating with external systems.
Parameters
override_content_type Optional[str]Custom MIME type to set in the response headers.
override_content_disposition Optional[str]Custom content disposition header value (e.g., “attachment; filename=myfile.bag”).
Return type
For a link, the URL is for the version of the target file that the link pins.
Parameters
override_content_type Optional[str]override_content_disposition Optional[str]Returns
Signed URL string that provides temporary access to the file.
Raises
This file is a link whose target, at the pinned version, no longer exists.
Caller lacks permission to access the file, or a link’s target.
Usage
file = File.from_id("file_abc123")
url = file.get_signed_url()
print(f"Direct access URL: {url}")# Force download with custom filename
download_url = file.get_signed_url(override_content_disposition="attachment; filename=data.bag")File.get_topic()
Get a specific topic from this file by name.
Retrieves a topic with the specified name that is associated with this file. Topics contain the structured data extracted from the file during ingestion.
Parameters
topic_name strName of the topic to retrieve (e.g., “/camera/image”, “/imu/data”).
Returns
Topic instance for the specified topic name.
Raises
Topic with the given name does not exist in this file.
Caller lacks permission to access the topic.
Usage
file = File.from_id("file_abc123")
camera_topic = file.get_topic("/camera/image")
print(f"Topic schema: {camera_topic.schema}")# Access topic data
for record in camera_topic.get_data():
print(f"Timestamp: {record['timestamp']}")File.get_topics()
Get all topics associated with this file, with optional filtering.
Retrieves all topics that were extracted from this file during ingestion. Topics can be filtered by name using include/exclude patterns.
Parameters
include Optional[collections.If provided, only topics with names in this sequence are yielded.
exclude Optional[collections.If provided, topics with names in this sequence are skipped.
Yields
Topic instances associated with this file, filtered according to the parameters.
Return type
Usage
file = File.from_id("file_abc123")
for topic in file.get_topics():
print(f"Topic: {topic.name}")
# Topic: /camera/image
# Topic: /imu/data
# Topic: /gps/fix# Only get camera topics
camera_topics = list(file.get_topics(include=["/camera/image", "/camera/info"]))
print(f"Found {len(camera_topics)} camera topics")# Exclude diagnostic topics
data_topics = list(file.get_topics(exclude=["/diagnostics"]))File.import_batch()
Import files from customer S3 bring-your-own buckets into Roboto datasets.
This is the ingress point for importing data stored in customer-owned S3 buckets that have been registered as read-only bring-your-own bucket (BYOB) integrations with Roboto. Files remain in their original S3 locations while metadata is registered with Roboto for discovery, processing, and analysis.
This method only works with S3 URIs from buckets that have been properly registered as BYOB integrations for your organization. It performs batch operations to efficiently import multiple files in a single API call, reducing overhead and improving performance.
Parameters
requests collections.Sequence of import requests, each specifying file details and metadata.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
caller_org_id Optional[str]Organization ID of the caller. Required for multi-org users.
Returns
Sequence of File objects representing the imported files.
Raises
If any URI is not a valid S3 URI, if the batch exceeds 500 items, or if bucket integrations are not properly configured.
If the caller lacks upload permissions for target datasets or if buckets don’t belong to the caller’s organization.
Notes
- Only works with S3 URIs from registered read-only BYOB integrations
- Files are not copied; only metadata is imported into Roboto
- Batch size is limited to 500 items per request
- All S3 buckets must be registered to the caller’s organization
Usage
from roboto.domain.files import ImportFileRequest
requests = [
ImportFileRequest(
dataset_id="ds_abc123",
relative_path="logs/session1.bag",
uri="s3://my-bucket/data/session1.bag",
size=1024000,
),
ImportFileRequest(
dataset_id="ds_abc123",
relative_path="logs/session2.bag",
uri="s3://my-bucket/data/session2.bag",
size=2048000,
),
]
files = File.import_batch(requests)
print(f"Imported {len(files)} files")
# Imported 2 filesFile.import_one()
Import a single file from an external bucket into a Roboto dataset. This currently only supports AWS S3.
This is a convenience method for importing a single file from customer-owned buckets that have been registered as bring-your-own bucket (BYOB) integrations with Roboto. Unlike import_batch(), this method automatically determines the file size by querying the object store and verifies that the object actually exists before importing, providing additional validation and convenience for single-file operations.
The file remains in its original location while metadata is registered with Roboto for discovery, processing, and analysis. This method currently only works with S3 URIs from buckets that have been properly registered as BYOB integrations for your organization.
Parameters
dataset_id strID of the dataset to import the file into.
relative_path strPath of the file relative to the dataset root (e.g., logs/session1.bag).
uri strURI where the file is located (e.g., s3://my-bucket/path/to/file.bag). Must be from a registered BYOB integration.
description Optional[str]Optional human-readable description of the file.
tags Optional[list[str]]Optional list of tags for file discovery and organization.
metadata Optional[dict[str, Any]]Optional key-value metadata pairs to associate with the file.
device_id Optional[str]Optional identifier of the device that generated this data.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
Returns
File object representing the imported file.
Raises
If the URI is not a valid URI or if the bucket integration is not properly configured.
If the specified object does not exist.
If the caller lacks upload permissions for the target dataset or if the bucket doesn’t belong to the caller’s organization.
Notes
- Only works with S3 URIs from registered BYOB integrations
- File size is automatically determined from the object metadata
- The file is not copied; only metadata is imported into Roboto
- For importing multiple files efficiently, use
import_batch()instead
Usage
Import a single ROS bag file:
from roboto.domain.files import File
file = File.import_one(
dataset_id="ds_abc123", relative_path="logs/session1.bag", uri="s3://my-bucket/data/session1.bag"
)
print(f"Imported file: {file.relative_path}")
# Imported file: logs/session1.bagImport a file with metadata and tags:
file = File.import_one(
dataset_id="ds_abc123",
relative_path="sensors/lidar_data.pcd",
uri="s3://my-bucket/sensors/lidar_data.pcd",
description="LiDAR point cloud from highway test",
tags=["lidar", "highway", "test"],
metadata={"sensor_type": "Velodyne", "resolution": "high"},
)
print(f"File size: {file.size} bytes")Properties
File.ingestion_status
Current ingestion status of this file.
Returns the status indicating whether this file has been processed and its data extracted into topics. Used to track ingestion pipeline progress.
File.is_link
Whether this file is a link to one version of another file.
A link sits at its own path under its own dataset, device, or org, and stores no object. download() and get_signed_url() fetch the target at the version the link pins.
File.mark_ingested()
Mark this file as fully ingested and ready for post-processing.
Updates the file’s ingestion status to indicate that all data has been successfully processed and extracted into topics. This enables triggers and other automated workflows that depend on complete ingestion.
Returns
Updated File instance with ingestion status set to Ingested.
Raises
Caller lacks permission to update the file.
Notes
This method is typically called by ingestion actions after they have successfully processed all data in the file. Once marked as ingested, the file becomes eligible for additional post-processing actions.
Usage
file = File.from_id("file_abc123")
print(file.ingestion_status)
# IngestionStatus.NotIngested
updated_file = file.mark_ingested()
print(updated_file.ingestion_status)
# IngestionStatus.IngestedProperties
File.metadata
Custom metadata associated with this file.
Returns the file’s metadata dictionary containing arbitrary key-value pairs for storing custom information. Supports nested structures and dot notation for accessing nested fields.
File.modified
Timestamp when this file was last modified.
Returns the UTC datetime when this file’s metadata, tags, or other properties were most recently updated. The file content itself is immutable, but metadata can be modified.
File.modified_by
Identifier of the user who last modified this file.
Returns the user ID or identifier of the person who most recently updated this file’s metadata, tags, or other mutable properties.
File.org_id
Organization identifier that owns this file.
Returns the unique identifier of the organization that owns and has primary access control over this file.
File.put_metadata()
Add or update metadata fields for this file.
Adds new metadata fields or updates existing ones. Existing fields not specified in the metadata dict are preserved.
Parameters
metadata dict[str, Any]Dictionary of metadata key-value pairs to add or update.
Returns
Updated File instance with the new metadata.
Raises
Caller lacks permission to update the file.
Usage
file = File.from_id("file_abc123")
updated_file = file.put_metadata(
{"vehicle_id": "vehicle_001", "session_type": "highway_driving", "weather": "sunny"}
)
print(updated_file.metadata["vehicle_id"])
# 'vehicle_001'File.put_tags()
Add or update tags for this file.
Replaces the file’s current tags with the provided list. To add tags while preserving existing ones, retrieve current tags first and combine them.
Parameters
tags list[str]List of tag strings to set on the file.
Returns
Updated File instance with the new tags.
Raises
Caller lacks permission to update the file.
Usage
file = File.from_id("file_abc123")
updated_file = file.put_tags(["sensor-data", "highway", "sunny"])
print(updated_file.tags)
# ['sensor-data', 'highway', 'sunny']File.query()
Query files using a specification with filters and pagination.
Searches for files matching the provided query specification. Results are returned as a generator that automatically handles pagination, yielding File instances as they are retrieved from the API.
Parameters
spec Optional[roboto.Query specification with filters, sorting, and pagination options. If None, returns all accessible files.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
owner_org_id Optional[str]Organization ID to scope the query. If None, uses caller’s org.
Yields
File instances matching the query specification.
Raises
ValueErrorQuery specification references unknown file attributes.
Caller lacks permission to query files.
Return type
Usage
from roboto.query import Comparator, Condition, QuerySpecification
spec = QuerySpecification(
condition=Condition(field="tags", comparator=Comparator.Contains, value="sensor-data")
)
for file in File.query(spec):
print(f"Found file: {file.relative_path}")
# Found file: logs/sensors_2024_01_01.bag
# Found file: logs/sensors_2024_01_02.bag# Query with metadata filter
spec = QuerySpecification(
condition=Condition(field="metadata.vehicle_id", comparator=Comparator.Equals, value="vehicle_001")
)
files = list(File.query(spec))
print(f"Found {len(files)} files for vehicle_001")Properties
File.record
Underlying data record for this file.
Returns the raw FileRecord that contains all the file’s data fields. This provides access to the complete file state as stored in the platform.
File.refresh()
Refresh this file instance with the latest data from the platform.
Fetches the current state of the file from the Roboto platform and updates this instance’s data. Useful when the file may have been modified by other processes or users.
Returns
This File instance with refreshed data.
Raises
File no longer exists.
Caller lacks permission to access the file.
Usage
file = File.from_id("file_abc123")
# File may have been updated by another process
refreshed_file = file.refresh()
print(f"Current version: {refreshed_file.version}")Properties
File.relative_path
Path of this file relative to the root of its association’s files.
Uses forward slashes as separators regardless of the operating system. This path uniquely identifies the file among the files of its dataset, device, or org.
File.rename_file()
Rename this file to a new path within its dataset, device, or org.
Changes the relative path of the file among the files of its association. This updates the file’s location identifier but does not move the actual file content.
Parameters
file_id strFile ID (currently unused, kept for API compatibility).
new_path strNew relative path for the file, relative to the root of its association’s files.
Returns
Updated FileRecord with the new path.
Raises
Caller lacks permission to rename the file.
New path is invalid or conflicts with existing file.
Usage
file = File.from_id("file_abc123")
print(file.relative_path)
# 'old_logs/session1.bag'
updated_record = file.rename_file("file_abc123", "logs/session1.bag")
print(updated_record.relative_path)
# 'logs/session1.bag'File.set_device_id()
Set the device ID for this file.
Parameters
device_id strThe device ID to set for this file.
Returns
Updated File instance with the new device ID.
Raises
Caller lacks permission to update the file.
The specified device ID does not exist.
Usage
file = File.from_id("file_abc123")
updated_file = file.set_device_id("device_xyz789")File.set_representations()
Replace the representations the named topics’ data on this File is read from.
This File is the one the topics were declared on, through declare_topics() or a Session. Each topic listed ends up with exactly the representations listed, over the part of this File its entry’s data_range names; its other representations there are removed, whether they cover the whole topic or one field of it. Topics and slices not listed keep theirs. Nothing else about the topics changes: their schemas, timeline sources, bounds, anchors and slices stay as declared, and so do the time bounds of every Session holding this File, which do not depend on which files the data is read from.
Use it for what redeclaring a topic cannot do:
- Remove a representation.
- Replace a representation with one that names another file and differs from it in what it covers, its storage format, its content format or its transformations. Those four identify a representation, as
RepresentationDeclarationdescribes, so declaring the new one adds it beside the first. - Stop a topic being read at all, by listing no representations for it.
The platform checks every entry before writing any, and one refused entry refuses the whole call.
Parameters
topics collections.One entry per topic and slice, at most MAX_FILES_AND_TOPICS_PER_REQUEST; split a larger set across several calls. An empty sequence returns without contacting the platform.
Raises
pydantic.ValidationErrortopics is longer than the cap, lists one topic and slice twice, or lists representations naming one file in two storage formats. Raised before any request is made. The rules for one topic’s own representations are enforced earlier, when the caller builds its TopicRepresentations.
This File does not exist, a topic is not declared on it, an entry’s data_range is not one the topic is declared over on it, or a representation names a file that does not exist in this File’s org or whose status is not Available. Nothing is written.
A topic declared with a timeline source other than SchemaFieldSource (MCAP log or publish time, MP4 presentation time) would get a PARQUET representation, a topic declared over a data_range would be left with a representation that cannot be read by row position and none that can covering the same fields, as transformations describes, or a representation’s field_path names no field of the schema the topic is declared under on this File. Nothing is written.
The caller cannot edit this File or a file a listed representation names.
Return type
Usage
Replace a camera topic’s representation re-encoded as JPEG with one downsampled and re-encoded as PNG, keeping the untransformed one that names the recording, and stop reading a debug topic at all:
from roboto.domain.files import File
from roboto.domain.topics import RepresentationStorageFormat
from roboto.experimental.ingest import RepresentationDeclaration, TopicRepresentations
recording = File.from_id("fl_recording_0412_mcap")
recording.set_representations(
[
TopicRepresentations(
topic_name="/camera/front/image_raw",
representations=[
RepresentationDeclaration(
file_id=recording.file_id,
storage_format=RepresentationStorageFormat.MCAP,
),
RepresentationDeclaration(
file_id="fl_front_png",
storage_format=RepresentationStorageFormat.MCAP,
content_format="png",
transformations=["downsample:0.5", "encode:png"],
),
],
),
TopicRepresentations(topic_name="/debug/raw_dump", representations=[]),
]
)File.set_timeline_offset()
Calibrate this file’s timeline to Unix-epoch wall-clock, optionally scoped to a topic and/or source.
Contract:
- The offset, in nanoseconds, is added to stored partition time to produce session wall-clock:
session_time_ns = stored_time_ns + offset_ns. An offset given as an instant, such as adatetime, is the nanoseconds since the Unix epoch at which stored time 0 occurred. topic/topic_namescopes the update to a single topic in this file;timeline_source/timeline_source_namescopes it to a single source. With no selectors, the offset applies to every timeline on the file.
Use set_timeline_offsets() to send several offsets in one atomic request.
Parameters
offset roboto.Offset to apply: an int of nanoseconds, or any other Time, read as to_epoch_nanoseconds() reads it (a datetime or ISO 8601 string is that instant; a float, Decimal, or numeric string is seconds). Must not be negative, and must fit in a signed 64-bit integer of nanoseconds.
topic Optional[roboto.Topic to scope the update to. Mutually exclusive with topic_name.
topic_name Optional[str]Topic name to scope the update to (e.g. "/imu/raw"). Mutually exclusive with topic.
timeline_source Optional[roboto.Source record to scope the update to. Mutually exclusive with timeline_source_name.
timeline_source_name Optional[str]Source name to scope the update to (e.g. "header.stamp"). Mutually exclusive with timeline_source.
Returns
The updated TimelineExtentRecord objects returned by the server.
Raises
TypeErroroffset is not one of the Time types.
ValueErroroffset is a boolean, a negative number (an int, float, Decimal, or numeric string), or a string that is neither seconds nor ISO 8601; or both of a mutually exclusive pair of selectors are given. Raised before any request is made.
OverflowErroroffset is an infinite float, Decimal, or string, such as "inf". Raised before any request is made.
pydantic.ValidationErroroffset converts to a negative number of nanoseconds, or to more than a signed 64-bit integer holds. A subclass of ValueError, raised before any request is made.
The caller cannot edit this file.
The file carries no timeline data, or the selectors match none of it.
The offset would place the data it reaches, or a session time range declared over that data, before the Unix epoch or past the largest storable Unix-epoch nanosecond value. Nothing is written.
Usage
Apply a file-wide offset:
file = File.from_id("file_abc123")
file.set_timeline_offset(1_700_000_000_000_000_000)Apply the same offset as a datetime, the instant stored time 0 occurred:
import datetime
file.set_timeline_offset(datetime.datetime(2023, 11, 14, 22, 13, 20, tzinfo=datetime.timezone.utc))Apply an offset to a single topic by name:
file.set_timeline_offset(1_700_000_000_000_000_000, topic_name="/imu/raw")Apply an offset to a specific source on a topic:
file.set_timeline_offset(
500_000_000,
topic_name="data",
timeline_source_name="ts",
)File.set_timeline_offsets()
Apply multiple timeline offsets to this file in one atomic request.
Each entry carries a unix_epoch_offset_ns and optional selectors (topic_name, timeline_source_id, timeline_source_name) that narrow where the offset is applied. An entry with no selectors targets every timeline on the file.
Use set_timeline_offset() for the single-offset convenience form.
Parameters
offsets list[roboto.Offset entries to apply, each with its own selectors.
Returns
The updated TimelineExtentRecord objects returned by the server.
Raises
pydantic.ValidationErroroffsets is empty. Raised before any request is made.
The caller cannot edit this file.
The file carries no timeline data, or the selectors of every entry together match none of it.
An entry’s offset would place the data it reaches, or a session time range declared over that data, before the Unix epoch or past the largest storable Unix-epoch nanosecond value. The whole request is refused and nothing is written.
Usage
Apply per-topic offsets in a single request:
from roboto.domain.topics import TimelineOffsetEntry
file = File.from_id("file_abc123")
file.set_timeline_offsets(
[
TimelineOffsetEntry(unix_epoch_offset_ns=1_700_000_000_000_000_000, topic_name="/imu/raw"),
TimelineOffsetEntry(unix_epoch_offset_ns=1_700_000_000_000_000_000, topic_name="/camera/image"),
]
)Properties
File.tags
List of tags associated with this file.
Returns the list of string tags that have been applied to this file for categorization and filtering purposes.
File.to_association()
Convert this file to an Association reference.
Creates an Association object that can be used to reference this file in other contexts, such as when creating collections or specifying action inputs.
Returns
Association object referencing this file and its current version.
Usage
file = File.from_id("file_abc123")
association = file.to_association()
print(f"Association: {association.association_type}:{association.association_id}")
# Association: file:file_abc123File.to_dict()
Convert this file to a dictionary representation.
Returns the file’s data as a JSON-serializable dictionary containing all file attributes and metadata.
Returns
Dictionary representation of the file data.
Usage
file = File.from_id("file_abc123")
file_dict = file.to_dict()
print(file_dict["relative_path"])
# 'logs/session1.bag'
print(file_dict["metadata"])
# {'vehicle_id': 'vehicle_001', 'session_type': 'highway'}File.update()
Update this file’s properties.
Updates various properties of the file including description, metadata, and ingestion status. Only specified parameters are updated; others remain unchanged.
Parameters
description Optional[Union[str, roboto.New description for the file. Use NotSet to leave unchanged.
metadata_changeset Union[roboto.Metadata changes to apply (add, update, or remove fields/tags). Use NotSet to leave metadata unchanged.
ingestion_complete Union[Literal[True], roboto.Set to True to mark the file as fully ingested. Use NotSet to leave ingestion status unchanged.
device_id Optional[Union[str, roboto.New device ID for the file. Use NotSet to leave unchanged.
Returns
Updated File instance with the new properties.
Raises
Caller lacks permission to update the file.
Usage
file = File.from_id("file_abc123")
updated_file = file.update(description="Updated sensor data from highway test")
print(updated_file.description)
# 'Updated sensor data from highway test'# Update metadata and mark as ingested
from roboto.updates import MetadataChangeset
changeset = MetadataChangeset(put_fields={"processed": True})
updated_file = file.update(metadata_changeset=changeset, ingestion_complete=True)Properties
File.uri
Storage URI for this file’s content.
Returns the storage location URI where the file’s actual content is stored. This is typically an S3 URI or similar cloud storage reference.
File.version
Version number of this file.
Returns the version number that increments each time the file’s metadata or properties are updated. The file content itself is immutable, but metadata changes create new versions.
FileRecord
Bases: pydantic.BaseModel
Wire-transmissible representation of a file in the Roboto platform.
FileRecord contains all the metadata and properties associated with a file, including its location, status, ingestion state, and user-defined metadata. This is the data structure used for API communication and persistence.
FileRecord instances are typically created by the platform during file import or upload operations, and are updated as files are processed and modified. The File domain class wraps FileRecord to provide a more convenient interface for file operations.
Parameters
data AnyAttributes
FileRecord.association_id
Properties
FileRecord.bucket
Name of the bucket holding this file’s object.
Raises
This record is a link, which stores no object.
Return type
Attributes
FileRecord.created
FileRecord.created_by
FileRecord.description
FileRecord.device_id
FileRecord.file_id
FileRecord.ingestable
Whether this file is meant to be ingested: its path matched one of its org’s ingestion rules when this version was created or when the file was last renamed or moved, or it has since been partly or fully ingested. A file that is ingestable and ingestion_status not_ingested is awaiting ingestion.
FileRecord.ingestion_status
Properties
FileRecord.is_link
Whether this record is a link to another file rather than a file with an object of its own.
FileRecord.key
Key of this file’s object within bucket.
Raises
This record is a link, which stores no object.
Return type
Attributes
FileRecord.metadata
FileRecord.modified
FileRecord.modified_by
FileRecord.name
FileRecord.org_id
FileRecord.origination
FileRecord.parent_id
FileRecord.relative_path
FileRecord.size
FileRecord.status
FileRecord.storage_type
FileRecord.tags
FileRecord.upload_id
FileRecord.uri
FileRecord.version
FileRecordRequest
Bases: pydantic.BaseModel
Request payload for upserting a file record.
Used to create or update file metadata records in the platform. This is typically used during file import or metadata update operations.
Parameters
data AnyAttributes
FileStatus
Bases: roboto.compat.StrEnum
Enumeration of possible file status values in the Roboto platform.
File status tracks the lifecycle state of a file from initial upload through to availability for use. This status is managed automatically by the platform and affects file visibility and accessibility.
The typical file lifecycle is: Reserved → Available → (optionally) Deleted.
Attributes
FileStatus.Available
File upload is complete and the file is ready for use.
Files with this status are visible in dataset listings, searchable through the query system, and available for download and processing by actions.
FileStatus.Deleted
File is marked for deletion and is no longer accessible.
Files with this status are not visible in listings and cannot be accessed. This status may be temporary during the deletion process.
FileStatus.Reserved
File upload has been initiated but not yet completed.
Files with this status are not yet available for use and are not visible in dataset listings. This is the initial status when an upload begins.
FileStorageType
Bases: roboto.compat.StrEnum
Enumeration of file storage types in the Roboto platform.
Storage type indicates how the file was added to the platform and affects access patterns and permissions. This information is used internally for credential management and access control.
Attributes
FileStorageType.S3Imported
File was imported from a read-only customer-managed S3 bucket.
These files remain in the customer’s bucket and are accessed using customer-provided credentials. The customer retains full control over the file storage and access permissions.
FileStorageType.S3Uploaded
File was uploaded to a Roboto-managed or customer read/write bucket.
These files were explicitly uploaded through the Roboto platform to either a Roboto-managed bucket or a customer’s bring-your-own read/write bucket. Access is managed through Roboto’s credential system.
FileSystem
The files and directories under one association: a dataset, a device, or the org itself.
A FileSystem holds that association’s directory tree and the operations on it: listing, uploading, downloading, renaming, and deleting files, and creating and renaming directories. Datasets, devices, and orgs each return one as files, files, and files.
Every relative path a method takes or returns is relative to the root of the association’s tree, which files of other associations do not share.
Parameters
association roboto.roboto_client Optional[roboto.file_service Optional[roboto.org_id Optional[str]Properties
FileSystem.association
The dataset, device, or org whose files this object works on.
FileSystem.create_directory()
Create a directory among the association’s files.
Parameters
name strName of the directory to create.
error_if_exists boolIf True, raises an exception if the directory already exists.
parent_path Optional[pathlib.Path of the parent directory. If None, creates the directory at the root of the association’s files.
origination Optional[str]Optional string describing the source or context of the directory creation.
create_intermediate_dirs boolIf True, creates intermediate directories in the path if they don’t exist. If False, requires all parent directories to already exist.
Raises
If the directory already exists and error_if_exists is True.
If the caller lacks permission to create the directory.
If the directory name is invalid or the parent path does not exist (when create_intermediate_dirs is False).
Returns
DirectoryRecord of the created directory.
Usage
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
directory = device.files.create_directory("calib")
print(directory.relative_path)
# calibdirectory = device.files.create_directory(
name="final",
parent_path=pathlib.Path("path/to/deep"),
create_intermediate_dirs=True,
)
print(directory.relative_path)
# path/to/deep/finalFileSystem.create_link()
Put a link to another file at relative_path among the association’s files.
A link lets one file, such as a URDF in the org’s own files, appear in many devices’ files without being copied. It pins one version of its target: passing a File pins that file’s version, and passing a file ID pins the target’s current version. Later versions of the target do not move the link; create the link again at the same path to re-point it, which adds a version to the link unless it already pins that target version. Missing parent directories are created. Downloading the link, or asking it for a signed URL, fetches the pinned version of the target.
Parameters
target Union[roboto.The file to link to, or its ID. It must be a file in the same org, not a link or a directory.
relative_path strWhere the link sits, relative to the root of the association’s files.
Returns
The link, whose is_link is True.
Raises
A file or a directory already occupies relative_path. The reverse is refused too: uploading a file to a link’s path is a conflict until the link is deleted.
The target is a link or a directory, is in another org, or does not exist at the version to pin.
The caller cannot edit the association’s files or view the target.
Usage
from roboto.domain import devices, orgs
urdf = orgs.Org.from_id("og_abc123").files.get_file_by_path("urdf/rover/rover.urdf")
device = devices.Device.from_id("rover-01")
link = device.files.create_link(urdf, "urdf/rover.urdf")
link.download(pathlib.Path("/tmp/rover.urdf"))FileSystem.delete_files()
Delete the association’s files that match the given patterns.
Deletes files that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.
Parameters
include_patterns Optional[list[str]]List of gitignore-style patterns for files to include. If None or empty, all files are considered for deletion. An empty list is treated as no filter (all files), not as “include nothing”.
exclude_patterns Optional[list[str]]List of gitignore-style patterns for files to exclude from deletion. Takes precedence over include patterns. If None or empty, no files are excluded.
Raises
Caller lacks permission to delete files.
Return type
Notes
Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.
Usage
from roboto.domain import orgs
org = orgs.Org.from_id("og_abc123")
org.files.delete_files(include_patterns=["**/*.png"], exclude_patterns=["**/back_camera/**"])FileSystem.download_files()
Download the association’s files to a local directory.
Downloads files that match the specified patterns to the given local directory. The files’ directory structure is preserved in the download location. If the output directory doesn’t exist, it will be created. Files are found with list_files(), which does not return links yet, so no link is downloaded; download one with download().
Parameters
out_path pathlib.Local directory path where files should be downloaded.
include_patterns Optional[list[str]]List of gitignore-style patterns for files to include. If None or empty, all files are downloaded. An empty list is treated as no filter (all files), not as “include nothing”.
exclude_patterns Optional[list[str]]List of gitignore-style patterns for files to exclude from download. Takes precedence over include patterns. If None or empty, no files are excluded.
print_progress boolWhether to show a progress bar during download.
Returns
List of tuples containing (FileRecord, local_path) for each downloaded file.
Raises
A selected file’s path resolves outside out_path; nothing is downloaded.
Caller lacks permission to download files.
Notes
Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.
Usage
import pathlib
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
downloaded = device.files.download_files(pathlib.Path("/tmp/rover-01"), include_patterns=["calib/**"])
print(f"Downloaded {len(downloaded)} files")
# Downloaded 2 filesFileSystem.get_file_by_path()
Get a File instance for the association’s file at the specified path.
Parameters
relative_path Union[str, pathlib.Path of the file relative to the root of the association’s files.
version_id Optional[int]Specific version of the file to retrieve. If None, gets the latest version.
Returns
File instance representing the file at the specified path.
Raises
The association has no file at the given path.
Caller lacks permission to access the file.
Usage
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
file = device.files.get_file_by_path("manifest.json")
print(file.file_id)
# fl_xyz789old_file = device.files.get_file_by_path("manifest.json", version_id=1)
print(old_file.version)
# 1FileSystem.list_directories()
Yield every directory among the association’s files, at any depth.
Usage
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
for directory in device.files.list_directories():
print(directory.relative_path)
# calib
# urdfReturn type
FileSystem.list_files()
List the association’s files with optional pattern-based filtering.
Returns all of the association’s files that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.
Parameters
include_patterns Optional[list[str]]List of gitignore-style patterns for files to include. If None or empty, all files are considered. An empty list is treated as no filter (all files), not as “include nothing”.
exclude_patterns Optional[list[str]]List of gitignore-style patterns for files to exclude. Takes precedence over include patterns. If None or empty, no files are excluded.
Yields
File instances that match the specified patterns.
Raises
Caller lacks permission to list files.
Return type
Notes
Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.
Files appear in this list shortly after their upload completes, not instantly.
Usage
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
for file in device.files.list_files():
print(file.relative_path)
# manifest.json
# calib/front_cam.yamlfor file in device.files.list_files(include_patterns=["calib/**"], exclude_patterns=["**/*.bak"]):
print(file.relative_path)
# calib/front_cam.yamlFileSystem.rename_directory()
Rename or move a directory among the association’s files.
Both old_path and new_path are relative to the root of the association’s files. Pass a new_path with fewer path components to move the directory up the tree, or a different leaf name at the same depth to rename in place.
Parameters
old_path strCurrent relative path of the directory (e.g. "logs/session1").
new_path strTarget relative path of the directory (e.g. "session1" to move up one level).
Returns
Updated DirectoryRecord reflecting the new path.
Raises
No directory exists at old_path.
new_path conflicts with an existing node or contains a cycle.
Usage
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
device.files.rename_directory("calib/front", "front_calib")FileSystem.rename_file()
Rename or move a file among the association’s files.
new_path is relative to the root of the association’s files. Pass a path with fewer components to move the file up the tree, a different name at the same depth to rename in place, or a path under a different directory to move sideways.
The file’s storage URI is unchanged; only its relative path changes.
Parameters
file_id strID of the file to rename or move.
new_path strTarget relative path for the file (e.g. "file.bag" to move to the root, or "other_dir/file.bag" to move into an existing directory).
Returns
Updated FileRecord reflecting the new path.
Raises
No file with file_id exists.
new_path conflicts with an existing file, the parent directory does not exist, or the move would create a cycle.
Usage
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
record = device.files.rename_file("fl_xyz789", "manifest.json")
record.relative_path
# 'manifest.json'FileSystem.upload_directory()
Upload all files and directories recursively from the specified directory path.
Use include_patterns and exclude_patterns to control what files and directories are uploaded, and delete_after_upload to clean up your local filesystem after the uploads succeed.
Parameters
directory_path pathlib.Local directory whose contents are uploaded, keeping its layout.
include_patterns Optional[list[str]]gitignore-style patterns for files to include. If None, every file is included.
exclude_patterns Optional[list[str]]gitignore-style patterns for files to exclude. Takes precedence over include_patterns.
delete_after_upload boolIf True, each uploaded local file is deleted once the uploads succeed.
max_batch_size intMaximum number of files per upload transaction.
print_progress boolWhether to display an upload progress bar.
device_id Optional[str]Optional identifier of the device that generated this data.
Return type
Notes
Both pattern lists follow the gitignore pattern format described in https://git-scm.com/docs/gitignore#_pattern_format.
Usage
import pathlib
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
device.files.upload_directory(
pathlib.Path("/path/to/calibration"),
exclude_patterns=["**/*.log"],
)FileSystem.upload_file()
Upload a single file associated with association.
Parameters
file_path pathlib.Local file to upload.
file_destination_path Optional[str]Destination path among the association’s files. Defaults to the file’s own name at the root.
print_progress boolWhether to display an upload progress bar.
device_id Optional[str]Optional identifier of the device that generated this data.
Returns
The uploaded file. Its record is fetched from the platform the first time it is read.
Raises
The upload reported success without reporting a file ID.
Usage
import pathlib
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
device.files.upload_file(pathlib.Path("/path/to/manifest.json"))FileSystem.upload_files()
Upload multiple files associated with association.
Parameters
files collections.Local files to upload.
file_destination_paths collections.Mapping from local path to destination path among the association’s files. Files not in the mapping upload to the root under their own name.
max_batch_size intMaximum number of files per upload transaction.
print_progress boolWhether to display an upload progress bar.
device_id Optional[str]Optional identifier of the device that generated this data.
Returns
Mapping from each uploaded local path to the ID of the file record it created.
Usage
import pathlib
from roboto.domain import devices
device = devices.Device.from_id("rover-01")
file_ids = device.files.upload_files(
[pathlib.Path("/path/to/front_cam.yaml")],
file_destination_paths={pathlib.Path("/path/to/front_cam.yaml"): "calib/front_cam.yaml"},
)
file_ids[pathlib.Path("/path/to/front_cam.yaml")]
# 'fl_0123456789abcdef'FileTag
Bases: enum.Enum
Enumeration of system-defined file tag types.
These tags are used internally by the platform for indexing and organizing files. They are automatically applied during file operations and should not be manually modified by users.
Attributes
FileTag.AssociationId
Tag containing the ID of the dataset, device, or org a file is associated with.
FileTag.CommonPrefix
Tag containing the common path prefix for files in a batch operation.
FileTag.DatasetId
Tag containing the ID of the dataset that contains this file.
Deprecated in favour of AssociationId, which names a file’s dataset, device, or org alike. The platform still sets it on its own server-side copies of dataset files.
FileTag.TransactionId
Tag containing the transaction ID for files uploaded in a batch.
ImportFileRequest
Bases: pydantic.BaseModel
Request payload for importing an existing file into a dataset.
Used to register files that already exist in storage (such as customer S3 buckets) with the Roboto platform. The file content remains in its original location while metadata is stored in Roboto for discovery and processing.
Parameters
data AnyAttributes
ImportFileRequest.description
Optional human-readable description of the file.
ImportFileRequest.device_id
Optional identifier of the device that generated this data.
ImportFileRequest.metadata
Optional key-value metadata pairs to associate with the file.
ImportFileRequest.relative_path
Path of the file relative to the dataset root (e.g., logs/session1.bag).
ImportFileRequest.size
Size of the file in bytes. When importing a single file, you can omit the size, as Roboto will look up the size from the object store. When calling import_batch, you must provide the size explicitly.
ImportFileRequest.tags
Optional list of tags for file discovery and organization.
ImportFileRequest.uri
Storage URI where the file is located (e.g., s3://bucket/path/to/file.bag).
IngestionStatus
Bases: roboto.compat.StrEnum
Enumeration of file ingestion status values in the Roboto platform.
Ingestion status tracks whether a file’s data has been processed and extracted into topics for analysis and visualization. This status determines what platform features are available for the file and whether it can trigger automated workflows.
File ingestion happens as a post-upload processing step. Roboto supports many common robotics log formats (ROS bags, MCAP files, ULOG files, etc.) out-of-the-box. Custom ingestion actions can be written for other formats.
When writing custom ingestion actions, be sure to update the file’s ingestion status to mark it as fully ingested. This enables triggers and other automated workflows that depend on complete ingestion.
Ingested files have first-class visualization support and can be queried through the topic data system.
Attributes
IngestionStatus.Ingested
All topics from this file have been fully processed and recorded.
Files with this status have complete topic data available for visualization, analysis, and querying. They are eligible for post-ingestion triggers and automated workflows that depend on complete data extraction.
IngestionStatus.NotIngested
No topics from this file have been processed or recorded.
Files with this status have not undergone data extraction. They cannot be visualized through the topic system and are not eligible for topic-based triggers or analysis workflows.
IngestionStatus.PartlyIngested
Some but not all topics from this file have been processed.
Files with this status have at least one topic record but ingestion is incomplete. Some visualization and analysis features may be available, but the file is not yet eligible for post-ingestion triggers.
LazyLookupFile
Bases: roboto.domain.files.file.File
A File subclass that defers instantiation (hydration) of the real File until any non‐internal attribute is first accessed.
This is useful for scenarios where we know how to dereference a File (e.g., by ID), and we want to return a handle in case the caller wants to work with it, but we don’t want to pay the cost of dereferencing it unless necessary.
Parameters
hydrate_fn Callable[[], roboto.QueryDatasetFilesRequest
Bases: pydantic.BaseModel
Request payload for listing the files associated with a dataset, an org, or a device.
Supports gitignore-style patterns for flexible file selection and pagination. Despite the name, the same body lists the files of any association type.
Parameters
data AnyAttributes
QueryDatasetFilesRequest.exclude_patterns
List of gitignore-style patterns for files to exclude from results.
QueryDatasetFilesRequest.include_patterns
List of gitignore-style patterns for files to include in results.
QueryDatasetFilesRequest.page_token
Token for retrieving the next page of results in paginated queries.
QueryDatasetFilesRequest.sort_by
Field to sort results by. Defaults to ‘created’.
QueryDatasetFilesRequest.sort_direction
Sort direction (‘ASC’ or ‘DESC’). Defaults to ‘DESC’.
QueryFilesRequest
Bases: pydantic.BaseModel
Request payload for querying files with filters.
Used to search for files based on various criteria such as metadata, tags, ingestion status, and other file properties. The filters are applied server-side to efficiently return matching files.
Parameters
data AnyRenameDirectoryRequest
Bases: pydantic.BaseModel
Request payload for renaming a directory among the files of one association.
Changes the path of a directory and all its contained files. This updates the logical organization without moving actual file content.
Parameters
data AnyRenameFileRequest
Bases: pydantic.BaseModel
Request payload for renaming a file within its dataset.
Changes the relative path of a file within its dataset. This updates the file’s logical location but does not move the actual file content in storage.
Parameters
data AnySignedUrlResponse
Bases: pydantic.BaseModel
Response containing a signed URL for direct file access.
Provides a time-limited URL that allows direct access to file content without requiring Roboto authentication. Used for file downloads and integration with external systems.
Parameters
data AnyAttributes
UpdateFileRecordRequest
Bases: pydantic.BaseModel
Request payload for updating file record properties.
Used to modify file metadata, description, and ingestion status. Only specified fields are updated; others remain unchanged. Uses NotSet sentinel values to distinguish between explicit None values and fields that should not be modified.
Parameters
data AnyAttributes
UpdateFileRecordRequest.description
New description for the file, or NotSet to leave unchanged.
UpdateFileRecordRequest.device_id
New device ID for the file, or NotSet to leave unchanged.
UpdateFileRecordRequest.ingestion_complete
Set to True to mark file as fully ingested, or NotSet to leave unchanged.
UpdateFileRecordRequest.metadata_changeset
Metadata changes to apply (add, update, or remove fields/tags), or NotSet to leave unchanged.
UpdateFileRecordRequest.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
is_directory()
Parameters
record Union[FileRecord, DirectoryRecord]Return type
is_file()
Parameters
record Union[FileRecord, DirectoryRecord]Return type