Skip to content
Roboto
Esc
↑↓navigate↵open⌘Jpreview
On this page

roboto.domain.datasets

Submodules

Package Contents

CreateDatasetIfNotExistsRequest

class roboto.domain.datasets.CreateDatasetIfNotExistsRequest(/, **data)#View Source

Bases: pydantic.BaseModel

Request payload to create a dataset if no existing dataset matches the specified query.

Searches for existing datasets using the provided RoboQL query. If a matching dataset is found, returns that dataset. If no match is found, creates a new dataset with the specified properties and returns it.

Parameters

data Any

Attributes

CreateDatasetIfNotExistsRequest.create_request

create_request CreateDatasetRequest #

CreateDatasetIfNotExistsRequest.match_roboql_query

match_roboql_query str #

CreateDatasetRequest

class roboto.domain.datasets.CreateDatasetRequest(**data)#View Source

Bases: pydantic.BaseModel

Request payload for creating a new dataset.

Used to specify the initial properties of a dataset during creation, including optional metadata, tags, name, and description.

Attributes

CreateDatasetRequest.custom_fields

custom_fields dict[str, Any] | None = None #

Initial values for Ready custom fields on this dataset.

Each key must be the name of a CustomField that is Ready for the caller’s org and the Dataset entity type; each value must satisfy the field’s declared type. Names that are undefined or not Ready, and values that don’t match the field’s type, are rejected with a structured error.

CreateDatasetRequest.description

description str | None = None #

Optional human-readable description of the dataset.

CreateDatasetRequest.device_id

device_id str | None = None #

Optional identifier of the device that generated this data.

CreateDatasetRequest.metadata

metadata dict[str, Any] = None #

Key-value metadata pairs to associate with the dataset for discovery and search.

CreateDatasetRequest.name

name str | None = None #

Optional short name for the dataset (max 120 characters).

CreateDatasetRequest.tags

tags list[str] = None #

List of tags for dataset discovery and organization.

CreateDirectoryRequest

class roboto.domain.datasets.CreateDirectoryRequest(/, **data)#View Source

Bases: pydantic.BaseModel

Request payload to create a directory among the files of one association.

Parameters

data Any

Attributes

CreateDirectoryRequest.create_intermediate_dirs

create_intermediate_dirs bool = False #

If True, creates intermediate directories in the path if they don’t exist. If False, requires all parent directories to already exist.

CreateDirectoryRequest.error_if_exists

error_if_exists bool = False #

CreateDirectoryRequest.name

name str #

CreateDirectoryRequest.origination

origination str | None = None #

CreateDirectoryRequest.parent_path

parent_path str | None = None #

Dataset

class roboto.domain.datasets.Dataset(record, roboto_client=None, file_service=None, content_mode=None)#View Source

Represents a dataset within the Roboto platform.

A dataset is a logical container for files organized in a directory structure. Datasets are the primary organizational unit in Roboto, typically containing files from a single robot activity such as a drone flight, autonomous vehicle mission, or sensor data collection session. However, datasets are versatile enough to serve as a general-purpose assembly of files.

Datasets provide functionality for:

  • File upload and download operations
  • Metadata and tag management
  • File organization and directory operations
  • Topic data access and analysis
  • AI-powered content summarization
  • Integration with automated workflows and triggers

Files within a dataset can be processed by actions, visualized in the web interface, and searched using the query system. Datasets inherit access permissions from their organization and can be shared with other users and systems.

The Dataset class serves as the primary interface for dataset operations in the Roboto SDK, providing methods for file management, metadata operations, and content analysis.

Parameters

Dataset.clear_custom_field()

clear_custom_field(name)#View Source

Clear a single custom-field value on this dataset to None.

Parameters

name str

Return type

Dataset.clear_custom_fields()

clear_custom_fields(names)#View Source

Clear multiple custom-field values on this dataset to None.

Parameters

names collections.abc.Sequence[str]

Return type

Dataset.create()

classmethod create(description=None, metadata=None, name=None, tags=None, device_id=None, custom_fields=None, caller_org_id=None, roboto_client=None, create_device_if_missing=False)#View Source

Create a new dataset in the Roboto platform.

Creates a new dataset with the specified properties and returns a Dataset instance for interacting with it. The dataset will be created in the caller’s organization unless a different organization is specified.

Parameters

description Optional[str]

Optional human-readable description of the dataset.

metadata Optional[dict[str, Any]]

Optional key-value metadata pairs to associate with the dataset.

name Optional[str]

Optional short name for the dataset (max 120 characters).

tags Optional[list[str]]

Optional list of tags for dataset discovery and organization.

device_id Optional[str]

Optional identifier of the device that generated this data.

custom_fields Optional[dict[str, Any]]

Optional initial values for Ready custom fields defined on Datasets in the caller’s org. Keys must match Ready field names; values must satisfy each field’s declared type.

caller_org_id Optional[str]

Organization ID to create the dataset in. Required for multi-org users.

roboto_client Optional[roboto.http.RobotoClient]

HTTP client for API communication. If None, uses the default client.

create_device_if_missing bool

If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.

Returns

Dataset instance representing the newly created dataset.

Raises

A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.

Invalid dataset parameters.

Caller lacks permission to create datasets.

Usage

dataset = Dataset.create(
    name="Highway Test Session",
    description="Autonomous vehicle highway driving test data",
    tags=["highway", "autonomous", "test"],
    metadata={"vehicle_id": "vehicle_001", "test_type": "highway"},
)
print(dataset.dataset_id)
# ds_abc123
# Create minimal dataset
dataset = Dataset.create()
print(f"Created dataset: {dataset.dataset_id}")

Dataset.create_directory()

create_directory(name, error_if_exists=False, create_intermediate_dirs=False, parent_path=None, origination=None)#View Source

Create a directory within the dataset.

Parameters

name str

Name of the directory to create.

error_if_exists bool

If True, raises an exception if the directory already exists.

parent_path Optional[pathlib.Path]

Path of the parent directory. If None, creates the directory in the root of the dataset.

origination Optional[str]

Optional string describing the source or context of the directory creation.

create_intermediate_dirs bool

If True, creates intermediate directories in the path if they don’t exist. If False, requires all parent directories to already exist.

Raises

If the directory already exists and error_if_exists is True.

If the caller lacks permission to create the directory.

If the directory name is invalid or the parent path does not exist (when create_intermediate_dirs is False).

Returns

DirectoryRecord of the created directory.

Usage

Create a simple directory:

from roboto.domain import datasets
dataset = datasets.Dataset.from_id(...)
directory = dataset.create_directory("foo")
print(directory.relative_path)
# foo

Create a directory with intermediate directories:

directory = dataset.create_directory(
    name="final",
    parent_path=pathlib.Path("path/to/deep"),
    create_intermediate_dirs=True,
)
print(directory.relative_path)
# path/to/deep/final

Dataset.create_if_not_exists()

classmethod create_if_not_exists(match_roboql_query, description=None, metadata=None, name=None, tags=None, device_id=None, custom_fields=None, caller_org_id=None, roboto_client=None, create_device_if_missing=False)#View Source

Create a dataset if no existing dataset matches the specified query.

Searches for existing datasets using the provided RoboQL query. If a matching dataset is found, returns that dataset. If no match is found, creates a new dataset with the specified properties and returns it.

Concurrent calls with the same match_roboql_query in one organization create one dataset between them: the service runs them one at a time from the search through the create. Calls with different queries are not serialized, even when both queries would match the same dataset.

The dataset created must match match_roboql_query. If name, tags or metadata describe a dataset the query does not match, every later call creates another one. When several datasets match, which one is returned is not defined unless the query ends with a SORT BY clause.

Parameters

match_roboql_query str

RoboQL query string to search for existing datasets. If this query matches any dataset, that dataset will be returned instead of creating a new one.

description Optional[str]

Optional human-readable description of the dataset.

metadata Optional[dict[str, Any]]

Optional key-value metadata pairs to associate with the dataset.

name Optional[str]

Optional short name for the dataset (max 120 characters).

tags Optional[list[str]]

Optional list of tags for dataset discovery and organization.

device_id Optional[str]

Optional identifier of the device that generated this data.

custom_fields Optional[dict[str, Any]]

Optional initial values for Ready custom fields defined on Datasets in caller_org_id. Keys must match Ready field names; values must satisfy each field’s declared type. Ignored when an existing dataset matches match_roboql_query — the existing record is returned unchanged.

caller_org_id Optional[str]

Organization ID to create the dataset in. Required for multi-org users.

roboto_client Optional[roboto.http.RobotoClient]

HTTP client for API communication. If None, uses the default client.

create_device_if_missing bool

If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.

Returns

Dataset instance representing either the existing matched dataset or the newly created dataset.

Raises

A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.

Invalid dataset parameters or malformed RoboQL query.

Other calls with the same query kept this one waiting for more than 10 seconds, after the SDK’s own retries. Calling again is safe.

Caller lacks permission to create datasets or search existing ones.

Usage

Create a dataset only if no dataset with specific metadata exists:

dataset = Dataset.create_if_not_exists(
    match_roboql_query="dataset.metadata.vehicle_id = 'vehicle_001'",
    name="Vehicle 001 Test Session",
    description="Test data for vehicle 001",
    metadata={"vehicle_id": "vehicle_001", "test_type": "highway"},
    tags=["vehicle_001", "highway"],
)
print(dataset.dataset_id)
# ds_abc123

Create a dataset only if no dataset with specific tags exists:

dataset = Dataset.create_if_not_exists(
    match_roboql_query="dataset.tags CONTAINS 'unique_session_id_xyz'",
    name="Unique Test Session",
    tags=["unique_session_id_xyz", "test"],
)
# If a dataset with tag 'unique_session_id_xyz' already exists,
# that dataset is returned instead of creating a new one

Dataset.create_session()

create_session(name=None, *, device_ids=None, include_patterns=None, exclude_patterns=None, description=None, metadata=None, tags=None, custom_fields=None)#View Source

Create a Session populated with files from this dataset.

By default, every file in the dataset is added to the new Session. include_patterns / exclude_patterns narrow that set using the same gitignore-style syntax as Dataset.list_files().

Parameters

name Optional[str]

Short display name for the Session (max 120 characters).

device_ids Optional[collections.abc.Sequence[str]]

Devices to attach to the Session. Defaults to no devices.

include_patterns Optional[list[str]]

Gitignore-style patterns selecting which files to include. Same syntax as Dataset.list_files().

exclude_patterns Optional[list[str]]

Gitignore-style patterns selecting which files to exclude. Takes precedence over include_patterns.

description Optional[str]

Optional description of the Session.

metadata Optional[dict[str, Any]]

Optional initial metadata. Sessions are not filterable or sortable by metadata keys; for queryable structured attributes, define a custom field on the Session entity type.

tags Optional[collections.abc.Sequence[str]]

Optional initial tags. Sessions can be filtered by tag membership but are not sortable by tag.

custom_fields Optional[dict[str, Any]]

Optional initial values for Ready custom fields defined on Sessions in this dataset’s org. Keys must match Ready field names; values must satisfy each field’s declared type.

Returns

The newly created Session.

Usage

dataset = Dataset.from_id("ds_abc123")
session = dataset.create_session("flight-2026-04-23-001")

Notes

Convenience wrapper around Session.create() followed by Session.add_files(). A dataset may hold more files than one add request accepts: the files are sent in consecutive requests of at most MAX_FILES_AND_TOPICS_PER_REQUEST each. Every file goes in or none does: if any add request fails, or the platform refuses any one file, the partially populated Session is deleted before the exception propagates, so the call is safe to retry. If that cleanup fails too (e.g. a transient network error), the Session is left behind holding whatever files did go in; the cleanup failure is logged, and the original exception is what the caller sees.

Properties

Dataset.created

created datetime.datetime #

Timestamp when this dataset was created.

Returns the UTC datetime when this dataset was first created in the Roboto platform. This property is immutable.

Return type: datetime.datetime

Dataset.created_by

created_by str #

Identifier of the user who created this dataset.

Returns the identifier of the person or service which originally created this dataset in the Roboto platform.

Return type: str

Dataset.custom_fields

custom_fields dict[str, Any] #

Custom-field values defined on Datasets in this org.

Every Ready CustomField defined for (org_id, Dataset) appears as a key. Values that have not been set on this dataset surface as None rather than being absent. Empty when no custom fields are defined for the org.

A Timestamp value is returned as an ISO 8601 string.

Return type: dict[str, Any]

Dataset.dataset_id

dataset_id str #

Unique identifier for this dataset.

Returns the globally unique identifier assigned to this dataset when it was created. This ID is immutable and used to reference the dataset across the Roboto platform. It is always prefixed with ‘ds_’ to distinguish it from other Roboto resource IDs.

Return type: str

Dataset.delete()

delete()#View Source

Delete this dataset from the Roboto platform.

Permanently removes the dataset and all its associated files, metadata, and topics. This operation cannot be undone.

If a dataset’s files are hosted in Roboto managed S3 buckets or customer read/write bring-your-own-buckets, the files in this dataset will be deleted from S3 as well. For files hosted in customer read-only buckets, the files will not be deleted from S3, but the dataset record and all associated metadata will be deleted.

Raises

Dataset does not exist or has already been deleted.

Caller lacks permission to delete the dataset.

Return type

None

Usage

dataset = Dataset.from_id("ds_abc123")
dataset.delete()
# # Dataset and all its files are now permanently deleted

Dataset.delete_files()

delete_files(include_patterns=None, exclude_patterns=None)#View Source

Delete files from this dataset based on pattern matching.

Deletes files that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.

Parameters

include_patterns Optional[list[str]]

List of gitignore-style patterns for files to include. If None or empty, all files are considered for deletion. An empty list is treated as no filter (all files), not as “include nothing”.

exclude_patterns Optional[list[str]]

List of gitignore-style patterns for files to exclude from deletion. Takes precedence over include patterns. If None or empty, no files are excluded.

Raises

Caller lacks permission to delete files.

Return type

None

Notes

Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.

Usage

dataset = Dataset.from_id("ds_abc123")
# Delete all PNG files except those in back_camera directory
dataset.delete_files(include_patterns=["**/*.png"], exclude_patterns=["**/back_camera/**"])
# Delete all log files
dataset.delete_files(include_patterns=["**/*.log"])

Properties

Dataset.description

description str | None #

Human-readable description of this dataset.

Returns the optional description text that provides details about the dataset’s contents, purpose, or context. Can be None if no description was provided.

Return type: Optional[str]

Dataset.device_id

device_id str | None #

Identifier of the device that generated this data.

Returns the optional identifier of the device that generated the data contained within this dataset. Can be None if the dataset was not generated by a device.

Return type: Optional[str]

Dataset.download_files()

download_files(out_path, include_patterns=None, exclude_patterns=None, print_progress=True)#View Source

Download files from this dataset to a local directory.

Downloads files that match the specified patterns to the given local directory. The directory structure from the dataset is preserved in the download location. If the output directory doesn’t exist, it will be created.

Parameters

out_path pathlib.Path

Local directory path where files should be downloaded.

include_patterns Optional[list[str]]

List of gitignore-style patterns for files to include. If None or empty, all files are downloaded. An empty list is treated as no filter (all files), not as “include nothing”.

exclude_patterns Optional[list[str]]

List of gitignore-style patterns for files to exclude from download. Takes precedence over include patterns. If None or empty, no files are excluded.

print_progress bool

Whether to show a progress bar during download.

Returns

list[tuple[roboto.domain.files.FileRecord, pathlib.Path]]

List of tuples containing (FileRecord, local_path) for each downloaded file.

Raises

Caller lacks permission to download files.

Notes

Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.

Usage

import pathlib
dataset = Dataset.from_id("ds_abc123")
downloaded = dataset.download_files(
    pathlib.Path("/tmp/dataset_download"),
    include_patterns=["**/*.bag"],
    exclude_patterns=["**/test/**"],
)
print(f"Downloaded {len(downloaded)} files")
# Downloaded 5 files
# Download all files
all_files = dataset.download_files(pathlib.Path("/tmp/all_files"))

Properties

Dataset.files

The files associated with this dataset.

The file methods on Dataset delegate here, so dataset.upload_files(...) and dataset.files.upload_files(...) are the same call.

Dataset.from_id()

classmethod from_id(dataset_id, roboto_client=None)#View Source

Create a Dataset instance from a dataset ID.

Retrieves dataset information from the Roboto platform using the provided dataset ID and returns a Dataset instance for interacting with it.

Parameters

dataset_id str

Unique identifier for the dataset.

roboto_client Optional[roboto.http.RobotoClient]

HTTP client for API communication. If None, uses the default client.

Returns

Dataset instance representing the requested dataset.

Raises

Dataset with the given ID does not exist.

Caller lacks permission to access the dataset.

Usage

dataset = Dataset.from_id("ds_abc123")
print(dataset.name)
# 'Highway Test Session'
print(len(list(dataset.list_files())))
# 42

Dataset.generate_summary()

generate_summary()#View Source

Generate a new AI summary for this dataset.

Creates a new AI-generated summary that analyzes the dataset’s content, structure, and metadata. The summary generation is asynchronous and can be monitored through the returned StreamingAISummary object.

Returns

StreamingAISummary object that provides access to the summary as it is being generated. The summary starts in pending status and can be monitored for completion.

Raises

Caller lacks permission to generate summaries for this dataset.

Usage

Generate a summary and wait for completion:

dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
complete_text = summary.complete_text
print(complete_text)
# 'This dataset contains 42 files with sensor data from highway driving tests...'

Generate a summary and stream the text as it’s generated:

dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
for text_chunk in summary.text_stream():
    print(text_chunk, end="", flush=True)

Check summary status without blocking:

dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
if summary.current and summary.current.status == AISummaryStatus.Complete:
    print("Summary is ready!")

Dataset.get_file_by_path()

get_file_by_path(relative_path, version_id=None)#View Source

Get a File instance for a file at the specified path in this dataset.

Retrieves a file by its relative path within the dataset. Optionally retrieves a specific version of the file.

Parameters

relative_path Union[str, pathlib.Path]

Path of the file relative to the dataset root.

version_id Optional[int]

Specific version of the file to retrieve. If None, gets the latest version.

Returns

File instance representing the file at the specified path.

Raises

File at the given path does not exist in the dataset.

Caller lacks permission to access the file.

Usage

dataset = Dataset.from_id("ds_abc123")
file = dataset.get_file_by_path("logs/session1.bag")
print(file.file_id)
# file_xyz789
# Get specific version
old_file = dataset.get_file_by_path("data/sensors.csv", version_id=1)
print(old_file.version)
# 1

Dataset.get_metadata()

get_metadata()#View Source

Return custom metadata associated with this dataset.

Returns a copy of the dataset’s metadata dictionary containing arbitrary key-value pairs for storing custom information. Supports nested structures and dot notation for accessing nested fields.

Return type

dict[str, Any]

Dataset.get_sessions()

get_sessions()#View Source

Iterate over Sessions that include at least one file from this dataset.

A Session may draw files from one or more datasets; this method yields every Session whose current-version file list intersects this dataset.

Yields

Each matching Session. Pagination is handled automatically.

Return type

collections.abc.Generator[roboto.experimental.sessions.Session, None, None]

Usage

dataset = Dataset.from_id("ds_abc123")
for session in dataset.get_sessions():
    print(session.session_id, session.name)

Dataset.get_summary()

get_summary()#View Source

Retrieve this dataset’s existing AI summary.

Returns the dataset’s current AI summary if one exists. Reading never generates a summary as a side effect: a dataset that has never been summarized raises RobotoNotFoundException rather than implicitly kicking off — and paying for — generation. Call generate_summary() to create one explicitly.

Returns

StreamingAISummary wrapping the dataset’s existing summary. If a generation kicked off elsewhere is still in flight, the returned summary is Pending; poll it via await_completion or text_stream.

Raises

This dataset has no AI summary yet. Call generate_summary() to create one.

Caller lacks permission to access summaries for this dataset.

Usage

Get the existing summary, generating one first if there is none:

from roboto.exceptions import RobotoNotFoundException
dataset = Dataset.from_id("ds_abc123")
try:
    summary = dataset.get_summary()
except RobotoNotFoundException:
    summary = dataset.generate_summary()
print(summary.complete_text)
# 'This dataset contains 42 files with sensor data from highway driving tests...'

Check whether a summary exists without generating one:

from roboto.exceptions import RobotoNotFoundException
dataset = Dataset.from_id("ds_abc123")
try:
    summary = dataset.get_summary()
    print(summary.complete_text)
except RobotoNotFoundException:
    print("No summary yet — call generate_summary() to create one.")

Dataset.get_topic_time_bounds()

get_topic_time_bounds()#View Source

Get the earliest start and latest end across every topic in this dataset.

The same aggregate you would reach by folding start_time and end_time over get_topics(), computed server-side in one request instead of one per page of topics. Reach for it when you want the dataset’s time extent and not the topics themselves.

Returns

Bounds in nanoseconds since the Unix epoch. Both fields are None for a dataset whose files hold no topics, and either is None when no topic in the dataset carries that timestamp.

Raises

Dataset does not exist.

Caller lacks permission to access the dataset.

Usage

dataset = Dataset.from_id("ds_abc123")
bounds = dataset.get_topic_time_bounds()
print(bounds.start_time, bounds.end_time)
# 1722870127699468923 1722870187004821001

Dataset.get_topics()

get_topics(include=None, exclude=None)#View Source

Get all topics associated with files in this dataset, with optional filtering.

Retrieves all topics that were extracted from files in this dataset during ingestion. If multiple files have topics with the same name (e.g., chunked files with the same schema), they are returned as separate topic objects.

Topics can be filtered by name using include/exclude patterns. Topics specified on both the inclusion and exclusion lists will be excluded.

Parameters

include Optional[collections.abc.Sequence[str]]

If provided, only topics with names in this sequence are yielded.

exclude Optional[collections.abc.Sequence[str]]

If provided, topics with names in this sequence are skipped. Takes precedence over include list.

Yields

Topic instances associated with files in this dataset, filtered according to the parameters.

Return type

collections.abc.Generator[roboto.domain.topics.Topic, None, None]

Usage

dataset = Dataset.from_id("ds_abc123")
for topic in dataset.get_topics():
    print(f"Topic: {topic.name}")
# Topic: /camera/image
# Topic: /imu/data
# Topic: /gps/fix
# Only get camera topics
camera_topics = list(dataset.get_topics(include=["/camera/image", "/camera/info"]))
print(f"Found {len(camera_topics)} camera topics")
# Exclude diagnostic topics
data_topics = list(dataset.get_topics(exclude=["/diagnostics"]))

Dataset.get_topics_by_file()

get_topics_by_file(relative_path)#View Source

Get all topics associated with a specific file in this dataset.

Retrieves all topics that were extracted from the specified file during ingestion. This is a convenience method that combines file lookup and topic retrieval.

Parameters

relative_path Union[str, pathlib.Path]

Path of the file relative to the dataset root.

Yields

Topic instances associated with the specified file.

Raises

File at the given path does not exist in the dataset.

Caller lacks permission to access the file or its topics.

Return type

collections.abc.Generator[roboto.domain.topics.Topic, None, None]

Usage

dataset = Dataset.from_id("ds_abc123")
for topic in dataset.get_topics_by_file("logs/session1.bag"):
    print(f"Topic: {topic.name}")
# Topic: /camera/image
# Topic: /imu/data
# Topic: /gps/fix

Dataset.list_directories()

list_directories()#View Source

Yield every directory in this dataset, at any depth.

Usage

dataset = Dataset.from_id("ds_abc123")
for directory in dataset.list_directories():
    print(directory.relative_path)
# logs
# logs/session1

Return type

collections.abc.Generator[roboto.domain.files.DirectoryRecord, None, None]

Dataset.list_files()

list_files(include_patterns=None, exclude_patterns=None)#View Source

List files in this dataset with optional pattern-based filtering.

Returns all files in the dataset that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.

Parameters

include_patterns Optional[list[str]]

List of gitignore-style patterns for files to include. If None or empty, all files are considered. An empty list is treated as no filter (all files), not as “include nothing”.

exclude_patterns Optional[list[str]]

List of gitignore-style patterns for files to exclude. Takes precedence over include patterns. If None or empty, no files are excluded.

Yields

File instances that match the specified patterns.

Raises

Caller lacks permission to list files.

Return type

collections.abc.Generator[roboto.domain.files.File, None, None]

Notes

Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.

Usage

dataset = Dataset.from_id("ds_abc123")
for file in dataset.list_files():
    print(file.relative_path)
# logs/session1.bag
# data/sensors.csv
# images/camera_001.jpg
# List only image files, excluding back camera
for file in dataset.list_files(
    include_patterns=["**/*.png", "**/*.jpg"], exclude_patterns=["**/back_camera/**"]
):
    print(file.relative_path)
# images/front_camera_001.jpg
# images/side_camera_001.jpg

Properties

Dataset.metadata

metadata dict[str, Any] #

Custom metadata associated with this dataset.

Returns a copy of the dataset’s metadata dictionary containing arbitrary key-value pairs for storing custom information. Supports nested structures and dot notation for accessing nested fields.

Note: this attribute is kept for backward compatibility. Prefer get_metadata(), since metadata may need to be loaded on-demand from the server.

Return type: dict[str, Any]

Dataset.modified

modified datetime.datetime #

Timestamp when this dataset was last modified.

Returns the UTC datetime when this dataset was most recently updated. This includes changes to metadata, tags, description, or other properties.

Return type: datetime.datetime

Dataset.modified_by

modified_by str #

Identifier of the user or service which last modified this dataset.

Returns the identifier of the person or service which most recently updated this dataset’s metadata, tags, description, or other properties.

Return type: str

Dataset.name

name str | None #

Human-readable name of this dataset.

Returns the optional display name for this dataset. Can be None if no name was provided during creation. For users whose organizations have their own idiomatic internal dataset IDs, it’s recommended to set the name to the organization’s internal dataset ID, since the Roboto dataset_id property is randomly generated.

Return type: Optional[str]

Dataset.org_id

org_id str #

Organization identifier that owns this dataset.

Returns the unique identifier of the organization that owns and has primary access control over this dataset.

Return type: str

Dataset.put_metadata()

put_metadata(metadata)#View Source

Add or update metadata fields for this dataset.

Sets each key-value pair in the provided dictionary as dataset metadata. If a key doesn’t exist, it will be created. If it exists, the value will be overwritten. Keys must be strings and dot notation is supported for nested keys.

Parameters

metadata dict[str, Any]

Dictionary of metadata key-value pairs to add or update.

Raises

Caller lacks permission to update the dataset.

Return type

None

Usage

dataset = Dataset.from_id("ds_abc123")
dataset.put_metadata(
    {
        "vehicle_id": "vehicle_001",
        "test_type": "highway_driving",
        "weather.condition": "sunny",
        "weather.temperature": 25,
    }
)
print(dataset.metadata["vehicle_id"])
# 'vehicle_001'
print(dataset.metadata["weather"]["condition"])
# 'sunny'

Dataset.put_tags()

put_tags(tags)#View Source

Add or update tags for this dataset.

Adds each tag in the provided sequence to the dataset. If a tag already exists, it will not be duplicated. This operation replaces the current tag list with the provided tags.

Parameters

Sequence of tag strings to set on the dataset.

Raises

Caller lacks permission to update the dataset.

Return type

None

Usage

dataset = Dataset.from_id("ds_abc123")
dataset.put_tags(["highway", "autonomous", "test", "sunny"])
print(dataset.tags)
# ['highway', 'autonomous', 'test', 'sunny']

Dataset.query()

classmethod query(spec=None, roboto_client=None, owner_org_id=None)#View Source

Query datasets using a specification with filters and pagination.

Searches for datasets matching the provided query specification. Results are returned as a generator that automatically handles pagination, yielding Dataset instances as they are retrieved from the API.

Parameters

Query specification with filters, sorting, and pagination options. If None, returns all accessible datasets.

roboto_client Optional[roboto.http.RobotoClient]

HTTP client for API communication. If None, uses the default client.

owner_org_id Optional[str]

Organization ID to scope the query. If None, uses caller’s org.

Yields

Dataset instances matching the query specification.

Raises

ValueError

Query specification references unknown dataset attributes.

Caller lacks permission to query datasets.

Return type

collections.abc.Generator[Dataset, None, None]

Usage

from roboto.query import Comparator, Condition, QuerySpecification
spec = QuerySpecification(
    condition=Condition(field="name", comparator=Comparator.Contains, value="Roboto")
)
for dataset in Dataset.query(spec):
    print(f"Found dataset: {dataset.name}")
# Found dataset: Roboto Test
# Found dataset: Other Roboto Test

Properties

Dataset.record

Underlying data record for this dataset.

Returns the raw DatasetRecord that contains all the dataset’s data fields. This provides access to the complete dataset state as stored in the platform.

Dataset.refresh()

refresh()#View Source

Refresh this dataset instance with the latest data from the platform.

Fetches the current state of the dataset from the Roboto platform and updates this instance’s data. Useful when the dataset may have been modified by other processes or users.

Returns

This Dataset instance with refreshed data.

Raises

Dataset no longer exists.

Caller lacks permission to access the dataset.

Usage

dataset = Dataset.from_id("ds_abc123")
# Dataset may have been updated by another process
refreshed_dataset = dataset.refresh()
print(f"Current file count: {len(list(refreshed_dataset.list_files()))}")

Dataset.remove_metadata()

remove_metadata(metadata)#View Source

Remove each key in this sequence from dataset metadata if it exists. Keys must be strings. Dot notation is supported for nested keys.

Usage

from roboto.domain import datasets
dataset = datasets.Dataset(...)
dataset.remove_metadata(["foo", "baz.qux"])

Parameters

Return type

None

Dataset.remove_tags()

remove_tags(tags)#View Source

Remove each tag in this sequence if it exists

Return type

None

Dataset.rename_directory()

rename_directory(old_path, new_path)#View Source

Rename or move a directory within this dataset.

Both old_path and new_path are relative to the dataset root. Pass a new_path with fewer path components to move the directory up the tree, or a different leaf name at the same depth to rename in place.

Parameters

old_path str

Current relative path of the directory (e.g. "logs/session1").

new_path str

Target relative path of the directory (e.g. "session1" to move up one level).

Returns

Raises

No directory exists at old_path.

new_path conflicts with an existing node or contains a cycle.

Usage

dataset = Dataset.from_id("ds_abc123")
dataset.rename_directory("logs/session1", "session1")

Dataset.rename_file()

rename_file(file_id, new_path)#View Source

Rename or move a file within this dataset.

new_path is relative to the dataset root. Pass a path with fewer components to move the file up the tree, a different name at the same depth to rename in place, or a path under a different directory to move sideways.

The file’s storage URI is unchanged; only the logical location in the dataset hierarchy moves.

Parameters

file_id str

ID of the file to rename or move.

new_path str

Target relative path for the file within this dataset (e.g. "file.bag" to move to the root, or "other_dir/file.bag" to move into an existing directory).

Returns

Updated FileRecord reflecting the new path.

Raises

No file with file_id exists.

new_path conflicts with an existing file, the parent directory does not exist, or the move would create a cycle.

Usage

dataset = Dataset.from_id("ds_abc123")
record = dataset.rename_file("file_xyz789", "file.bag")
record.relative_path
# 'file.bag'

Dataset.set_custom_field()

set_custom_field(name, value)#View Source

Set a single custom-field value on this dataset.

name must be the name of a Ready custom field for this dataset’s org and the Dataset entity type; value must satisfy the field’s declared type.

Parameters

name str
value Any

Return type

Dataset.set_custom_fields()

set_custom_fields(fields)#View Source

Set or overwrite multiple custom-field values on this dataset.

Each key must name a Ready custom field for this dataset’s org and the Dataset entity type; each value must satisfy the field’s declared type.

Parameters

fields dict[str, Any]

Return type

Dataset.set_device_id()

set_device_id(device_id, create_device_if_missing=False)#View Source

Set the device ID for this dataset.

Parameters

device_id Optional[str]

The device ID to set for this dataset. If None, the device association will be cleared.

create_device_if_missing bool

If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.

Returns

This Dataset instance with refreshed data.

Raises

A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization this dataset is being created in.

Dataset.set_summary()

set_summary(summary)#View Source

Explicitly set the AI summary text for this dataset.

This method is intended to be used in cases where an action or other active component is able to generate a more specialized summary than Dataset::generate_summary would, and you want to make that summary canonical from the perspective of the UI and Dataset::get_summary.

Parameters

summary str

The summary text to set for this dataset. This text will be rendered as Markdown, and can include

roboto specialized

// entity links for rich UI linking.

Returns

This Dataset instance for method chaining.

Properties

Dataset.tags

tags list[str] #

List of tags associated with this dataset.

Returns a copy of the list of string tags that have been applied to this dataset for categorization and filtering purposes.

Return type: list[str]

Dataset.to_association()

to_association()#View Source

Dataset.to_dict()

to_dict()#View Source

Convert this dataset to a dictionary representation.

Returns the dataset’s data as a JSON-serializable dictionary containing all dataset attributes and metadata.

Returns

dict[str, Any]

Dictionary representation of the dataset data.

Usage

dataset = Dataset.from_id("ds_abc123")
dataset_dict = dataset.to_dict()
print(dataset_dict["name"])
# 'Highway Test Session'
print(dataset_dict["metadata"])
# {'vehicle_id': 'vehicle_001', 'test_type': 'highway'}

Dataset.update()

update(description=NotSet, device_id=NotSet, metadata_changeset=NotSet, name=NotSet, create_device_if_missing=False, custom_fields_changeset=None)#View Source

Update this dataset’s properties.

Updates various properties of the dataset including name, description, and metadata. Only specified parameters are updated; others remain unchanged.

Parameters

description Optional[Union[str, roboto.sentinels.NotSetType]]

New description for the dataset. Set to None to clear the description.

device_id Optional[Union[str, roboto.sentinels.NotSetType]]

New device ID for the dataset. Set to None to clear the device association.

Metadata changes to apply (add, update, or remove fields/tags).

name Optional[Union[str, roboto.sentinels.NotSetType]]

New name for the dataset. Set to None to clear the name.

create_device_if_missing bool

If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.

custom_fields_changeset Optional[roboto.updates.CustomFieldChangeset]

Changes to apply to Ready custom-field values on this dataset. Field names not referenced by the changeset are left unchanged.

Returns

Updated Dataset instance with the new properties.

Raises

A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.

Caller lacks permission to update the dataset.

Usage

dataset = Dataset.from_id("ds_abc123")
updated_dataset = dataset.update(
    name="Updated Highway Test Session", description="Updated description with more details"
)
print(updated_dataset.name)
# 'Updated Highway Test Session'
# Update with metadata changes
from roboto.updates import MetadataChangeset
changeset = MetadataChangeset(put_fields={"processed": True})
updated_dataset = dataset.update(metadata_changeset=changeset)
# Clear the device association
updated_dataset = dataset.update(device_id=None)
# Clear the description
updated_dataset = dataset.update(description=None)

Dataset.upload_directory()

upload_directory(directory_path, include_patterns=None, exclude_patterns=None, delete_after_upload=False, max_batch_size=MAX_FILES_PER_MANIFEST, print_progress=True, device_id=None)#View Source

Uploads all files and directories recursively from the specified directory path. You can use include_patterns and exclude_patterns to control what files and directories are uploaded, and can use delete_after_upload to clean up your local filesystem after the uploads succeed.

Usage

from roboto import Dataset
dataset = Dataset(...)
dataset.upload_directory(
    pathlib.Path("/path/to/directory"),
    exclude_patterns=[
        "__pycache__/",
        "*.pyc",
        "node_modules/",
        "**/*.log",
    ],
)

Notes

  • Both include_patterns and exclude_patterns follow the ‘gitignore’ pattern format described in https://git-scm.com/docs/gitignore#_pattern_format.
  • If both include_patterns and exclude_patterns are provided, files matching exclude_patterns will be excluded even if they match include_patterns.

Parameters

directory_path pathlib.Path
include_patterns Optional[list[str]]
exclude_patterns Optional[list[str]]
delete_after_upload bool
max_batch_size int
print_progress bool
device_id Optional[str]

Return type

None

Dataset.upload_file()

upload_file(file_path, file_destination_path=None, print_progress=True, device_id=None)#View Source

Upload a single file to the dataset. If file_destination_path is not provided, the file will be uploaded to the top-level of the dataset.

Parameters

file_path pathlib.Path

Local file to upload.

file_destination_path Optional[str]

Destination path within the dataset. Defaults to the file’s own name at the dataset’s top level.

print_progress bool

Whether to display an upload progress bar.

device_id Optional[str]

Optional identifier of the device that generated this data.

Returns

The file record the upload created.

Raises

The upload reported success without reporting a file ID.

Usage

from roboto.domain import datasets
dataset = datasets.Dataset(...)
dataset.upload_file(
    pathlib.Path("/path/to/file.txt"),
    file_destination_path="foo/bar.txt",
)

Dataset.upload_files()

upload_files(files, file_destination_paths={}, max_batch_size=MAX_FILES_PER_MANIFEST, print_progress=True, device_id=None)#View Source

Upload multiple files to the dataset.

If file_destination_paths is not provided, files will be uploaded to the top-level of the dataset.

Parameters

files collections.abc.Iterable[pathlib.Path]

Local files to upload.

file_destination_paths collections.abc.Mapping[pathlib.Path, str]

Mapping from local path to destination path within the dataset. Files not in the mapping upload to the dataset’s top level under their own name.

max_batch_size int

Maximum number of files per upload transaction.

print_progress bool

Whether to display an upload progress bar.

device_id Optional[str]

Optional identifier of the device that generated this data.

Returns

dict[pathlib.Path, str]

Mapping from each uploaded local path to the ID of the file record it created.

Usage

import pathlib
from roboto.domain import datasets
dataset = datasets.Dataset.from_id("ds_abc123")
file_ids = dataset.upload_files(
    [pathlib.Path("/path/to/file.txt")],
    file_destination_paths={
        pathlib.Path("/path/to/file.txt"): "foo/bar.txt",
    },
)
file_ids[pathlib.Path("/path/to/file.txt")]
# 'fl_0123456789abcdef'

DatasetRecord

class roboto.domain.datasets.DatasetRecord(/, **data)#View Source

Bases: pydantic.BaseModel

Wire-transmissible representation of a dataset in the Roboto platform.

DatasetRecord contains all the metadata and properties associated with a dataset, including its identification, timestamps, metadata, tags, and organizational information. This is the data structure used for API communication and persistence.

DatasetRecord instances are typically created by the platform during dataset creation operations and are updated as datasets are modified. The Dataset domain class wraps DatasetRecord to provide a more convenient interface for dataset operations.

The record includes audit information (created/modified timestamps and users), organizational context, and user-defined metadata and tags for discovery and organization purposes.

Parameters

data Any

Attributes

DatasetRecord.administrator

administrator str = 'Roboto' #

Deprecated field maintained for backwards compatibility. Always defaults to ‘Roboto’.

DatasetRecord.created

created datetime.datetime #

Timestamp when this dataset was created in the Roboto platform.

DatasetRecord.created_by

created_by str #

User ID or service account that created this dataset.

DatasetRecord.custom_fields

custom_fields dict[str, Any] = None #

Values for the custom fields defined on Datasets in this org.

Every Ready custom field defined for (org_id, Dataset) appears as a key — values that have not been set surface as None rather than being absent. Empty when no custom fields are defined for the org.

DatasetRecord.dataset_id

dataset_id str #

Unique identifier for this dataset within the Roboto platform.

DatasetRecord.description

description str | None = None #

Human-readable description of the dataset’s contents and purpose.

DatasetRecord.device_id

device_id str | None = None #

Optional identifier of the device that generated this dataset’s data.

DatasetRecord.metadata

metadata dict[str, Any] = None #

User-defined key-value pairs for storing additional dataset information.

DatasetRecord.modified

modified datetime.datetime #

Timestamp when this dataset was last modified.

DatasetRecord.modified_by

modified_by str #

User ID or service account that last modified this dataset.

DatasetRecord.name

name str | None = None #

A short name for this dataset. This may be an org-specific unique ID that’s more meaningful than the dataset_id, or a short summary of the dataset’s contents. If provided, must be 120 characters or less.

DatasetRecord.org_id

org_id str #

Organization ID that owns this dataset.

DatasetRecord.roboto_record_version

roboto_record_version int = 0 #

Internal version number for this record, automatically incremented on updates.

DatasetRecord.storage_ctx

storage_ctx dict[str, Any] = None #

Deprecated storage context field maintained for backwards compatibility with SDK versions prior to 0.10.0.

DatasetRecord.storage_location

storage_location str = 'S3' #

Deprecated storage location field maintained for backwards compatibility. Always defaults to ‘S3’.

DatasetRecord.tags

tags list[str] = None #

List of tags for categorizing and discovering this dataset.

DeleteDirectoriesRequest

class roboto.domain.datasets.DeleteDirectoriesRequest(/, **data)#View Source

Bases: pydantic.BaseModel

Request payload for deleting directories within a dataset.

Used to remove entire directory structures and all contained files from a dataset. This is a bulk operation that affects multiple files.

Parameters

data Any

Attributes

DeleteDirectoriesRequest.directory_paths

directory_paths list[str] #

List of directory paths to delete from the dataset.

QueryDatasetFilesRequest

class roboto.domain.datasets.QueryDatasetFilesRequest(/, **data)#View Source

Bases: pydantic.BaseModel

Request payload for listing the files associated with a dataset, an org, or a device.

Supports gitignore-style patterns for flexible file selection and pagination. Despite the name, the same body lists the files of any association type.

Parameters

data Any

Attributes

QueryDatasetFilesRequest.exclude_patterns

exclude_patterns list[str] | None = None #

List of gitignore-style patterns for files to exclude from results.

QueryDatasetFilesRequest.include_patterns

include_patterns list[str] | None = None #

List of gitignore-style patterns for files to include in results.

QueryDatasetFilesRequest.limit

limit int | None = None #

Maximum number of files to return per page.

QueryDatasetFilesRequest.page_token

page_token str | None = None #

Token for retrieving the next page of results in paginated queries.

QueryDatasetFilesRequest.sort_by

sort_by str | None = None #

Field to sort results by. Defaults to ‘created’.

QueryDatasetFilesRequest.sort_direction

sort_direction str | None = None #

Sort direction (‘ASC’ or ‘DESC’). Defaults to ‘DESC’.

QueryDatasetsRequest

class roboto.domain.datasets.QueryDatasetsRequest(/, **data)#View Source

Bases: pydantic.BaseModel

Request payload for querying datasets with filters.

Used to search for datasets based on various criteria such as metadata, tags, and other dataset properties. The filters are applied server-side to efficiently return matching datasets.

Parameters

data Any

Attributes

QueryDatasetsRequest.filters

filters dict[str, Any] = None #

Dictionary of filter criteria to apply when searching for datasets.

QueryDatasetsRequest.model_config

model_config #

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

RenameDirectoryRequest

class roboto.domain.datasets.RenameDirectoryRequest(/, **data)#View Source

Bases: pydantic.BaseModel

Request payload for renaming a directory among the files of one association.

Changes the path of a directory and all its contained files. This updates the logical organization without moving actual file content.

Parameters

data Any

Attributes

RenameDirectoryRequest.new_path

new_path str #

New path for the directory.

RenameDirectoryRequest.old_path

old_path str #

Current path of the directory to rename.

UpdateDatasetRequest

class roboto.domain.datasets.UpdateDatasetRequest(/, **data)#View Source

Bases: pydantic.BaseModel

Request payload for updating dataset properties.

Used to modify dataset metadata, description, name, and other properties. Only specified fields will be updated; others remain unchanged.

Parameters

data Any

Attributes

UpdateDatasetRequest.custom_fields_changeset

custom_fields_changeset roboto.updates.CustomFieldChangeset | None = None #

Changes to apply to Ready custom-field values on this dataset.

Each referenced field name must be a Ready custom field for this dataset’s org and the Dataset entity type; each set_fields value must satisfy the field’s declared type. Names that are undefined or not Ready are rejected with a structured error. Field names not mentioned by the changeset are left unchanged.

UpdateDatasetRequest.description

description str | roboto.sentinels.NotSetType | None #

New description for the dataset. Set to None to clear the description.

UpdateDatasetRequest.device_id

device_id str | roboto.sentinels.NotSetType | None #

New device ID for the dataset. Set to None to clear the device association.

UpdateDatasetRequest.metadata_changeset

Metadata changes to apply (add, update, or remove fields/tags).

UpdateDatasetRequest.model_config

model_config #

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

UpdateDatasetRequest.name

name Annotated[str, pydantic.StringConstraints(max_length=120)] | roboto.sentinels.NotSetType | None #

New name for the dataset (max 120 characters). Set to None to clear the name.

Was this page helpful?