roboto.domain.datasets
Submodules
Package Contents
CreateDatasetIfNotExistsRequest
Bases: pydantic.BaseModel
Request payload to create a dataset if no existing dataset matches the specified query.
Searches for existing datasets using the provided RoboQL query. If a matching dataset is found, returns that dataset. If no match is found, creates a new dataset with the specified properties and returns it.
Parameters
data AnyAttributes
CreateDatasetIfNotExistsRequest.create_request
CreateDatasetIfNotExistsRequest.match_roboql_query
CreateDatasetRequest
Bases: pydantic.BaseModel
Request payload for creating a new dataset.
Used to specify the initial properties of a dataset during creation, including optional metadata, tags, name, and description.
Attributes
CreateDatasetRequest.custom_fields
Initial values for Ready custom fields on this dataset.
Each key must be the name of a CustomField that is Ready for the caller’s org and the Dataset entity type; each value must satisfy the field’s declared type. Names that are undefined or not Ready, and values that don’t match the field’s type, are rejected with a structured error.
CreateDatasetRequest.description
Optional human-readable description of the dataset.
CreateDatasetRequest.device_id
Optional identifier of the device that generated this data.
CreateDatasetRequest.metadata
Key-value metadata pairs to associate with the dataset for discovery and search.
CreateDatasetRequest.name
Optional short name for the dataset (max 120 characters).
CreateDatasetRequest.tags
List of tags for dataset discovery and organization.
CreateDirectoryRequest
Bases: pydantic.BaseModel
Request payload to create a directory among the files of one association.
Parameters
data AnyAttributes
CreateDirectoryRequest.create_intermediate_dirs
If True, creates intermediate directories in the path if they don’t exist. If False, requires all parent directories to already exist.
CreateDirectoryRequest.error_if_exists
CreateDirectoryRequest.name
CreateDirectoryRequest.origination
CreateDirectoryRequest.parent_path
Dataset
Represents a dataset within the Roboto platform.
A dataset is a logical container for files organized in a directory structure. Datasets are the primary organizational unit in Roboto, typically containing files from a single robot activity such as a drone flight, autonomous vehicle mission, or sensor data collection session. However, datasets are versatile enough to serve as a general-purpose assembly of files.
Datasets provide functionality for:
- File upload and download operations
- Metadata and tag management
- File organization and directory operations
- Topic data access and analysis
- AI-powered content summarization
- Integration with automated workflows and triggers
Files within a dataset can be processed by actions, visualized in the web interface, and searched using the query system. Datasets inherit access permissions from their organization and can be shared with other users and systems.
The Dataset class serves as the primary interface for dataset operations in the Roboto SDK, providing methods for file management, metadata operations, and content analysis.
Parameters
roboto_client Optional[roboto.file_service Optional[roboto.content_mode Optional[roboto.Dataset.clear_custom_field()
Clear a single custom-field value on this dataset to None.
Parameters
name strReturn type
Dataset.clear_custom_fields()
Clear multiple custom-field values on this dataset to None.
Parameters
names collections.Return type
Dataset.create()
Create a new dataset in the Roboto platform.
Creates a new dataset with the specified properties and returns a Dataset instance for interacting with it. The dataset will be created in the caller’s organization unless a different organization is specified.
Parameters
description Optional[str]Optional human-readable description of the dataset.
metadata Optional[dict[str, Any]]Optional key-value metadata pairs to associate with the dataset.
name Optional[str]Optional short name for the dataset (max 120 characters).
tags Optional[list[str]]Optional list of tags for dataset discovery and organization.
device_id Optional[str]Optional identifier of the device that generated this data.
custom_fields Optional[dict[str, Any]]Optional initial values for Ready custom fields defined on Datasets in the caller’s org. Keys must match Ready field names; values must satisfy each field’s declared type.
caller_org_id Optional[str]Organization ID to create the dataset in. Required for multi-org users.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
create_device_if_missing boolIf True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.
Returns
Dataset instance representing the newly created dataset.
Raises
A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.
Invalid dataset parameters.
Caller lacks permission to create datasets.
Usage
dataset = Dataset.create(
name="Highway Test Session",
description="Autonomous vehicle highway driving test data",
tags=["highway", "autonomous", "test"],
metadata={"vehicle_id": "vehicle_001", "test_type": "highway"},
)
print(dataset.dataset_id)
# ds_abc123# Create minimal dataset
dataset = Dataset.create()
print(f"Created dataset: {dataset.dataset_id}")Dataset.create_directory()
Create a directory within the dataset.
Parameters
name strName of the directory to create.
error_if_exists boolIf True, raises an exception if the directory already exists.
parent_path Optional[pathlib.Path of the parent directory. If None, creates the directory in the root of the dataset.
origination Optional[str]Optional string describing the source or context of the directory creation.
create_intermediate_dirs boolIf True, creates intermediate directories in the path if they don’t exist. If False, requires all parent directories to already exist.
Raises
If the directory already exists and error_if_exists is True.
If the caller lacks permission to create the directory.
If the directory name is invalid or the parent path does not exist (when create_intermediate_dirs is False).
Returns
DirectoryRecord of the created directory.
Usage
Create a simple directory:
from roboto.domain import datasets
dataset = datasets.Dataset.from_id(...)
directory = dataset.create_directory("foo")
print(directory.relative_path)
# fooCreate a directory with intermediate directories:
directory = dataset.create_directory(
name="final",
parent_path=pathlib.Path("path/to/deep"),
create_intermediate_dirs=True,
)
print(directory.relative_path)
# path/to/deep/finalDataset.create_if_not_exists()
Create a dataset if no existing dataset matches the specified query.
Searches for existing datasets using the provided RoboQL query. If a matching dataset is found, returns that dataset. If no match is found, creates a new dataset with the specified properties and returns it.
Concurrent calls with the same match_roboql_query in one organization create one dataset between them: the service runs them one at a time from the search through the create. Calls with different queries are not serialized, even when both queries would match the same dataset.
The dataset created must match match_roboql_query. If name, tags or metadata describe a dataset the query does not match, every later call creates another one. When several datasets match, which one is returned is not defined unless the query ends with a SORT BY clause.
Parameters
match_roboql_query strRoboQL query string to search for existing datasets. If this query matches any dataset, that dataset will be returned instead of creating a new one.
description Optional[str]Optional human-readable description of the dataset.
metadata Optional[dict[str, Any]]Optional key-value metadata pairs to associate with the dataset.
name Optional[str]Optional short name for the dataset (max 120 characters).
tags Optional[list[str]]Optional list of tags for dataset discovery and organization.
device_id Optional[str]Optional identifier of the device that generated this data.
custom_fields Optional[dict[str, Any]]Optional initial values for Ready custom fields defined on Datasets in caller_org_id. Keys must match Ready field names; values must satisfy each field’s declared type. Ignored when an existing dataset matches match_roboql_query — the existing record is returned unchanged.
caller_org_id Optional[str]Organization ID to create the dataset in. Required for multi-org users.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
create_device_if_missing boolIf True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.
Returns
Dataset instance representing either the existing matched dataset or the newly created dataset.
Raises
A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.
Invalid dataset parameters or malformed RoboQL query.
Other calls with the same query kept this one waiting for more than 10 seconds, after the SDK’s own retries. Calling again is safe.
Caller lacks permission to create datasets or search existing ones.
Usage
Create a dataset only if no dataset with specific metadata exists:
dataset = Dataset.create_if_not_exists(
match_roboql_query="dataset.metadata.vehicle_id = 'vehicle_001'",
name="Vehicle 001 Test Session",
description="Test data for vehicle 001",
metadata={"vehicle_id": "vehicle_001", "test_type": "highway"},
tags=["vehicle_001", "highway"],
)
print(dataset.dataset_id)
# ds_abc123Create a dataset only if no dataset with specific tags exists:
dataset = Dataset.create_if_not_exists(
match_roboql_query="dataset.tags CONTAINS 'unique_session_id_xyz'",
name="Unique Test Session",
tags=["unique_session_id_xyz", "test"],
)
# If a dataset with tag 'unique_session_id_xyz' already exists,
# that dataset is returned instead of creating a new oneDataset.create_session()
Create a Session populated with files from this dataset.
By default, every file in the dataset is added to the new Session. include_patterns / exclude_patterns narrow that set using the same gitignore-style syntax as Dataset.list_files().
Parameters
name Optional[str]Short display name for the Session (max 120 characters).
device_ids Optional[collections.Devices to attach to the Session. Defaults to no devices.
include_patterns Optional[list[str]]Gitignore-style patterns selecting which files to include. Same syntax as Dataset.list_files().
exclude_patterns Optional[list[str]]Gitignore-style patterns selecting which files to exclude. Takes precedence over include_patterns.
description Optional[str]Optional description of the Session.
metadata Optional[dict[str, Any]]Optional initial metadata. Sessions are not filterable or sortable by metadata keys; for queryable structured attributes, define a custom field on the Session entity type.
tags Optional[collections.Optional initial tags. Sessions can be filtered by tag membership but are not sortable by tag.
custom_fields Optional[dict[str, Any]]Optional initial values for Ready custom fields defined on Sessions in this dataset’s org. Keys must match Ready field names; values must satisfy each field’s declared type.
Returns
The newly created Session.
Usage
dataset = Dataset.from_id("ds_abc123")
session = dataset.create_session("flight-2026-04-23-001")Notes
Convenience wrapper around Session.create() followed by Session.add_files(). A dataset may hold more files than one add request accepts: the files are sent in consecutive requests of at most MAX_FILES_AND_TOPICS_PER_REQUEST each. Every file goes in or none does: if any add request fails, or the platform refuses any one file, the partially populated Session is deleted before the exception propagates, so the call is safe to retry. If that cleanup fails too (e.g. a transient network error), the Session is left behind holding whatever files did go in; the cleanup failure is logged, and the original exception is what the caller sees.
Properties
Dataset.created
Timestamp when this dataset was created.
Returns the UTC datetime when this dataset was first created in the Roboto platform. This property is immutable.
Dataset.created_by
Identifier of the user who created this dataset.
Returns the identifier of the person or service which originally created this dataset in the Roboto platform.
Dataset.custom_fields
Custom-field values defined on Datasets in this org.
Every Ready CustomField defined for (org_id, Dataset) appears as a key. Values that have not been set on this dataset surface as None rather than being absent. Empty when no custom fields are defined for the org.
A Timestamp value is returned as an ISO 8601 string.
Dataset.dataset_id
Unique identifier for this dataset.
Returns the globally unique identifier assigned to this dataset when it was created. This ID is immutable and used to reference the dataset across the Roboto platform. It is always prefixed with ‘ds_’ to distinguish it from other Roboto resource IDs.
Dataset.delete()
Delete this dataset from the Roboto platform.
Permanently removes the dataset and all its associated files, metadata, and topics. This operation cannot be undone.
If a dataset’s files are hosted in Roboto managed S3 buckets or customer read/write bring-your-own-buckets, the files in this dataset will be deleted from S3 as well. For files hosted in customer read-only buckets, the files will not be deleted from S3, but the dataset record and all associated metadata will be deleted.
Raises
Dataset does not exist or has already been deleted.
Caller lacks permission to delete the dataset.
Return type
Usage
dataset = Dataset.from_id("ds_abc123")
dataset.delete()
# # Dataset and all its files are now permanently deletedDataset.delete_files()
Delete files from this dataset based on pattern matching.
Deletes files that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.
Parameters
include_patterns Optional[list[str]]List of gitignore-style patterns for files to include. If None or empty, all files are considered for deletion. An empty list is treated as no filter (all files), not as “include nothing”.
exclude_patterns Optional[list[str]]List of gitignore-style patterns for files to exclude from deletion. Takes precedence over include patterns. If None or empty, no files are excluded.
Raises
Caller lacks permission to delete files.
Return type
Notes
Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.
Usage
dataset = Dataset.from_id("ds_abc123")
# Delete all PNG files except those in back_camera directory
dataset.delete_files(include_patterns=["**/*.png"], exclude_patterns=["**/back_camera/**"])# Delete all log files
dataset.delete_files(include_patterns=["**/*.log"])Properties
Dataset.description
Human-readable description of this dataset.
Returns the optional description text that provides details about the dataset’s contents, purpose, or context. Can be None if no description was provided.
Dataset.device_id
Identifier of the device that generated this data.
Returns the optional identifier of the device that generated the data contained within this dataset. Can be None if the dataset was not generated by a device.
Dataset.download_files()
Download files from this dataset to a local directory.
Downloads files that match the specified patterns to the given local directory. The directory structure from the dataset is preserved in the download location. If the output directory doesn’t exist, it will be created.
Parameters
out_path pathlib.Local directory path where files should be downloaded.
include_patterns Optional[list[str]]List of gitignore-style patterns for files to include. If None or empty, all files are downloaded. An empty list is treated as no filter (all files), not as “include nothing”.
exclude_patterns Optional[list[str]]List of gitignore-style patterns for files to exclude from download. Takes precedence over include patterns. If None or empty, no files are excluded.
print_progress boolWhether to show a progress bar during download.
Returns
List of tuples containing (FileRecord, local_path) for each downloaded file.
Raises
Caller lacks permission to download files.
Notes
Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.
Usage
import pathlib
dataset = Dataset.from_id("ds_abc123")
downloaded = dataset.download_files(
pathlib.Path("/tmp/dataset_download"),
include_patterns=["**/*.bag"],
exclude_patterns=["**/test/**"],
)
print(f"Downloaded {len(downloaded)} files")
# Downloaded 5 files# Download all files
all_files = dataset.download_files(pathlib.Path("/tmp/all_files"))Properties
Dataset.files
The files associated with this dataset.
The file methods on Dataset delegate here, so dataset.upload_files(...) and dataset.files.upload_files(...) are the same call.
Dataset.from_id()
Create a Dataset instance from a dataset ID.
Retrieves dataset information from the Roboto platform using the provided dataset ID and returns a Dataset instance for interacting with it.
Parameters
dataset_id strUnique identifier for the dataset.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
Returns
Dataset instance representing the requested dataset.
Raises
Dataset with the given ID does not exist.
Caller lacks permission to access the dataset.
Usage
dataset = Dataset.from_id("ds_abc123")
print(dataset.name)
# 'Highway Test Session'
print(len(list(dataset.list_files())))
# 42Dataset.generate_summary()
Generate a new AI summary for this dataset.
Creates a new AI-generated summary that analyzes the dataset’s content, structure, and metadata. The summary generation is asynchronous and can be monitored through the returned StreamingAISummary object.
Returns
StreamingAISummary object that provides access to the summary as it is being generated. The summary starts in pending status and can be monitored for completion.
Raises
Caller lacks permission to generate summaries for this dataset.
Usage
Generate a summary and wait for completion:
dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
complete_text = summary.complete_text
print(complete_text)
# 'This dataset contains 42 files with sensor data from highway driving tests...'Generate a summary and stream the text as it’s generated:
dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
for text_chunk in summary.text_stream():
print(text_chunk, end="", flush=True)Check summary status without blocking:
dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
if summary.current and summary.current.status == AISummaryStatus.Complete:
print("Summary is ready!")Dataset.get_file_by_path()
Get a File instance for a file at the specified path in this dataset.
Retrieves a file by its relative path within the dataset. Optionally retrieves a specific version of the file.
Parameters
relative_path Union[str, pathlib.Path of the file relative to the dataset root.
version_id Optional[int]Specific version of the file to retrieve. If None, gets the latest version.
Returns
File instance representing the file at the specified path.
Raises
File at the given path does not exist in the dataset.
Caller lacks permission to access the file.
Usage
dataset = Dataset.from_id("ds_abc123")
file = dataset.get_file_by_path("logs/session1.bag")
print(file.file_id)
# file_xyz789# Get specific version
old_file = dataset.get_file_by_path("data/sensors.csv", version_id=1)
print(old_file.version)
# 1Dataset.get_metadata()
Return custom metadata associated with this dataset.
Returns a copy of the dataset’s metadata dictionary containing arbitrary key-value pairs for storing custom information. Supports nested structures and dot notation for accessing nested fields.
Return type
Dataset.get_sessions()
Iterate over Sessions that include at least one file from this dataset.
A Session may draw files from one or more datasets; this method yields every Session whose current-version file list intersects this dataset.
Yields
Each matching Session. Pagination is handled automatically.
Return type
Usage
dataset = Dataset.from_id("ds_abc123")
for session in dataset.get_sessions():
print(session.session_id, session.name)Dataset.get_summary()
Retrieve this dataset’s existing AI summary.
Returns the dataset’s current AI summary if one exists. Reading never generates a summary as a side effect: a dataset that has never been summarized raises RobotoNotFoundException rather than implicitly kicking off — and paying for — generation. Call generate_summary() to create one explicitly.
Returns
StreamingAISummary wrapping the dataset’s existing summary. If a generation kicked off elsewhere is still in flight, the returned summary is Pending; poll it via await_completion or text_stream.
Raises
This dataset has no AI summary yet. Call generate_summary() to create one.
Caller lacks permission to access summaries for this dataset.
Usage
Get the existing summary, generating one first if there is none:
from roboto.exceptions import RobotoNotFoundException
dataset = Dataset.from_id("ds_abc123")
try:
summary = dataset.get_summary()
except RobotoNotFoundException:
summary = dataset.generate_summary()
print(summary.complete_text)
# 'This dataset contains 42 files with sensor data from highway driving tests...'Check whether a summary exists without generating one:
from roboto.exceptions import RobotoNotFoundException
dataset = Dataset.from_id("ds_abc123")
try:
summary = dataset.get_summary()
print(summary.complete_text)
except RobotoNotFoundException:
print("No summary yet — call generate_summary() to create one.")Dataset.get_topic_time_bounds()
Get the earliest start and latest end across every topic in this dataset.
The same aggregate you would reach by folding start_time and end_time over get_topics(), computed server-side in one request instead of one per page of topics. Reach for it when you want the dataset’s time extent and not the topics themselves.
Returns
Bounds in nanoseconds since the Unix epoch. Both fields are None for a dataset whose files hold no topics, and either is None when no topic in the dataset carries that timestamp.
Raises
Dataset does not exist.
Caller lacks permission to access the dataset.
Usage
dataset = Dataset.from_id("ds_abc123")
bounds = dataset.get_topic_time_bounds()
print(bounds.start_time, bounds.end_time)
# 1722870127699468923 1722870187004821001Dataset.get_topics()
Get all topics associated with files in this dataset, with optional filtering.
Retrieves all topics that were extracted from files in this dataset during ingestion. If multiple files have topics with the same name (e.g., chunked files with the same schema), they are returned as separate topic objects.
Topics can be filtered by name using include/exclude patterns. Topics specified on both the inclusion and exclusion lists will be excluded.
Parameters
include Optional[collections.If provided, only topics with names in this sequence are yielded.
exclude Optional[collections.If provided, topics with names in this sequence are skipped. Takes precedence over include list.
Yields
Topic instances associated with files in this dataset, filtered according to the parameters.
Return type
Usage
dataset = Dataset.from_id("ds_abc123")
for topic in dataset.get_topics():
print(f"Topic: {topic.name}")
# Topic: /camera/image
# Topic: /imu/data
# Topic: /gps/fix# Only get camera topics
camera_topics = list(dataset.get_topics(include=["/camera/image", "/camera/info"]))
print(f"Found {len(camera_topics)} camera topics")# Exclude diagnostic topics
data_topics = list(dataset.get_topics(exclude=["/diagnostics"]))Dataset.get_topics_by_file()
Get all topics associated with a specific file in this dataset.
Retrieves all topics that were extracted from the specified file during ingestion. This is a convenience method that combines file lookup and topic retrieval.
Parameters
relative_path Union[str, pathlib.Path of the file relative to the dataset root.
Yields
Topic instances associated with the specified file.
Raises
File at the given path does not exist in the dataset.
Caller lacks permission to access the file or its topics.
Return type
Usage
dataset = Dataset.from_id("ds_abc123")
for topic in dataset.get_topics_by_file("logs/session1.bag"):
print(f"Topic: {topic.name}")
# Topic: /camera/image
# Topic: /imu/data
# Topic: /gps/fixDataset.list_directories()
Yield every directory in this dataset, at any depth.
Usage
dataset = Dataset.from_id("ds_abc123")
for directory in dataset.list_directories():
print(directory.relative_path)
# logs
# logs/session1Return type
Dataset.list_files()
List files in this dataset with optional pattern-based filtering.
Returns all files in the dataset that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.
Parameters
include_patterns Optional[list[str]]List of gitignore-style patterns for files to include. If None or empty, all files are considered. An empty list is treated as no filter (all files), not as “include nothing”.
exclude_patterns Optional[list[str]]List of gitignore-style patterns for files to exclude. Takes precedence over include patterns. If None or empty, no files are excluded.
Yields
File instances that match the specified patterns.
Raises
Caller lacks permission to list files.
Return type
Notes
Pattern matching follows gitignore syntax. See https://git-scm.com/docs/gitignore for detailed pattern format documentation.
Usage
dataset = Dataset.from_id("ds_abc123")
for file in dataset.list_files():
print(file.relative_path)
# logs/session1.bag
# data/sensors.csv
# images/camera_001.jpg# List only image files, excluding back camera
for file in dataset.list_files(
include_patterns=["**/*.png", "**/*.jpg"], exclude_patterns=["**/back_camera/**"]
):
print(file.relative_path)
# images/front_camera_001.jpg
# images/side_camera_001.jpgProperties
Dataset.metadata
Custom metadata associated with this dataset.
Returns a copy of the dataset’s metadata dictionary containing arbitrary key-value pairs for storing custom information. Supports nested structures and dot notation for accessing nested fields.
Note: this attribute is kept for backward compatibility. Prefer get_metadata(), since metadata may need to be loaded on-demand from the server.
Dataset.modified
Timestamp when this dataset was last modified.
Returns the UTC datetime when this dataset was most recently updated. This includes changes to metadata, tags, description, or other properties.
Dataset.modified_by
Identifier of the user or service which last modified this dataset.
Returns the identifier of the person or service which most recently updated this dataset’s metadata, tags, description, or other properties.
Dataset.name
Human-readable name of this dataset.
Returns the optional display name for this dataset. Can be None if no name was provided during creation. For users whose organizations have their own idiomatic internal dataset IDs, it’s recommended to set the name to the organization’s internal dataset ID, since the Roboto dataset_id property is randomly generated.
Dataset.org_id
Organization identifier that owns this dataset.
Returns the unique identifier of the organization that owns and has primary access control over this dataset.
Dataset.put_metadata()
Add or update metadata fields for this dataset.
Sets each key-value pair in the provided dictionary as dataset metadata. If a key doesn’t exist, it will be created. If it exists, the value will be overwritten. Keys must be strings and dot notation is supported for nested keys.
Parameters
metadata dict[str, Any]Dictionary of metadata key-value pairs to add or update.
Raises
Caller lacks permission to update the dataset.
Return type
Usage
dataset = Dataset.from_id("ds_abc123")
dataset.put_metadata(
{
"vehicle_id": "vehicle_001",
"test_type": "highway_driving",
"weather.condition": "sunny",
"weather.temperature": 25,
}
)
print(dataset.metadata["vehicle_id"])
# 'vehicle_001'
print(dataset.metadata["weather"]["condition"])
# 'sunny'Dataset.put_tags()
Add or update tags for this dataset.
Adds each tag in the provided sequence to the dataset. If a tag already exists, it will not be duplicated. This operation replaces the current tag list with the provided tags.
Parameters
Sequence of tag strings to set on the dataset.
Raises
Caller lacks permission to update the dataset.
Return type
Usage
dataset = Dataset.from_id("ds_abc123")
dataset.put_tags(["highway", "autonomous", "test", "sunny"])
print(dataset.tags)
# ['highway', 'autonomous', 'test', 'sunny']Dataset.query()
Query datasets using a specification with filters and pagination.
Searches for datasets matching the provided query specification. Results are returned as a generator that automatically handles pagination, yielding Dataset instances as they are retrieved from the API.
Parameters
spec Optional[roboto.Query specification with filters, sorting, and pagination options. If None, returns all accessible datasets.
roboto_client Optional[roboto.HTTP client for API communication. If None, uses the default client.
owner_org_id Optional[str]Organization ID to scope the query. If None, uses caller’s org.
Yields
Dataset instances matching the query specification.
Raises
ValueErrorQuery specification references unknown dataset attributes.
Caller lacks permission to query datasets.
Return type
Usage
from roboto.query import Comparator, Condition, QuerySpecification
spec = QuerySpecification(
condition=Condition(field="name", comparator=Comparator.Contains, value="Roboto")
)
for dataset in Dataset.query(spec):
print(f"Found dataset: {dataset.name}")
# Found dataset: Roboto Test
# Found dataset: Other Roboto TestProperties
Dataset.record
Underlying data record for this dataset.
Returns the raw DatasetRecord that contains all the dataset’s data fields. This provides access to the complete dataset state as stored in the platform.
Dataset.refresh()
Refresh this dataset instance with the latest data from the platform.
Fetches the current state of the dataset from the Roboto platform and updates this instance’s data. Useful when the dataset may have been modified by other processes or users.
Returns
This Dataset instance with refreshed data.
Raises
Dataset no longer exists.
Caller lacks permission to access the dataset.
Usage
dataset = Dataset.from_id("ds_abc123")
# Dataset may have been updated by another process
refreshed_dataset = dataset.refresh()
print(f"Current file count: {len(list(refreshed_dataset.list_files()))}")Dataset.remove_metadata()
Remove each key in this sequence from dataset metadata if it exists. Keys must be strings. Dot notation is supported for nested keys.
Usage
from roboto.domain import datasets
dataset = datasets.Dataset(...)
dataset.remove_metadata(["foo", "baz.qux"])Parameters
metadata roboto.Return type
Dataset.remove_tags()
Remove each tag in this sequence if it exists
Parameters
Return type
Dataset.rename_directory()
Rename or move a directory within this dataset.
Both old_path and new_path are relative to the dataset root. Pass a new_path with fewer path components to move the directory up the tree, or a different leaf name at the same depth to rename in place.
Parameters
old_path strCurrent relative path of the directory (e.g. "logs/session1").
new_path strTarget relative path of the directory (e.g. "session1" to move up one level).
Returns
Updated DirectoryRecord reflecting the new path.
Raises
No directory exists at old_path.
new_path conflicts with an existing node or contains a cycle.
Usage
dataset = Dataset.from_id("ds_abc123")
dataset.rename_directory("logs/session1", "session1")Dataset.rename_file()
Rename or move a file within this dataset.
new_path is relative to the dataset root. Pass a path with fewer components to move the file up the tree, a different name at the same depth to rename in place, or a path under a different directory to move sideways.
The file’s storage URI is unchanged; only the logical location in the dataset hierarchy moves.
Parameters
file_id strID of the file to rename or move.
new_path strTarget relative path for the file within this dataset (e.g. "file.bag" to move to the root, or "other_dir/file.bag" to move into an existing directory).
Returns
Updated FileRecord reflecting the new path.
Raises
No file with file_id exists.
new_path conflicts with an existing file, the parent directory does not exist, or the move would create a cycle.
Usage
dataset = Dataset.from_id("ds_abc123")
record = dataset.rename_file("file_xyz789", "file.bag")
record.relative_path
# 'file.bag'Dataset.set_custom_field()
Dataset.set_custom_fields()
Dataset.set_device_id()
Set the device ID for this dataset.
Parameters
device_id Optional[str]The device ID to set for this dataset. If None, the device association will be cleared.
create_device_if_missing boolIf True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.
Returns
This Dataset instance with refreshed data.
Raises
A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization this dataset is being created in.
Dataset.set_summary()
Explicitly set the AI summary text for this dataset.
This method is intended to be used in cases where an action or other active component is able to generate a more specialized summary than Dataset::generate_summary would, and you want to make that summary canonical from the perspective of the UI and Dataset::get_summary.
Parameters
summary strThe summary text to set for this dataset. This text will be rendered as Markdown, and can include
roboto specialized// entity links for rich UI linking.
Returns
This Dataset instance for method chaining.
Properties
Dataset.tags
List of tags associated with this dataset.
Returns a copy of the list of string tags that have been applied to this dataset for categorization and filtering purposes.
Dataset.to_association()
Return type
Dataset.to_dict()
Convert this dataset to a dictionary representation.
Returns the dataset’s data as a JSON-serializable dictionary containing all dataset attributes and metadata.
Returns
Dictionary representation of the dataset data.
Usage
dataset = Dataset.from_id("ds_abc123")
dataset_dict = dataset.to_dict()
print(dataset_dict["name"])
# 'Highway Test Session'
print(dataset_dict["metadata"])
# {'vehicle_id': 'vehicle_001', 'test_type': 'highway'}Dataset.update()
Update this dataset’s properties.
Updates various properties of the dataset including name, description, and metadata. Only specified parameters are updated; others remain unchanged.
Parameters
description Optional[Union[str, roboto.New description for the dataset. Set to None to clear the description.
device_id Optional[Union[str, roboto.New device ID for the dataset. Set to None to clear the device association.
metadata_changeset Union[roboto.Metadata changes to apply (add, update, or remove fields/tags).
name Optional[Union[str, roboto.New name for the dataset. Set to None to clear the name.
create_device_if_missing boolIf True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.
custom_fields_changeset Optional[roboto.Changes to apply to Ready custom-field values on this dataset. Field names not referenced by the changeset are left unchanged.
Returns
Updated Dataset instance with the new properties.
Raises
A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.
Caller lacks permission to update the dataset.
Usage
dataset = Dataset.from_id("ds_abc123")
updated_dataset = dataset.update(
name="Updated Highway Test Session", description="Updated description with more details"
)
print(updated_dataset.name)
# 'Updated Highway Test Session'# Update with metadata changes
from roboto.updates import MetadataChangeset
changeset = MetadataChangeset(put_fields={"processed": True})
updated_dataset = dataset.update(metadata_changeset=changeset)# Clear the device association
updated_dataset = dataset.update(device_id=None)# Clear the description
updated_dataset = dataset.update(description=None)Dataset.upload_directory()
Uploads all files and directories recursively from the specified directory path. You can use include_patterns and exclude_patterns to control what files and directories are uploaded, and can use delete_after_upload to clean up your local filesystem after the uploads succeed.
Usage
from roboto import Dataset
dataset = Dataset(...)
dataset.upload_directory(
pathlib.Path("/path/to/directory"),
exclude_patterns=[
"__pycache__/",
"*.pyc",
"node_modules/",
"**/*.log",
],
)Notes
- Both include_patterns and exclude_patterns follow the ‘gitignore’ pattern format described in https://git-scm.com/docs/gitignore#_pattern_format.
- If both include_patterns and exclude_patterns are provided, files matching exclude_patterns will be excluded even if they match include_patterns.
Parameters
directory_path pathlib.include_patterns Optional[list[str]]exclude_patterns Optional[list[str]]delete_after_upload boolmax_batch_size intprint_progress booldevice_id Optional[str]Return type
Dataset.upload_file()
Upload a single file to the dataset. If file_destination_path is not provided, the file will be uploaded to the top-level of the dataset.
Parameters
file_path pathlib.Local file to upload.
file_destination_path Optional[str]Destination path within the dataset. Defaults to the file’s own name at the dataset’s top level.
print_progress boolWhether to display an upload progress bar.
device_id Optional[str]Optional identifier of the device that generated this data.
Returns
The file record the upload created.
Raises
The upload reported success without reporting a file ID.
Usage
from roboto.domain import datasets
dataset = datasets.Dataset(...)
dataset.upload_file(
pathlib.Path("/path/to/file.txt"),
file_destination_path="foo/bar.txt",
)Dataset.upload_files()
Upload multiple files to the dataset.
If file_destination_paths is not provided, files will be uploaded to the top-level of the dataset.
Parameters
files collections.Local files to upload.
file_destination_paths collections.Mapping from local path to destination path within the dataset. Files not in the mapping upload to the dataset’s top level under their own name.
max_batch_size intMaximum number of files per upload transaction.
print_progress boolWhether to display an upload progress bar.
device_id Optional[str]Optional identifier of the device that generated this data.
Returns
Mapping from each uploaded local path to the ID of the file record it created.
Usage
import pathlib
from roboto.domain import datasets
dataset = datasets.Dataset.from_id("ds_abc123")
file_ids = dataset.upload_files(
[pathlib.Path("/path/to/file.txt")],
file_destination_paths={
pathlib.Path("/path/to/file.txt"): "foo/bar.txt",
},
)
file_ids[pathlib.Path("/path/to/file.txt")]
# 'fl_0123456789abcdef'DatasetRecord
Bases: pydantic.BaseModel
Wire-transmissible representation of a dataset in the Roboto platform.
DatasetRecord contains all the metadata and properties associated with a dataset, including its identification, timestamps, metadata, tags, and organizational information. This is the data structure used for API communication and persistence.
DatasetRecord instances are typically created by the platform during dataset creation operations and are updated as datasets are modified. The Dataset domain class wraps DatasetRecord to provide a more convenient interface for dataset operations.
The record includes audit information (created/modified timestamps and users), organizational context, and user-defined metadata and tags for discovery and organization purposes.
Parameters
data AnyAttributes
DatasetRecord.administrator
Deprecated field maintained for backwards compatibility. Always defaults to ‘Roboto’.
DatasetRecord.created
Timestamp when this dataset was created in the Roboto platform.
DatasetRecord.custom_fields
Values for the custom fields defined on Datasets in this org.
Every Ready custom field defined for (org_id, Dataset) appears as a key — values that have not been set surface as None rather than being absent. Empty when no custom fields are defined for the org.
DatasetRecord.dataset_id
Unique identifier for this dataset within the Roboto platform.
DatasetRecord.description
Human-readable description of the dataset’s contents and purpose.
DatasetRecord.device_id
Optional identifier of the device that generated this dataset’s data.
DatasetRecord.metadata
User-defined key-value pairs for storing additional dataset information.
DatasetRecord.modified_by
User ID or service account that last modified this dataset.
DatasetRecord.name
A short name for this dataset. This may be an org-specific unique ID that’s more meaningful than the dataset_id, or a short summary of the dataset’s contents. If provided, must be 120 characters or less.
DatasetRecord.roboto_record_version
Internal version number for this record, automatically incremented on updates.
DatasetRecord.storage_ctx
Deprecated storage context field maintained for backwards compatibility with SDK versions prior to 0.10.0.
DatasetRecord.storage_location
Deprecated storage location field maintained for backwards compatibility. Always defaults to ‘S3’.
DatasetRecord.tags
List of tags for categorizing and discovering this dataset.
DeleteDirectoriesRequest
Bases: pydantic.BaseModel
Request payload for deleting directories within a dataset.
Used to remove entire directory structures and all contained files from a dataset. This is a bulk operation that affects multiple files.
Parameters
data AnyAttributes
DeleteDirectoriesRequest.directory_paths
List of directory paths to delete from the dataset.
QueryDatasetFilesRequest
Bases: pydantic.BaseModel
Request payload for listing the files associated with a dataset, an org, or a device.
Supports gitignore-style patterns for flexible file selection and pagination. Despite the name, the same body lists the files of any association type.
Parameters
data AnyAttributes
QueryDatasetFilesRequest.exclude_patterns
List of gitignore-style patterns for files to exclude from results.
QueryDatasetFilesRequest.include_patterns
List of gitignore-style patterns for files to include in results.
QueryDatasetFilesRequest.page_token
Token for retrieving the next page of results in paginated queries.
QueryDatasetFilesRequest.sort_by
Field to sort results by. Defaults to ‘created’.
QueryDatasetFilesRequest.sort_direction
Sort direction (‘ASC’ or ‘DESC’). Defaults to ‘DESC’.
QueryDatasetsRequest
Bases: pydantic.BaseModel
Request payload for querying datasets with filters.
Used to search for datasets based on various criteria such as metadata, tags, and other dataset properties. The filters are applied server-side to efficiently return matching datasets.
Parameters
data AnyRenameDirectoryRequest
Bases: pydantic.BaseModel
Request payload for renaming a directory among the files of one association.
Changes the path of a directory and all its contained files. This updates the logical organization without moving actual file content.
Parameters
data AnyUpdateDatasetRequest
Bases: pydantic.BaseModel
Request payload for updating dataset properties.
Used to modify dataset metadata, description, name, and other properties. Only specified fields will be updated; others remain unchanged.
Parameters
data AnyAttributes
UpdateDatasetRequest.custom_fields_changeset
Changes to apply to Ready custom-field values on this dataset.
Each referenced field name must be a Ready custom field for this dataset’s org and the Dataset entity type; each set_fields value must satisfy the field’s declared type. Names that are undefined or not Ready are rejected with a structured error. Field names not mentioned by the changeset are left unchanged.
UpdateDatasetRequest.description
New description for the dataset. Set to None to clear the description.
UpdateDatasetRequest.device_id
New device ID for the dataset. Set to None to clear the device association.
UpdateDatasetRequest.metadata_changeset
Metadata changes to apply (add, update, or remove fields/tags).
UpdateDatasetRequest.model_config
model_config #Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
UpdateDatasetRequest.name
name Annotated[str, pydantic.New name for the dataset (max 120 characters). Set to None to clear the name.