---
sidebar:
  hidden: true
title: roboto.domain.datasets.dataset
---
## Module Contents

### Dataset

```python
class roboto.domain.datasets.dataset.Dataset(
    record: roboto.domain.datasets.record.DatasetRecord,
    roboto_client: Optional[roboto.http.RobotoClient] = None,
    file_service: Optional[roboto.storage.FileService] = None,
    content_mode: Optional[roboto.query.QueryContentMode] = None,
)
```

`from roboto import Dataset`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L61-L1601)

Represents a dataset within the Roboto platform.

A dataset is a logical container for files organized in a directory structure. Datasets are the primary organizational unit in Roboto, typically containing files from a single robot activity such as a drone flight, autonomous vehicle mission, or sensor data collection session. However, datasets are versatile enough to serve as a general-purpose assembly of files.

Datasets provide functionality for:

- File upload and download operations
- Metadata and tag management
- File organization and directory operations
- Topic data access and analysis
- AI-powered content summarization
- Integration with automated workflows and triggers

Files within a dataset can be processed by actions, visualized in the web interface, and searched using the query system. Datasets inherit access permissions from their organization and can be shared with other users and systems.

The Dataset class serves as the primary interface for dataset operations in the Roboto SDK, providing methods for file management, metadata operations, and content analysis.

**Parameters**

- **record** (`roboto.domain.datasets.record.DatasetRecord`)
- **roboto_client** (`Optional[roboto.http.RobotoClient]`)
- **file_service** (`Optional[roboto.storage.FileService]`)
- **content_mode** (`Optional[roboto.query.QueryContentMode]`)

**Properties**

- **Dataset.created** (`datetime.datetime`): Timestamp when this dataset was created.

  Returns the UTC datetime when this dataset was first created in the Roboto platform. This property is immutable.

- **Dataset.created_by** (`str`): Identifier of the user who created this dataset.

  Returns the identifier of the person or service which originally created this dataset in the Roboto platform.

- **Dataset.custom_fields** (`dict[str, Any]`): Custom-field values defined on Datasets in this org.

  Every `Ready` [`CustomField`](/reference/python-sdk/roboto/domain/custom_fields/custom_field#roboto.domain.custom_fields.custom_field.CustomField) defined for `(org_id, Dataset)` appears as a key. Values that have not been set on this dataset surface as `None` rather than being absent. Empty when no custom fields are defined for the org.

  A [`Timestamp`](/reference/python-sdk/roboto/domain/custom_fields/record#roboto.domain.custom_fields.record.CustomFieldType.Timestamp) value is returned as an ISO 8601 string.

- **Dataset.dataset_id** (`str`): Unique identifier for this dataset.

  Returns the globally unique identifier assigned to this dataset when it was created. This ID is immutable and used to reference the dataset across the Roboto platform. It is always prefixed with 'ds\_' to distinguish it from other Roboto resource IDs.

- **Dataset.description** (`str | None`): Human-readable description of this dataset.

  Returns the optional description text that provides details about the dataset's contents, purpose, or context. Can be None if no description was provided.

  Return type: `Optional[str]`

- **Dataset.device_id** (`str | None`): Identifier of the device that generated this data.

  Returns the optional identifier of the device that generated the data contained within this dataset. Can be None if the dataset was not generated by a device.

  Return type: `Optional[str]`

- **Dataset.files** (`roboto.domain.files.FileSystem`): The files associated with this dataset.

  The file methods on `Dataset` delegate here, so `dataset.upload_files(...)` and `dataset.files.upload_files(...)` are the same call.

- **Dataset.metadata** (`dict[str, Any]`): Custom metadata associated with this dataset.

  Returns a copy of the dataset's metadata dictionary containing arbitrary key-value pairs for storing custom information. Supports nested structures and dot notation for accessing nested fields.

  Note: this attribute is kept for backward compatibility. Prefer [`get_metadata()`](/reference/python-sdk/roboto/domain/datasets/dataset#roboto.domain.datasets.dataset.Dataset.get_metadata), since metadata may need to be loaded on-demand from the server.

- **Dataset.modified** (`datetime.datetime`): Timestamp when this dataset was last modified.

  Returns the UTC datetime when this dataset was most recently updated. This includes changes to metadata, tags, description, or other properties.

- **Dataset.modified_by** (`str`): Identifier of the user or service which last modified this dataset.

  Returns the identifier of the person or service which most recently updated this dataset's metadata, tags, description, or other properties.

- **Dataset.name** (`str | None`): Human-readable name of this dataset.

  Returns the optional display name for this dataset. Can be None if no name was provided during creation. For users whose organizations have their own idiomatic internal dataset IDs, it's recommended to set the name to the organization's internal dataset ID, since the Roboto dataset_id property is randomly generated.

  Return type: `Optional[str]`

- **Dataset.org_id** (`str`): Organization identifier that owns this dataset.

  Returns the unique identifier of the organization that owns and has primary access control over this dataset.

- **Dataset.record** (`roboto.domain.datasets.record.DatasetRecord`): Underlying data record for this dataset.

  Returns the raw [`DatasetRecord`](/reference/python-sdk/roboto/domain/datasets/record#roboto.domain.datasets.record.DatasetRecord) that contains all the dataset's data fields. This provides access to the complete dataset state as stored in the platform.

- **Dataset.tags** (`list[str]`): List of tags associated with this dataset.

  Returns a copy of the list of string tags that have been applied to this dataset for categorization and filtering purposes.

#### Dataset.clear_custom_field()

```python
def clear_custom_field(name: str) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1259-L1261)

Clear a single custom-field value on this dataset to `None`.

**Parameters**

- **name** (`str`)

**Returns**

- `Dataset`

#### Dataset.clear_custom_fields()

```python
def clear_custom_fields(names: collections.abc.Sequence[str]) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1274-L1276)

Clear multiple custom-field values on this dataset to `None`.

**Parameters**

- **names** (`collections.abc.Sequence[str]`)

**Returns**

- `Dataset`

#### Dataset.create()

```python
@classmethod
def create(
    description: Optional[str] = None,
    metadata: Optional[dict[str, Any]] = None,
    name: Optional[str] = None,
    tags: Optional[list[str]] = None,
    device_id: Optional[str] = None,
    custom_fields: Optional[dict[str, Any]] = None,
    caller_org_id: Optional[str] = None,
    roboto_client: Optional[roboto.http.RobotoClient] = None,
    create_device_if_missing: bool = False,
) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L94-L178)

Create a new dataset in the Roboto platform.

Creates a new dataset with the specified properties and returns a Dataset instance for interacting with it. The dataset will be created in the caller's organization unless a different organization is specified.

**Parameters**

- **description** (`Optional[str]`): Optional human-readable description of the dataset.
- **metadata** (`Optional[dict[str, Any]]`): Optional key-value metadata pairs to associate with the dataset.
- **name** (`Optional[str]`): Optional short name for the dataset (max 120 characters).
- **tags** (`Optional[list[str]]`): Optional list of tags for dataset discovery and organization.
- **device_id** (`Optional[str]`): Optional identifier of the device that generated this data.
- **custom_fields** (`Optional[dict[str, Any]]`): Optional initial values for Ready custom fields defined on Datasets in the caller's org. Keys must match Ready field names; values must satisfy each field's declared type.
- **caller_org_id** (`Optional[str]`): Organization ID to create the dataset in. Required for multi-org users.
- **roboto_client** (`Optional[roboto.http.RobotoClient]`): HTTP client for API communication. If None, uses the default client.
- **create_device_if_missing** (`bool`): If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.

**Returns**

- `Dataset`: Dataset instance representing the newly created dataset.

**Raises**

- [`RobotoDeviceNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoDeviceNotFoundException): A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.
- [`RobotoInvalidRequestException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoInvalidRequestException): Invalid dataset parameters.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to create datasets.

**Usage**

```python
dataset = Dataset.create(
    name="Highway Test Session",
    description="Autonomous vehicle highway driving test data",
    tags=["highway", "autonomous", "test"],
    metadata={"vehicle_id": "vehicle_001", "test_type": "highway"},
)
print(dataset.dataset_id)
# ds_abc123
```

```python
# Create minimal dataset
dataset = Dataset.create()
print(f"Created dataset: {dataset.dataset_id}")
```

#### Dataset.create_directory()

```python
def create_directory(
    name: str,
    error_if_exists: bool = False,
    create_intermediate_dirs: bool = False,
    parent_path: Optional[pathlib.Path] = None,
    origination: Optional[str] = None,
) -> roboto.domain.files.DirectoryRecord
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L582-L635)

Create a directory within the dataset.

**Parameters**

- **name** (`str`): Name of the directory to create.
- **error_if_exists** (`bool`): If True, raises an exception if the directory already exists.
- **parent_path** (`Optional[pathlib.Path]`): Path of the parent directory. If None, creates the directory in the root of the dataset.
- **origination** (`Optional[str]`): Optional string describing the source or context of the directory creation.
- **create_intermediate_dirs** (`bool`): If True, creates intermediate directories in the path if they don't exist. If False, requires all parent directories to already exist.

**Raises**

- [`RobotoConflictException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoConflictException): If the directory already exists and error_if_exists is True.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): If the caller lacks permission to create the directory.
- [`RobotoInvalidRequestException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoInvalidRequestException): If the directory name is invalid or the parent path does not exist (when create_intermediate_dirs is False).

**Returns**

- `roboto.domain.files.DirectoryRecord`: DirectoryRecord of the created directory.

**Usage**

Create a simple directory:

```python
from roboto.domain import datasets
dataset = datasets.Dataset.from_id(...)
directory = dataset.create_directory("foo")
print(directory.relative_path)
# foo
```

Create a directory with intermediate directories:

```python
directory = dataset.create_directory(
    name="final",
    parent_path=pathlib.Path("path/to/deep"),
    create_intermediate_dirs=True,
)
print(directory.relative_path)
# path/to/deep/final
```

#### Dataset.create_if_not_exists()

```python
@classmethod
def create_if_not_exists(
    match_roboql_query: str,
    description: Optional[str] = None,
    metadata: Optional[dict[str, Any]] = None,
    name: Optional[str] = None,
    tags: Optional[list[str]] = None,
    device_id: Optional[str] = None,
    custom_fields: Optional[dict[str, Any]] = None,
    caller_org_id: Optional[str] = None,
    roboto_client: Optional[roboto.http.RobotoClient] = None,
    create_device_if_missing: bool = False,
) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L181-L298)

Create a dataset if no existing dataset matches the specified query.

Searches for existing datasets using the provided RoboQL query. If a matching dataset is found, returns that dataset. If no match is found, creates a new dataset with the specified properties and returns it.

Concurrent calls with the same `match_roboql_query` in one organization create one dataset between them: the service runs them one at a time from the search through the create. Calls with different queries are not serialized, even when both queries would match the same dataset.

The dataset created must match `match_roboql_query`. If `name`, `tags` or `metadata` describe a dataset the query does not match, every later call creates another one. When several datasets match, which one is returned is not defined unless the query ends with a `SORT BY` clause.

**Parameters**

- **match_roboql_query** (`str`): RoboQL query string to search for existing datasets. If this query matches any dataset, that dataset will be returned instead of creating a new one.
- **description** (`Optional[str]`): Optional human-readable description of the dataset.
- **metadata** (`Optional[dict[str, Any]]`): Optional key-value metadata pairs to associate with the dataset.
- **name** (`Optional[str]`): Optional short name for the dataset (max 120 characters).
- **tags** (`Optional[list[str]]`): Optional list of tags for dataset discovery and organization.
- **device_id** (`Optional[str]`): Optional identifier of the device that generated this data.
- **custom_fields** (`Optional[dict[str, Any]]`): Optional initial values for `Ready` custom fields defined on Datasets in `caller_org_id`. Keys must match Ready field names; values must satisfy each field's declared type. Ignored when an existing dataset matches `match_roboql_query` — the existing record is returned unchanged.
- **caller_org_id** (`Optional[str]`): Organization ID to create the dataset in. Required for multi-org users.
- **roboto_client** (`Optional[roboto.http.RobotoClient]`): HTTP client for API communication. If None, uses the default client.
- **create_device_if_missing** (`bool`): If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.

**Returns**

- `Dataset`: Dataset instance representing either the existing matched dataset or the newly created dataset.

**Raises**

- [`RobotoDeviceNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoDeviceNotFoundException): A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.
- [`RobotoInvalidRequestException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoInvalidRequestException): Invalid dataset parameters or malformed RoboQL query.
- [`RobotoServiceUnavailableException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoServiceUnavailableException): Other calls with the same query kept this one waiting for more than 10 seconds, after the SDK's own retries. Calling again is safe.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to create datasets or search existing ones.

**Usage**

Create a dataset only if no dataset with specific metadata exists:

```python
dataset = Dataset.create_if_not_exists(
    match_roboql_query="dataset.metadata.vehicle_id = 'vehicle_001'",
    name="Vehicle 001 Test Session",
    description="Test data for vehicle 001",
    metadata={"vehicle_id": "vehicle_001", "test_type": "highway"},
    tags=["vehicle_001", "highway"],
)
print(dataset.dataset_id)
# ds_abc123
```

Create a dataset only if no dataset with specific tags exists:

```python
dataset = Dataset.create_if_not_exists(
    match_roboql_query="dataset.tags CONTAINS 'unique_session_id_xyz'",
    name="Unique Test Session",
    tags=["unique_session_id_xyz", "test"],
)
# If a dataset with tag 'unique_session_id_xyz' already exists,
# that dataset is returned instead of creating a new one
```

#### Dataset.create_session()

```python
def create_session(
    name: Optional[str] = None,
    *,
    device_ids: Optional[collections.abc.Sequence[str]] = None,
    include_patterns: Optional[list[str]] = None,
    exclude_patterns: Optional[list[str]] = None,
    description: Optional[str] = None,
    metadata: Optional[dict[str, Any]] = None,
    tags: Optional[collections.abc.Sequence[str]] = None,
    custom_fields: Optional[dict[str, Any]] = None,
) -> roboto.experimental.sessions.Session
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L638-L722)

Create a Session populated with files from this dataset.

By default, every file in the dataset is added to the new Session. `include_patterns` / `exclude_patterns` narrow that set using the same gitignore-style syntax as [`Dataset.list_files()`](/reference/python-sdk/roboto/domain/datasets/dataset#roboto.domain.datasets.dataset.Dataset.list_files).

**Parameters**

- **name** (`Optional[str]`): Short display name for the Session (max 120 characters).
- **device_ids** (`Optional[collections.abc.Sequence[str]]`): Devices to attach to the Session. Defaults to no devices.
- **include_patterns** (`Optional[list[str]]`): Gitignore-style patterns selecting which files to include. Same syntax as [`Dataset.list_files()`](/reference/python-sdk/roboto/domain/datasets/dataset#roboto.domain.datasets.dataset.Dataset.list_files).
- **exclude_patterns** (`Optional[list[str]]`): Gitignore-style patterns selecting which files to exclude. Takes precedence over `include_patterns`.
- **description** (`Optional[str]`): Optional description of the Session.
- **metadata** (`Optional[dict[str, Any]]`): Optional initial metadata. Sessions are not filterable or sortable by `metadata` keys; for queryable structured attributes, define a custom field on the `Session` entity type.
- **tags** (`Optional[collections.abc.Sequence[str]]`): Optional initial tags. Sessions can be filtered by tag membership but are not sortable by tag.
- **custom_fields** (`Optional[dict[str, Any]]`): Optional initial values for Ready custom fields defined on Sessions in this dataset's org. Keys must match Ready field names; values must satisfy each field's declared type.

**Returns**

- `roboto.experimental.sessions.Session`: The newly created Session.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
session = dataset.create_session("flight-2026-04-23-001")
```

**Notes**

Convenience wrapper around [`Session.create()`](/reference/python-sdk/roboto/experimental/sessions/session#roboto.experimental.sessions.session.Session.create) followed by [`Session.add_files()`](/reference/python-sdk/roboto/experimental/sessions/session#roboto.experimental.sessions.session.Session.add_files). A dataset may hold more files than one add request accepts: the files are sent in consecutive requests of at most [`MAX_FILES_AND_TOPICS_PER_REQUEST`](/reference/python-sdk/roboto/experimental/ingest/operations#roboto.experimental.ingest.operations.MAX_FILES_AND_TOPICS_PER_REQUEST) each. Every file goes in or none does: if any add request fails, or the platform refuses any one file, the partially populated Session is deleted before the exception propagates, so the call is safe to retry. If that cleanup fails too (e.g. a transient network error), the Session is left behind holding whatever files did go in; the cleanup failure is logged, and the original exception is what the caller sees.

#### Dataset.delete()

```python
def delete() -> None
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L724-L743)

Delete this dataset from the Roboto platform.

Permanently removes the dataset and all its associated files, metadata, and topics. This operation cannot be undone.

If a dataset's files are hosted in Roboto managed S3 buckets or customer read/write bring-your-own-buckets, the files in this dataset will be deleted from S3 as well. For files hosted in customer read-only buckets, the files will not be deleted from S3, but the dataset record and all associated metadata will be deleted.

**Raises**

- [`RobotoDatasetNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoDatasetNotFoundException): Dataset does not exist or has already been deleted.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to delete the dataset.

**Returns**

- `None`

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
dataset.delete()
# # Dataset and all its files are now permanently deleted
```

#### Dataset.delete_files()

```python
def delete_files(
    include_patterns: Optional[list[str]] = None,
    exclude_patterns: Optional[list[str]] = None,
) -> None
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L745-L779)

Delete files from this dataset based on pattern matching.

Deletes files that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.

**Parameters**

- **include_patterns** (`Optional[list[str]]`): List of gitignore-style patterns for files to include. If None or empty, all files are considered for deletion. An empty list is treated as no filter (all files), not as "include nothing".
- **exclude_patterns** (`Optional[list[str]]`): List of gitignore-style patterns for files to exclude from deletion. Takes precedence over include patterns. If None or empty, no files are excluded.

**Raises**

- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to delete files.

**Returns**

- `None`

**Notes**

Pattern matching follows gitignore syntax. See [https://git-scm.com/docs/gitignore](https://git-scm.com/docs/gitignore) for detailed pattern format documentation.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
# Delete all PNG files except those in back_camera directory
dataset.delete_files(include_patterns=["**/*.png"], exclude_patterns=["**/back_camera/**"])
```

```python
# Delete all log files
dataset.delete_files(include_patterns=["**/*.log"])
```

#### Dataset.download_files()

```python
def download_files(
    out_path: pathlib.Path,
    include_patterns: Optional[list[str]] = None,
    exclude_patterns: Optional[list[str]] = None,
    print_progress: bool = True,
) -> list[tuple[roboto.domain.files.FileRecord, pathlib.Path]]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L781-L828)

Download files from this dataset to a local directory.

Downloads files that match the specified patterns to the given local directory. The directory structure from the dataset is preserved in the download location. If the output directory doesn't exist, it will be created.

**Parameters**

- **out_path** (`pathlib.Path`): Local directory path where files should be downloaded.
- **include_patterns** (`Optional[list[str]]`): List of gitignore-style patterns for files to include. If None or empty, all files are downloaded. An empty list is treated as no filter (all files), not as "include nothing".
- **exclude_patterns** (`Optional[list[str]]`): List of gitignore-style patterns for files to exclude from download. Takes precedence over include patterns. If None or empty, no files are excluded.
- **print_progress** (`bool`): Whether to show a progress bar during download.

**Returns**

- `list[tuple[roboto.domain.files.FileRecord, pathlib.Path]]`: List of tuples containing (FileRecord, local_path) for each downloaded file.

**Raises**

- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to download files.

**Notes**

Pattern matching follows gitignore syntax. See [https://git-scm.com/docs/gitignore](https://git-scm.com/docs/gitignore) for detailed pattern format documentation.

**Usage**

```python
import pathlib
dataset = Dataset.from_id("ds_abc123")
downloaded = dataset.download_files(
    pathlib.Path("/tmp/dataset_download"),
    include_patterns=["**/*.bag"],
    exclude_patterns=["**/test/**"],
)
print(f"Downloaded {len(downloaded)} files")
# Downloaded 5 files
```

```python
# Download all files
all_files = dataset.download_files(pathlib.Path("/tmp/all_files"))
```

#### Dataset.from_id()

```python
@classmethod
def from_id(
    dataset_id: str,
    roboto_client: Optional[roboto.http.RobotoClient] = None,
) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L301-L327)

Create a Dataset instance from a dataset ID.

Retrieves dataset information from the Roboto platform using the provided dataset ID and returns a Dataset instance for interacting with it.

**Parameters**

- **dataset_id** (`str`): Unique identifier for the dataset.
- **roboto_client** (`Optional[roboto.http.RobotoClient]`): HTTP client for API communication. If None, uses the default client.

**Returns**

- `Dataset`: Dataset instance representing the requested dataset.

**Raises**

- [`RobotoDatasetNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoDatasetNotFoundException): Dataset with the given ID does not exist.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to access the dataset.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
print(dataset.name)
# 'Highway Test Session'
print(len(list(dataset.list_files())))
# 42
```

#### Dataset.generate_summary()

```python
def generate_summary() -> roboto.ai.summary.StreamingAISummary
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L864-L909)

Generate a new AI summary for this dataset.

Creates a new AI-generated summary that analyzes the dataset's content, structure, and metadata. The summary generation is asynchronous and can be monitored through the returned StreamingAISummary object.

**Returns**

- `roboto.ai.summary.StreamingAISummary`: StreamingAISummary object that provides access to the summary as it is being generated. The summary starts in pending status and can be monitored for completion.

**Raises**

- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to generate summaries for this dataset.

**Usage**

Generate a summary and wait for completion:

```python
dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
complete_text = summary.complete_text
print(complete_text)
# 'This dataset contains 42 files with sensor data from highway driving tests...'
```

Generate a summary and stream the text as it's generated:

```python
dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
for text_chunk in summary.text_stream():
    print(text_chunk, end="", flush=True)
```

Check summary status without blocking:

```python
dataset = Dataset.from_id("ds_abc123")
summary = dataset.generate_summary()
if summary.current and summary.current.status == AISummaryStatus.Complete:
    print("Summary is ready!")
```

#### Dataset.get_file_by_path()

```python
def get_file_by_path(
    relative_path: Union[str, pathlib.Path],
    version_id: Optional[int] = None,
) -> roboto.domain.files.File
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L830-L862)

Get a File instance for a file at the specified path in this dataset.

Retrieves a file by its relative path within the dataset. Optionally retrieves a specific version of the file.

**Parameters**

- **relative_path** (`Union[str, pathlib.Path]`): Path of the file relative to the dataset root.
- **version_id** (`Optional[int]`): Specific version of the file to retrieve. If None, gets the latest version.

**Returns**

- `roboto.domain.files.File`: File instance representing the file at the specified path.

**Raises**

- [`RobotoNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoNotFoundException): File at the given path does not exist in the dataset.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to access the file.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
file = dataset.get_file_by_path("logs/session1.bag")
print(file.file_id)
# file_xyz789
```

```python
# Get specific version
old_file = dataset.get_file_by_path("data/sensors.csv", version_id=1)
print(old_file.version)
# 1
```

#### Dataset.get_metadata()

```python
def get_metadata() -> dict[str, Any]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L511-L524)

Return custom metadata associated with this dataset.

Returns a copy of the dataset's metadata dictionary containing arbitrary key-value pairs for storing custom information. Supports nested structures and dot notation for accessing nested fields.

**Returns**

- `dict[str, Any]`

#### Dataset.get_sessions()

```python
def get_sessions() -> collections.abc.Generator[roboto.experimental.sessions.Session, None, None]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L957-L972)

Iterate over Sessions that include at least one file from this dataset.

A Session may draw files from one or more datasets; this method yields every Session whose current-version file list intersects this dataset.

**Yields**

- Each matching Session. Pagination is handled automatically.

**Returns**

- `collections.abc.Generator[roboto.experimental.sessions.Session, None, None]`

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
for session in dataset.get_sessions():
    print(session.session_id, session.name)
```

#### Dataset.get_summary()

```python
def get_summary() -> roboto.ai.summary.StreamingAISummary
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L911-L954)

Retrieve this dataset's existing AI summary.

Returns the dataset's current AI summary if one exists. Reading never generates a summary as a side effect: a dataset that has never been summarized raises [`RobotoNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoNotFoundException) rather than implicitly kicking off — and paying for — generation. Call [`generate_summary()`](/reference/python-sdk/roboto/domain/datasets/dataset#roboto.domain.datasets.dataset.Dataset.generate_summary) to create one explicitly.

**Returns**

- `roboto.ai.summary.StreamingAISummary`: StreamingAISummary wrapping the dataset's existing summary. If a generation kicked off elsewhere is still in flight, the returned summary is `Pending`; poll it via `await_completion` or `text_stream`.

**Raises**

- [`RobotoNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoNotFoundException): This dataset has no AI summary yet. Call [`generate_summary()`](/reference/python-sdk/roboto/domain/datasets/dataset#roboto.domain.datasets.dataset.Dataset.generate_summary) to create one.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to access summaries for this dataset.

**Usage**

Get the existing summary, generating one first if there is none:

```python
from roboto.exceptions import RobotoNotFoundException
dataset = Dataset.from_id("ds_abc123")
try:
    summary = dataset.get_summary()
except RobotoNotFoundException:
    summary = dataset.generate_summary()
print(summary.complete_text)
# 'This dataset contains 42 files with sensor data from highway driving tests...'
```

Check whether a summary exists without generating one:

```python
from roboto.exceptions import RobotoNotFoundException
dataset = Dataset.from_id("ds_abc123")
try:
    summary = dataset.get_summary()
    print(summary.complete_text)
except RobotoNotFoundException:
    print("No summary yet — call generate_summary() to create one.")
```

#### Dataset.get_topic_time_bounds()

```python
def get_topic_time_bounds() -> roboto.domain.topics.TopicTimeBounds
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1026-L1049)

Get the earliest start and latest end across every topic in this dataset.

The same aggregate you would reach by folding `start_time` and `end_time` over [`get_topics()`](/reference/python-sdk/roboto/domain/datasets/dataset#roboto.domain.datasets.dataset.Dataset.get_topics), computed server-side in one request instead of one per page of topics. Reach for it when you want the dataset's time extent and not the topics themselves.

**Returns**

- `roboto.domain.topics.TopicTimeBounds`: Bounds in nanoseconds since the Unix epoch. Both fields are None for a dataset whose files hold no topics, and either is None when no topic in the dataset carries that timestamp.

**Raises**

- [`RobotoNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoNotFoundException): Dataset does not exist.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to access the dataset.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
bounds = dataset.get_topic_time_bounds()
print(bounds.start_time, bounds.end_time)
# 1722870127699468923 1722870187004821001
```

#### Dataset.get_topics()

```python
def get_topics(
    include: Optional[collections.abc.Sequence[str]] = None,
    exclude: Optional[collections.abc.Sequence[str]] = None,
) -> collections.abc.Generator[roboto.domain.topics.Topic, None, None]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L974-L1024)

Get all topics associated with files in this dataset, with optional filtering.

Retrieves all topics that were extracted from files in this dataset during ingestion. If multiple files have topics with the same name (e.g., chunked files with the same schema), they are returned as separate topic objects.

Topics can be filtered by name using include/exclude patterns. Topics specified on both the inclusion and exclusion lists will be excluded.

> **Note**
>
> This method returns topics WITHOUT message_paths populated (for performance). The message_paths list will be empty. If you need message_paths, iterate through files and call file.get_topics() instead.

**Parameters**

- **include** (`Optional[collections.abc.Sequence[str]]`): If provided, only topics with names in this sequence are yielded.
- **exclude** (`Optional[collections.abc.Sequence[str]]`): If provided, topics with names in this sequence are skipped. Takes precedence over include list.

**Yields**

- Topic instances associated with files in this dataset, filtered according to the parameters.

**Returns**

- `collections.abc.Generator[roboto.domain.topics.Topic, None, None]`

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
for topic in dataset.get_topics():
    print(f"Topic: {topic.name}")
# Topic: /camera/image
# Topic: /imu/data
# Topic: /gps/fix
```

```python
# Only get camera topics
camera_topics = list(dataset.get_topics(include=["/camera/image", "/camera/info"]))
print(f"Found {len(camera_topics)} camera topics")
```

```python
# Exclude diagnostic topics
data_topics = list(dataset.get_topics(exclude=["/diagnostics"]))
```

#### Dataset.get_topics_by_file()

```python
def get_topics_by_file(
    relative_path: Union[str, pathlib.Path],
) -> collections.abc.Generator[roboto.domain.topics.Topic, None, None]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1051-L1079)

Get all topics associated with a specific file in this dataset.

Retrieves all topics that were extracted from the specified file during ingestion. This is a convenience method that combines file lookup and topic retrieval.

**Parameters**

- **relative_path** (`Union[str, pathlib.Path]`): Path of the file relative to the dataset root.

**Yields**

- Topic instances associated with the specified file.

**Raises**

- [`RobotoNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoNotFoundException): File at the given path does not exist in the dataset.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to access the file or its topics.

**Returns**

- `collections.abc.Generator[roboto.domain.topics.Topic, None, None]`

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
for topic in dataset.get_topics_by_file("logs/session1.bag"):
    print(f"Topic: {topic.name}")
# Topic: /camera/image
# Topic: /imu/data
# Topic: /gps/fix
```

#### Dataset.list_directories()

```python
def list_directories() -> collections.abc.Generator[roboto.domain.files.DirectoryRecord, None, None]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1081-L1093)

Yield every directory in this dataset, at any depth.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
for directory in dataset.list_directories():
    print(directory.relative_path)
# logs
# logs/session1
```

**Returns**

- `collections.abc.Generator[roboto.domain.files.DirectoryRecord, None, None]`

#### Dataset.list_files()

```python
def list_files(
    include_patterns: Optional[list[str]] = None,
    exclude_patterns: Optional[list[str]] = None,
) -> collections.abc.Generator[roboto.domain.files.File, None, None]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1095-L1139)

List files in this dataset with optional pattern-based filtering.

Returns all files in the dataset that match the specified include patterns while excluding those that match exclude patterns. Uses gitignore-style pattern matching for flexible file selection.

**Parameters**

- **include_patterns** (`Optional[list[str]]`): List of gitignore-style patterns for files to include. If None or empty, all files are considered. An empty list is treated as no filter (all files), not as "include nothing".
- **exclude_patterns** (`Optional[list[str]]`): List of gitignore-style patterns for files to exclude. Takes precedence over include patterns. If None or empty, no files are excluded.

**Yields**

- File instances that match the specified patterns.

**Raises**

- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to list files.

**Returns**

- `collections.abc.Generator[roboto.domain.files.File, None, None]`

**Notes**

Pattern matching follows gitignore syntax. See [https://git-scm.com/docs/gitignore](https://git-scm.com/docs/gitignore) for detailed pattern format documentation.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
for file in dataset.list_files():
    print(file.relative_path)
# logs/session1.bag
# data/sensors.csv
# images/camera_001.jpg
```

```python
# List only image files, excluding back camera
for file in dataset.list_files(
    include_patterns=["**/*.png", "**/*.jpg"], exclude_patterns=["**/back_camera/**"]
):
    print(file.relative_path)
# images/front_camera_001.jpg
# images/side_camera_001.jpg
```

#### Dataset.put_metadata()

```python
def put_metadata(metadata: dict[str, Any]) -> None
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1141-L1173)

Add or update metadata fields for this dataset.

Sets each key-value pair in the provided dictionary as dataset metadata. If a key doesn't exist, it will be created. If it exists, the value will be overwritten. Keys must be strings and dot notation is supported for nested keys.

**Parameters**

- **metadata** (`dict[str, Any]`): Dictionary of metadata key-value pairs to add or update.

**Raises**

- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to update the dataset.

**Returns**

- `None`

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
dataset.put_metadata(
    {
        "vehicle_id": "vehicle_001",
        "test_type": "highway_driving",
        "weather.condition": "sunny",
        "weather.temperature": 25,
    }
)
print(dataset.metadata["vehicle_id"])
# 'vehicle_001'
print(dataset.metadata["weather"]["condition"])
# 'sunny'
```

#### Dataset.put_tags()

```python
def put_tags(tags: roboto.updates.StrSequence) -> None
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1175-L1199)

Add or update tags for this dataset.

Adds each tag in the provided sequence to the dataset. If a tag already exists, it will not be duplicated. This operation replaces the current tag list with the provided tags.

**Parameters**

- **tags** (`roboto.updates.StrSequence`): Sequence of tag strings to set on the dataset.

**Raises**

- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to update the dataset.

**Returns**

- `None`

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
dataset.put_tags(["highway", "autonomous", "test", "sunny"])
print(dataset.tags)
# ['highway', 'autonomous', 'test', 'sunny']
```

#### Dataset.query()

```python
@classmethod
def query(
    spec: Optional[roboto.query.QuerySpecification] = None,
    roboto_client: Optional[roboto.http.RobotoClient] = None,
    owner_org_id: Optional[str] = None,
) -> collections.abc.Generator[Dataset, None, None]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L330-L397)

Query datasets using a specification with filters and pagination.

Searches for datasets matching the provided query specification. Results are returned as a generator that automatically handles pagination, yielding Dataset instances as they are retrieved from the API.

**Parameters**

- **spec** (`Optional[roboto.query.QuerySpecification]`): Query specification with filters, sorting, and pagination options. If None, returns all accessible datasets.
- **roboto_client** (`Optional[roboto.http.RobotoClient]`): HTTP client for API communication. If None, uses the default client.
- **owner_org_id** (`Optional[str]`): Organization ID to scope the query. If None, uses caller's org.

**Yields**

- Dataset instances matching the query specification.

**Raises**

- `ValueError`: Query specification references unknown dataset attributes.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to query datasets.

**Returns**

- `collections.abc.Generator[Dataset, None, None]`

**Usage**

```python
from roboto.query import Comparator, Condition, QuerySpecification
spec = QuerySpecification(
    condition=Condition(field="name", comparator=Comparator.Contains, value="Roboto")
)
for dataset in Dataset.query(spec):
    print(f"Found dataset: {dataset.name}")
# Found dataset: Roboto Test
# Found dataset: Other Roboto Test
```

#### Dataset.refresh()

```python
def refresh() -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1201-L1222)

Refresh this dataset instance with the latest data from the platform.

Fetches the current state of the dataset from the Roboto platform and updates this instance's data. Useful when the dataset may have been modified by other processes or users.

**Returns**

- `Dataset`: This Dataset instance with refreshed data.

**Raises**

- [`RobotoDatasetNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoDatasetNotFoundException): Dataset no longer exists.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to access the dataset.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
# Dataset may have been updated by another process
refreshed_dataset = dataset.refresh()
print(f"Current file count: {len(list(refreshed_dataset.list_files()))}")
```

#### Dataset.remove_metadata()

```python
def remove_metadata(metadata: roboto.updates.StrSequence) -> None
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1224-L1237)

Remove each key in this sequence from dataset metadata if it exists. Keys must be strings. Dot notation is supported for nested keys.

**Usage**

```python
from roboto.domain import datasets
dataset = datasets.Dataset(...)
dataset.remove_metadata(["foo", "baz.qux"])
```

**Parameters**

- **metadata** (`roboto.updates.StrSequence`)

**Returns**

- `None`

#### Dataset.remove_tags()

```python
def remove_tags(tags: roboto.updates.StrSequence) -> None
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1239-L1244)

Remove each tag in this sequence if it exists

**Parameters**

- **tags** (`roboto.updates.StrSequence`)

**Returns**

- `None`

#### Dataset.rename_directory()

```python
def rename_directory(
    old_path: str,
    new_path: str,
) -> roboto.domain.files.DirectoryRecord
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1428-L1450)

Rename or move a directory within this dataset.

Both `old_path` and `new_path` are relative to the dataset root. Pass a `new_path` with fewer path components to move the directory up the tree, or a different leaf name at the same depth to rename in place.

**Parameters**

- **old_path** (`str`): Current relative path of the directory (e.g. `"logs/session1"`).
- **new_path** (`str`): Target relative path of the directory (e.g. `"session1"` to move up one level).

**Returns**

- `roboto.domain.files.DirectoryRecord`: Updated [`DirectoryRecord`](/reference/python-sdk/roboto/domain/files/record#roboto.domain.files.record.DirectoryRecord) reflecting the new path.

**Raises**

- [`RobotoNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoNotFoundException): No directory exists at `old_path`.
- [`RobotoInvalidRequestException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoInvalidRequestException): `new_path` conflicts with an existing node or contains a cycle.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
dataset.rename_directory("logs/session1", "session1")
```

#### Dataset.rename_file()

```python
def rename_file(file_id: str, new_path: str) -> roboto.domain.files.FileRecord
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1452-L1482)

Rename or move a file within this dataset.

`new_path` is relative to the dataset root. Pass a path with fewer components to move the file up the tree, a different name at the same depth to rename in place, or a path under a different directory to move sideways.

The file's storage URI is unchanged; only the logical location in the dataset hierarchy moves.

**Parameters**

- **file_id** (`str`): ID of the file to rename or move.
- **new_path** (`str`): Target relative path for the file within this dataset (e.g. `"file.bag"` to move to the root, or `"other_dir/file.bag"` to move into an existing directory).

**Returns**

- `roboto.domain.files.FileRecord`: Updated [`FileRecord`](/reference/python-sdk/roboto/domain/files/record#roboto.domain.files.record.FileRecord) reflecting the new path.

**Raises**

- [`RobotoNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoNotFoundException): No file with `file_id` exists.
- [`RobotoInvalidRequestException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoInvalidRequestException): `new_path` conflicts with an existing file, the parent directory does not exist, or the move would create a cycle.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
record = dataset.rename_file("file_xyz789", "file.bag")
record.relative_path
# 'file.bag'
```

#### Dataset.set_custom_field()

```python
def set_custom_field(name: str, value: Any) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1247-L1256)

Set a single custom-field value on this dataset.

`name` must be the name of a [`Ready`](/reference/python-sdk/roboto/domain/custom_fields/record#roboto.domain.custom_fields.record.CustomFieldStatus.Ready) custom field for this dataset's org and the [`Dataset`](/reference/python-sdk/roboto/domain/custom_fields/record#roboto.domain.custom_fields.record.TargetEntityType.Dataset) entity type; `value` must satisfy the field's declared type.

**Parameters**

- **name** (`str`)
- **value** (`Any`)

**Returns**

- `Dataset`

#### Dataset.set_custom_fields()

```python
def set_custom_fields(fields: dict[str, Any]) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1264-L1271)

Set or overwrite multiple custom-field values on this dataset.

Each key must name a Ready custom field for this dataset's org and the [`Dataset`](/reference/python-sdk/roboto/domain/custom_fields/record#roboto.domain.custom_fields.record.TargetEntityType.Dataset) entity type; each value must satisfy the field's declared type.

**Parameters**

- **fields** (`dict[str, Any]`)

**Returns**

- `Dataset`

#### Dataset.set_device_id()

```python
def set_device_id(
    device_id: Optional[str],
    create_device_if_missing: bool = False,
) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1278-L1299)

Set the device ID for this dataset.

**Parameters**

- **device_id** (`Optional[str]`): The device ID to set for this dataset. If None, the device association will be cleared.
- **create_device_if_missing** (`bool`): If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.

**Returns**

- `Dataset`: This Dataset instance with refreshed data.

**Raises**

- [`RobotoDeviceNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoDeviceNotFoundException): A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization this dataset is being created in.

#### Dataset.set_summary()

```python
def set_summary(summary: str) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1301-L1319)

Explicitly set the AI summary text for this dataset.

This method is intended to be used in cases where an action or other active component is able to generate a more specialized summary than Dataset::generate_summary would, and you want to make that summary canonical from the perspective of the UI and Dataset::get_summary.

**Parameters**

- **summary** (`str`): The summary text to set for this dataset. This text will be rendered as Markdown, and can include
- **roboto** (`specialized`): // entity links for rich UI linking.

**Returns**

- `Dataset`: This Dataset instance for method chaining.

#### Dataset.to_association()

```python
def to_association() -> roboto.association.Association
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1321-L1322)

**Returns**

- `roboto.association.Association`

#### Dataset.to_dict()

```python
def to_dict() -> dict[str, Any]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1324-L1341)

Convert this dataset to a dictionary representation.

Returns the dataset's data as a JSON-serializable dictionary containing all dataset attributes and metadata.

**Returns**

- `dict[str, Any]`: Dictionary representation of the dataset data.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
dataset_dict = dataset.to_dict()
print(dataset_dict["name"])
# 'Highway Test Session'
print(dataset_dict["metadata"])
# {'vehicle_id': 'vehicle_001', 'test_type': 'highway'}
```

#### Dataset.update()

```python
def update(
    description: Optional[Union[str, roboto.sentinels.NotSetType]] = NotSet,
    device_id: Optional[Union[str, roboto.sentinels.NotSetType]] = NotSet,
    metadata_changeset: Union[roboto.updates.MetadataChangeset, roboto.sentinels.NotSetType] = NotSet,
    name: Optional[Union[str, roboto.sentinels.NotSetType]] = NotSet,
    create_device_if_missing: bool = False,
    custom_fields_changeset: Optional[roboto.updates.CustomFieldChangeset] = None,
) -> Dataset
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1343-L1426)

Update this dataset's properties.

Updates various properties of the dataset including name, description, and metadata. Only specified parameters are updated; others remain unchanged.

**Parameters**

- **description** (`Optional[Union[str, roboto.sentinels.NotSetType]]`): New description for the dataset. Set to None to clear the description.
- **device_id** (`Optional[Union[str, roboto.sentinels.NotSetType]]`): New device ID for the dataset. Set to None to clear the device association.
- **metadata_changeset** (`Union[roboto.updates.MetadataChangeset, roboto.sentinels.NotSetType]`): Metadata changes to apply (add, update, or remove fields/tags).
- **name** (`Optional[Union[str, roboto.sentinels.NotSetType]]`): New name for the dataset. Set to None to clear the name.
- **create_device_if_missing** (`bool`): If True, and a device_id is provided that does not exist in the organization, a new device will be created automatically. If False, and a device_id is provided that does not exist, a RobotoDeviceNotFoundException will be raised.
- **custom_fields_changeset** (`Optional[roboto.updates.CustomFieldChangeset]`): Changes to apply to Ready custom-field values on this dataset. Field names not referenced by the changeset are left unchanged.

**Returns**

- `Dataset`: Updated Dataset instance with the new properties.

**Raises**

- [`RobotoDeviceNotFoundException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoDeviceNotFoundException): A device_id has been provided in this request, but was not found as a device registered with Roboto for the organization.
- [`RobotoUnauthorizedException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoUnauthorizedException): Caller lacks permission to update the dataset.

**Usage**

```python
dataset = Dataset.from_id("ds_abc123")
updated_dataset = dataset.update(
    name="Updated Highway Test Session", description="Updated description with more details"
)
print(updated_dataset.name)
# 'Updated Highway Test Session'
```

```python
# Update with metadata changes
from roboto.updates import MetadataChangeset
changeset = MetadataChangeset(put_fields={"processed": True})
updated_dataset = dataset.update(metadata_changeset=changeset)
```

```python
# Clear the device association
updated_dataset = dataset.update(device_id=None)
```

```python
# Clear the description
updated_dataset = dataset.update(description=None)
```

#### Dataset.upload_directory()

```python
def upload_directory(
    directory_path: pathlib.Path,
    include_patterns: Optional[list[str]] = None,
    exclude_patterns: Optional[list[str]] = None,
    delete_after_upload: bool = False,
    max_batch_size: int = MAX_FILES_PER_MANIFEST,
    print_progress: bool = True,
    device_id: Optional[str] = None,
) -> None
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1484-L1526)

Uploads all files and directories recursively from the specified directory path. You can use include_patterns and exclude_patterns to control what files and directories are uploaded, and can use delete_after_upload to clean up your local filesystem after the uploads succeed.

**Usage**

```python
from roboto import Dataset
dataset = Dataset(...)
dataset.upload_directory(
    pathlib.Path("/path/to/directory"),
    exclude_patterns=[
        "__pycache__/",
        "*.pyc",
        "node_modules/",
        "**/*.log",
    ],
)
```

**Notes**

- Both include_patterns and exclude_patterns follow the 'gitignore' pattern format described in [https://git-scm.com/docs/gitignore#\_pattern_format](https://git-scm.com/docs/gitignore#_pattern_format).
- If both include_patterns and exclude_patterns are provided, files matching exclude_patterns will be excluded even if they match include_patterns.

**Parameters**

- **directory_path** (`pathlib.Path`)
- **include_patterns** (`Optional[list[str]]`)
- **exclude_patterns** (`Optional[list[str]]`)
- **delete_after_upload** (`bool`)
- **max_batch_size** (`int`)
- **print_progress** (`bool`)
- **device_id** (`Optional[str]`)

**Returns**

- `None`

#### Dataset.upload_file()

```python
def upload_file(
    file_path: pathlib.Path,
    file_destination_path: Optional[str] = None,
    print_progress: bool = True,
    device_id: Optional[str] = None,
) -> roboto.domain.files.File
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1528-L1560)

Upload a single file to the dataset. If file_destination_path is not provided, the file will be uploaded to the top-level of the dataset.

**Parameters**

- **file_path** (`pathlib.Path`): Local file to upload.
- **file_destination_path** (`Optional[str]`): Destination path within the dataset. Defaults to the file's own name at the dataset's top level.
- **print_progress** (`bool`): Whether to display an upload progress bar.
- **device_id** (`Optional[str]`): Optional identifier of the device that generated this data.

**Returns**

- `roboto.domain.files.File`: The file record the upload created.

**Raises**

- [`RobotoInternalException`](/reference/python-sdk/roboto/exceptions/domain#roboto.exceptions.domain.RobotoInternalException): The upload reported success without reporting a file ID.

**Usage**

```python
from roboto.domain import datasets
dataset = datasets.Dataset(...)
dataset.upload_file(
    pathlib.Path("/path/to/file.txt"),
    file_destination_path="foo/bar.txt",
)
```

#### Dataset.upload_files()

```python
def upload_files(
    files: collections.abc.Iterable[pathlib.Path],
    file_destination_paths: collections.abc.Mapping[pathlib.Path, str] = {},
    max_batch_size: int = MAX_FILES_PER_MANIFEST,
    print_progress: bool = True,
    device_id: Optional[str] = None,
) -> dict[pathlib.Path, str]
```

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L1562-L1598)

Upload multiple files to the dataset.

If `file_destination_paths` is not provided, files will be uploaded to the top-level of the dataset.

**Parameters**

- **files** (`collections.abc.Iterable[pathlib.Path]`): Local files to upload.
- **file_destination_paths** (`collections.abc.Mapping[pathlib.Path, str]`): Mapping from local path to destination path within the dataset. Files not in the mapping upload to the dataset's top level under their own name.
- **max_batch_size** (`int`): Maximum number of files per upload transaction.
- **print_progress** (`bool`): Whether to display an upload progress bar.
- **device_id** (`Optional[str]`): Optional identifier of the device that generated this data.

**Returns**

- `dict[pathlib.Path, str]`: Mapping from each uploaded local path to the ID of the file record it created.

**Usage**

```python
import pathlib
from roboto.domain import datasets
dataset = datasets.Dataset.from_id("ds_abc123")
file_ids = dataset.upload_files(
    [pathlib.Path("/path/to/file.txt")],
    file_destination_paths={
        pathlib.Path("/path/to/file.txt"): "foo/bar.txt",
    },
)
file_ids[pathlib.Path("/path/to/file.txt")]
# 'fl_0123456789abcdef'
```

### logger

```python
roboto.domain.datasets.dataset.logger
```

`from roboto.domain.datasets.dataset import logger`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/domain/datasets/dataset.py#L58-L58)
