---
sidebar:
  hidden: true
title: roboto.formats.parquet.table_transforms
---
## Module Contents

### compute_time_filter_mask()

```python
def roboto.formats.parquet.table_transforms.compute_time_filter_mask(
    timestamps: pyarrow.Array,
    start_time: Optional[int] = None,
    end_time: Optional[int] = None,
) -> Optional[pyarrow.BooleanArray]
```

`from roboto.formats.parquet import compute_time_filter_mask`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L23-L50)

Compute a boolean mask indicating which rows fall within the specified time range. Returns None if no time filtering is needed (both start_time and end_time are None).

**Parameters**

- **timestamps** (`pyarrow.Array`)
- **start_time** (`Optional[int]`)
- **end_time** (`Optional[int]`)

**Returns**

- `Optional[pyarrow.BooleanArray]`

### extract_timestamp_field()

```python
def roboto.formats.parquet.table_transforms.extract_timestamp_field(
    schema: pyarrow.Schema,
    timestamp_field: roboto.formats.fields.FieldSelection,
    unit_hint: Optional[str],
) -> roboto.formats.parquet.timestamp.Timestamp
```

`from roboto.formats.parquet import extract_timestamp_field`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L114-L155)

Find a Parquet file's timestamp field and describe it as a [`Timestamp`](/reference/python-sdk/roboto/formats/parquet/timestamp#roboto.formats.parquet.timestamp.Timestamp).

The field is found by walking `timestamp_field.path_in_schema` one component at a time, first among the schema's top-level fields and then through struct children, so a nested field and a top-level column whose name contains dots are never confused. `schema` must be the file's `ParquetFile.schema_arrow`: the timestamp's column index is counted over the schema's leaves, which match the file's leaf columns one for one.

`unit_hint` is the unit of the stored values, used when the field's Arrow type does not carry one (an integer, floating-point or decimal field). Callers take it from the `Unit` metadata of the timestamp's [`MessagePathRecord`](/reference/python-sdk/roboto/domain/topics/record#roboto.domain.topics.record.MessagePathRecord), or from [`unit`](/reference/python-sdk/roboto/experimental/topics/read_plan#roboto.experimental.topics.read_plan.ReadPlanTimestamp.unit).

**Parameters**

- **schema** (`pyarrow.Schema`)
- **timestamp_field** (`roboto.formats.fields.FieldSelection`)
- **unit_hint** (`Optional[str]`)

**Raises**

- `KeyError`: A path component is not in the schema, or names a child of a field that is not a struct.

**Returns**

- `roboto.formats.parquet.timestamp.Timestamp`

### extract_timestamps()

```python
def roboto.formats.parquet.table_transforms.extract_timestamps(
    table: pyarrow.Table,
    timestamp: roboto.formats.parquet.timestamp.Timestamp,
) -> pyarrow.Int64Array
```

`from roboto.formats.parquet import extract_timestamps`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L53-L111)

Extract timestamps in nanoseconds since Unix epoch from the table's timestamp field.

The field is found by walking `timestamp.path` through the table's struct columns. A row whose enclosing struct is null has a null timestamp.

**Parameters**

- **table** (`pyarrow.Table`)
- **timestamp** (`roboto.formats.parquet.timestamp.Timestamp`)

**Returns**

- `pyarrow.Int64Array`

### narrow_list_nested_fields()

```python
def roboto.formats.parquet.table_transforms.narrow_list_nested_fields(
    table: pyarrow.Table,
    schema: pyarrow.Schema,
    fields: collections.abc.Iterable[roboto.formats.fields.FieldSelection],
) -> pyarrow.Table
```

`from roboto.formats.parquet import narrow_list_nested_fields`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L395-L446)

Prune list-of-struct columns to the projected leaves inside each element.

PyArrow's prefix-based nested column selection cannot reach through list wrapper nodes, so [`resolve_columns()`](/reference/python-sdk/roboto/formats/parquet/table_transforms#roboto.formats.parquet.table_transforms.resolve_columns) reads a list-nested leaf's whole list ancestor column — every element keeps all of its struct fields. This Arrow-native post-read pass narrows each such element down to the requested leaves, leaving every other read path byte-identical.

A top-level root is narrowed iff at least one of its projected paths has a list ancestor; otherwise the table is returned unchanged (pure struct, scalar, and scalar-list reads never enter the rebuild). Per root, a trie is built from its paths with the root component stripped so non-list-nested siblings the projection also keeps are preserved. Every struct keeps its fields in the order the file stores them.

No field of `fields` may lie inside another: the outer one would be narrowed to the inner one rather than kept whole.

**Parameters**

- **table** (`pyarrow.Table`)
- **schema** (`pyarrow.Schema`)
- **fields** (`collections.abc.Iterable[roboto.formats.fields.FieldSelection]`)

**Returns**

- `pyarrow.Table`

### resolve_columns()

```python
def roboto.formats.parquet.table_transforms.resolve_columns(
    schema: pyarrow.Schema,
    fields: collections.abc.Iterable[roboto.formats.fields.FieldSelection],
) -> list[str]
```

`from roboto.formats.parquet import resolve_columns`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L276-L309)

Build a deduplicated list of column names safe for `read_row_group(columns=...)`.

Children of list-type columns are replaced by their list ancestor's column name because PyArrow's prefix-based nested column selection does not work through list wrapper nodes in the physical Parquet schema. Selecting the parent list column already returns its full nested structure.

This is important because the projected fields contain only *leaf* paths. For a column like `points: list<struct<x, y>>`, only `points.x` and `points.y` are selected — the parent `points` field is absent. This function derives the correct parent column name from the child's `path_in_schema`.

Children of struct-type columns are preserved because PyArrow can resolve them via dot-separated prefix matching (e.g. `"position.x"` selects the `x` child of the `position` struct).

Each name joins path components with dots, the form in which `read_row_group` takes columns. Two fields whose paths join to the same name are both selected by it: `"header.stamp"` selects both a `stamp` field nested in a `header` struct and a top-level column of that name. [`select_fields()`](/reference/python-sdk/roboto/formats/parquet/table_transforms#roboto.formats.parquet.table_transforms.select_fields) removes the fields that were not projected.

**Parameters**

- **schema** (`pyarrow.Schema`)
- **fields** (`collections.abc.Iterable[roboto.formats.fields.FieldSelection]`)

**Returns**

- `list[str]`

### select_fields()

```python
def roboto.formats.parquet.table_transforms.select_fields(
    table: pyarrow.Table,
    fields: collections.abc.Iterable[roboto.formats.fields.FieldSelection],
) -> pyarrow.Table
```

`from roboto.formats.parquet import select_fields`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L537-L590)

Keep only the fields `fields` names, found by walking their path components, in the table's order.

Reading columns by their dot-joined names ([`resolve_columns()`](/reference/python-sdk/roboto/formats/parquet/table_transforms#roboto.formats.parquet.table_transforms.resolve_columns)) can bring in more than was projected, such as a top-level column named `header.stamp` read along with a `header` struct's `stamp` field, or a timestamp field read only to filter rows. This removes them:

- A top-level column that no field's path starts with is dropped.
- A field that names a column or struct keeps it whole, even when other fields name some of its children.
- Otherwise a struct keeps only the children that some field's path runs through, with its own null rows.
- Lists and maps, and everything below them, are kept as read ([`narrow_list_nested_fields()`](/reference/python-sdk/roboto/formats/parquet/table_transforms#roboto.formats.parquet.table_transforms.narrow_list_nested_fields) narrows the structs inside a list).
- A field whose path the table lacks is skipped.

A column from which nothing is removed is returned as read. With no fields at all, the result has no columns and keeps the table's row count.

**Parameters**

- **table** (`pyarrow.Table`)
- **fields** (`collections.abc.Iterable[roboto.formats.fields.FieldSelection]`)

**Returns**

- `pyarrow.Table`

### should_narrow_list_nested_fields()

```python
def roboto.formats.parquet.table_transforms.should_narrow_list_nested_fields(
    schema: pyarrow.Schema,
    fields: collections.abc.Iterable[roboto.formats.fields.FieldSelection],
) -> bool
```

`from roboto.formats.parquet import should_narrow_list_nested_fields`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L378-L392)

Return whether [`narrow_list_nested_fields()`](/reference/python-sdk/roboto/formats/parquet/table_transforms#roboto.formats.parquet.table_transforms.narrow_list_nested_fields) would change the table.

True iff at least one projected field addresses a leaf *inside* a list (its path has a list ancestor). When False, every projected field resolves through structs and scalars alone, so PyArrow's column selection already returns the narrowed shape and the post-read prune is a no-op — callers can skip it.

Cheap enough to evaluate once per file and hoist the per-row-group narrowing decision out of the decode loop.

**Parameters**

- **schema** (`pyarrow.Schema`)
- **fields** (`collections.abc.Iterable[roboto.formats.fields.FieldSelection]`)

**Returns**

- `bool`

### should_read_row_group()

```python
def roboto.formats.parquet.table_transforms.should_read_row_group(
    row_group_metadata: pyarrow.parquet.RowGroupMetaData,
    timestamp: roboto.formats.parquet.timestamp.Timestamp,
    start_time: Optional[int] = None,
    end_time: Optional[int] = None,
) -> bool
```

`from roboto.formats.parquet import should_read_row_group`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L201-L227)

Determine whether a Parquet row group contains data within the requested time range. Used to short-circuit requesting column chunks from the given row group if not relevant.

**Parameters**

- **row_group_metadata** (`pyarrow.parquet.RowGroupMetaData`)
- **timestamp** (`roboto.formats.parquet.timestamp.Timestamp`)
- **start_time** (`Optional[int]`)
- **end_time** (`Optional[int]`)

**Returns**

- `bool`

### timestamp_statistics()

```python
def roboto.formats.parquet.table_transforms.timestamp_statistics(
    row_group_metadata: pyarrow.parquet.RowGroupMetaData,
    timestamp: roboto.formats.parquet.timestamp.Timestamp,
) -> Optional[pyarrow.parquet.Statistics]
```

`from roboto.formats.parquet import timestamp_statistics`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/table_transforms.py#L178-L198)

The statistics of the timestamp's column chunk in a row group, or `None` when there are none to use.

A nested field and a top-level column whose name contains dots can share a column chunk's `path_in_schema` (`header.stamp` names both a `header` struct's `stamp` field and a column of that name), so the chunk is found by its position, `timestamp.column_index`. Returns `None` when the row group has no chunk at that index, when that chunk's `path_in_schema` is not the timestamp's path joined with dots (`timestamp` was found in a schema other than this file's), or when the chunk has no statistics.

**Parameters**

- **row_group_metadata** (`pyarrow.parquet.RowGroupMetaData`)
- **timestamp** (`roboto.formats.parquet.timestamp.Timestamp`)

**Returns**

- `Optional[pyarrow.parquet.Statistics]`
