---
sidebar:
  hidden: true
title: roboto.formats.parquet.arrow_to_roboto
---
## Module Contents

### arrow_type_to_canonical_type()

```python
def roboto.formats.parquet.arrow_to_roboto.arrow_type_to_canonical_type(
    arrow_type: pyarrow.DataType,
) -> roboto.domain.topics.record.CanonicalDataType
```

`from roboto.formats.parquet.arrow_to_roboto import arrow_type_to_canonical_type`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L28-L66)

**Parameters**

- **arrow_type** (`pyarrow.DataType`)

**Returns**

- `roboto.domain.topics.record.CanonicalDataType`

### compute_boolean_statistics()

```python
def roboto.formats.parquet.arrow_to_roboto.compute_boolean_statistics(
    data: Union[pyarrow.Array, pyarrow.ChunkedArray],
) -> dict[str, Any]
```

`from roboto.formats.parquet.arrow_to_roboto import compute_boolean_statistics`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L111-L121)

**Parameters**

- **data** (`Union[pyarrow.Array, pyarrow.ChunkedArray]`)

**Returns**

- `dict[str, Any]`

### compute_dictionary_metadata()

```python
def roboto.formats.parquet.arrow_to_roboto.compute_dictionary_metadata(
    column_name: str,
    data: Union[pyarrow.Array, pyarrow.ChunkedArray],
    max_dictionary_size: int = 2048,
) -> dict[str, Any]
```

`from roboto.formats.parquet.arrow_to_roboto import compute_dictionary_metadata`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L124-L156)

**Parameters**

- **column_name** (`str`)
- **data** (`Union[pyarrow.Array, pyarrow.ChunkedArray]`)
- **max_dictionary_size** (`int`)

**Returns**

- `dict[str, Any]`

### compute_field_metadata()

```python
def roboto.formats.parquet.arrow_to_roboto.compute_field_metadata(
    parser: roboto.formats.parquet.parquet_parser.ParquetParser,
    column_name: str,
    field_path: list[str],
    canonical_data_type: roboto.domain.topics.record.CanonicalDataType,
    is_inside_list: bool = False,
) -> dict[str, Any]
```

`from roboto.formats.parquet.arrow_to_roboto import compute_field_metadata`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L223-L285)

Compute metadata including statistics for a field.

Handles both top-level and nested fields, extracting data appropriately based on the field's location in the schema hierarchy.

**Parameters**

- **parser** (`roboto.formats.parquet.parquet_parser.ParquetParser`): ParquetParser instance to read data from.
- **column_name** (`str`): Name of the top-level column.
- **field_path** (`list[str]`): List of field names to traverse (empty for top-level fields).
- **canonical_data_type** (`roboto.domain.topics.record.CanonicalDataType`): The canonical type of the field.
- **is_inside_list** (`bool`): Whether this field is inside a list (affects data extraction).

**Returns**

- `dict[str, Any]`: Metadata dictionary with statistics if applicable.

### compute_numeric_statistics()

```python
def roboto.formats.parquet.arrow_to_roboto.compute_numeric_statistics(
    data: Union[pyarrow.Array, pyarrow.ChunkedArray],
) -> dict[str, Any]
```

`from roboto.formats.parquet.arrow_to_roboto import compute_numeric_statistics`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L73-L108)

**Parameters**

- **data** (`Union[pyarrow.Array, pyarrow.ChunkedArray]`)

**Returns**

- `dict[str, Any]`

### generate_message_path_requests()

```python
def roboto.formats.parquet.arrow_to_roboto.generate_message_path_requests(
    parser: roboto.formats.parquet.parquet_parser.ParquetParser,
    timestamp: roboto.formats.parquet.timestamp.TimestampInfo,
    max_depth: int = 10,
) -> Generator[roboto.domain.topics.operations.AddMessagePathRequest, None, None]
```

`from roboto.formats.parquet import generate_message_path_requests`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L422-L465)

Generate AddMessagePathRequest objects for all fields in a Parquet schema.

Traverses the schema recursively to generate message paths for nested types (structs, lists) in addition to top-level fields.

**Parameters**

- **parser** (`roboto.formats.parquet.parquet_parser.ParquetParser`): ParquetParser instance containing the schema and data.
- **timestamp** (`roboto.formats.parquet.timestamp.TimestampInfo`): Timestamp information for the topic.
- **max_depth** (`int`): Maximum recursion depth for nested types (default: 10).

**Yields**

- AddMessagePathRequest objects for each field and nested field in the schema.

**Returns**

- `Generator[roboto.domain.topics.operations.AddMessagePathRequest, None, None]`

**Usage**

For a schema with a struct column \`position: struct\<x: float, y: float>\`: - Yields position (Object) - Yields position.x (Number) - Yields position.y (Number)

For a schema with \`values: list\<float64>\`: - Yields values (NumberArray)

For a schema with \`points: list\<struct\<x: float, y: float>>\`: - Yields points (Array) - Yields points.x (Number) - Yields points.y (Number)

### get_list_element_data()

```python
def roboto.formats.parquet.arrow_to_roboto.get_list_element_data(
    parser: roboto.formats.parquet.parquet_parser.ParquetParser,
    column_name: str,
    field_path: list[str],
) -> Union[pyarrow.Array, pyarrow.ChunkedArray]
```

`from roboto.formats.parquet.arrow_to_roboto import get_list_element_data`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L191-L220)

Extract flattened data from list columns for statistics computation.

For list\<primitive> columns, flattens all list elements into a single array. For list\<struct> columns, flattens and then accesses the struct field.

**Parameters**

- **parser** (`roboto.formats.parquet.parquet_parser.ParquetParser`): ParquetParser instance to read data from.
- **column_name** (`str`): Name of the top-level column.
- **field_path** (`list[str]`): List of field names to traverse after flattening the list.

**Returns**

- `Union[pyarrow.Array, pyarrow.ChunkedArray]`: The flattened Array or ChunkedArray suitable for statistics computation.

### get_nested_column_data()

```python
def roboto.formats.parquet.arrow_to_roboto.get_nested_column_data(
    parser: roboto.formats.parquet.parquet_parser.ParquetParser,
    column_name: str,
    field_path: list[str],
) -> Union[pyarrow.Array, pyarrow.ChunkedArray]
```

`from roboto.formats.parquet.arrow_to_roboto import get_nested_column_data`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L159-L188)

Extract data for nested fields from a PyArrow table.

Navigates through struct fields using the provided field path to extract the data for a nested field.

**Parameters**

- **parser** (`roboto.formats.parquet.parquet_parser.ParquetParser`): ParquetParser instance to read data from.
- **column_name** (`str`): Name of the top-level column.
- **field_path** (`list[str]`): List of field names to traverse (excluding the column name).

**Returns**

- `Union[pyarrow.Array, pyarrow.ChunkedArray]`: The extracted Array or ChunkedArray for the nested field.

**Raises**

- `KeyError`: If a field in the path does not exist.

### logger

```python
roboto.formats.parquet.arrow_to_roboto.logger
```

`from roboto.formats.parquet.arrow_to_roboto import logger`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L25-L25)

### sanitize_column_name()

```python
def roboto.formats.parquet.arrow_to_roboto.sanitize_column_name(
    field: pyarrow.Field,
) -> str
```

`from roboto.formats.parquet.arrow_to_roboto import sanitize_column_name`

[Source](https://github.com/roboto-ai/roboto-python-sdk/blob/main/src/roboto/formats/parquet/arrow_to_roboto.py#L69-L70)

**Parameters**

- **field** (`pyarrow.Field`)

**Returns**

- `str`
