roboto.formats.parquet.arrow_to_roboto
Module Contents
arrow_type_to_canonical_type()
Parameters
arrow_type pyarrow.compute_boolean_statistics()
Parameters
data Union[pyarrow.Return type
compute_dictionary_metadata()
Parameters
column_name strdata Union[pyarrow.max_dictionary_size intReturn type
compute_field_metadata()
Compute metadata including statistics for a field.
Handles both top-level and nested fields, extracting data appropriately based on the field’s location in the schema hierarchy.
Parameters
ParquetParser instance to read data from.
column_name strName of the top-level column.
field_path list[str]List of field names to traverse (empty for top-level fields).
canonical_data_type roboto.The canonical type of the field.
is_inside_list boolWhether this field is inside a list (affects data extraction).
Returns
Metadata dictionary with statistics if applicable.
compute_numeric_statistics()
Parameters
data Union[pyarrow.Return type
generate_message_path_requests()
Generate AddMessagePathRequest objects for all fields in a Parquet schema.
Traverses the schema recursively to generate message paths for nested types (structs, lists) in addition to top-level fields.
Parameters
ParquetParser instance containing the schema and data.
Timestamp information for the topic.
max_depth intMaximum recursion depth for nested types (default: 10).
Yields
AddMessagePathRequest objects for each field and nested field in the schema.
Return type
Usage
For a schema with a struct column `position: struct<x: float, y: float>`: - Yields position (Object) - Yields position.x (Number) - Yields position.y (Number)
For a schema with `values: list<float64>`: - Yields values (NumberArray)
For a schema with `points: list<struct<x: float, y: float>>`: - Yields points (Array) - Yields points.x (Number) - Yields points.y (Number)
get_list_element_data()
Extract flattened data from list columns for statistics computation.
For list<primitive> columns, flattens all list elements into a single array. For list<struct> columns, flattens and then accesses the struct field.
Parameters
ParquetParser instance to read data from.
column_name strName of the top-level column.
field_path list[str]List of field names to traverse after flattening the list.
Returns
The flattened Array or ChunkedArray suitable for statistics computation.
get_nested_column_data()
Extract data for nested fields from a PyArrow table.
Navigates through struct fields using the provided field path to extract the data for a nested field.
Parameters
ParquetParser instance to read data from.
column_name strName of the top-level column.
field_path list[str]List of field names to traverse (excluding the column name).
Returns
The extracted Array or ChunkedArray for the nested field.
Raises
KeyErrorIf a field in the path does not exist.
logger
sanitize_column_name()
Parameters
field pyarrow.Return type