pola-rs/polars · error · TypeError
Iceberg source field
Error message
Iceberg source field {source_id} has non-struct parent What it means
While building the source expression for an Iceberg partition/sort key, _iceberg_source_expr walks a nested accessor chain; every intermediate field must be a StructType to keep descending. A non-struct parent field was encountered mid-path, so polars raises TypeError.
Solutions
- Fix the nested key path so every parent along it is a struct: verify with the table schema before writing.
- Flatten needed values into top-level columns in the DataFrame and partition on those instead.
- If the schema changed, update the sink's partition/sort specification to match the new types.
Example fix
// before
partition_by=["address.city"] # 'address' is a list, not a struct
// after
df = df.with_columns(pl.col("address").struct.field("city").alias("city"))
# partition_by=["city"] Defensive patterns
Strategy: validation
Validate before calling
schema = tbl.schema()
for key in partition_by:
f = schema
for part in key.split('.'):
f = f[part]
if not isinstance(f.field_type, StructType) and part != key.split('.')[-1]:
raise ValueError(f"{part} is not a struct; fix nested key '{key}'") Type guard
def is_struct(field) -> bool:
return isinstance(field.field_type, StructType) Try / catch
try:
df.sink_iceberg(target, partition_by=keys)
except TypeError as e:
if "non-struct parent" in str(e):
df.sink_iceberg(target, partition_by=[keys[0].split('.')[0]])
else:
raise Prevention
- Validate nested partition/sort key paths against the Iceberg schema before sinking.
- Flatten nested values you partition on into top-level columns.
- Re-check partition specs after any schema evolution.
When it happens
Trigger: Defining an Iceberg sink with partition_by or sort keys referencing nested columns (e.g. 'a.b.c') where an intermediate field is a list, map, or primitive instead of a struct.
Common situations: Typos in nested key paths, schema evolution turning a struct field into a list/primitive, or partition specs copied from a table with a different schema shape.
Understand the failure class
Background: Type mismatch errors: IllegalArgumentException, TypeError and type guards across 150 open-source libraries — this error's family across 150 libraries.
Related errors
- Iceberg sort order is no longer available
- native sink row count
- partition source field
- sink to Iceberg table with partition field
- the remote engine can only sink to a URI or a…
AI-assisted analysis of pola-rs/polars@fe841f959e (2026-09-18).
Data as JSON: /api/errors/dd1d3ed07aaea306.
Report an issue: GitHub.
Appendix: source
Thrown at py-polars/src/polars/io/iceberg/_sink.py:75
if schema.accessor_for_field(field.source_id).inner is not None
}
def _iceberg_source_expr(schema: Schema, source_id: int) -> pl.Expr:
from pyiceberg.types import StructType
import polars as pl
accessor = schema.accessor_for_field(source_id)
source_field = schema.fields[accessor.position]
expr = pl.col(source_field.name)
while accessor.inner is not None:
accessor = accessor.inner
source_type = source_field.field_type
if not isinstance(source_type, StructType):
msg = f"Iceberg source field {source_id} has non-struct parent"
raise TypeError(msg)
source_field = source_type.fields[accessor.position]
expr = expr.struct.field(source_field.name)
return expr
def _infer_partition_from_statistics(
statistics: DataFileStatistics, spec: PartitionSpec, schema: Schema
) -> Record:
from pyiceberg.partitioning import partition_record_value
from pyiceberg.typedef import Record
partition_values: list[Any] = []
for field in spec.fields:
aggregate = statistics.column_aggregates.get(field.source_id)
if aggregate is None:
partition_values.append(None)
continueView on GitHub (pinned to fe841f959e)