pola-rs/polars · error · ValueError

IcebergScanResolver: requested snapshot

Error message

IcebergScanResolver: requested snapshot {self.snapshot_id} did not contain a schema ID

What it means

The requested snapshot exists but its metadata carries schema_id = None, so IcebergScanResolver.schema() cannot map the snapshot to a schema in the table's schemas() map. polars raises ValueError because a schema is mandatory to build the Arrow schema for the scan.

Solutions

  1. Use a different snapshot (one that has a schema_id) or omit snapshot_id to use the current schema.
  2. Rewrite the table metadata (e.g. with pyiceberg's rewrite or a rewrite_data_files/migrate operation) so snapshots carry a schema_id.
  3. As a workaround, read the table's current schema via table.schema() and construct the scan without snapshot pinning.

Example fix

// before
resolver.schema()  # snapshot has schema_id == None

// after
snapshot = tbl.snapshot_by_id(resolver.snapshot_id)
if snapshot.schema_id is None:
    snapshot = tbl.current_snapshot()  # or another snapshot with a schema_id
Defensive patterns

Strategy: try-catch

Validate before calling

snapshot = tbl.snapshot_by_id(snapshot_id)
if snapshot is not None and snapshot.schema_id is None:
    snapshot_id = None  # fall back to current schema

Type guard

def has_schema_id(tbl, snapshot_id: int | None) -> bool:
    if snapshot_id is None:
        return True
    snap = tbl.snapshot_by_id(snapshot_id)
    return snap is not None and snap.schema_id is not None

Try / catch

try:
    resolver.schema()
except ValueError as e:
    if "did not contain a schema ID" in str(e):
        resolver = IcebergScanResolver(..., snapshot_id=None)
        return resolver.schema()
    raise

Prevention

When it happens

Trigger: Calling IcebergScanResolver.schema() on a snapshot whose Snapshot.schema_id is None — typically snapshots produced by writers that did not record a schema-id field in snapshot metadata.

Common situations: Tables written by older or non-standard Iceberg writers; snapshots imported from legacy formats; hand-built/patched table metadata in testing.

Understand the failure class

Background: "missing required argument" and "the following required arguments were not provided": what required-argument errors mean and how to fix them — this error's family across 20 libraries.

Related errors


AI-assisted analysis of pola-rs/polars@fe841f959e (2026-09-18). Data as JSON: /api/errors/da0dcce2ce69aacb. Report an issue: GitHub.

Appendix: source

Thrown at py-polars/src/polars/io/iceberg/_dataset.py:244

        from pyiceberg.io.pyarrow import schema_to_pyarrow

        if self.snapshot_id is None:
            return self.table.arrow_schema()

        snapshot = self.table.get().snapshot_by_id(self.snapshot_id)

        if snapshot is None:
            msg = f"iceberg snapshot ID not found: {self.snapshot_id}"
            raise ValueError(msg)

        schema_id = snapshot.schema_id

        if schema_id is None:
            msg = (
                f"IcebergScanResolver: requested snapshot {self.snapshot_id} "
                "did not contain a schema ID"
            )
            raise ValueError(msg)
        return schema_to_pyarrow(self.table.get().schemas()[schema_id])

    def to_dataset_scan(
        self,
        *,
        existing_resolved_version_key: str | None = None,
        limit: int | None = None,
        projection: list[str] | None = None,
        filter_columns: list[str] | None = None,
        pyarrow_predicate: str | None = None,
    ) -> tuple[LazyFrame, str] | None:
        """Construct a LazyFrame scan."""
        if (
            scan_data := self._to_dataset_scan_impl(
                existing_resolved_version_key=existing_resolved_version_key,
                limit=limit,
                projection=projection,
                filter_columns=filter_columns,

View on GitHub (pinned to fe841f959e)