pola-rs/polars · error · ValueError

iceberg snapshot ID not found

Error message

iceberg snapshot ID not found: {self.snapshot_id}

What it means

IcebergScanResolver.schema() was asked to resolve the table schema against a specific snapshot ID, but pyiceberg's snapshot_by_id() returned None: the table's metadata contains no snapshot with that ID. polars raises ValueError rather than silently falling back to the current schema.

Solutions

  1. List valid snapshot IDs via table.metadata.snapshots (or table.snapshots()) and use one that exists.
  2. Omit the snapshot_id / use the current snapshot instead of a pinned one.
  3. If the snapshot was expired, restore it from the table's history or re-resolve the scan against the current snapshot.

Example fix

// before
df = pl.scan_iceberg(tbl, snapshot_id=1234567890).collect()  # expired snapshot

// after
valid = {s.snapshot_id for s in tbl.snapshots()}
sid = 1234567890 if 1234567890 in valid else None  # fall back to current
 df = pl.scan_iceberg(tbl, snapshot_id=sid).collect()
Defensive patterns

Strategy: validation

Validate before calling

valid_ids = {s.snapshot_id for s in tbl.snapshots()}
if snapshot_id is not None and snapshot_id not in valid_ids:
    snapshot_id = None  # or pick latest
pl.scan_iceberg(uri, snapshot_id=snapshot_id)

Type guard

def snapshot_exists(tbl, snapshot_id: int) -> bool:
    return snapshot_id is not None and tbl.snapshot_by_id(snapshot_id) is not None

Try / catch

try:
    df = pl.scan_iceberg(uri, snapshot_id=sid).collect()
except ValueError as e:
    if "snapshot ID not found" in str(e):
        df = pl.scan_iceberg(uri).collect()  # current snapshot
    else:
        raise

Prevention

When it happens

Trigger: Constructing IcebergScanResolver with a snapshot_id (e.g. from a previously cached scan version key) that no longer exists in the table metadata, then calling .schema().

Common situations: Scanning a table at an old snapshot ID after the table was expired (expire_snapshots) or rewritten; passing a snapshot ID from a different table; a typo'd or integer-vs-string coerced snapshot ID.

Understand the failure class

Background: Record Not Found Errors: "not found", RecordNotFound, and "was not found" — what they mean and how to fix them — this error's family across 28 libraries.

Related errors


AI-assisted analysis of pola-rs/polars@fe841f959e (2026-09-18). Data as JSON: /api/errors/33679e9d3511e044. Report an issue: GitHub.

Appendix: source

Thrown at py-polars/src/polars/io/iceberg/_dataset.py:235

    fast_deletion_count: bool
    use_pyiceberg_filter: bool

    #
    # PythonDatasetProvider interface functions
    #

    def schema(self) -> pa.schema:
        """Fetch the schema of the table."""
        from pyiceberg.io.pyarrow import schema_to_pyarrow

        if self.snapshot_id is None:
            return self.table.arrow_schema()

        snapshot = self.table.get().snapshot_by_id(self.snapshot_id)

        if snapshot is None:
            msg = f"iceberg snapshot ID not found: {self.snapshot_id}"
            raise ValueError(msg)

        schema_id = snapshot.schema_id

        if schema_id is None:
            msg = (
                f"IcebergScanResolver: requested snapshot {self.snapshot_id} "
                "did not contain a schema ID"
            )
            raise ValueError(msg)
        return schema_to_pyarrow(self.table.get().schemas()[schema_id])

    def to_dataset_scan(
        self,
        *,
        existing_resolved_version_key: str | None = None,
        limit: int | None = None,
        projection: list[str] | None = None,
        filter_columns: list[str] | None = None,

View on GitHub (pinned to fe841f959e)