pola-rs/polars · error · TypeError

DataFrame should contain only String repr data; found

Error message

DataFrame should contain only String repr data; found {tp!r}

What it means

_cast_repr_strings_with_schema converts a DataFrame built from polars object-repr text back into typed values, and it requires every column to be String typed. If any column has a non-String dtype, it raises TypeError naming the offending dtype. It's an internal precondition on repr-parsing input.

Solutions

  1. Cast the DataFrame to all-String before parsing: df.cast({c: pl.String for c in df.columns}).
  2. Build the DataFrame with an explicit schema of pl.String for every column.
  3. Ensure stringified values: convert each value with str(v) rather than relying on inference.

Example fix

// before
df = pl.DataFrame({'a': [1, 2]})  # Int64 columns
// after
df = pl.DataFrame({'a': ['1', '2']}, schema={'a': pl.String})
Defensive patterns

Strategy: validation

Validate before calling

if df.schema and not all(tp == pl.String for tp in df.schema.values()):
    df = df.cast({c: pl.String for c in df.columns})

Type guard

def is_all_string(df) -> bool:
    return all(tp == pl.String for tp in df.schema.values())

Try / catch

try:
    out = _cast_repr_strings_with_schema(df, schema)
except TypeError as e:
    if 'only String repr data' in str(e):
        out = _cast_repr_strings_with_schema(df.cast({c: pl.String for c in df.columns}), schema)

Prevention

When it happens

Trigger: Constructing a DataFrame for _from_dataframe_repr/_from_series_repr parsing where a column was created with a non-String dtype (e.g. from_dict with ints/None-inferring dtypes), then passing it to the repr parser.

Common situations: Building repr DataFrames manually in tests or tooling and letting polars infer dtypes (numbers, nulls), instead of forcing pl.String on all columns.

Related errors


AI-assisted analysis of pola-rs/polars@fe841f959e (2026-09-18). Data as JSON: /api/errors/1aaa2dee1a9ce3e5. Report an issue: GitHub.

Appendix: source

Thrown at py-polars/src/polars/_utils/various.py:337

    Parameters
    ----------
    df
        Dataframe containing string-repr column data.
    schema
        DataFrame schema containing the desired end-state types.

    Notes
    -----
    Table repr strings are less strict (or different) than equivalent CSV data, so need
    special handling; as this function is only used for reprs, parsing is flexible.
    """
    tp: PolarsDataType | None
    if not df.is_empty():
        for tp in df.schema.values():
            if tp != String:
                msg = f"DataFrame should contain only String repr data; found {tp!r}"
                raise TypeError(msg)

    special_floats = {"-inf", "+inf", "inf", "nan"}

    # duration string scaling
    ns_sec = 1_000_000_000
    duration_scaling = {
        "ns": 1,
        "us": 1_000,
        "µs": 1_000,
        "ms": 1_000_000,
        "s": ns_sec,
        "m": ns_sec * 60,
        "h": ns_sec * 60 * 60,
        "d": ns_sec * 3_600 * 24,
        "w": ns_sec * 3_600 * 24 * 7,
    }

    # identify duration units and convert to nanoseconds

View on GitHub (pinned to fe841f959e)