apache/seatunnel · warning

Failed to read parquet file

Error message

Failed to read parquet file [{}] with Avro reader due to illegal Avro field name, fallback to native parquet reader

What it means

Warning in ParquetReadStrategy.read: the Avro-based parquet reader failed because the parquet schema contains field names that are illegal Avro field names (e.g. starting with a digit, containing characters Avro disallows). The reader detects this via isIllegalAvroFieldNameException and falls back to the native parquet reader instead of failing.

Solutions

  1. No action needed — the fallback to the native parquet reader is automatic; verify output correctness.
  2. Rename offending parquet columns at write time to valid Avro identifiers ([A-Za-z_][A-Za-z0-9_]*).
  3. Sanitize upstream field names (e.g. in Spark before writing) if Avro-based reading is preferred for performance.
  4. If the fallback misbehaves, force native parquet reading or upgrade to a version with improved field-name handling.

Example fix

// before (parquet schema)
field "1st_column"
// after
field "first_column"
Defensive patterns

Strategy: fallback

Validate before calling

// validate field names before writing parquet
Pattern p = Pattern.compile("[A-Za-z_][A-Za-z0-9_]*");
for (String col : columns) {
    if (!p.matcher(col).matches()) { /* rename column */ }
}

Type guard

boolean isAvroSafeName(String name) {
    return name != null && name.matches("[A-Za-z_][A-Za-z0-9_]*");
}

Try / catch

try {
    readWithAvro(split, output);
} catch (RuntimeException e) {
    if (!isIllegalAvroFieldNameException(e)) throw e;
    readWithNativeParquet(split, output);
}

Prevention

When it happens

Trigger: readWithAvro throws a RuntimeException classified as an illegal-Avro-field-name error — parquet files written by systems (Spark/Flink/pandas) that allow field names like '1col', 'my field', 'col-1' which Avro's schema rules reject.

Common situations: Reading parquet produced by Spark with columns created from arbitrary JSON keys or user data; case-sensitive or unicode column names; migrating data from engines with laxer naming rules.

Understand the failure class

Background: "invalid id" errors: invalid identifier format — why libraries reject IDs before lookup, and how to fix them — this error's family across 37 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/9bf07b9f39b0d10d. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/ParquetReadStrategy.java:109

    private static final long JULIAN_DAY_NUMBER_FOR_UNIX_EPOCH = 2440588;
    private static final String PARQUET = "Parquet";

    @Override
    public void read(String path, String tableId, Collector<SeaTunnelRow> output)
            throws FileConnectorException, IOException {
        this.read(new FileSourceSplit(path), output);
    }

    @Override
    public void read(FileSourceSplit split, Collector<SeaTunnelRow> output)
            throws IOException, FileConnectorException {
        try {
            readWithAvro(split, output);
        } catch (RuntimeException e) {
            if (!isIllegalAvroFieldNameException(e)) {
                throw e;
            }
            log.warn(
                    "Failed to read parquet file [{}] with Avro reader due to illegal Avro field"
                            + " name, fallback to native parquet reader",
                    split.getFilePath(),
                    e);
            readWithNativeParquet(split, output);
        }
    }

    private void readWithAvro(FileSourceSplit split, Collector<SeaTunnelRow> output)
            throws IOException, FileConnectorException {
        String tableId = split.getTableId();
        String path = split.getFilePath();
        if (Boolean.FALSE.equals(checkFileType(path))) {
            String errorMsg =
                    String.format(
                            "This file [%s] is not a parquet file, please check the format of this file",
                            path);
            throw new FileConnectorException(FileConnectorErrorCode.FILE_TYPE_INVALID, errorMsg);

View on GitHub (pinned to cf67b549a7)