{"record":{"id":"9bf07b9f39b0d10d","repo":"apache/seatunnel","slug":"failed-to-read-parquet-file-with-avro-reader","errorCode":null,"errorMessage":"Failed to read parquet file [{}] with Avro reader due to illegal Avro field name, fallback to native parquet reader","messagePattern":"Failed to read parquet file \\[(.+?)\\] with Avro reader due to illegal Avro field name, fallback to native parquet reader","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/ParquetReadStrategy.java","lineNumber":109,"sourceCode":"    private static final long JULIAN_DAY_NUMBER_FOR_UNIX_EPOCH = 2440588;\n    private static final String PARQUET = \"Parquet\";\n\n    @Override\n    public void read(String path, String tableId, Collector<SeaTunnelRow> output)\n            throws FileConnectorException, IOException {\n        this.read(new FileSourceSplit(path), output);\n    }\n\n    @Override\n    public void read(FileSourceSplit split, Collector<SeaTunnelRow> output)\n            throws IOException, FileConnectorException {\n        try {\n            readWithAvro(split, output);\n        } catch (RuntimeException e) {\n            if (!isIllegalAvroFieldNameException(e)) {\n                throw e;\n            }\n            log.warn(\n                    \"Failed to read parquet file [{}] with Avro reader due to illegal Avro field\"\n                            + \" name, fallback to native parquet reader\",\n                    split.getFilePath(),\n                    e);\n            readWithNativeParquet(split, output);\n        }\n    }\n\n    private void readWithAvro(FileSourceSplit split, Collector<SeaTunnelRow> output)\n            throws IOException, FileConnectorException {\n        String tableId = split.getTableId();\n        String path = split.getFilePath();\n        if (Boolean.FALSE.equals(checkFileType(path))) {\n            String errorMsg =\n                    String.format(\n                            \"This file [%s] is not a parquet file, please check the format of this file\",\n                            path);\n            throw new FileConnectorException(FileConnectorErrorCode.FILE_TYPE_INVALID, errorMsg);","sourceCodeStart":91,"sourceCodeEnd":127,"githubUrl":"https://github.com/apache/seatunnel/blob/cf67b549a7a6c35fa0beb12d83c62892427ea919/seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/ParquetReadStrategy.java#L91-L127","documentation":"Warning in ParquetReadStrategy.read: the Avro-based parquet reader failed because the parquet schema contains field names that are illegal Avro field names (e.g. starting with a digit, containing characters Avro disallows). The reader detects this via isIllegalAvroFieldNameException and falls back to the native parquet reader instead of failing.","triggerScenarios":"readWithAvro throws a RuntimeException classified as an illegal-Avro-field-name error — parquet files written by systems (Spark/Flink/pandas) that allow field names like '1col', 'my field', 'col-1' which Avro's schema rules reject.","commonSituations":"Reading parquet produced by Spark with columns created from arbitrary JSON keys or user data; case-sensitive or unicode column names; migrating data from engines with laxer naming rules.","solutions":["No action needed — the fallback to the native parquet reader is automatic; verify output correctness.","Rename offending parquet columns at write time to valid Avro identifiers ([A-Za-z_][A-Za-z0-9_]*).","Sanitize upstream field names (e.g. in Spark before writing) if Avro-based reading is preferred for performance.","If the fallback misbehaves, force native parquet reading or upgrade to a version with improved field-name handling."],"exampleFix":"// before (parquet schema)\nfield \"1st_column\"\n// after\nfield \"first_column\"","handlingStrategy":"fallback","validationCode":"// validate field names before writing parquet\nPattern p = Pattern.compile(\"[A-Za-z_][A-Za-z0-9_]*\");\nfor (String col : columns) {\n    if (!p.matcher(col).matches()) { /* rename column */ }\n}","typeGuard":"boolean isAvroSafeName(String name) {\n    return name != null && name.matches(\"[A-Za-z_][A-Za-z0-9_]*\");\n}","tryCatchPattern":"try {\n    readWithAvro(split, output);\n} catch (RuntimeException e) {\n    if (!isIllegalAvroFieldNameException(e)) throw e;\n    readWithNativeParquet(split, output);\n}","preventionTips":["Sanitize column names to Avro-valid identifiers at data production time.","Avoid raw user/JSON keys as parquet column names.","Treat the fallback warning as informational; native reader handles these files."],"tags":["parquet","avro","schema","fallback"],"backgroundTag":"invalid-identifier-format","analyzedSha":"cf67b549a7a6c35fa0beb12d83c62892427ea919","analyzedAt":"2026-09-10T21:44:55.265Z","contentChangedAt":"2026-09-10T21:44:55.265Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}