apache/iceberg · error · IllegalArgumentException

Could not parse Avro schema string.

Error message

Could not parse Avro schema string.

What it means

AvroSchemaConverter.convertToTypeInfo parses an Avro schema given as a JSON string before converting it to Flink TypeInformation. If Schema.Parser().parse fails (malformed JSON, invalid Avro constructs), it throws IllegalArgumentException('Could not parse Avro schema string.') with the SchemaParseException as cause.

Solutions

  1. Validate the string with a JSON linter / Schema.Parser before passing it in; inspect the cause SchemaParseException for the exact position.
  2. Ensure the schema is read completely from its source (registry, file, resource) and not truncated.
  3. Fix escaping when the schema is embedded in Java strings or YAML/properties config.
  4. Generate the schema programmatically via AvroSchemaConverter.convertToSchema(LogicalType) instead of hand-writing JSON.

Example fix

// before
String schema = "{\"type\":\"record\",\"name\":\"R\",\"fields\":[{"name":"a","type":"int"}"; // truncated
// after
String schema = "{\"type\":\"record\",\"name\":\"R\",\"fields\":[{\"name\":\"a\",\"type\":\"int\"}]}";
Defensive patterns

Strategy: validation

Validate before calling

try {
  new Schema.Parser().parse(avroSchemaString);
} catch (SchemaParseException e) {
  throw new IllegalArgumentException("Invalid Avro schema supplied to pipeline: " + e.getMessage(), e);
}

Try / catch

try { ti = AvroSchemaConverter.convertToTypeInfo(schemaString, legacy); } catch (IllegalArgumentException e) { log.error("Schema parse failed: {}", e.getCause() != null ? e.getCause().getMessage() : e.getMessage()); throw e; }

Prevention

When it happens

Trigger: Calling AvroSchemaConverter.convertToTypeInfo(avroSchemaString, legacyTimestampMapping) with a string that is not valid Avro schema JSON: syntax errors, unknown types, truncated output, or non-JSON content.

Common situations: Reading the schema from a file/registry and getting HTML or a connection-error page; string truncated by logging; escaping issues when the schema is embedded in config or code; single vs double quote mixups.

Understand the failure class

Background: JSON parse error: "Unexpected token" / "not valid JSON" / "failed to parse" — what JSON parsers are really complaining about — this error's family across 45 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/1508997d06df56b6. Report an issue: GitHub.

Appendix: source

Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/formats/avro/typeutils/AvroSchemaConverter.java:127

  }

  /**
   * Converts an Avro schema string into a nested row structure with deterministic field order and
   * data types that are compatible with Flink's Table & SQL API.
   *
   * @param avroSchemaString Avro schema definition string
   * @param legacyTimestampMapping legacy mapping of timestamp types
   * @return type information matching the schema
   */
  @SuppressWarnings("unchecked")
  public static <T> TypeInformation<T> convertToTypeInfo(
      String avroSchemaString, boolean legacyTimestampMapping) {
    Preconditions.checkNotNull(avroSchemaString, "Avro schema must not be null.");
    final Schema schema;
    try {
      schema = new Schema.Parser().parse(avroSchemaString);
    } catch (SchemaParseException e) {
      throw new IllegalArgumentException("Could not parse Avro schema string.", e);
    }
    return (TypeInformation<T>) convertToTypeInfo(schema, legacyTimestampMapping);
  }

  private static TypeInformation<?> convertToTypeInfo(
      Schema schema, boolean legacyTimestampMapping) {
    switch (schema.getType()) {
      case RECORD:
        final List<Schema.Field> fields = schema.getFields();

        final TypeInformation<?>[] types = new TypeInformation<?>[fields.size()];
        final String[] names = new String[fields.size()];
        for (int i = 0; i < fields.size(); i++) {
          final Schema.Field field = fields.get(i);
          types[i] = convertToTypeInfo(field.schema(), legacyTimestampMapping);
          names[i] = field.name();
        }
        return Types.ROW_NAMED(names, types);

View on GitHub (pinned to 86d9c8fc54)