apache/iceberg · error · IllegalArgumentException

Encountered an unsupported ORC type during a write from…

Error message

Encountered an unsupported ORC type during a write from Spark.

What it means

While creating the field getter for a column, SparkOrcWriter switches on the ORC type category (struct, list, map, primitive). An ORC type outside the supported categories triggers this IllegalArgumentException, indicating the ORC schema contains a type the Spark write path cannot read values from.

Solutions

  1. Remove or rewrite ORC union-typed columns; Iceberg does not support ORC unions.
  2. Regenerate files with an Iceberg-supported schema.
  3. Check the file schema with orc-tools (meta) to identify the unsupported category.
Defensive patterns

Strategy: validation

Validate before calling

for (TypeDescription child : fileSchema.getChildren()) {
  switch (child.getCategory()) {
    case STRUCT: case LIST: case MAP: case UNION: -> { if (child.getCategory() == TypeDescription.Category.UNION) throw new IllegalStateException("ORC union unsupported: " + child); }
    default -> {}
  }
}

Type guard

boolean orcTypeSupported(TypeDescription t) {
  return t.getCategory() != TypeDescription.Category.UNION;
}

Try / catch

try {
  sparkOrcWriter.write(rows);
} catch (IllegalArgumentException e) {
  if (e.getMessage().contains("unsupported ORC type")) {
    // rewrite files without union columns
  }
}

Prevention

When it happens

Trigger: Writing to ORC where a column's ORC TypeDescription category is not STRUCT/LIST/MAP/PRIMITIVE (e.g. UNION type in the ORC schema).

Common situations: Reading/writing ORC files produced by non-Iceberg tools that use ORC union types; hand-edited ORC schemas.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/9907399377c06db1. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/data/SparkOrcWriter.java:221

            (row, ordinal) ->
                row.getDecimal(ordinal, fieldType.getPrecision(), fieldType.getScale());
        break;
      case STRING:
      case CHAR:
      case VARCHAR:
        fieldGetter = SpecializedGetters::getUTF8String;
        break;
      case STRUCT:
        fieldGetter = (row, ordinal) -> row.getStruct(ordinal, fieldType.getChildren().size());
        break;
      case LIST:
        fieldGetter = SpecializedGetters::getArray;
        break;
      case MAP:
        fieldGetter = SpecializedGetters::getMap;
        break;
      default:
        throw new IllegalArgumentException(
            "Encountered an unsupported ORC type during a write from Spark.");
    }

    return (row, ordinal) -> {
      if (row.isNullAt(ordinal)) {
        return null;
      }
      return fieldGetter.getFieldOrNull(row, ordinal);
    };
  }

  interface FieldGetter<T> extends Serializable {

    /**
     * Returns a value from a complex Spark data holder such ArrayData, InternalRow, etc... Calls
     * the appropriate getter for the expected data type.
     *
     * @param row Spark's data representation

View on GitHub (pinned to 86d9c8fc54)