apache/iceberg · error · IllegalStateException

Configuration was not serialized on purpose but was not set…

Error message

Configuration was not serialized on purpose but was not set manually either

What it means

Iceberg's mapreduce IcebergInputFormat wraps the Hadoop Configuration in a NonSerializingConfig that deliberately refuses to serialize the Configuration object (Configurations are large and often unserializable). The wrapper only holds a reference to a Configuration injected after deserialization; calling get() before that injection returns null and triggers this IllegalStateException.

Solutions

  1. Set the Configuration manually after deserialization before calling get() (e.g. implement Configurable/setConf or re-inject via new NonSerializingConfig(conf)).
  2. Do not capture the Configuration in objects that get serialized into distributed tasks; fetch it inside the task from the job's context (context.getConfiguration()).
  3. If you own the code path, ensure the wrapper is initialized with a non-null Configuration before use.

Example fix

// before
Configuration conf = serializedHolder.getConfig().get(); // throws IllegalStateException
// after
Configuration conf = context.getConfiguration(); // re-inject in the task
serializedHolder.setConfig(conf);
Configuration resolved = serializedHolder.getConfig().get();
Defensive patterns

Strategy: validation

Validate before calling

if (configRef == null || configRef.get() == null) {
  throw new IllegalArgumentException("Configuration must be set before use; inject via context.getConfiguration()");
}

Type guard

boolean hasConfig(NonSerializingConfig c) { return c != null && c.get() != null; } // guard in try/catch or via null-check wrapper if exposed

Try / catch

try {
  Configuration conf = nonSerializingConfig.get();
} catch (IllegalStateException e) {
  conf = context.getConfiguration(); // re-inject and retry
  nonSerializingConfig.set(conf);
}

Prevention

When it happens

Trigger: Calling NonSerializingConfig.get() when the underlying Configuration reference is null — i.e. code path that obtained the config via deserialization (e.g. a mapper/reducer or lambda serialized into a task) and no one called setConf/manually re-injected a Configuration before use.

Common situations: Shipping a closure or object referencing the input format's config into a MapReduce/Spark task without re-injecting the Hadoop Configuration; testing code that constructs NonSerializingConfig with a null conf.

Understand the failure class

Background: "X is required", "must be set", "cannot be empty": the missing-required-config error family, from Vertex AI project/location to WeChat keys — this error's family across 18 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/670f2dfe9ac84a20. Report an issue: GitHub.

Appendix: source

Thrown at mr/src/main/java/org/apache/iceberg/mr/mapreduce/IcebergInputFormat.java:218

    }
  }

  @Override
  public RecordReader<Void, T> createRecordReader(InputSplit split, TaskAttemptContext context) {
    return new IcebergRecordReader<>();
  }

  private static class NonSerializingConfig implements Serializable {

    private final transient Configuration conf;

    NonSerializingConfig(Configuration conf) {
      this.conf = conf;
    }

    public Configuration get() {
      if (conf == null) {
        throw new IllegalStateException(
            "Configuration was not serialized on purpose but was not set manually either");
      }

      return conf;
    }
  }

  private static final class IcebergRecordReader<T> extends RecordReader<Void, T> {

    private TaskAttemptContext context;
    private Schema tableSchema;
    private Schema expectedSchema;
    private String nameMapping;
    private boolean reuseContainers;
    private boolean caseSensitive;
    private Iterator<FileScanTask> tasks;
    private T current;
    private CloseableIterator<T> currentIterator;

View on GitHub (pinned to 86d9c8fc54)