apache/iceberg · error · IllegalArgumentException
Cannot parse predicates in where option:
Error message
Cannot parse predicates in where option:
What it means
Iceberg procedures accepting a 'where' option parse the predicate via Spark (collecting a resolved expression from a query against the table). If parsing raises AnalysisException — the predicate is invalid SQL or references unknown columns — it is rethrown as IllegalArgumentException with the offending where string.
Solutions
- Test the predicate in a Spark SQL SELECT first (SELECT * FROM table WHERE <expr>) to confirm it resolves.
- Fix column names/types in the where clause to match the table schema (df.printSchema()).
- Escape single quotes correctly in the where string within the CALL statement.
Example fix
// before call remove_orphan_files(table => 'db.t', where => "event_time < '2024-01-01'") -- no such column // after call remove_orphan_files(table => 'db.t', where => "ts < '2024-01-01'")
Defensive patterns
Strategy: validation
Validate before calling
spark.sql(s"SELECT * FROM $table WHERE $where").limit(1).collect() // probe before CALL
Try / catch
try {
spark.sql(s"CALL cat.system.remove_orphan_files(table => '$t', where => '$w')")
} catch {
case e: IllegalArgumentException if e.getMessage.startsWith("Cannot parse predicates") =>
logger.error(s"Bad where clause '$w': ${e.getCause}")
} Prevention
- Test predicates in a plain SELECT first
- Match column names/types against the table schema
- Beware quote escaping inside CALL statements
When it happens
Trigger: Calling a procedure (remove_orphan_files, rewrite_data_files, expire_snapshots) with where => '<expr>' where expr fails Spark analysis: syntax error, non-existent column, or type-incompatible comparison.
Common situations: Referencing partition columns by wrong names; using Spark SQL functions unsupported in the expression; quoting issues in SQL strings embedded in procedure calls.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- ALTER TABLE contains multiple distribution clauses
- ALTER TABLE contains multiple distribution clauses
- ALTER TABLE contains multiple ordering clauses
- ALTER TABLE contains multiple ordering clauses
- ALTER TABLE has no changes: missing both distribution and…
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/af9e1b473f6e6bf5.
Report an issue: GitHub.
Appendix: source
Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/procedures/BaseProcedure.java:198
String tableName = Spark3Util.quotedFullIdentifier(tableCatalog().name(), tableIdent);
return spark().read().options(options).table(tableName);
}
protected void refreshSparkCache(Identifier ident, Table table) {
CacheManager cacheManager = spark.sharedState().cacheManager();
DataSourceV2Relation relation =
DataSourceV2Relation.create(table, Option.apply(tableCatalog), Option.apply(ident));
cacheManager.recacheByPlan(spark, relation);
}
protected Expression filterExpression(Identifier ident, String where) {
try {
String name = Spark3Util.quotedFullIdentifier(tableCatalog.name(), ident);
org.apache.spark.sql.catalyst.expressions.Expression expression =
SparkExpressionConverter.collectResolvedSparkExpression(spark, name, where);
return SparkExpressionConverter.convertToIcebergExpression(expression);
} catch (AnalysisException e) {
throw new IllegalArgumentException("Cannot parse predicates in where option: " + where, e);
}
}
protected InternalRow newInternalRow(Object... values) {
return new GenericInternalRow(values);
}
protected static class Result implements LocalScan {
private final StructType readSchema;
private final InternalRow[] rows;
public Result(StructType readSchema, InternalRow[] rows) {
this.readSchema = readSchema;
this.rows = rows;
}
@Override
public StructType readSchema() {View on GitHub (pinned to 86d9c8fc54)