apache/iceberg · error · IllegalArgumentException

Cannot translate Spark expression: $sparkExpression to data…

Error message

Cannot translate Spark expression: $sparkExpression to data source filter

What it means

When a Spark expression cannot be translated into any DataSource V2 filter by Spark's `translateFilterV2` (returns None), Iceberg cannot represent it for pushdown and throws this IllegalArgumentException. It guards the first stage of the two-stage expression→filter→Iceberg-expression conversion.

Solutions

  1. Rewrite the filter with built-in operators supported by V2 pushdown
  2. Materialize the subquery or join separately before filtering
  3. Compute the predicate value first, then pass it as a literal
  4. Keep unsupported predicates outside the pushed-down filter (post-scan filter)

Example fix

// before
.where("col IN (SELECT id FROM other)")
// after
val ids = spark.table("other").collect().map(_.getInt(0)).toSeq
.where(col("col").isin(ids: _*))
Defensive patterns

Strategy: try-catch

Validate before calling

// avoid subqueries/UDFs in pushed filters; precompute literals
val threshold = spark.table("other").agg(max("id")).first().getInt(0)
val df2 = df.filter(col("id") < threshold)

Try / catch

try { df.filter(expr) } catch { case e: IllegalArgumentException if e.getMessage.contains("Cannot translate Spark expression") => /* rewrite or compute post-scan */ }

Prevention

When it happens

Trigger: Calling `SparkExpressionConverter.convertToIcebergExpression` with a predicate Spark cannot lower to a V2 filter — e.g. subqueries, nondeterministic UDFs, unsupported expression types.

Common situations: Filters containing subqueries/EXISTS; deterministic-vs-nondeterministic UDF predicates; Spark version differences in translateFilterV2 coverage.

Understand the failure class

Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/8dedb3a2d1bc8865. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.2/spark/src/main/scala/org/apache/spark/sql/execution/datasources/SparkExpressionConverter.scala:49

object SparkExpressionConverter {

  def convertToIcebergExpression(
      sparkExpression: Expression): org.apache.iceberg.expressions.Expression = {
    // Currently, it is a double conversion as we are converting Spark expression to Spark predicate
    // and then converting Spark predicate to Iceberg expression.
    // But these two conversions already exist and well tested. So, we are going with this approach.
    DataSourceV2Strategy.translateFilterV2(sparkExpression) match {
      case Some(filter) =>
        val converted = SparkV2Filters.convert(filter)
        if (converted == null) {
          throw new IllegalArgumentException(
            s"Cannot convert Spark filter: $filter to Iceberg expression")
        }

        converted
      case _ =>
        throw new IllegalArgumentException(
          s"Cannot translate Spark expression: $sparkExpression to data source filter")
    }
  }

  @throws[IcebergAnalysisException]
  def collectResolvedSparkExpression(
      session: SparkSession,
      tableName: String,
      where: String): Expression = {
    val tableAttrs = session.table(tableName).queryExecution.analyzed.output
    val unresolvedExpression = session.sessionState.sqlParser.parseExpression(where)
    val filter = Filter(unresolvedExpression, DummyRelation(tableAttrs))
    val optimizedLogicalPlan = session.sessionState.executePlan(filter).optimizedPlan
    optimizedLogicalPlan
      .collectFirst {
        case filter: Filter => filter.condition
        case _: DummyRelation => Literal.TrueLiteral
        case _: LocalRelation => Literal.FalseLiteral

View on GitHub (pinned to 86d9c8fc54)