apache/iceberg · error · AnalysisException

CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS

CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS

Error message

CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS (viewName: %s, viewColumns: %s, dataColumns: %s)

What it means

When creating or replacing an Iceberg view, if the declared column list has FEWER columns than the view's query produces, Spark's structured-error CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS is thrown with the view name, declared columns, and query output columns. A view with an explicit column list must name every column of the query.

Solutions

  1. Expand the column list to exactly match the query output count, or drop the column list entirely.
  2. Replace SELECT * with an explicit projection so the arity is stable.
  3. Alias columns in the query (SELECT x AS a, y AS b, z AS c) and omit the column list.
  4. Check the 'dataColumns' value in the error message and add the missing view columns.

Example fix

// before
CREATE VIEW db.v (id, name) AS SELECT id, name, created_at FROM db.t
// after
CREATE VIEW db.v (id, name, created_at) AS SELECT id, name, created_at FROM db.t
Defensive patterns

Strategy: validation

Validate before calling

val out = spark.sql(query).schema.length
require(columns.isEmpty || columns.length == out, s"view column list (${columns.length}) must match query output ($out)")

Try / catch

try {
  spark.sql(createViewSql)
} catch {
  case e: AnalysisException if e.getMessage.contains("CREATE_VIEW_COLUMN_ARITY_MISMATCH") =>
    // fall back to letting the query define the columns
    spark.sql(createViewSql.replaceAll("\\([^)]*\\) AS", "AS"))
}

Prevention

When it happens

Trigger: Executing CREATE VIEW v (a, b) AS SELECT x, y, z FROM t — a column list shorter than the query's output arity — evaluated in CheckViews.verifyColumnCount when columns.nonEmpty and columns.length < query.output.length (the NOT_ENOUGH branch fires when fewer view columns than data columns).

Common situations: Hand-writing a column list for a wide SELECT and miscounting; SELECT * picking up newly added upstream columns after the view SQL was written; code generation that fixes the column list but lets the query drift.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/e9e84891f1e06b41. Report an issue: GitHub.

Appendix: source

Thrown at spark/v3.5/spark-extensions/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckViews.scala:77

            resolvedIdent.catalog.name() +: resolvedIdent.identifier.asMultipartIdentifier
          checkCyclicViewReference(viewIdent, query, Seq(viewIdent))
        }

      case AlterViewAs(ResolvedV2View(_, _), _, _) =>
        throw new AnalysisException(
          "ALTER VIEW <viewName> AS is not supported. Use CREATE OR REPLACE VIEW instead")

      case _ => // OK
    }
  }

  private def verifyColumnCount(
      ident: ResolvedIdentifier,
      columns: Seq[String],
      query: LogicalPlan): Unit = {
    if (columns.nonEmpty) {
      if (columns.length > query.output.length) {
        throw new AnalysisException(
          errorClass = "CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS",
          messageParameters = Map(
            "viewName" -> String.format("%s.%s", ident.catalog.name(), ident.identifier),
            "viewColumns" -> columns.mkString(", "),
            "dataColumns" -> query.output.map(c => c.name).mkString(", ")))
      } else if (columns.length < query.output.length) {
        throw new AnalysisException(
          errorClass = "CREATE_VIEW_COLUMN_ARITY_MISMATCH.TOO_MANY_DATA_COLUMNS",
          messageParameters = Map(
            "viewName" -> String.format("%s.%s", ident.catalog.name(), ident.identifier),
            "viewColumns" -> columns.mkString(", "),
            "dataColumns" -> query.output.map(c => c.name).mkString(", ")))
      }
    }
  }

  private def checkCyclicViewReference(
      viewIdent: Seq[String],

View on GitHub (pinned to 86d9c8fc54)