apache/iceberg · error · AnalysisException

CREATE_VIEW_COLUMN_ARITY_MISMATCH.TOO_MANY_DATA_COLUMNS

CREATE_VIEW_COLUMN_ARITY_MISMATCH.TOO_MANY_DATA_COLUMNS

Error message

CREATE_VIEW_COLUMN_ARITY_MISMATCH.TOO_MANY_DATA_COLUMNS (viewName: %s, viewColumns: %s, dataColumns: %s)

What it means

Thrown by Iceberg's CheckViews rule when CREATE VIEW lists a different number of column names than the view's query produces. Specifically, the query outputs MORE columns than the explicit column list provides (columns.length < query.output.length). Spark/Iceberg requires an exact 1:1 mapping between declared view columns and query output columns.

Solutions

  1. Add or remove column aliases so the number of declared view columns exactly equals the number of SELECT output columns
  2. Drop the explicit column list entirely and let the view inherit the query's column names
  3. If the source table gained columns unintentionally, replace SELECT * with an explicit column list matching the intended view columns

Example fix

// before
CREATE VIEW catalog.db.v (id) AS SELECT id, name, ts FROM catalog.db.t
// after
CREATE VIEW catalog.db.v (id, name, ts) AS SELECT id, name, ts FROM catalog.db.t
Defensive patterns

Strategy: validation

Validate before calling

val queryOutput = df.queryExecution.analyzed.output.map(_.name)
require(viewColumns.length == queryOutput.length,
  s"view declares ${viewColumns.length} columns but query outputs ${queryOutput.length}")

Try / catch

try { spark.sql(createViewSql) } catch { case e: AnalysisException if e.getErrorClass == "CREATE_VIEW_COLUMN_ARITY_MISMATCH" => // fix column list }

Prevention

When it happens

Trigger: Executing CREATE VIEW catalog.db.v (col1) AS SELECT a, b, c FROM tbl — the SELECT outputs 3 columns but only 1 view column name was declared.

Common situations: Typos or omissions when renaming view columns; copying a view definition across table versions where the source table gained columns; hand-written DDL where the SELECT list was widened later.

Understand the failure class

Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/6af940dd8daff7b4. Report an issue: GitHub.

Appendix: source

Thrown at spark/v3.5/spark-extensions/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckViews.scala:84

      case _ => // OK
    }
  }

  private def verifyColumnCount(
      ident: ResolvedIdentifier,
      columns: Seq[String],
      query: LogicalPlan): Unit = {
    if (columns.nonEmpty) {
      if (columns.length > query.output.length) {
        throw new AnalysisException(
          errorClass = "CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS",
          messageParameters = Map(
            "viewName" -> String.format("%s.%s", ident.catalog.name(), ident.identifier),
            "viewColumns" -> columns.mkString(", "),
            "dataColumns" -> query.output.map(c => c.name).mkString(", ")))
      } else if (columns.length < query.output.length) {
        throw new AnalysisException(
          errorClass = "CREATE_VIEW_COLUMN_ARITY_MISMATCH.TOO_MANY_DATA_COLUMNS",
          messageParameters = Map(
            "viewName" -> String.format("%s.%s", ident.catalog.name(), ident.identifier),
            "viewColumns" -> columns.mkString(", "),
            "dataColumns" -> query.output.map(c => c.name).mkString(", ")))
      }
    }
  }

  private def checkCyclicViewReference(
      viewIdent: Seq[String],
      plan: LogicalPlan,
      cyclePath: Seq[Seq[String]]): Unit = {
    plan match {
      case sub @ SubqueryAlias(_, Project(_, _)) =>
        val currentViewIdent: Seq[String] = sub.identifier.qualifier :+ sub.identifier.name
        checkIfRecursiveView(viewIdent, currentViewIdent, cyclePath, sub.children)
      case v1View: View =>

View on GitHub (pinned to 86d9c8fc54)