{"record":{"id":"e9e84891f1e06b41","repo":"apache/iceberg","slug":"create-view-column-arity-mismatch-not-enough-data","errorCode":"CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS","errorMessage":"CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS (viewName: %s, viewColumns: %s, dataColumns: %s)","messagePattern":"CREATE_VIEW_COLUMN_ARITY_MISMATCH\\.NOT_ENOUGH_DATA_COLUMNS \\(viewName: (.+?), viewColumns: (.+?), dataColumns: (.+?)\\)","errorType":"error_code","errorClass":"AnalysisException","httpStatus":null,"severity":"error","filePath":"spark/v3.5/spark-extensions/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckViews.scala","lineNumber":77,"sourceCode":"            resolvedIdent.catalog.name() +: resolvedIdent.identifier.asMultipartIdentifier\n          checkCyclicViewReference(viewIdent, query, Seq(viewIdent))\n        }\n\n      case AlterViewAs(ResolvedV2View(_, _), _, _) =>\n        throw new AnalysisException(\n          \"ALTER VIEW <viewName> AS is not supported. Use CREATE OR REPLACE VIEW instead\")\n\n      case _ => // OK\n    }\n  }\n\n  private def verifyColumnCount(\n      ident: ResolvedIdentifier,\n      columns: Seq[String],\n      query: LogicalPlan): Unit = {\n    if (columns.nonEmpty) {\n      if (columns.length > query.output.length) {\n        throw new AnalysisException(\n          errorClass = \"CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS\",\n          messageParameters = Map(\n            \"viewName\" -> String.format(\"%s.%s\", ident.catalog.name(), ident.identifier),\n            \"viewColumns\" -> columns.mkString(\", \"),\n            \"dataColumns\" -> query.output.map(c => c.name).mkString(\", \")))\n      } else if (columns.length < query.output.length) {\n        throw new AnalysisException(\n          errorClass = \"CREATE_VIEW_COLUMN_ARITY_MISMATCH.TOO_MANY_DATA_COLUMNS\",\n          messageParameters = Map(\n            \"viewName\" -> String.format(\"%s.%s\", ident.catalog.name(), ident.identifier),\n            \"viewColumns\" -> columns.mkString(\", \"),\n            \"dataColumns\" -> query.output.map(c => c.name).mkString(\", \")))\n      }\n    }\n  }\n\n  private def checkCyclicViewReference(\n      viewIdent: Seq[String],","sourceCodeStart":59,"sourceCodeEnd":95,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v3.5/spark-extensions/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckViews.scala#L59-L95","documentation":"When creating or replacing an Iceberg view, if the declared column list has FEWER columns than the view's query produces, Spark's structured-error CREATE_VIEW_COLUMN_ARITY_MISMATCH.NOT_ENOUGH_DATA_COLUMNS is thrown with the view name, declared columns, and query output columns. A view with an explicit column list must name every column of the query.","triggerScenarios":"Executing CREATE VIEW v (a, b) AS SELECT x, y, z FROM t — a column list shorter than the query's output arity — evaluated in CheckViews.verifyColumnCount when columns.nonEmpty and columns.length < query.output.length (the NOT_ENOUGH branch fires when fewer view columns than data columns).","commonSituations":"Hand-writing a column list for a wide SELECT and miscounting; SELECT * picking up newly added upstream columns after the view SQL was written; code generation that fixes the column list but lets the query drift.","solutions":["Expand the column list to exactly match the query output count, or drop the column list entirely.","Replace SELECT * with an explicit projection so the arity is stable.","Alias columns in the query (SELECT x AS a, y AS b, z AS c) and omit the column list.","Check the 'dataColumns' value in the error message and add the missing view columns."],"exampleFix":"// before\nCREATE VIEW db.v (id, name) AS SELECT id, name, created_at FROM db.t\n// after\nCREATE VIEW db.v (id, name, created_at) AS SELECT id, name, created_at FROM db.t","handlingStrategy":"validation","validationCode":"val out = spark.sql(query).schema.length\nrequire(columns.isEmpty || columns.length == out, s\"view column list (${columns.length}) must match query output ($out)\")","typeGuard":null,"tryCatchPattern":"try {\n  spark.sql(createViewSql)\n} catch {\n  case e: AnalysisException if e.getMessage.contains(\"CREATE_VIEW_COLUMN_ARITY_MISMATCH\") =>\n    // fall back to letting the query define the columns\n    spark.sql(createViewSql.replaceAll(\"\\\\([^)]*\\\\) AS\", \"AS\"))\n}","preventionTips":["Avoid explicit column lists; alias columns inside the SELECT instead.","Never pair a column list with SELECT * — upstream schema changes will break it.","Count query output columns against the declared list before running CREATE VIEW."],"tags":["spark","sql","views","arity-mismatch"],"backgroundTag":"shape-mismatch","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}