{"record":{"id":"2c978650a1bb8175","repo":"apache/iceberg","slug":"the-average-length-of-the-rows-appears-to-be-zero","errorCode":null,"errorMessage":"The average length of the rows appears to be zero.","messagePattern":"The average length of the rows appears to be zero\\.","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"orc/src/main/java/org/apache/iceberg/orc/OrcFileAppender.java","lineNumber":78,"sourceCode":"      Schema schema,\n      OutputFile file,\n      BiFunction<Schema, TypeDescription, OrcRowWriter<?>> createWriterFunc,\n      Configuration conf,\n      Map<String, byte[]> metadata,\n      int batchSize,\n      MetricsConfig metricsConfig) {\n    this.file = file;\n    this.batchSize = batchSize;\n    this.metricsConfig = metricsConfig;\n\n    TypeDescription orcSchema = ORCSchemaUtil.convert(schema);\n\n    this.avgRowByteSize =\n        OrcSchemaVisitor.visitSchema(orcSchema, new EstimateOrcAvgWidthVisitor()).stream()\n            .reduce(Integer::sum)\n            .orElse(0);\n    if (avgRowByteSize == 0) {\n      LOG.warn(\"The average length of the rows appears to be zero.\");\n    }\n\n    this.batch = orcSchema.createRowBatch(this.batchSize);\n\n    OrcFile.WriterOptions options = OrcFile.writerOptions(conf).useUTCTimestamp(true);\n    if (file instanceof HadoopOutputFile) {\n      options.fileSystem(((HadoopOutputFile) file).getFileSystem());\n    }\n    options.setSchema(orcSchema);\n    this.writer = ORC.newFileWriter(file, options, metadata);\n    this.valueWriter = newOrcRowWriter(schema, orcSchema, createWriterFunc);\n  }\n\n  @Override\n  public void add(D datum) {\n    try {\n      valueWriter.write(datum, batch);\n      if (batch.size == this.batchSize) {","sourceCodeStart":60,"sourceCodeEnd":96,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/orc/src/main/java/org/apache/iceberg/orc/OrcFileAppender.java#L60-L96","documentation":"During OrcFileAppender construction, the estimated average row byte size computed by EstimateOrcAvgWidthVisitor summed to zero. The appender only logs a warning (writing continues), but memory estimation for the writer will be inaccurate, potentially affecting row-group sizing decisions.","triggerScenarios":"Creating an ORC writer for a schema where all columns' estimated widths are 0 — e.g. a schema with only columns the width visitor does not estimate (certain nested/complex types), or an effectively empty/odd orcSchema.","commonSituations":"Writing tables whose schema contains only types unsupported by the width estimator; writing empty schemas; misconfigured ORC schema conversion producing an all-unknown schema.","solutions":["Inspect the ORC schema being written; ensure it contains estimable primitive columns","Upgrade Iceberg — the width estimator may have been extended for more types","Ignore the warning if data is written correctly; it only affects size estimates"],"exampleFix":null,"handlingStrategy":"validation","validationCode":"if (schema.columns().stream().allMatch(c -> c.type().isNestedType())) { LOG.warn(\"ORC avg-width estimate may be zero for this schema\"); }","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Include at least primitive-typed columns in written schemas where possible","Verify written ORC files open/read correctly when this warning appears","Upgrade Iceberg to benefit from improved width estimators"],"tags":["orc","writer","size-estimation"],"backgroundTag":"internal-invariant-violation","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}