influxdata/influxdb · error · CodecError

failed to build parquet file

Error message

failed to build parquet file: {0}

What it means

CodecError::Writer, a #[from] ParquetError wrapper raised when the underlying Parquet writer fails while building the file (schema mismatch, I/O error on the sink, buffer exhaustion, etc.). Note per docs: a ResourcesExhausted error likely means the in-memory buffer used for parquet data grew too large.

Solutions

  1. If the message mentions ResourcesExhausted, reduce batch size / file size or increase the memory budget for parquet writing
  2. Verify the Arrow record batch schema exactly matches the schema given to the Parquet writer
  3. Check the sink (file/object store) health and permissions; surface the inner ParquetError details
  4. Retry the write if the failure was transient I/O

Example fix

// before: one giant batch exhausting the buffer
let batch = concat_batches(&schema, all_batches)?;
write(batch)?;
// after: write in bounded chunks
for chunk in all_batches.chunks(64) {
    let batch = concat_batches(&schema, chunk)?;
    write(batch)?;
}
Defensive patterns

Strategy: try-catch

Validate before calling

if batch.schema() != expected_schema {
    return Err(format!("batch schema {:?} != writer schema {:?}",
        batch.schema(), expected_schema));
}

Try / catch

match res {
    Err(CodecError::Writer(ParquetError::General(msg)))
        if msg.contains("ResourcesExhausted") =>
    {
        warn!("parquet buffer exhausted; reduce file size or raise memory budget")
    }
    other => other?,
}

Prevention

When it happens

Trigger: RecordBatch serialization hits ParquetError during arrow-to-parquet write: schema/Arrow type mismatch with the writer schema, sink write failures, or memory budget (ResourcesExhausted) exceeded.

Common situations: Writing very large single files exhausting the parquet buffer memory budget; Arrow schema not matching the schema the ParquetWriter was created with; underlying object-store/sink failing mid-write.

Understand the failure class

Background: "failed to write file", "Could not save figure", "Error saving remote file" — file write failed: causes and fixes across languages and libraries — this error's family across 38 libraries.

Related errors


AI-assisted analysis of influxdata/influxdb@06200ef96b (2026-09-19). Data as JSON: /api/errors/efed3d3ba813bc78. Report an issue: GitHub.

Appendix: source

Thrown at core/parquet_file/src/serialize.rs:80

    /// This would result in an empty file being uploaded to object store.
    ///
    /// [`RecordBatch`]: arrow::record_batch::RecordBatch
    #[error("no rows to serialise")]
    NoRows,

    /// A DataFusion error during the plan execution.
    ///
    /// Of note: a ResourcesExhaused error likely means the buffer
    /// used for parquet data became too large.
    #[error(transparent)]
    DataFusion(Box<DataFusionError>),

    /// Serialising the [`IoxMetadata`] to protobuf-encoded bytes failed.
    #[error("failed to serialize iox metadata: {0}")]
    MetadataSerialisation(#[from] prost::EncodeError),

    /// Writing the parquet file failed with the specified error.
    #[error("failed to build parquet file: {0}")]
    Writer(#[from] ParquetError),

    /// Attempting to clone a handle to the provided write sink failed.
    #[error("failed to obtain writer handle clone: {0}")]
    CloneSink(std::io::Error),
}

impl From<CodecError> for DataFusionError {
    fn from(value: CodecError) -> Self {
        match value {
            e @ (CodecError::NoRecordBatches
            | CodecError::NoRows
            | CodecError::MetadataSerialisation(_)
            | CodecError::CloneSink(_)) => Self::External(Box::new(e)),
            CodecError::Writer(e) => Self::ParquetError(Box::new(e)),
            CodecError::DataFusion(e) => *e,
        }
    }

View on GitHub (pinned to 06200ef96b)