influxdata/influxdb · error · CodecError
failed to build parquet file
Error message
failed to build parquet file: {0} What it means
CodecError::Writer, a #[from] ParquetError wrapper raised when the underlying Parquet writer fails while building the file (schema mismatch, I/O error on the sink, buffer exhaustion, etc.). Note per docs: a ResourcesExhausted error likely means the in-memory buffer used for parquet data grew too large.
Solutions
- If the message mentions ResourcesExhausted, reduce batch size / file size or increase the memory budget for parquet writing
- Verify the Arrow record batch schema exactly matches the schema given to the Parquet writer
- Check the sink (file/object store) health and permissions; surface the inner ParquetError details
- Retry the write if the failure was transient I/O
Example fix
// before: one giant batch exhausting the buffer
let batch = concat_batches(&schema, all_batches)?;
write(batch)?;
// after: write in bounded chunks
for chunk in all_batches.chunks(64) {
let batch = concat_batches(&schema, chunk)?;
write(batch)?;
} Defensive patterns
Strategy: try-catch
Validate before calling
if batch.schema() != expected_schema {
return Err(format!("batch schema {:?} != writer schema {:?}",
batch.schema(), expected_schema));
} Try / catch
match res {
Err(CodecError::Writer(ParquetError::General(msg)))
if msg.contains("ResourcesExhausted") =>
{
warn!("parquet buffer exhausted; reduce file size or raise memory budget")
}
other => other?,
} Prevention
- Bound batch/file sizes so the parquet buffer cannot exhaust memory
- Assert record batch schema equals the writer schema before writing
- Monitor sink (object store/disk) health for mid-write failures
When it happens
Trigger: RecordBatch serialization hits ParquetError during arrow-to-parquet write: schema/Arrow type mismatch with the writer schema, sink write failures, or memory budget (ResourcesExhausted) exceeded.
Common situations: Writing very large single files exhausting the parquet buffer memory budget; Arrow schema not matching the schema the ParquetWriter was created with; underlying object-store/sink failing mid-write.
Understand the failure class
Background: "failed to write file", "Could not save figure", "Error saving remote file" — file write failed: causes and fixes across languages and libraries — this error's family across 38 libraries.
Related errors
- failed to allocate buffer while writing parquet
- failed to obtain writer handle clone
- failed to write parquet file
- cannot write parquet to a terminal, use `--output
- Could not convert Parquet schema to Arrow schema
AI-assisted analysis of influxdata/influxdb@06200ef96b (2026-09-19).
Data as JSON: /api/errors/efed3d3ba813bc78.
Report an issue: GitHub.
Appendix: source
Thrown at core/parquet_file/src/serialize.rs:80
/// This would result in an empty file being uploaded to object store.
///
/// [`RecordBatch`]: arrow::record_batch::RecordBatch
#[error("no rows to serialise")]
NoRows,
/// A DataFusion error during the plan execution.
///
/// Of note: a ResourcesExhaused error likely means the buffer
/// used for parquet data became too large.
#[error(transparent)]
DataFusion(Box<DataFusionError>),
/// Serialising the [`IoxMetadata`] to protobuf-encoded bytes failed.
#[error("failed to serialize iox metadata: {0}")]
MetadataSerialisation(#[from] prost::EncodeError),
/// Writing the parquet file failed with the specified error.
#[error("failed to build parquet file: {0}")]
Writer(#[from] ParquetError),
/// Attempting to clone a handle to the provided write sink failed.
#[error("failed to obtain writer handle clone: {0}")]
CloneSink(std::io::Error),
}
impl From<CodecError> for DataFusionError {
fn from(value: CodecError) -> Self {
match value {
e @ (CodecError::NoRecordBatches
| CodecError::NoRows
| CodecError::MetadataSerialisation(_)
| CodecError::CloneSink(_)) => Self::External(Box::new(e)),
CodecError::Writer(e) => Self::ParquetError(Box::new(e)),
CodecError::DataFusion(e) => *e,
}
}View on GitHub (pinned to 06200ef96b)