influxdata/influxdb · error

Key column, , of last_batch has no data

Error message

Key column, {name}, of last_batch has no data

What it means

During deduplication (deduplicate / last_batch_with_no_same_sort_key), the sort-key column is taken from the last batch to seed the algorithm. If that column exists in the schema but its RecordBatch array is empty (zero rows), the algorithm cannot establish ordering state and panics naming the key column.

Solutions

  1. Filter out empty RecordBatches before calling deduplicate.
  2. Ensure the last batch passed has at least one row, or reorder so a non-empty batch is last.
  3. Check upstream operators (filter, limit) that may produce zero-row batches and drop them.
  4. If you maintain a fork, return an error instead of panicking on the empty key column.

Example fix

// before
deduplicate(&batches, &sort_key)?;
// after
let batches: Vec<_> = batches.into_iter().filter(|b| b.num_rows() > 0).collect();
deduplicate(&batches, &sort_key)?;
Defensive patterns

Strategy: validation

Validate before calling

let batches: Vec<RecordBatch> = batches.into_iter().filter(|b| b.num_rows() > 0).collect();
assert!(!batches.is_empty(), "no non-empty batches to deduplicate");

Try / catch

// deduplicate panics; filter empty batches before invoking
let kept: Vec<_> = batches.into_iter().filter(|b| b.num_rows() > 0).collect();
deduplicate(&kept, &sort_key)?;

Prevention

When it happens

Trigger: Calling deduplicate with a batch list whose final batch has zero rows while the schema declares a sort key on that column.

Common situations: Feeding empty trailing batches produced by upstream filters/limits; concatenated chunk scans where the last chunk returned no rows; tests constructing batches with one empty record batch.

Understand the failure class

Background: "must not be empty", "cannot be empty" — required-field validation errors across open-source libraries — this error's family across 41 libraries.

Related errors


AI-assisted analysis of influxdata/influxdb@06200ef96b (2026-09-19). Data as JSON: /api/errors/6f2cb56c5d84a1ca. Report an issue: GitHub.

Appendix: source

Thrown at core/iox_query/src/provider/deduplicate/algo.rs:122

        // Take the previous batch, if any, out of it storage self.last_batch
        if let Some(last_batch) = self.last_batch.take() {
            // Build sorted columns for last_batch and current one
            let schema = last_batch.schema();
            // is_sort_key[col_idx] = true if it is present in sort keys
            let mut is_sort_key: Vec<bool> = vec![false; last_batch.columns().len()];
            let last_batch_key_columns = self
                .sort_keys
                .iter()
                .map(|skey| {
                    // figure out the index of the key columns
                    let name = get_col_name(skey.expr.as_ref());
                    let index = schema.index_of(name).unwrap();
                    is_sort_key[index] = true;

                    // Key column of last_batch of this index
                    let last_batch_array = last_batch.column(index);
                    if last_batch_array.is_empty() {
                        panic!("Key column, {name}, of last_batch has no data");
                    }
                    DedupSortColumn {
                        array: last_batch_array,
                        options: skey.options,
                    }
                })
                .collect::<Vec<_>>();

            // Build sorted columns for current batch
            // Schema of both batches are the same
            let batch_key_columns = self
                .sort_keys
                .iter()
                .map(|skey| {
                    // figure out the index of the key columns
                    let name = get_col_name(skey.expr.as_ref());
                    let index = schema.index_of(name).unwrap();

View on GitHub (pinned to 06200ef96b)