influxdata/influxdb · error · CacheError

must pass a non-empty set of column ids

Error message

must pass a non-empty set of column ids

What it means

CacheError::EmptyColumnSet is returned when creating a distinct value cache with an empty set of column ids. A distinct cache must track at least one column, otherwise it has nothing to deduplicate, so the constructor rejects the request upfront.

Solutions

  1. Supply at least one column id when creating the cache
  2. Validate the column list length before invoking the cache creation API
  3. If columns are derived dynamically, return a user-facing error early when the derived list is empty

Example fix

// before
provider.new_cache(table_id, columns) // columns is empty
// after
assert!(!columns.is_empty(), "distinct cache requires at least one column");
provider.new_cache(table_id, columns)
Defensive patterns

Strategy: validation

Validate before calling

fn valid_column_set(col_ids: &[ColumnId]) -> bool {
    !col_ids.is_empty()
}

Try / catch

// Rust
match provider.new_cache(table_id, cols) {
    Err(ProviderError::Cache(CacheError::EmptyColumnSet)) => bail!("distinct cache requires at least one column id"),
    other => other,
}

Prevention

When it happens

Trigger: Calling the distinct cache creation API (CREATE DISTINCT CACHE or the provider's create_cache/new_cache call) with an empty column id list.

Common situations: SQL statements where the column list is built dynamically from empty user input, or JSON config where the 'columns' array was omitted/empty.

Understand the failure class

Background: "must not be empty", "cannot be empty" — required-field validation errors across open-source libraries — this error's family across 41 libraries.

Related errors


AI-assisted analysis of influxdata/influxdb@06200ef96b (2026-09-19). Data as JSON: /api/errors/4c0a6736fd130487. Report an issue: GitHub.

Appendix: source

Thrown at influxdb3_cache/src/distinct_cache/cache.rs:24

use anyhow::Context;
use arrow::{
    array::{ArrayRef, RecordBatch, StringViewBuilder},
    datatypes::{DataType, Field, SchemaBuilder, SchemaRef},
    error::ArrowError,
};
use indexmap::IndexMap;
use influxdb3_catalog::catalog::legacy;
use influxdb3_catalog::catalog::{MaxAge, MaxCardinality, TableDefinition};
use influxdb3_id::{ColumnId, ColumnIdentifier};
use influxdb3_wal::{FieldData, Row};
use iox_time::TimeProvider;
use observability_deps::tracing::debug;
use schema::{InfluxColumnType, InfluxFieldType};

#[derive(Debug, thiserror::Error)]
pub enum CacheError {
    #[error("must pass a non-empty set of column ids")]
    EmptyColumnSet,
    #[error(
        "cannot use a column of type {attempted} in a distinct value cache, only \
                    tags and string fields can be used"
    )]
    NonTagOrStringColumn { attempted: InfluxColumnType },
    #[error("cannot overwrite an an existing cache: {message}")]
    ConfigurationMismatch { message: String },
    #[error("unexpected error: {0}")]
    Unexpected(#[from] anyhow::Error),
}

/// A cache for storing distinct values for a set of columns in a table
#[derive(Debug)]
pub(crate) struct DistinctCache {
    time_provider: Arc<dyn TimeProvider>,
    /// The maximum number of unique value combinations in the cache
    max_cardinality: usize,

View on GitHub (pinned to 06200ef96b)