tursodatabase/turso · error · ValueError

: missing benchmark columns

Error message

{path}: missing benchmark columns

What it means

read_file parses a benchmark results CSV with csv.DictReader and validates that every required column is present: the CONFIGURATION columns plus engine, mode, run, query, and queries. It raises this ValueError naming the file when the header is missing any required column, because downstream grouping and plotting depend on them.

Solutions

  1. Regenerate the CSV with the current benchmark script so the header includes all required columns (CONFIGURATION fields, engine, mode, run, query, queries).
  2. Add the missing column(s) to the CSV header and fill the values if the data is recoverable.
  3. Point plot-fts.py at the correct, current results file instead of a stale export.

Example fix

# before (stale header)
engine,mode,run,query,elapsed_ms
# after (add required columns)
sample_budget,fixture,engine,mode,run,query,queries,elapsed_ms
Defensive patterns

Strategy: validation

Validate before calling

import csv
with open(path, newline='') as f:
    header = set(csv.DictReader(f).fieldnames or [])
required = set(CONFIGURATION) | {'engine', 'mode', 'run', 'query', 'queries'}
assert required <= header, f'{path}: missing columns {sorted(required - header)}'

Try / catch

try:
    config, runs = read_file(path, percentile, samples)
except ValueError as e:
    if 'missing benchmark columns' in str(e):
        print(f'Fix: {e} - regenerate the CSV with the current benchmark script')
    else:
        raise

Prevention

When it happens

Trigger: Calling read_file (via read_runs) on a CSV whose header lacks any of the required columns, e.g. an older results format without a 'queries' column or a renamed 'engine' column.

Common situations: Plotting CSVs produced by an older benchmark script version; manually edited or truncated headers; mixing in unrelated CSVs that happen to sit in the results directory.

Understand the failure class

Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.

Related errors


AI-assisted analysis of tursodatabase/turso@8d4a589f8d (2026-09-20). Data as JSON: /api/errors/d97a33bbaecfa668. Report an issue: GitHub.

Appendix: source

Thrown at perf/fts/plot/plot-fts.py:97

            if expected_series is None:
                expected_series = set(run)
            if set(run) != expected_series or any(seen != set(QUERIES) for seen in run.values()):
                raise ValueError(f"{path}: each run must contain the same series and all six query cases")
    if configuration is None:
        raise ValueError("no results found")
    configuration = dict(zip(CONFIGURATION, configuration))
    configuration["percentile"] = percentile
    return configuration, samples


def read_file(path, percentile, samples):
    configuration = None
    runs = {}
    with path.open(newline="") as stream:
        reader = csv.DictReader(stream)
        required = set(CONFIGURATION) | {"engine", "mode", "run", "query", "queries"}
        if not required <= set(reader.fieldnames or []):
            raise ValueError(f"{path}: missing benchmark columns")
        for row in reader:
            if row.get("profiled", "false") != "false":
                raise ValueError("profiled timings must not be used for benchmark comparisons")
            current = tuple(row[key] for key in CONFIGURATION)
            if configuration is None:
                configuration = current
            if configuration != current:
                raise ValueError("plot only one configuration at a time")
            key = (row["engine"], row["mode"])
            if key not in SERIES:
                raise ValueError(f"unsupported engine/mode: {key}")
            series = samples.setdefault(key, {query: [] for query in QUERIES})
            query = row["query"]
            seen = runs.setdefault(row["run"], {}).setdefault(key, set())
            if query not in QUERIES or query in seen:
                raise ValueError(f"{path}: unknown or duplicate query {query}")
            seen.add(query)
            series[query].append(read_measurement(row, percentile))

View on GitHub (pinned to 8d4a589f8d)