tursodatabase/turso · error · ValueError
: missing benchmark columns
Error message
{path}: missing benchmark columns What it means
read_file parses a benchmark results CSV with csv.DictReader and validates that every required column is present: the CONFIGURATION columns plus engine, mode, run, query, and queries. It raises this ValueError naming the file when the header is missing any required column, because downstream grouping and plotting depend on them.
Solutions
- Regenerate the CSV with the current benchmark script so the header includes all required columns (CONFIGURATION fields, engine, mode, run, query, queries).
- Add the missing column(s) to the CSV header and fill the values if the data is recoverable.
- Point plot-fts.py at the correct, current results file instead of a stale export.
Example fix
# before (stale header) engine,mode,run,query,elapsed_ms # after (add required columns) sample_budget,fixture,engine,mode,run,query,queries,elapsed_ms
Defensive patterns
Strategy: validation
Validate before calling
import csv
with open(path, newline='') as f:
header = set(csv.DictReader(f).fieldnames or [])
required = set(CONFIGURATION) | {'engine', 'mode', 'run', 'query', 'queries'}
assert required <= header, f'{path}: missing columns {sorted(required - header)}' Try / catch
try:
config, runs = read_file(path, percentile, samples)
except ValueError as e:
if 'missing benchmark columns' in str(e):
print(f'Fix: {e} - regenerate the CSV with the current benchmark script')
else:
raise Prevention
- Always generate result CSVs with the current version of the benchmark script.
- Check the header against the required column set before feeding files to plot-fts.py.
- Do not hand-edit or rename columns in benchmark output CSVs.
When it happens
Trigger: Calling read_file (via read_runs) on a CSV whose header lacks any of the required columns, e.g. an older results format without a 'queries' column or a renamed 'engine' column.
Common situations: Plotting CSVs produced by an older benchmark script version; manually edited or truncated headers; mixing in unrelated CSVs that happen to sit in the results directory.
Understand the failure class
Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.
Related errors
- each sweep file must have a distinct positive connection…
- plot only one configuration at a time
- has columns , has
- sweep files must match benchmark, sample budget, fixture…
- batch mode must be one of
AI-assisted analysis of tursodatabase/turso@8d4a589f8d (2026-09-20).
Data as JSON: /api/errors/d97a33bbaecfa668.
Report an issue: GitHub.
Appendix: source
Thrown at perf/fts/plot/plot-fts.py:97
if expected_series is None:
expected_series = set(run)
if set(run) != expected_series or any(seen != set(QUERIES) for seen in run.values()):
raise ValueError(f"{path}: each run must contain the same series and all six query cases")
if configuration is None:
raise ValueError("no results found")
configuration = dict(zip(CONFIGURATION, configuration))
configuration["percentile"] = percentile
return configuration, samples
def read_file(path, percentile, samples):
configuration = None
runs = {}
with path.open(newline="") as stream:
reader = csv.DictReader(stream)
required = set(CONFIGURATION) | {"engine", "mode", "run", "query", "queries"}
if not required <= set(reader.fieldnames or []):
raise ValueError(f"{path}: missing benchmark columns")
for row in reader:
if row.get("profiled", "false") != "false":
raise ValueError("profiled timings must not be used for benchmark comparisons")
current = tuple(row[key] for key in CONFIGURATION)
if configuration is None:
configuration = current
if configuration != current:
raise ValueError("plot only one configuration at a time")
key = (row["engine"], row["mode"])
if key not in SERIES:
raise ValueError(f"unsupported engine/mode: {key}")
series = samples.setdefault(key, {query: [] for query in QUERIES})
query = row["query"]
seen = runs.setdefault(row["run"], {}).setdefault(key, set())
if query not in QUERIES or query in seen:
raise ValueError(f"{path}: unknown or duplicate query {query}")
seen.add(query)
series[query].append(read_measurement(row, percentile))View on GitHub (pinned to 8d4a589f8d)