{"record":{"id":"d97a33bbaecfa668","repo":"tursodatabase/turso","slug":"path-missing-benchmark-columns","errorCode":null,"errorMessage":"{path}: missing benchmark columns","messagePattern":"(.+?): missing benchmark columns","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"perf/fts/plot/plot-fts.py","lineNumber":97,"sourceCode":"            if expected_series is None:\n                expected_series = set(run)\n            if set(run) != expected_series or any(seen != set(QUERIES) for seen in run.values()):\n                raise ValueError(f\"{path}: each run must contain the same series and all six query cases\")\n    if configuration is None:\n        raise ValueError(\"no results found\")\n    configuration = dict(zip(CONFIGURATION, configuration))\n    configuration[\"percentile\"] = percentile\n    return configuration, samples\n\n\ndef read_file(path, percentile, samples):\n    configuration = None\n    runs = {}\n    with path.open(newline=\"\") as stream:\n        reader = csv.DictReader(stream)\n        required = set(CONFIGURATION) | {\"engine\", \"mode\", \"run\", \"query\", \"queries\"}\n        if not required <= set(reader.fieldnames or []):\n            raise ValueError(f\"{path}: missing benchmark columns\")\n        for row in reader:\n            if row.get(\"profiled\", \"false\") != \"false\":\n                raise ValueError(\"profiled timings must not be used for benchmark comparisons\")\n            current = tuple(row[key] for key in CONFIGURATION)\n            if configuration is None:\n                configuration = current\n            if configuration != current:\n                raise ValueError(\"plot only one configuration at a time\")\n            key = (row[\"engine\"], row[\"mode\"])\n            if key not in SERIES:\n                raise ValueError(f\"unsupported engine/mode: {key}\")\n            series = samples.setdefault(key, {query: [] for query in QUERIES})\n            query = row[\"query\"]\n            seen = runs.setdefault(row[\"run\"], {}).setdefault(key, set())\n            if query not in QUERIES or query in seen:\n                raise ValueError(f\"{path}: unknown or duplicate query {query}\")\n            seen.add(query)\n            series[query].append(read_measurement(row, percentile))","sourceCodeStart":79,"sourceCodeEnd":115,"githubUrl":"https://github.com/tursodatabase/turso/blob/8d4a589f8d13ac184700d2a8f724f27e1995be3b/perf/fts/plot/plot-fts.py#L79-L115","documentation":"read_file parses a benchmark results CSV with csv.DictReader and validates that every required column is present: the CONFIGURATION columns plus engine, mode, run, query, and queries. It raises this ValueError naming the file when the header is missing any required column, because downstream grouping and plotting depend on them.","triggerScenarios":"Calling read_file (via read_runs) on a CSV whose header lacks any of the required columns, e.g. an older results format without a 'queries' column or a renamed 'engine' column.","commonSituations":"Plotting CSVs produced by an older benchmark script version; manually edited or truncated headers; mixing in unrelated CSVs that happen to sit in the results directory.","solutions":["Regenerate the CSV with the current benchmark script so the header includes all required columns (CONFIGURATION fields, engine, mode, run, query, queries).","Add the missing column(s) to the CSV header and fill the values if the data is recoverable.","Point plot-fts.py at the correct, current results file instead of a stale export."],"exampleFix":"# before (stale header)\nengine,mode,run,query,elapsed_ms\n# after (add required columns)\nsample_budget,fixture,engine,mode,run,query,queries,elapsed_ms","handlingStrategy":"validation","validationCode":"import csv\nwith open(path, newline='') as f:\n    header = set(csv.DictReader(f).fieldnames or [])\nrequired = set(CONFIGURATION) | {'engine', 'mode', 'run', 'query', 'queries'}\nassert required <= header, f'{path}: missing columns {sorted(required - header)}'","typeGuard":null,"tryCatchPattern":"try:\n    config, runs = read_file(path, percentile, samples)\nexcept ValueError as e:\n    if 'missing benchmark columns' in str(e):\n        print(f'Fix: {e} - regenerate the CSV with the current benchmark script')\n    else:\n        raise","preventionTips":["Always generate result CSVs with the current version of the benchmark script.","Check the header against the required column set before feeding files to plot-fts.py.","Do not hand-edit or rename columns in benchmark output CSVs."],"tags":["python","csv","input-validation"],"backgroundTag":"schema-validation-failed","analyzedSha":"8d4a589f8d13ac184700d2a8f724f27e1995be3b","analyzedAt":"2026-09-20T13:18:14.658Z","contentChangedAt":"2026-09-20T13:18:14.658Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}