tursodatabase/turso · error · ValueError

profiled timings must not be used for benchmark comparisons

Error message

profiled timings must not be used for benchmark comparisons

What it means

read_file() validates every CSV row produced by the FTS benchmark harness before plotting. If a row has profiled=true, its timings were collected under a profiler and are not comparable to clean runs, so the script refuses to plot them. This protects users from publishing skewed benchmark comparisons.

Solutions

  1. Regenerate the benchmark CSV with profiling disabled so the 'profiled' column is false for all rows
  2. Remove profiled rows from the CSV, keeping only rows where profiled=false
  3. Keep profiled data separate and plot it with a different (profiling-aware) tool

Example fix

// before (CSV row)
engine,turso,...,profiled=true
// after
engine,turso,...,profiled=false
Defensive patterns

Strategy: validation

Validate before calling

def validate_not_profiled(row):
    if row.get("profiled", "false") != "false":
        raise ValueError("profiled timings must not be used for benchmark comparisons")

Type guard

def is_clean_timing(row) -> bool:
    return row.get("profiled", "false") == "false"

Try / catch

try:
    configuration, runs = read_runs(path, percentile)
except ValueError as e:
    print(f"skipping {path}: {e}")

Prevention

When it happens

Trigger: A results CSV whose 'profiled' column is anything other than 'false' is passed to read_runs()/read_file(), e.g. plotting output collected with the profiler enabled.

Common situations: Developer forgot to disable profiling when generating benchmark results; a mixed CSV contains both profiled and clean rows; a stale results file from a profiling session was reused for plotting.

Understand the failure class

Background: "Invalid value" and "allowed values are" config errors: what your library rejected and how to fix it — this error's family across 41 libraries.

Related errors


AI-assisted analysis of tursodatabase/turso@8d4a589f8d (2026-09-20). Data as JSON: /api/errors/805a917e1ace8d76. Report an issue: GitHub.

Appendix: source

Thrown at perf/fts/plot/plot-fts.py:100

                raise ValueError(f"{path}: each run must contain the same series and all six query cases")
    if configuration is None:
        raise ValueError("no results found")
    configuration = dict(zip(CONFIGURATION, configuration))
    configuration["percentile"] = percentile
    return configuration, samples


def read_file(path, percentile, samples):
    configuration = None
    runs = {}
    with path.open(newline="") as stream:
        reader = csv.DictReader(stream)
        required = set(CONFIGURATION) | {"engine", "mode", "run", "query", "queries"}
        if not required <= set(reader.fieldnames or []):
            raise ValueError(f"{path}: missing benchmark columns")
        for row in reader:
            if row.get("profiled", "false") != "false":
                raise ValueError("profiled timings must not be used for benchmark comparisons")
            current = tuple(row[key] for key in CONFIGURATION)
            if configuration is None:
                configuration = current
            if configuration != current:
                raise ValueError("plot only one configuration at a time")
            key = (row["engine"], row["mode"])
            if key not in SERIES:
                raise ValueError(f"unsupported engine/mode: {key}")
            series = samples.setdefault(key, {query: [] for query in QUERIES})
            query = row["query"]
            seen = runs.setdefault(row["run"], {}).setdefault(key, set())
            if query not in QUERIES or query in seen:
                raise ValueError(f"{path}: unknown or duplicate query {query}")
            seen.add(query)
            series[query].append(read_measurement(row, percentile))
    if not runs:
        raise ValueError(f"{path}: no results found")
    return configuration, runs

View on GitHub (pinned to 8d4a589f8d)