abhigyanpatwari/GitNexus · critical

GraphEmitSink: streamed CSV writer(s) hit an IO error…

Error message

GraphEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error (disk-full / out-of-fds) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: ${first instanceof Error ? first.message : String(first)}

What it means

At finalize time, GraphEmitSink closes every per-relationship-pair CSV writer and collects each writer's recorded `poison` error (an async write failure captured during streaming). If any writer hit an IO error — the message names disk-full and out-of-file-descriptors — the persisted CSVs would be truncated, so the run is failed instead of COPYing a partial graph into LadybugDB. This is a data-integrity refusal: a partial graph would look like a valid index.

Solutions

  1. Free disk space (or point GitNexus storage at a larger volume) and re-run analyze — staging CSVs are regenerated on the next run
  2. Raise the file-descriptor limit before running: `ulimit -n 4096` (or the container/systemd equivalent) if the error is fd exhaustion
  3. Shrink the graph with .gitnexusignore excludes for directories you don't need indexed, reducing both CSV size and writer count

Example fix

# before
$ ulimit -n 256 && gitnexus analyze .
# after
$ ulimit -n 4096 && gitnexus analyze .
Defensive patterns

Strategy: retry

Validate before calling

import fs from 'node:fs';

// Pre-flight before a big analyze: disk headroom and fd budget
const stat = fs.statfsSync(storageDir);
const freeGB = (stat.bavail * stat.bsize) / 1024 ** 3;
if (freeGB < 5) throw new Error(`Only ${freeGB.toFixed(1)}GB free on ${storageDir} — free space before analyze`);
const [softFd] = [ulimit?]; // `ulimit -n` equivalent
// e.g. run `ulimit -n` and require >= a few thousand for many relationship pairs

Try / catch

try {
  const { totalRows } = await sink.finalize();
} catch (err) {
  if (err instanceof Error && err.message.includes('streamed CSV writer(s) hit an IO error')) {
    // free disk / raise `ulimit -n`, then re-run analyze from scratch — staging CSVs regenerate;
    // never continue with the partial CSVs
  }
  throw err;
}

Prevention

When it happens

Trigger: The storage volume fills up mid-emit while writers flush relationship CSVs; or the process exhausts its fd limit because the sink keeps one open CSV writer per relationship-type pair and a graph produces many distinct pairs.

Common situations: Very large repositories generating many relationship kinds (call/contains/imports/…) on CI runners with low `ulimit -n`; small tmpfs or shared disks that fill during big analyzes; other processes on the host consuming fds/disk concurrently.

Related errors


AI-assisted analysis of abhigyanpatwari/GitNexus@ac9a4e9abd (2026-08-20). Data as JSON: /api/errors/1b837f71f89c4a2d. Report an issue: GitHub.

Appendix: source

Thrown at gitnexus/src/core/lbug/graph-emit-sink.ts:524

  finalize(): GraphEmitManifest {
    if (this.finalized) throw new Error('GraphEmitSink.finalize() called twice');
    this.finalized = true;

    const errors: unknown[] = [];
    if (this.openFailure !== undefined) errors.push(this.openFailure);

    const relsByPair = new Map<string, { csvPath: string; rows: number }>();
    let totalRows = 0;
    for (const [pairKey, writer] of this.relWriters) {
      writer.close();
      if (writer.poison !== undefined) errors.push(writer.poison);
      relsByPair.set(pairKey, { csvPath: writer.csvPath, rows: writer.rows });
      totalRows += writer.rows;
    }

    if (errors.length > 0) {
      const first = errors[0];
      throw new Error(
        `GraphEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error ` +
          `(disk-full / out-of-fds) during the emit — the persisted graph would be ` +
          `truncated, so the run is failed rather than COPYing a partial CSV: ${
            first instanceof Error ? first.message : String(first)
          }`,
      );
    }

    return { relsByPair, totalRows, structuralRows: this.structuralRows };
  }

  /** Best-effort fd release for the error path — when the pipeline throws
   *  before {@link finalize} runs, the caller's `finally` calls this so the
   *  per-pair fds never leak. Idempotent with finalize via `finalized`. */
  close(): void {
    if (this.finalized) return;
    this.finalized = true;
    for (const writer of this.relWriters.values()) {

View on GitHub (pinned to ac9a4e9abd)