abhigyanpatwari/GitNexus · critical

PdgEmitSink: streamed CSV writer(s) hit an IO error…

Error message

PdgEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error (disk-full / out-of-fds) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: ${first.message}

What it means

PdgEmitSink.finalize (gitnexus/src/core/lbug/pdg-emit-sink.ts:239) closes every streamed CSV writer and checks for poison. Synchronous write faults (fs.writeSync ENOSPC/EIO) and writer-open failures (EMFILE — out of file descriptors) are swallowed by the emit loop's per-file try/catch, so finalize is the backstop: it fails the whole run instead of handing a truncated CSV to the bulk COPY.

Solutions

  1. Free disk space on the volume holding the GitNexus storage/csv directory and re-run the analyze
  2. If the message cites EMFILE, raise the process fd limit (ulimit -n 4096, or LimitNOFILE= in systemd) and retry
  3. Point GitNexus storage at a volume with headroom larger than the expected graph CSV set
  4. Re-run the analyze — the aborted run left the dirty flag so the next run rebuilds cleanly
Defensive patterns

Strategy: retry

Validate before calling

// before a streamed analyze, verify headroom and fd budget
import { statfs } from 'node:fs/promises';
const { bavail, bsize } = await statfs(csvDir);
const freeBytes = Number(bavail) * Number(bsize);
if (freeBytes < MIN_REQUIRED_BYTES) throw new Error(`only ${freeBytes} bytes free for CSV emit`);

Type guard

const isPdgSinkIoError = (e: unknown): boolean =>
  e instanceof Error && e.message.startsWith('PdgEmitSink:') && e.message.includes('IO error');

Try / catch

try {
  const manifest = sink.finalize();
} catch (e) {
  if (isPdgSinkIoError(e)) {
    // truncated CSVs must never be COPYed — free space / raise ulimit, then re-run the analyze
    throw new Error(`streamed emit failed on IO; freeing resources and re-running analyze is required`);
  }
  throw e;
}

Prevention

When it happens

Trigger: A streamed --pdg analyze on a large repo when the disk fills mid-emit, an fs.writeSync hits an IO error, or the process exhausts file descriptors opening the BasicBlock and per-pair rel CSV writers.

Common situations: Indexing monorepo or kernel-scale trees on small temp volumes; containers with low nofile limits; CI runners with quota-limited scratch space.

Related errors


AI-assisted analysis of abhigyanpatwari/GitNexus@ac9a4e9abd (2026-08-20). Data as JSON: /api/errors/b1a1bf3bc3be2a26. Report an issue: GitHub.

Appendix: source

Thrown at gitnexus/src/core/lbug/pdg-emit-sink.ts:239

    if (this.bbWriter !== undefined) {
      this.bbWriter.close();
      if (this.bbWriter.poison !== undefined) errors.push(this.bbWriter.poison);
      nodeFiles.set('BasicBlock' as NodeTableName, {
        csvPath: this.bbWriter.csvPath,
        rows: this.bbWriter.rows,
      });
    }

    const relsByPair = new Map<string, { csvPath: string; rows: number }>();
    for (const [pairKey, writer] of this.relWriters) {
      writer.close();
      if (writer.poison !== undefined) errors.push(writer.poison);
      relsByPair.set(pairKey, { csvPath: writer.csvPath, rows: writer.rows });
    }

    if (errors.length > 0) {
      const first = errors[0];
      throw new Error(
        `PdgEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error ` +
          `(disk-full / out-of-fds) during the emit — the persisted graph would ` +
          `be truncated, so the run is failed rather than COPYing a partial CSV: ${
            first instanceof Error ? first.message : String(first)
          }`,
      );
    }

    return { nodeFiles, relsByPair };
  }

  /**
   * Best-effort fd release for the error path — when a language pass throws
   * before {@link finalize} runs, the caller's `finally` calls this so the
   * BasicBlock + per-pair fds never leak. Idempotent with finalize via the
   * `finalized` flag; close errors are swallowed because the run is already
   * failing.
   */

View on GitHub (pinned to ac9a4e9abd)