abhigyanpatwari/GitNexus · critical
PdgEmitSink: streamed CSV writer(s) hit an IO error…
Error message
PdgEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error (disk-full / out-of-fds) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: ${first.message} What it means
PdgEmitSink.finalize (gitnexus/src/core/lbug/pdg-emit-sink.ts:239) closes every streamed CSV writer and checks for poison. Synchronous write faults (fs.writeSync ENOSPC/EIO) and writer-open failures (EMFILE — out of file descriptors) are swallowed by the emit loop's per-file try/catch, so finalize is the backstop: it fails the whole run instead of handing a truncated CSV to the bulk COPY.
Solutions
- Free disk space on the volume holding the GitNexus storage/csv directory and re-run the analyze
- If the message cites EMFILE, raise the process fd limit (ulimit -n 4096, or LimitNOFILE= in systemd) and retry
- Point GitNexus storage at a volume with headroom larger than the expected graph CSV set
- Re-run the analyze — the aborted run left the dirty flag so the next run rebuilds cleanly
Defensive patterns
Strategy: retry
Validate before calling
// before a streamed analyze, verify headroom and fd budget
import { statfs } from 'node:fs/promises';
const { bavail, bsize } = await statfs(csvDir);
const freeBytes = Number(bavail) * Number(bsize);
if (freeBytes < MIN_REQUIRED_BYTES) throw new Error(`only ${freeBytes} bytes free for CSV emit`); Type guard
const isPdgSinkIoError = (e: unknown): boolean =>
e instanceof Error && e.message.startsWith('PdgEmitSink:') && e.message.includes('IO error'); Try / catch
try {
const manifest = sink.finalize();
} catch (e) {
if (isPdgSinkIoError(e)) {
// truncated CSVs must never be COPYed — free space / raise ulimit, then re-run the analyze
throw new Error(`streamed emit failed on IO; freeing resources and re-running analyze is required`);
}
throw e;
} Prevention
- Provision scratch space larger than the expected graph CSV set before large --pdg indexes
- Raise nofile limits for the analyze process in containers and CI (ulimit -n / LimitNOFILE)
- Treat any emit abort as requiring a fresh run — partial CSVs are never loaded
When it happens
Trigger: A streamed --pdg analyze on a large repo when the disk fills mid-emit, an fs.writeSync hits an IO error, or the process exhausts file descriptors opening the BasicBlock and per-pair rel CSV writers.
Common situations: Indexing monorepo or kernel-scale trees on small temp volumes; containers with low nofile limits; CI runners with quota-limited scratch space.
Related errors
- Streaming PDG manifest collides with a structural node CSV…
- Streaming PDG manifest collides with a structural…
- content filter triggered mid-stream. The generated content…
- GraphEmitSink: streamed CSV writer(s) hit an IO error…
- [lbug-load] node COPY also failed while relationship emit…
AI-assisted analysis of abhigyanpatwari/GitNexus@ac9a4e9abd (2026-08-20).
Data as JSON: /api/errors/b1a1bf3bc3be2a26.
Report an issue: GitHub.
Appendix: source
Thrown at gitnexus/src/core/lbug/pdg-emit-sink.ts:239
if (this.bbWriter !== undefined) {
this.bbWriter.close();
if (this.bbWriter.poison !== undefined) errors.push(this.bbWriter.poison);
nodeFiles.set('BasicBlock' as NodeTableName, {
csvPath: this.bbWriter.csvPath,
rows: this.bbWriter.rows,
});
}
const relsByPair = new Map<string, { csvPath: string; rows: number }>();
for (const [pairKey, writer] of this.relWriters) {
writer.close();
if (writer.poison !== undefined) errors.push(writer.poison);
relsByPair.set(pairKey, { csvPath: writer.csvPath, rows: writer.rows });
}
if (errors.length > 0) {
const first = errors[0];
throw new Error(
`PdgEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error ` +
`(disk-full / out-of-fds) during the emit — the persisted graph would ` +
`be truncated, so the run is failed rather than COPYing a partial CSV: ${
first instanceof Error ? first.message : String(first)
}`,
);
}
return { nodeFiles, relsByPair };
}
/**
* Best-effort fd release for the error path — when a language pass throws
* before {@link finalize} runs, the caller's `finally` calls this so the
* BasicBlock + per-pair fds never leak. Idempotent with finalize via the
* `finalized` flag; close errors are swallowed because the run is already
* failing.
*/View on GitHub (pinned to ac9a4e9abd)