abhigyanpatwari/GitNexus · critical
GraphEmitSink: streamed CSV writer(s) hit an IO error…
Error message
GraphEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error (disk-full / out-of-fds) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: ${first instanceof Error ? first.message : String(first)} What it means
At finalize time, GraphEmitSink closes every per-relationship-pair CSV writer and collects each writer's recorded `poison` error (an async write failure captured during streaming). If any writer hit an IO error — the message names disk-full and out-of-file-descriptors — the persisted CSVs would be truncated, so the run is failed instead of COPYing a partial graph into LadybugDB. This is a data-integrity refusal: a partial graph would look like a valid index.
Solutions
- Free disk space (or point GitNexus storage at a larger volume) and re-run analyze — staging CSVs are regenerated on the next run
- Raise the file-descriptor limit before running: `ulimit -n 4096` (or the container/systemd equivalent) if the error is fd exhaustion
- Shrink the graph with .gitnexusignore excludes for directories you don't need indexed, reducing both CSV size and writer count
Example fix
# before $ ulimit -n 256 && gitnexus analyze . # after $ ulimit -n 4096 && gitnexus analyze .
Defensive patterns
Strategy: retry
Validate before calling
import fs from 'node:fs';
// Pre-flight before a big analyze: disk headroom and fd budget
const stat = fs.statfsSync(storageDir);
const freeGB = (stat.bavail * stat.bsize) / 1024 ** 3;
if (freeGB < 5) throw new Error(`Only ${freeGB.toFixed(1)}GB free on ${storageDir} — free space before analyze`);
const [softFd] = [ulimit?]; // `ulimit -n` equivalent
// e.g. run `ulimit -n` and require >= a few thousand for many relationship pairs Try / catch
try {
const { totalRows } = await sink.finalize();
} catch (err) {
if (err instanceof Error && err.message.includes('streamed CSV writer(s) hit an IO error')) {
// free disk / raise `ulimit -n`, then re-run analyze from scratch — staging CSVs regenerate;
// never continue with the partial CSVs
}
throw err;
} Prevention
- Run large analyzes with a raised fd limit (`ulimit -n 4096`) — one CSV writer stays open per relationship-type pair
- Monitor disk space on the storage volume before and during analyze; big repos emit multi-GB CSVs
- Use .gitnexusignore to keep unneeded directories out of the graph, shrinking both CSV size and writer count
When it happens
Trigger: The storage volume fills up mid-emit while writers flush relationship CSVs; or the process exhausts its fd limit because the sink keeps one open CSV writer per relationship-type pair and a graph produces many distinct pairs.
Common situations: Very large repositories generating many relationship kinds (call/contains/imports/…) on CI runners with low `ulimit -n`; small tmpfs or shared disks that fill during big analyzes; other processes on the host consuming fds/disk concurrently.
Related errors
- Cannot safely encode CSV string-list item
- [lbug-load] node COPY also failed while relationship emit…
- PdgEmitSink: streamed CSV writer(s) hit an IO error…
- Analysis did not finalize for
- Analysis did not finalize for
AI-assisted analysis of abhigyanpatwari/GitNexus@ac9a4e9abd (2026-08-20).
Data as JSON: /api/errors/1b837f71f89c4a2d.
Report an issue: GitHub.
Appendix: source
Thrown at gitnexus/src/core/lbug/graph-emit-sink.ts:524
finalize(): GraphEmitManifest {
if (this.finalized) throw new Error('GraphEmitSink.finalize() called twice');
this.finalized = true;
const errors: unknown[] = [];
if (this.openFailure !== undefined) errors.push(this.openFailure);
const relsByPair = new Map<string, { csvPath: string; rows: number }>();
let totalRows = 0;
for (const [pairKey, writer] of this.relWriters) {
writer.close();
if (writer.poison !== undefined) errors.push(writer.poison);
relsByPair.set(pairKey, { csvPath: writer.csvPath, rows: writer.rows });
totalRows += writer.rows;
}
if (errors.length > 0) {
const first = errors[0];
throw new Error(
`GraphEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error ` +
`(disk-full / out-of-fds) during the emit — the persisted graph would be ` +
`truncated, so the run is failed rather than COPYing a partial CSV: ${
first instanceof Error ? first.message : String(first)
}`,
);
}
return { relsByPair, totalRows, structuralRows: this.structuralRows };
}
/** Best-effort fd release for the error path — when the pipeline throws
* before {@link finalize} runs, the caller's `finally` calls this so the
* per-pair fds never leak. Idempotent with finalize via `finalized`. */
close(): void {
if (this.finalized) return;
this.finalized = true;
for (const writer of this.relWriters.values()) {View on GitHub (pinned to ac9a4e9abd)