{"record":{"id":"1b837f71f89c4a2d","repo":"abhigyanpatwari/GitNexus","slug":"graphemitsink-errors-length-streamed-csv-write","errorCode":null,"errorMessage":"GraphEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error (disk-full / out-of-fds) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: ${first instanceof Error ? first.message : String(first)}","messagePattern":"GraphEmitSink: (.+?) streamed CSV writer\\(s\\) hit an IO error \\(disk-full / out-of-fds\\) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: (.+?)","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"critical","filePath":"gitnexus/src/core/lbug/graph-emit-sink.ts","lineNumber":524,"sourceCode":"  finalize(): GraphEmitManifest {\n    if (this.finalized) throw new Error('GraphEmitSink.finalize() called twice');\n    this.finalized = true;\n\n    const errors: unknown[] = [];\n    if (this.openFailure !== undefined) errors.push(this.openFailure);\n\n    const relsByPair = new Map<string, { csvPath: string; rows: number }>();\n    let totalRows = 0;\n    for (const [pairKey, writer] of this.relWriters) {\n      writer.close();\n      if (writer.poison !== undefined) errors.push(writer.poison);\n      relsByPair.set(pairKey, { csvPath: writer.csvPath, rows: writer.rows });\n      totalRows += writer.rows;\n    }\n\n    if (errors.length > 0) {\n      const first = errors[0];\n      throw new Error(\n        `GraphEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error ` +\n          `(disk-full / out-of-fds) during the emit — the persisted graph would be ` +\n          `truncated, so the run is failed rather than COPYing a partial CSV: ${\n            first instanceof Error ? first.message : String(first)\n          }`,\n      );\n    }\n\n    return { relsByPair, totalRows, structuralRows: this.structuralRows };\n  }\n\n  /** Best-effort fd release for the error path — when the pipeline throws\n   *  before {@link finalize} runs, the caller's `finally` calls this so the\n   *  per-pair fds never leak. Idempotent with finalize via `finalized`. */\n  close(): void {\n    if (this.finalized) return;\n    this.finalized = true;\n    for (const writer of this.relWriters.values()) {","sourceCodeStart":506,"sourceCodeEnd":542,"githubUrl":"https://github.com/abhigyanpatwari/GitNexus/blob/ac9a4e9abd8fd3058c070b72c23402a4f887929a/gitnexus/src/core/lbug/graph-emit-sink.ts#L506-L542","documentation":"At finalize time, GraphEmitSink closes every per-relationship-pair CSV writer and collects each writer's recorded `poison` error (an async write failure captured during streaming). If any writer hit an IO error — the message names disk-full and out-of-file-descriptors — the persisted CSVs would be truncated, so the run is failed instead of COPYing a partial graph into LadybugDB. This is a data-integrity refusal: a partial graph would look like a valid index.","triggerScenarios":"The storage volume fills up mid-emit while writers flush relationship CSVs; or the process exhausts its fd limit because the sink keeps one open CSV writer per relationship-type pair and a graph produces many distinct pairs.","commonSituations":"Very large repositories generating many relationship kinds (call/contains/imports/…) on CI runners with low `ulimit -n`; small tmpfs or shared disks that fill during big analyzes; other processes on the host consuming fds/disk concurrently.","solutions":["Free disk space (or point GitNexus storage at a larger volume) and re-run analyze — staging CSVs are regenerated on the next run","Raise the file-descriptor limit before running: `ulimit -n 4096` (or the container/systemd equivalent) if the error is fd exhaustion","Shrink the graph with .gitnexusignore excludes for directories you don't need indexed, reducing both CSV size and writer count"],"exampleFix":"# before\n$ ulimit -n 256 && gitnexus analyze .\n# after\n$ ulimit -n 4096 && gitnexus analyze .","handlingStrategy":"retry","validationCode":"import fs from 'node:fs';\n\n// Pre-flight before a big analyze: disk headroom and fd budget\nconst stat = fs.statfsSync(storageDir);\nconst freeGB = (stat.bavail * stat.bsize) / 1024 ** 3;\nif (freeGB < 5) throw new Error(`Only ${freeGB.toFixed(1)}GB free on ${storageDir} — free space before analyze`);\nconst [softFd] = [ulimit?]; // `ulimit -n` equivalent\n// e.g. run `ulimit -n` and require >= a few thousand for many relationship pairs","typeGuard":null,"tryCatchPattern":"try {\n  const { totalRows } = await sink.finalize();\n} catch (err) {\n  if (err instanceof Error && err.message.includes('streamed CSV writer(s) hit an IO error')) {\n    // free disk / raise `ulimit -n`, then re-run analyze from scratch — staging CSVs regenerate;\n    // never continue with the partial CSVs\n  }\n  throw err;\n}","preventionTips":["Run large analyzes with a raised fd limit (`ulimit -n 4096`) — one CSV writer stays open per relationship-type pair","Monitor disk space on the storage volume before and during analyze; big repos emit multi-GB CSVs","Use .gitnexusignore to keep unneeded directories out of the graph, shrinking both CSV size and writer count"],"tags":["csv","io","disk-full","file-descriptors","emit","data-integrity"],"backgroundTag":"disk-full","analyzedSha":"ac9a4e9abd8fd3058c070b72c23402a4f887929a","analyzedAt":"2026-08-20T23:29:22.980Z","contentChangedAt":"2026-08-20T23:29:22.980Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}