{"record":{"id":"b1a1bf3bc3be2a26","repo":"abhigyanpatwari/GitNexus","slug":"pdgemitsink-errors-length-streamed-csv-writer","errorCode":null,"errorMessage":"PdgEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error (disk-full / out-of-fds) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: ${first.message}","messagePattern":"PdgEmitSink: (.+?) streamed CSV writer\\(s\\) hit an IO error \\(disk-full / out-of-fds\\) during the emit — the persisted graph would be truncated, so the run is failed rather than COPYing a partial CSV: (.+?)","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"critical","filePath":"gitnexus/src/core/lbug/pdg-emit-sink.ts","lineNumber":239,"sourceCode":"    if (this.bbWriter !== undefined) {\n      this.bbWriter.close();\n      if (this.bbWriter.poison !== undefined) errors.push(this.bbWriter.poison);\n      nodeFiles.set('BasicBlock' as NodeTableName, {\n        csvPath: this.bbWriter.csvPath,\n        rows: this.bbWriter.rows,\n      });\n    }\n\n    const relsByPair = new Map<string, { csvPath: string; rows: number }>();\n    for (const [pairKey, writer] of this.relWriters) {\n      writer.close();\n      if (writer.poison !== undefined) errors.push(writer.poison);\n      relsByPair.set(pairKey, { csvPath: writer.csvPath, rows: writer.rows });\n    }\n\n    if (errors.length > 0) {\n      const first = errors[0];\n      throw new Error(\n        `PdgEmitSink: ${errors.length} streamed CSV writer(s) hit an IO error ` +\n          `(disk-full / out-of-fds) during the emit — the persisted graph would ` +\n          `be truncated, so the run is failed rather than COPYing a partial CSV: ${\n            first instanceof Error ? first.message : String(first)\n          }`,\n      );\n    }\n\n    return { nodeFiles, relsByPair };\n  }\n\n  /**\n   * Best-effort fd release for the error path — when a language pass throws\n   * before {@link finalize} runs, the caller's `finally` calls this so the\n   * BasicBlock + per-pair fds never leak. Idempotent with finalize via the\n   * `finalized` flag; close errors are swallowed because the run is already\n   * failing.\n   */","sourceCodeStart":221,"sourceCodeEnd":257,"githubUrl":"https://github.com/abhigyanpatwari/GitNexus/blob/ac9a4e9abd8fd3058c070b72c23402a4f887929a/gitnexus/src/core/lbug/pdg-emit-sink.ts#L221-L257","documentation":"PdgEmitSink.finalize (gitnexus/src/core/lbug/pdg-emit-sink.ts:239) closes every streamed CSV writer and checks for poison. Synchronous write faults (fs.writeSync ENOSPC/EIO) and writer-open failures (EMFILE — out of file descriptors) are swallowed by the emit loop's per-file try/catch, so finalize is the backstop: it fails the whole run instead of handing a truncated CSV to the bulk COPY.","triggerScenarios":"A streamed --pdg analyze on a large repo when the disk fills mid-emit, an fs.writeSync hits an IO error, or the process exhausts file descriptors opening the BasicBlock and per-pair rel CSV writers.","commonSituations":"Indexing monorepo or kernel-scale trees on small temp volumes; containers with low nofile limits; CI runners with quota-limited scratch space.","solutions":["Free disk space on the volume holding the GitNexus storage/csv directory and re-run the analyze","If the message cites EMFILE, raise the process fd limit (ulimit -n 4096, or LimitNOFILE= in systemd) and retry","Point GitNexus storage at a volume with headroom larger than the expected graph CSV set","Re-run the analyze — the aborted run left the dirty flag so the next run rebuilds cleanly"],"exampleFix":null,"handlingStrategy":"retry","validationCode":"// before a streamed analyze, verify headroom and fd budget\nimport { statfs } from 'node:fs/promises';\nconst { bavail, bsize } = await statfs(csvDir);\nconst freeBytes = Number(bavail) * Number(bsize);\nif (freeBytes < MIN_REQUIRED_BYTES) throw new Error(`only ${freeBytes} bytes free for CSV emit`);","typeGuard":"const isPdgSinkIoError = (e: unknown): boolean =>\n  e instanceof Error && e.message.startsWith('PdgEmitSink:') && e.message.includes('IO error');","tryCatchPattern":"try {\n  const manifest = sink.finalize();\n} catch (e) {\n  if (isPdgSinkIoError(e)) {\n    // truncated CSVs must never be COPYed — free space / raise ulimit, then re-run the analyze\n    throw new Error(`streamed emit failed on IO; freeing resources and re-running analyze is required`);\n  }\n  throw e;\n}","preventionTips":["Provision scratch space larger than the expected graph CSV set before large --pdg indexes","Raise nofile limits for the analyze process in containers and CI (ulimit -n / LimitNOFILE)","Treat any emit abort as requiring a fresh run — partial CSVs are never loaded"],"tags":["gitnexus","pdg","streaming","disk-full","emfile","csv"],"backgroundTag":"disk-full-write-failure","analyzedSha":"ac9a4e9abd8fd3058c070b72c23402a4f887929a","analyzedAt":"2026-08-20T23:29:22.980Z","contentChangedAt":"2026-08-20T23:29:22.980Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}