{"record":{"id":"707506397f5e1f6b","repo":"Graphify-Labs/graphify","slug":"error-graph-is-empty-extraction-produced-no-nod","errorCode":null,"errorMessage":"ERROR: Graph is empty - extraction produced no nodes.","messagePattern":"ERROR: Graph is empty - extraction produced no nodes\\.","errorType":"console","errorClass":"SystemExit","httpStatus":null,"severity":"error","filePath":"graphify/skill.md","lineNumber":419,"sourceCode":"from graphify.build import build_from_json\nfrom graphify.cluster import cluster, score_all\nfrom graphify.analyze import god_nodes, surprising_connections, suggest_questions\nfrom graphify.report import generate\nfrom graphify.export import to_json\nfrom pathlib import Path\n\nextraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\\\"utf-8\\\"))\ndetection  = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\\\"utf-8\\\"))\n\n# root= mirrors the --update runbook (#1361): relativize source_file to the same\n# base so the full build and incremental --update never drift apart on re-extract.\nG = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)\n# Guard BEFORE any write: an empty extraction must not clobber a good graph.json /\n# GRAPH_REPORT.md / analysis sidecar. Check immediately after build (#1392).\nif G.number_of_nodes() == 0:\n    print('ERROR: Graph is empty - extraction produced no nodes.')\n    print('Possible causes: all files were skipped, binary-only corpus, or extraction failed.')\n    raise SystemExit(1)\ncommunities = cluster(G)\ncohesion = score_all(G, communities)\ntokens = {'input': extraction.get('input_tokens', 0), 'output': extraction.get('output_tokens', 0)}\ngods = god_nodes(G)\nsurprises = surprising_connections(G, communities)\nlabels = {cid: 'Community ' + str(cid) for cid in communities}\n# Placeholder questions - regenerated with real labels in Step 5\nquestions = suggest_questions(G, communities, labels)\n\n# Export FIRST and honor the #479 shrink-guard: to_json returns False (writing\n# nothing) when the new graph is smaller than the existing graph.json. Only write\n# GRAPH_REPORT.md + the analysis sidecar when the graph was actually written, so\n# they never describe a graph that graph.json doesn't contain (#1392).\nwrote = to_json(G, communities, 'graphify-out/graph.json')\nif not wrote:\n    print('ERROR: refused to shrink graphify-out/graph.json (existing graph has more nodes; #479).')\n    print('If this shrink is intentional (you deleted files), re-run a full build with --force.')\n    raise SystemExit(1)","sourceCodeStart":401,"sourceCodeEnd":437,"githubUrl":"https://github.com/Graphify-Labs/graphify/blob/7fe58b0b0f3873be9a21c30106b8b8527c353aa6/graphify/skill.md#L401-L437","documentation":"This is not library code but the guarded build snippet embedded in graphify/skill.md (Step 4). After build_from_json() constructs the graph from .graphify_extract.json, the script checks G.number_of_nodes() and exits before ANY write if the graph is empty — so a failed or all-skipped extraction cannot clobber an existing good graph.json / GRAPH_REPORT.md / analysis sidecar (issue #1392). The message names the typical causes: everything skipped, a binary-only corpus, or extraction failure.","triggerScenarios":"Running the skill.md Step 4 build block when .graphify_extract.json contains no nodes: all candidate files were ignored/skipped (.graphifyignore covering everything, binary-only corpus), the LLM extraction step failed but wrote partial output, or the extraction JSON is from a different path that matched nothing.","commonSituations":"First run against a directory whose files are all ignored or non-source; extraction JSON left over from a run on a different INPUT_PATH; LLM extraction returned empty entities for every file; path passed with a trailing typo so nothing matched.","solutions":["Open graphify-out/.graphify_extract.json and confirm it has entities/nodes — if not, re-run the extraction step (Step 3) for the same INPUT_PATH","Check .graphifyignore patterns are not excluding everything, and confirm the input path actually contains source files","If extraction genuinely found nothing to index (binary-only corpus), point /graphify at a directory with extractable text/code","If extraction output looks truncated after an API failure, delete .graphify_extract.json and re-run the full extract"],"exampleFix":"before (unguarded):\nG = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)\nto_json(G, communities, 'graphify-out/graph.json')  # empty graph overwrites good graph.json\n\nafter (as in skill.md):\nG = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)\nif G.number_of_nodes() == 0:\n    print('ERROR: Graph is empty - extraction produced no nodes.')\n    raise SystemExit(1)","handlingStrategy":"validation","validationCode":"extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding='utf-8'))\nnode_count = len(extraction.get('nodes', extraction.get('entities', [])))\nif node_count == 0:\n    raise SystemExit('extraction empty — fix ignore rules / corpus before building')\nG = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)","typeGuard":"def extraction_has_nodes(extraction: dict) -> bool:\n    \"\"\"True when the extraction JSON carries at least one node/entity.\"\"\"\n    return bool(extraction.get('nodes') or extraction.get('entities'))","tryCatchPattern":null,"preventionTips":["Run extraction and build in the same session; never reuse an extract JSON from a different INPUT_PATH","Smoke-test a new corpus with a tiny directory first to confirm ignore rules do not exclude everything","Fail extraction loudly (non-zero exit) when it produces zero nodes so the build step never starts"],"tags":["graphify","guard","data-integrity","empty-state","python"],"backgroundTag":null,"analyzedSha":"7fe58b0b0f3873be9a21c30106b8b8527c353aa6","analyzedAt":"2026-08-14T19:23:21.323Z","schemaVersion":2},"datasetVersion":"2026-08-15T22:17:37.221Z"}