{"record":{"id":"5f81bae23268df94","repo":"666ghj/MiroFish","slug":"the-corpus-did-not-produce-enough-artifacts-to-exe","errorCode":null,"errorMessage":"the corpus did not produce enough artifacts to exercise pagination","messagePattern":"the corpus did not produce enough artifacts to exercise pagination","errorType":"exception","errorClass":"AssertionError","httpStatus":null,"severity":"warning","filePath":"backend/scripts/validate_zep_cloud_integration.py","lineNumber":561,"sourceCode":"            \"item_pages_at_size_3\": batch_pages,\n            \"episode_uuids\": baseline_episode_uuids,\n        }\n        print(f\"[zep-deep] batch completed pages={batch_pages}\", flush=True)\n\n        baseline_nodes = fetch_all_nodes(client, graph_id, page_size=2)\n        baseline_edges = fetch_all_edges(client, graph_id, page_size=2)\n        raw_nodes, node_pages = _raw_pages(\n            client.graph.node.with_raw_response.get_by_graph_id, graph_id, page_size=2\n        )\n        raw_edges, edge_pages = _raw_pages(\n            client.graph.edge.with_raw_response.get_by_graph_id, graph_id, page_size=2\n        )\n        if {_uuid(item) for item in baseline_nodes} != {_uuid(item) for item in raw_nodes}:\n            raise AssertionError(\"production node pagination did not match raw cursor traversal\")\n        if {_uuid(item) for item in baseline_edges} != {_uuid(item) for item in raw_edges}:\n            raise AssertionError(\"production edge pagination did not match raw cursor traversal\")\n        if len(baseline_nodes) <= 2 or len(baseline_edges) <= 2:\n            raise AssertionError(\"the corpus did not produce enough artifacts to exercise pagination\")\n\n        baseline_names = {_uuid(node): node.name for node in baseline_nodes}\n        baseline_ceo = client.graph.search(\n            graph_id=graph_id,\n            query=\"截至2026年4月底，谁担任澜舟科技首席执行官？\",\n            scope=\"edges\",\n            reranker=\"cross_encoder\",\n            limit=10,\n        )\n        result[\"baseline\"] = {\n            \"node_count\": len(baseline_nodes),\n            \"edge_count\": len(baseline_edges),\n            \"node_pages_at_size_2\": node_pages,\n            \"edge_pages_at_size_2\": edge_pages,\n            \"invalidated_edge_count\": sum(bool(edge.invalid_at) for edge in baseline_edges),\n            \"ceo_search\": _search_view(baseline_ceo, baseline_names),\n        }\n        print(","sourceCodeStart":543,"sourceCodeEnd":579,"githubUrl":"https://github.com/666ghj/MiroFish/blob/b5b53acc57189a4a42e44a23e149dc655c98fe82/backend/scripts/validate_zep_cloud_integration.py#L543-L579","documentation":"An AssertionError requiring len(baseline_nodes) > 2 and len(baseline_edges) > 2 (the check is <= 2 fails). Because both traversals use page_size=2, having at most 2 nodes or 2 edges would mean multi-page pagination was never exercised — the validation would silently prove nothing. This is a test-fixture adequacy check on the BASELINE_EPISODES corpus, not a production failure.","triggerScenarios":"The ingested baseline episodes produced 2 or fewer nodes or edges — corpus too small, episodes not fully processed at fetch time, or Zep's entity/fact extraction finding too little in the fixture text.","commonSituations":"Race with incomplete episode processing (see the TimeoutError from _wait_for_episode); a trimmed-down BASELINE_EPISODES fixture; short/generic fixture text yielding few entities and facts.","solutions":["Confirm every baseline episode reached processed=True before the fetch step — edges appear asynchronously after node extraction.","Enlarge BASELINE_EPISODES or make the fixture text richer in named entities and relationships so extraction yields >2 nodes and >2 edges.","Treat the assertion as a fixture-quality signal: fix the corpus, do not skip the check."],"exampleFix":null,"handlingStrategy":"validation","validationCode":"def corpus_exercises_pagination(nodes: list[Any], edges: list[Any]) -> bool:\n    return len(nodes) > 2 and len(edges) > 2","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Keep BASELINE_EPISODES text rich in named entities and relationships so extraction yields many nodes and edges.","Confirm all episodes report processed=True before fetching artifacts.","Treat this assertion as a fixture-quality signal, not a flake — fix the corpus, don't skip the check."],"tags":["zep","validation-script","test-fixture","assertion","data-quality"],"backgroundTag":null,"analyzedSha":"b5b53acc57189a4a42e44a23e149dc655c98fe82","analyzedAt":"2026-08-14T22:29:33.146Z","schemaVersion":2},"datasetVersion":"2026-08-15T22:17:37.221Z"}