666ghj/MiroFish · error · AssertionError

production edge pagination did not match raw cursor…

Error message

production edge pagination did not match raw cursor traversal

What it means

The edge counterpart of the node cross-check: the UUID set from fetch_all_edges must exactly match the set from _raw_pages over graph.edge.with_raw_response.get_by_graph_id. A mismatch means the production edge paginator and the manual raw-cursor walk observed different edges — dropped pages, duplicates, or edge churn between the two traversals.

Solutions

  1. Wait for every episode to reach processed=True and stop all writers before running the validation.
  2. Upgrade zep-cloud and rerun on a quiesced graph.
  3. If reproducible, diff both edge UUID sets to identify dropped/extra edges, then escalate to Zep support.
Defensive patterns

Strategy: retry

Try / catch

try:
    assert_equal_uuid_sets(baseline_edges, raw_edges)
except AssertionError as e:
    if "edge pagination did not match" in str(e):
        time.sleep(10)  # let async edge derivation settle
        baseline_edges = fetch_all_edges(client, graph_id, page_size=2)
        raw_edges, _ = _raw_pages(client.graph.edge.with_raw_response.get_by_graph_id, graph_id, page_size=2)
        assert_equal_uuid_sets(baseline_edges, raw_edges)
    else:
        raise

Prevention

When it happens

Trigger: Edges being created/invalidated concurrently (Zep derives edges asynchronously from episodes still processing); one pagination path mis-handling the zep-next-cursor for edges.

Common situations: A still-processing episode adds edges mid-check; the updater runs during validation; SDK version drift affecting edge pagination only.

Related errors


AI-assisted analysis of 666ghj/MiroFish@b5b53acc57 (2026-08-14). Data as JSON: /api/errors/4edc9ee9dd337ddb. Report an issue: GitHub.

Appendix: source

Thrown at backend/scripts/validate_zep_cloud_integration.py:559

            "status": client.batch.get(batch_id=batch_id).status,
            "item_count": len(listed_items),
            "item_pages_at_size_3": batch_pages,
            "episode_uuids": baseline_episode_uuids,
        }
        print(f"[zep-deep] batch completed pages={batch_pages}", flush=True)

        baseline_nodes = fetch_all_nodes(client, graph_id, page_size=2)
        baseline_edges = fetch_all_edges(client, graph_id, page_size=2)
        raw_nodes, node_pages = _raw_pages(
            client.graph.node.with_raw_response.get_by_graph_id, graph_id, page_size=2
        )
        raw_edges, edge_pages = _raw_pages(
            client.graph.edge.with_raw_response.get_by_graph_id, graph_id, page_size=2
        )
        if {_uuid(item) for item in baseline_nodes} != {_uuid(item) for item in raw_nodes}:
            raise AssertionError("production node pagination did not match raw cursor traversal")
        if {_uuid(item) for item in baseline_edges} != {_uuid(item) for item in raw_edges}:
            raise AssertionError("production edge pagination did not match raw cursor traversal")
        if len(baseline_nodes) <= 2 or len(baseline_edges) <= 2:
            raise AssertionError("the corpus did not produce enough artifacts to exercise pagination")

        baseline_names = {_uuid(node): node.name for node in baseline_nodes}
        baseline_ceo = client.graph.search(
            graph_id=graph_id,
            query="截至2026年4月底,谁担任澜舟科技首席执行官?",
            scope="edges",
            reranker="cross_encoder",
            limit=10,
        )
        result["baseline"] = {
            "node_count": len(baseline_nodes),
            "edge_count": len(baseline_edges),
            "node_pages_at_size_2": node_pages,
            "edge_pages_at_size_2": edge_pages,
            "invalidated_edge_count": sum(bool(edge.invalid_at) for edge in baseline_edges),
            "ceo_search": _search_view(baseline_ceo, baseline_names),

View on GitHub (pinned to b5b53acc57)