risingwavelabs/risingwave · error

table fragment of job

Error message

table fragment of job {job_id} not found

What it means

`get_fragments` asks the meta service for the table fragments of a streaming job and expects at least one entry. An empty map means the meta service has no fragment record for that job id, so the handler bails.

Solutions

  1. Verify the job is still running: `SHOW JOBS` / `SHOW MATERIALIZED VIEWS` and check the job status
  2. Retry after a short wait if DDL was just submitted (job not yet scheduled)
  3. If the job is finished or failed, its fragments no longer exist — inspect the successor MV/index or re-create the job
  4. Check meta service logs for the job id if state seems inconsistent
Defensive patterns

Strategy: retry

Validate before calling

-- ensure the job exists and is running before listing fragments
SELECT * FROM rw_catalog.rw_ddl_progress;
SHOW JOBS;

Try / catch

match list_fragments(job_id) {
    Err(e) if e.to_string().contains("table fragment of job") => {
        // wait and retry once; else treat job as finished/missing
    }
    other => other?,
}

Prevention

When it happens

Trigger: `EXPLAIN ANALYZE` / stream-job inspection for a job id whose fragments were never created, were dropped (job finished/canceled/background DDL cleaned up), or before the streaming job was scheduled on workers.

Common situations: Querying fragments right after DDL submission before the job is scheduled; job already finished (e.g. materialized view creation completed and fragments cleaned) or failed; stale frontend cache pointing at an old job id; meta service state lost after cluster restore.

Understand the failure class

Background: Record Not Found Errors: "not found", RecordNotFound, and "was not found" — what they mean and how to fix them — this error's family across 28 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/2470795a4154cc13. Report an issue: GitHub.

Appendix: source

Thrown at src/frontend/src/handler/explain_analyze_stream_job.rs:202

            .filter(|node| {
                node.property
                    .as_ref()
                    .map(|p| p.is_streaming)
                    .unwrap_or_else(|| false)
            })
            .collect::<Vec<_>>();
        Ok(stream_worker_nodes)
    }

    // TODO(kwannoel): Only fetch the names, actor_ids and graph of the fragments
    pub(super) async fn get_fragments(
        meta_client: &dyn FrontendMetaClient,
        job_id: JobId,
    ) -> Result<Vec<FragmentInfo>> {
        let mut fragment_map = meta_client.list_table_fragments(&[job_id]).await?;
        let mut table_fragments = fragment_map.drain();
        let Some((fragment_job_id, table_fragment_info)) = table_fragments.next() else {
            bail!("table fragment of job {job_id} not found");
        };
        assert_eq!(
            table_fragments.next(),
            None,
            "expected only at most one fragment"
        );
        assert_eq!(fragment_job_id, job_id);
        Ok(table_fragment_info.fragments)
    }

    pub(super) async fn get_executor_stats(
        handler_args: &HandlerArgs,
        worker_nodes: &[WorkerNode],
        executor_ids: &HashSet<ExecutorId>,
        dispatcher_fragment_ids: &HashSet<FragmentId>,
        profiling_duration: Duration,
    ) -> Result<ExecutorStats> {
        let dispatcher_fragment_ids = dispatcher_fragment_ids.iter().copied().collect::<Vec<_>>();

View on GitHub (pinned to 6469eb736d)