janhq/jan · error

API request failed with status

Error message

API request failed with status ${response.status}: ${JSON.stringify(errorData)}

What it means

Thrown by the llamacpp extension's streaming chat-completion request wrapper when the llama.cpp server returns a non-2xx HTTP response. The extension reads the JSON error body (tolerating a missing/invalid body) and embeds both the status code and the raw payload into the message so the caller can see exactly what the inference server rejected.

Solutions

  1. Check the llama.cpp server logs for the root cause (load failure, OOM, bad request).
  2. Verify the model is loaded before sending requests (loadModel / check session health).
  3. Confirm the base URL/port matches the running llama.cpp server instance.
  4. Reduce prompt/context size or ctx length if the error indicates allocation failure.

Example fix

// before
const res = await fetch(url, { body })
const reader = res.body.getReader() // crashes or throws generic error
// after
const res = await fetch(url, { body })
if (!res.ok) {
  const err = await res.json().catch(() => null)
  throw new Error(`llamacpp request failed (${res.status}): ${JSON.stringify(err)}`)
}
const reader = res.body.getReader()
Defensive patterns

Strategy: try-catch

Validate before calling

const ok = await fetch(`${base}/health`).then(r => r.ok).catch(() => false)
if (!ok) throw new Error('llamacpp server is not reachable')
if (!(await isModelLoaded(modelId))) await loadModel(modelId)

Type guard

function isApiErrorResponse(e: unknown): e is Error & { status?: number } {
  return e instanceof Error && /API request failed with status \d+/.test(e.message)
}

Try / catch

try {
  for await (const chunk of stream) { /* ... */ }
} catch (e) {
  if (isApiErrorResponse(e)) {
    const status = Number(e.message.match(/status (\d+)/)?.[1])
    if (status >= 500) await restartEngineAndRetry()
    else showUserError(e.message)
  } else throw e
}

Prevention

When it happens

Trigger: Any call to the streaming completion path where response.ok is false — e.g. model not loaded (404/400), OOM or server crash (500), invalid request body, or the server port not actually serving llama.cpp.

Common situations: Model failed to load or was unloaded while a request was in flight; context length exceeded by prompt; malformed chat template; llama.cpp server returning 500 on internal errors; pointing at the wrong port where another service answers.

Understand the failure class

Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.

Related errors


AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17). Data as JSON: /api/errors/c950d57c1b565a83. Report an issue: GitHub.

Appendix: source

Thrown at extensions/llamacpp-extension/src/index.ts:2422

        combinedController.abort(abortController.signal.reason)
      } else {
        abortController.signal.addEventListener(
          'abort',
          () => combinedController.abort(abortController.signal.reason),
          { once: true }
        )
      }
    }
    const response = await fetch(url, {
      method: 'POST',
      headers,
      body,
      connectTimeout: Number(this.timeout) * 1000, // default 10 minutes
      signal: combinedController.signal,
    }).finally(() => clearTimeout(timeoutId))
    if (!response.ok) {
      const errorData = await response.json().catch(() => null)
      throw new Error(
        `API request failed with status ${response.status}: ${JSON.stringify(
          errorData
        )}`
      )
    }

    if (!response.body) {
      throw new Error('Response body is null')
    }

    const reader = response.body.getReader()
    const decoder = new TextDecoder('utf-8')
    let buffer = ''
    let jsonStr = ''
    try {
      while (true) {
        const { done, value } = await reader.read()

View on GitHub (pinned to 7205d770c1)