janhq/jan · error

the request exceeds the available context size.

Error message

the request exceeds the available context size.

What it means

After a successful non-streaming completion, the extension inspects finish_reason; when the MLX server reports 'length', the prompt+response exceeded the model's context window and the response was truncated, so the extension throws OUT_OF_CONTEXT_SIZE instead of returning a cut-off answer.

Solutions

  1. Reduce the input: trim or summarize older conversation messages before sending
  2. Lower max_tokens in the request or in app settings
  3. Reload the model with a larger context size (n_ctx / context length setting)
  4. Switch to a model with a larger context window

Example fix

// before
await chat(allMessages, model)
// after
const trimmed = allMessages.slice(-10) // keep recent turns only
await chat(trimmed, model)
Defensive patterns

Strategy: validation

Validate before calling

function estimateTokens(messages) {
  return messages.reduce((n, m) => n + Math.ceil((m.content?.length ?? 0) / 4), 0)
}
const LIMIT = 2048 // model context size
if (estimateTokens(messages) + (maxTokens ?? 1024) >= LIMIT) {
  messages = [messages[0], ...messages.slice(-6)]
}

Type guard

null

Try / catch

try {
  return await chat(messages, model)
} catch (e) {
  if (String(e.message).includes('context size')) {
    return chat([messages[0], ...messages.slice(-6)], model)
  }
  throw e
}

Prevention

When it happens

Trigger: Calling chat() with messages whose total token count plus max_tokens exceeds the model's context size, causing the server to stop generation with finish_reason === 'length'.

Common situations: Long conversation histories sent in full every turn, very large system prompts, models loaded with a small context (e.g. 2048) while users paste large documents.

Related errors


AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17). Data as JSON: /api/errors/82d7c555460bfd26. Report an issue: GitHub.

Appendix: source

Thrown at extensions/mlx-extension/src/index.ts:441

    const response = await fetch(url, {
      method: 'POST',
      headers,
      body,
      signal: abortController?.signal,
    })

    if (!response.ok) {
      const errorData = await response.json().catch(() => null)
      throw new Error(
        `MLX API request failed with status ${response.status}: ${JSON.stringify(errorData)}`
      )
    }

    const completionResponse = (await response.json()) as chatCompletion

    if (completionResponse.choices?.[0]?.finish_reason === 'length') {
      throw new Error(OUT_OF_CONTEXT_SIZE)
    }

    return completionResponse
  }

  private async *handleStreamingResponse(
    url: string,
    headers: HeadersInit,
    body: string,
    abortController?: AbortController
  ): AsyncIterable<chatCompletionChunk> {
    // AbortSignal.any() is not available in all runtimes (e.g. WebKit/JavaScriptCore),
    // so we manually combine the timeout and external abort signals.
    const combinedController = new AbortController()
    const timeoutId = setTimeout(
      () => combinedController.abort(new Error('Request timed out')),
      this.timeout * 1000
    )

View on GitHub (pinned to 7205d770c1)