janhq/jan · error
the request exceeds the available context size.
Error message
the request exceeds the available context size.
What it means
After a successful non-streaming completion, the extension inspects finish_reason; when the MLX server reports 'length', the prompt+response exceeded the model's context window and the response was truncated, so the extension throws OUT_OF_CONTEXT_SIZE instead of returning a cut-off answer.
Solutions
- Reduce the input: trim or summarize older conversation messages before sending
- Lower max_tokens in the request or in app settings
- Reload the model with a larger context size (n_ctx / context length setting)
- Switch to a model with a larger context window
Example fix
// before await chat(allMessages, model) // after const trimmed = allMessages.slice(-10) // keep recent turns only await chat(trimmed, model)
Defensive patterns
Strategy: validation
Validate before calling
function estimateTokens(messages) {
return messages.reduce((n, m) => n + Math.ceil((m.content?.length ?? 0) / 4), 0)
}
const LIMIT = 2048 // model context size
if (estimateTokens(messages) + (maxTokens ?? 1024) >= LIMIT) {
messages = [messages[0], ...messages.slice(-6)]
} Type guard
null
Try / catch
try {
return await chat(messages, model)
} catch (e) {
if (String(e.message).includes('context size')) {
return chat([messages[0], ...messages.slice(-6)], model)
}
throw e
} Prevention
- Track running token count across turns and compact history proactively
- Set max_tokens below the model context minus the prompt size
- Load models with the largest context your RAM allows
- Summarize long documents before including them in prompts
When it happens
Trigger: Calling chat() with messages whose total token count plus max_tokens exceeds the model's context size, causing the server to stop generation with finish_reason === 'length'.
Common situations: Long conversation histories sent in full every turn, very large system prompts, models loaded with a small context (e.g. 2048) while users paste large documents.
Related errors
- error.message
- IO error
- MLX API request failed with status
- MLX model appears to have crashed! Please reload!
- MLX model has crashed! Please reload!
AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17).
Data as JSON: /api/errors/82d7c555460bfd26.
Report an issue: GitHub.
Appendix: source
Thrown at extensions/mlx-extension/src/index.ts:441
const response = await fetch(url, {
method: 'POST',
headers,
body,
signal: abortController?.signal,
})
if (!response.ok) {
const errorData = await response.json().catch(() => null)
throw new Error(
`MLX API request failed with status ${response.status}: ${JSON.stringify(errorData)}`
)
}
const completionResponse = (await response.json()) as chatCompletion
if (completionResponse.choices?.[0]?.finish_reason === 'length') {
throw new Error(OUT_OF_CONTEXT_SIZE)
}
return completionResponse
}
private async *handleStreamingResponse(
url: string,
headers: HeadersInit,
body: string,
abortController?: AbortController
): AsyncIterable<chatCompletionChunk> {
// AbortSignal.any() is not available in all runtimes (e.g. WebKit/JavaScriptCore),
// so we manually combine the timeout and external abort signals.
const combinedController = new AbortController()
const timeoutId = setTimeout(
() => combinedController.abort(new Error('Request timed out')),
this.timeout * 1000
)View on GitHub (pinned to 7205d770c1)