{"record":{"id":"3742ab0d21bab956","repo":"Mintplex-Labs/anything-llm","slug":"e-message-3742ab","errorCode":null,"errorMessage":"e.message","messagePattern":"e\\.message","errorType":"exception","errorClass":"Error","httpStatus":null,"severity":"error","filePath":"server/utils/AiProviders/foundry/index.js","lineNumber":324,"sourceCode":"    if (!this.model)\n      throw new Error(\n        `Foundry chat: ${this.model} is not valid or defined model for chat completion!`\n      );\n\n    // max_completion_tokens is required by Foundry (it caps output at 1024\n    // otherwise), so the window has to be resolved before the request is built.\n    await this.assertModelContextLimits();\n    await this.assertModelLoaded();\n    const result = await LLMPerformanceMonitor.measureAsyncFunction(\n      this.openai.chat.completions\n        .create({\n          model: this.model,\n          messages,\n          temperature,\n          max_completion_tokens: this.promptWindowLimit(),\n        })\n        .catch((e) => {\n          throw new Error(e.message);\n        })\n    );\n\n    if (\n      !result.output.hasOwnProperty(\"choices\") ||\n      result.output.choices.length === 0\n    )\n      return null;\n\n    return {\n      textResponse: result.output.choices[0].message.content,\n      metrics: {\n        prompt_tokens: result.output.usage.prompt_tokens || 0,\n        completion_tokens: result.output.usage.completion_tokens || 0,\n        total_tokens: result.output.usage.total_tokens || 0,\n        outputTps: result.output.usage.completion_tokens / result.duration,\n        duration: result.duration,\n        model: this.model,","sourceCodeStart":306,"sourceCodeEnd":342,"githubUrl":"https://github.com/Mintplex-Labs/anything-llm/blob/3aec848f2885144aa8f1e53b9731a04310d5d558/server/utils/AiProviders/foundry/index.js#L306-L342","documentation":"FoundryLLM.getChatCompletion wraps the OpenAI-compatible chat.completions.create call in a .catch that rethrows `new Error(e.message)`, flattening whatever the SDK produced. This is a passthrough: the real failure is a Foundry Local HTTP/transport error (model crashed mid-request, malformed request, context overflow of max_completion_tokens, connection dropped). The rethrow exists to normalize SDK APIError objects into plain Errors for AnythingLLM's handler, but it discards status codes and the original stack.","triggerScenarios":"The HTTP POST to Foundry Local's /chat/completions returning 4xx/5xx: requesting more max_completion_tokens than the loaded model supports, a request payload the local service rejects, the model process dying mid-inference, or the socket closing after an eviction ('Premature close' — see #handleStreamFailure which treats this case).","commonSituations":"Context larger than the local model's window; model evicted between the load check and the request; Foundry Local version mismatch producing a schema the SDK rejects; laptop under memory pressure so inference crashes.","solutions":["Read the inner message — it is the literal Foundry Local error (e.g. 'max_completion_tokens too large', 'model not found') and dictates the fix","Reduce the prompt size / lower GENERIC token limits or clear chat history if the context overflows the small default 16K window","Reload or pick a smaller model if inference crashed (FoundryLLM forgets evicted models, so a retry may reload it)","Update Foundry Local and restart the service if the error indicates a schema/protocol problem"],"exampleFix":null,"handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"try {\n  return await llm.getChatCompletion(messages);\n} catch (err) {\n  // err.message is the raw Foundry Local error text; branch on its content\n  if (/context|token/i.test(err.message)) return respondTrimmedAndRetry(messages);\n  if (/Premature close|socket/i.test(err.message)) return retryOnce(); // model eviction\n  return respond(`Foundry inference failed: ${err.message}`);\n}","preventionTips":["Cap prompt size to the model's real window (FoundryLLM defaults to 16K) before sending","Retry once on socket/premature-close errors — the class already forgets evicted models so the retry reloads","Watch Foundry Local logs alongside your own to correlate the rethrown message with a status code"],"tags":["foundry-local","openai-compat","api-error-passthrough","chat-completions"],"backgroundTag":"openai-api-error","analyzedSha":"3aec848f2885144aa8f1e53b9731a04310d5d558","analyzedAt":"2026-08-18T10:02:21.017Z","contentChangedAt":"2026-08-18T10:02:21.017Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}