spring-projects/spring-ai · error · RuntimeException
ai.onnxruntime.OrtException
Error message
ai.onnxruntime.OrtException
What it means
TransformersEmbeddingModel.call() runs sentence embedding through the ONNX Runtime session. Any OrtException raised by the ONNX runtime during inference (bad input tensor shape, tokenizer/model mismatch, corrupted or incompatible model file, native runtime failure) is caught and rethrown as a RuntimeException wrapping the OrtException.
Solutions
- Read the wrapped OrtException cause message — it names the failing tensor/shape/operator.
- Clear the spring-ai model cache directory and let the library re-download clean model files.
- Ensure the tokenizer and ONNX model come from the same model repository/version.
- Verify ONNX Runtime native libraries load correctly on your platform/architecture.
Example fix
// before
model.setMetadata(new TransformersEmbeddingModel.Metadata("model.onnx", "tokenizer.json from another model"));
// after
model.setMetadata(new TransformersEmbeddingModel.Metadata(
"https://huggingface.co/Xenova/all-MiniLM-L6-v2/onnx/model.onnx",
"https://huggingface.co/Xenova/all-MiniLM-L6-v2/tokenizer.json")); Defensive patterns
Strategy: try-catch
Try / catch
try {
float[][] embeddings = embeddingModel.embed(docs);
} catch (RuntimeException e) {
if (e.getCause() instanceof OrtException ort) {
throw new IllegalStateException("ONNX inference failed: " + ort.getMessage(), ort);
}
throw e;
} Prevention
- Keep tokenizer and ONNX model files from the same model repo/version.
- Clear and re-download the model cache if files may be corrupted.
- Verify ONNX Runtime native libs support your platform/arch.
- Keep input text lengths within the model's sequence limit.
When it happens
Trigger: Calling embed()/call() when the ONNX session cannot process the input: input tensor shape does not match model expectations, wrong tokenizer/model pairing, corrupted cached ONNX model, or ONNX Runtime native library problems.
Common situations: Manually replaced or partially downloaded cached model files; using a model whose input signature differs from what the code feeds (e.g. different max sequence length); mixing model and tokenizer resources from different models; failing ONNX Runtime native extraction in exotic environments.
Understand the failure class
Background: Tensor shape mismatch errors ("must have shape", "expected shape ... got ..."): when tensor dimensions disagree with what an op or layer was told to expect — this error's family across 6 libraries.
Related errors
- EmbeddingModel must be provided
- Empty embedding vector returned for input at index +…
- Failed to obtain the embedding dimensions from the…
- Failed to obtain the embedding dimensions from the…
- Failed to obtain the embedding dimensions from the…
AI-assisted analysis of spring-projects/spring-ai@98a7beda4f (2026-09-11).
Data as JSON: /api/errors/8d5bbec4c7daf1ee.
Report an issue: GitHub.
Appendix: source
Thrown at models/spring-ai-transformers/src/main/java/org/springframework/ai/transformers/TransformersEmbeddingModel.java:379
// 1 - sequence_length (128)
// 2 - embedding dimensions (384)
float[][][] tokenEmbeddings = (float[][][]) lastHiddenState.getValue();
try (NDManager manager = NDManager.newBaseManager()) {
NDArray ndTokenEmbeddings = create(tokenEmbeddings, manager);
NDArray ndAttentionMask = manager.create(attention_mask0);
NDArray embedding = meanPooling(ndTokenEmbeddings, ndAttentionMask);
for (int i = 0; i < embedding.size(0); i++) {
resultEmbeddings.add(embedding.get(i).toFloatArray());
}
}
}
}
}
catch (OrtException ex) {
throw new RuntimeException(ex);
}
var indexCounter = new AtomicInteger(0);
EmbeddingResponse embeddingResponse = new EmbeddingResponse(
resultEmbeddings.stream().map(e -> new Embedding(e, indexCounter.incrementAndGet())).toList());
observationContext.setResponse(embeddingResponse);
return embeddingResponse;
});
}
private Map<String, OnnxTensor> removeUnknownModelInputs(Map<String, OnnxTensor> modelInputs) {
return modelInputs.entrySet()
.stream()
.filter(a -> this.onnxModelInputs.contains(a.getKey()))
.collect(Collectors.toMap(Map.Entry::getKey, Map.Entry::getValue));View on GitHub (pinned to 98a7beda4f)