spring-projects/spring-ai · error · RuntimeException

ai.onnxruntime.OrtException

Error message

ai.onnxruntime.OrtException

What it means

TransformersEmbeddingModel.call() runs sentence embedding through the ONNX Runtime session. Any OrtException raised by the ONNX runtime during inference (bad input tensor shape, tokenizer/model mismatch, corrupted or incompatible model file, native runtime failure) is caught and rethrown as a RuntimeException wrapping the OrtException.

Solutions

  1. Read the wrapped OrtException cause message — it names the failing tensor/shape/operator.
  2. Clear the spring-ai model cache directory and let the library re-download clean model files.
  3. Ensure the tokenizer and ONNX model come from the same model repository/version.
  4. Verify ONNX Runtime native libraries load correctly on your platform/architecture.

Example fix

// before
model.setMetadata(new TransformersEmbeddingModel.Metadata("model.onnx", "tokenizer.json from another model"));
// after
model.setMetadata(new TransformersEmbeddingModel.Metadata(
    "https://huggingface.co/Xenova/all-MiniLM-L6-v2/onnx/model.onnx",
    "https://huggingface.co/Xenova/all-MiniLM-L6-v2/tokenizer.json"));
Defensive patterns

Strategy: try-catch

Try / catch

try {
    float[][] embeddings = embeddingModel.embed(docs);
} catch (RuntimeException e) {
    if (e.getCause() instanceof OrtException ort) {
        throw new IllegalStateException("ONNX inference failed: " + ort.getMessage(), ort);
    }
    throw e;
}

Prevention

When it happens

Trigger: Calling embed()/call() when the ONNX session cannot process the input: input tensor shape does not match model expectations, wrong tokenizer/model pairing, corrupted cached ONNX model, or ONNX Runtime native library problems.

Common situations: Manually replaced or partially downloaded cached model files; using a model whose input signature differs from what the code feeds (e.g. different max sequence length); mixing model and tokenizer resources from different models; failing ONNX Runtime native extraction in exotic environments.

Understand the failure class

Background: Tensor shape mismatch errors ("must have shape", "expected shape ... got ..."): when tensor dimensions disagree with what an op or layer was told to expect — this error's family across 6 libraries.

Related errors


AI-assisted analysis of spring-projects/spring-ai@98a7beda4f (2026-09-11). Data as JSON: /api/errors/8d5bbec4c7daf1ee. Report an issue: GitHub.

Appendix: source

Thrown at models/spring-ai-transformers/src/main/java/org/springframework/ai/transformers/TransformersEmbeddingModel.java:379

							// 1 - sequence_length (128)
							// 2 - embedding dimensions (384)
							float[][][] tokenEmbeddings = (float[][][]) lastHiddenState.getValue();

							try (NDManager manager = NDManager.newBaseManager()) {
								NDArray ndTokenEmbeddings = create(tokenEmbeddings, manager);
								NDArray ndAttentionMask = manager.create(attention_mask0);

								NDArray embedding = meanPooling(ndTokenEmbeddings, ndAttentionMask);

								for (int i = 0; i < embedding.size(0); i++) {
									resultEmbeddings.add(embedding.get(i).toFloatArray());
								}
							}
						}
					}
				}
				catch (OrtException ex) {
					throw new RuntimeException(ex);
				}

				var indexCounter = new AtomicInteger(0);

				EmbeddingResponse embeddingResponse = new EmbeddingResponse(
						resultEmbeddings.stream().map(e -> new Embedding(e, indexCounter.incrementAndGet())).toList());
				observationContext.setResponse(embeddingResponse);

				return embeddingResponse;
			});
	}

	private Map<String, OnnxTensor> removeUnknownModelInputs(Map<String, OnnxTensor> modelInputs) {

		return modelInputs.entrySet()
			.stream()
			.filter(a -> this.onnxModelInputs.contains(a.getKey()))
			.collect(Collectors.toMap(Map.Entry::getKey, Map.Entry::getValue));

View on GitHub (pinned to 98a7beda4f)