dotnet/machinelearning · error · InvalidDataException
The tokenizer.json model 'unk_id
Error message
The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces. What it means
A numeric unk_id must index a valid entry in the parsed vocabulary. If unk_id is negative or >= the number of vocab pieces, it cannot refer to any piece, so InvalidDataException is thrown before the tokenizer is constructed.
Solutions
- Set unk_id to the actual index of the '<unk>' piece within model.vocab (0-based).
- If the unknown token was intentionally removed, set unk_id to null and enable byte_fallback instead.
- Re-export tokenizer.json from the original model so unk_id and vocab stay consistent.
Example fix
// before "unk_id": 3, "vocab": [["<unk>", 0.0], ["a", -1.0]] // only 2 pieces // after "unk_id": 0, "vocab": [["<unk>", 0.0], ["a", -1.0]]
Defensive patterns
Strategy: validation
Validate before calling
int vocabCount = model.GetProperty("vocab").GetArrayLength();
var unkId = model.GetProperty("unk_id");
if (unkId.ValueKind == JsonValueKind.Number &&
(unkId.GetInt32() < 0 || unkId.GetInt32() >= vocabCount))
throw new InvalidOperationException($"unk_id {unkId.GetInt32()} out of range for {vocabCount} pieces."); Try / catch
try { var tok = SentencePieceTokenizer.CreateFromTokenizerJson(stream); }
catch (InvalidDataException ex) when (ex.Message.Contains("out of range")) { /* fix unk_id or restore full vocab */ } Prevention
- Update unk_id whenever vocab is edited or reordered.
- Never mix model sections from different tokenizer.json files.
- Assert unk_id < vocab length in a pre-load sanity check.
When it happens
Trigger: CreateFromTokenizerJson where model.unk_id is, e.g., 3 but vocab contains only 2 pieces, or unk_id is -1 while not null (note: the loader uses -1 internally only for null unk_id).
Common situations: Editing the vocab (removing <unk>) without updating unk_id; tokenizer.json files whose unk_id was written for a different/larger vocabulary; mixing the model section of one file with the vocab of another.
Understand the failure class
Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.
Related errors
- A piece string in 'model.vocab' is null.
- Each entry in 'model.vocab' must be a [piece, score] array.
- Each entry in 'model.vocab' must be a [string piece, number…
- Expected model type 'Unigram' but found
- The tokenizer.json model does not contain a valid 'vocab'…
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/18dfa7b886fd3cd8.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/SentencePieceTokenizer.cs:661
throw new InvalidDataException("A piece string in 'model.vocab' is null.");
}
vocab.Add((piece, entry[1].GetSingle()));
}
if (unkIsNull)
{
// Without an unknown token the only way to represent out-of-vocabulary input is byte fallback; a model
// with neither cannot encode OOV text, so reject that combination up front rather than emitting an
// invalid token id at encode time.
if (!byteFallback)
{
throw new NotSupportedException("The tokenizer.json model has a null 'unk_id' but does not enable 'byte_fallback'; a Unigram model without an unknown token is only supported when byte_fallback is enabled.");
}
}
else if (unkId < 0 || unkId >= vocab.Count)
{
throw new InvalidDataException($"The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces.");
}
// Extract normalizer settings
byte[]? precompiledCharsMap = null;
bool addDummyPrefix = true;
// HF tokenizer.json has no remove_extra_whitespaces flag; SpmConverter encodes that behavior as
// explicit normalizer steps (a right-Strip plus a Replace collapsing runs of spaces). Deduce it from
// those steps, defaulting to false when absent to match the HF fast-tokenizer runtime.
bool removeExtraWhitespaces = false;
// When the normalizer has content-modifying steps that the charsmap + removeExtraWhitespaces
// approximation cannot represent (per-character Replace, Lowercase, Unicode normalization, ...),
// apply the full normalizer chain (charsmap included) before the Metaspace pass instead.
SentencePieceNormalizationStep? chainNormalizer = null;
if (root.TryGetProperty("normalizer", out JsonElement normalizerElement) &&
normalizerElement.ValueKind == JsonValueKind.Object)
{
if (SentencePieceNormalizationStep.HasRichSteps(normalizerElement))
{View on GitHub (pinned to 7b76e69cf9)