dotnet/machinelearning · error · InvalidDataException

The tokenizer.json model 'unk_id

Error message

The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces.

What it means

A numeric unk_id must index a valid entry in the parsed vocabulary. If unk_id is negative or >= the number of vocab pieces, it cannot refer to any piece, so InvalidDataException is thrown before the tokenizer is constructed.

Solutions

  1. Set unk_id to the actual index of the '<unk>' piece within model.vocab (0-based).
  2. If the unknown token was intentionally removed, set unk_id to null and enable byte_fallback instead.
  3. Re-export tokenizer.json from the original model so unk_id and vocab stay consistent.

Example fix

// before
"unk_id": 3, "vocab": [["<unk>", 0.0], ["a", -1.0]]  // only 2 pieces
// after
"unk_id": 0, "vocab": [["<unk>", 0.0], ["a", -1.0]]
Defensive patterns

Strategy: validation

Validate before calling

int vocabCount = model.GetProperty("vocab").GetArrayLength();
var unkId = model.GetProperty("unk_id");
if (unkId.ValueKind == JsonValueKind.Number &&
    (unkId.GetInt32() < 0 || unkId.GetInt32() >= vocabCount))
    throw new InvalidOperationException($"unk_id {unkId.GetInt32()} out of range for {vocabCount} pieces.");

Try / catch

try { var tok = SentencePieceTokenizer.CreateFromTokenizerJson(stream); }
catch (InvalidDataException ex) when (ex.Message.Contains("out of range")) { /* fix unk_id or restore full vocab */ }

Prevention

When it happens

Trigger: CreateFromTokenizerJson where model.unk_id is, e.g., 3 but vocab contains only 2 pieces, or unk_id is -1 while not null (note: the loader uses -1 internally only for null unk_id).

Common situations: Editing the vocab (removing <unk>) without updating unk_id; tokenizer.json files whose unk_id was written for a different/larger vocabulary; mixing the model section of one file with the vocab of another.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/18dfa7b886fd3cd8. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/SentencePieceTokenizer.cs:661

                    throw new InvalidDataException("A piece string in 'model.vocab' is null.");
                }

                vocab.Add((piece, entry[1].GetSingle()));
            }

            if (unkIsNull)
            {
                // Without an unknown token the only way to represent out-of-vocabulary input is byte fallback; a model
                // with neither cannot encode OOV text, so reject that combination up front rather than emitting an
                // invalid token id at encode time.
                if (!byteFallback)
                {
                    throw new NotSupportedException("The tokenizer.json model has a null 'unk_id' but does not enable 'byte_fallback'; a Unigram model without an unknown token is only supported when byte_fallback is enabled.");
                }
            }
            else if (unkId < 0 || unkId >= vocab.Count)
            {
                throw new InvalidDataException($"The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces.");
            }

            // Extract normalizer settings
            byte[]? precompiledCharsMap = null;
            bool addDummyPrefix = true;
            // HF tokenizer.json has no remove_extra_whitespaces flag; SpmConverter encodes that behavior as
            // explicit normalizer steps (a right-Strip plus a Replace collapsing runs of spaces). Deduce it from
            // those steps, defaulting to false when absent to match the HF fast-tokenizer runtime.
            bool removeExtraWhitespaces = false;
            // When the normalizer has content-modifying steps that the charsmap + removeExtraWhitespaces
            // approximation cannot represent (per-character Replace, Lowercase, Unicode normalization, ...),
            // apply the full normalizer chain (charsmap included) before the Metaspace pass instead.
            SentencePieceNormalizationStep? chainNormalizer = null;
            if (root.TryGetProperty("normalizer", out JsonElement normalizerElement) &&
                normalizerElement.ValueKind == JsonValueKind.Object)
            {
                if (SentencePieceNormalizationStep.HasRichSteps(normalizerElement))
                {

View on GitHub (pinned to 7b76e69cf9)