{"record":{"id":"18dfa7b886fd3cd8","repo":"dotnet/machinelearning","slug":"the-tokenizer-json-model-unk-id-unkid-is-out","errorCode":null,"errorMessage":"The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces.","messagePattern":"The tokenizer\\.json model 'unk_id' \\((.+?)\\) is out of range for a vocabulary of (.+?) pieces\\.","errorType":"validation","errorClass":"InvalidDataException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/SentencePieceTokenizer.cs","lineNumber":661,"sourceCode":"                    throw new InvalidDataException(\"A piece string in 'model.vocab' is null.\");\n                }\n\n                vocab.Add((piece, entry[1].GetSingle()));\n            }\n\n            if (unkIsNull)\n            {\n                // Without an unknown token the only way to represent out-of-vocabulary input is byte fallback; a model\n                // with neither cannot encode OOV text, so reject that combination up front rather than emitting an\n                // invalid token id at encode time.\n                if (!byteFallback)\n                {\n                    throw new NotSupportedException(\"The tokenizer.json model has a null 'unk_id' but does not enable 'byte_fallback'; a Unigram model without an unknown token is only supported when byte_fallback is enabled.\");\n                }\n            }\n            else if (unkId < 0 || unkId >= vocab.Count)\n            {\n                throw new InvalidDataException($\"The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces.\");\n            }\n\n            // Extract normalizer settings\n            byte[]? precompiledCharsMap = null;\n            bool addDummyPrefix = true;\n            // HF tokenizer.json has no remove_extra_whitespaces flag; SpmConverter encodes that behavior as\n            // explicit normalizer steps (a right-Strip plus a Replace collapsing runs of spaces). Deduce it from\n            // those steps, defaulting to false when absent to match the HF fast-tokenizer runtime.\n            bool removeExtraWhitespaces = false;\n            // When the normalizer has content-modifying steps that the charsmap + removeExtraWhitespaces\n            // approximation cannot represent (per-character Replace, Lowercase, Unicode normalization, ...),\n            // apply the full normalizer chain (charsmap included) before the Metaspace pass instead.\n            SentencePieceNormalizationStep? chainNormalizer = null;\n            if (root.TryGetProperty(\"normalizer\", out JsonElement normalizerElement) &&\n                normalizerElement.ValueKind == JsonValueKind.Object)\n            {\n                if (SentencePieceNormalizationStep.HasRichSteps(normalizerElement))\n                {","sourceCodeStart":643,"sourceCodeEnd":679,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/SentencePieceTokenizer.cs#L643-L679","documentation":"A numeric unk_id must index a valid entry in the parsed vocabulary. If unk_id is negative or >= the number of vocab pieces, it cannot refer to any piece, so InvalidDataException is thrown before the tokenizer is constructed.","triggerScenarios":"CreateFromTokenizerJson where model.unk_id is, e.g., 3 but vocab contains only 2 pieces, or unk_id is -1 while not null (note: the loader uses -1 internally only for null unk_id).","commonSituations":"Editing the vocab (removing <unk>) without updating unk_id; tokenizer.json files whose unk_id was written for a different/larger vocabulary; mixing the model section of one file with the vocab of another.","solutions":["Set unk_id to the actual index of the '<unk>' piece within model.vocab (0-based).","If the unknown token was intentionally removed, set unk_id to null and enable byte_fallback instead.","Re-export tokenizer.json from the original model so unk_id and vocab stay consistent."],"exampleFix":"// before\n\"unk_id\": 3, \"vocab\": [[\"<unk>\", 0.0], [\"a\", -1.0]]  // only 2 pieces\n// after\n\"unk_id\": 0, \"vocab\": [[\"<unk>\", 0.0], [\"a\", -1.0]]","handlingStrategy":"validation","validationCode":"int vocabCount = model.GetProperty(\"vocab\").GetArrayLength();\nvar unkId = model.GetProperty(\"unk_id\");\nif (unkId.ValueKind == JsonValueKind.Number &&\n    (unkId.GetInt32() < 0 || unkId.GetInt32() >= vocabCount))\n    throw new InvalidOperationException($\"unk_id {unkId.GetInt32()} out of range for {vocabCount} pieces.\");","typeGuard":null,"tryCatchPattern":"try { var tok = SentencePieceTokenizer.CreateFromTokenizerJson(stream); }\ncatch (InvalidDataException ex) when (ex.Message.Contains(\"out of range\")) { /* fix unk_id or restore full vocab */ }","preventionTips":["Update unk_id whenever vocab is edited or reordered.","Never mix model sections from different tokenizer.json files.","Assert unk_id < vocab length in a pre-load sanity check."],"tags":["tokenizer","json-validation","out-of-range"],"backgroundTag":"value-out-of-range","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}