{"record":{"id":"c4dda0101f8557e8","repo":"dotnet/machinelearning","slug":"the-unknown-token-options-unknowntoken-is-not","errorCode":null,"errorMessage":"The unknown token '{options.UnknownToken}' is not in the vocabulary.","messagePattern":"The unknown token '(.+?)' is not in the vocabulary\\.","errorType":"exception","errorClass":"ArgumentException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs","lineNumber":64,"sourceCode":"\n            if (options.UnknownToken is null)\n            {\n                throw new ArgumentNullException(nameof(options.UnknownToken));\n            }\n\n            if (options.ContinuingSubwordPrefix is null)\n            {\n                throw new ArgumentNullException(nameof(options.ContinuingSubwordPrefix));\n            }\n\n            if (options.MaxInputCharsPerWord <= 0)\n            {\n                throw new ArgumentOutOfRangeException(nameof(options.MaxInputCharsPerWord), \"The maximum number of characters per word must be greater than zero.\");\n            }\n\n            if (!vocab!.TryGetValue(options.UnknownToken, out int id))\n            {\n                throw new ArgumentException($\"The unknown token '{options.UnknownToken}' is not in the vocabulary.\");\n            }\n\n            UnknownToken = options.UnknownToken;\n            UnknownTokenId = id;\n            ContinuingSubwordPrefix = options.ContinuingSubwordPrefix;\n            MaxInputCharsPerWord = options.MaxInputCharsPerWord;\n\n            _preTokenizer = options.PreTokenizer ?? PreTokenizer.CreateWhiteSpace(options.SpecialTokens);\n            _normalizer = options.Normalizer;\n        }\n\n        /// <summary>\n        /// Gets the unknown token ID.\n        /// A token that is not in the vocabulary cannot be converted to an ID and is set to be this token instead.\n        /// </summary>\n        public int UnknownTokenId { get; }\n\n        /// <summary>","sourceCodeStart":46,"sourceCodeEnd":82,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs#L46-L82","documentation":"WordPieceTokenizer must map words it cannot tokenize to a known [UNK]-style token; the constructor looks up options.UnknownToken in the supplied vocab and throws ArgumentException if it is absent. Every WordPiece vocabulary must contain the unknown token id.","triggerScenarios":"Creating a WordPieceTokenizer whose WordPieceTokenizerOptions.UnknownToken (default \"[UNK]\") does not exist as a key in the provided vocab dictionary — e.g. a vocab using a different unknown marker like \"<unk>\" or a truncated vocab file.","commonSituations":"Loading a third-party BERT vocab where the unknown token is spelled differently; building a vocab programmatically and forgetting the unknown entry; passing the wrong vocab file to the options.","solutions":["Ensure the vocab contains options.UnknownToken exactly (default \"[UNK]\"), or set UnknownToken to the string your vocab actually uses (e.g. \"<unk>\")","Verify you loaded the vocab file matching the tokenizer's expected format","If building vocab in code, add options.UnknownToken as a key before constructing the tokenizer"],"exampleFix":"// before\nvar tok = new WordPieceTokenizer(new WordPieceTokenizerOptions { Vocab = vocab }); // vocab lacks \"[UNK]\"\n// after\nvar tok = new WordPieceTokenizer(new WordPieceTokenizerOptions { Vocab = vocab, UnknownToken = \"<unk>\" });","handlingStrategy":"validation","validationCode":"if (!vocab.ContainsKey(options.UnknownToken)) throw new ArgumentException($\"Vocab is missing unknown token '{options.UnknownToken}'\");","typeGuard":"bool VocabHasUnknownToken(WordPieceTokenizerOptions o) => o.Vocab is not null && o.Vocab.ContainsKey(o.UnknownToken);","tryCatchPattern":"try { var tok = new WordPieceTokenizer(options); } catch (ArgumentException ex) when (ex.Message.Contains(\"is not in the vocabulary\")) { log.LogError(ex, \"Vocab missing unknown token {Token}\", options.UnknownToken); }","preventionTips":["Verify the vocab file contains the unknown token marker before loading","Match UnknownToken to your model's convention ([UNK] vs <unk>)","Keep vocab and options paired together per model, not mixed across models"],"tags":["tokenizers","vocabulary","wordpiece"],"backgroundTag":"resource-not-found","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}