dotnet/machinelearning · error · ArgumentException
The unknown token ' ' is not in the vocabulary.
Error message
The unknown token '{options.UnknownToken}' is not in the vocabulary. What it means
WordPieceTokenizer must map words it cannot tokenize to a known [UNK]-style token; the constructor looks up options.UnknownToken in the supplied vocab and throws ArgumentException if it is absent. Every WordPiece vocabulary must contain the unknown token id.
Solutions
- Ensure the vocab contains options.UnknownToken exactly (default "[UNK]"), or set UnknownToken to the string your vocab actually uses (e.g. "<unk>")
- Verify you loaded the vocab file matching the tokenizer's expected format
- If building vocab in code, add options.UnknownToken as a key before constructing the tokenizer
Example fix
// before
var tok = new WordPieceTokenizer(new WordPieceTokenizerOptions { Vocab = vocab }); // vocab lacks "[UNK]"
// after
var tok = new WordPieceTokenizer(new WordPieceTokenizerOptions { Vocab = vocab, UnknownToken = "<unk>" }); Defensive patterns
Strategy: validation
Validate before calling
if (!vocab.ContainsKey(options.UnknownToken)) throw new ArgumentException($"Vocab is missing unknown token '{options.UnknownToken}'"); Type guard
bool VocabHasUnknownToken(WordPieceTokenizerOptions o) => o.Vocab is not null && o.Vocab.ContainsKey(o.UnknownToken);
Try / catch
try { var tok = new WordPieceTokenizer(options); } catch (ArgumentException ex) when (ex.Message.Contains("is not in the vocabulary")) { log.LogError(ex, "Vocab missing unknown token {Token}", options.UnknownToken); } Prevention
- Verify the vocab file contains the unknown token marker before loading
- Match UnknownToken to your model's convention ([UNK] vs <unk>)
- Keep vocab and options paired together per model, not mixed across models
When it happens
Trigger: Creating a WordPieceTokenizer whose WordPieceTokenizerOptions.UnknownToken (default "[UNK]") does not exist as a key in the provided vocab dictionary — e.g. a vocab using a different unknown marker like "<unk>" or a truncated vocab file.
Common situations: Loading a third-party BERT vocab where the unknown token is spelled differently; building a vocab programmatically and forgetting the unknown entry; passing the wrong vocab file to the options.
Understand the failure class
Background: 'Could not be found', 'does not exist', 'not found in database': the resource-not-found family when an ID, slug, key, or URI lookup comes back empty — this error's family across 20 libraries.
Related errors
- The maximum number of characters per word must be greater…
- The vocabulary cannot be null.
- Unknown Token Out Of Vocabulary.
- Unknown Token ' ' was not present in ' '.
- A tokenizer.json normalizer entry must be a JSON object.
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/c4dda0101f8557e8.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs:64
if (options.UnknownToken is null)
{
throw new ArgumentNullException(nameof(options.UnknownToken));
}
if (options.ContinuingSubwordPrefix is null)
{
throw new ArgumentNullException(nameof(options.ContinuingSubwordPrefix));
}
if (options.MaxInputCharsPerWord <= 0)
{
throw new ArgumentOutOfRangeException(nameof(options.MaxInputCharsPerWord), "The maximum number of characters per word must be greater than zero.");
}
if (!vocab!.TryGetValue(options.UnknownToken, out int id))
{
throw new ArgumentException($"The unknown token '{options.UnknownToken}' is not in the vocabulary.");
}
UnknownToken = options.UnknownToken;
UnknownTokenId = id;
ContinuingSubwordPrefix = options.ContinuingSubwordPrefix;
MaxInputCharsPerWord = options.MaxInputCharsPerWord;
_preTokenizer = options.PreTokenizer ?? PreTokenizer.CreateWhiteSpace(options.SpecialTokens);
_normalizer = options.Normalizer;
}
/// <summary>
/// Gets the unknown token ID.
/// A token that is not in the vocabulary cannot be converted to an ID and is set to be this token instead.
/// </summary>
public int UnknownTokenId { get; }
/// <summary>View on GitHub (pinned to 7b76e69cf9)