dotnet/machinelearning · error · ArgumentOutOfRangeException
The maximum number of characters per word must be greater…
Error message
The maximum number of characters per word must be greater than zero.
What it means
The WordPieceTokenizer constructor (via WordPieceTokenizerOptions) requires MaxInputCharsPerWord to be a positive integer because wordpieces longer than this limit are replaced by the unknown token. A value of zero or negative is nonsensical and throws ArgumentOutOfRangeException naming options.MaxInputCharsPerWord.
Solutions
- Set WordPieceTokenizerOptions.MaxInputCharsPerWord to a positive value (BERT uses 100; the library default is 200)
- If constructing options dynamically, initialize the property explicitly rather than relying on default(int)
- Clamp with Math.Max(1, configuredValue) at config load time
Example fix
// before
var options = new WordPieceTokenizerOptions { Vocab = vocab }; // MaxInputCharsPerWord defaults to 0
// after
var options = new WordPieceTokenizerOptions { Vocab = vocab, MaxInputCharsPerWord = 200 }; Defensive patterns
Strategy: validation
Validate before calling
if (options.MaxInputCharsPerWord <= 0) throw new ArgumentException("MaxInputCharsPerWord must be positive", nameof(options)); Type guard
bool HasValidMaxChars(WordPieceTokenizerOptions o) => o.MaxInputCharsPerWord > 0;
Try / catch
try { var tok = new WordPieceTokenizer(options); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "options.MaxInputCharsPerWord") { options.MaxInputCharsPerWord = 200; /* retry or fail fast */ } Prevention
- Always set MaxInputCharsPerWord explicitly (BERT: 100, library default: 200)
- Never rely on default(int) for required positive options
- Validate tokenizer options at config-load time
When it happens
Trigger: Constructing new WordPieceTokenizer(...) or new WordPieceTokenizerOptions with MaxInputCharsPerWord = 0 or a negative number, often left at default(int) when using object initializers on a hand-built options object.
Common situations: Forgetting to set MaxInputCharsPerWord on a manually instantiated options object (default 0); copying config from BERT code where the limit was configured as 0/disabled; off-by-one when intending a large limit.
Related errors
- The max token count must be greater than 0.
- The maximum number of tokens must be greater than zero.
- The maximum number of tokens must be greater than zero.
- The unknown token ' ' is not in the vocabulary.
- A tokenizer.json normalizer entry must be a JSON object.
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/1cb8657eed061740.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs:59
options ??= new();
SpecialTokens = options.SpecialTokens;
SpecialTokensReverse = options.SpecialTokens is not null ? options.SpecialTokens.GroupBy(kvp => kvp.Value).ToDictionary(g => g.Key, g => g.First().Key) : null;
if (options.UnknownToken is null)
{
throw new ArgumentNullException(nameof(options.UnknownToken));
}
if (options.ContinuingSubwordPrefix is null)
{
throw new ArgumentNullException(nameof(options.ContinuingSubwordPrefix));
}
if (options.MaxInputCharsPerWord <= 0)
{
throw new ArgumentOutOfRangeException(nameof(options.MaxInputCharsPerWord), "The maximum number of characters per word must be greater than zero.");
}
if (!vocab!.TryGetValue(options.UnknownToken, out int id))
{
throw new ArgumentException($"The unknown token '{options.UnknownToken}' is not in the vocabulary.");
}
UnknownToken = options.UnknownToken;
UnknownTokenId = id;
ContinuingSubwordPrefix = options.ContinuingSubwordPrefix;
MaxInputCharsPerWord = options.MaxInputCharsPerWord;
_preTokenizer = options.PreTokenizer ?? PreTokenizer.CreateWhiteSpace(options.SpecialTokens);
_normalizer = options.Normalizer;
}
/// <summary>
/// Gets the unknown token ID.View on GitHub (pinned to 7b76e69cf9)