dotnet/machinelearning · error · ArgumentOutOfRangeException
The max token count must be greater than 0.
Error message
The max token count must be greater than 0.
What it means
WordPieceTokenizer.GetIndexByTokenCount (and the index-of-token-count overloads) validate settings.MaxTokenCount > 0 and throw ArgumentOutOfRangeException with this message otherwise. It computes the character index at which encoding should stop, requiring a valid positive token budget.
Solutions
- Pass a positive MaxTokenCount in the EncodeSettings
- Fix the upstream limit computation so it yields at least 1
- Early-return in your own code when the configured limit is 0 instead of calling the tokenizer
Example fix
// before
var idx = tokenizer.GetIndexByTokenCount(text, new EncodeSettings { MaxTokenCount = 0 }, out _, out _);
// after
var idx = tokenizer.GetIndexByTokenCount(text, new EncodeSettings { MaxTokenCount = 512 }, out _, out _); Defensive patterns
Strategy: validation
Validate before calling
if (settings.MaxTokenCount <= 0) throw new ArgumentException("MaxTokenCount must be positive before calling GetIndexByTokenCount"); Type guard
bool IsValidMaxTokenCount(int n) => n > 0;
Try / catch
try { var idx = tokenizer.GetIndexByTokenCount(text, settings, fromEnd, out _, out _); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "settings.MaxTokenCount") { /* clamp limit and retry */ } Prevention
- Clamp computed truncation limits to Math.Max(1, limit)
- Skip the call entirely when the limit is 0 (return index 0)
- Share one validated EncodeSettings instance across encode/count/index calls
When it happens
Trigger: Calling tokenizer.GetIndexByTokenCount / GetIndexFromEndByTokenCount with a settings object whose MaxTokenCount is 0 or negative.
Common situations: Truncating text to fit a max length where the limit came back as 0 from an empty config field; copying settings objects without copying MaxTokenCount; passing 0 to mean 'return immediately'.
Related errors
- The maximum number of characters per word must be greater…
- The maximum number of tokens must be greater than zero.
- The maximum number of tokens must be greater than zero.
- A tokenizer.json normalizer entry must be a JSON object.
- ArgumentNullException
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/0e052221335a41e7.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs:606
/// </summary>
/// <param name="text">The text to encode.</param>
/// <param name="textSpan">The span of the text to encode which will be used if the <paramref name="text"/> is <see langword="null"/>.</param>
/// <param name="settings">The settings used to encode the text.</param>
/// <param name="fromEnd">Indicate whether to find the index from the end of the text.</param>
/// <param name="normalizedText">If the tokenizer's normalization is enabled or <paramRef name="settings" /> has <see cref="EncodeSettings.ConsiderNormalization"/> is <see langword="false"/>, this will be set to <paramRef name="text" /> in its normalized form; otherwise, this value will be set to <see langword="null"/>.</param>
/// <param name="tokenCount">The token count can be generated which should be smaller than the maximum token count.</param>
/// <returns>
/// The index of the maximum encoding capacity within the processed text without surpassing the token limit.
/// If <paramRef name="fromEnd" /> is <see langword="false"/>, it represents the index immediately following the last character to be included. In cases where no tokens fit, the result will be 0; conversely,
/// if all tokens fit, the result will be length of the input text or the <paramref name="normalizedText"/> if the normalization is enabled.
/// If <paramRef name="fromEnd" /> is <see langword="true"/>, it represents the index of the first character to be included. In cases where no tokens fit, the result will be the text length; conversely,
/// if all tokens fit, the result will be zero.
/// </returns>
protected override int GetIndexByTokenCount(string? text, ReadOnlySpan<char> textSpan, EncodeSettings settings, bool fromEnd, out string? normalizedText, out int tokenCount)
{
if (settings.MaxTokenCount <= 0)
{
throw new ArgumentOutOfRangeException(nameof(settings.MaxTokenCount), "The max token count must be greater than 0.");
}
if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
{
normalizedText = null;
tokenCount = 0;
return 0;
}
IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(
text,
textSpan,
settings.ConsiderNormalization,
settings.ConsiderNormalization,
_normalizer,
_preTokenizer,
out normalizedText,
out ReadOnlySpan<char> textSpanToEncode,View on GitHub (pinned to 7b76e69cf9)