dotnet/machinelearning · error · ArgumentOutOfRangeException
The maximum number of tokens must be greater than zero.
Error message
The maximum number of tokens must be greater than zero.
What it means
SentencePieceBpeModel.EncodeToIds validates maxTokenCount and throws ArgumentOutOfRangeException when it is zero or negative, since at least one token must be allowed. maxTokenCount defaults to int.MaxValue, so this only happens when explicitly set.
Solutions
- Pass maxTokenCount >= 1 (or omit it to use the int.MaxValue default)
- Clamp computed budgets: Math.Max(1, remaining)
- Validate configuration values for max token length at startup
Example fix
// before var ids = tok.EncodeToIds(text, maxTokenCount: remaining); // remaining == 0 // after var ids = tok.EncodeToIds(text, maxTokenCount: Math.Max(1, remaining));
Defensive patterns
Strategy: validation
Validate before calling
if (maxTokenCount <= 0) throw new InvalidOperationException("maxTokenCount must be at least 1."); Try / catch
try { var ids = tok.EncodeToIds(text, maxTokenCount: limit); } catch (ArgumentOutOfRangeException ex) { log.LogError(ex, "maxTokenCount must be > 0"); throw; } Prevention
- Clamp computed budgets to at least 1
- Validate configured max-length values at startup
- Rely on the int.MaxValue default unless truncation is intended
When it happens
Trigger: Calling EncodeToIds with maxTokenCount: 0 or a negative value, often from a computed limit (e.g. budget left = maxTokens - usedTokens which hit 0 or went negative).
Common situations: Config-driven max-length settings that are 0 or unset; arithmetic computing a remaining budget that underflows; UI inputs defaulting to 0.
Related errors
- The maximum number of tokens must be greater than zero.
- The max token count must be greater than 0.
- The maximum number of characters per word must be greater…
- A tokenizer.json normalizer entry must be a JSON object.
- ArgumentNullException
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/67aa6d7cc5bda622.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/SentencePieceBpeModel.cs:286
if (id.Type != (byte)ModelProto.Types.SentencePiece.Types.Type.Unused ||
revMerge is null ||
!revMerge.TryGetValue((pieceSpan.Index, pieceSpan.Length), out (int LeftIndex, int LeftLen, int RightIndex, int RightLen) merge))
{
tokens.Add(new EncodedToken(id.Id, text.Slice(pieceSpan.Index, pieceSpan.Length).ToString(), new Range(pieceSpan.Index, pieceSpan.Index + pieceSpan.Length)));
return;
}
Segment((merge.LeftIndex, merge.LeftLen), text);
Segment((merge.RightIndex, merge.RightLen), text);
}
}
public override IReadOnlyList<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, bool addBeginningOfSentence, bool addEndOfSentence, bool considerNormalization,
out string? normalizedText, out int charsConsumed, int maxTokenCount = int.MaxValue)
{
if (maxTokenCount <= 0)
{
throw new ArgumentOutOfRangeException(nameof(maxTokenCount), "The maximum number of tokens must be greater than zero.");
}
if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
{
normalizedText = null;
charsConsumed = 0;
return [];
}
return EncodeToIds(text is null ? textSpan : text.AsSpan(), addBeginningOfSentence, addEndOfSentence, considerNormalization, out normalizedText, out charsConsumed, maxTokenCount);
}
/// <summary>
/// Encodes input text to token Ids up to maximum number of tokens.
/// </summary>
/// <param name="text">The text to encode.</param>
/// <param name="addBeginningOfSentence">Indicate emitting the beginning of sentence token during the encoding.</param>
/// <param name="addEndOfSentence">Indicate emitting the end of sentence token during the encoding.</param>View on GitHub (pinned to 7b76e69cf9)