dotnet/machinelearning · error · ArgumentOutOfRangeException
The maximum number of tokens must be greater than zero.
Error message
The maximum number of tokens must be greater than zero.
What it means
WordPieceTokenizer.EncodeToIds validates settings.MaxTokenCount and throws ArgumentOutOfRangeException when it is <= 0, since a non-positive cap on output tokens has no meaning. This mirrors validation in all encoder paths of the tokenizer.
Solutions
- Pass a positive maxTokenCount, or int.MaxValue for unlimited encoding
- Fix the configuration/source so the limit defaults to a positive value
- Guard the call site: if (limit <= 0) limit = int.MaxValue;
Example fix
// before var ids = tokenizer.EncodeToIds(text, 0); // after var ids = tokenizer.EncodeToIds(text, int.MaxValue);
Defensive patterns
Strategy: validation
Validate before calling
if (maxTokenCount <= 0) maxTokenCount = int.MaxValue; // unlimited
Type guard
bool IsValidMaxTokenCount(int n) => n > 0;
Try / catch
try { var res = tokenizer.EncodeToIds(text, maxTokenCount); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "settings.MaxTokenCount") { var res2 = tokenizer.EncodeToIds(text, int.MaxValue); } Prevention
- Use int.MaxValue, never 0, to mean 'no limit'
- Validate configured max-length values at startup
- Wrap encode calls in a helper that normalizes the limit
When it happens
Trigger: Calling tokenizer.EncodeToIds(text, maxTokenCount: 0) or EncodeToIds(text, settings) with settings.MaxTokenCount = 0 or negative — commonly the default int value when a MaxTokenCount variable was never assigned.
Common situations: Passing 0 intending 'no limit' (must instead use int.MaxValue); reading maxTokenCount from config where it defaulted to 0; uninitialized field in a wrapper class.
Related errors
- The maximum number of tokens must be greater than zero.
- The max token count must be greater than 0.
- The maximum number of characters per word must be greater…
- A tokenizer.json normalizer entry must be a JSON object.
- ArgumentNullException
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/cdf140b768652180.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs:394
if (arrayPool is not null)
{
ArrayPool<char>.Shared.Return(arrayPool);
}
}
/// <summary>
/// Encodes input text to token Ids.
/// </summary>
/// <param name="text">The text to encode.</param>
/// <param name="textSpan">The span of the text to encode which will be used if the <paramref name="text"/> is <see langword="null"/>.</param>
/// <param name="settings">The settings used to encode the text.</param>
/// <returns>The encoded results containing the list of encoded Ids.</returns>
protected override EncodeResults<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, EncodeSettings settings)
{
int maxTokenCount = settings.MaxTokenCount;
if (maxTokenCount <= 0)
{
throw new ArgumentOutOfRangeException(nameof(settings.MaxTokenCount), "The maximum number of tokens must be greater than zero.");
}
if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
{
return new EncodeResults<int> { NormalizedText = null, Tokens = [], CharsConsumed = 0 };
}
IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(
text,
textSpan,
settings.ConsiderPreTokenization,
settings.ConsiderNormalization,
_normalizer,
_preTokenizer,
out string? normalizedText,
out ReadOnlySpan<char> textSpanToEncode,
out int charsConsumed);
View on GitHub (pinned to 7b76e69cf9)