dotnet/machinelearning · error · ArgumentOutOfRangeException

The maximum number of tokens must be greater than zero.

Error message

The maximum number of tokens must be greater than zero.

What it means

WordPieceTokenizer.EncodeToIds validates settings.MaxTokenCount and throws ArgumentOutOfRangeException when it is <= 0, since a non-positive cap on output tokens has no meaning. This mirrors validation in all encoder paths of the tokenizer.

Solutions

  1. Pass a positive maxTokenCount, or int.MaxValue for unlimited encoding
  2. Fix the configuration/source so the limit defaults to a positive value
  3. Guard the call site: if (limit <= 0) limit = int.MaxValue;

Example fix

// before
var ids = tokenizer.EncodeToIds(text, 0);
// after
var ids = tokenizer.EncodeToIds(text, int.MaxValue);
Defensive patterns

Strategy: validation

Validate before calling

if (maxTokenCount <= 0) maxTokenCount = int.MaxValue; // unlimited

Type guard

bool IsValidMaxTokenCount(int n) => n > 0;

Try / catch

try { var res = tokenizer.EncodeToIds(text, maxTokenCount); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "settings.MaxTokenCount") { var res2 = tokenizer.EncodeToIds(text, int.MaxValue); }

Prevention

When it happens

Trigger: Calling tokenizer.EncodeToIds(text, maxTokenCount: 0) or EncodeToIds(text, settings) with settings.MaxTokenCount = 0 or negative — commonly the default int value when a MaxTokenCount variable was never assigned.

Common situations: Passing 0 intending 'no limit' (must instead use int.MaxValue); reading maxTokenCount from config where it defaulted to 0; uninitialized field in a wrapper class.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/cdf140b768652180. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs:394

            if (arrayPool is not null)
            {
                ArrayPool<char>.Shared.Return(arrayPool);
            }
        }

        /// <summary>
        /// Encodes input text to token Ids.
        /// </summary>
        /// <param name="text">The text to encode.</param>
        /// <param name="textSpan">The span of the text to encode which will be used if the <paramref name="text"/> is <see langword="null"/>.</param>
        /// <param name="settings">The settings used to encode the text.</param>
        /// <returns>The encoded results containing the list of encoded Ids.</returns>
        protected override EncodeResults<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, EncodeSettings settings)
        {
            int maxTokenCount = settings.MaxTokenCount;
            if (maxTokenCount <= 0)
            {
                throw new ArgumentOutOfRangeException(nameof(settings.MaxTokenCount), "The maximum number of tokens must be greater than zero.");
            }

            if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
            {
                return new EncodeResults<int> { NormalizedText = null, Tokens = [], CharsConsumed = 0 };
            }

            IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(
                                                                text,
                                                                textSpan,
                                                                settings.ConsiderPreTokenization,
                                                                settings.ConsiderNormalization,
                                                                _normalizer,
                                                                _preTokenizer,
                                                                out string? normalizedText,
                                                                out ReadOnlySpan<char> textSpanToEncode,
                                                                out int charsConsumed);

View on GitHub (pinned to 7b76e69cf9)