dotnet/machinelearning · error · ArgumentOutOfRangeException

The maximum number of tokens must be greater than zero.

Error message

The maximum number of tokens must be greater than zero.

What it means

EncodeToIds validates that maxTokenCount, when specified, is greater than zero, throwing ArgumentOutOfRangeException otherwise. This limit caps how many token IDs are produced, so a non-positive value is meaningless.

Solutions

  1. Pass int.MaxValue (the default) when no limit is intended instead of 0.
  2. Validate that the configured maxTokenCount >= 1 before calling EncodeToIds.
  3. If the value comes from settings, clamp it: max = Math.Max(1, configuredMax).

Example fix

// before
var ids = tokenizer.EncodeToIds(text, new EncoderSettings { MaxTokenCount = 0 });
// after
int max = Math.Max(1, configuredMax); // or int.MaxValue for unlimited
var ids = tokenizer.EncodeToIds(text, new EncoderSettings { MaxTokenCount = max });
Defensive patterns

Strategy: validation

Validate before calling

int max = configuredMax;
if (max <= 0) max = int.MaxValue; // API has no '0 = unlimited' semantics
if (max <= 0) throw new InvalidOperationException("MaxTokenCount must be >= 1");

Try / catch

try { var ids = tokenizer.EncodeToIds(text, settings); }
catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "maxTokenCount") { throw new InvalidOperationException("EncoderSettings.MaxTokenCount must be positive", ex); }

Prevention

When it happens

Trigger: Calling any public EncodeToIds overload (including EncodeToIds(text, settings) via EncoderSettings.MaxTokenCount) with maxTokenCount <= 0, e.g. 0, -1, or a computed value that underflowed.

Common situations: Computing MaxTokenCount from config where a sentinel 0 means 'unlimited' by the developer's convention but the API expects int.MaxValue for unlimited; user-supplied limit from UI/query string not validated; integer arithmetic that produced a negative value.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/92c8a4fa94eec223. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs:414

            ArrayPool<int>.Shared.Return(indexMapping);
            return result;
        }

        /// <summary>
        /// Encodes input text to token Ids.
        /// </summary>
        /// <param name="text">The text to encode.</param>
        /// <param name="textSpan">The span of the text to encode which will be used if the <paramref name="text"/> is <see langword="null"/>.</param>
        /// <param name="settings">The settings used to encode the text.</param>
        /// <returns>The encoded results containing the list of encoded Ids.</returns>
        protected override EncodeResults<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, EncodeSettings settings)
            => EncodeToIds(text, textSpan, settings.ConsiderPreTokenization, settings.ConsiderNormalization, settings.MaxTokenCount);

        private EncodeResults<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, bool considerPreTokenization, bool considerNormalization, int maxTokenCount = int.MaxValue)
        {
            if (maxTokenCount <= 0)
            {
                throw new ArgumentOutOfRangeException(nameof(maxTokenCount), "The maximum number of tokens must be greater than zero.");
            }

            if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
            {
                return new EncodeResults<int> { Tokens = [], NormalizedText = null, CharsConsumed = 0 };
            }

            IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(
                                                                text,
                                                                textSpan,
                                                                considerPreTokenization,
                                                                considerNormalization,
                                                                _normalizer,
                                                                _preTokenizer,
                                                                out string? normalizedText,
                                                                out ReadOnlySpan<char> textSpanToEncode,
                                                                out _);

View on GitHub (pinned to 7b76e69cf9)