dotnet/machinelearning · error · ArgumentOutOfRangeException

The maximum number of tokens must be greater than zero.

Error message

The maximum number of tokens must be greater than zero.

What it means

SentencePieceBpeModel.EncodeToIds validates maxTokenCount and throws ArgumentOutOfRangeException when it is zero or negative, since at least one token must be allowed. maxTokenCount defaults to int.MaxValue, so this only happens when explicitly set.

Solutions

  1. Pass maxTokenCount >= 1 (or omit it to use the int.MaxValue default)
  2. Clamp computed budgets: Math.Max(1, remaining)
  3. Validate configuration values for max token length at startup

Example fix

// before
var ids = tok.EncodeToIds(text, maxTokenCount: remaining); // remaining == 0
// after
var ids = tok.EncodeToIds(text, maxTokenCount: Math.Max(1, remaining));
Defensive patterns

Strategy: validation

Validate before calling

if (maxTokenCount <= 0) throw new InvalidOperationException("maxTokenCount must be at least 1.");

Try / catch

try { var ids = tok.EncodeToIds(text, maxTokenCount: limit); } catch (ArgumentOutOfRangeException ex) { log.LogError(ex, "maxTokenCount must be > 0"); throw; }

Prevention

When it happens

Trigger: Calling EncodeToIds with maxTokenCount: 0 or a negative value, often from a computed limit (e.g. budget left = maxTokens - usedTokens which hit 0 or went negative).

Common situations: Config-driven max-length settings that are 0 or unset; arithmetic computing a remaining budget that underflows; UI inputs defaulting to 0.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/67aa6d7cc5bda622. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/SentencePieceBpeModel.cs:286

                if (id.Type != (byte)ModelProto.Types.SentencePiece.Types.Type.Unused ||
                    revMerge is null ||
                    !revMerge.TryGetValue((pieceSpan.Index, pieceSpan.Length), out (int LeftIndex, int LeftLen, int RightIndex, int RightLen) merge))
                {
                    tokens.Add(new EncodedToken(id.Id, text.Slice(pieceSpan.Index, pieceSpan.Length).ToString(), new Range(pieceSpan.Index, pieceSpan.Index + pieceSpan.Length)));
                    return;
                }

                Segment((merge.LeftIndex, merge.LeftLen), text);
                Segment((merge.RightIndex, merge.RightLen), text);
            }
        }

        public override IReadOnlyList<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, bool addBeginningOfSentence, bool addEndOfSentence, bool considerNormalization,
                                        out string? normalizedText, out int charsConsumed, int maxTokenCount = int.MaxValue)
        {
            if (maxTokenCount <= 0)
            {
                throw new ArgumentOutOfRangeException(nameof(maxTokenCount), "The maximum number of tokens must be greater than zero.");
            }

            if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
            {
                normalizedText = null;
                charsConsumed = 0;
                return [];
            }

            return EncodeToIds(text is null ? textSpan : text.AsSpan(), addBeginningOfSentence, addEndOfSentence, considerNormalization, out normalizedText, out charsConsumed, maxTokenCount);
        }

        /// <summary>
        /// Encodes input text to token Ids up to maximum number of tokens.
        /// </summary>
        /// <param name="text">The text to encode.</param>
        /// <param name="addBeginningOfSentence">Indicate emitting the beginning of sentence token during the encoding.</param>
        /// <param name="addEndOfSentence">Indicate emitting the end of sentence token during the encoding.</param>

View on GitHub (pinned to 7b76e69cf9)