dotnet/machinelearning · error · ArgumentOutOfRangeException

The max token count must be greater than 0.

Error message

The max token count must be greater than 0.

What it means

LastIndexOf (used by GetIndexByTokenCount when searching backwards) requires maxTokenCount > 0 and throws ArgumentOutOfRangeException otherwise. Note the message wording differs slightly from the EncodeToIds/CountTokens validation ('must be greater than 0').

Solutions

  1. Ensure the token count passed is at least 1 before calling GetIndexByTokenCount.
  2. Handle the zero case at the call site (there is nothing to search when the limit is 0) instead of calling the API.
  3. Clamp: max = Math.Max(1, requestedCount).

Example fix

// before
int idx = tokenizer.GetIndexByTokenCount(text, false, false, out _, maxTokenCount: remaining); // remaining could be 0
// after
if (remaining <= 0) return -1; // nothing to search
int idx = tokenizer.GetIndexByTokenCount(text, false, false, out _, maxTokenCount: remaining);
Defensive patterns

Strategy: validation

Validate before calling

if (maxTokenCount <= 0)
{
    if (maxTokenCount == 0) return -1; // nothing to search
    throw new InvalidOperationException("maxTokenCount must be >= 1");
}

Try / catch

try { int idx = tokenizer.GetIndexByTokenCount(text, false, false, out _, maxTokenCount: n); }
catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "maxTokenCount") { throw new InvalidOperationException("Backward token count must be positive", ex); }

Prevention

When it happens

Trigger: Calling GetIndexByTokenCount with a text-splitting/token-count value of 0 or negative, which is forwarded to LastIndexOf as maxTokenCount.

Common situations: Computing the 'last N tokens' N from a remainder calculation that yielded 0; config value of 0 meant as 'disable' but the API has no such semantics; off-by-one in a loop decrementing the count.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/0118cbf8977cdc36. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs:546

                    if (length < split.Length || count >= maxTokenCount)
                    {
                        break;
                    }
                }
            }
            else
            {
                count += EncodeToIdsInternal(textSpanToEncode, null, out charsConsumed, maxTokenCount);
            }

            return count;
        }

        private int LastIndexOf(string? text, ReadOnlySpan<char> textSpan, int maxTokenCount, bool considerPreTokenization, bool considerNormalization, out string? normalizedText, out int tokenCount)
        {
            if (maxTokenCount <= 0)
            {
                throw new ArgumentOutOfRangeException(nameof(maxTokenCount), "The max token count must be greater than 0.");
            }

            if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
            {
                normalizedText = null;
                tokenCount = 0;
                return 0;
            }

            IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(
                                                                text,
                                                                textSpan,
                                                                considerPreTokenization,
                                                                considerNormalization,
                                                                _normalizer,
                                                                _preTokenizer,
                                                                out normalizedText,
                                                                out ReadOnlySpan<char> textSpanToEncode,

View on GitHub (pinned to 7b76e69cf9)