dotnet/machinelearning · error · ArgumentOutOfRangeException

The max token count must be greater than 0.

Error message

The max token count must be greater than 0.

What it means

WordPieceTokenizer.GetIndexByTokenCount (and the index-of-token-count overloads) validate settings.MaxTokenCount > 0 and throw ArgumentOutOfRangeException with this message otherwise. It computes the character index at which encoding should stop, requiring a valid positive token budget.

Solutions

  1. Pass a positive MaxTokenCount in the EncodeSettings
  2. Fix the upstream limit computation so it yields at least 1
  3. Early-return in your own code when the configured limit is 0 instead of calling the tokenizer

Example fix

// before
var idx = tokenizer.GetIndexByTokenCount(text, new EncodeSettings { MaxTokenCount = 0 }, out _, out _);
// after
var idx = tokenizer.GetIndexByTokenCount(text, new EncodeSettings { MaxTokenCount = 512 }, out _, out _);
Defensive patterns

Strategy: validation

Validate before calling

if (settings.MaxTokenCount <= 0) throw new ArgumentException("MaxTokenCount must be positive before calling GetIndexByTokenCount");

Type guard

bool IsValidMaxTokenCount(int n) => n > 0;

Try / catch

try { var idx = tokenizer.GetIndexByTokenCount(text, settings, fromEnd, out _, out _); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "settings.MaxTokenCount") { /* clamp limit and retry */ }

Prevention

When it happens

Trigger: Calling tokenizer.GetIndexByTokenCount / GetIndexFromEndByTokenCount with a settings object whose MaxTokenCount is 0 or negative.

Common situations: Truncating text to fit a max length where the limit came back as 0 from an empty config field; copying settings objects without copying MaxTokenCount; passing 0 to mean 'return immediately'.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/0e052221335a41e7. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs:606

        /// </summary>
        /// <param name="text">The text to encode.</param>
        /// <param name="textSpan">The span of the text to encode which will be used if the <paramref name="text"/> is <see langword="null"/>.</param>
        /// <param name="settings">The settings used to encode the text.</param>
        /// <param name="fromEnd">Indicate whether to find the index from the end of the text.</param>
        /// <param name="normalizedText">If the tokenizer's normalization is enabled or <paramRef name="settings" /> has <see cref="EncodeSettings.ConsiderNormalization"/> is <see langword="false"/>, this will be set to <paramRef name="text" /> in its normalized form; otherwise, this value will be set to <see langword="null"/>.</param>
        /// <param name="tokenCount">The token count can be generated which should be smaller than the maximum token count.</param>
        /// <returns>
        /// The index of the maximum encoding capacity within the processed text without surpassing the token limit.
        /// If <paramRef name="fromEnd" /> is <see langword="false"/>, it represents the index immediately following the last character to be included. In cases where no tokens fit, the result will be 0; conversely,
        /// if all tokens fit, the result will be length of the input text or the <paramref name="normalizedText"/> if the normalization is enabled.
        /// If <paramRef name="fromEnd" /> is <see langword="true"/>, it represents the index of the first character to be included. In cases where no tokens fit, the result will be the text length; conversely,
        /// if all tokens fit, the result will be zero.
        /// </returns>
        protected override int GetIndexByTokenCount(string? text, ReadOnlySpan<char> textSpan, EncodeSettings settings, bool fromEnd, out string? normalizedText, out int tokenCount)
        {
            if (settings.MaxTokenCount <= 0)
            {
                throw new ArgumentOutOfRangeException(nameof(settings.MaxTokenCount), "The max token count must be greater than 0.");
            }

            if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)
            {
                normalizedText = null;
                tokenCount = 0;
                return 0;
            }

            IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(
                                                                text,
                                                                textSpan,
                                                                settings.ConsiderNormalization,
                                                                settings.ConsiderNormalization,
                                                                _normalizer,
                                                                _preTokenizer,
                                                                out normalizedText,
                                                                out ReadOnlySpan<char> textSpanToEncode,

View on GitHub (pinned to 7b76e69cf9)