dotnet/machinelearning · error · ArgumentOutOfRangeException

The maximum number of characters per word must be greater…

Error message

The maximum number of characters per word must be greater than zero.

What it means

The WordPieceTokenizer constructor (via WordPieceTokenizerOptions) requires MaxInputCharsPerWord to be a positive integer because wordpieces longer than this limit are replaced by the unknown token. A value of zero or negative is nonsensical and throws ArgumentOutOfRangeException naming options.MaxInputCharsPerWord.

Solutions

  1. Set WordPieceTokenizerOptions.MaxInputCharsPerWord to a positive value (BERT uses 100; the library default is 200)
  2. If constructing options dynamically, initialize the property explicitly rather than relying on default(int)
  3. Clamp with Math.Max(1, configuredValue) at config load time

Example fix

// before
var options = new WordPieceTokenizerOptions { Vocab = vocab }; // MaxInputCharsPerWord defaults to 0
// after
var options = new WordPieceTokenizerOptions { Vocab = vocab, MaxInputCharsPerWord = 200 };
Defensive patterns

Strategy: validation

Validate before calling

if (options.MaxInputCharsPerWord <= 0) throw new ArgumentException("MaxInputCharsPerWord must be positive", nameof(options));

Type guard

bool HasValidMaxChars(WordPieceTokenizerOptions o) => o.MaxInputCharsPerWord > 0;

Try / catch

try { var tok = new WordPieceTokenizer(options); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == "options.MaxInputCharsPerWord") { options.MaxInputCharsPerWord = 200; /* retry or fail fast */ }

Prevention

When it happens

Trigger: Constructing new WordPieceTokenizer(...) or new WordPieceTokenizerOptions with MaxInputCharsPerWord = 0 or a negative number, often left at default(int) when using object initializers on a hand-built options object.

Common situations: Forgetting to set MaxInputCharsPerWord on a manually instantiated options object (default 0); copying config from BERT code where the limit was configured as 0/disabled; off-by-one when intending a large limit.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/1cb8657eed061740. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs:59

            options ??= new();

            SpecialTokens = options.SpecialTokens;
            SpecialTokensReverse = options.SpecialTokens is not null ? options.SpecialTokens.GroupBy(kvp => kvp.Value).ToDictionary(g => g.Key, g => g.First().Key) : null;

            if (options.UnknownToken is null)
            {
                throw new ArgumentNullException(nameof(options.UnknownToken));
            }

            if (options.ContinuingSubwordPrefix is null)
            {
                throw new ArgumentNullException(nameof(options.ContinuingSubwordPrefix));
            }

            if (options.MaxInputCharsPerWord <= 0)
            {
                throw new ArgumentOutOfRangeException(nameof(options.MaxInputCharsPerWord), "The maximum number of characters per word must be greater than zero.");
            }

            if (!vocab!.TryGetValue(options.UnknownToken, out int id))
            {
                throw new ArgumentException($"The unknown token '{options.UnknownToken}' is not in the vocabulary.");
            }

            UnknownToken = options.UnknownToken;
            UnknownTokenId = id;
            ContinuingSubwordPrefix = options.ContinuingSubwordPrefix;
            MaxInputCharsPerWord = options.MaxInputCharsPerWord;

            _preTokenizer = options.PreTokenizer ?? PreTokenizer.CreateWhiteSpace(options.SpecialTokens);
            _normalizer = options.Normalizer;
        }

        /// <summary>
        /// Gets the unknown token ID.

View on GitHub (pinned to 7b76e69cf9)