{"record":{"id":"1cb8657eed061740","repo":"dotnet/machinelearning","slug":"the-maximum-number-of-characters-per-word-must-be","errorCode":null,"errorMessage":"The maximum number of characters per word must be greater than zero.","messagePattern":"The maximum number of characters per word must be greater than zero\\.","errorType":"exception","errorClass":"ArgumentOutOfRangeException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs","lineNumber":59,"sourceCode":"\n            options ??= new();\n\n            SpecialTokens = options.SpecialTokens;\n            SpecialTokensReverse = options.SpecialTokens is not null ? options.SpecialTokens.GroupBy(kvp => kvp.Value).ToDictionary(g => g.Key, g => g.First().Key) : null;\n\n            if (options.UnknownToken is null)\n            {\n                throw new ArgumentNullException(nameof(options.UnknownToken));\n            }\n\n            if (options.ContinuingSubwordPrefix is null)\n            {\n                throw new ArgumentNullException(nameof(options.ContinuingSubwordPrefix));\n            }\n\n            if (options.MaxInputCharsPerWord <= 0)\n            {\n                throw new ArgumentOutOfRangeException(nameof(options.MaxInputCharsPerWord), \"The maximum number of characters per word must be greater than zero.\");\n            }\n\n            if (!vocab!.TryGetValue(options.UnknownToken, out int id))\n            {\n                throw new ArgumentException($\"The unknown token '{options.UnknownToken}' is not in the vocabulary.\");\n            }\n\n            UnknownToken = options.UnknownToken;\n            UnknownTokenId = id;\n            ContinuingSubwordPrefix = options.ContinuingSubwordPrefix;\n            MaxInputCharsPerWord = options.MaxInputCharsPerWord;\n\n            _preTokenizer = options.PreTokenizer ?? PreTokenizer.CreateWhiteSpace(options.SpecialTokens);\n            _normalizer = options.Normalizer;\n        }\n\n        /// <summary>\n        /// Gets the unknown token ID.","sourceCodeStart":41,"sourceCodeEnd":77,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs#L41-L77","documentation":"The WordPieceTokenizer constructor (via WordPieceTokenizerOptions) requires MaxInputCharsPerWord to be a positive integer because wordpieces longer than this limit are replaced by the unknown token. A value of zero or negative is nonsensical and throws ArgumentOutOfRangeException naming options.MaxInputCharsPerWord.","triggerScenarios":"Constructing new WordPieceTokenizer(...) or new WordPieceTokenizerOptions with MaxInputCharsPerWord = 0 or a negative number, often left at default(int) when using object initializers on a hand-built options object.","commonSituations":"Forgetting to set MaxInputCharsPerWord on a manually instantiated options object (default 0); copying config from BERT code where the limit was configured as 0/disabled; off-by-one when intending a large limit.","solutions":["Set WordPieceTokenizerOptions.MaxInputCharsPerWord to a positive value (BERT uses 100; the library default is 200)","If constructing options dynamically, initialize the property explicitly rather than relying on default(int)","Clamp with Math.Max(1, configuredValue) at config load time"],"exampleFix":"// before\nvar options = new WordPieceTokenizerOptions { Vocab = vocab }; // MaxInputCharsPerWord defaults to 0\n// after\nvar options = new WordPieceTokenizerOptions { Vocab = vocab, MaxInputCharsPerWord = 200 };","handlingStrategy":"validation","validationCode":"if (options.MaxInputCharsPerWord <= 0) throw new ArgumentException(\"MaxInputCharsPerWord must be positive\", nameof(options));","typeGuard":"bool HasValidMaxChars(WordPieceTokenizerOptions o) => o.MaxInputCharsPerWord > 0;","tryCatchPattern":"try { var tok = new WordPieceTokenizer(options); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == \"options.MaxInputCharsPerWord\") { options.MaxInputCharsPerWord = 200; /* retry or fail fast */ }","preventionTips":["Always set MaxInputCharsPerWord explicitly (BERT: 100, library default: 200)","Never rely on default(int) for required positive options","Validate tokenizer options at config-load time"],"tags":["tokenizers","argument-out-of-range","wordpiece"],"backgroundTag":"argument-out-of-range","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}