{"record":{"id":"0e052221335a41e7","repo":"dotnet/machinelearning","slug":"the-max-token-count-must-be-greater-than-0-0e0522","errorCode":null,"errorMessage":"The max token count must be greater than 0.","messagePattern":"The max token count must be greater than 0\\.","errorType":"validation","errorClass":"ArgumentOutOfRangeException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs","lineNumber":606,"sourceCode":"        /// </summary>\n        /// <param name=\"text\">The text to encode.</param>\n        /// <param name=\"textSpan\">The span of the text to encode which will be used if the <paramref name=\"text\"/> is <see langword=\"null\"/>.</param>\n        /// <param name=\"settings\">The settings used to encode the text.</param>\n        /// <param name=\"fromEnd\">Indicate whether to find the index from the end of the text.</param>\n        /// <param name=\"normalizedText\">If the tokenizer's normalization is enabled or <paramRef name=\"settings\" /> has <see cref=\"EncodeSettings.ConsiderNormalization\"/> is <see langword=\"false\"/>, this will be set to <paramRef name=\"text\" /> in its normalized form; otherwise, this value will be set to <see langword=\"null\"/>.</param>\n        /// <param name=\"tokenCount\">The token count can be generated which should be smaller than the maximum token count.</param>\n        /// <returns>\n        /// The index of the maximum encoding capacity within the processed text without surpassing the token limit.\n        /// If <paramRef name=\"fromEnd\" /> is <see langword=\"false\"/>, it represents the index immediately following the last character to be included. In cases where no tokens fit, the result will be 0; conversely,\n        /// if all tokens fit, the result will be length of the input text or the <paramref name=\"normalizedText\"/> if the normalization is enabled.\n        /// If <paramRef name=\"fromEnd\" /> is <see langword=\"true\"/>, it represents the index of the first character to be included. In cases where no tokens fit, the result will be the text length; conversely,\n        /// if all tokens fit, the result will be zero.\n        /// </returns>\n        protected override int GetIndexByTokenCount(string? text, ReadOnlySpan<char> textSpan, EncodeSettings settings, bool fromEnd, out string? normalizedText, out int tokenCount)\n        {\n            if (settings.MaxTokenCount <= 0)\n            {\n                throw new ArgumentOutOfRangeException(nameof(settings.MaxTokenCount), \"The max token count must be greater than 0.\");\n            }\n\n            if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)\n            {\n                normalizedText = null;\n                tokenCount = 0;\n                return 0;\n            }\n\n            IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(\n                                                                text,\n                                                                textSpan,\n                                                                settings.ConsiderNormalization,\n                                                                settings.ConsiderNormalization,\n                                                                _normalizer,\n                                                                _preTokenizer,\n                                                                out normalizedText,\n                                                                out ReadOnlySpan<char> textSpanToEncode,","sourceCodeStart":588,"sourceCodeEnd":624,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/WordPieceTokenizer.cs#L588-L624","documentation":"WordPieceTokenizer.GetIndexByTokenCount (and the index-of-token-count overloads) validate settings.MaxTokenCount > 0 and throw ArgumentOutOfRangeException with this message otherwise. It computes the character index at which encoding should stop, requiring a valid positive token budget.","triggerScenarios":"Calling tokenizer.GetIndexByTokenCount / GetIndexFromEndByTokenCount with a settings object whose MaxTokenCount is 0 or negative.","commonSituations":"Truncating text to fit a max length where the limit came back as 0 from an empty config field; copying settings objects without copying MaxTokenCount; passing 0 to mean 'return immediately'.","solutions":["Pass a positive MaxTokenCount in the EncodeSettings","Fix the upstream limit computation so it yields at least 1","Early-return in your own code when the configured limit is 0 instead of calling the tokenizer"],"exampleFix":"// before\nvar idx = tokenizer.GetIndexByTokenCount(text, new EncodeSettings { MaxTokenCount = 0 }, out _, out _);\n// after\nvar idx = tokenizer.GetIndexByTokenCount(text, new EncodeSettings { MaxTokenCount = 512 }, out _, out _);","handlingStrategy":"validation","validationCode":"if (settings.MaxTokenCount <= 0) throw new ArgumentException(\"MaxTokenCount must be positive before calling GetIndexByTokenCount\");","typeGuard":"bool IsValidMaxTokenCount(int n) => n > 0;","tryCatchPattern":"try { var idx = tokenizer.GetIndexByTokenCount(text, settings, fromEnd, out _, out _); } catch (ArgumentOutOfRangeException ex) when (ex.ParamName == \"settings.MaxTokenCount\") { /* clamp limit and retry */ }","preventionTips":["Clamp computed truncation limits to Math.Max(1, limit)","Skip the call entirely when the limit is 0 (return index 0)","Share one validated EncodeSettings instance across encode/count/index calls"],"tags":["tokenizers","argument-out-of-range","truncation"],"backgroundTag":"argument-out-of-range","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}