{"record":{"id":"92c8a4fa94eec223","repo":"dotnet/machinelearning","slug":"the-maximum-number-of-tokens-must-be-greater-than-92c8a4","errorCode":null,"errorMessage":"The maximum number of tokens must be greater than zero.","messagePattern":"The maximum number of tokens must be greater than zero\\.","errorType":"exception","errorClass":"ArgumentOutOfRangeException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs","lineNumber":414,"sourceCode":"            ArrayPool<int>.Shared.Return(indexMapping);\n            return result;\n        }\n\n        /// <summary>\n        /// Encodes input text to token Ids.\n        /// </summary>\n        /// <param name=\"text\">The text to encode.</param>\n        /// <param name=\"textSpan\">The span of the text to encode which will be used if the <paramref name=\"text\"/> is <see langword=\"null\"/>.</param>\n        /// <param name=\"settings\">The settings used to encode the text.</param>\n        /// <returns>The encoded results containing the list of encoded Ids.</returns>\n        protected override EncodeResults<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, EncodeSettings settings)\n            => EncodeToIds(text, textSpan, settings.ConsiderPreTokenization, settings.ConsiderNormalization, settings.MaxTokenCount);\n\n        private EncodeResults<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, bool considerPreTokenization, bool considerNormalization, int maxTokenCount = int.MaxValue)\n        {\n            if (maxTokenCount <= 0)\n            {\n                throw new ArgumentOutOfRangeException(nameof(maxTokenCount), \"The maximum number of tokens must be greater than zero.\");\n            }\n\n            if (string.IsNullOrEmpty(text) && textSpan.IsEmpty)\n            {\n                return new EncodeResults<int> { Tokens = [], NormalizedText = null, CharsConsumed = 0 };\n            }\n\n            IEnumerable<(int Offset, int Length)>? splits = InitializeForEncoding(\n                                                                text,\n                                                                textSpan,\n                                                                considerPreTokenization,\n                                                                considerNormalization,\n                                                                _normalizer,\n                                                                _preTokenizer,\n                                                                out string? normalizedText,\n                                                                out ReadOnlySpan<char> textSpanToEncode,\n                                                                out _);\n","sourceCodeStart":396,"sourceCodeEnd":432,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs#L396-L432","documentation":"EncodeToIds validates that maxTokenCount, when specified, is greater than zero, throwing ArgumentOutOfRangeException otherwise. This limit caps how many token IDs are produced, so a non-positive value is meaningless.","triggerScenarios":"Calling any public EncodeToIds overload (including EncodeToIds(text, settings) via EncoderSettings.MaxTokenCount) with maxTokenCount <= 0, e.g. 0, -1, or a computed value that underflowed.","commonSituations":"Computing MaxTokenCount from config where a sentinel 0 means 'unlimited' by the developer's convention but the API expects int.MaxValue for unlimited; user-supplied limit from UI/query string not validated; integer arithmetic that produced a negative value.","solutions":["Pass int.MaxValue (the default) when no limit is intended instead of 0.","Validate that the configured maxTokenCount >= 1 before calling EncodeToIds.","If the value comes from settings, clamp it: max = Math.Max(1, configuredMax)."],"exampleFix":"// before\nvar ids = tokenizer.EncodeToIds(text, new EncoderSettings { MaxTokenCount = 0 });\n// after\nint max = Math.Max(1, configuredMax); // or int.MaxValue for unlimited\nvar ids = tokenizer.EncodeToIds(text, new EncoderSettings { MaxTokenCount = max });","handlingStrategy":"validation","validationCode":"int max = configuredMax;\nif (max <= 0) max = int.MaxValue; // API has no '0 = unlimited' semantics\nif (max <= 0) throw new InvalidOperationException(\"MaxTokenCount must be >= 1\");","typeGuard":null,"tryCatchPattern":"try { var ids = tokenizer.EncodeToIds(text, settings); }\ncatch (ArgumentOutOfRangeException ex) when (ex.ParamName == \"maxTokenCount\") { throw new InvalidOperationException(\"EncoderSettings.MaxTokenCount must be positive\", ex); }","preventionTips":["Never pass 0 as MaxTokenCount; use int.MaxValue for unlimited.","Validate user/config-supplied limits at the boundary (>= 1).","Clamp computed limits with Math.Max(1, value).","Add unit tests covering maxTokenCount boundary values."],"tags":["argument-validation","tokenizer","csharp"],"backgroundTag":"argument-out-of-range","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}