{"record":{"id":"4af853514ab3c86a","repo":"dotnet/machinelearning","slug":"can-t-parse-line-substring-spaceindex-to-intege","errorCode":null,"errorMessage":"Can't parse {line.Substring(spaceIndex)} to integer","messagePattern":"Can't parse (.+?) to integer","errorType":"exception","errorClass":"FormatException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs","lineNumber":222,"sourceCode":"                }\n\n                while (line is not null)\n                {\n                    if (line.Length > 0)\n                    {\n                        int spaceIndex = line.IndexOf(' ');\n                        if (spaceIndex <= 0 || spaceIndex >= line.Length - 1 || line.IndexOf(' ', spaceIndex + 1) >= 0)\n                        {\n                            throw new FormatException($\"Invalid format in the BPE vocab file stream\");\n                        }\n\n                        if (Helpers.TryParseInt32(line, spaceIndex + 1, out int rank))\n                        {\n                            AddData(Helpers.FromBase64String(line, 0, spaceIndex), rank);\n                        }\n                        else\n                        {\n                            throw new FormatException($\"Can't parse {line.Substring(spaceIndex)} to integer\");\n                        }\n\n                        line = useAsync ?\n                            await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) :\n                            reader.ReadLine();\n                    }\n                }\n            }\n            catch (Exception ex)\n            {\n                throw new InvalidOperationException($\"Failed to load from BPE vocab file stream: {ex.Message}\", ex);\n            }\n\n            return (encoder, vocab, decoder);\n\n            void AddData(byte[] tokenBytes, int rank)\n            {\n                encoder[tokenBytes] = rank;","sourceCodeStart":204,"sourceCodeEnd":240,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs#L204-L240","documentation":"After locating the single space in a vocab line, the substring after it must parse as a 32-bit integer rank. If not, a FormatException 'Can't parse <suffix> to integer' is thrown, identifying the offending text.","triggerScenarios":"TiktokenTokenizer.CreateAsync/CreateForModel/TiktokenTokenizer ctor reading a vocab stream (LoadTiktokenBpeAsync) where a line like 'dGhl abc' or 'dGhl 999999999999' has a non-numeric or out-of-int32-range rank after the space.","commonSituations":"Corrupted or partially-written vocab file; ranks pasted as hex ('0x1a') or with signs/floats; line truncated mid-rank by a bad download.","solutions":["Fix the offending line so the text after the single space is a valid decimal integer rank.","Re-download or regenerate the vocab file (the message names the exact bad substring to locate it).","Ensure ranks are within Int32 range and written in plain decimal.","Use the official data package files instead of hand-modified vocab streams."],"exampleFix":"// before\ndGhl 12ab\n// after\ndGhl 12","handlingStrategy":"validation","validationCode":"static void ValidateBpeRank(string line)\n{\n    int sp = line.IndexOf(' ');\n    if (sp > 0 && !int.TryParse(line.AsSpan(sp + 1), out _))\n        throw new FormatException($\"Rank is not an integer: '{line.Substring(sp)}'\");\n}","typeGuard":"static bool HasValidRank(string line)\n{\n    int sp = line.IndexOf(' ');\n    return sp > 0 && int.TryParse(line.AsSpan(sp + 1), out _);\n}","tryCatchPattern":"try { var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens); }\ncatch (InvalidOperationException ex) when (ex.InnerException is FormatException fe)\n{ throw new InvalidDataException($\"Vocab rank parse failure: {fe.Message}\", fe); }","preventionTips":["Keep ranks as plain decimal integers within Int32 range.","Re-verify downloaded vocab files against upstream checksums.","Fail fast in CI by parsing the full vocab file once during build."],"tags":["tokenizer","format-exception","bpe-vocab","integer-parsing"],"backgroundTag":"invalid-argument-format","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}