{"record":{"id":"d4f2480d2cab7d82","repo":"dotnet/machinelearning","slug":"cannot-parse-the-line-line","errorCode":null,"errorMessage":"Cannot parse the line: '{line}'.","messagePattern":"Cannot parse the line: '(.+?)'\\.","errorType":"exception","errorClass":"ArgumentException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs","lineNumber":1267,"sourceCode":"            using StreamReader reader = new StreamReader(stream);\n\n            while (reader.Peek() >= 0)\n            {\n                string? line = reader.ReadLine();\n                if (line is null)\n                {\n                    continue;\n                }\n\n                var splitLine = line.Trim().Split(' ');\n                if (splitLine.Length != 2)\n                {\n                    throw new ArgumentException(\"Incorrect vocabulary format, expected \\\"<token> <cnt>\\\"\");\n                }\n\n                if (!int.TryParse(splitLine[1], out int occurrenceScore))\n                {\n                    throw new ArgumentException($\"Cannot parse the line: '{line}'.\");\n                }\n\n                if (!int.TryParse(splitLine[0], out var id))\n                {\n                    ReserveStringSymbolSlot(splitLine[0], occurrenceScore);\n                }\n                else\n                {\n                    AddSymbol(id, occurrenceScore);\n                }\n            }\n        }\n    }\n}\n","sourceCodeStart":1249,"sourceCodeEnd":1282,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs#L1249-L1282","documentation":"After confirming the line has two fields, the loader parses the second field as the occurrence count with int.TryParse. If the second field is not an integer, this ArgumentException is thrown naming the offending line.","triggerScenarios":"Loading an EnglishRobertaTokenizer vocabulary where a line's second token is non-numeric (e.g. 'token abc', float values, or a merges file where the second column is another token).","commonSituations":"Using the BPE merges file (two tokens per line) as if it were the counts vocab; localized number formats; corrupted downloads.","solutions":["Ensure the second field of every vocab line is a plain integer occurrence count","Check you are not passing the merges.txt file where the counts vocab is expected","Re-download or regenerate the vocabulary file from the model source"],"exampleFix":"// before\n// vocab line: \"Ġthe 1.5\" -> not an int\n// after\n// vocab line: \"Ġthe 124\" -> parses as occurrenceScore 124","handlingStrategy":"validation","validationCode":"bool IsValidVocabLine(string line) { var p = line.Trim().Split(' '); return p.Length == 2 && int.TryParse(p[1], out _); }","typeGuard":null,"tryCatchPattern":"try { var tok = new EnglishRobertaTokenizer(vocabStream, ...); } catch (ArgumentException ex) { log.LogError(ex, \"Malformed vocab line: non-integer count\"); throw; }","preventionTips":["Confirm the second column of every vocab line parses as int","Make sure you are not passing merges.txt as the counts vocab","Verify downloaded files with a hash"],"tags":["tokenizers","argument-exception","parse-failure"],"backgroundTag":"invalid-argument-format","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}