{"record":{"id":"44153b540a38bba7","repo":"dotnet/machinelearning","slug":"incorrect-vocabulary-format-expected-token-cn","errorCode":null,"errorMessage":"Incorrect vocabulary format, expected \"<token> <cnt>\"","messagePattern":"Incorrect vocabulary format, expected \"<token> <cnt>\"","errorType":"exception","errorClass":"ArgumentException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs","lineNumber":1262,"sourceCode":"        /// Loads a pre-existing vocabulary from a text stream and adds its symbols to this instance.\n        /// </summary>\n        public void AddFromStream(Stream stream)\n        {\n            Debug.Assert(stream is not null);\n            using StreamReader reader = new StreamReader(stream);\n\n            while (reader.Peek() >= 0)\n            {\n                string? line = reader.ReadLine();\n                if (line is null)\n                {\n                    continue;\n                }\n\n                var splitLine = line.Trim().Split(' ');\n                if (splitLine.Length != 2)\n                {\n                    throw new ArgumentException(\"Incorrect vocabulary format, expected \\\"<token> <cnt>\\\"\");\n                }\n\n                if (!int.TryParse(splitLine[1], out int occurrenceScore))\n                {\n                    throw new ArgumentException($\"Cannot parse the line: '{line}'.\");\n                }\n\n                if (!int.TryParse(splitLine[0], out var id))\n                {\n                    ReserveStringSymbolSlot(splitLine[0], occurrenceScore);\n                }\n                else\n                {\n                    AddSymbol(id, occurrenceScore);\n                }\n            }\n        }\n    }","sourceCodeStart":1244,"sourceCodeEnd":1280,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs#L1244-L1280","documentation":"EnglishRobertaTokenizer vocabulary loading parses each vocabulary file line as '<token> <count>'. A line that does not split into exactly two space-separated fields is rejected with this ArgumentException, because the vocab file is malformed or not a Roberta vocab file.","triggerScenarios":"Calling the EnglishRobertaTokenizer constructor/Create with a stream whose text contains a line with no space or more than one space-separated token (e.g. BPE vocab files with plain token lines, JSON vocab, or trailing formatting).","commonSituations":"Feeding a GPT-2 style vocab.json or merges.txt instead of the Roberta vocab format; hand-edited vocab files; downloading the wrong model file for the tokenizer.","solutions":["Verify the vocab file is the English Roberta 'vocab.bpe'-style file where each line is '<token> <count>'","Open the vocab file and fix or remove the malformed line so every line has exactly two space-separated fields","Re-download the correct vocabulary asset from the original model repository instead of a converted one"],"exampleFix":"// before\nstream = File.OpenRead(\"vocab.json\"); // wrong format, one token per line\n// after\nstream = File.OpenRead(\"roberta-vocab.txt\"); // lines like \"token 123\"","handlingStrategy":"validation","validationCode":"bool IsValidRobertaVocabLine(string line) => line.Trim().Split(' ').Length == 2;","typeGuard":null,"tryCatchPattern":"try { var tok = new EnglishRobertaTokenizer(vocabStream, ...); } catch (ArgumentException ex) { log.LogError(ex, \"Vocabulary file is not in '<token> <cnt>' format\"); throw; }","preventionTips":["Validate a few sample lines of the vocab file before loading","Keep model assets checksummed and downloaded from a known source","Do not hand-edit vocab files"],"tags":["tokenizers","argument-exception","vocab-format"],"backgroundTag":"invalid-argument-format","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}