dotnet/machinelearning · error · ArgumentException

Incorrect vocabulary format, expected

Error message

Incorrect vocabulary format, expected "<token> <cnt>"

What it means

EnglishRobertaTokenizer vocabulary loading parses each vocabulary file line as '<token> <count>'. A line that does not split into exactly two space-separated fields is rejected with this ArgumentException, because the vocab file is malformed or not a Roberta vocab file.

Solutions

  1. Verify the vocab file is the English Roberta 'vocab.bpe'-style file where each line is '<token> <count>'
  2. Open the vocab file and fix or remove the malformed line so every line has exactly two space-separated fields
  3. Re-download the correct vocabulary asset from the original model repository instead of a converted one

Example fix

// before
stream = File.OpenRead("vocab.json"); // wrong format, one token per line
// after
stream = File.OpenRead("roberta-vocab.txt"); // lines like "token 123"
Defensive patterns

Strategy: validation

Validate before calling

bool IsValidRobertaVocabLine(string line) => line.Trim().Split(' ').Length == 2;

Try / catch

try { var tok = new EnglishRobertaTokenizer(vocabStream, ...); } catch (ArgumentException ex) { log.LogError(ex, "Vocabulary file is not in '<token> <cnt>' format"); throw; }

Prevention

When it happens

Trigger: Calling the EnglishRobertaTokenizer constructor/Create with a stream whose text contains a line with no space or more than one space-separated token (e.g. BPE vocab files with plain token lines, JSON vocab, or trailing formatting).

Common situations: Feeding a GPT-2 style vocab.json or merges.txt instead of the Roberta vocab format; hand-edited vocab files; downloading the wrong model file for the tokenizer.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/44153b540a38bba7. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs:1262

        /// Loads a pre-existing vocabulary from a text stream and adds its symbols to this instance.
        /// </summary>
        public void AddFromStream(Stream stream)
        {
            Debug.Assert(stream is not null);
            using StreamReader reader = new StreamReader(stream);

            while (reader.Peek() >= 0)
            {
                string? line = reader.ReadLine();
                if (line is null)
                {
                    continue;
                }

                var splitLine = line.Trim().Split(' ');
                if (splitLine.Length != 2)
                {
                    throw new ArgumentException("Incorrect vocabulary format, expected \"<token> <cnt>\"");
                }

                if (!int.TryParse(splitLine[1], out int occurrenceScore))
                {
                    throw new ArgumentException($"Cannot parse the line: '{line}'.");
                }

                if (!int.TryParse(splitLine[0], out var id))
                {
                    ReserveStringSymbolSlot(splitLine[0], occurrenceScore);
                }
                else
                {
                    AddSymbol(id, occurrenceScore);
                }
            }
        }
    }

View on GitHub (pinned to 7b76e69cf9)