dotnet/machinelearning · error · ArgumentException
Incorrect vocabulary format, expected
Error message
Incorrect vocabulary format, expected "<token> <cnt>"
What it means
EnglishRobertaTokenizer vocabulary loading parses each vocabulary file line as '<token> <count>'. A line that does not split into exactly two space-separated fields is rejected with this ArgumentException, because the vocab file is malformed or not a Roberta vocab file.
Solutions
- Verify the vocab file is the English Roberta 'vocab.bpe'-style file where each line is '<token> <count>'
- Open the vocab file and fix or remove the malformed line so every line has exactly two space-separated fields
- Re-download the correct vocabulary asset from the original model repository instead of a converted one
Example fix
// before
stream = File.OpenRead("vocab.json"); // wrong format, one token per line
// after
stream = File.OpenRead("roberta-vocab.txt"); // lines like "token 123" Defensive patterns
Strategy: validation
Validate before calling
bool IsValidRobertaVocabLine(string line) => line.Trim().Split(' ').Length == 2; Try / catch
try { var tok = new EnglishRobertaTokenizer(vocabStream, ...); } catch (ArgumentException ex) { log.LogError(ex, "Vocabulary file is not in '<token> <cnt>' format"); throw; } Prevention
- Validate a few sample lines of the vocab file before loading
- Keep model assets checksummed and downloaded from a known source
- Do not hand-edit vocab files
When it happens
Trigger: Calling the EnglishRobertaTokenizer constructor/Create with a stream whose text contains a line with no space or more than one space-separated token (e.g. BPE vocab files with plain token lines, JSON vocab, or trailing formatting).
Common situations: Feeding a GPT-2 style vocab.json or merges.txt instead of the Roberta vocab format; hand-edited vocab files; downloading the wrong model file for the tokenizer.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- Cannot parse the line
- and must be different
- A tokenizer.json normalizer entry must be a JSON object.
- ArgumentNullException
- ArgumentNullException
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/44153b540a38bba7.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs:1262
/// Loads a pre-existing vocabulary from a text stream and adds its symbols to this instance.
/// </summary>
public void AddFromStream(Stream stream)
{
Debug.Assert(stream is not null);
using StreamReader reader = new StreamReader(stream);
while (reader.Peek() >= 0)
{
string? line = reader.ReadLine();
if (line is null)
{
continue;
}
var splitLine = line.Trim().Split(' ');
if (splitLine.Length != 2)
{
throw new ArgumentException("Incorrect vocabulary format, expected \"<token> <cnt>\"");
}
if (!int.TryParse(splitLine[1], out int occurrenceScore))
{
throw new ArgumentException($"Cannot parse the line: '{line}'.");
}
if (!int.TryParse(splitLine[0], out var id))
{
ReserveStringSymbolSlot(splitLine[0], occurrenceScore);
}
else
{
AddSymbol(id, occurrenceScore);
}
}
}
}View on GitHub (pinned to 7b76e69cf9)