dotnet/machinelearning · error · ArgumentException
Cannot parse the line
Error message
Cannot parse the line: '{line}'. What it means
After confirming the line has two fields, the loader parses the second field as the occurrence count with int.TryParse. If the second field is not an integer, this ArgumentException is thrown naming the offending line.
Solutions
- Ensure the second field of every vocab line is a plain integer occurrence count
- Check you are not passing the merges.txt file where the counts vocab is expected
- Re-download or regenerate the vocabulary file from the model source
Example fix
// before // vocab line: "Ġthe 1.5" -> not an int // after // vocab line: "Ġthe 124" -> parses as occurrenceScore 124
Defensive patterns
Strategy: validation
Validate before calling
bool IsValidVocabLine(string line) { var p = line.Trim().Split(' '); return p.Length == 2 && int.TryParse(p[1], out _); } Try / catch
try { var tok = new EnglishRobertaTokenizer(vocabStream, ...); } catch (ArgumentException ex) { log.LogError(ex, "Malformed vocab line: non-integer count"); throw; } Prevention
- Confirm the second column of every vocab line parses as int
- Make sure you are not passing merges.txt as the counts vocab
- Verify downloaded files with a hash
When it happens
Trigger: Loading an EnglishRobertaTokenizer vocabulary where a line's second token is non-numeric (e.g. 'token abc', float values, or a merges file where the second column is another token).
Common situations: Using the BPE merges file (two tokens per line) as if it were the counts vocab; localized number formats; corrupted downloads.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- Incorrect vocabulary format, expected
- and must be different
- A tokenizer.json normalizer entry must be a JSON object.
- ArgumentNullException
- ArgumentNullException
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/d4f2480d2cab7d82.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs:1267
using StreamReader reader = new StreamReader(stream);
while (reader.Peek() >= 0)
{
string? line = reader.ReadLine();
if (line is null)
{
continue;
}
var splitLine = line.Trim().Split(' ');
if (splitLine.Length != 2)
{
throw new ArgumentException("Incorrect vocabulary format, expected \"<token> <cnt>\"");
}
if (!int.TryParse(splitLine[1], out int occurrenceScore))
{
throw new ArgumentException($"Cannot parse the line: '{line}'.");
}
if (!int.TryParse(splitLine[0], out var id))
{
ReserveStringSymbolSlot(splitLine[0], occurrenceScore);
}
else
{
AddSymbol(id, occurrenceScore);
}
}
}
}
}
View on GitHub (pinned to 7b76e69cf9)