dotnet/machinelearning · error · ArgumentException

Cannot parse the line

Error message

Cannot parse the line: '{line}'.

What it means

After confirming the line has two fields, the loader parses the second field as the occurrence count with int.TryParse. If the second field is not an integer, this ArgumentException is thrown naming the offending line.

Solutions

  1. Ensure the second field of every vocab line is a plain integer occurrence count
  2. Check you are not passing the merges.txt file where the counts vocab is expected
  3. Re-download or regenerate the vocabulary file from the model source

Example fix

// before
// vocab line: "Ġthe 1.5" -> not an int
// after
// vocab line: "Ġthe 124" -> parses as occurrenceScore 124
Defensive patterns

Strategy: validation

Validate before calling

bool IsValidVocabLine(string line) { var p = line.Trim().Split(' '); return p.Length == 2 && int.TryParse(p[1], out _); }

Try / catch

try { var tok = new EnglishRobertaTokenizer(vocabStream, ...); } catch (ArgumentException ex) { log.LogError(ex, "Malformed vocab line: non-integer count"); throw; }

Prevention

When it happens

Trigger: Loading an EnglishRobertaTokenizer vocabulary where a line's second token is non-numeric (e.g. 'token abc', float values, or a merges file where the second column is another token).

Common situations: Using the BPE merges file (two tokens per line) as if it were the counts vocab; localized number formats; corrupted downloads.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/d4f2480d2cab7d82. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/EnglishRobertaTokenizer.cs:1267

            using StreamReader reader = new StreamReader(stream);

            while (reader.Peek() >= 0)
            {
                string? line = reader.ReadLine();
                if (line is null)
                {
                    continue;
                }

                var splitLine = line.Trim().Split(' ');
                if (splitLine.Length != 2)
                {
                    throw new ArgumentException("Incorrect vocabulary format, expected \"<token> <cnt>\"");
                }

                if (!int.TryParse(splitLine[1], out int occurrenceScore))
                {
                    throw new ArgumentException($"Cannot parse the line: '{line}'.");
                }

                if (!int.TryParse(splitLine[0], out var id))
                {
                    ReserveStringSymbolSlot(splitLine[0], occurrenceScore);
                }
                else
                {
                    AddSymbol(id, occurrenceScore);
                }
            }
        }
    }
}

View on GitHub (pinned to 7b76e69cf9)