dotnet/machinelearning · error · FormatException
Can't parse to integer
Error message
Can't parse {line.Substring(spaceIndex)} to integer What it means
After locating the single space in a vocab line, the substring after it must parse as a 32-bit integer rank. If not, a FormatException 'Can't parse <suffix> to integer' is thrown, identifying the offending text.
Solutions
- Fix the offending line so the text after the single space is a valid decimal integer rank.
- Re-download or regenerate the vocab file (the message names the exact bad substring to locate it).
- Ensure ranks are within Int32 range and written in plain decimal.
- Use the official data package files instead of hand-modified vocab streams.
Example fix
// before dGhl 12ab // after dGhl 12
Defensive patterns
Strategy: validation
Validate before calling
static void ValidateBpeRank(string line)
{
int sp = line.IndexOf(' ');
if (sp > 0 && !int.TryParse(line.AsSpan(sp + 1), out _))
throw new FormatException($"Rank is not an integer: '{line.Substring(sp)}'");
} Type guard
static bool HasValidRank(string line)
{
int sp = line.IndexOf(' ');
return sp > 0 && int.TryParse(line.AsSpan(sp + 1), out _);
} Try / catch
try { var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens); }
catch (InvalidOperationException ex) when (ex.InnerException is FormatException fe)
{ throw new InvalidDataException($"Vocab rank parse failure: {fe.Message}", fe); } Prevention
- Keep ranks as plain decimal integers within Int32 range.
- Re-verify downloaded vocab files against upstream checksums.
- Fail fast in CI by parsing the full vocab file once during build.
When it happens
Trigger: TiktokenTokenizer.CreateAsync/CreateForModel/TiktokenTokenizer ctor reading a vocab stream (LoadTiktokenBpeAsync) where a line like 'dGhl abc' or 'dGhl 999999999999' has a non-numeric or out-of-int32-range rank after the space.
Common situations: Corrupted or partially-written vocab file; ranks pasted as hex ('0x1a') or with signs/floats; line truncated mid-rank by a bad download.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- Invalid format in the BPE vocab file stream
- Failed to load from BPE vocab file stream
- A piece string in 'model.vocab' is null.
- A post_processor template 'SpecialToken.id' must be a…
- An 'added_tokens' entry must have a string 'content' and a…
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/4af853514ab3c86a.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs:222
}
while (line is not null)
{
if (line.Length > 0)
{
int spaceIndex = line.IndexOf(' ');
if (spaceIndex <= 0 || spaceIndex >= line.Length - 1 || line.IndexOf(' ', spaceIndex + 1) >= 0)
{
throw new FormatException($"Invalid format in the BPE vocab file stream");
}
if (Helpers.TryParseInt32(line, spaceIndex + 1, out int rank))
{
AddData(Helpers.FromBase64String(line, 0, spaceIndex), rank);
}
else
{
throw new FormatException($"Can't parse {line.Substring(spaceIndex)} to integer");
}
line = useAsync ?
await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) :
reader.ReadLine();
}
}
}
catch (Exception ex)
{
throw new InvalidOperationException($"Failed to load from BPE vocab file stream: {ex.Message}", ex);
}
return (encoder, vocab, decoder);
void AddData(byte[] tokenBytes, int rank)
{
encoder[tokenBytes] = rank;View on GitHub (pinned to 7b76e69cf9)