dotnet/machinelearning · error · InvalidOperationException
Failed to load from BPE vocab file stream
Error message
Failed to load from BPE vocab file stream: {ex.Message} What it means
LoadTiktokenBpeAsync wraps any exception thrown while parsing the vocab stream (including 660-662) in an InvalidOperationException 'Failed to load from BPE vocab file stream: <inner message>', preserving the inner exception. This is the error callers of CreateAsync/CreateForModel actually observe.
Solutions
- Read the InnerException to get the root cause (FormatException, IOException, etc.) and fix accordingly.
- Validate the vocab file independently: every line 'base64 rank', optional 'Capacity: N' header, integer ranks.
- Re-download the vocab file and verify its checksum/size against the official source.
- Catch InvalidOperationException around tokenizer creation and surface the inner message to diagnose.
Example fix
// before
var tok = TiktokenTokenizer.CreateForModel("gpt-4o");
// after
try { var tok = TiktokenTokenizer.CreateForModel("gpt-4o"); }
catch (InvalidOperationException ex) { Console.Error.WriteLine(ex.InnerException?.Message ?? ex.Message); throw; } Defensive patterns
Strategy: try-catch
Try / catch
try
{
var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens);
}
catch (InvalidOperationException ex)
{
// InnerException carries the real cause (FormatException, IOException, ...)
logger.LogError(ex.InnerException, "BPE vocab load failed: {Detail}", ex.Message);
throw;
} Prevention
- Always inspect InnerException when this wrapper fires — it names the exact bad line or header.
- Cache tokenizer instances after successful creation to avoid re-parsing.
- Validate vocab streams once at startup rather than per-request.
- Pin data-package versions and verify embedded resources exist before deployment.
When it happens
Trigger: Any malformed vocab stream (bad Capacity header, wrong line format, unparseable rank) or an unexpected exception (IOException, OutOfMemory, etc.) during LoadTiktokenBpeAsync triggered via TiktokenTokenizer ctor, CreateForModel, or CreateAsync.
Common situations: Wrong file passed as vocab stream; network download truncated the file; stream disposal/closure mid-read; corrupted embedded resource in a data package.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Can't parse to integer
- Invalid format in the BPE vocab file stream
- A piece string in 'model.vocab' is null.
- A post_processor template 'SpecialToken.id' must be a…
- An 'added_tokens' entry must have a string 'content' and a…
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/e62fd7354f94507b.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs:233
if (Helpers.TryParseInt32(line, spaceIndex + 1, out int rank))
{
AddData(Helpers.FromBase64String(line, 0, spaceIndex), rank);
}
else
{
throw new FormatException($"Can't parse {line.Substring(spaceIndex)} to integer");
}
line = useAsync ?
await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) :
reader.ReadLine();
}
}
}
catch (Exception ex)
{
throw new InvalidOperationException($"Failed to load from BPE vocab file stream: {ex.Message}", ex);
}
return (encoder, vocab, decoder);
void AddData(byte[] tokenBytes, int rank)
{
encoder[tokenBytes] = rank;
decoder[rank] = tokenBytes;
string decodedToken = Encoding.UTF8.GetString(tokenBytes);
if (decodedToken.IndexOf('\uFFFD') < 0)
{
vocab[new StringSpanOrdinalKey(decodedToken)] = (rank, decodedToken);
}
}
}
View on GitHub (pinned to 7b76e69cf9)