dotnet/machinelearning · error · FormatException
Invalid format in the BPE vocab file stream
Error message
Invalid format in the BPE vocab file stream
What it means
LoadTiktokenBpeAsync reads an optional 'Capacity: N' header line from the BPE vocab stream; if the line starts with 'Capacity: ' but the remainder is not a valid 32-bit integer, a FormatException is thrown. It means the vocab file stream is malformed and cannot be used to build the tokenizer encoder.
Solutions
- Fix or regenerate the vocab file so the 'Capacity: ' line is followed by a valid integer (Helpers.TryParseInt32 parses from offset 10).
- Remove the 'Capacity: ' header line entirely — it is optional; the parser then uses a default capacity of 0.
- Verify the stream contents point to a genuine tiktoken BPE vocab file, not an HTML/error page or another format.
- If the stream comes from the Microsoft.ML.Tokenizers data package (e.g. Microsoft.ML.Tokenizers.Data.O200kBase), reference the correct package instead of a hand-built file.
Example fix
// before (corrupt header) Capacity: 12x3 dGhl 1 // after Capacity: 100000 dGhl 1
Defensive patterns
Strategy: validation
Validate before calling
static void ValidateBpeStream(Stream s)
{
using var sr = new StreamReader(s, leaveOpen: true);
var first = sr.ReadLine();
if (first != null && first.StartsWith("Capacity: ", StringComparison.Ordinal)
&& !int.TryParse(first.AsSpan("Capacity: ".Length), out _))
throw new FormatException($"Bad Capacity header: '{first}'");
s.Seek(0, SeekOrigin.Begin);
} Type guard
static bool HasValidCapacityHeader(string? line) =>
line is null || !line.StartsWith("Capacity: ", StringComparison.Ordinal)
|| int.TryParse(line.AsSpan("Capacity: ".Length), out _); Try / catch
try { var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens); }
catch (InvalidOperationException ex) when (ex.InnerException is FormatException fe)
{ logger.LogError(fe, "Malformed BPE vocab stream"); throw new InvalidDataException("Vocab stream is not valid tiktoken BPE format", fe); } Prevention
- Never hand-edit vocab files; generate them from official tiktoken assets.
- Remove or fix the optional 'Capacity: N' header if you build custom vocab files.
- Sanity-check the stream begins with expected content before constructing a tokenizer.
When it happens
Trigger: Calling TiktokenTokenizer.CreateAsync/CreateForModel (which route to LoadTiktokenBpeAsync) with a custom vocab stream whose first line begins with 'Capacity: ' but is followed by non-numeric text (e.g. 'Capacity: abc', 'Capacity: ' with nothing after, or an empty token).
Common situations: Hand-edited or truncated tiktoken .model files; a custom vocab file generated by a different tool version that writes a corrupt capacity header; passing an HTML error page or wrong file as the vocab stream.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- Can't parse to integer
- Failed to load from BPE vocab file stream
- A piece string in 'model.vocab' is null.
- A post_processor template 'SpecialToken.id' must be a…
- An 'added_tokens' entry must have a string 'content' and a…
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/80194008e80396a9.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs:176
Stream vocabStream, bool useAsync, CancellationToken cancellationToken = default)
{
Dictionary<ReadOnlyMemory<byte>, int> encoder;
Dictionary<StringSpanOrdinalKey, (int Id, string Token)> vocab;
Dictionary<int, ReadOnlyMemory<byte>> decoder;
try
{
// Don't dispose the reader as it will dispose the underlying stream vocabStream. The caller is responsible for disposing the stream.
StreamReader reader = new StreamReader(vocabStream);
string? line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();
const string capacity = "Capacity: ";
int suggestedCapacity = 0; // default capacity
if (line is not null && line.StartsWith(capacity, StringComparison.Ordinal))
{
if (!Helpers.TryParseInt32(line, capacity.Length, out suggestedCapacity))
{
throw new FormatException($"Invalid format in the BPE vocab file stream");
}
line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();
}
encoder = new Dictionary<ReadOnlyMemory<byte>, int>(suggestedCapacity, ReadOnlyMemoryByteComparer.Instance);
vocab = new Dictionary<StringSpanOrdinalKey, (int Id, string Token)>(suggestedCapacity);
decoder = new Dictionary<int, ReadOnlyMemory<byte>>(suggestedCapacity);
// skip empty lines
while (line is not null && line.Length == 0)
{
line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();
}
if (line is not null && line.IndexOf(' ') < 0)
{
// We generate the ranking using the line numberView on GitHub (pinned to 7b76e69cf9)