dotnet/machinelearning · error · FormatException

Invalid format in the BPE vocab file stream

Error message

Invalid format in the BPE vocab file stream

What it means

LoadTiktokenBpeAsync reads an optional 'Capacity: N' header line from the BPE vocab stream; if the line starts with 'Capacity: ' but the remainder is not a valid 32-bit integer, a FormatException is thrown. It means the vocab file stream is malformed and cannot be used to build the tokenizer encoder.

Solutions

  1. Fix or regenerate the vocab file so the 'Capacity: ' line is followed by a valid integer (Helpers.TryParseInt32 parses from offset 10).
  2. Remove the 'Capacity: ' header line entirely — it is optional; the parser then uses a default capacity of 0.
  3. Verify the stream contents point to a genuine tiktoken BPE vocab file, not an HTML/error page or another format.
  4. If the stream comes from the Microsoft.ML.Tokenizers data package (e.g. Microsoft.ML.Tokenizers.Data.O200kBase), reference the correct package instead of a hand-built file.

Example fix

// before (corrupt header)
Capacity: 12x3
dGhl 1
// after
Capacity: 100000
dGhl 1
Defensive patterns

Strategy: validation

Validate before calling

static void ValidateBpeStream(Stream s)
{
    using var sr = new StreamReader(s, leaveOpen: true);
    var first = sr.ReadLine();
    if (first != null && first.StartsWith("Capacity: ", StringComparison.Ordinal)
        && !int.TryParse(first.AsSpan("Capacity: ".Length), out _))
        throw new FormatException($"Bad Capacity header: '{first}'");
    s.Seek(0, SeekOrigin.Begin);
}

Type guard

static bool HasValidCapacityHeader(string? line) =>
    line is null || !line.StartsWith("Capacity: ", StringComparison.Ordinal)
    || int.TryParse(line.AsSpan("Capacity: ".Length), out _);

Try / catch

try { var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens); }
catch (InvalidOperationException ex) when (ex.InnerException is FormatException fe)
{ logger.LogError(fe, "Malformed BPE vocab stream"); throw new InvalidDataException("Vocab stream is not valid tiktoken BPE format", fe); }

Prevention

When it happens

Trigger: Calling TiktokenTokenizer.CreateAsync/CreateForModel (which route to LoadTiktokenBpeAsync) with a custom vocab stream whose first line begins with 'Capacity: ' but is followed by non-numeric text (e.g. 'Capacity: abc', 'Capacity: ' with nothing after, or an empty token).

Common situations: Hand-edited or truncated tiktoken .model files; a custom vocab file generated by a different tool version that writes a corrupt capacity header; passing an HTML error page or wrong file as the vocab stream.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/80194008e80396a9. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs:176

            Stream vocabStream, bool useAsync, CancellationToken cancellationToken = default)
        {
            Dictionary<ReadOnlyMemory<byte>, int> encoder;
            Dictionary<StringSpanOrdinalKey, (int Id, string Token)> vocab;
            Dictionary<int, ReadOnlyMemory<byte>> decoder;

            try
            {
                // Don't dispose the reader as it will dispose the underlying stream vocabStream. The caller is responsible for disposing the stream.
                StreamReader reader = new StreamReader(vocabStream);
                string? line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();

                const string capacity = "Capacity: ";
                int suggestedCapacity = 0; // default capacity
                if (line is not null && line.StartsWith(capacity, StringComparison.Ordinal))
                {
                    if (!Helpers.TryParseInt32(line, capacity.Length, out suggestedCapacity))
                    {
                        throw new FormatException($"Invalid format in the BPE vocab file stream");
                    }

                    line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();
                }

                encoder = new Dictionary<ReadOnlyMemory<byte>, int>(suggestedCapacity, ReadOnlyMemoryByteComparer.Instance);
                vocab = new Dictionary<StringSpanOrdinalKey, (int Id, string Token)>(suggestedCapacity);
                decoder = new Dictionary<int, ReadOnlyMemory<byte>>(suggestedCapacity);

                // skip empty lines
                while (line is not null && line.Length == 0)
                {
                    line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();
                }

                if (line is not null && line.IndexOf(' ') < 0)
                {
                    // We generate the ranking using the line number

View on GitHub (pinned to 7b76e69cf9)