{"record":{"id":"80194008e80396a9","repo":"dotnet/machinelearning","slug":"invalid-format-in-the-bpe-vocab-file-stream","errorCode":null,"errorMessage":"Invalid format in the BPE vocab file stream","messagePattern":"Invalid format in the BPE vocab file stream","errorType":"exception","errorClass":"FormatException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs","lineNumber":176,"sourceCode":"            Stream vocabStream, bool useAsync, CancellationToken cancellationToken = default)\n        {\n            Dictionary<ReadOnlyMemory<byte>, int> encoder;\n            Dictionary<StringSpanOrdinalKey, (int Id, string Token)> vocab;\n            Dictionary<int, ReadOnlyMemory<byte>> decoder;\n\n            try\n            {\n                // Don't dispose the reader as it will dispose the underlying stream vocabStream. The caller is responsible for disposing the stream.\n                StreamReader reader = new StreamReader(vocabStream);\n                string? line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();\n\n                const string capacity = \"Capacity: \";\n                int suggestedCapacity = 0; // default capacity\n                if (line is not null && line.StartsWith(capacity, StringComparison.Ordinal))\n                {\n                    if (!Helpers.TryParseInt32(line, capacity.Length, out suggestedCapacity))\n                    {\n                        throw new FormatException($\"Invalid format in the BPE vocab file stream\");\n                    }\n\n                    line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();\n                }\n\n                encoder = new Dictionary<ReadOnlyMemory<byte>, int>(suggestedCapacity, ReadOnlyMemoryByteComparer.Instance);\n                vocab = new Dictionary<StringSpanOrdinalKey, (int Id, string Token)>(suggestedCapacity);\n                decoder = new Dictionary<int, ReadOnlyMemory<byte>>(suggestedCapacity);\n\n                // skip empty lines\n                while (line is not null && line.Length == 0)\n                {\n                    line = useAsync ? await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) : reader.ReadLine();\n                }\n\n                if (line is not null && line.IndexOf(' ') < 0)\n                {\n                    // We generate the ranking using the line number","sourceCodeStart":158,"sourceCodeEnd":194,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs#L158-L194","documentation":"LoadTiktokenBpeAsync reads an optional 'Capacity: N' header line from the BPE vocab stream; if the line starts with 'Capacity: ' but the remainder is not a valid 32-bit integer, a FormatException is thrown. It means the vocab file stream is malformed and cannot be used to build the tokenizer encoder.","triggerScenarios":"Calling TiktokenTokenizer.CreateAsync/CreateForModel (which route to LoadTiktokenBpeAsync) with a custom vocab stream whose first line begins with 'Capacity: ' but is followed by non-numeric text (e.g. 'Capacity: abc', 'Capacity: ' with nothing after, or an empty token).","commonSituations":"Hand-edited or truncated tiktoken .model files; a custom vocab file generated by a different tool version that writes a corrupt capacity header; passing an HTML error page or wrong file as the vocab stream.","solutions":["Fix or regenerate the vocab file so the 'Capacity: ' line is followed by a valid integer (Helpers.TryParseInt32 parses from offset 10).","Remove the 'Capacity: ' header line entirely — it is optional; the parser then uses a default capacity of 0.","Verify the stream contents point to a genuine tiktoken BPE vocab file, not an HTML/error page or another format.","If the stream comes from the Microsoft.ML.Tokenizers data package (e.g. Microsoft.ML.Tokenizers.Data.O200kBase), reference the correct package instead of a hand-built file."],"exampleFix":"// before (corrupt header)\nCapacity: 12x3\ndGhl 1\n// after\nCapacity: 100000\ndGhl 1","handlingStrategy":"validation","validationCode":"static void ValidateBpeStream(Stream s)\n{\n    using var sr = new StreamReader(s, leaveOpen: true);\n    var first = sr.ReadLine();\n    if (first != null && first.StartsWith(\"Capacity: \", StringComparison.Ordinal)\n        && !int.TryParse(first.AsSpan(\"Capacity: \".Length), out _))\n        throw new FormatException($\"Bad Capacity header: '{first}'\");\n    s.Seek(0, SeekOrigin.Begin);\n}","typeGuard":"static bool HasValidCapacityHeader(string? line) =>\n    line is null || !line.StartsWith(\"Capacity: \", StringComparison.Ordinal)\n    || int.TryParse(line.AsSpan(\"Capacity: \".Length), out _);","tryCatchPattern":"try { var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens); }\ncatch (InvalidOperationException ex) when (ex.InnerException is FormatException fe)\n{ logger.LogError(fe, \"Malformed BPE vocab stream\"); throw new InvalidDataException(\"Vocab stream is not valid tiktoken BPE format\", fe); }","preventionTips":["Never hand-edit vocab files; generate them from official tiktoken assets.","Remove or fix the optional 'Capacity: N' header if you build custom vocab files.","Sanity-check the stream begins with expected content before constructing a tokenizer."],"tags":["tokenizer","format-exception","bpe-vocab","file-parsing"],"backgroundTag":"invalid-argument-format","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}