{"record":{"id":"e62fd7354f94507b","repo":"dotnet/machinelearning","slug":"failed-to-load-from-bpe-vocab-file-stream-ex-mes","errorCode":null,"errorMessage":"Failed to load from BPE vocab file stream: {ex.Message}","messagePattern":"Failed to load from BPE vocab file stream: (.+?)","errorType":"exception","errorClass":"InvalidOperationException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs","lineNumber":233,"sourceCode":"\n                        if (Helpers.TryParseInt32(line, spaceIndex + 1, out int rank))\n                        {\n                            AddData(Helpers.FromBase64String(line, 0, spaceIndex), rank);\n                        }\n                        else\n                        {\n                            throw new FormatException($\"Can't parse {line.Substring(spaceIndex)} to integer\");\n                        }\n\n                        line = useAsync ?\n                            await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) :\n                            reader.ReadLine();\n                    }\n                }\n            }\n            catch (Exception ex)\n            {\n                throw new InvalidOperationException($\"Failed to load from BPE vocab file stream: {ex.Message}\", ex);\n            }\n\n            return (encoder, vocab, decoder);\n\n            void AddData(byte[] tokenBytes, int rank)\n            {\n                encoder[tokenBytes] = rank;\n                decoder[rank] = tokenBytes;\n\n                string decodedToken = Encoding.UTF8.GetString(tokenBytes);\n\n                if (decodedToken.IndexOf('\\uFFFD') < 0)\n                {\n                    vocab[new StringSpanOrdinalKey(decodedToken)] = (rank, decodedToken);\n                }\n            }\n        }\n","sourceCodeStart":215,"sourceCodeEnd":251,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs#L215-L251","documentation":"LoadTiktokenBpeAsync wraps any exception thrown while parsing the vocab stream (including 660-662) in an InvalidOperationException 'Failed to load from BPE vocab file stream: <inner message>', preserving the inner exception. This is the error callers of CreateAsync/CreateForModel actually observe.","triggerScenarios":"Any malformed vocab stream (bad Capacity header, wrong line format, unparseable rank) or an unexpected exception (IOException, OutOfMemory, etc.) during LoadTiktokenBpeAsync triggered via TiktokenTokenizer ctor, CreateForModel, or CreateAsync.","commonSituations":"Wrong file passed as vocab stream; network download truncated the file; stream disposal/closure mid-read; corrupted embedded resource in a data package.","solutions":["Read the InnerException to get the root cause (FormatException, IOException, etc.) and fix accordingly.","Validate the vocab file independently: every line 'base64 rank', optional 'Capacity: N' header, integer ranks.","Re-download the vocab file and verify its checksum/size against the official source.","Catch InvalidOperationException around tokenizer creation and surface the inner message to diagnose."],"exampleFix":"// before\nvar tok = TiktokenTokenizer.CreateForModel(\"gpt-4o\");\n// after\ntry { var tok = TiktokenTokenizer.CreateForModel(\"gpt-4o\"); }\ncatch (InvalidOperationException ex) { Console.Error.WriteLine(ex.InnerException?.Message ?? ex.Message); throw; }","handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"try\n{\n    var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens);\n}\ncatch (InvalidOperationException ex)\n{\n    // InnerException carries the real cause (FormatException, IOException, ...)\n    logger.LogError(ex.InnerException, \"BPE vocab load failed: {Detail}\", ex.Message);\n    throw;\n}","preventionTips":["Always inspect InnerException when this wrapper fires — it names the exact bad line or header.","Cache tokenizer instances after successful creation to avoid re-parsing.","Validate vocab streams once at startup rather than per-request.","Pin data-package versions and verify embedded resources exist before deployment."],"tags":["tokenizer","invalid-operation-exception","bpe-vocab","wrapped-exception"],"backgroundTag":"file-read-failed","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}