{"record":{"id":"27c4792ba898de2c","repo":"dotnet/machinelearning","slug":"a-tokenizer-json-normalizer-entry-must-be-a-json-o","errorCode":null,"errorMessage":"A tokenizer.json normalizer entry must be a JSON object.","messagePattern":"A tokenizer\\.json normalizer entry must be a JSON object\\.","errorType":"validation","errorClass":"InvalidDataException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs","lineNumber":131,"sourceCode":"                    // Replace (literal punctuation splitting, character substitution) is content-modifying.\n                    return !ReplaceIsWhitespaceCollapse(normalizer);\n\n                default:\n                    // Lowercase, NFC/NFD/NFKC/NFKD, Nmt, Prepend, ... all change content.\n                    return true;\n            }\n        }\n\n        /// <summary>\n        /// Builds the managed normalizer chain from a <c>tokenizer.json</c> normalizer element. Throws\n        /// <see cref=\"NotSupportedException\"/> for normalizer types that are not modeled so callers fail loudly\n        /// rather than silently mis-tokenizing.\n        /// </summary>\n        public static SentencePieceNormalizationStep Build(JsonElement normalizer)\n        {\n            if (normalizer.ValueKind != JsonValueKind.Object)\n            {\n                throw new InvalidDataException(\"A tokenizer.json normalizer entry must be a JSON object.\");\n            }\n\n            string? type = normalizer.TryGetProperty(\"type\", out JsonElement typeElement) && typeElement.ValueKind == JsonValueKind.String\n                ? typeElement.GetString() : null;\n            switch (type)\n            {\n                case \"Sequence\":\n                    var children = new List<SentencePieceNormalizationStep>();\n                    if (normalizer.TryGetProperty(\"normalizers\", out JsonElement steps) && steps.ValueKind == JsonValueKind.Array)\n                    {\n                        foreach (JsonElement step in steps.EnumerateArray())\n                        {\n                            children.Add(Build(step));\n                        }\n                    }\n                    return new SequenceStep(children);\n\n                case \"Precompiled\":","sourceCodeStart":113,"sourceCodeEnd":149,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs#L113-L149","documentation":"SentencePieceNormalizationStep.Build parses the 'normalizer' element of a Hugging Face tokenizer.json and requires it to be a JSON object. If the element is an array, string, null, or missing kind (e.g. a malformed file), InvalidDataException is thrown because the normalizer structure cannot be interpreted.","triggerScenarios":"Loading a tokenizer.json (via SentencePieceTokenizer/Unigram model loading) where the normalizer field is not a JSON object — e.g. an array of normalizers, a string, or a corrupted/truncated file.","commonSituations":"tokenizer.json files saved by newer/older tokenizers versions with a different normalizer schema; hand-edited tokenizer.json; passing the wrong JSON node (whole file instead of the normalizer section) to a custom loader.","solutions":["Validate the tokenizer.json: normalizer must be a JSON object ({\"type\": ...}), re-export with tokenizers if it isn't","Fix hand-edits so normalizer is an object (or remove it entirely if no normalization is wanted)","Ensure you're passing the normalizer element, not the root document, to Build"],"exampleFix":"// before\n\"normalizer\": [ {\"type\": \"NFD\"}, {\"type\": \"Lowercase\"} ]\n// after\n\"normalizer\": { \"type\": \"Sequence\", \"normalizers\": [ {\"type\": \"NFD\"}, {\"type\": \"Lowercase\"} ] }","handlingStrategy":"validation","validationCode":"using var doc = JsonDocument.Parse(tokenizerJson);\nif (doc.RootElement.TryGetProperty(\"normalizer\", out var n) && n.ValueKind != JsonValueKind.Object && n.ValueKind != JsonValueKind.Null)\n    throw new InvalidDataException(\"tokenizer.json 'normalizer' must be a JSON object\");","typeGuard":"bool IsValidNormalizerNode(JsonElement e) => e.ValueKind == JsonValueKind.Object || e.ValueKind == JsonValueKind.Null;","tryCatchPattern":"try { LoadTokenizer(tokenizerJson); } catch (InvalidDataException ex) when (ex.Message.Contains(\"normalizer\")) { log.LogError(ex, \"Malformed normalizer in tokenizer.json\"); }","preventionTips":["Re-export tokenizer.json with an up-to-date Hugging Face tokenizers version","Never hand-edit tokenizer.json without re-validating it","Store tokenizer files with checksums and verify before loading"],"tags":["tokenizers","json","sentencepiece"],"backgroundTag":"schema-validation-failed","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}