dotnet/machinelearning · error · InvalidDataException

A tokenizer.json normalizer entry must be a JSON object.

Error message

A tokenizer.json normalizer entry must be a JSON object.

What it means

SentencePieceNormalizationStep.Build parses the 'normalizer' element of a Hugging Face tokenizer.json and requires it to be a JSON object. If the element is an array, string, null, or missing kind (e.g. a malformed file), InvalidDataException is thrown because the normalizer structure cannot be interpreted.

Solutions

  1. Validate the tokenizer.json: normalizer must be a JSON object ({"type": ...}), re-export with tokenizers if it isn't
  2. Fix hand-edits so normalizer is an object (or remove it entirely if no normalization is wanted)
  3. Ensure you're passing the normalizer element, not the root document, to Build

Example fix

// before
"normalizer": [ {"type": "NFD"}, {"type": "Lowercase"} ]
// after
"normalizer": { "type": "Sequence", "normalizers": [ {"type": "NFD"}, {"type": "Lowercase"} ] }
Defensive patterns

Strategy: validation

Validate before calling

using var doc = JsonDocument.Parse(tokenizerJson);
if (doc.RootElement.TryGetProperty("normalizer", out var n) && n.ValueKind != JsonValueKind.Object && n.ValueKind != JsonValueKind.Null)
    throw new InvalidDataException("tokenizer.json 'normalizer' must be a JSON object");

Type guard

bool IsValidNormalizerNode(JsonElement e) => e.ValueKind == JsonValueKind.Object || e.ValueKind == JsonValueKind.Null;

Try / catch

try { LoadTokenizer(tokenizerJson); } catch (InvalidDataException ex) when (ex.Message.Contains("normalizer")) { log.LogError(ex, "Malformed normalizer in tokenizer.json"); }

Prevention

When it happens

Trigger: Loading a tokenizer.json (via SentencePieceTokenizer/Unigram model loading) where the normalizer field is not a JSON object — e.g. an array of normalizers, a string, or a corrupted/truncated file.

Common situations: tokenizer.json files saved by newer/older tokenizers versions with a different normalizer schema; hand-edited tokenizer.json; passing the wrong JSON node (whole file instead of the normalizer section) to a custom loader.

Understand the failure class

Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/27c4792ba898de2c. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs:131

                    // Replace (literal punctuation splitting, character substitution) is content-modifying.
                    return !ReplaceIsWhitespaceCollapse(normalizer);

                default:
                    // Lowercase, NFC/NFD/NFKC/NFKD, Nmt, Prepend, ... all change content.
                    return true;
            }
        }

        /// <summary>
        /// Builds the managed normalizer chain from a <c>tokenizer.json</c> normalizer element. Throws
        /// <see cref="NotSupportedException"/> for normalizer types that are not modeled so callers fail loudly
        /// rather than silently mis-tokenizing.
        /// </summary>
        public static SentencePieceNormalizationStep Build(JsonElement normalizer)
        {
            if (normalizer.ValueKind != JsonValueKind.Object)
            {
                throw new InvalidDataException("A tokenizer.json normalizer entry must be a JSON object.");
            }

            string? type = normalizer.TryGetProperty("type", out JsonElement typeElement) && typeElement.ValueKind == JsonValueKind.String
                ? typeElement.GetString() : null;
            switch (type)
            {
                case "Sequence":
                    var children = new List<SentencePieceNormalizationStep>();
                    if (normalizer.TryGetProperty("normalizers", out JsonElement steps) && steps.ValueKind == JsonValueKind.Array)
                    {
                        foreach (JsonElement step in steps.EnumerateArray())
                        {
                            children.Add(Build(step));
                        }
                    }
                    return new SequenceStep(children);

                case "Precompiled":

View on GitHub (pinned to 7b76e69cf9)