dotnet/machinelearning · error · InvalidDataException
A tokenizer.json normalizer entry must be a JSON object.
Error message
A tokenizer.json normalizer entry must be a JSON object.
What it means
SentencePieceNormalizationStep.Build parses the 'normalizer' element of a Hugging Face tokenizer.json and requires it to be a JSON object. If the element is an array, string, null, or missing kind (e.g. a malformed file), InvalidDataException is thrown because the normalizer structure cannot be interpreted.
Solutions
- Validate the tokenizer.json: normalizer must be a JSON object ({"type": ...}), re-export with tokenizers if it isn't
- Fix hand-edits so normalizer is an object (or remove it entirely if no normalization is wanted)
- Ensure you're passing the normalizer element, not the root document, to Build
Example fix
// before
"normalizer": [ {"type": "NFD"}, {"type": "Lowercase"} ]
// after
"normalizer": { "type": "Sequence", "normalizers": [ {"type": "NFD"}, {"type": "Lowercase"} ] } Defensive patterns
Strategy: validation
Validate before calling
using var doc = JsonDocument.Parse(tokenizerJson);
if (doc.RootElement.TryGetProperty("normalizer", out var n) && n.ValueKind != JsonValueKind.Object && n.ValueKind != JsonValueKind.Null)
throw new InvalidDataException("tokenizer.json 'normalizer' must be a JSON object"); Type guard
bool IsValidNormalizerNode(JsonElement e) => e.ValueKind == JsonValueKind.Object || e.ValueKind == JsonValueKind.Null;
Try / catch
try { LoadTokenizer(tokenizerJson); } catch (InvalidDataException ex) when (ex.Message.Contains("normalizer")) { log.LogError(ex, "Malformed normalizer in tokenizer.json"); } Prevention
- Re-export tokenizer.json with an up-to-date Hugging Face tokenizers version
- Never hand-edit tokenizer.json without re-validating it
- Store tokenizer files with checksums and verify before loading
When it happens
Trigger: Loading a tokenizer.json (via SentencePieceTokenizer/Unigram model loading) where the normalizer field is not a JSON object — e.g. an array of normalizers, a string, or a corrupted/truncated file.
Common situations: tokenizer.json files saved by newer/older tokenizers versions with a different normalizer schema; hand-edited tokenizer.json; passing the wrong JSON node (whole file instead of the normalizer section) to a custom loader.
Understand the failure class
Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.
Related errors
- The Precompiled normalizer 'precompiled_charsmap' must be a…
- The Prepend normalizer 'prepend' must be a string.
- ArgumentNullException
- Normalization ' ' is not supported.
- The Metaspace 'replacement
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/27c4792ba898de2c.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs:131
// Replace (literal punctuation splitting, character substitution) is content-modifying.
return !ReplaceIsWhitespaceCollapse(normalizer);
default:
// Lowercase, NFC/NFD/NFKC/NFKD, Nmt, Prepend, ... all change content.
return true;
}
}
/// <summary>
/// Builds the managed normalizer chain from a <c>tokenizer.json</c> normalizer element. Throws
/// <see cref="NotSupportedException"/> for normalizer types that are not modeled so callers fail loudly
/// rather than silently mis-tokenizing.
/// </summary>
public static SentencePieceNormalizationStep Build(JsonElement normalizer)
{
if (normalizer.ValueKind != JsonValueKind.Object)
{
throw new InvalidDataException("A tokenizer.json normalizer entry must be a JSON object.");
}
string? type = normalizer.TryGetProperty("type", out JsonElement typeElement) && typeElement.ValueKind == JsonValueKind.String
? typeElement.GetString() : null;
switch (type)
{
case "Sequence":
var children = new List<SentencePieceNormalizationStep>();
if (normalizer.TryGetProperty("normalizers", out JsonElement steps) && steps.ValueKind == JsonValueKind.Array)
{
foreach (JsonElement step in steps.EnumerateArray())
{
children.Add(Build(step));
}
}
return new SequenceStep(children);
case "Precompiled":View on GitHub (pinned to 7b76e69cf9)