dotnet/machinelearning · error · InvalidDataException
The Precompiled normalizer 'precompiled_charsmap' must be a…
Error message
The Precompiled normalizer 'precompiled_charsmap' must be a string.
What it means
When the tokenizer.json normalizer type is 'Precompiled' (SentencePiece's precompiled charsmap), the 'precompiled_charsmap' property, if present, must be a JSON string (or null) containing base64 data. Any other JSON kind throws InvalidDataException.
Solutions
- Re-export the tokenizer.json with Hugging Face tokenizers so precompiled_charsmap is a base64 string
- Convert the charsmap bytes to base64 in the JSON if hand-authoring
- Remove the precompiled_charsmap property if no precompilation is needed (it is treated as default)
Example fix
// before "precompiled_charsmap": [210, 128, 1] // after "precompiled_charsmap": "4oKwAQA="
Defensive patterns
Strategy: validation
Validate before calling
if (normalizer.TryGetProperty("precompiled_charsmap", out var m) && m.ValueKind != JsonValueKind.String && m.ValueKind != JsonValueKind.Null)
throw new InvalidDataException("precompiled_charsmap must be a base64 string"); Type guard
bool IsValidCharsMap(JsonElement e) => !e.TryGetProperty("precompiled_charsmap", out var m) || m.ValueKind == JsonValueKind.String || m.ValueKind == JsonValueKind.Null; Try / catch
try { LoadTokenizer(tokenizerJson); } catch (InvalidDataException ex) when (ex.Message.Contains("precompiled_charsmap")) { log.LogError(ex, "Invalid Precompiled normalizer charsmap"); } Prevention
- Always export precompiled_charsmap as a base64 string
- Convert sentencepiece proto bytes to base64 when converting models manually
- Diff generated tokenizer.json against a known-good export
When it happens
Trigger: Loading a tokenizer.json whose normalizer is {"type": "Precompiled", "precompiled_charsmap": <non-string>} — e.g. the charsmap serialized as an array of bytes or a number instead of a base64 string.
Common situations: tokenizer.json produced by a tool that serializes charsmap differently; manual conversion of a sentencepiece model proto where the bytes were emitted as an array; corrupted file.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- A tokenizer.json normalizer entry must be a JSON object.
- The Prepend normalizer 'prepend' must be a string.
- ArgumentNullException
- Normalization ' ' is not supported.
- The Metaspace 'replacement
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/7e9c9cae54b63f69.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs:156
case "Sequence":
var children = new List<SentencePieceNormalizationStep>();
if (normalizer.TryGetProperty("normalizers", out JsonElement steps) && steps.ValueKind == JsonValueKind.Array)
{
foreach (JsonElement step in steps.EnumerateArray())
{
children.Add(Build(step));
}
}
return new SequenceStep(children);
case "Precompiled":
{
string? charsMap = null;
if (normalizer.TryGetProperty("precompiled_charsmap", out JsonElement mapElement))
{
if (mapElement.ValueKind != JsonValueKind.String && mapElement.ValueKind != JsonValueKind.Null)
{
throw new InvalidDataException("The Precompiled normalizer 'precompiled_charsmap' must be a string.");
}
charsMap = mapElement.GetString();
}
return new PrecompiledStep(string.IsNullOrEmpty(charsMap) ? default : DecodePrecompiledCharsMap(charsMap!));
}
case "Replace":
return ReplaceStep.Create(normalizer);
case "Strip":
return new StripStep(
stripLeft: !normalizer.TryGetProperty("strip_left", out JsonElement left) || left.ValueKind != JsonValueKind.False,
stripRight: !normalizer.TryGetProperty("strip_right", out JsonElement right) || right.ValueKind != JsonValueKind.False);
case "Lowercase":
return LowercaseStep.Instance;View on GitHub (pinned to 7b76e69cf9)