dotnet/machinelearning · error · InvalidDataException
The Prepend normalizer 'prepend' must be a string.
Error message
The Prepend normalizer 'prepend' must be a string.
What it means
For the 'Prepend' normalizer in a tokenizer.json, the optional 'prepend' property must be a JSON string (or null); any other kind throws InvalidDataException. Prepend inserts a literal string (e.g. '▁') at the start of the text during normalization.
Solutions
- Make the prepend value a quoted string, e.g. "prepend": "▁"
- Remove the prepend property entirely (it then defaults to empty string)
- Validate the tokenizer.json with Hugging Face tokenizers before loading it in .NET
Example fix
// before
{"type": "Prepend", "prepend": ['▁']}
// after
{"type": "Prepend", "prepend": "▁"} Defensive patterns
Strategy: validation
Validate before calling
if (normalizer.TryGetProperty("prepend", out var p) && p.ValueKind != JsonValueKind.String && p.ValueKind != JsonValueKind.Null)
throw new InvalidDataException("Prepend normalizer 'prepend' must be a string"); Type guard
bool IsValidPrepend(JsonElement e) => !e.TryGetProperty("prepend", out var p) || p.ValueKind == JsonValueKind.String || p.ValueKind == JsonValueKind.Null; Try / catch
try { LoadTokenizer(tokenizerJson); } catch (InvalidDataException ex) when (ex.Message.Contains("'prepend'")) { log.LogError(ex, "Invalid Prepend normalizer in tokenizer.json"); } Prevention
- Quote string values in hand-authored tokenizer.json
- Validate the whole tokenizer.json schema before deployment
- Use the Hugging Face tokenizers library to generate, not hand-write, normalizer entries
When it happens
Trigger: Loading a tokenizer.json with normalizer {"type": "Prepend", "prepend": <non-string>} such as a number, array, or object.
Common situations: Hand-written tokenizer.json where prepend was given as a char array or omitted quotes; a JSON export tool writing a different representation; copy/paste introducing wrong types.
Understand the failure class
Background: Type mismatch errors: IllegalArgumentException, TypeError and type guards across 150 open-source libraries — this error's family across 150 libraries.
Related errors
- A tokenizer.json normalizer entry must be a JSON object.
- The Precompiled normalizer 'precompiled_charsmap' must be a…
- ArgumentNullException
- Normalization ' ' is not supported.
- The Metaspace 'replacement
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/a06cfa82a6f82464.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs:193
case "StripAccents":
return StripAccentsStep.Instance;
case "NFC":
return new UnicodeStep(NormalizationForm.FormC);
case "NFD":
return new UnicodeStep(NormalizationForm.FormD);
case "NFKC":
return new UnicodeStep(NormalizationForm.FormKC);
case "NFKD":
return new UnicodeStep(NormalizationForm.FormKD);
case "Prepend":
{
if (normalizer.TryGetProperty("prepend", out JsonElement prependElement) &&
prependElement.ValueKind != JsonValueKind.String && prependElement.ValueKind != JsonValueKind.Null)
{
throw new InvalidDataException("The Prepend normalizer 'prepend' must be a string.");
}
string prepend = prependElement.ValueKind == JsonValueKind.String ? prependElement.GetString() ?? "" : "";
return new PrependStep(prepend);
}
case "Nmt":
return NmtStep.Instance;
default:
throw new NotSupportedException(
$"Unigram normalizer type '{type ?? "<missing>"}' is not supported when loading a tokenizer.json with content-modifying normalizer steps.");
}
}
// Decodes a base64 'precompiled_charsmap' value, surfacing malformed input as InvalidDataException so callers
// get a consistent, diagnostic failure for bad tokenizer.json files instead of a raw FormatException.
internal static byte[] DecodePrecompiledCharsMap(string base64)View on GitHub (pinned to 7b76e69cf9)