dotnet/machinelearning · error · InvalidDataException

The Prepend normalizer 'prepend' must be a string.

Error message

The Prepend normalizer 'prepend' must be a string.

What it means

For the 'Prepend' normalizer in a tokenizer.json, the optional 'prepend' property must be a JSON string (or null); any other kind throws InvalidDataException. Prepend inserts a literal string (e.g. '▁') at the start of the text during normalization.

Solutions

  1. Make the prepend value a quoted string, e.g. "prepend": "▁"
  2. Remove the prepend property entirely (it then defaults to empty string)
  3. Validate the tokenizer.json with Hugging Face tokenizers before loading it in .NET

Example fix

// before
{"type": "Prepend", "prepend": ['▁']}
// after
{"type": "Prepend", "prepend": "▁"}
Defensive patterns

Strategy: validation

Validate before calling

if (normalizer.TryGetProperty("prepend", out var p) && p.ValueKind != JsonValueKind.String && p.ValueKind != JsonValueKind.Null)
    throw new InvalidDataException("Prepend normalizer 'prepend' must be a string");

Type guard

bool IsValidPrepend(JsonElement e) => !e.TryGetProperty("prepend", out var p) || p.ValueKind == JsonValueKind.String || p.ValueKind == JsonValueKind.Null;

Try / catch

try { LoadTokenizer(tokenizerJson); } catch (InvalidDataException ex) when (ex.Message.Contains("'prepend'")) { log.LogError(ex, "Invalid Prepend normalizer in tokenizer.json"); }

Prevention

When it happens

Trigger: Loading a tokenizer.json with normalizer {"type": "Prepend", "prepend": <non-string>} such as a number, array, or object.

Common situations: Hand-written tokenizer.json where prepend was given as a char array or omitted quotes; a JSON export tool writing a different representation; copy/paste introducing wrong types.

Understand the failure class

Background: Type mismatch errors: IllegalArgumentException, TypeError and type guards across 150 open-source libraries — this error's family across 150 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/a06cfa82a6f82464. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs:193

                case "StripAccents":
                    return StripAccentsStep.Instance;

                case "NFC":
                    return new UnicodeStep(NormalizationForm.FormC);
                case "NFD":
                    return new UnicodeStep(NormalizationForm.FormD);
                case "NFKC":
                    return new UnicodeStep(NormalizationForm.FormKC);
                case "NFKD":
                    return new UnicodeStep(NormalizationForm.FormKD);

                case "Prepend":
                    {
                        if (normalizer.TryGetProperty("prepend", out JsonElement prependElement) &&
                            prependElement.ValueKind != JsonValueKind.String && prependElement.ValueKind != JsonValueKind.Null)
                        {
                            throw new InvalidDataException("The Prepend normalizer 'prepend' must be a string.");
                        }

                        string prepend = prependElement.ValueKind == JsonValueKind.String ? prependElement.GetString() ?? "" : "";
                        return new PrependStep(prepend);
                    }

                case "Nmt":
                    return NmtStep.Instance;

                default:
                    throw new NotSupportedException(
                        $"Unigram normalizer type '{type ?? "<missing>"}' is not supported when loading a tokenizer.json with content-modifying normalizer steps.");
            }
        }

        // Decodes a base64 'precompiled_charsmap' value, surfacing malformed input as InvalidDataException so callers
        // get a consistent, diagnostic failure for bad tokenizer.json files instead of a raw FormatException.
        internal static byte[] DecodePrecompiledCharsMap(string base64)

View on GitHub (pinned to 7b76e69cf9)