{"record":{"id":"a06cfa82a6f82464","repo":"dotnet/machinelearning","slug":"the-prepend-normalizer-prepend-must-be-a-string","errorCode":null,"errorMessage":"The Prepend normalizer 'prepend' must be a string.","messagePattern":"The Prepend normalizer 'prepend' must be a string\\.","errorType":"validation","errorClass":"InvalidDataException","httpStatus":null,"severity":"error","filePath":"src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs","lineNumber":193,"sourceCode":"\n                case \"StripAccents\":\n                    return StripAccentsStep.Instance;\n\n                case \"NFC\":\n                    return new UnicodeStep(NormalizationForm.FormC);\n                case \"NFD\":\n                    return new UnicodeStep(NormalizationForm.FormD);\n                case \"NFKC\":\n                    return new UnicodeStep(NormalizationForm.FormKC);\n                case \"NFKD\":\n                    return new UnicodeStep(NormalizationForm.FormKD);\n\n                case \"Prepend\":\n                    {\n                        if (normalizer.TryGetProperty(\"prepend\", out JsonElement prependElement) &&\n                            prependElement.ValueKind != JsonValueKind.String && prependElement.ValueKind != JsonValueKind.Null)\n                        {\n                            throw new InvalidDataException(\"The Prepend normalizer 'prepend' must be a string.\");\n                        }\n\n                        string prepend = prependElement.ValueKind == JsonValueKind.String ? prependElement.GetString() ?? \"\" : \"\";\n                        return new PrependStep(prepend);\n                    }\n\n                case \"Nmt\":\n                    return NmtStep.Instance;\n\n                default:\n                    throw new NotSupportedException(\n                        $\"Unigram normalizer type '{type ?? \"<missing>\"}' is not supported when loading a tokenizer.json with content-modifying normalizer steps.\");\n            }\n        }\n\n        // Decodes a base64 'precompiled_charsmap' value, surfacing malformed input as InvalidDataException so callers\n        // get a consistent, diagnostic failure for bad tokenizer.json files instead of a raw FormatException.\n        internal static byte[] DecodePrecompiledCharsMap(string base64)","sourceCodeStart":175,"sourceCodeEnd":211,"githubUrl":"https://github.com/dotnet/machinelearning/blob/7b76e69cf964daeca3f1377af6bc5543284d56c6/src/Microsoft.ML.Tokenizers/Normalizer/SentencePieceNormalizationStep.cs#L175-L211","documentation":"For the 'Prepend' normalizer in a tokenizer.json, the optional 'prepend' property must be a JSON string (or null); any other kind throws InvalidDataException. Prepend inserts a literal string (e.g. '▁') at the start of the text during normalization.","triggerScenarios":"Loading a tokenizer.json with normalizer {\"type\": \"Prepend\", \"prepend\": <non-string>} such as a number, array, or object.","commonSituations":"Hand-written tokenizer.json where prepend was given as a char array or omitted quotes; a JSON export tool writing a different representation; copy/paste introducing wrong types.","solutions":["Make the prepend value a quoted string, e.g. \"prepend\": \"▁\"","Remove the prepend property entirely (it then defaults to empty string)","Validate the tokenizer.json with Hugging Face tokenizers before loading it in .NET"],"exampleFix":"// before\n{\"type\": \"Prepend\", \"prepend\": ['▁']}\n// after\n{\"type\": \"Prepend\", \"prepend\": \"▁\"}","handlingStrategy":"validation","validationCode":"if (normalizer.TryGetProperty(\"prepend\", out var p) && p.ValueKind != JsonValueKind.String && p.ValueKind != JsonValueKind.Null)\n    throw new InvalidDataException(\"Prepend normalizer 'prepend' must be a string\");","typeGuard":"bool IsValidPrepend(JsonElement e) => !e.TryGetProperty(\"prepend\", out var p) || p.ValueKind == JsonValueKind.String || p.ValueKind == JsonValueKind.Null;","tryCatchPattern":"try { LoadTokenizer(tokenizerJson); } catch (InvalidDataException ex) when (ex.Message.Contains(\"'prepend'\")) { log.LogError(ex, \"Invalid Prepend normalizer in tokenizer.json\"); }","preventionTips":["Quote string values in hand-authored tokenizer.json","Validate the whole tokenizer.json schema before deployment","Use the Hugging Face tokenizers library to generate, not hand-write, normalizer entries"],"tags":["tokenizers","json","sentencepiece"],"backgroundTag":"type-mismatch","analyzedSha":"7b76e69cf964daeca3f1377af6bc5543284d56c6","analyzedAt":"2026-09-11T12:35:38.930Z","contentChangedAt":"2026-09-11T12:35:38.930Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}