dotnet/machinelearning · error · NotSupportedException
The model ' ' is not supported.
Error message
The model '{modelName ?? modelEncoding.ToString()}' is not supported. What it means
When no explicit modelName is given, the constructor switches on the ModelEncoding enum; a value outside the known cases (default arm) throws NotSupportedException 'The model X is not supported.' This happens with invalid casts or enum values from a different (older/newer) library version.
Solutions
- Validate the ModelEncoding value before constructing (Enum.IsDefined) and reject unknown values at the config boundary.
- Align all Microsoft.ML.Tokenizers package versions (core + Data packages) so the enum members match.
- Use TiktokenTokenizer.CreateForModel with a string model name, which has broader name-based resolution.
- If unknown, fall back to creating the tokenizer from an explicit vocab stream + special tokens.
Example fix
// before
var enc = (ModelEncoding)int.Parse(cfg["encoding"]);
var tok = TiktokenTokenizer.Create(enc);
// after
var enc = (ModelEncoding)int.Parse(cfg["encoding"]);
if (!Enum.IsDefined(enc)) throw new InvalidOperationException($"Unknown encoding: {cfg["encoding"]}");
var tok = TiktokenTokenizer.Create(enc); Defensive patterns
Strategy: type-guard
Validate before calling
if (!Enum.IsDefined(modelEncoding))
throw new InvalidDataException($"Unknown ModelEncoding value: {(int)modelEncoding}"); Type guard
static bool IsDefinedEncoding(ModelEncoding e) => Enum.IsDefined(e);
Try / catch
try { var tok = TiktokenTokenizer.Create(encoding); }
catch (NotSupportedException ex)
{ logger.LogError(ex, "Unsupported ModelEncoding {V}", encoding); throw new InvalidDataException("Configured encoding is not supported by this library version", ex); } Prevention
- Never cast raw ints/strings to ModelEncoding without Enum.IsDefined/Enum.TryParse.
- Keep the core package and Data packages on the same version to avoid enum drift.
- Validate encoding values read from config at startup.
When it happens
Trigger: TiktokenTokenizer ctor / Create overload with modelEncoding set to an undefined enum value (e.g. (ModelEncoding)999 from config deserialization), or a ModelEncoding defined only in a newer/older assembly version.
Common situations: Deserializing a numeric encoding from config produced by a different library version; assembly version mismatch where a Data package references newer enum members; manual enum casting from external input.
Understand the failure class
Background: Invalid enum value errors: "Unknown type", "Invalid scope", "must be one of" — when a string is not on the library's allowed list — this error's family across 23 libraries.
Related errors
- The model ' ' is not supported.
- A piece string in 'model.vocab' is null.
- A post_processor template 'SpecialToken.id' must be a…
- An 'added_tokens' entry must have a string 'content' and a…
- argument should not be null.
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/1bdb8b6f4943f8f8.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs:1238
case ModelEncoding.O200kBase:
return (new Dictionary<string, int> { { EndOfText, 199999 }, { EndOfPrompt, 200018 } }, O200kBaseRegex(), O200kBaseFile, Type.GetType(O200kBaseTypeName), O200kBasePackageName);
case ModelEncoding.P50kBase:
return (new Dictionary<string, int> { { EndOfText, 50256 } }, P50kBaseRegex(), P50RanksFile, Type.GetType(P50kBaseTypeName), P50kBasePackageName);
case ModelEncoding.P50kEdit:
return (new Dictionary<string, int>
{ { EndOfText, 50256 }, { FimPrefix, 50281 }, { FimMiddle, 50282 }, { FimSuffix, 50283 } }, P50kBaseRegex(), P50RanksFile, Type.GetType(P50kBaseTypeName), P50kBasePackageName);
case ModelEncoding.R50kBase:
return (new Dictionary<string, int> { { EndOfText, 50256 } }, P50kBaseRegex(), R50RanksFile, Type.GetType(R50kBaseTypeName), R50kBasePackageName);
case ModelEncoding.O200kHarmony:
return (CreateHarmonyEncodingSpecialTokens(), O200kBaseRegex(), O200kBaseFile, Type.GetType(O200kBaseTypeName), O200kBasePackageName);
default:
throw new NotSupportedException($"The model '{modelName ?? modelEncoding.ToString()}' is not supported.");
}
}
// Regex patterns based on https://github.com/openai/tiktoken/blob/main/tiktoken_ext/openai_public.py
private const string Cl100kBaseRegexPattern = /*lang=regex*/ @"'(?i:[sdmt]|ll|ve|re)|(?>[^\r\n\p{L}\p{N}]?)(?>\p{L}+)|(?>\p{N}{1,3})| ?(?>[^\s\p{L}\p{N}]+)(?>[\r\n]*)|(?>\s+)$|\s*[\r\n]|\s+(?!\S)|\s";
private const string P50kBaseRegexPattern = /*lang=regex*/ @"'(?:[sdmt]|ll|ve|re)| ?(?>\p{L}+)| ?(?>\p{N}+)| ?(?>[^\s\p{L}\p{N}]+)|(?>\s+)$|\s+(?!\S)|\s";
private const string O200kBaseRegexPattern = /*lang=regex*/ @"[^\r\n\p{L}\p{N}]?[\p{Lu}\p{Lt}\p{Lm}\p{Lo}\p{M}]*[\p{Ll}\p{Lm}\p{Lo}\p{M}]+(?i:'s|'t|'re|'ve|'m|'ll|'d)?|[^\r\n\p{L}\p{N}]?[\p{Lu}\p{Lt}\p{Lm}\p{Lo}\p{M}]+[\p{Ll}\p{Lm}\p{Lo}\p{M}]*(?i:'s|'t|'re|'ve|'m|'ll|'d)?|\p{N}{1,3}| ?[^\s\p{L}\p{N}]+[\r\n/]*|\s*[\r\n]+|\s+(?!\S)|\s+";
private const string Cl100kBaseVocabFile = "cl100k_base.tiktoken.deflate"; // "https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken"
private const string P50RanksFile = "p50k_base.tiktoken.deflate"; // "https://openaipublic.blob.core.windows.net/encodings/p50k_base.tiktoken"
private const string R50RanksFile = "r50k_base.tiktoken.deflate"; // "https://openaipublic.blob.core.windows.net/encodings/r50k_base.tiktoken"
private const string GPT2File = "gpt2.tiktoken.deflate"; // "https://openaipublic.blob.core.windows.net/encodings/r50k_base.tiktoken". Gpt2 is using the same encoding as R50kBase
private const string O200kBaseFile = "o200k_base.tiktoken.deflate"; // "https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken"
internal const string Cl100kBaseEncodingName = "cl100k_base";
internal const string P50kBaseEncodingName = "p50k_base";
internal const string P50kEditEncodingName = "p50k_edit";View on GitHub (pinned to 7b76e69cf9)