dotnet/machinelearning · error · NotSupportedException

The model ' ' is not supported.

Error message

The model '{modelName ?? modelEncoding.ToString()}' is not supported.

What it means

When no explicit modelName is given, the constructor switches on the ModelEncoding enum; a value outside the known cases (default arm) throws NotSupportedException 'The model X is not supported.' This happens with invalid casts or enum values from a different (older/newer) library version.

Solutions

  1. Validate the ModelEncoding value before constructing (Enum.IsDefined) and reject unknown values at the config boundary.
  2. Align all Microsoft.ML.Tokenizers package versions (core + Data packages) so the enum members match.
  3. Use TiktokenTokenizer.CreateForModel with a string model name, which has broader name-based resolution.
  4. If unknown, fall back to creating the tokenizer from an explicit vocab stream + special tokens.

Example fix

// before
var enc = (ModelEncoding)int.Parse(cfg["encoding"]);
var tok = TiktokenTokenizer.Create(enc);
// after
var enc = (ModelEncoding)int.Parse(cfg["encoding"]);
if (!Enum.IsDefined(enc)) throw new InvalidOperationException($"Unknown encoding: {cfg["encoding"]}");
var tok = TiktokenTokenizer.Create(enc);
Defensive patterns

Strategy: type-guard

Validate before calling

if (!Enum.IsDefined(modelEncoding))
    throw new InvalidDataException($"Unknown ModelEncoding value: {(int)modelEncoding}");

Type guard

static bool IsDefinedEncoding(ModelEncoding e) => Enum.IsDefined(e);

Try / catch

try { var tok = TiktokenTokenizer.Create(encoding); }
catch (NotSupportedException ex)
{ logger.LogError(ex, "Unsupported ModelEncoding {V}", encoding); throw new InvalidDataException("Configured encoding is not supported by this library version", ex); }

Prevention

When it happens

Trigger: TiktokenTokenizer ctor / Create overload with modelEncoding set to an undefined enum value (e.g. (ModelEncoding)999 from config deserialization), or a ModelEncoding defined only in a newer/older assembly version.

Common situations: Deserializing a numeric encoding from config produced by a different library version; assembly version mismatch where a Data package references newer enum members; manual enum casting from external input.

Understand the failure class

Background: Invalid enum value errors: "Unknown type", "Invalid scope", "must be one of" — when a string is not on the library's allowed list — this error's family across 23 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/1bdb8b6f4943f8f8. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs:1238

                case ModelEncoding.O200kBase:
                    return (new Dictionary<string, int> { { EndOfText, 199999 }, { EndOfPrompt, 200018 } }, O200kBaseRegex(), O200kBaseFile, Type.GetType(O200kBaseTypeName), O200kBasePackageName);

                case ModelEncoding.P50kBase:
                    return (new Dictionary<string, int> { { EndOfText, 50256 } }, P50kBaseRegex(), P50RanksFile, Type.GetType(P50kBaseTypeName), P50kBasePackageName);

                case ModelEncoding.P50kEdit:
                    return (new Dictionary<string, int>
                        { { EndOfText, 50256 }, { FimPrefix, 50281 }, { FimMiddle, 50282 }, { FimSuffix, 50283 } }, P50kBaseRegex(), P50RanksFile, Type.GetType(P50kBaseTypeName), P50kBasePackageName);

                case ModelEncoding.R50kBase:
                    return (new Dictionary<string, int> { { EndOfText, 50256 } }, P50kBaseRegex(), R50RanksFile, Type.GetType(R50kBaseTypeName), R50kBasePackageName);

                case ModelEncoding.O200kHarmony:
                    return (CreateHarmonyEncodingSpecialTokens(), O200kBaseRegex(), O200kBaseFile, Type.GetType(O200kBaseTypeName), O200kBasePackageName);

                default:
                    throw new NotSupportedException($"The model '{modelName ?? modelEncoding.ToString()}' is not supported.");
            }
        }

        // Regex patterns based on https://github.com/openai/tiktoken/blob/main/tiktoken_ext/openai_public.py

        private const string Cl100kBaseRegexPattern = /*lang=regex*/ @"'(?i:[sdmt]|ll|ve|re)|(?>[^\r\n\p{L}\p{N}]?)(?>\p{L}+)|(?>\p{N}{1,3})| ?(?>[^\s\p{L}\p{N}]+)(?>[\r\n]*)|(?>\s+)$|\s*[\r\n]|\s+(?!\S)|\s";
        private const string P50kBaseRegexPattern = /*lang=regex*/ @"'(?:[sdmt]|ll|ve|re)| ?(?>\p{L}+)| ?(?>\p{N}+)| ?(?>[^\s\p{L}\p{N}]+)|(?>\s+)$|\s+(?!\S)|\s";
        private const string O200kBaseRegexPattern = /*lang=regex*/ @"[^\r\n\p{L}\p{N}]?[\p{Lu}\p{Lt}\p{Lm}\p{Lo}\p{M}]*[\p{Ll}\p{Lm}\p{Lo}\p{M}]+(?i:'s|'t|'re|'ve|'m|'ll|'d)?|[^\r\n\p{L}\p{N}]?[\p{Lu}\p{Lt}\p{Lm}\p{Lo}\p{M}]+[\p{Ll}\p{Lm}\p{Lo}\p{M}]*(?i:'s|'t|'re|'ve|'m|'ll|'d)?|\p{N}{1,3}| ?[^\s\p{L}\p{N}]+[\r\n/]*|\s*[\r\n]+|\s+(?!\S)|\s+";

        private const string Cl100kBaseVocabFile = "cl100k_base.tiktoken.deflate";  // "https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken"
        private const string P50RanksFile = "p50k_base.tiktoken.deflate";           // "https://openaipublic.blob.core.windows.net/encodings/p50k_base.tiktoken"
        private const string R50RanksFile = "r50k_base.tiktoken.deflate";           // "https://openaipublic.blob.core.windows.net/encodings/r50k_base.tiktoken"
        private const string GPT2File = "gpt2.tiktoken.deflate";                    // "https://openaipublic.blob.core.windows.net/encodings/r50k_base.tiktoken". Gpt2 is using the same encoding as R50kBase
        private const string O200kBaseFile = "o200k_base.tiktoken.deflate";         // "https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken"

        internal const string Cl100kBaseEncodingName = "cl100k_base";
        internal const string P50kBaseEncodingName = "p50k_base";
        internal const string P50kEditEncodingName = "p50k_edit";

View on GitHub (pinned to 7b76e69cf9)