zylon-ai/private-gpt · error · ValueError

MistralTokenizer is not available with the given…

Error message

MistralTokenizer is not available with the given configuration.

What it means

The 'mistral' entry in the tokenizer registry calls MistralTokenizer.is_available(**kwargs) first; that check simply requires a non-None model_id. If it is missing, the registry raises ValueError before attempting to load, telling you the mistral tokenizer cannot be built with the given configuration.

Solutions

  1. Pass a model_id when requesting the mistral tokenizer: TokenizerRegistry.get_tokenizer('mistral', model_id='mistralai/Mistral-Small-2412').
  2. Ensure the settings feeding the tokenizer (e.g. llm model field) are populated in every environment that selects tokenizer_mode='mistral'.
  3. If you have no model id, use a mode that does not need one ('estimator') or rely on 'default', which falls back automatically.

Example fix

# before
tok = TokenizerRegistry.get_tokenizer('mistral', model_id=settings.model)  # settings.model is None

# after
assert settings.model, 'tokenizer_mode=mistral requires a model id'
tok = TokenizerRegistry.get_tokenizer('mistral', model_id=settings.model)
Defensive patterns

Strategy: validation

Validate before calling

def mistral_mode_config_ok(settings) -> bool:
    if getattr(settings, 'tokenizer_mode', None) != 'mistral':
        return True
    return bool(getattr(settings, 'model', None))

assert mistral_mode_config_ok(settings), 'tokenizer_mode=mistral requires a model id'

Try / catch

try:
    tok = TokenizerRegistry.get_tokenizer('mistral', model_id=model_id)
except ValueError as e:
    if 'not available with the given configuration' in str(e):
        raise ConfigurationError('mistral tokenizer needs model_id') from e
    raise

Prevention

When it happens

Trigger: TokenizerRegistry.get_tokenizer('mistral', ...) without a model_id kwarg (or model_id=None) — e.g. a component resolving the tokenizer from settings where llm.model was never set, or a remote-mode config that carries no model path.

Common situations: Switching tokenizer_mode to 'mistral' while the LLM settings only define an API base/key and no local/Hub model; optional-model DI wiring where model is None for remote backends; env-specific settings files omitting the model key.

Related errors


AI-assisted analysis of zylon-ai/private-gpt@4a030776a3 (2026-08-15). Data as JSON: /api/errors/24b990fd2855577f. Report an issue: GitHub.

Appendix: source

Thrown at private_gpt/components/llm/tokenizers/registry.py:25

from private_gpt.components.llm.tokenizers.tokenizer_base import TokenizerBase

TokenizerProvider = Callable[..., TokenizerBase]

_EXTERNAL_TOKENIZER_FACTORIES: dict[str, TokenizerProvider] = {}


def register_tokenizer_factory(
    tokenizer_mode: str,
    factory: TokenizerProvider,
) -> None:
    _EXTERNAL_TOKENIZER_FACTORIES[tokenizer_mode] = factory


def _build_mistral_tokenizer(**kwargs: Any) -> TokenizerBase:
    from private_gpt.components.llm.tokenizers.mistral import MistralTokenizer

    if not MistralTokenizer.is_available(**kwargs):
        raise ValueError(
            "MistralTokenizer is not available with the given configuration."
        )

    return MistralTokenizer.from_pretrained(**kwargs)


def _build_tiktoken_tokenizer(**kwargs: Any) -> TokenizerBase:
    from private_gpt.components.llm.tokenizers.tiktoken import TikTokenTokenizer

    return TikTokenTokenizer.from_pretrained(**kwargs)


def _build_estimator_tokenizer(**kwargs: Any) -> TokenizerBase:
    from private_gpt.components.llm.tokenizers.estimator import EstimatorTokenizer

    return EstimatorTokenizer.from_pretrained(**kwargs)

View on GitHub (pinned to 4a030776a3)