dotnet/machinelearning · error · InvalidDataException

A piece string in 'model.vocab' is null.

Error message

A piece string in 'model.vocab' is null.

What it means

A defensive check after GetString(): although the ValueKind check should have caught it, a piece that deserializes to a null string is rejected with InvalidDataException to prevent null pieces entering the vocabulary.

Solutions

  1. Replace the null piece with a valid non-empty piece string.
  2. Regenerate tokenizer.json from the original model so all pieces are populated.
  3. Pre-validate that every vocab entry has a non-null string in slot 0.

Example fix

// before
"vocab": [[null, 0.0]]
// after
"vocab": [["<unk>", 0.0]]
Defensive patterns

Strategy: validation

Validate before calling

foreach (var e in model.GetProperty("vocab").EnumerateArray())
    if (e[0].ValueKind != JsonValueKind.String || e[0].GetString() is null)
        throw new InvalidOperationException("Vocab piece must be a non-null string.");

Try / catch

try { return SentencePieceTokenizer.CreateFromTokenizerJson(stream); }
catch (InvalidDataException ex) when (ex.Message.Contains("piece string")) { /* regenerate tokenizer.json */ }

Prevention

When it happens

Trigger: CreateFromTokenizerJson with a vocab entry whose piece element is JSON null or a value that GetString() yields null for.

Common situations: Programmatically generated tokenizer.json with missing piece values; corrupted or partially written vocab arrays.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/d57d7f5c051dc9af. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/SentencePieceTokenizer.cs:643

            }

            List<(string Piece, float Score)> vocab = new List<(string Piece, float Score)>(vocabElement.GetArrayLength());
            foreach (JsonElement entry in vocabElement.EnumerateArray())
            {
                if (entry.ValueKind != JsonValueKind.Array || entry.GetArrayLength() < 2)
                {
                    throw new InvalidDataException("Each entry in 'model.vocab' must be a [piece, score] array.");
                }

                if (entry[0].ValueKind != JsonValueKind.String || entry[1].ValueKind != JsonValueKind.Number)
                {
                    throw new InvalidDataException("Each entry in 'model.vocab' must be a [string piece, number score] pair.");
                }

                string? piece = entry[0].GetString();
                if (piece is null)
                {
                    throw new InvalidDataException("A piece string in 'model.vocab' is null.");
                }

                vocab.Add((piece, entry[1].GetSingle()));
            }

            if (unkIsNull)
            {
                // Without an unknown token the only way to represent out-of-vocabulary input is byte fallback; a model
                // with neither cannot encode OOV text, so reject that combination up front rather than emitting an
                // invalid token id at encode time.
                if (!byteFallback)
                {
                    throw new NotSupportedException("The tokenizer.json model has a null 'unk_id' but does not enable 'byte_fallback'; a Unigram model without an unknown token is only supported when byte_fallback is enabled.");
                }
            }
            else if (unkId < 0 || unkId >= vocab.Count)
            {
                throw new InvalidDataException($"The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces.");

View on GitHub (pinned to 7b76e69cf9)