dotnet/machinelearning · error · InvalidDataException
A piece string in 'model.vocab' is null.
Error message
A piece string in 'model.vocab' is null.
What it means
A defensive check after GetString(): although the ValueKind check should have caught it, a piece that deserializes to a null string is rejected with InvalidDataException to prevent null pieces entering the vocabulary.
Solutions
- Replace the null piece with a valid non-empty piece string.
- Regenerate tokenizer.json from the original model so all pieces are populated.
- Pre-validate that every vocab entry has a non-null string in slot 0.
Example fix
// before "vocab": [[null, 0.0]] // after "vocab": [["<unk>", 0.0]]
Defensive patterns
Strategy: validation
Validate before calling
foreach (var e in model.GetProperty("vocab").EnumerateArray())
if (e[0].ValueKind != JsonValueKind.String || e[0].GetString() is null)
throw new InvalidOperationException("Vocab piece must be a non-null string."); Try / catch
try { return SentencePieceTokenizer.CreateFromTokenizerJson(stream); }
catch (InvalidDataException ex) when (ex.Message.Contains("piece string")) { /* regenerate tokenizer.json */ } Prevention
- Avoid hand-generating vocab arrays; export from the trained model.
- Check for null/missing fields after any JSON manipulation.
When it happens
Trigger: CreateFromTokenizerJson with a vocab entry whose piece element is JSON null or a value that GetString() yields null for.
Common situations: Programmatically generated tokenizer.json with missing piece values; corrupted or partially written vocab arrays.
Related errors
- Each entry in 'model.vocab' must be a [piece, score] array.
- Each entry in 'model.vocab' must be a [string piece, number…
- Expected model type 'Unigram' but found
- The tokenizer.json model does not contain a valid 'vocab'…
- The tokenizer.json model does not contain an 'unk_id'…
AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11).
Data as JSON: /api/errors/d57d7f5c051dc9af.
Report an issue: GitHub.
Appendix: source
Thrown at src/Microsoft.ML.Tokenizers/Model/SentencePieceTokenizer.cs:643
}
List<(string Piece, float Score)> vocab = new List<(string Piece, float Score)>(vocabElement.GetArrayLength());
foreach (JsonElement entry in vocabElement.EnumerateArray())
{
if (entry.ValueKind != JsonValueKind.Array || entry.GetArrayLength() < 2)
{
throw new InvalidDataException("Each entry in 'model.vocab' must be a [piece, score] array.");
}
if (entry[0].ValueKind != JsonValueKind.String || entry[1].ValueKind != JsonValueKind.Number)
{
throw new InvalidDataException("Each entry in 'model.vocab' must be a [string piece, number score] pair.");
}
string? piece = entry[0].GetString();
if (piece is null)
{
throw new InvalidDataException("A piece string in 'model.vocab' is null.");
}
vocab.Add((piece, entry[1].GetSingle()));
}
if (unkIsNull)
{
// Without an unknown token the only way to represent out-of-vocabulary input is byte fallback; a model
// with neither cannot encode OOV text, so reject that combination up front rather than emitting an
// invalid token id at encode time.
if (!byteFallback)
{
throw new NotSupportedException("The tokenizer.json model has a null 'unk_id' but does not enable 'byte_fallback'; a Unigram model without an unknown token is only supported when byte_fallback is enabled.");
}
}
else if (unkId < 0 || unkId >= vocab.Count)
{
throw new InvalidDataException($"The tokenizer.json model 'unk_id' ({unkId}) is out of range for a vocabulary of {vocab.Count} pieces.");View on GitHub (pinned to 7b76e69cf9)