dotnet/machinelearning · error · FormatException

Can't parse to integer

Error message

Can't parse {line.Substring(spaceIndex)} to integer

What it means

After locating the single space in a vocab line, the substring after it must parse as a 32-bit integer rank. If not, a FormatException 'Can't parse <suffix> to integer' is thrown, identifying the offending text.

Solutions

  1. Fix the offending line so the text after the single space is a valid decimal integer rank.
  2. Re-download or regenerate the vocab file (the message names the exact bad substring to locate it).
  3. Ensure ranks are within Int32 range and written in plain decimal.
  4. Use the official data package files instead of hand-modified vocab streams.

Example fix

// before
dGhl 12ab
// after
dGhl 12
Defensive patterns

Strategy: validation

Validate before calling

static void ValidateBpeRank(string line)
{
    int sp = line.IndexOf(' ');
    if (sp > 0 && !int.TryParse(line.AsSpan(sp + 1), out _))
        throw new FormatException($"Rank is not an integer: '{line.Substring(sp)}'");
}

Type guard

static bool HasValidRank(string line)
{
    int sp = line.IndexOf(' ');
    return sp > 0 && int.TryParse(line.AsSpan(sp + 1), out _);
}

Try / catch

try { var tok = await TiktokenTokenizer.CreateAsync(vocabStream, specialTokens); }
catch (InvalidOperationException ex) when (ex.InnerException is FormatException fe)
{ throw new InvalidDataException($"Vocab rank parse failure: {fe.Message}", fe); }

Prevention

When it happens

Trigger: TiktokenTokenizer.CreateAsync/CreateForModel/TiktokenTokenizer ctor reading a vocab stream (LoadTiktokenBpeAsync) where a line like 'dGhl abc' or 'dGhl 999999999999' has a non-numeric or out-of-int32-range rank after the space.

Common situations: Corrupted or partially-written vocab file; ranks pasted as hex ('0x1a') or with signs/floats; line truncated mid-rank by a bad download.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of dotnet/machinelearning@7b76e69cf9 (2026-09-11). Data as JSON: /api/errors/4af853514ab3c86a. Report an issue: GitHub.

Appendix: source

Thrown at src/Microsoft.ML.Tokenizers/Model/TiktokenTokenizer.cs:222

                }

                while (line is not null)
                {
                    if (line.Length > 0)
                    {
                        int spaceIndex = line.IndexOf(' ');
                        if (spaceIndex <= 0 || spaceIndex >= line.Length - 1 || line.IndexOf(' ', spaceIndex + 1) >= 0)
                        {
                            throw new FormatException($"Invalid format in the BPE vocab file stream");
                        }

                        if (Helpers.TryParseInt32(line, spaceIndex + 1, out int rank))
                        {
                            AddData(Helpers.FromBase64String(line, 0, spaceIndex), rank);
                        }
                        else
                        {
                            throw new FormatException($"Can't parse {line.Substring(spaceIndex)} to integer");
                        }

                        line = useAsync ?
                            await Helpers.ReadLineAsync(reader, cancellationToken).ConfigureAwait(false) :
                            reader.ReadLine();
                    }
                }
            }
            catch (Exception ex)
            {
                throw new InvalidOperationException($"Failed to load from BPE vocab file stream: {ex.Message}", ex);
            }

            return (encoder, vocab, decoder);

            void AddData(byte[] tokenBytes, int rank)
            {
                encoder[tokenBytes] = rank;

View on GitHub (pinned to 7b76e69cf9)