BoundaryML/baml · error

span.start is too large

Error message

span.start is too large

What it means

The lexer's token collector converts each span offset to TextSize (u32) via u32::try_from(span.start).expect("span.start is too large"). It panics when a token's start offset exceeds u32::MAX, i.e. a source file larger than 4 GiB (or a span computed from an oversized/incorrect offset). rust-analyzer-style TextSize is deliberately 32-bit.

Solutions

  1. Keep source files below 4 GiB; split very large generated files into modules.
  2. If offsets come from concatenating sources, verify the base offset math so spans don't accumulate past the file size.
  3. Handle the conversion explicitly: skip or error on oversized spans with a proper diagnostic instead of panicking.
  4. Check for upstream data (e.g. an external index) reporting byte offsets in units other than bytes.

Example fix

// before
TextSize::from(u32::try_from(span.start).expect("span.start is too large")),
// after
let start = u32::try_from(span.start).map_err(|_| LexError::SpanTooLarge(span.start))?;
Defensive patterns

Strategy: validation

Validate before calling

if span.start > u32::MAX as usize {
    return Err(LexError::SpanTooLarge(span.start));
}

Type guard

fn fits_u32(n: usize) -> Option<u32> { u32::try_from(n).ok() }

Try / catch

// This is a panic, not a Result; pre-validate offsets. If wrapping a whole lex run:
let toks = std::panic::catch_unwind(|| lex(src)).map_err(|_| LexError::SourceTooLarge)?;

Prevention

When it happens

Trigger: Lexing a file whose byte size (or the token's absolute start offset) is > 4,294,967,295 bytes; or feeding Token::new a span whose start was computed incorrectly (e.g. accumulated on the wrong base) yielding a value beyond u32::MAX.

Common situations: Real-world: pointing the compiler at a generated or concatenated BAML file over 4 GiB; or a bug in offset bookkeeping when splicing multiple source roots so offsets are added twice.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of BoundaryML/baml@bd85ce9dee (2026-09-12). Data as JSON: /api/errors/d9f4a6b93196f49b. Report an issue: GitHub.

Appendix: source

Thrown at baml_language/crates/baml_compiler_lexer/src/tokens.rs:529

///
/// This tokenizes the entire input including whitespace and comments,
/// allowing perfect source reconstruction.
pub fn lex_lossless(input: &str, file_id: FileId) -> Vec<Token> {
    let mut tokens = Vec::new();
    let mut lexer = TokenKind::lexer(input);

    while let Some(result) = lexer.next() {
        let kind = result.unwrap_or(TokenKind::Error);
        let span = lexer.span();
        let text = lexer.slice().to_string();

        tokens.push(Token {
            kind,
            text,
            span: Span::new(
                file_id,
                TextRange::new(
                    TextSize::from(u32::try_from(span.start).expect("span.start is too large")),
                    TextSize::from(u32::try_from(span.end).expect("span.end is too large")),
                ),
            ),
        });
    }

    tokens
}

/// Reconstruct source from tokens (for testing losslessness).
pub fn reconstruct_source(tokens: &[Token]) -> String {
    tokens.iter().map(|t| t.text.as_str()).collect()
}

#[cfg(test)]
mod tests {
    use baml_base::FileId;

View on GitHub (pinned to bd85ce9dee)