BoundaryML/baml · error

span.end is too large

Error message

span.end is too large

What it means

Same 32-bit TextSize conversion as the start offset, but for span.end: u32::try_from(span.end).expect("span.end is too large") panics when a token's end offset exceeds u32::MAX. This typically means the token spans a file larger than 4 GiB or the end offset was computed from bad bookkeeping.

Solutions

  1. Keep files under 4 GiB or split them into smaller modules.
  2. Validate that span.end equals the byte offset after the token and derives from the same buffer as span.start.
  3. Convert explicitly with error handling instead of expect so oversized spans produce a diagnostic rather than a panic.
  4. Audit any code adding lengths to offsets for double-accumulation bugs.

Example fix

// before
TextSize::from(u32::try_from(span.end).expect("span.end is too large")),
// after
let end = u32::try_from(span.end).map_err(|_| LexError::SpanTooLarge(span.end))?;
Defensive patterns

Strategy: validation

Validate before calling

if span.end > u32::MAX as usize || span.end > source_len {
    return Err(LexError::SpanTooLarge(span.end));
}

Type guard

fn fits_u32(n: usize) -> Option<u32> { u32::try_from(n).ok() }

Try / catch

// Panic-based; validate span.end against u32::MAX and the buffer length before constructing Token.

Prevention

When it happens

Trigger: Lexing input whose end offset (file length or token end) is > 4,294,967,295 bytes; or a span constructed with an end value not derived from the actual byte length.

Common situations: Giant generated BAML sources over 4 GiB; duplicated offset accumulation across concatenated source roots; external tooling supplying character counts instead of byte offsets.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of BoundaryML/baml@bd85ce9dee (2026-09-12). Data as JSON: /api/errors/0ce7350bc7ad2ee5. Report an issue: GitHub.

Appendix: source

Thrown at baml_language/crates/baml_compiler_lexer/src/tokens.rs:530

/// This tokenizes the entire input including whitespace and comments,
/// allowing perfect source reconstruction.
pub fn lex_lossless(input: &str, file_id: FileId) -> Vec<Token> {
    let mut tokens = Vec::new();
    let mut lexer = TokenKind::lexer(input);

    while let Some(result) = lexer.next() {
        let kind = result.unwrap_or(TokenKind::Error);
        let span = lexer.span();
        let text = lexer.slice().to_string();

        tokens.push(Token {
            kind,
            text,
            span: Span::new(
                file_id,
                TextRange::new(
                    TextSize::from(u32::try_from(span.start).expect("span.start is too large")),
                    TextSize::from(u32::try_from(span.end).expect("span.end is too large")),
                ),
            ),
        });
    }

    tokens
}

/// Reconstruct source from tokens (for testing losslessness).
pub fn reconstruct_source(tokens: &[Token]) -> String {
    tokens.iter().map(|t| t.text.as_str()).collect()
}

#[cfg(test)]
mod tests {
    use baml_base::FileId;

    use super::*;

View on GitHub (pinned to bd85ce9dee)