BoundaryML/baml · error
span.start is too large
Error message
span.start is too large
What it means
The lexer's token collector converts each span offset to TextSize (u32) via u32::try_from(span.start).expect("span.start is too large"). It panics when a token's start offset exceeds u32::MAX, i.e. a source file larger than 4 GiB (or a span computed from an oversized/incorrect offset). rust-analyzer-style TextSize is deliberately 32-bit.
Solutions
- Keep source files below 4 GiB; split very large generated files into modules.
- If offsets come from concatenating sources, verify the base offset math so spans don't accumulate past the file size.
- Handle the conversion explicitly: skip or error on oversized spans with a proper diagnostic instead of panicking.
- Check for upstream data (e.g. an external index) reporting byte offsets in units other than bytes.
Example fix
// before
TextSize::from(u32::try_from(span.start).expect("span.start is too large")),
// after
let start = u32::try_from(span.start).map_err(|_| LexError::SpanTooLarge(span.start))?; Defensive patterns
Strategy: validation
Validate before calling
if span.start > u32::MAX as usize {
return Err(LexError::SpanTooLarge(span.start));
} Type guard
fn fits_u32(n: usize) -> Option<u32> { u32::try_from(n).ok() } Try / catch
// This is a panic, not a Result; pre-validate offsets. If wrapping a whole lex run: let toks = std::panic::catch_unwind(|| lex(src)).map_err(|_| LexError::SourceTooLarge)?;
Prevention
- Reject source files larger than 4 GiB before lexing.
- Verify offset bookkeeping when concatenating or splicing sources.
- Ensure spans are byte offsets (not char indices) from the same buffer.
When it happens
Trigger: Lexing a file whose byte size (or the token's absolute start offset) is > 4,294,967,295 bytes; or feeding Token::new a span whose start was computed incorrectly (e.g. accumulated on the wrong base) yielding a value beyond u32::MAX.
Common situations: Real-world: pointing the compiler at a generated or concatenated BAML file over 4 GiB; or a bug in offset bookkeeping when splicing multiple source roots so offsets are added twice.
Understand the failure class
Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.
Related errors
- span.end is too large
- ai.Prompt._data must contain baml_builtins2::PromptAst
- ai.Prompt.messages receiver must be an ai.Prompt instance
- array access should be either map or array.
- ======================================== BAML Internal…
AI-assisted analysis of BoundaryML/baml@bd85ce9dee (2026-09-12).
Data as JSON: /api/errors/d9f4a6b93196f49b.
Report an issue: GitHub.
Appendix: source
Thrown at baml_language/crates/baml_compiler_lexer/src/tokens.rs:529
///
/// This tokenizes the entire input including whitespace and comments,
/// allowing perfect source reconstruction.
pub fn lex_lossless(input: &str, file_id: FileId) -> Vec<Token> {
let mut tokens = Vec::new();
let mut lexer = TokenKind::lexer(input);
while let Some(result) = lexer.next() {
let kind = result.unwrap_or(TokenKind::Error);
let span = lexer.span();
let text = lexer.slice().to_string();
tokens.push(Token {
kind,
text,
span: Span::new(
file_id,
TextRange::new(
TextSize::from(u32::try_from(span.start).expect("span.start is too large")),
TextSize::from(u32::try_from(span.end).expect("span.end is too large")),
),
),
});
}
tokens
}
/// Reconstruct source from tokens (for testing losslessness).
pub fn reconstruct_source(tokens: &[Token]) -> String {
tokens.iter().map(|t| t.text.as_str()).collect()
}
#[cfg(test)]
mod tests {
use baml_base::FileId;
View on GitHub (pinned to bd85ce9dee)