{"record":{"id":"d9f4a6b93196f49b","repo":"BoundaryML/baml","slug":"span-start-is-too-large","errorCode":null,"errorMessage":"span.start is too large","messagePattern":"span\\.start is too large","errorType":"panic","errorClass":null,"httpStatus":null,"severity":"error","filePath":"baml_language/crates/baml_compiler_lexer/src/tokens.rs","lineNumber":529,"sourceCode":"///\n/// This tokenizes the entire input including whitespace and comments,\n/// allowing perfect source reconstruction.\npub fn lex_lossless(input: &str, file_id: FileId) -> Vec<Token> {\n    let mut tokens = Vec::new();\n    let mut lexer = TokenKind::lexer(input);\n\n    while let Some(result) = lexer.next() {\n        let kind = result.unwrap_or(TokenKind::Error);\n        let span = lexer.span();\n        let text = lexer.slice().to_string();\n\n        tokens.push(Token {\n            kind,\n            text,\n            span: Span::new(\n                file_id,\n                TextRange::new(\n                    TextSize::from(u32::try_from(span.start).expect(\"span.start is too large\")),\n                    TextSize::from(u32::try_from(span.end).expect(\"span.end is too large\")),\n                ),\n            ),\n        });\n    }\n\n    tokens\n}\n\n/// Reconstruct source from tokens (for testing losslessness).\npub fn reconstruct_source(tokens: &[Token]) -> String {\n    tokens.iter().map(|t| t.text.as_str()).collect()\n}\n\n#[cfg(test)]\nmod tests {\n    use baml_base::FileId;\n","sourceCodeStart":511,"sourceCodeEnd":547,"githubUrl":"https://github.com/BoundaryML/baml/blob/bd85ce9dee1463ff04d27efd20531013a4ff46c1/baml_language/crates/baml_compiler_lexer/src/tokens.rs#L511-L547","documentation":"The lexer's token collector converts each span offset to TextSize (u32) via u32::try_from(span.start).expect(\"span.start is too large\"). It panics when a token's start offset exceeds u32::MAX, i.e. a source file larger than 4 GiB (or a span computed from an oversized/incorrect offset). rust-analyzer-style TextSize is deliberately 32-bit.","triggerScenarios":"Lexing a file whose byte size (or the token's absolute start offset) is > 4,294,967,295 bytes; or feeding Token::new a span whose start was computed incorrectly (e.g. accumulated on the wrong base) yielding a value beyond u32::MAX.","commonSituations":"Real-world: pointing the compiler at a generated or concatenated BAML file over 4 GiB; or a bug in offset bookkeeping when splicing multiple source roots so offsets are added twice.","solutions":["Keep source files below 4 GiB; split very large generated files into modules.","If offsets come from concatenating sources, verify the base offset math so spans don't accumulate past the file size.","Handle the conversion explicitly: skip or error on oversized spans with a proper diagnostic instead of panicking.","Check for upstream data (e.g. an external index) reporting byte offsets in units other than bytes."],"exampleFix":"// before\nTextSize::from(u32::try_from(span.start).expect(\"span.start is too large\")),\n// after\nlet start = u32::try_from(span.start).map_err(|_| LexError::SpanTooLarge(span.start))?;","handlingStrategy":"validation","validationCode":"if span.start > u32::MAX as usize {\n    return Err(LexError::SpanTooLarge(span.start));\n}","typeGuard":"fn fits_u32(n: usize) -> Option<u32> { u32::try_from(n).ok() }","tryCatchPattern":"// This is a panic, not a Result; pre-validate offsets. If wrapping a whole lex run:\nlet toks = std::panic::catch_unwind(|| lex(src)).map_err(|_| LexError::SourceTooLarge)?;","preventionTips":["Reject source files larger than 4 GiB before lexing.","Verify offset bookkeeping when concatenating or splicing sources.","Ensure spans are byte offsets (not char indices) from the same buffer."],"tags":["rust","lexer","spans","value-out-of-range","panic"],"backgroundTag":"value-out-of-range","analyzedSha":"bd85ce9dee1463ff04d27efd20531013a4ff46c1","analyzedAt":"2026-09-12T03:38:25.718Z","contentChangedAt":"2026-09-12T03:38:25.718Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}