{"record":{"id":"0ce7350bc7ad2ee5","repo":"BoundaryML/baml","slug":"span-end-is-too-large","errorCode":null,"errorMessage":"span.end is too large","messagePattern":"span\\.end is too large","errorType":"panic","errorClass":null,"httpStatus":null,"severity":"error","filePath":"baml_language/crates/baml_compiler_lexer/src/tokens.rs","lineNumber":530,"sourceCode":"/// This tokenizes the entire input including whitespace and comments,\n/// allowing perfect source reconstruction.\npub fn lex_lossless(input: &str, file_id: FileId) -> Vec<Token> {\n    let mut tokens = Vec::new();\n    let mut lexer = TokenKind::lexer(input);\n\n    while let Some(result) = lexer.next() {\n        let kind = result.unwrap_or(TokenKind::Error);\n        let span = lexer.span();\n        let text = lexer.slice().to_string();\n\n        tokens.push(Token {\n            kind,\n            text,\n            span: Span::new(\n                file_id,\n                TextRange::new(\n                    TextSize::from(u32::try_from(span.start).expect(\"span.start is too large\")),\n                    TextSize::from(u32::try_from(span.end).expect(\"span.end is too large\")),\n                ),\n            ),\n        });\n    }\n\n    tokens\n}\n\n/// Reconstruct source from tokens (for testing losslessness).\npub fn reconstruct_source(tokens: &[Token]) -> String {\n    tokens.iter().map(|t| t.text.as_str()).collect()\n}\n\n#[cfg(test)]\nmod tests {\n    use baml_base::FileId;\n\n    use super::*;","sourceCodeStart":512,"sourceCodeEnd":548,"githubUrl":"https://github.com/BoundaryML/baml/blob/bd85ce9dee1463ff04d27efd20531013a4ff46c1/baml_language/crates/baml_compiler_lexer/src/tokens.rs#L512-L548","documentation":"Same 32-bit TextSize conversion as the start offset, but for span.end: u32::try_from(span.end).expect(\"span.end is too large\") panics when a token's end offset exceeds u32::MAX. This typically means the token spans a file larger than 4 GiB or the end offset was computed from bad bookkeeping.","triggerScenarios":"Lexing input whose end offset (file length or token end) is > 4,294,967,295 bytes; or a span constructed with an end value not derived from the actual byte length.","commonSituations":"Giant generated BAML sources over 4 GiB; duplicated offset accumulation across concatenated source roots; external tooling supplying character counts instead of byte offsets.","solutions":["Keep files under 4 GiB or split them into smaller modules.","Validate that span.end equals the byte offset after the token and derives from the same buffer as span.start.","Convert explicitly with error handling instead of expect so oversized spans produce a diagnostic rather than a panic.","Audit any code adding lengths to offsets for double-accumulation bugs."],"exampleFix":"// before\nTextSize::from(u32::try_from(span.end).expect(\"span.end is too large\")),\n// after\nlet end = u32::try_from(span.end).map_err(|_| LexError::SpanTooLarge(span.end))?;","handlingStrategy":"validation","validationCode":"if span.end > u32::MAX as usize || span.end > source_len {\n    return Err(LexError::SpanTooLarge(span.end));\n}","typeGuard":"fn fits_u32(n: usize) -> Option<u32> { u32::try_from(n).ok() }","tryCatchPattern":"// Panic-based; validate span.end against u32::MAX and the buffer length before constructing Token.","preventionTips":["Enforce a max input size (< 4 GiB) at the project-loading layer.","Check that span.end is derived from the same byte buffer and never accumulates twice.","Convert spans with checked conversion and emit a diagnostic instead of panicking."],"tags":["rust","lexer","spans","value-out-of-range","panic"],"backgroundTag":"value-out-of-range","analyzedSha":"bd85ce9dee1463ff04d27efd20531013a4ff46c1","analyzedAt":"2026-09-12T03:38:25.718Z","contentChangedAt":"2026-09-12T03:38:25.718Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}