{"record":{"id":"6aa571d55550b4f0","repo":"BoundaryML/baml","slug":"http-body-is-not-utf-8","errorCode":null,"errorMessage":"HTTP body is not UTF-8: {}","messagePattern":"HTTP body is not UTF-8: (.+?)","errorType":"validation","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"engine/baml-lib/baml-types/src/tracing/events.rs","lineNumber":290,"sourceCode":"                Err(_) => format!(\"[{} bytes]\", self.raw.len()),\n            }\n        };\n\n        f.debug_struct(\"HTTPBody\").field(\"raw\", &preview).finish()\n    }\n}\n\nimpl HTTPBody {\n    pub fn new(body: Vec<u8>) -> Self {\n        Self { raw: body }\n    }\n\n    pub fn raw(&self) -> &[u8] {\n        &self.raw\n    }\n\n    pub fn text(&self) -> anyhow::Result<&str> {\n        std::str::from_utf8(&self.raw).map_err(|e| anyhow::anyhow!(\"HTTP body is not UTF-8: {}\", e))\n    }\n\n    pub fn json(&self) -> anyhow::Result<serde_json::Value> {\n        serde_json::from_str(self.text()?)\n            .map_err(|e| anyhow::anyhow!(\"HTTP body is not JSON: {}\", e))\n    }\n\n    /// Returns the HTTP body as a [`serde_json::Value`].\n    ///\n    /// If the body is not UTF-8 or JSON, it is returned as an array of bytes.\n    /// Used as input for [`serde_json::to_string_pretty`].\n    pub fn as_serde_value(&self) -> serde_json::Value {\n        self.json()\n            .or_else(|_e| self.text().map(|s| serde_json::Value::String(s.into())))\n            .unwrap_or_else(|_e| {\n                serde_json::Value::Array(\n                    self.raw()\n                        .iter()","sourceCodeStart":272,"sourceCodeEnd":308,"githubUrl":"https://github.com/BoundaryML/baml/blob/bd85ce9dee1463ff04d27efd20531013a4ff46c1/engine/baml-lib/baml-types/src/tracing/events.rs#L272-L308","documentation":"HttpBody::text decodes the recorded raw HTTP body bytes as UTF-8. If the bytes are not valid UTF-8 (binary content, compressed bodies, other encodings), the underlying from_utf8 error is wrapped as \"HTTP body is not UTF-8: {e}\". The library only exposes body text when it is valid UTF-8.","triggerScenarios":"Calling HttpBody::text() (directly or via json()) on a body containing binary data — e.g. gzip-compressed responses, image bytes, or non-UTF-8 character sets like latin-1.","commonSituations":"Tracing/event inspection of LLM HTTP responses where a provider returned compressed or binary payloads; proxies returning binary error pages; responses with charset other than UTF-8.","solutions":["Use raw() to get the bytes and decode with the correct encoding (e.g. inflate gzip, convert latin-1).","Check the Content-Encoding/Content-Type headers before calling text().","Call json() only for bodies known to be UTF-8 JSON; otherwise use the byte-array fallback path."],"exampleFix":"// before\nlet s = body.text()?;\n// after\nlet bytes = body.raw();\nlet s = String::from_utf8_lossy(bytes); // or gunzip first if Content-Encoding: gzip","handlingStrategy":"fallback","validationCode":"let is_utf8 = std::str::from_utf8(body.raw()).is_ok();","typeGuard":"fn is_text_body(b: &HttpBody) -> bool { std::str::from_utf8(b.raw()).is_ok() }","tryCatchPattern":"let text = match body.text() { Ok(t) => t.to_string(), Err(_) => String::from_utf8_lossy(body.raw()).into_owned() };","preventionTips":["Check Content-Encoding and gunzip before decoding","Use from_utf8_lossy for logging/inspection paths","Don't assume provider bodies are UTF-8; keep raw bytes available"],"tags":["http","encoding","utf-8"],"backgroundTag":"invalid-utf8-body","analyzedSha":"bd85ce9dee1463ff04d27efd20531013a4ff46c1","analyzedAt":"2026-09-12T03:38:25.718Z","contentChangedAt":"2026-09-12T03:38:25.718Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}