{"record":{"id":"550c591534ca1cf7","repo":"tonhowtf/omniget","slug":"html-para-markdown","errorCode":null,"errorMessage":"HTML para Markdown: {}","messagePattern":"HTML para Markdown: (.+?)","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"src-tauri/omniget-core/src/core/tools/arxiv.rs","lineNumber":1236,"sourceCode":"}\n\nasync fn fetch_source(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<SourceBundle> {\n    let resp = client.get(r.source_url()).send().await?;\n    if !resp.status().is_success() {\n        return Err(anyhow!(\"HTTP {}\", resp.status()));\n    }\n    let bytes = resp.bytes().await?;\n    extract_source(&bytes)\n}\n\nasync fn fetch_html(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<String> {\n    let resp = client.get(r.html_url()).send().await?;\n    if !resp.status().is_success() {\n        return Err(anyhow!(\"HTTP {}\", resp.status()));\n    }\n    let html = resp.text().await?;\n    let html = mathml_to_tex(&html);\n    let md = htmd::convert(&html).map_err(|e| anyhow!(\"HTML para Markdown: {}\", e))?;\n    if md.trim().is_empty() {\n        return Err(anyhow!(\"pagina HTML vazia\"));\n    }\n    Ok(md)\n}\n\n#[cfg(test)]\nmod tests {\n    use super::*;\n\n    #[test]\n    fn reconhece_as_formas_de_id() {\n        let cases = [\n            (\"2401.12345\", \"2401.12345\", None),\n            (\"2401.12345v2\", \"2401.12345\", Some(2)),\n            (\"arXiv:2401.12345\", \"2401.12345\", None),\n            (\"arxiv.org/abs/2401.12345\", \"2401.12345\", None),\n            (\"https://arxiv.org/abs/2401.12345v3\", \"2401.12345\", Some(3)),","sourceCodeStart":1218,"sourceCodeEnd":1254,"githubUrl":"https://github.com/tonhowtf/omniget/blob/8600b91f4246848bac346874daa9e61c1fc5677a/src-tauri/omniget-core/src/core/tools/arxiv.rs#L1218-L1254","documentation":"After successfully downloading the HTML, fetch_html converts it to Markdown with htmd::convert; if the conversion returns an error it is wrapped as \"HTML para Markdown: {e}\". This indicates the HTML->Markdown converter itself failed, not the network fetch.","triggerScenarios":"fetch -> fetch_html on HTML whose structure crashes htmd::convert: malformed/invalid HTML, extremely large documents, or unexpected encodings after mathml_to_tex preprocessing.","commonSituations":"Papers with heavily malformed ar5iv markup; converter crate version incompatibilities; pathological or empty-but-nonzero HTML bodies.","solutions":["Log the underlying htmd error (it is already embedded in the message) to identify the malformed input","Sanitize/repair the HTML before conversion (e.g. run an HTML tidier)","Fall back to fetch_source LaTeX extraction when HTML conversion fails","Pin/upgrade the htmd dependency to a version handling the offending markup"],"exampleFix":"// before\nlet md = htmd::convert(&html).map_err(|e| anyhow!(\"HTML para Markdown: {}\", e))?;\n// after\nlet md = match htmd::convert(&html) {\n    Ok(md) => md,\n    Err(e) => {\n        tracing::warn!(\"htmd convert failed: {e}; falling back to source extraction\");\n        return fetch_source(client, r).map(|_| unreachable!()); // or dedicated fallback path\n    }\n};","handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"let md = htmd::convert(&html)\n    .map_err(|e| anyhow!(\"HTML para Markdown: {}\", e))\n    .inspect_err(|e| log::warn!(\"htmd failure on {}: {e}\", r.html_url()))?;","preventionTips":["Sanitize/tidy malformed HTML before conversion","Keep the htmd crate updated","Test conversion against a corpus of real ar5iv pages","Fall back to LaTeX source extraction on conversion failure"],"tags":["html","markdown","conversion","rust"],"backgroundTag":"invalid-json-response","analyzedSha":"8600b91f4246848bac346874daa9e61c1fc5677a","analyzedAt":"2026-09-12T14:29:19.317Z","contentChangedAt":"2026-09-12T14:29:19.317Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}