tonhowtf/omniget · error
HTML para Markdown
Error message
HTML para Markdown: {} What it means
After successfully downloading the HTML, fetch_html converts it to Markdown with htmd::convert; if the conversion returns an error it is wrapped as "HTML para Markdown: {e}". This indicates the HTML->Markdown converter itself failed, not the network fetch.
Solutions
- Log the underlying htmd error (it is already embedded in the message) to identify the malformed input
- Sanitize/repair the HTML before conversion (e.g. run an HTML tidier)
- Fall back to fetch_source LaTeX extraction when HTML conversion fails
- Pin/upgrade the htmd dependency to a version handling the offending markup
Example fix
// before
let md = htmd::convert(&html).map_err(|e| anyhow!("HTML para Markdown: {}", e))?;
// after
let md = match htmd::convert(&html) {
Ok(md) => md,
Err(e) => {
tracing::warn!("htmd convert failed: {e}; falling back to source extraction");
return fetch_source(client, r).map(|_| unreachable!()); // or dedicated fallback path
}
}; Defensive patterns
Strategy: try-catch
Try / catch
let md = htmd::convert(&html)
.map_err(|e| anyhow!("HTML para Markdown: {}", e))
.inspect_err(|e| log::warn!("htmd failure on {}: {e}", r.html_url()))?; Prevention
- Sanitize/tidy malformed HTML before conversion
- Keep the htmd crate updated
- Test conversion against a corpus of real ar5iv pages
- Fall back to LaTeX source extraction on conversion failure
When it happens
Trigger: fetch -> fetch_html on HTML whose structure crashes htmd::convert: malformed/invalid HTML, extremely large documents, or unexpected encodings after mathml_to_tex preprocessing.
Common situations: Papers with heavily malformed ar5iv markup; converter crate version incompatibilities; pathological or empty-but-nonzero HTML bodies.
Understand the failure class
Background: "Invalid JSON response" and "Failed to parse response" errors: when an API answers 200 but the body isn't the JSON your library expected — this error's family across 28 libraries.
Related errors
- pagina HTML vazia
- a API do TikTok devolveu resposta vazia — normalmente é a…
- a API do TikTok respondeu HTTP
- a capa não é PNG, JPEG, GIF ou BMP
- a chave local do SponsorBlock está corrompida
AI-assisted analysis of tonhowtf/omniget@8600b91f42 (2026-09-12).
Data as JSON: /api/errors/550c591534ca1cf7.
Report an issue: GitHub.
Appendix: source
Thrown at src-tauri/omniget-core/src/core/tools/arxiv.rs:1236
}
async fn fetch_source(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<SourceBundle> {
let resp = client.get(r.source_url()).send().await?;
if !resp.status().is_success() {
return Err(anyhow!("HTTP {}", resp.status()));
}
let bytes = resp.bytes().await?;
extract_source(&bytes)
}
async fn fetch_html(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<String> {
let resp = client.get(r.html_url()).send().await?;
if !resp.status().is_success() {
return Err(anyhow!("HTTP {}", resp.status()));
}
let html = resp.text().await?;
let html = mathml_to_tex(&html);
let md = htmd::convert(&html).map_err(|e| anyhow!("HTML para Markdown: {}", e))?;
if md.trim().is_empty() {
return Err(anyhow!("pagina HTML vazia"));
}
Ok(md)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn reconhece_as_formas_de_id() {
let cases = [
("2401.12345", "2401.12345", None),
("2401.12345v2", "2401.12345", Some(2)),
("arXiv:2401.12345", "2401.12345", None),
("arxiv.org/abs/2401.12345", "2401.12345", None),
("https://arxiv.org/abs/2401.12345v3", "2401.12345", Some(3)),View on GitHub (pinned to 8600b91f42)