tonhowtf/omniget · error

HTML para Markdown

Error message

HTML para Markdown: {}

What it means

After successfully downloading the HTML, fetch_html converts it to Markdown with htmd::convert; if the conversion returns an error it is wrapped as "HTML para Markdown: {e}". This indicates the HTML->Markdown converter itself failed, not the network fetch.

Solutions

  1. Log the underlying htmd error (it is already embedded in the message) to identify the malformed input
  2. Sanitize/repair the HTML before conversion (e.g. run an HTML tidier)
  3. Fall back to fetch_source LaTeX extraction when HTML conversion fails
  4. Pin/upgrade the htmd dependency to a version handling the offending markup

Example fix

// before
let md = htmd::convert(&html).map_err(|e| anyhow!("HTML para Markdown: {}", e))?;
// after
let md = match htmd::convert(&html) {
    Ok(md) => md,
    Err(e) => {
        tracing::warn!("htmd convert failed: {e}; falling back to source extraction");
        return fetch_source(client, r).map(|_| unreachable!()); // or dedicated fallback path
    }
};
Defensive patterns

Strategy: try-catch

Try / catch

let md = htmd::convert(&html)
    .map_err(|e| anyhow!("HTML para Markdown: {}", e))
    .inspect_err(|e| log::warn!("htmd failure on {}: {e}", r.html_url()))?;

Prevention

When it happens

Trigger: fetch -> fetch_html on HTML whose structure crashes htmd::convert: malformed/invalid HTML, extremely large documents, or unexpected encodings after mathml_to_tex preprocessing.

Common situations: Papers with heavily malformed ar5iv markup; converter crate version incompatibilities; pathological or empty-but-nonzero HTML bodies.

Understand the failure class

Background: "Invalid JSON response" and "Failed to parse response" errors: when an API answers 200 but the body isn't the JSON your library expected — this error's family across 28 libraries.

Related errors


AI-assisted analysis of tonhowtf/omniget@8600b91f42 (2026-09-12). Data as JSON: /api/errors/550c591534ca1cf7. Report an issue: GitHub.

Appendix: source

Thrown at src-tauri/omniget-core/src/core/tools/arxiv.rs:1236

}

async fn fetch_source(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<SourceBundle> {
    let resp = client.get(r.source_url()).send().await?;
    if !resp.status().is_success() {
        return Err(anyhow!("HTTP {}", resp.status()));
    }
    let bytes = resp.bytes().await?;
    extract_source(&bytes)
}

async fn fetch_html(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<String> {
    let resp = client.get(r.html_url()).send().await?;
    if !resp.status().is_success() {
        return Err(anyhow!("HTTP {}", resp.status()));
    }
    let html = resp.text().await?;
    let html = mathml_to_tex(&html);
    let md = htmd::convert(&html).map_err(|e| anyhow!("HTML para Markdown: {}", e))?;
    if md.trim().is_empty() {
        return Err(anyhow!("pagina HTML vazia"));
    }
    Ok(md)
}

#[cfg(test)]
mod tests {
    use super::*;

    #[test]
    fn reconhece_as_formas_de_id() {
        let cases = [
            ("2401.12345", "2401.12345", None),
            ("2401.12345v2", "2401.12345", Some(2)),
            ("arXiv:2401.12345", "2401.12345", None),
            ("arxiv.org/abs/2401.12345", "2401.12345", None),
            ("https://arxiv.org/abs/2401.12345v3", "2401.12345", Some(3)),

View on GitHub (pinned to 8600b91f42)