tonhowtf/omniget · error

pagina HTML vazia

Error message

pagina HTML vazia

What it means

fetch_html rejects the result when htmd's Markdown output is entirely whitespace, returning "pagina HTML vazia" (empty HTML page). This guards against treating a technically-successful conversion of an empty/useless page as a valid paper text.

Solutions

  1. Fall back to fetch_source (LaTeX) extraction when HTML yields empty Markdown
  2. Log/inspect the raw HTML when this fires to detect interstitial pages
  3. Check Content-Type and response length before converting
  4. Consider parsing the PDF as a last resort

Example fix

// before
if md.trim().is_empty() {
    return Err(anyhow!("pagina HTML vazia"));
}
// after
if md.trim().is_empty() {
    tracing::warn!("HTML rendition empty for {}, falling back to source", r.id);
    return fetch_source(client, r).and_then(|_| Ok(String::new())); // wire real fallback
}
Defensive patterns

Strategy: fallback

Validate before calling

// Cheap pre-check: ensure the response body has real content
anyhow::ensure!(html.len() > 512, "HTML page too small ({} bytes); likely an interstitial", html.len());

Try / catch

match fetch_html(&client, &r).await {
    Ok(md) if !md.trim().is_empty() => md,
    _ => fetch_source_text(&client, &r).await?, // fallback to LaTeX extraction
}

Prevention

When it happens

Trigger: fetch -> fetch_html where the downloaded HTML contains no convertible body content: consent/error interstitials, pages that are only scripts/styles, or papers whose HTML rendition is an empty shell.

Common situations: arXiv HTML page that is a redirect/notice page; cloudflare or maintenance interstitial served with 200; content behind JavaScript not present in raw HTML.

Understand the failure class

Background: EmptyResultError / "no results found": when an API or scraper succeeds but returns zero rows — this error's family across 9 libraries.

Related errors


AI-assisted analysis of tonhowtf/omniget@8600b91f42 (2026-09-12). Data as JSON: /api/errors/3aef347113d279cd. Report an issue: GitHub.

Appendix: source

Thrown at src-tauri/omniget-core/src/core/tools/arxiv.rs:1238

async fn fetch_source(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<SourceBundle> {
    let resp = client.get(r.source_url()).send().await?;
    if !resp.status().is_success() {
        return Err(anyhow!("HTTP {}", resp.status()));
    }
    let bytes = resp.bytes().await?;
    extract_source(&bytes)
}

async fn fetch_html(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<String> {
    let resp = client.get(r.html_url()).send().await?;
    if !resp.status().is_success() {
        return Err(anyhow!("HTTP {}", resp.status()));
    }
    let html = resp.text().await?;
    let html = mathml_to_tex(&html);
    let md = htmd::convert(&html).map_err(|e| anyhow!("HTML para Markdown: {}", e))?;
    if md.trim().is_empty() {
        return Err(anyhow!("pagina HTML vazia"));
    }
    Ok(md)
}

#[cfg(test)]
mod tests {
    use super::*;

    #[test]
    fn reconhece_as_formas_de_id() {
        let cases = [
            ("2401.12345", "2401.12345", None),
            ("2401.12345v2", "2401.12345", Some(2)),
            ("arXiv:2401.12345", "2401.12345", None),
            ("arxiv.org/abs/2401.12345", "2401.12345", None),
            ("https://arxiv.org/abs/2401.12345v3", "2401.12345", Some(3)),
            (
                "https://arxiv.org/pdf/2401.12345v1.pdf",

View on GitHub (pinned to 8600b91f42)