tonhowtf/omniget · error
pagina HTML vazia
Error message
pagina HTML vazia
What it means
fetch_html rejects the result when htmd's Markdown output is entirely whitespace, returning "pagina HTML vazia" (empty HTML page). This guards against treating a technically-successful conversion of an empty/useless page as a valid paper text.
Solutions
- Fall back to fetch_source (LaTeX) extraction when HTML yields empty Markdown
- Log/inspect the raw HTML when this fires to detect interstitial pages
- Check Content-Type and response length before converting
- Consider parsing the PDF as a last resort
Example fix
// before
if md.trim().is_empty() {
return Err(anyhow!("pagina HTML vazia"));
}
// after
if md.trim().is_empty() {
tracing::warn!("HTML rendition empty for {}, falling back to source", r.id);
return fetch_source(client, r).and_then(|_| Ok(String::new())); // wire real fallback
} Defensive patterns
Strategy: fallback
Validate before calling
// Cheap pre-check: ensure the response body has real content
anyhow::ensure!(html.len() > 512, "HTML page too small ({} bytes); likely an interstitial", html.len()); Try / catch
match fetch_html(&client, &r).await {
Ok(md) if !md.trim().is_empty() => md,
_ => fetch_source_text(&client, &r).await?, // fallback to LaTeX extraction
} Prevention
- Check Content-Type is text/html before parsing
- Detect interstitial/notice pages by markers before conversion
- Always provide a secondary extraction path
- Log raw HTML snippets when conversion output is empty
When it happens
Trigger: fetch -> fetch_html where the downloaded HTML contains no convertible body content: consent/error interstitials, pages that are only scripts/styles, or papers whose HTML rendition is an empty shell.
Common situations: arXiv HTML page that is a redirect/notice page; cloudflare or maintenance interstitial served with 200; content behind JavaScript not present in raw HTML.
Understand the failure class
Background: EmptyResultError / "no results found": when an API or scraper succeeds but returns zero rows — this error's family across 9 libraries.
Related errors
- HTML para Markdown
- a API do TikTok devolveu resposta vazia — normalmente é a…
- a API do TikTok respondeu HTTP
- a capa não é PNG, JPEG, GIF ou BMP
- a chave local do SponsorBlock está corrompida
AI-assisted analysis of tonhowtf/omniget@8600b91f42 (2026-09-12).
Data as JSON: /api/errors/3aef347113d279cd.
Report an issue: GitHub.
Appendix: source
Thrown at src-tauri/omniget-core/src/core/tools/arxiv.rs:1238
async fn fetch_source(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<SourceBundle> {
let resp = client.get(r.source_url()).send().await?;
if !resp.status().is_success() {
return Err(anyhow!("HTTP {}", resp.status()));
}
let bytes = resp.bytes().await?;
extract_source(&bytes)
}
async fn fetch_html(client: &reqwest::Client, r: &ArxivRef) -> anyhow::Result<String> {
let resp = client.get(r.html_url()).send().await?;
if !resp.status().is_success() {
return Err(anyhow!("HTTP {}", resp.status()));
}
let html = resp.text().await?;
let html = mathml_to_tex(&html);
let md = htmd::convert(&html).map_err(|e| anyhow!("HTML para Markdown: {}", e))?;
if md.trim().is_empty() {
return Err(anyhow!("pagina HTML vazia"));
}
Ok(md)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn reconhece_as_formas_de_id() {
let cases = [
("2401.12345", "2401.12345", None),
("2401.12345v2", "2401.12345", Some(2)),
("arXiv:2401.12345", "2401.12345", None),
("arxiv.org/abs/2401.12345", "2401.12345", None),
("https://arxiv.org/abs/2401.12345v3", "2401.12345", Some(3)),
(
"https://arxiv.org/pdf/2401.12345v1.pdf",View on GitHub (pinned to 8600b91f42)