siyuan-note/siyuan · error
HTTP
Error message
HTTP %d
What it means
WebFetch treats any HTTP status >= 400 as a failure and returns a compact 'HTTP <code>' error without the response body. It deliberately does not differentiate 4xx/5xx beyond the numeric code, so the caller must interpret the status themselves.
Solutions
- Parse the status code from the 'HTTP <code>' message and handle 404/403/429 distinctly
- For 403/429, the site likely blocks the client — use an official API or fetch with proper credentials elsewhere
- Retry with backoff on 429/5xx; do not retry on 4xx client errors like 404
- Verify the URL in a browser; follow redirects manually if the final target differs
Example fix
// before
content, err := WebFetch(url, "markdown")
// after
content, err := WebFetch(url, "markdown")
if err != nil {
var code int
if _, perr := fmt.Sscanf(err.Error(), "HTTP %d", &code); perr == nil && code == 429 {
time.Sleep(backoff) // retry rate-limited fetch
}
} Defensive patterns
Strategy: try-catch
Validate before calling
resp, err := http.Head(url)
if err == nil && resp.StatusCode >= 400 {
return fmt.Errorf("skipping: preflight status %d", resp.StatusCode)
} Try / catch
content, err := WebFetch(url, format)
if err != nil {
var code int
if n, _ := fmt.Sscanf(err.Error(), "HTTP %d", &code); n == 1 {
switch {
case code == 429 || code >= 500: // retry with backoff
case code == 404: // drop the link permanently
default: // log and skip
}
}
} Prevention
- HEAD-probe links before batch fetching
- Expect 403/429 from bot-protected sites and use official APIs where available
- Classify retryable (429, 5xx) vs permanent (4xx) statuses by parsing the code
- Keep link lists fresh; prune 404s
When it happens
Trigger: The remote server returned 404 (page removed), 403 (bot blocking/paywall), 429 (rate limited), 500 (server error), or any other >= 400 status for the fetched URL.
Common situations: Scraping a URL that now redirects to a login page; fetching pages that block non-browser User-Agents; expired links in documents; APIs behind auth being fetched as public pages; hitting a site's rate limiter.
Understand the failure class
Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.
Related errors
- unexpected status code
- authentication probe returned HTTP " + response.status
- boot progress request returned HTTP " + response.status
- discover OIDC provider failed
- download custom emoji failed
AI-assisted analysis of siyuan-note/siyuan@9f775e8a12 (2026-09-19).
Data as JSON: /api/errors/3e56f6f56e550d58.
Report an issue: GitHub.
Appendix: source
Thrown at kernel/util/webfetch.go:60
if err != nil || (u.Scheme != "http" && u.Scheme != "https") {
return "", errors.New("URL must start with http:// or https://")
}
if u.Host == "" {
return "", errors.New("URL has no host")
}
if err := CheckHostSSRF(u.Hostname()); err != nil {
return "", err
}
resp, err := ssrfSafeClient.Get(rawURL)
if err != nil {
return "", errors.New("fetch failed: " + err.Error())
}
defer resp.Body.Close()
if resp.StatusCode >= 400 {
return "", fmt.Errorf("HTTP %d", resp.StatusCode)
}
contentType := resp.Header.Get("Content-Type")
maxReadBytes := int64(maxWebFetchBytes)
if !strings.HasPrefix(contentType, "text/html") && !strings.HasPrefix(contentType, "text/plain") {
maxReadBytes = maxWebFetchFileBytes
}
if resp.ContentLength > maxReadBytes {
return "", errors.New("response too large")
}
body, err := io.ReadAll(io.LimitReader(resp.Body, maxReadBytes))
if err != nil {
return "", errors.New("read body failed: " + err.Error())
}
if !strings.HasPrefix(contentType, "text/html") && !strings.HasPrefix(contentType, "text/plain") {
importDir := filepath.Join(TempDir, "import")View on GitHub (pinned to 9f775e8a12)