siyuan-note/siyuan · error

HTTP

Error message

HTTP %d

What it means

WebFetch treats any HTTP status >= 400 as a failure and returns a compact 'HTTP <code>' error without the response body. It deliberately does not differentiate 4xx/5xx beyond the numeric code, so the caller must interpret the status themselves.

Solutions

  1. Parse the status code from the 'HTTP <code>' message and handle 404/403/429 distinctly
  2. For 403/429, the site likely blocks the client — use an official API or fetch with proper credentials elsewhere
  3. Retry with backoff on 429/5xx; do not retry on 4xx client errors like 404
  4. Verify the URL in a browser; follow redirects manually if the final target differs

Example fix

// before
content, err := WebFetch(url, "markdown")

// after
content, err := WebFetch(url, "markdown")
if err != nil {
    var code int
    if _, perr := fmt.Sscanf(err.Error(), "HTTP %d", &code); perr == nil && code == 429 {
        time.Sleep(backoff) // retry rate-limited fetch
    }
}
Defensive patterns

Strategy: try-catch

Validate before calling

resp, err := http.Head(url)
if err == nil && resp.StatusCode >= 400 {
    return fmt.Errorf("skipping: preflight status %d", resp.StatusCode)
}

Try / catch

content, err := WebFetch(url, format)
if err != nil {
    var code int
    if n, _ := fmt.Sscanf(err.Error(), "HTTP %d", &code); n == 1 {
        switch {
        case code == 429 || code >= 500: // retry with backoff
        case code == 404:            // drop the link permanently
        default:                     // log and skip
        }
    }
}

Prevention

When it happens

Trigger: The remote server returned 404 (page removed), 403 (bot blocking/paywall), 429 (rate limited), 500 (server error), or any other >= 400 status for the fetched URL.

Common situations: Scraping a URL that now redirects to a login page; fetching pages that block non-browser User-Agents; expired links in documents; APIs behind auth being fetched as public pages; hitting a site's rate limiter.

Understand the failure class

Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.

Related errors


AI-assisted analysis of siyuan-note/siyuan@9f775e8a12 (2026-09-19). Data as JSON: /api/errors/3e56f6f56e550d58. Report an issue: GitHub.

Appendix: source

Thrown at kernel/util/webfetch.go:60

	if err != nil || (u.Scheme != "http" && u.Scheme != "https") {
		return "", errors.New("URL must start with http:// or https://")
	}
	if u.Host == "" {
		return "", errors.New("URL has no host")
	}

	if err := CheckHostSSRF(u.Hostname()); err != nil {
		return "", err
	}

	resp, err := ssrfSafeClient.Get(rawURL)
	if err != nil {
		return "", errors.New("fetch failed: " + err.Error())
	}
	defer resp.Body.Close()

	if resp.StatusCode >= 400 {
		return "", fmt.Errorf("HTTP %d", resp.StatusCode)
	}

	contentType := resp.Header.Get("Content-Type")
	maxReadBytes := int64(maxWebFetchBytes)
	if !strings.HasPrefix(contentType, "text/html") && !strings.HasPrefix(contentType, "text/plain") {
		maxReadBytes = maxWebFetchFileBytes
	}
	if resp.ContentLength > maxReadBytes {
		return "", errors.New("response too large")
	}

	body, err := io.ReadAll(io.LimitReader(resp.Body, maxReadBytes))
	if err != nil {
		return "", errors.New("read body failed: " + err.Error())
	}

	if !strings.HasPrefix(contentType, "text/html") && !strings.HasPrefix(contentType, "text/plain") {
		importDir := filepath.Join(TempDir, "import")

View on GitHub (pinned to 9f775e8a12)