{"record":{"id":"065f52931c69ff16","repo":"siyuan-note/siyuan","slug":"url-has-no-host-065f52","errorCode":null,"errorMessage":"URL has no host","messagePattern":"URL has no host","errorType":"validation","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"kernel/util/webfetch.go","lineNumber":46,"sourceCode":"\t\"strings\"\n\n\t\"github.com/88250/gulu\"\n\t\"github.com/88250/lute\"\n)\n\nconst (\n\tmaxWebFetchBytes     = 5 * 1024 * 1024  // text/html, text/plain\n\tmaxWebFetchFileBytes = 10 * 1024 * 1024 // file/image download\n\tmaxWebFetchChars     = 50000\n)\n\nfunc WebFetch(rawURL, format string) (string, error) {\n\tu, err := url.Parse(rawURL)\n\tif err != nil || (u.Scheme != \"http\" && u.Scheme != \"https\") {\n\t\treturn \"\", errors.New(\"URL must start with http:// or https://\")\n\t}\n\tif u.Host == \"\" {\n\t\treturn \"\", errors.New(\"URL has no host\")\n\t}\n\n\tif err := CheckHostSSRF(u.Hostname()); err != nil {\n\t\treturn \"\", err\n\t}\n\n\tresp, err := ssrfSafeClient.Get(rawURL)\n\tif err != nil {\n\t\treturn \"\", errors.New(\"fetch failed: \" + err.Error())\n\t}\n\tdefer resp.Body.Close()\n\n\tif resp.StatusCode >= 400 {\n\t\treturn \"\", fmt.Errorf(\"HTTP %d\", resp.StatusCode)\n\t}\n\n\tcontentType := resp.Header.Get(\"Content-Type\")\n\tmaxReadBytes := int64(maxWebFetchBytes)","sourceCodeStart":28,"sourceCodeEnd":64,"githubUrl":"https://github.com/siyuan-note/siyuan/blob/9f775e8a12daef8255556097396f9b2739078892/kernel/util/webfetch.go#L28-L64","documentation":"WebFetch parses the raw URL with net/url.Parse and requires an http(s) scheme; this error is returned when the parsed URL has an empty Host component (e.g. 'https://' or 'file:///tmp/x' style inputs that still pass the scheme check). It is a fail-fast input validation guard before any network activity or SSRF checks run.","triggerScenarios":"Calling WebFetch with 'https://', 'http://', a scheme-relative URL like '//example.com/x' (no scheme so rejected earlier — but 'https:' alone parses with empty Host), or a malformed URL whose host part is dropped by url.Parse.","commonSituations":"Building the URL by string concatenation where the host variable is empty; reading a URL from config or a document where the host was stripped; passing a path-only or protocol-relative link from user input.","solutions":["Check the URL includes a host part before calling WebFetch: u, _ := url.Parse(raw); require u.Host != \"\"","Ensure the scheme is included ('https://example.com', not 'example.com' or '//example.com')","Trim whitespace and hidden characters from the input URL; log the exact string passed in","If the URL comes from user content, normalize/repair it (prepend https:// when scheme is missing) before calling"],"exampleFix":"// before\nWebFetch(\"example.com/page\", \"markdown\") // rejected: no scheme/host\n\n// after\nWebFetch(\"https://example.com/page\", \"markdown\")","handlingStrategy":"validation","validationCode":"u, err := url.Parse(raw)\nif err != nil || (u.Scheme != \"http\" && u.Scheme != \"https\") || u.Host == \"\" {\n    return errors.New(\"skipping: URL must be an absolute http(s) URL with a host\")\n}","typeGuard":"func isFetchableURL(raw string) bool {\n    u, err := url.Parse(strings.TrimSpace(raw))\n    return err == nil && (u.Scheme == \"http\" || u.Scheme == \"https\") && u.Host != \"\"\n}","tryCatchPattern":null,"preventionTips":["Always store and pass absolute URLs including the scheme","Normalize user-supplied links (trim spaces, prepend https:// when scheme missing) before fetching","Add a pre-call url.Parse check in any wrapper around WebFetch"],"tags":["url","validation","input"],"backgroundTag":"invalid-url","analyzedSha":"9f775e8a12daef8255556097396f9b2739078892","analyzedAt":"2026-09-19T03:17:15.984Z","contentChangedAt":"2026-09-19T03:17:15.984Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}