{"record":{"id":"44e6463079423a08","repo":"siyuan-note/siyuan","slug":"html-to-markdown-panicked-v","errorCode":null,"errorMessage":"HTML to Markdown panicked: %v","messagePattern":"HTML to Markdown panicked: (.+?)","errorType":"panic","errorClass":null,"httpStatus":null,"severity":"error","filePath":"kernel/util/webfetch.go","lineNumber":124,"sourceCode":"\tdefault: // markdown\n\t\tmd, mdErr := safeHTML2Markdown(engine, htmlStr)\n\t\tif mdErr != nil {\n\t\t\treturn \"\", errors.New(\"HTML to Markdown conversion failed: \" + mdErr.Error())\n\t\t}\n\t\tresult = md\n\t}\n\n\tif result == \"\" {\n\t\treturn htmlStr, nil\n\t}\n\n\treturn truncateRunes(result, maxWebFetchChars), nil\n}\n\nfunc safeHTML2Markdown(engine *lute.Lute, htmlStr string) (result string, err error) {\n\tdefer func() {\n\t\tif r := recover(); r != nil {\n\t\t\terr = fmt.Errorf(\"HTML to Markdown panicked: %v\", r)\n\t\t}\n\t}()\n\tresult, err = engine.HTML2Markdown(htmlStr)\n\treturn\n}\n\nfunc safeHTML2Text(engine *lute.Lute, htmlStr string) (result string, err error) {\n\tdefer func() {\n\t\tif r := recover(); r != nil {\n\t\t\terr = fmt.Errorf(\"HTML to text panicked: %v\", r)\n\t\t}\n\t}()\n\tresult = engine.HTML2Text(htmlStr)\n\treturn\n}\n\nfunc truncateRunes(s string, maxChars int) string {\n\trunes := []rune(s)","sourceCodeStart":106,"sourceCodeEnd":142,"githubUrl":"https://github.com/siyuan-note/siyuan/blob/9f775e8a12daef8255556097396f9b2739078892/kernel/util/webfetch.go#L106-L142","documentation":"safeHTML2Markdown wraps lute's HTML2Markdown with a recover() so a panic inside the conversion engine (index out of range, nil deref on malformed input) is converted into this error instead of crashing the kernel. It signals an engine bug or input that trips an unguarded path in Lute.","triggerScenarios":"HTML2Markdown panics on a specific HTML payload — typically deeply nested or malformed markup that hits an unguarded code path in the Lute engine.","commonSituations":"Fetching attacker-controlled or user-generated HTML with pathological nesting; Lute regression where a previously fine document panics after an engine update; fuzz-like input from scraped web pages.","solutions":["The kernel already survives the panic — handle the returned error and fall back to raw HTML (WebFetch itself returns htmlStr when result is empty)","Reproduce with the same HTML and report/trim the offending markup","Retry with format=\"text\" to bypass the Markdown converter path","Check for a Lute update if a previously-working URL started panicking"],"exampleFix":"// before\nmd, err := WebFetch(url, \"markdown\")\n\n// after\nmd, err := WebFetch(url, \"markdown\")\nif err != nil && strings.Contains(err.Error(), \"panicked\") {\n    md, _ = WebFetch(url, \"text\") // avoid the panicking converter path\n}","handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"md, err := WebFetch(url, \"markdown\")\nif err != nil && strings.Contains(err.Error(), \"panicked\") {\n    // engine panicked on this input; fall back to text or raw HTML\n    md, _ = WebFetch(url, \"text\")\n}","preventionTips":["Treat panics as engine-input bugs: capture the URL and offending HTML for upstream reports","Always handle the error — the panic is already converted, so a plain error check suffices","Keep Lute updated; panics on specific markup are often fixed upstream","Favor text extraction for untrusted, heavily malformed HTML"],"tags":["panic","recovery","lute","conversion"],"backgroundTag":"recovered-panic","analyzedSha":"9f775e8a12daef8255556097396f9b2739078892","analyzedAt":"2026-09-19T03:17:15.984Z","contentChangedAt":"2026-09-19T03:17:15.984Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}