{"record":{"id":"99f259849840d4e7","repo":"siyuan-note/siyuan","slug":"html-to-text-panicked-v","errorCode":null,"errorMessage":"HTML to text panicked: %v","messagePattern":"HTML to text panicked: (.+?)","errorType":"panic","errorClass":null,"httpStatus":null,"severity":"error","filePath":"kernel/util/webfetch.go","lineNumber":134,"sourceCode":"\t}\n\n\treturn truncateRunes(result, maxWebFetchChars), nil\n}\n\nfunc safeHTML2Markdown(engine *lute.Lute, htmlStr string) (result string, err error) {\n\tdefer func() {\n\t\tif r := recover(); r != nil {\n\t\t\terr = fmt.Errorf(\"HTML to Markdown panicked: %v\", r)\n\t\t}\n\t}()\n\tresult, err = engine.HTML2Markdown(htmlStr)\n\treturn\n}\n\nfunc safeHTML2Text(engine *lute.Lute, htmlStr string) (result string, err error) {\n\tdefer func() {\n\t\tif r := recover(); r != nil {\n\t\t\terr = fmt.Errorf(\"HTML to text panicked: %v\", r)\n\t\t}\n\t}()\n\tresult = engine.HTML2Text(htmlStr)\n\treturn\n}\n\nfunc truncateRunes(s string, maxChars int) string {\n\trunes := []rune(s)\n\tif len(runes) <= maxChars {\n\t\treturn s\n\t}\n\treturn string(runes[:maxChars]) + \"\\n\\n...content truncated, total length \" + fmt.Sprintf(\"%d\", len(runes)) + \" characters...\"\n}\n\nfunc extractFilename(rawURL, contentType string) string {\n\tu, err := url.Parse(rawURL)\n\tif err != nil {\n\t\treturn gulu.Rand.String(7) + extByContentType(contentType)","sourceCodeStart":116,"sourceCodeEnd":152,"githubUrl":"https://github.com/siyuan-note/siyuan/blob/9f775e8a12daef8255556097396f9b2739078892/kernel/util/webfetch.go#L116-L152","documentation":"safeHTML2Text wraps lute's HTML2Text with a recover() so a panic inside the plain-text extraction engine is converted into the error 'HTML to text panicked: <value>' instead of crashing the kernel. Like its Markdown counterpart, it indicates an unguarded failure inside Lute on the given HTML input.","triggerScenarios":"engine.HTML2Text panics while processing the fetched HTML — e.g. malformed structure that triggers an index-out-of-range or nil dereference inside the engine.","commonSituations":"Pathological HTML from scraped pages; engine regressions after a Lute update; documents with broken encoding that the text extractor mishandles.","solutions":["Treat the returned error and use the original HTML string (WebFetch returns htmlStr when the converted result is empty)","Sanitize the HTML before fetching, or fetch a cleaner representation of the same content","Check for a Lute engine update if the same page previously extracted fine","Report the panicking input upstream to the Lute project with the HTML that triggers it"],"exampleFix":"// before\ntext, err := WebFetch(url, \"text\")\nif err != nil {\n    return err // aborts even though raw HTML is available\n}\n\n// after\ntext, err := WebFetch(url, \"text\")\nif err != nil && strings.Contains(err.Error(), \"panicked\") {\n    text = fallbackRawHTML // keep going with unconverted content\n}","handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"text, err := WebFetch(url, \"text\")\nif err != nil && strings.Contains(err.Error(), \"panicked\") {\n    // engine panic on this HTML; degrade to raw HTML or skip the source\n}","preventionTips":["Handle the returned error rather than letting unusual input break your pipeline","Retain the raw HTML fallback — WebFetch already returns it when conversion yields empty","Keep the Lute engine current to pick up panic fixes","Sanitize or reject pathological HTML sources instead of repeatedly fetching them"],"tags":["panic","recovery","lute","text-extraction"],"backgroundTag":"recovered-panic","analyzedSha":"9f775e8a12daef8255556097396f9b2739078892","analyzedAt":"2026-09-19T03:17:15.984Z","contentChangedAt":"2026-09-19T03:17:15.984Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}