siyuan-note/siyuan · error
HTML to text panicked
Error message
HTML to text panicked: %v
What it means
safeHTML2Text wraps lute's HTML2Text with a recover() so a panic inside the plain-text extraction engine is converted into the error 'HTML to text panicked: <value>' instead of crashing the kernel. Like its Markdown counterpart, it indicates an unguarded failure inside Lute on the given HTML input.
Solutions
- Treat the returned error and use the original HTML string (WebFetch returns htmlStr when the converted result is empty)
- Sanitize the HTML before fetching, or fetch a cleaner representation of the same content
- Check for a Lute engine update if the same page previously extracted fine
- Report the panicking input upstream to the Lute project with the HTML that triggers it
Example fix
// before
text, err := WebFetch(url, "text")
if err != nil {
return err // aborts even though raw HTML is available
}
// after
text, err := WebFetch(url, "text")
if err != nil && strings.Contains(err.Error(), "panicked") {
text = fallbackRawHTML // keep going with unconverted content
} Defensive patterns
Strategy: try-catch
Try / catch
text, err := WebFetch(url, "text")
if err != nil && strings.Contains(err.Error(), "panicked") {
// engine panic on this HTML; degrade to raw HTML or skip the source
} Prevention
- Handle the returned error rather than letting unusual input break your pipeline
- Retain the raw HTML fallback — WebFetch already returns it when conversion yields empty
- Keep the Lute engine current to pick up panic fixes
- Sanitize or reject pathological HTML sources instead of repeatedly fetching them
When it happens
Trigger: engine.HTML2Text panics while processing the fetched HTML — e.g. malformed structure that triggers an index-out-of-range or nil dereference inside the engine.
Common situations: Pathological HTML from scraped pages; engine regressions after a Lute update; documents with broken encoding that the text extractor mishandles.
Related errors
- HTML to Markdown panicked
- task executor panicked
- 347
- agent session has an uncommitted turn
- attribute view rich text normalization did not converge
AI-assisted analysis of siyuan-note/siyuan@9f775e8a12 (2026-09-19).
Data as JSON: /api/errors/99f259849840d4e7.
Report an issue: GitHub.
Appendix: source
Thrown at kernel/util/webfetch.go:134
}
return truncateRunes(result, maxWebFetchChars), nil
}
func safeHTML2Markdown(engine *lute.Lute, htmlStr string) (result string, err error) {
defer func() {
if r := recover(); r != nil {
err = fmt.Errorf("HTML to Markdown panicked: %v", r)
}
}()
result, err = engine.HTML2Markdown(htmlStr)
return
}
func safeHTML2Text(engine *lute.Lute, htmlStr string) (result string, err error) {
defer func() {
if r := recover(); r != nil {
err = fmt.Errorf("HTML to text panicked: %v", r)
}
}()
result = engine.HTML2Text(htmlStr)
return
}
func truncateRunes(s string, maxChars int) string {
runes := []rune(s)
if len(runes) <= maxChars {
return s
}
return string(runes[:maxChars]) + "\n\n...content truncated, total length " + fmt.Sprintf("%d", len(runes)) + " characters..."
}
func extractFilename(rawURL, contentType string) string {
u, err := url.Parse(rawURL)
if err != nil {
return gulu.Rand.String(7) + extByContentType(contentType)View on GitHub (pinned to 9f775e8a12)