{"record":{"id":"e23dc87a05cc9819","repo":"siyuan-note/siyuan","slug":"344","errorCode":"344","errorMessage":"Markdown [%s] is not valid UTF-8","messagePattern":"Markdown \\[(.+?)\\] is not valid UTF-8","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"kernel/model/import_obsidian.go","lineNumber":862,"sourceCode":"}\n\nfunc analyzeObsidianDocuments(ctx context.Context, vault *obsidianVaultContext, progress func(int, string)) error {\n\tsourceCount := countObsidianSourceDocs(vault.Docs)\n\tprocessed := 0\n\tfor _, doc := range vault.Docs {\n\t\tif doc.Synthetic {\n\t\t\tcontinue\n\t\t}\n\t\tif err := ctx.Err(); err != nil {\n\t\t\treturn err\n\t\t}\n\t\tdata, err := readStableObsidianFile(doc.Source)\n\t\tif err != nil {\n\t\t\treturn newObsidianReadUserError(doc.Source, err)\n\t\t}\n\t\tif !utf8.Valid(data) {\n\t\t\treturn newObsidianUserError(344, doc.Source.RelPath,\n\t\t\t\tfmt.Errorf(\"Markdown [%s] is not valid UTF-8\", doc.Source.RelPath))\n\t\t}\n\t\tscan := scanObsidianSource(data)\n\t\tfor _, blockID := range scan.BlockIDs {\n\t\t\tif scan.Duplicates[blockID] {\n\t\t\t\tdoc.DuplicateBlocks[blockID] = true\n\t\t\t}\n\t\t\tif doc.BlockIDs[blockID] == \"\" {\n\t\t\t\tdoc.BlockIDs[blockID] = ast.NewNodeID()\n\t\t\t}\n\t\t}\n\t\ttree, _, _, _ := parseStdMd(data)\n\t\tif tree == nil {\n\t\t\treturn newObsidianUserError(347, doc.Source.RelPath,\n\t\t\t\tfmt.Errorf(\"parse Markdown [%s] failed\", doc.Source.RelPath))\n\t\t}\n\t\tbuildObsidianHeadingIndex(doc, tree)\n\t\tvault.Analysis.WikiLinkCount += countObsidianNonEmbedTokens(scan.Wikis)\n\t\tvault.Analysis.EmbedCount += countObsidianEmbedTokens(scan.Wikis)","sourceCodeStart":844,"sourceCodeEnd":880,"githubUrl":"https://github.com/siyuan-note/siyuan/blob/9f775e8a12daef8255556097396f9b2739078892/kernel/model/import_obsidian.go#L844-L880","documentation":"A Markdown file's bytes are not valid UTF-8, so the importer refuses to process it (error code 344). SiYuan stores note content as UTF-8 text; importing non-UTF-8 bytes would corrupt block content, so the offending file is reported with its vault-relative path and the import of that document fails.","triggerScenarios":"readStableObsidianFile returns bytes for a `.md` source and utf8.Valid(data) is false — the file contains legacy encodings (GBK, Latin-1, UTF-16) or binary junk.","commonSituations":"Old notes created with Windows editors in a legacy codepage; files converted from Evernote/Word without re-encoding; UTF-16 files produced by PowerShell redirection; a `.md` file that is actually binary.","solutions":["Re-encode the reported file to UTF-8 (iconv -f GBK -t UTF-8 file.md, or VS Code 'Save with Encoding → UTF-8')","Detect the actual encoding with file/chardet before converting","If the file is UTF-16 with BOM, convert to UTF-8 explicitly","Exclude or delete the non-text file if it is not a real note"],"exampleFix":"$ iconv -f GB18030 -t UTF-8 notes/old.md > notes/old.utf8.md\n$ mv notes/old.utf8.md notes/old.md","handlingStrategy":"validation","validationCode":"const fs = require('fs');\nconst buf = fs.readFileSync(mdFile);\nconst decoded = new TextDecoder('utf-8', { fatal: true }).decode(buf); // throws on invalid UTF-8","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Convert legacy-encoded notes (GBK, Latin-1, UTF-16) to UTF-8 before import","Run a bulk check (e.g. isutf8 from moreutils) across the vault","Avoid creating .md files via PowerShell redirection (produces UTF-16)"],"tags":["obsidian-import","encoding","utf-8","markdown"],"backgroundTag":"invalid-argument-format","analyzedSha":"9f775e8a12daef8255556097396f9b2739078892","analyzedAt":"2026-09-19T03:17:15.984Z","contentChangedAt":"2026-09-19T03:17:15.984Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}