{"record":{"id":"38b51a23578533dd","repo":"janhq/jan","slug":"parseerror","errorCode":"ParseError","errorMessage":"Failed to parse document: {0}","messagePattern":"Failed to parse document: (.+?)","errorType":"error_code","errorClass":"RagError","httpStatus":null,"severity":"error","filePath":"src-tauri/plugins/tauri-plugin-rag/src/error.rs","lineNumber":5,"sourceCode":"use serde::{Deserialize, Serialize};\n\n#[derive(Debug, thiserror::Error, Serialize, Deserialize)]\npub enum RagError {\n    #[error(\"Failed to parse document: {0}\")]\n    ParseError(String),\n\n    #[error(\"Unsupported file type: {0}\")]\n    UnsupportedFileType(String),\n\n    #[error(\"IO error: {0}\")]\n    IoError(String),\n}\n\nimpl From<std::io::Error> for RagError {\n    fn from(err: std::io::Error) -> Self {\n        RagError::IoError(err.to_string())\n    }\n}\n\n","sourceCodeStart":1,"sourceCodeEnd":21,"githubUrl":"https://github.com/janhq/jan/blob/7205d770c1e097c3daf35a911176410e93bc5564/src-tauri/plugins/tauri-plugin-rag/src/error.rs#L1-L21","documentation":"RagError::ParseError is raised when the RAG plugin fails to parse a document's contents (displayed as \"Failed to parse document: {0}\" with the underlying reason). Document ingestion (PDF, text, etc.) must yield readable text before chunking/embedding; unparsable or damaged files abort the pipeline with this error.","triggerScenarios":"Ingesting a file into the RAG index whose contents can't be decoded — corrupted PDFs, password-protected documents, binary files mislabeled as text, or encoding failures.","commonSituations":"Uploading scanned/image-only PDFs without an OCR layer; files with wrong extensions; documents saved with unexpected encodings; truncated uploads.","solutions":["Read the wrapped message to see which parser stage failed.","Verify the file opens correctly in its native application; re-export or repair if not.","Convert problematic documents (e.g. re-save the PDF, normalize to UTF-8 text) before ingestion.","For image-only PDFs, run OCR first or reject the file with a friendly user message."],"exampleFix":"// before\nrag.ingest(path).unwrap();\n\n// after\nmatch rag.ingest(path) {\n    Err(RagError::ParseError(msg)) => {\n        eprintln!(\"Skipping {path:?}: not a readable document ({msg})\");\n    }\n    other => other.expect(\"ingest failed\"),\n}","handlingStrategy":"validation","validationCode":"// Rust: pre-check file readability before RAG ingestion\nfn looks_parseable(path: &Path) -> bool {\n    std::fs::read(path).map(|bytes| !bytes.is_empty() && !bytes.starts_with(&[0u8; 4][..])).unwrap_or(false)\n}","typeGuard":"fn as_parse_error(err: &RagError) -> Option<&str> {\n    match err { RagError::ParseError(m) => Some(m), _ => None }\n}","tryCatchPattern":"match rag.ingest(doc) {\n    Err(RagError::ParseError(msg)) => {\n        notify_user(format!(\"Could not read document: {msg}\"));\n    }\n    other => other?,\n}","preventionTips":["Validate files open in their native format before ingestion.","Run OCR on image-only PDFs before indexing.","Normalize text encodings to UTF-8 during upload.","Quarantine files that repeatedly fail parsing for user review."],"tags":["rag","rust","document-parsing"],"backgroundTag":"json-parse-error","analyzedSha":"7205d770c1e097c3daf35a911176410e93bc5564","analyzedAt":"2026-09-17T14:27:30.100Z","contentChangedAt":"2026-09-17T14:27:30.100Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}