{"record":{"id":"b13028e72398d6d1","repo":"apache/seatunnel","slug":"file-read-failed-b13028","errorCode":"FILE_READ_FAILED","errorMessage":"Failed to parse PDF document [%s]: %s","messagePattern":"Failed to parse PDF document \\[(.+?)\\]: (.+?)","errorType":"exception","errorClass":"FileConnectorException","httpStatus":null,"severity":"error","filePath":"seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java","lineNumber":159,"sourceCode":"\n                log.info(\n                        \"PDF file '{}' processed successfully, generated {} elements\",\n                        path,\n                        elements.size());\n\n                int chunkIndex = 1;\n                for (DocumentElement element : elements) {\n                    output.collect(\n                            toSeaTunnelRow(\n                                    element,\n                                    sourceUri,\n                                    documentId,\n                                    chunkIndex++,\n                                    pdfRagMetadataEnabled));\n                }\n            }\n        } catch (Exception e) {\n            throw new FileConnectorException(\n                    FileConnectorErrorCode.FILE_READ_FAILED,\n                    String.format(\"Failed to parse PDF document [%s]: %s\", path, e.getMessage()),\n                    e);\n        } finally {\n            if (tempPdfPath != null) {\n                deleteTempPdfPath(tempPdfPath);\n            }\n        }\n    }\n\n    Path createTempPdfPath() throws IOException {\n        return Files.createTempFile(\"seatunnel-pdf-read-\", \".pdf\");\n    }\n\n    void deleteTempPdfPath(Path tempPdfPath) throws IOException {\n        Files.deleteIfExists(tempPdfPath);\n    }\n","sourceCodeStart":141,"sourceCodeEnd":177,"githubUrl":"https://github.com/apache/seatunnel/blob/cf67b549a7a6c35fa0beb12d83c62892427ea919/seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java#L141-L177","documentation":"PdfReadStrategy.read parses a PDF (optionally extracting RAG metadata) and wraps any parsing exception into FileConnectorException(FILE_READ_FAILED) with the document path and the underlying message. It means the PDF could not be opened or processed, including download-to-temp or decryption issues.","triggerScenarios":"read() on a split whose file is corrupt, encrypted/password-protected, not actually a PDF, or whose remote stream fails mid-download; any Exception raised by the underlying PDF library while iterating pages/elements.","commonSituations":"Misconfigured file_format=pdf pointing at non-PDF files; corrupted uploads; PDFs with restrictive permissions; temporary download failures from remote storage; malformed RAG metadata extraction.","solutions":["Open the failing PDF locally (path shown in the message) to confirm it is valid and unencrypted","If password-protected, remove the password or use a library-configured credential path if available","Re-copy/re-upload the file to rule out truncation or corruption","Check the wrapped cause (e) for the exact underlying library error and address it specifically"],"exampleFix":"// before\n// reading file.pdf which is actually a renamed .docx\n// after\nfile file.pdf | grep -i pdf  # verify real content type before configuring file_format = pdf","handlingStrategy":"try-catch","validationCode":"// pre-check before configuring pdf format\nPath p = Paths.get(localOrRemotePath);\nbyte[] head = Files.readAllBytes(p); // or read first 5 bytes\nif (!new String(head, 0, 5, StandardCharsets.US_ASCII).startsWith(\"%PDF-\"))\n    throw new IllegalArgumentException(\"Not a PDF: \" + p);","typeGuard":"boolean looksLikePdf(byte[] head) {\n    return head != null && head.length >= 5\n        && head[0]=='%' && head[1]=='P' && head[2]=='D' && head[3]=='F' && head[4]=='-';\n}","tryCatchPattern":"try {\n    strategy.read(split, reader, output);\n} catch (FileConnectorException e) {\n    if (FileConnectorErrorCode.FILE_READ_FAILED.equals(e.getErrorCode())) {\n        LOG.error(\"PDF parse failed for {}: cause={}\", split.getFilePath(), e.getCause(), e);\n        // quarantine the file instead of failing the whole job\n    } else throw e;\n}","preventionTips":["Validate files start with %PDF- before ingesting","Reject or preprocess password-protected PDFs","Verify download integrity (checksums) for remote PDFs"],"tags":["pdf","parsing","file"],"backgroundTag":"file-read-failed","analyzedSha":"cf67b549a7a6c35fa0beb12d83c62892427ea919","analyzedAt":"2026-09-10T21:44:55.265Z","contentChangedAt":"2026-09-10T21:44:55.265Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}