apache/seatunnel · error · FileConnectorException

FILE_READ_FAILED

FILE_READ_FAILED

Error message

Failed to parse PDF document [%s]: %s

What it means

PdfReadStrategy.read parses a PDF (optionally extracting RAG metadata) and wraps any parsing exception into FileConnectorException(FILE_READ_FAILED) with the document path and the underlying message. It means the PDF could not be opened or processed, including download-to-temp or decryption issues.

Solutions

  1. Open the failing PDF locally (path shown in the message) to confirm it is valid and unencrypted
  2. If password-protected, remove the password or use a library-configured credential path if available
  3. Re-copy/re-upload the file to rule out truncation or corruption
  4. Check the wrapped cause (e) for the exact underlying library error and address it specifically

Example fix

// before
// reading file.pdf which is actually a renamed .docx
// after
file file.pdf | grep -i pdf  # verify real content type before configuring file_format = pdf
Defensive patterns

Strategy: try-catch

Validate before calling

// pre-check before configuring pdf format
Path p = Paths.get(localOrRemotePath);
byte[] head = Files.readAllBytes(p); // or read first 5 bytes
if (!new String(head, 0, 5, StandardCharsets.US_ASCII).startsWith("%PDF-"))
    throw new IllegalArgumentException("Not a PDF: " + p);

Type guard

boolean looksLikePdf(byte[] head) {
    return head != null && head.length >= 5
        && head[0]=='%' && head[1]=='P' && head[2]=='D' && head[3]=='F' && head[4]=='-';
}

Try / catch

try {
    strategy.read(split, reader, output);
} catch (FileConnectorException e) {
    if (FileConnectorErrorCode.FILE_READ_FAILED.equals(e.getErrorCode())) {
        LOG.error("PDF parse failed for {}: cause={}", split.getFilePath(), e.getCause(), e);
        // quarantine the file instead of failing the whole job
    } else throw e;
}

Prevention

When it happens

Trigger: read() on a split whose file is corrupt, encrypted/password-protected, not actually a PDF, or whose remote stream fails mid-download; any Exception raised by the underlying PDF library while iterating pages/elements.

Common situations: Misconfigured file_format=pdf pointing at non-PDF files; corrupted uploads; PDFs with restrictive permissions; temporary download failures from remote storage; malformed RAG metadata extraction.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/b13028e72398d6d1. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java:159

                log.info(
                        "PDF file '{}' processed successfully, generated {} elements",
                        path,
                        elements.size());

                int chunkIndex = 1;
                for (DocumentElement element : elements) {
                    output.collect(
                            toSeaTunnelRow(
                                    element,
                                    sourceUri,
                                    documentId,
                                    chunkIndex++,
                                    pdfRagMetadataEnabled));
                }
            }
        } catch (Exception e) {
            throw new FileConnectorException(
                    FileConnectorErrorCode.FILE_READ_FAILED,
                    String.format("Failed to parse PDF document [%s]: %s", path, e.getMessage()),
                    e);
        } finally {
            if (tempPdfPath != null) {
                deleteTempPdfPath(tempPdfPath);
            }
        }
    }

    Path createTempPdfPath() throws IOException {
        return Files.createTempFile("seatunnel-pdf-read-", ".pdf");
    }

    void deleteTempPdfPath(Path tempPdfPath) throws IOException {
        Files.deleteIfExists(tempPdfPath);
    }

View on GitHub (pinned to cf67b549a7)