apache/seatunnel · warning

Could not extract coordinates for outline item

Error message

Could not extract coordinates for outline item: {}

What it means

Warning in PdfReadStrategy.extractCoordinateFromOutline: when extracting the (page, x, y) coordinates for a PDF outline item, an IOException occurs. The method returns null, so that outline item simply contributes no coordinate information rather than aborting the read.

Solutions

  1. Repair/re-export the PDF to regenerate valid outline destinations.
  2. If coordinates (for heading-based paragraph grouping) are essential, use PDFs with standard GoTo page destinations.
  3. Ignore if coordinate extraction is optional — extraction continues for other outline items.
  4. Validate the PDF (e.g. with qpdf --check) and fix structural issues before ingestion.
Defensive patterns

Strategy: fallback

Validate before calling

// verify destination resolvability before coordinate extraction
PDDestination d = item.getDestination();
if (!(d instanceof PDPageXYZDestination)) { /* skip coordinate extraction */ }

Type guard

boolean hasXyzDestination(PDOutlineItem item) {
    try { return item.getDestination() instanceof PDPageXYZDestination; }
    catch (IOException e) { return false; }
}

Try / catch

try {
    // resolve page + xyz coordinates
} catch (IOException e) {
    log.warn("Could not extract coordinates for outline item: {}", item.getTitle());
    return null;
}

Prevention

When it happens

Trigger: Resolving the outline item's destination page or XYZ coordinates throws IOException — unresolvable destination, malformed XYZ destination, or document access errors while navigating pages.

Common situations: PDFs with outline entries pointing to actions rather than page destinations; broken cross-reference tables; documents produced by generators with non-standard destinations; partially corrupt downloads.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/8412c1f286f75456. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java:402

                // Get page number
                int pageNumber = xyzDest.retrievePageNumber();
                if (pageNumber < 0) {
                    PDPage page = xyzDest.getPage();
                    if (page != null) {
                        pageNumber = document.getPages().indexOf(page);
                    }
                }

                // Get coordinates - note that PDF coordinates have origin at bottom-left
                float x = xyzDest.getLeft();
                float y = xyzDest.getTop();

                return new CoordinateInfo(pageNumber, x, y, elementId);
            }

        } catch (IOException e) {
            log.warn("Could not extract coordinates for outline item: {}", item.getTitle());
        }

        return null;
    }

    /** Extract paragraphs for each heading using coordinate-based approach */
    private List<DocumentElement> extractParagraphsForHeadings(
            PDDocument document,
            List<DocumentElement> headings,
            Map<String, CoordinateInfo> coordMap)
            throws IOException {

        List<DocumentElement> paragraphs = new ArrayList<>();

        for (int i = 0; i < headings.size(); i++) {
            DocumentElement heading = headings.get(i);
            DocumentElement nextHeading = (i + 1 < headings.size()) ? headings.get(i + 1) : null;

View on GitHub (pinned to cf67b549a7)