apache/seatunnel · warning

Could not determine page number for heading

Error message

Could not determine page number for heading: {}

What it means

Warning in PdfReadStrategy.convertOutlineToHeading: when resolving a PDF outline/bookmark entry to a page number, resolving the destination or computing the page index throws an IOException. The heading is still returned but without a page number (pageNumber not set).

Solutions

  1. Open the PDF in a viewer and fix/rebuild the bookmarks, then re-ingest.
  2. Re-export the PDF from the source tool to regenerate a clean outline.
  3. If page numbers are optional for your use case, ignore the warning — headings are still extracted.
  4. Pre-process with a PDF tool (e.g. qpdf/ghostscript) to normalize or strip the outline.
Defensive patterns

Strategy: fallback

Validate before calling

// validate outline destinations before processing
PDOutlineItem item = ...;
PDDestination dest = item.getDestination();
if (dest == null || document.getPageNumber(...) < 0) { /* skip or repair outline */ }

Try / catch

try {
    PDPage destPage = resolver.resolveDestinationPage(item);
    int idx = document.getPages().indexOf(destPage);
    if (idx >= 0) heading.setPageNumber(idx + 1);
} catch (IOException e) {
    log.warn("Could not determine page number for heading: {}", title);
}

Prevention

When it happens

Trigger: A PDF outline destination cannot be resolved to a page — corrupted or non-standard outline destinations, named destinations that don't resolve, or I/O errors while parsing the document's page tree during outline conversion.

Common situations: PDFs generated by tools that emit broken/unresolvable bookmarks; scanned or OCR'd PDFs with malformed outline trees; encrypted or partially corrupted PDFs; PDFBox failing on named destinations.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/4043f90585643548. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java:349

        }

        String title = item.getTitle().trim();
        DocumentElement heading = new DocumentElement("heading", title);
        heading.setHeadingLevel(Math.min(level, 6)); // Limit to max level 6
        heading.setPositionIndex(positionIndex);
        heading.setParentId(parentId);

        // Get page number from destination
        try {
            PDPage destinationPage = item.findDestinationPage(document);
            if (destinationPage != null) {
                int pageIndex = document.getPages().indexOf(destinationPage);
                if (pageIndex >= 0) {
                    heading.setPageNumber(pageIndex + 1);
                }
            }
        } catch (IOException e) {
            log.warn("Could not determine page number for heading: {}", title);
        }

        return heading;
    }

    /** Extract coordinate information from PDF outline item */
    private CoordinateInfo extractCoordinateFromOutline(
            PDDocument document, PDOutlineItem item, String elementId) {
        try {
            PDDestination destination = item.getDestination();

            if (destination == null) {
                // Check for action-based destinations
                PDAction action = item.getAction();
                if (action instanceof PDActionGoTo) {
                    destination = ((PDActionGoTo) action).getDestination();
                }
            }

View on GitHub (pinned to cf67b549a7)