{"record":{"id":"4043f90585643548","repo":"apache/seatunnel","slug":"could-not-determine-page-number-for-heading","errorCode":null,"errorMessage":"Could not determine page number for heading: {}","messagePattern":"Could not determine page number for heading: (.+?)","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java","lineNumber":349,"sourceCode":"        }\n\n        String title = item.getTitle().trim();\n        DocumentElement heading = new DocumentElement(\"heading\", title);\n        heading.setHeadingLevel(Math.min(level, 6)); // Limit to max level 6\n        heading.setPositionIndex(positionIndex);\n        heading.setParentId(parentId);\n\n        // Get page number from destination\n        try {\n            PDPage destinationPage = item.findDestinationPage(document);\n            if (destinationPage != null) {\n                int pageIndex = document.getPages().indexOf(destinationPage);\n                if (pageIndex >= 0) {\n                    heading.setPageNumber(pageIndex + 1);\n                }\n            }\n        } catch (IOException e) {\n            log.warn(\"Could not determine page number for heading: {}\", title);\n        }\n\n        return heading;\n    }\n\n    /** Extract coordinate information from PDF outline item */\n    private CoordinateInfo extractCoordinateFromOutline(\n            PDDocument document, PDOutlineItem item, String elementId) {\n        try {\n            PDDestination destination = item.getDestination();\n\n            if (destination == null) {\n                // Check for action-based destinations\n                PDAction action = item.getAction();\n                if (action instanceof PDActionGoTo) {\n                    destination = ((PDActionGoTo) action).getDestination();\n                }\n            }","sourceCodeStart":331,"sourceCodeEnd":367,"githubUrl":"https://github.com/apache/seatunnel/blob/cf67b549a7a6c35fa0beb12d83c62892427ea919/seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java#L331-L367","documentation":"Warning in PdfReadStrategy.convertOutlineToHeading: when resolving a PDF outline/bookmark entry to a page number, resolving the destination or computing the page index throws an IOException. The heading is still returned but without a page number (pageNumber not set).","triggerScenarios":"A PDF outline destination cannot be resolved to a page — corrupted or non-standard outline destinations, named destinations that don't resolve, or I/O errors while parsing the document's page tree during outline conversion.","commonSituations":"PDFs generated by tools that emit broken/unresolvable bookmarks; scanned or OCR'd PDFs with malformed outline trees; encrypted or partially corrupted PDFs; PDFBox failing on named destinations.","solutions":["Open the PDF in a viewer and fix/rebuild the bookmarks, then re-ingest.","Re-export the PDF from the source tool to regenerate a clean outline.","If page numbers are optional for your use case, ignore the warning — headings are still extracted.","Pre-process with a PDF tool (e.g. qpdf/ghostscript) to normalize or strip the outline."],"exampleFix":null,"handlingStrategy":"fallback","validationCode":"// validate outline destinations before processing\nPDOutlineItem item = ...;\nPDDestination dest = item.getDestination();\nif (dest == null || document.getPageNumber(...) < 0) { /* skip or repair outline */ }","typeGuard":null,"tryCatchPattern":"try {\n    PDPage destPage = resolver.resolveDestinationPage(item);\n    int idx = document.getPages().indexOf(destPage);\n    if (idx >= 0) heading.setPageNumber(idx + 1);\n} catch (IOException e) {\n    log.warn(\"Could not determine page number for heading: {}\", title);\n}","preventionTips":["Repair or rebuild bookmarks in PDFs before ingestion.","Use qpdf/ghostscript to normalize malformed outlines.","Treat missing page numbers as optional metadata in downstream processing."],"tags":["pdf","outline","bookmark","pdfbox"],"backgroundTag":"file-read-failed","analyzedSha":"cf67b549a7a6c35fa0beb12d83c62892427ea919","analyzedAt":"2026-09-10T21:44:55.265Z","contentChangedAt":"2026-09-10T21:44:55.265Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}