{"record":{"id":"8412c1f286f75456","repo":"apache/seatunnel","slug":"could-not-extract-coordinates-for-outline-item","errorCode":null,"errorMessage":"Could not extract coordinates for outline item: {}","messagePattern":"Could not extract coordinates for outline item: (.+?)","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java","lineNumber":402,"sourceCode":"\n                // Get page number\n                int pageNumber = xyzDest.retrievePageNumber();\n                if (pageNumber < 0) {\n                    PDPage page = xyzDest.getPage();\n                    if (page != null) {\n                        pageNumber = document.getPages().indexOf(page);\n                    }\n                }\n\n                // Get coordinates - note that PDF coordinates have origin at bottom-left\n                float x = xyzDest.getLeft();\n                float y = xyzDest.getTop();\n\n                return new CoordinateInfo(pageNumber, x, y, elementId);\n            }\n\n        } catch (IOException e) {\n            log.warn(\"Could not extract coordinates for outline item: {}\", item.getTitle());\n        }\n\n        return null;\n    }\n\n    /** Extract paragraphs for each heading using coordinate-based approach */\n    private List<DocumentElement> extractParagraphsForHeadings(\n            PDDocument document,\n            List<DocumentElement> headings,\n            Map<String, CoordinateInfo> coordMap)\n            throws IOException {\n\n        List<DocumentElement> paragraphs = new ArrayList<>();\n\n        for (int i = 0; i < headings.size(); i++) {\n            DocumentElement heading = headings.get(i);\n            DocumentElement nextHeading = (i + 1 < headings.size()) ? headings.get(i + 1) : null;\n","sourceCodeStart":384,"sourceCodeEnd":420,"githubUrl":"https://github.com/apache/seatunnel/blob/cf67b549a7a6c35fa0beb12d83c62892427ea919/seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java#L384-L420","documentation":"Warning in PdfReadStrategy.extractCoordinateFromOutline: when extracting the (page, x, y) coordinates for a PDF outline item, an IOException occurs. The method returns null, so that outline item simply contributes no coordinate information rather than aborting the read.","triggerScenarios":"Resolving the outline item's destination page or XYZ coordinates throws IOException — unresolvable destination, malformed XYZ destination, or document access errors while navigating pages.","commonSituations":"PDFs with outline entries pointing to actions rather than page destinations; broken cross-reference tables; documents produced by generators with non-standard destinations; partially corrupt downloads.","solutions":["Repair/re-export the PDF to regenerate valid outline destinations.","If coordinates (for heading-based paragraph grouping) are essential, use PDFs with standard GoTo page destinations.","Ignore if coordinate extraction is optional — extraction continues for other outline items.","Validate the PDF (e.g. with qpdf --check) and fix structural issues before ingestion."],"exampleFix":null,"handlingStrategy":"fallback","validationCode":"// verify destination resolvability before coordinate extraction\nPDDestination d = item.getDestination();\nif (!(d instanceof PDPageXYZDestination)) { /* skip coordinate extraction */ }","typeGuard":"boolean hasXyzDestination(PDOutlineItem item) {\n    try { return item.getDestination() instanceof PDPageXYZDestination; }\n    catch (IOException e) { return false; }\n}","tryCatchPattern":"try {\n    // resolve page + xyz coordinates\n} catch (IOException e) {\n    log.warn(\"Could not extract coordinates for outline item: {}\", item.getTitle());\n    return null;\n}","preventionTips":["Ensure PDFs use standard GoTo/XYZ destinations, not actions.","Run qpdf --check and repair corrupt PDFs before ingestion.","Design downstream logic to tolerate null coordinates for some outline items."],"tags":["pdf","outline","coordinates","pdfbox"],"backgroundTag":"file-read-failed","analyzedSha":"cf67b549a7a6c35fa0beb12d83c62892427ea919","analyzedAt":"2026-09-10T21:44:55.265Z","contentChangedAt":"2026-09-10T21:44:55.265Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}