apache/seatunnel · warning
Could not extract coordinates for outline item
Error message
Could not extract coordinates for outline item: {} What it means
Warning in PdfReadStrategy.extractCoordinateFromOutline: when extracting the (page, x, y) coordinates for a PDF outline item, an IOException occurs. The method returns null, so that outline item simply contributes no coordinate information rather than aborting the read.
Solutions
- Repair/re-export the PDF to regenerate valid outline destinations.
- If coordinates (for heading-based paragraph grouping) are essential, use PDFs with standard GoTo page destinations.
- Ignore if coordinate extraction is optional — extraction continues for other outline items.
- Validate the PDF (e.g. with qpdf --check) and fix structural issues before ingestion.
Defensive patterns
Strategy: fallback
Validate before calling
// verify destination resolvability before coordinate extraction
PDDestination d = item.getDestination();
if (!(d instanceof PDPageXYZDestination)) { /* skip coordinate extraction */ } Type guard
boolean hasXyzDestination(PDOutlineItem item) {
try { return item.getDestination() instanceof PDPageXYZDestination; }
catch (IOException e) { return false; }
} Try / catch
try {
// resolve page + xyz coordinates
} catch (IOException e) {
log.warn("Could not extract coordinates for outline item: {}", item.getTitle());
return null;
} Prevention
- Ensure PDFs use standard GoTo/XYZ destinations, not actions.
- Run qpdf --check and repair corrupt PDFs before ingestion.
- Design downstream logic to tolerate null coordinates for some outline items.
When it happens
Trigger: Resolving the outline item's destination page or XYZ coordinates throws IOException — unresolvable destination, malformed XYZ destination, or document access errors while navigating pages.
Common situations: PDFs with outline entries pointing to actions rather than page destinations; broken cross-reference tables; documents produced by generators with non-standard destinations; partially corrupt downloads.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10).
Data as JSON: /api/errors/8412c1f286f75456.
Report an issue: GitHub.
Appendix: source
Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java:402
// Get page number
int pageNumber = xyzDest.retrievePageNumber();
if (pageNumber < 0) {
PDPage page = xyzDest.getPage();
if (page != null) {
pageNumber = document.getPages().indexOf(page);
}
}
// Get coordinates - note that PDF coordinates have origin at bottom-left
float x = xyzDest.getLeft();
float y = xyzDest.getTop();
return new CoordinateInfo(pageNumber, x, y, elementId);
}
} catch (IOException e) {
log.warn("Could not extract coordinates for outline item: {}", item.getTitle());
}
return null;
}
/** Extract paragraphs for each heading using coordinate-based approach */
private List<DocumentElement> extractParagraphsForHeadings(
PDDocument document,
List<DocumentElement> headings,
Map<String, CoordinateInfo> coordMap)
throws IOException {
List<DocumentElement> paragraphs = new ArrayList<>();
for (int i = 0; i < headings.size(); i++) {
DocumentElement heading = headings.get(i);
DocumentElement nextHeading = (i + 1 < headings.size()) ? headings.get(i + 1) : null;
View on GitHub (pinned to cf67b549a7)