apache/seatunnel · warning
Could not determine page number for heading
Error message
Could not determine page number for heading: {} What it means
Warning in PdfReadStrategy.convertOutlineToHeading: when resolving a PDF outline/bookmark entry to a page number, resolving the destination or computing the page index throws an IOException. The heading is still returned but without a page number (pageNumber not set).
Solutions
- Open the PDF in a viewer and fix/rebuild the bookmarks, then re-ingest.
- Re-export the PDF from the source tool to regenerate a clean outline.
- If page numbers are optional for your use case, ignore the warning — headings are still extracted.
- Pre-process with a PDF tool (e.g. qpdf/ghostscript) to normalize or strip the outline.
Defensive patterns
Strategy: fallback
Validate before calling
// validate outline destinations before processing
PDOutlineItem item = ...;
PDDestination dest = item.getDestination();
if (dest == null || document.getPageNumber(...) < 0) { /* skip or repair outline */ } Try / catch
try {
PDPage destPage = resolver.resolveDestinationPage(item);
int idx = document.getPages().indexOf(destPage);
if (idx >= 0) heading.setPageNumber(idx + 1);
} catch (IOException e) {
log.warn("Could not determine page number for heading: {}", title);
} Prevention
- Repair or rebuild bookmarks in PDFs before ingestion.
- Use qpdf/ghostscript to normalize malformed outlines.
- Treat missing page numbers as optional metadata in downstream processing.
When it happens
Trigger: A PDF outline destination cannot be resolved to a page — corrupted or non-standard outline destinations, named destinations that don't resolve, or I/O errors while parsing the document's page tree during outline conversion.
Common situations: PDFs generated by tools that emit broken/unresolvable bookmarks; scanned or OCR'd PDFs with malformed outline trees; encrypted or partially corrupted PDFs; PDFBox failing on named destinations.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10).
Data as JSON: /api/errors/4043f90585643548.
Report an issue: GitHub.
Appendix: source
Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/PdfReadStrategy.java:349
}
String title = item.getTitle().trim();
DocumentElement heading = new DocumentElement("heading", title);
heading.setHeadingLevel(Math.min(level, 6)); // Limit to max level 6
heading.setPositionIndex(positionIndex);
heading.setParentId(parentId);
// Get page number from destination
try {
PDPage destinationPage = item.findDestinationPage(document);
if (destinationPage != null) {
int pageIndex = document.getPages().indexOf(destinationPage);
if (pageIndex >= 0) {
heading.setPageNumber(pageIndex + 1);
}
}
} catch (IOException e) {
log.warn("Could not determine page number for heading: {}", title);
}
return heading;
}
/** Extract coordinate information from PDF outline item */
private CoordinateInfo extractCoordinateFromOutline(
PDDocument document, PDOutlineItem item, String elementId) {
try {
PDDestination destination = item.getDestination();
if (destination == null) {
// Check for action-based destinations
PDAction action = item.getAction();
if (action instanceof PDActionGoTo) {
destination = ((PDActionGoTo) action).getDestination();
}
}View on GitHub (pinned to cf67b549a7)