santifer/career-ops · error · Error

Could not determine the rendered PDF page count from its…

Error message

Could not determine the rendered PDF page count from its page tree.

What it means

countRenderedPdfPages parses the PDF bytes Chromium produced, finds the /Type /Catalog object, follows its /Pages reference, and reads /Count from the page-tree root to get the page count. If any step fails — no catalog, no /Pages ref, no /Type /Pages object, or a missing/unparseable /Count — pageCount is 0 and this error is thrown rather than returning a bogus count.

Solutions

  1. Verify the PDF buffer is a real, complete PDF (starts with %PDF-, non-trivial size); if not, fix the render step (Playwright page.pdf) first
  2. Update Playwright/Chromium to a current version — old or mismatched browsers can emit PDF layouts the parser misses
  3. Re-render the PDF; a transient Chromium failure often resolves on retry
  4. If parsing custom PDFs, pre-flatten object streams (e.g. via qpdf) before counting pages
  5. Inspect the failing PDF manually (pdfinfo) to confirm it has a readable page tree

Example fix

// before
const pdf = await page.pdf({ format: 'A4' });
enforcePageBudget(countPages(pdf));
// after: guard against empty/failed renders
const pdf = await page.pdf({ format: 'A4' });
if (!pdf || pdf.length < 100 || !pdf.subarray(0, 5).toString('latin1').startsWith('%PDF-')) {
  throw new Error('Chromium produced an empty or invalid PDF; check the render step');
}
enforcePageBudget(countRenderedPdfPages(pdf));
Defensive patterns

Strategy: try-catch

Validate before calling

function isPlausiblePdf(buf) {
  return Buffer.isBuffer(buf) && buf.length > 100 && buf.subarray(0, 5).toString('latin1') === '%PDF-';
}
if (!isPlausiblePdf(pdfBuffer)) throw new Error('render produced no valid PDF');

Type guard

const hasPageTree = (buf) => /\/Type\s*\/Catalog\b[\s\S]*?\/Pages\s+\d+\s+\d+\s+R\b/.test(buf.toString('latin1'));

Try / catch

try {
  const pageCount = countRenderedPdfPages(pdfBuffer);
} catch (err) {
  if (err.message.includes('page tree')) {
    console.error('Unparseable or empty PDF — check Chromium/Playwright render output, retry, or inspect with pdfinfo.');
    throw err; // or fall back to re-rendering
  }
  throw err;
}

Prevention

When it happens

Trigger: The PDF buffer is empty or truncated (Chromium crashed or page.goto failed), the buffer is not actually a PDF (an HTML error page was saved), or the PDF uses an incremental/compressed object layout (object streams) the regex-based parser cannot read, so the catalog or page tree is not found.

Common situations: Outdated or broken Chromium/Playwright returning partial PDFs; a proxy or error page captured instead of a PDF; very large PDFs with cross-reference streams; version changes in the headless browser's PDF writer.

Related errors


AI-assisted analysis of santifer/career-ops@aac998c7ed (2026-09-16). Data as JSON: /api/errors/34bd01f3a554f486. Report an issue: GitHub.

Appendix: source

Thrown at generate-pdf.mjs:1070

  const objects = new Map();
  const objectPattern = /(?:^|[\r\n])(\d+)\s+(\d+)\s+obj\b([\s\S]*?)\bendobj\b/g;

  for (const match of pdf.matchAll(objectPattern)) {
    const streamIndex = match[3].search(/\bstream(?:\r?\n|\r)/);
    const dictionary = streamIndex === -1 ? match[3] : match[3].slice(0, streamIndex);
    objects.set(`${match[1]} ${match[2]}`, dictionary);
  }

  const catalog = [...objects.values()].find((body) => /\/Type\s*\/Catalog\b/.test(body));
  const pagesRef = catalog?.match(/\/Pages\s+(\d+)\s+(\d+)\s+R\b/);
  const pages = pagesRef ? objects.get(`${pagesRef[1]} ${pagesRef[2]}`) : null;
  const count = pages && /\/Type\s*\/Pages\b/.test(pages)
    ? pages.match(/\/Count\s+(\d+)\b/)
    : null;
  const pageCount = count ? Number(count[1]) : 0;

  if (!Number.isInteger(pageCount) || pageCount < 1) {
    throw new Error('Could not determine the rendered PDF page count from its page tree.');
  }
  return pageCount;
}

/**
 * Convert a path to a workspace-relative manifest entry, or blank if it is
 * unknown or outside the tracker-owned workspace.
 *
 * @param {string} pathValue - Absolute or cwd-relative filesystem path.
 * @param {string} [rootDir] - Workspace root used as the manifest base.
 * @returns {string} Workspace-relative path using forward slashes, or an empty string.
 */
export function workspaceRelativeManifestPath(pathValue, rootDir = currentWorkspaceRoot()) {
  if (!pathValue) return '';
  const rel = relative(rootDir, resolve(pathValue));
  if (rel === '' || rel === '..' || rel.startsWith(`..${sep}`) || isAbsolute(rel)) return '';
  return rel.split(sep).join('/');
}

View on GitHub (pinned to aac998c7ed)