n8n-io/n8n · error
PDF " " contains no extractable text (it may be a scanned…
Error message
PDF "${attachment.fileName}" contains no extractable text (it may be a scanned image). What it means
getText() succeeded but returned only whitespace — the PDF has no text layer, which almost always means it is a scan (raster images of pages). pdf-parse does not do OCR, so the extraction is reported as empty rather than fabricating content; the parser is properly destroyed in a finally block before this throw.
Solutions
- Run OCR on the PDF (e.g. with an OCR tool/service) and attach the text output instead
- Request a text-based (digital) version of the document
- For a few key pages, transcribe manually or use a vision-capable model on page images
Defensive patterns
Strategy: validation
When it happens
Trigger: Thrown at packages/@n8n/instance-ai/src/parsers/pdf-parser.ts:50 when the library encounters an invalid state.
Common situations: See trigger scenarios.
AI-assisted analysis of n8n-io/n8n@5ac6606e81 (2026-08-12).
Data as JSON: /api/errors/a505a7f8ecee9e85.
Report an issue: GitHub.
Appendix: source
Thrown at packages/@n8n/instance-ai/src/parsers/pdf-parser.ts:50
const { PDFParse } = await import('pdf-parse');
const parser = new PDFParse({ data: decoded });
let extractedText: string;
let totalPages: number;
try {
const result = await parser.getText();
extractedText = result.text;
totalPages = result.total;
} catch (error) {
const message = error instanceof Error ? error.message : 'unknown error';
throw new Error(`Failed to parse PDF "${attachment.fileName}": ${message}`);
} finally {
await parser.destroy();
}
const text = extractedText?.trim() ?? '';
if (!text) {
throw new Error(
`PDF "${attachment.fileName}" contains no extractable text (it may be a scanned image).`,
);
}
if (text.length > MAX_RESULT_CHARS) {
return {
text: text.slice(0, MAX_RESULT_CHARS),
pages: totalPages,
truncated: true,
};
}
return { text, pages: totalPages, truncated: false };
}
View on GitHub (pinned to 5ac6606e81)