santifer/career-ops · error
HTTP
Error message
HTTP ${res.status} ${res.statusText} What it means
upskill.mjs fetches JD URLs with fetch(..., { redirect: 'error' }) — it refuses to follow redirects (SSRF guard: validateUrlSecurity only vets the initial URL, so an unvetted Location hop is untrusted) — and throws `HTTP <status> <statusText>` whenever the response is not ok. The error is caught by the surrounding handler, which prints 'Fatal: Failed to fetch JD from URL' and exits 1.
Solutions
- Find and use the final, direct ATS posting URL (e.g. the Greenhouse/Lever/Ashby URL) instead of the redirecting link.
- For 404/410, the posting is gone — fetch an archived copy or use the report's archived JD.
- For 403, download the page in a browser and pass the saved file/text instead of the URL.
- For 429/5xx, wait and retry; the fetch also has a 30s timeout, so slow hosts may need the file path too.
- Note redirects are blocked by design (#1851); there is no flag to enable following them.
Example fix
// before node upskill.mjs --url https://www.linkedin.com/jobs/view/12345 # 302 → blocked // after node upskill.mjs --url https://boards.greenhouse.io/company/jobs/12345
Defensive patterns
Strategy: try-catch
Validate before calling
const res = await fetch(url, { method: 'HEAD', redirect: 'manual' });
if ([301,302,303,307,308].includes(res.status)) console.error('URL redirects; resolve to the final ATS URL first');
else if (!res.ok) console.error(`URL returns ${res.status}; use an archived copy or a direct link`); Try / catch
try {
await upskillFromUrl(url);
} catch (e) {
const m = e.message.match(/HTTP (\d+)/);
if (m && ['301','302','307','308'].includes(m[1])) console.error('blocked redirect: use the final posting URL');
else if (m && m[1] === '404') console.error('posting gone: use archived JD');
else if (m && ['429'].includes(m[1])) console.error('rate limited: retry later');
else throw e;
} Prevention
- Always pass the direct ATS (Greenhouse/Lever/Ashby) URL, not aggregator links that redirect.
- Remember redirect:'error' is intentional (#1851) — never try to follow redirects through the tool.
- Archive the JD at evaluation time so a later 404 doesn't block re-analysis.
- For bot-protected sites (403), save the page and use file input.
When it happens
Trigger: The initial (security-vetted) URL returns a non-2xx status: 301/302/307/308 redirects (turned into hard errors by redirect:'error'), 403 bot-blocking, 404 gone postings, 429 rate limits, or 5xx server errors.
Common situations: Job boards redirecting to an ATS (e.g. LinkedIn → company ATS) — the tool deliberately refuses to follow; Cloudflare or WAF returning 403 to non-browser clients; expired postings returning 404; transient 502/503 from the host.
Understand the failure class
Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.
Related errors
- Access denied: Egress guard blocked private target IP
- Access denied: Egress guard blocked private target IP
- Blocked request to restricted destination
- could not download the index
- Could not fetch job page
AI-assisted analysis of santifer/career-ops@aac998c7ed (2026-09-16).
Data as JSON: /api/errors/2b9758b3f49d251e.
Report an issue: GitHub.
Appendix: source
Thrown at upskill.mjs:982
} catch (err) {
console.error(`Security Violation on Redirect: ${err.message}`);
await route.abort('blockedbyclient');
process.exit(1);
}
});
await page.goto(secureUrl, { waitUntil: 'networkidle', timeout: 30000 });
targetText = await page.innerText('body');
} catch (err) {
console.warn('Playwright extraction failed or blocked, trying fallback WebFetch...', err.message);
try {
const secureUrl = await validateUrlSecurity(inputSource);
// validateUrlSecurity only vets the initial URL; a redirect could still
// steer the fetch at an internal host (SSRF). The Playwright path
// re-validates per hop, but this plain fetch must refuse redirects
// outright — fail closed rather than follow an unvetted Location (#1851).
const res = await fetch(secureUrl, { signal: AbortSignal.timeout(30000), redirect: 'error' });
if (!res.ok) throw new Error(`HTTP ${res.status} ${res.statusText}`);
targetText = await res.text();
} catch (fetchErr) {
console.error(`Fatal: Failed to fetch JD from URL: ${fetchErr.message}`);
process.exit(1);
}
} finally {
if (browser) await browser.close();
}
// Whitespace-collapse + length-cap the fetched page text. Use compactText
// (string -> string), NOT normalizeJd: normalizeJd expects the { title,
// text } DOM-read object and returns { url, title, text }, so feeding it
// the innerText/fetch STRING silently produced { text: '' } — destroying
// the JD and then throwing `text.matchAll is not a function` downstream
// (#1894). compactText is the string-in/string-out helper this wants.
try {
const { compactText } = await import('./browser-extract.mjs');
targetText = compactText(targetText);View on GitHub (pinned to aac998c7ed)