{"record":{"id":"86a8fbb17a13e3ba","repo":"santifer/career-ops","slug":"could-not-fetch-job-page-e-message","errorCode":null,"errorMessage":"Could not fetch job page: ${e.message}","messagePattern":"Could not fetch job page: (.+?)","errorType":"http","errorClass":"Error","httpStatus":null,"severity":"error","filePath":"openrouter-runner.mjs","lineNumber":458,"sourceCode":"      });\n      return text.slice(0, 16_000);\n    } catch (e) {\n      console.warn(`[fetch] Playwright error: ${e.message} — falling back to plain fetch.`);\n    } finally {\n      if (browser) await browser.close().catch(() => {});\n    }\n  }\n\n  // Plain HTTP fallback\n  try {\n    const r = await fetch(url, {\n      headers: { 'User-Agent': DEFAULT_USER_AGENT }\n    });\n    if (!r.ok) throw new Error(`HTTP ${r.status} ${r.statusText}`);\n    const html = await r.text();\n    return html.replace(/<[^>]+>/g, ' ').replace(/\\s+/g, ' ').trim().slice(0, 16_000);\n  } catch (e) {\n    throw new Error(`Could not fetch job page: ${e.message}`);\n  }\n}\n\n// ---------------------------------------------------------------------------\n// portals.yml parser — reads the canonical schema with js-yaml (same library and\n// field names as scan.mjs: `title_filter.positive/negative` + `tracked_companies`),\n// so it never drifts from the main scanner. The runner's no-CLI scan path covers\n// companies that expose a direct JSON `api:`; careers_url-only / Playwright /\n// search-query companies are handled by the full /career-ops scan pipeline.\n// `rawOverride` lets tests feed YAML text directly (see test-all.mjs drift guard).\n// ---------------------------------------------------------------------------\nexport function parsePortals(rawOverride) {\n  const raw = rawOverride ?? readFile('portals.yml');\n  if (!raw) throw new Error('portals.yml not found');\n  const config = yaml.load(raw) || {};\n\n  // The shared predicate rather than a second copy of the matching rules. This\n  // path kept its own `includes` loop, and the two had drifted three ways: an","sourceCodeStart":440,"sourceCodeEnd":476,"githubUrl":"https://github.com/santifer/career-ops/blob/aac998c7ed7248ea853b720ceeb1fdbeb322fc5d/openrouter-runner.mjs#L440-L476","documentation":"fetchJobPage fetches a job posting URL, strips HTML tags, and returns trimmed text. Any failure during the fetch/parse (network error, non-2xx HTTP status, invalid URL, body read failure) is caught and re-thrown as 'Could not fetch job page: <cause>' so callers get one consistent wrapper message.","triggerScenarios":"fetchJobPage(url) is called with an unreachable or dead URL, the server returns a non-ok status (the inner throw `HTTP ${r.status} ${r.statusText}`), the request is rejected (DNS failure, TLS error, timeout), or r.text() fails.","commonSituations":"Job posting has been taken down (404/410); corporate firewall or proxy blocks the request; site blocks the default User-Agent (403); typo'd or non-URL input; network is offline.","solutions":["Open the original message in e.message to identify the underlying cause (HTTP status vs network error)","Verify the URL is correct and the posting is still live in a browser","Retry with the full /career-ops scan pipeline (Playwright fallback) for JS-rendered or bot-protected pages","Check network/proxy connectivity and retry; for persistent 403, the site likely blocks the User-Agent"],"exampleFix":"// before\nconst html = await fetchJobPage('https://boards.greenhouse.io/acme/jobs/123');\n// after\ntry {\n  const html = await fetchJobPage(url);\n} catch (e) {\n  console.error(e.message); // e.g. 'Could not fetch job page: HTTP 404 Not Found'\n  // fall back to browser-based extraction or skip this posting\n}","handlingStrategy":"try-catch","validationCode":"function isLikelyJobUrl(u) { try { const p = new URL(u); return p.protocol === 'https:' || p.protocol === 'http:'; } catch { return false; } }","typeGuard":"const isHttpUrl = (u) => { try { const p = new URL(u); return ['http:','https:'].includes(p.protocol); } catch { return false; } };","tryCatchPattern":"try { const text = await fetchJobPage(url); } catch (e) { if (!e.message.startsWith('Could not fetch job page')) throw e; logger.warn({ url, cause: e.message }, 'job page fetch failed'); text = null; }","preventionTips":["Validate the URL is well-formed http(s) before calling","Check liveness with check-liveness.mjs before heavy fetches","Handle non-2xx statuses explicitly by parsing the wrapped e.message","Fall back to the Playwright-based scan pipeline for bot-protected pages"],"tags":["network","http","fetch"],"backgroundTag":"http-error-response","analyzedSha":"aac998c7ed7248ea853b720ceeb1fdbeb322fc5d","analyzedAt":"2026-09-16T06:35:29.214Z","contentChangedAt":"2026-09-16T06:35:29.214Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}