santifer/career-ops · error · Error

Could not fetch job page

Error message

Could not fetch job page: ${e.message}

What it means

fetchJobPage fetches a job posting URL, strips HTML tags, and returns trimmed text. Any failure during the fetch/parse (network error, non-2xx HTTP status, invalid URL, body read failure) is caught and re-thrown as 'Could not fetch job page: <cause>' so callers get one consistent wrapper message.

Solutions

  1. Open the original message in e.message to identify the underlying cause (HTTP status vs network error)
  2. Verify the URL is correct and the posting is still live in a browser
  3. Retry with the full /career-ops scan pipeline (Playwright fallback) for JS-rendered or bot-protected pages
  4. Check network/proxy connectivity and retry; for persistent 403, the site likely blocks the User-Agent

Example fix

// before
const html = await fetchJobPage('https://boards.greenhouse.io/acme/jobs/123');
// after
try {
  const html = await fetchJobPage(url);
} catch (e) {
  console.error(e.message); // e.g. 'Could not fetch job page: HTTP 404 Not Found'
  // fall back to browser-based extraction or skip this posting
}
Defensive patterns

Strategy: try-catch

Validate before calling

function isLikelyJobUrl(u) { try { const p = new URL(u); return p.protocol === 'https:' || p.protocol === 'http:'; } catch { return false; } }

Type guard

const isHttpUrl = (u) => { try { const p = new URL(u); return ['http:','https:'].includes(p.protocol); } catch { return false; } };

Try / catch

try { const text = await fetchJobPage(url); } catch (e) { if (!e.message.startsWith('Could not fetch job page')) throw e; logger.warn({ url, cause: e.message }, 'job page fetch failed'); text = null; }

Prevention

When it happens

Trigger: fetchJobPage(url) is called with an unreachable or dead URL, the server returns a non-ok status (the inner throw `HTTP ${r.status} ${r.statusText}`), the request is rejected (DNS failure, TLS error, timeout), or r.text() fails.

Common situations: Job posting has been taken down (404/410); corporate firewall or proxy blocks the request; site blocks the default User-Agent (403); typo'd or non-URL input; network is offline.

Understand the failure class

Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.

Related errors


AI-assisted analysis of santifer/career-ops@aac998c7ed (2026-09-16). Data as JSON: /api/errors/86a8fbb17a13e3ba. Report an issue: GitHub.

Appendix: source

Thrown at openrouter-runner.mjs:458

      });
      return text.slice(0, 16_000);
    } catch (e) {
      console.warn(`[fetch] Playwright error: ${e.message} — falling back to plain fetch.`);
    } finally {
      if (browser) await browser.close().catch(() => {});
    }
  }

  // Plain HTTP fallback
  try {
    const r = await fetch(url, {
      headers: { 'User-Agent': DEFAULT_USER_AGENT }
    });
    if (!r.ok) throw new Error(`HTTP ${r.status} ${r.statusText}`);
    const html = await r.text();
    return html.replace(/<[^>]+>/g, ' ').replace(/\s+/g, ' ').trim().slice(0, 16_000);
  } catch (e) {
    throw new Error(`Could not fetch job page: ${e.message}`);
  }
}

// ---------------------------------------------------------------------------
// portals.yml parser — reads the canonical schema with js-yaml (same library and
// field names as scan.mjs: `title_filter.positive/negative` + `tracked_companies`),
// so it never drifts from the main scanner. The runner's no-CLI scan path covers
// companies that expose a direct JSON `api:`; careers_url-only / Playwright /
// search-query companies are handled by the full /career-ops scan pipeline.
// `rawOverride` lets tests feed YAML text directly (see test-all.mjs drift guard).
// ---------------------------------------------------------------------------
export function parsePortals(rawOverride) {
  const raw = rawOverride ?? readFile('portals.yml');
  if (!raw) throw new Error('portals.yml not found');
  const config = yaml.load(raw) || {};

  // The shared predicate rather than a second copy of the matching rules. This
  // path kept its own `includes` loop, and the two had drifted three ways: an

View on GitHub (pinned to aac998c7ed)