santifer/career-ops · error

deutschebahn: cannot resolve db.jobs search id for

Error message

deutschebahn: cannot resolve db.jobs search id for ${entry.name}

What it means

The Deutsche Bahn provider scrapes DB's db.jobs portal, which requires a numeric search-config id in the URL path (/service/search/de-de/{searchId}). resolveConfig() must derive that id from the entry's `api:`/`careers_url`; when it returns null — because the value is missing, unparseable, or its host is not db.jobs (or a *.db.jobs subdomain) — fetch() throws this error. Note the id itself has a well-known fallback ('5441588'), so a null return means the HOST check failed, not that the id is absent.

Solutions

  1. Set the entry's `api:` to a db.jobs URL, e.g. https://db.jobs/service/search/de-de/5441588 — the search id is optional (well-known fallback) but the host must be db.jobs.
  2. If you only have the branded URL, follow its redirect once manually and copy the resulting db.jobs URL into the config.
  3. Check for hostname typos: the check is host === 'db.jobs' or host.endsWith('.db.jobs'), case-insensitive.
  4. If DB has moved off db.jobs, update the provider's host check and URL patterns to the new portal.

Example fix

// before (portals.yml)
- name: deutsche-bahn
  careers_url: https://jobs.deutschebahngroup.careers/search
// after
- name: deutsche-bahn
  careers_url: https://jobs.deutschebahngroup.careers/search
  api: https://db.jobs/service/search/de-de/5441588
Defensive patterns

Strategy: validation

Validate before calling

function isDbJobsEntry(entry) {
  const raw = entry.api || entry.careers_url || '';
  try {
    const u = new URL(raw);
    const host = u.host.toLowerCase();
    return u.protocol === 'https:' && (host === 'db.jobs' || host.endsWith('.db.jobs'));
  } catch { return false; }
}
if (!isDbJobsEntry(entry)) console.warn(`${entry.name}: use a db.jobs api: URL, not the branded careers front`);

Try / catch

try {
  await provider.fetch(entry, ctx);
} catch (e) {
  if (e.message.includes('cannot resolve db.jobs search id')) {
    // config problem, not transient — fix the entry's api: host, don't retry
    console.error(`Fix ${entry.name}: api: must be a db.jobs URL (e.g. https://db.jobs/service/search/de-de/5441588)`);
  } else throw e;
}

Prevention

When it happens

Trigger: fetch(entry) is called with an entry whose api/careers_url is (a) empty or not a valid URL, (b) not on host db.jobs / *.db.jobs (e.g. the branded jobs.deutschebahngroup.careers front, which only 302-redirects into db.jobs), or (c) an http/https URL on a foreign host — resolveConfig returns null in all cases and the error is thrown with the entry's name interpolated.

Common situations: Portals.yml configured with the branded careers URL jobs.deutschebahngroup.careers instead of a db.jobs URL; a typo in the hostname; an entry with no api/careers_url at all; DB migrating portals so the old host no longer matches the db.jobs check.

Understand the failure class

Background: "Invalid value" and "allowed values are" config errors: what your library rejected and how to fix it — this error's family across 41 libraries.

Related errors


AI-assisted analysis of santifer/career-ops@e7abd431fc (2026-09-16). Data as JSON: /api/errors/7561a8ca1efb0bfc. Report an issue: GitHub.

Appendix: source

Thrown at providers/deutschebahn.mjs:113

function resolveMaxPages(entry) {
  const v = entry?.max_pages;
  if (Number.isInteger(v) && v > 0) return Math.min(v, MAX_PAGES);
  return MAX_PAGES;
}

/** @type {Provider} */
export default {
  id: 'deutschebahn',

  detect(entry) {
    const url = entry.api || entry.careers_url || '';
    if (typeof url !== 'string') return null;
    return resolveConfig({ api: url }) ? { url } : null;
  },

  async fetch(entry, ctx) {
    const cfg = resolveConfig(entry);
    if (!cfg) throw new Error(`deutschebahn: cannot resolve db.jobs search id for ${entry.name}`);

    const wait = (ms) => (ctx.sleep ? ctx.sleep(ms) : new Promise((r) => setTimeout(r, ms)));
    const maxPages = resolveMaxPages(entry);
    const jobs = [];
    const seen = new Set();

    // Each page is a ~450KB HTML fragment; on a walk of up to 60 pages, an
    // occasional single-page timeout/abort is a transient blip, not a board
    // failure — fetchTextWithRetry absorbs it instead of failing the whole scan.

    for (let page = 0; page < maxPages; page++) {
      if (page > 0) await wait(PAGE_DELAY_MS);
      const url = `${cfg.searchBase}?qli=true&query=&sort=score&itemsPerPage=${ITEMS_PER_PAGE}&pageNum=${page}`;
      const html = await fetchTextWithRetry(ctx, url, { headers: { accept: 'text/html' }, redirect: 'error' });
      const rows = parseHits(html, cfg.origin);
      if (rows.length === 0) break; // past the last page

      let fresh = 0;

View on GitHub (pinned to e7abd431fc)