santifer/career-ops · error · Error

avature: still contains JobDetail links but no article…

Error message

avature: ${url} still contains JobDetail links but no article could be parsed — the listing markup changed

What it means

assertParsedSomething guards the Avature career-site scraper against silent-zero failures. When a fetched listing page still contains posting-shaped /JobDetail/ links but the parser produced no <article class="article--result"> blocks, the listing markup has changed and the provider throws instead of reporting an empty board. This distinction matters because a genuinely empty board (no JobDetail links at all) returns silently.

Solutions

  1. Inspect the fetched HTML at url and update the <article class="article article--result..."> regex in providers/avature.mjs parseArticles to match the new listing markup
  2. Verify the page is a real listing and not a bot-challenge/login interstitial; if so, add the tenant's challenge handling or headers
  3. Pin the entry's api: to the tenant's /careers/SearchJobs URL with current facet params and re-run to confirm the markup hypothesis
  4. If the tenant has permanently moved off Avature's article markup, switch its portals.yml entry to the correct provider

Example fix

// before: parser expects old markup only
const re = /<article class="article article--result[^"]*"[\s\S]*?<\/article>/g;
// after: accept a redesigned result container
class="listing-card job-result"' (adjust regex, e.g. /<article class="[^"]*(?:article--result|job-result)[^"]*"[\s\S]*?<\/article>/g)
Defensive patterns

Strategy: validation

Validate before calling

function looksLikeEmptyAvatureBoard(html) {
  return !/\/JobDetail\/[^"'\s]+/i.test(String(html ?? ''));
}
// if it contains JobDetail links, expect the parser to throw — surface that as a markup alert, not '0 jobs'

Type guard

function isAvatureListingHtml(html) {
  return typeof html === 'string' &&
    /<article class="article article--result[^"]*"[\s\S]*?<\/article>/g.test(html);
}

Try / catch

try {
  const jobs = await provider.fetch(entry, ctx);
} catch (e) {
  if (String(e.message).startsWith('avature:') && e.message.includes('listing markup changed')) {
    alertScrapeRegression(entry); // parser broke; do not treat as empty board
  } else throw e;
}

Prevention

When it happens

Trigger: Called after parsing the first page of an Avature SearchJobs listing; throws when the raw HTML matches /\/JobDetail\// but zero articles were extracted — i.e. Avature or a branded tenant renamed/restructured the article--result markup while still rendering job links.

Common situations: Avature ships a tenant-wide template redesign; a branded tenant (e.g. Siemens-style indexed classes) diverges from the expected article class; a WAF/challenge page injects JobDetail-like URLs without real result articles; scraping an HTML snapshot in tests after the live site changed.

Related errors


AI-assisted analysis of santifer/career-ops@e7abd431fc (2026-09-22). Data as JSON: /api/errors/ed1d4673b2c181b5. Report an issue: GitHub.

Appendix: source

Thrown at providers/avature.mjs:127

}

/**
 * A first page with zero parsed articles is either a genuinely empty board
 * or a markup change breaking the `<article class="article article--result">`
 * selector — those must not look the same to a caller (the whole risk of
 * scraping is a silent-zero failure reading as "no jobs" instead of "the
 * parser broke"). Throws when the raw HTML still carries posting-shaped
 * `JobDetail/` links (the page has rows; the selector just isn't finding
 * them); returns silently for a page carrying neither — a genuinely empty
 * board. Mirrors `providers/itviec.mjs`'s `assertParsedSomething`; called on
 * the first page only (a later empty/short page is just the end of a real
 * board, not a parser regression).
 * @param {string} html
 * @param {string} url
 */
export function assertParsedSomething(html, url) {
  if (!/\/JobDetail\/[^"'\s]+/i.test(String(html ?? ''))) return;
  throw new Error(
    `avature: ${url} still contains JobDetail links but no article could be parsed — the listing markup changed`,
  );
}

/** @param {string} htmlText @param {string} origin */
export function parseArticles(htmlText, origin) {
  const out = [];
  // Tenants vary the result class: Synopsys uses `article--result`, Siemens
  // appends a position index (`article--result 1`). Accept any suffix.
  const re = /<article class="article article--result[^"]*"[\s\S]*?<\/article>/g;
  let a;
  while ((a = re.exec(htmlText)) !== null) {
    const block = a[0];
    // JobDetail path may or may not sit under /careers/ (branded tenants vary),
    // so anchor on JobDetail/ itself rather than a fixed prefix. Prefer the
    // `class="link"` title anchor (most tenants); fall back to any JobDetail
    // anchor for tenants (e.g. Rohde & Schwarz) whose title link carries no
    // class. Share/mailto buttons url-encode the path (%2FJobDetail%2F) so they

View on GitHub (pinned to e7abd431fc)