{"record":{"id":"1e9ac712c7b2100e","repo":"santifer/career-ops","slug":"joinup-entry-name-failed-to-parse-next-data","errorCode":null,"errorMessage":"joinup: ${entry.name} failed to parse __NEXT_DATA__ — ${err.message}","messagePattern":"joinup: (.+?) failed to parse __NEXT_DATA__ — (.+?)","errorType":"exception","errorClass":"Error","httpStatus":null,"severity":"error","filePath":"providers/joinup.mjs","lineNumber":59,"sourceCode":"    try { host = new URL(entry.careers_url || '').hostname; } catch { return null; }\n    return /(^|\\.)joinup\\.ch$/i.test(host) ? { url: BROWSE_URL } : null;\n  },\n\n  async fetch(entry, ctx) {\n    // redirect:'error' — BROWSE_URL is pinned to joinup.ch (https); a 3xx must\n    // not be followed to a private/metadata IP (matches every other provider).\n    const html = await ctx.fetchText(BROWSE_URL, { redirect: 'error' });\n    const m = html.match(/<script id=\"__NEXT_DATA__\"[^>]*>([\\s\\S]*?)<\\/script>/);\n    // Fail closed: a missing/unparseable __NEXT_DATA__ is a scraper break, not an\n    // empty board — throw so the scan logs it instead of silently reporting zero.\n    if (!m) throw new Error(`joinup: ${entry.name} page is missing __NEXT_DATA__ (structure changed?)`);\n    let hits = [];\n    try {\n      const data = JSON.parse(m[1]);\n      const ir = data?.props?.pageProps?.serverState?.initialResults?.jobs?.results;\n      hits = Array.isArray(ir) && ir[0]?.hits ? ir[0].hits : [];\n    } catch (err) {\n      throw new Error(`joinup: ${entry.name} failed to parse __NEXT_DATA__ — ${err.message}`);\n    }\n    return hits\n      .filter(h => h && h.slug && (h.title || h.headline))\n      .map(h => ({\n        title: h.title || h.headline || '',\n        url: `https://joinup.ch/job/${h.slug}`,\n        company: h.startup || entry.name || '',\n        location: typeof h.location === 'string' ? h.location\n          : (h.location?.name || h.location?.city || ''),\n        postedAt: toEpochMs(h.created),\n      }));\n  },\n};\n","sourceCodeStart":41,"sourceCodeEnd":73,"githubUrl":"https://github.com/santifer/career-ops/blob/aac998c7ed7248ea853b720ceeb1fdbeb322fc5d/providers/joinup.mjs#L41-L73","documentation":"The joinup.ch provider found the __NEXT_DATA__ script but JSON.parse of its contents failed, so it rethrows with the underlying parser message. The script body is expected to be a JSON document whose props.pageProps.serverState.initialResults.jobs.results[0].holds the Typesense hits array; malformed JSON means the payload is truncated, HTML-escaped, or otherwise not parseable. As elsewhere, the provider fails closed rather than reporting an empty board.","triggerScenarios":"Raised in fetch() inside the try block when JSON.parse(m[1]) throws — the __NEXT_DATA__ body is cut off mid-document (truncated response), contains escaped/entities-mangled JSON, or the response is a challenge/error page that happens to include a __NEXT_DATA__-shaped script with non-JSON content.","commonSituations":"Network interruption or CDN edge truncates the multi-hundred-KB hydration payload; a security product injects content into the script body; joinup.ch starts HTML-escaping the JSON (e.g. \"</script>\" escaping changes); an intermediary proxy compresses/transcodes the body in a way the fetcher decodes incorrectly.","solutions":["Log the first ~200 chars of m[1] on failure to see whether the payload is truncated, HTML, or escaped JSON — that distinguishes network truncation from format change","Re-fetch and retry: truncation is usually transient, so wrap ctx.fetchText in a retry with backoff for large payloads","If the JSON contains escaped sequences (e.g. \\u003C), decode them before JSON.parse or adjust the extraction regex to capture the raw body correctly","Check response integrity (Content-Length vs actual bytes) in ctx.fetchText to detect truncation and treat it as a retryable failure","If joinup.ch changed escaping/format permanently, update the extraction in providers/joinup.mjs to match the new serialization"],"exampleFix":"// before: single parse attempt, raw error surfaced\nconst data = JSON.parse(m[1]);\n// after: unescape common HTML-safe sequences and surface payload context\nlet raw = m[1].replace(/\\\\u003C/g, '<').replace(/\\\\u003E/g, '>').replace(/\\\\u0026/g, '&').replace(/\\\\\"/g, '\"');\nlet data;\ntry {\n  data = JSON.parse(raw);\n} catch (err) {\n  throw new Error(`joinup: ${entry.name} failed to parse __NEXT_DATA__ — ${err.message}; head: ${raw.slice(0, 120)}`);\n}","handlingStrategy":"retry","validationCode":"function nextDataBodyLooksLikeJson(raw) {\n  if (typeof raw !== 'string') return false;\n  const t = raw.trim();\n  return t.startsWith('{') && t.endsWith('}');\n}\n// before JSON.parse: if (!nextDataBodyLooksLikeJson(m[1])) treat as truncated and retry;","typeGuard":"function parsesAsNextData(raw) {\n  try {\n    const d = JSON.parse(raw);\n    return !!d && typeof d === 'object'\n      && Array.isArray(d?.props?.pageProps?.serverState?.initialResults?.jobs?.results);\n  } catch { return false; }\n}","tryCatchPattern":"try {\n  const hits = fetchJoinupJobs();\n} catch (err) {\n  if (String(err.message).startsWith('joinup:') && err.message.includes('failed to parse __NEXT_DATA__')) {\n    console.error('joinup __NEXT_DATA__ JSON was malformed (likely truncated) — retrying with backoff');\n    return await withRetry(fetchJoinupJobs, 3);\n  }\n  throw err;\n}","preventionTips":["Retry large hydration payloads — truncation mid-document is the most common cause of JSON.parse failure here","Log a short prefix of the script body on parse failure to distinguish truncation from escaping changes","Handle HTML-safe escaping (\\u003C, \\u0026) before/inside parsing if joinup.ch changes serialization","Verify Content-Length matches received bytes in the fetch layer to detect and auto-retry truncated responses","Cache a known-good __NEXT_DATA__ fixture to detect format drift quickly when errors spike"],"tags":["scraper","json","truncation","html-parsing"],"backgroundTag":"json-parse-error","analyzedSha":"aac998c7ed7248ea853b720ceeb1fdbeb322fc5d","analyzedAt":"2026-09-16T06:35:29.214Z","contentChangedAt":"2026-09-16T06:35:29.214Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}