{"record":{"id":"c7c5058f94043105","repo":"abhigyanpatwari/GitNexus","slug":"source-must-not-contain-markdown-significant-ch","errorCode":null,"errorMessage":"${source} must not contain Markdown-significant characters (` * [ ] < >).","messagePattern":"(.+?) must not contain Markdown-significant characters \\(` \\* \\[ \\] < >\\)\\.","errorType":"validation","errorClass":"GitNexusRcError","httpStatus":null,"severity":"warning","filePath":"gitnexus/src/cli/analyze-config.ts","lineNumber":235,"sourceCode":"    }\n    case 'string': {\n      if (typeof value !== 'string') {\n        throw new GitNexusRcError(`${source} must be a string.`);\n      }\n      const trimmed = value.trim();\n      if (!trimmed) {\n        throw new GitNexusRcError(`${source} must not be empty.`);\n      }\n      assertNoHiddenChars(trimmed, source);\n      // `name` flows into the generated AGENTS.md/CLAUDE.md as `**${name}**` and\n      // inside `gitnexus://repo/${name}/…` code spans, so a Markdown-significant\n      // character would break those spans or inject emphasis/links/HTML into\n      // agent-instruction content (#1996 tri-review P1). `_` is intentionally\n      // allowed (legitimate in repo names; intraword `_` is not emphasis).\n      // embeddingDevice (the other `string`-kind option) only ever holds a\n      // fixed device token, so this guard never rejects a valid value there.\n      if (/[`*[\\]<>]/.test(trimmed)) {\n        throw new GitNexusRcError(\n          `${source} must not contain Markdown-significant characters (\\` * [ ] < >).`,\n        );\n      }\n      return trimmed;\n    }\n    case 'string-array': {\n      // Generic shared validator — `source` already names the config key, so\n      // messages here stay key-agnostic (no fetch-wrapper coupling in the\n      // shared normalizer; #1589/#1852 review F7).\n      if (!Array.isArray(value)) {\n        throw new GitNexusRcError(`${source} must be an array of strings.`);\n      }\n      const names: string[] = [];\n      for (const item of value) {\n        if (typeof item !== 'string') {\n          throw new GitNexusRcError(`${source} entries must all be strings.`);\n        }\n        const trimmed = item.trim();","sourceCodeStart":217,"sourceCodeEnd":253,"githubUrl":"https://github.com/abhigyanpatwari/GitNexus/blob/ac9a4e9abd8fd3058c070b72c23402a4f887929a/gitnexus/src/cli/analyze-config.ts#L217-L253","documentation":"The source text handed to tree-sitter contained embedded NUL bytes (\\0), which break tree-sitter parsing. safeParse replaces every NUL with a space (counting them, and tagging the log with the file label) before calling parser.parse, so parsing proceeds on sanitized text. It is a data-hygiene warning: offsets are preserved (1:1 replacement), but the file itself contains binary garbage or an encoding issue.","triggerScenarios":"sourceText.includes('\\0') at parse entry — files with stray binary bytes (mangled merges, embedded binary blobs), UTF-16 files read as UTF-8 (NUL-filled high bytes), or tooling that writes text with C-string terminators.","commonSituations":"Repos containing pseudo-binary files with text extensions (.ts, .json) that pass extension filters; files transcoded badly between encodings; fixtures generated from binary templates. Column positions survive because each NUL becomes exactly one space.","solutions":["Inspect the named file around the NUL bytes (e.g. grep -P '\\x00' or a hex view) and fix its encoding or truncate the binary prefix","If the file is genuinely binary, exclude it from indexing so it never reaches the parser","For UTF-16 sources, re-encode as UTF-8 before committing","No code change needed — sanitization is automatic and lossless in length"],"exampleFix":"// before: binary-ish file goes straight to the parser\nconst tree = safeParse(parser, sourceText);\n\n// after: pre-sanitize (or exclude) binary content\nconst sourceText = fileText.replaceAll('\\0', ' '); // what safeParse does internally\nconst tree = safeParse(parser, sourceText);","handlingStrategy":"fallback","validationCode":"// before parsing: detect and sanitize (1:1, length-preserving) like safeParse does\nconst sanitizeNuls = (text: string): { text: string; nulCount: number } => {\n  let nulCount = 0;\n  const clean = text.replaceAll('\\0', () => { nulCount += 1; return ' '; });\n  return { text: clean, nulCount };\n};","typeGuard":"const hasEmbeddedNuls = (buf: Buffer | string): boolean =>\n  typeof buf === 'string' ? buf.includes('\\0') : buf.includes(0); // 0 byte = binary suspect","tryCatchPattern":null,"preventionTips":["Filter binary files out of the parse pipeline (extension allowlist + NUL-byte sniff)","Re-encode UTF-16 sources to UTF-8 before committing","Expect this warn on mangled files; offsets stay valid because replacement is 1:1"],"tags":["tree-sitter","nul-byte","sanitization","encoding"],"backgroundTag":"nul-bytes-in-file","analyzedSha":"ac9a4e9abd8fd3058c070b72c23402a4f887929a","analyzedAt":"2026-08-20T23:29:22.980Z","contentChangedAt":"2026-08-20T23:29:22.980Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}