{"record":{"id":"edf45b5dd39e3462","repo":"stride3d/stride","slug":"while-parsing-a-tag-find-an-incorrect-leading-utf-8-octet","errorCode":null,"errorMessage":"While parsing a tag, find an incorrect leading UTF-8 octet.","messagePattern":"While parsing a tag, find an incorrect leading UTF-8 octet\\.","errorType":"exception","errorClass":"SyntaxErrorException","httpStatus":null,"severity":"error","filePath":"sources/core/Stride.Core.Yaml/Scanner.cs","lineNumber":2181,"sourceCode":"                    throw new SyntaxErrorException(start, mark, \"While parsing a tag, did not find URI escaped octet.\");\n                }\n\n                // Get the octet.\n\n                int octet = (analyzer.AsHex(1) << 4) + analyzer.AsHex(2);\n\n                // If it is the leading octet, determine the length of the UTF-8 sequence.\n\n                if (width == 0)\n                {\n                    width = (octet & 0x80) == 0x00 ? 1 :\n                        (octet & 0xE0) == 0xC0 ? 2 :\n                            (octet & 0xF0) == 0xE0 ? 3 :\n                                (octet & 0xF8) == 0xF0 ? 4 : 0;\n\n                    if (width == 0)\n                    {\n                        throw new SyntaxErrorException(start, mark, \"While parsing a tag, find an incorrect leading UTF-8 octet.\");\n                    }\n                }\n                else\n                {\n                    // Check if the trailing octet is correct.\n\n                    if ((octet & 0xC0) != 0x80)\n                    {\n                        throw new SyntaxErrorException(start, mark, \"While parsing a tag, find an incorrect trailing UTF-8 octet.\");\n                    }\n                }\n\n                // Copy the octet and move the pointers.\n\n                charBytes.Add((byte) octet);\n\n                Skip();\n                Skip();","sourceCodeStart":2163,"sourceCodeEnd":2199,"githubUrl":"https://github.com/stride3d/stride/blob/96fad776d210c221682aac1ccdf4c79dc046fc38/sources/core/Stride.Core.Yaml/Scanner.cs#L2163-L2199","documentation":"Scanner.ScanTagUri decodes URI-escaped octets (%XX) into UTF-8. When an escape's byte should begin a multi-byte UTF-8 sequence, its leading bits must match 0xC0/0xE0/0xF8 masks; a leading octet like 0x80-0xBF or 0xF8+ cannot start a character, so the scanner throws SyntaxErrorException.","triggerScenarios":"A tag contains a URI escape whose decoded first byte is not a valid UTF-8 leading octet, e.g. '!%80abc' or '!%FF...' — produced by ScanTagUri after ScanUriEscapes.","commonSituations":"Tags built by double-encoding or manually assembling bytes of invalid UTF-8; binary data pasted into a tag; encoders that percent-encode raw bytes without validating UTF-8 well-formedness.","solutions":["Fix the tag so its percent-encoded octets form well-formed UTF-8 (encode the Unicode characters, not raw bytes).","Re-encode the tag from the original string with proper UTF-8 percent-encoding (Uri.EscapeDataString).","Remove the invalid escape sequence if the tag content was not intentional.","Catch SyntaxErrorException and show the mark/position in the error message."],"exampleFix":"// before (input.yaml)\ntag: !%C3\n\n// after (input.yaml)\ntag: !%C3%A9  // decodes to 'é', a complete UTF-8 sequence","handlingStrategy":"validation","validationCode":"// C#\n// Ensure tag escapes decode as valid UTF-8 before parsing\nbyte[] bytes = System.Text.RegularExpressions.Regex.Matches(yamlText, \"%([0-9A-Fa-f]{2})\")\n    .Select(m => Convert.ToByte(m.Groups[1].Value, 16)).ToArray();\ntry { var s = System.Text.Encoding.UTF8.GetString(bytes); _ = s.Length; }\ncatch (Exception) { throw new FormatException(\"Tag contains invalid UTF-8 escape sequence.\"); }","typeGuard":null,"tryCatchPattern":"try { deserializer.Deserialize(reader, targetType); }\ncatch (SyntaxErrorException ex) { throw new YamlConfigException($\"Invalid UTF-8 in tag at line {ex.Start.Line}\", ex); }","preventionTips":["Encode Unicode characters (not raw bytes) when building tags: UTF-8 encode, then percent-encode.","Never slice strings byte-wise in the middle of a multi-byte character.","Keep tags ASCII to avoid UTF-8 decoding paths entirely."],"tags":["yaml","utf8","encoding"],"backgroundTag":"yaml-parse-error","analyzedSha":"96fad776d210c221682aac1ccdf4c79dc046fc38","analyzedAt":"2026-09-14T02:59:31.279Z","contentChangedAt":"2026-09-14T02:59:31.279Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}