SEO

JSON-LD Unescaping: One Pass, and Why Markup Breaks

JSON-LD Unescaping: One Pass, and Why Markup Breaks

Markup that validates in every editor and then shows up as nothing in Search Console is a hard fault to walk a site owner through, because every obvious check passes. The schema is in the page source. The types are correct. The required properties are present. What has gone wrong usually sits a layer below the schema, in how the payload was escaped on its way into the HTML document, and on 21 August 2026 Google narrowed the margin for error on exactly that.

One pass of unescaping, and no more

The change, announced by Google and reported on 21 August 2026, is that Googlebot’s JSON-LD extraction performs a single pass of HTML unescaping and then stops. Double-escaped entities are no longer unrolled. The behaviour that preceded it was more forgiving, in that badly escaped output was quietly corrected on the way in, which meant a good number of sites were shipping malformed JSON and never finding out. Google’s advice alongside the change was to move to standard JSON escapes or Unicode hexadecimal escapes such as \u0026, and Gary Illyes pointed at RFC 8259, section 7, the part of the JSON specification that defines how a string escapes a character. JSON already has an escaping mechanism, and HTML entities are not it.

One pass matters because escaping composes. A value escaped twice needs unescaping twice to come back. Give it one pass and what remains is a string carrying entity text where a character should be, which either fails the parse outright or stores the wrong value in a property that a rich result depends on. This is worth separating from the wider question of what schema markup actually does, because here the markup is right and only the transport is wrong.

Where double escaping comes from

Nobody writes double-escaped output deliberately. It accumulates at a seam between two layers that each believe they are the last one to touch the string:

  • A theme or template function escaping a value that a schema generator had already escaped before handing it over.
  • An editor or importer that escapes on save, so the stored value carries entity text before any template runs.
  • A translation or multilingual layer that round-trips strings through an escaping function on every request.
  • A page builder that stores block attributes as escaped strings and prints them without decoding first.

The pattern is identical in each case, and so is the trap. Because nothing is wrong with the schema itself, a correction applied to the schema will not hold; the next render puts the entities straight back.

JSON on a page is not structured data

Four days later, on 25 August 2026, Illyes made a related point worth holding alongside the first: Google’s crawlers do not parse JSON. They download bytes, and parsing happens downstream, in indexing. That explains why the extraction rule can afford to be strict, since nothing in the crawl path is positioned to repair a malformed document. It also draws a boundary that gets crossed often. A block of JSON in a data attribute, a JSON file fetched by a script, or an API response rendered into the DOM is not structured data. Only JSON-LD inside a script element typed application/ld+json is read as structured data at all. Where that markup is injected client side, whether it survives is really a question about what happens when Google renders your site, which is a different failure with different symptoms.

Reproduce the failure before you change anything

The test that settles it takes a minute. Request the raw HTML with curl rather than reading the browser’s DOM inspector, because the inspector shows a decoded view and hides the fault you are hunting. Pull the contents of the application/ld+json block out of that raw response and feed it to a JSON parser. A parse error says the document is malformed and the problem is entirely yours. A clean parse whose string values still contain visible entity text says the escaping ran one time too many. Only the second looks healthy in a validator that decodes before it checks.

Run it across a sample rather than a single URL. Double escaping tends to be conditional, appearing on titles that contain an ampersand or an apostrophe and nowhere else, which is how it survives in a template for a year. Product names and VideoObject markup descriptions are usually the first casualties, since those are the fields most likely to carry punctuation.

The fix belongs at the output layer

The correction is never in the schema definition, and it is never a find-and-replace across stored content. It belongs where the JSON string is printed into the document, and it consists of removing an escaping step rather than adding one. Google’s guidance on generating structured data with JavaScript makes the neighbouring point that the markup has to be present in the rendered DOM, and the JSON-LD specification is unambiguous that a JSON-LD document is a JSON document first. Encode once, using JSON’s own escapes, and let the HTML layer pass the result through untouched.

Rakibuzzaman Siam
Rakibuzzaman Siam Customer Experience Specialist at Rank Math, building AI automation projects on the side.