RC RANDOM CHAOS

The Ghost Characters Haunting Unicode Since a 1978 Japanese Typo

· via Hacker News

Original source

A spectre is haunting Unicode

Hacker News →

When Japan’s trade ministry standardized the JIS X 0208 character encoding in 1978, it accidentally enshrined a handful of characters that meant nothing. Dubbed “ghost characters” (幽霊文字), they had no discernible pronunciation, meaning, or verifiable source. A 1997 investigation finally traced most of them back to clerical errors made while compiling the standard from sprawling references like a seven-volume, ~6,000-page gazetteer of Japanese place names. In one case, the character 妛 was born when 山 and 女 were printed separately, physically cut out and pasted onto paper, then photocopied — the seam between the two paper scraps read as an extra stroke and got recorded as part of the glyph.

The investigators, by interviewing the original catalogers, resolved nearly all the mysteries, attributing them to transcription slips and misreadings. Only one character, 彁, resisted explanation entirely; the best guess is that it’s a corrupted reading of 彊, but no specific incident was ever pinned down. Because JIS became the foundation for later Japanese encodings, these fabricated characters were inherited wholesale by Unicode — which then spawned its own ghosts during the CJK unification process.

The episode is a case study in how errors calcify once they enter widely-adopted infrastructure. A few small mistakes went uncaught just long enough to be locked into a standard, and now they sit dormant in the character tables of effectively every computer on the planet — practically impossible to remove, and likely to persist indefinitely.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.