Japanese Font and Character Pitfalls — Handling JIS2004, Ideographic Variation Selectors, and Gaiji in Business Apps

· Updated: · · Japanese Fonts, JIS2004, Variant Characters, Gaiji, Character Encoding, Unicode, Business Applications, Reports, Windows

Revision history (2 updates, last updated Sep 8, 2026)

A log of the changes made to this article. Where a pre-update version was archived, it stays readable at a permanent DOI link.

Rewrote the search description as a complete sentence: character trouble in business systems sorts itself out once character codes (data) are separated from fonts (appearance). The article body is unchanged. Read the version before this update (DOI: 10.5281/zenodo.22652543)
Retranslated as a full translation of the current Japanese original. The previous English version was an abridgement that dropped subsections, tables, diagrams, and paragraphs; all of them have been restored to match the Japanese article, and the knowledge map section has been added where the Japanese article has one. The technical claims are the same as in the Japanese version. Read the version before this update (DOI: 10.5281/zenodo.22170883)
First published
Cite this article(DOI: 10.5281/zenodo.22170882)

This article is archived on Zenodo. Below are both the DOI that always resolves to the latest version and the DOI pinned to the version you are reading.

Go Komura (2026). Japanese Font and Character Pitfalls — Handling JIS2004, Ideographic Variation Selectors, and Gaiji in Business Apps. KomuraSoft LLC. https://doi.org/10.5281/zenodo.22170882 https://comcomponent.com/en/blog/japanese-fonts-jis2004-ivs-gaiji-business-apps/

DOI (latest version)
10.5281/zenodo.22170882
DOI (this version)
10.5281/zenodo.22652719

“The name is the same, but the shape of the character differs between the screen and the printed form.” “After we replaced the PC, a character that used to display became □.” Business systems that handle Japanese attract inquiries like these.

The first thing to check is whether the character data changed, or whether only the appearance of the same data changed. If you try to fix it as “mojibake” without separating the two, you enter the investigation through the wrong door.

This article starts from two inquiries: a 葛 in a customer list that looks different on screen and on the printed form, and a name on a document submitted to a government office that could no longer be displayed after a PC replacement.

Neither of these two inquiries is the mojibake that an encoding mismatch causes. In the first, not one bit of the data has changed and only the appearance did; in the second, a “gaiji” (an end-user-defined character) that existed only on that PC has been lost.

What the two common inquiries really areThe inquiry that 葛 looks different on screen and on the printed form is a case where only the appearance changed while the data stayed the same, the inquiry that a character became □ after a PC replacement is a case where a gaiji that existed only on that PC was lost, and both are a different problem from encoding-mismatch mojibakeInquiry 1: the shape differs on screen and on the formData unchanged; only the appearance changedInquiry 2: it became □ after the PC was replacedA gaiji that existed only on that PC was lostA different problem from encoding mojibake

Figure 1: The two inquiries that tend to be called “mojibake” are both a different problem from an encoding mismatch.

Separate the character code (the data layer) from the font (the appearance layer), and most Japanese character trouble sorts itself out. This article is written for business-system developers and IT staff, and walks in order from isolating the symptom, through how JIS2004, IVS, and gaiji work, to the range of characters to accept and the design of printed forms and PDFs.

The garbling that happens in conversion between Shift_JIS and UTF-8 is covered in existing articles, so this article concentrates on the problem where the codes round-trip correctly but the appearance, or whether the character can be displayed at all, is off.

1. The Bottom Line First

“Which byte sequence to store” is a data-design question; “how it looks” is a font-design question. When you decide how to respond, separate the following three things.

First, Confirm Whether the Data Is the Same

“Mojibake” is a data-layer problem: a byte sequence being interpreted wrongly. By contrast, the same Unicode code point can have a different glyph in a different font. JIS2004 changed the exemplar glyphs of 168 characters, including 葛, 辻, and 飴, and from Vista onward the JIS2004 glyphs are the default in Windows’ MS Gothic and MS Mincho.12

Do Not Confuse Specifying a Glyph with Carrying Gaiji Around

IVS is the standard means of specifying a glyph as data. It needs a supporting font and a supporting app, however, and in an environment without them the correct behavior is to ignore the selector and display the base character’s default glyph. Because what looks like one character can be up to four UTF-16 code units, it affects not only display but also character counts and substring extraction.34

Gaiji (EUDC, end-user-defined characters) are a separate matter: a Private Use Area number has no meaning shared across the world. The glyph in eudc.tte does not travel to the other party with the data, so a migration needs a survey and a mapping to substitute characters.56

Write the Accepted Characters and the Output Environment into the Specification

A system that handles personal names decides on the character set it accepts and states it explicitly. Where the system exchanges data with government, also keep an eye on the Standard Characters for Administrative Affairs, which build on the Koseki Unified Characters and the Character Information Platform.789 For printed forms and PDFs, the baseline is to use the same font as the screen, confirm the license, and embed the font. For long-term retention, consider PDF/A.1011

Also, do not casually apply NFKC normalization to the original of a personal name. Replacing fullwidth and halfwidth forms and compatibility characters loses distinctions that should be preserved.12

Pick Where to Read from Your Symptom or Goal

What you are struggling with or need to decide What to check first Where to read
Cannot tell whether it is mojibake or a glyph difference Whether the code points are the same. What the difference between “�” and “□” is Chapter 2: Data and appearance
The shape of 葛, 辻, and the like differs before and after a migration, or between screen and form The JIS90/JIS2004 glyph difference and the font in use Chapter 3: JIS2004, Chapter 7: Forms and PDFs
Want to distinguish name glyphs in the data itself Whether IVS support covers display, print, and every downstream system Chapter 4: IVS, Section 4.2: Implementation impact
A character displays only on an old PC Where the Private Use Area is used, and the original gaiji font Chapter 5: Gaiji, Section 5.1: Migration procedure
Need to decide how far to accept names The rules of downstream systems, and the handling of out-of-range characters Chapter 6: Character-set design
Only part of the text changes typeface, or turns into □ Whether the specified font and the fallback font have the glyph Chapter 8: Fallback
Want to check for gaps in an implementation or migration Each layer: input, normalization, storage, display, print, and integration Chapter 9: Checklist

To understand the whole picture, read from Chapter 2 in order; to investigate the symptom in front of you, start from the relevant section in the table. When you implement a countermeasure, check its effect on the other layers in Chapter 9 at the end.

In the diagram a solid line marks a relation that always holds and a dashed line marks a conditional one (the conditions are given per relation on the detail page). The full list of relations (14 in total, with evidence and certainty) and the definitions of the main concepts are collected on the knowledge map detail page (in Japanese). Data: JSON-LD / Turtle

2. Think of Data and Appearance Separately — Code Points and Glyphs

The Number and the Drawn Shape Are Different Things

In Unicode, a character is represented by a number called a code point. 葛 is U+845B, and this number is the same on every PC.

How that number is drawn on a screen or on paper, on the other hand, is decided by the glyph the font holds. It is normal for the same U+845B to differ in the details of its shape between font A and font B.

Separate the Layer to Investigate by Symptom

With these two layers as the premise, the symptoms you meet in the field can be split as follows.

Layer What goes wrong Typical symptoms Main remedy
Data layer (character encoding) Misinterpretation of an encoding, loss during conversion Garbling such as 縺ッ, replacement with ? or 〓, U+FFFD (�) Identify and fix the conversion path
Appearance layer (fonts) Glyph differences between fonts, missing glyphs Same data but a different shape; turns into □ (tofu) Unify or change the font, embed it

As a clue for the split, it helps to remember the difference between “�” and “□”.

“�” is the trace of a failed conversion. Once it has been replaced with U+FFFD (REPLACEMENT CHARACTER), the original character is already lost. Investigate the data layer.

“□” is, in most cases, a display of “glyph not found”. If the data is still there and the font merely lacks a glyph, changing the font may make it displayable. Investigate the appearance layer.

Splitting the symptom by � versus □When a character does not display correctly, � is the trace of a conversion failure at the data layer in which the original character has been lost, while □ means only that the data is still there but the font has no glyph, and changing the font may make it displayable� is visible□ is visibleA character does not display correctlyWhat do you see?Data-layer failureTrace of a failed conversion (the original character is lost)Appearance-layer failureThe font merely lacks a glyphChanging the font may make it displayable

Figure 2: � signals a data-layer failure and □ an appearance-layer failure, and the entry point of the investigation changes accordingly.

The basics of encodings themselves (CP932 and UTF-8, BOM, line endings) are covered in “An Introduction to Windows Text Encodings - The Mojibake That Happens When Integrating with Linux” and “Windows Text Encodings and Line Endings - The Basics of Mojibake and CRLF/LF”. From here on, the subject is the appearance layer and the problems that occur at its boundary.

3. From JIS90 to JIS2004 — The Glyph Changed While the Code Stayed the Same

The real cause behind the opening “葛 looks different on screen and on the form” is, in most cases, found here.

What Changed: The Font’s Default Glyphs

Following the Hyogai Kanji Jitaihyo (the table of character forms for kanji outside the Joyo list) reported by the National Language Council in 2000, the 2004 revision JIS X 0213:2004 (commonly called JIS2004) changed the exemplar glyphs of 168 kanji to the printing-standard forms, which are close to the so-called Kangxi Dictionary forms. 葛, 辻, 飴, 芦, 溢, and 餅 are representative examples.1

Windows followed suit and made the JIS2004 glyphs the default in MS Gothic and MS Mincho (and the newly introduced Meiryo) from Windows Vista onward. Current MS Gothic, too, has JIS2004-based default glyphs, with the JIS90-era glyphs reachable through the OpenType jp90 feature.21

How the glyphs of current MS Gothic are organizedMS Gothic from Vista onward has the JIS2004 glyphs as its default, and the JIS90-era glyphs can be reached through the OpenType jp90 featureMS Gothic (Vista onward)Default glyphs: JIS2004-basedVia the jp90 featureJIS90-era glyphs

Figure 3: Current MS Gothic defaults to the JIS2004 glyphs and can switch to the JIS90 glyphs with the jp90 feature.

What Did Not Change: The Character’s Code Point

What matters here is that only the font changed; the data did not change at all.

  • The code point of 葛 is U+845B on XP and on Windows 11 alike
  • XP (JIS90 glyphs) displays the form in which the inside of the 勹 is simplified to ヒ; Vista onward (JIS2004 glyphs) displays the form that writes 人 inside as well
  • Therefore a scanned image of a form printed by the old system and the screen of the new PC disagree in the shape of the character. A comparison of the data matches completely

Whether the shinnyo radical of 辻 has one dot or two, and the shape of the food radical in 飴, are the same kind of thing. Without knowing this history, an investigation tends to head in the wrong direction: “the migration corrupted the data”.

The order to investigate is: compare the code points before and after the migration → if they match, suspect a glyph difference between fonts. Do not conclude that the data is corrupted merely because it looks different.

The same code point, a different glyph depending on the fontThe code point U+845B of 葛 stays the same on XP and on Windows 11 alike, only the displayed shape differs between a font with JIS90 glyphs and a font with JIS2004 glyphs, and a comparison of the data matches completelyCode point U+845B for 葛Font with JIS90 glyphs (XP)Font with JIS2004 glyphs (Vista onward)Form with the inside of 勹 simplified to ヒPrinting-standard form that writes 人 insideData comparison matches completely

Figure 4: Only the font changed; the code point U+845B stays the same in every environment.

Note that because the character itself did not change, both glyphs are “the same character”. In personal names, however, the person or the government office sometimes insists on a particular shape, and answering the demand to distinguish that “as data” is the job of the next topic, IVS.

4. Ideographic Variation Selectors (IVS) — Specifying a Glyph as Data

This chapter looks separately at the mechanism that specifies a glyph, the conditions under which it can be displayed, and the impact on implementation.

An IVS (Ideographic Variation Sequence) is a mechanism that places an invisible code point called an “ideographic variation selector” immediately after a kanji to specify a glyph variant as data. The selectors used are U+E0100 through U+E01EF (VS17 through VS256).3

The Glyph Mapping Is Decided by IVD Registration

Which “base character + selector” sequence refers to which glyph is decided by a registry called the IVD (Ideographic Variation Database), managed by the Unicode Consortium.

The main collections are as follows.13

Collection Registered Origin and use
Adobe-Japan1 2007 Adobe’s Japanese character collection. The foundation for variant-glyph switching in commercial fonts
Hanyo-Denshi 2010 The General-Purpose Electronic Information Exchange Environment Development Program. Covers government characters such as those of family registers and the Basic Resident Register
Moji_Joho 2014 Corresponds to the Character Information Platform (MJ). Used by IPAmj Mincho. Additional registrations were made in August 2026 as well

For example, Microsoft’s documentation gives U+845B alone (葛) as the form used in the name of Nishi-Kasai Station, and U+845B followed by U+E0100 (VS17) as the form used in the name of Katsuragi City in Nara Prefecture. The same 葛, yet the data can say which glyph is meant.3

An example of distinguishing the same 葛 as data with IVS葛 as U+845B alone is used in the name of Nishi-Kasai Station, the sequence of U+845B followed by VS17 is used in the name of Katsuragi City, and which sequence refers to which glyph is decided by the IVD registryU+845B aloneGlyph used in the name of Nishi-Kasai StationU+845B + VS17Glyph used in the name of Katsuragi CityIVD (registry)

Figure 5: Even for the same 葛, the presence or absence of a selector lets the data say which glyph is meant.

4.1. Behavior in an Environment That Does Not Support It

On the font side, the mapping between an IVS and a glyph is implemented in the OpenType cmap table (format 14).4 When a supporting font (such as IPAmj Mincho) and a supporting app are both present, the specified glyph appears.

When they are not, the specified glyph is not guaranteed to survive. Separate the following two cases.

  • The correct behavior per the specification: the selector is ignored and the base character’s default glyph is displayed (the selector itself is invisible)
  • Older apps and some rendering stacks: the selector is treated as an independent unknown character, and an extra □ is displayed

In other words, IVS is designed so that “even when it degrades, the base character stays readable”, but whether “it always displays in the specified glyph” depends on the receiving environment.

Government resident-record and family-register systems use the combination of a Character Information Platform font plus IVS, but if a general business system accepts it casually, the glyph gets dropped somewhere along display, print, or a downstream system.

How IVS-bearing data is displayedWhen a supporting font and a supporting app are both present it displays in the specified glyph, when they are not the selector is ignored and the base character's default glyph is shown, and in older apps and some rendering stacks the selector is treated as an unknown character and an extra □ is displayedYesNoOlder apps or some rendering stacksBase character + ideographic variation selectorSupporting font and app both present?Displayed in the specified glyphSelector ignored; default glyph displayedAn extra □ is displayedThis is the correct behavior per the specification

Figure 6: IVS stays readable as the base character even when it degrades, but whether the specified glyph appears depends on the receiving environment.

4.2. Implementation Caveats — “One Character” Can Be up to Four Code Units

IVS selectors from U+E0100 onward are supplementary-plane code points, so in UTF-16 they are always a surrogate pair (two code units). If the base character is itself a supplementary-plane kanji (for example 𠮟 (U+20B9F), added in JIS2004), the base alone is two code units, and the sequence a user perceives as “one character” is up to four code units in UTF-16 and up to eight bytes in UTF-8.

Character Counts and Substrings: Do Not Split the Apparent Character

In C#, "葛󠄀" (葛 + VS17) has string.Length == 3. Substring and fixed-length slicing risk separating the base character from its selector.

Validate character counts and extract substrings in grapheme units, with APIs such as StringInfo, not in code units.

Storage: Check the Unit of the Column Length

SQL Server’s nvarchar(n) is measured in UTF-16 code units. If you accept IVS, plan DB column lengths at two to four times the apparent character count.

Search and Comparison: Decide Whether Selectors Are Distinguished

The presence or absence of a selector makes a different string. Whether a search for “葛” should hit “葛 + VS17” is something you have to decide as a requirement and implement.

One IVS-bearing character and UTF-16 code unitsThe sequence of a base character and an ideographic variation selector that a user perceives as one character has a selector that is always a surrogate pair, plus two more code units if the base character is a supplementary-plane kanji, for up to four code units in UTF-16One character as the user perceives itBase characterIdeographic variation selectorTwo code units if it is a supplementary-plane kanjiAlways a surrogate pair (two code units)Up to four code units in UTF-16Fixed-length slicing risks splitting them apart

Figure 7: One IVS-bearing character can be up to four UTF-16 code units, so slicing by code unit is dangerous.

5. Gaiji (EUDC) — Characters That Display Only on That PC

Gaiji are a mechanism by which a user assigns a glyph of their own to a code point in the Unicode Private Use Area (PUA: U+E000 through U+F8FF and other ranges). A Private Use Area code point has no meaning shared across the world; the same U+E000 can be assigned a different character on each PC and in each organization.5

The Glyph Lives in the Gaiji Font; Only the Number Stays in the Data

On Windows, you create the glyph in the Private Character Editor (eudcedit.exe), and it is saved in a font file called eudc.tte.

This file is installed as a hidden font and associated with each font through the HKEY_CURRENT_USER\EUDC registry key.6 In the Shift_JIS (CP932) era the gaiji range was 0xF040 through 0xF9FC, and conversion to Unicode maps it to the Private Use Area.

A number in the data does not mean the other party has the same glyph. This mechanism gives rise to the following problems.

  • eudc.tte belongs to that PC (that user) and does not travel to the other party with the data
  • The moment the data reaches email, a PDF, the web, or another system, the character turns into □ or looks like a different gaiji on the other side
  • If you forget to move eudc.tte during an OS migration or a PC replacement, “a character that displayed on the old PC no longer displays” happens

This is the real cause behind the second inquiry in the opening.

Why gaiji display only on that PCA glyph created in the Private Character Editor is saved in eudc.tte and associated with fonts through that PC's registry, so when only the Private Use Area code travels to email, a PDF, or another system, it turns into □ or looks like a different characterCreate the glyph in the Private Character EditorSaved in eudc.tteAssociated with fonts through the registryDisplays on that PCeudc.tte does not travel with the dataOnly the Private Use Area code reaches the other partyEmail, PDF, other systemsTurns into □ or looks like a different character

Figure 8: The glyph lives in eudc.tte and only a Private Use Area number stays in the data, so gaiji look broken once they leave the PC.

5.1. A Realistic Answer for a System That Has Already Taken In Gaiji

The problem is when data inherited from a legacy system already contains gaiji. What we recommend on migration engagements is a four-stage procedure: survey → identify → replace → block.

1. Survey: Collect the Numbers in Use and the Original Glyphs

Scan databases and files with a regular expression for the Private Use Area (U+E000 through U+F8FF) and inventory the gaiji codes in use and their counts. Collect eudc.tte from the PCs at each site and check the glyphs.

2. Identify: Build a Substitute-Character Mapping Table

For each gaiji, check whether it can be represented by a regular Unicode character, whether it can be represented with IVS, and whether the Character Information Platform (MJ) has a corresponding character, and build a substitute-character mapping table. In practice, the majority of cases turn out to be nothing more than an old character form that had been created as a JIS gaiji.

3. Replace: Replace the Data According to the Table

Replace the data according to the mapping table. Only when no corresponding character exists at all, keep the character as an image or attach a note to the record in question.

4. Block: Do Not Add New Gaiji

In the new system, reject Private Use Area input in validation and do not create new gaiji.

Procedure for migrating data that contains gaijiInventory the gaiji in use by scanning the Private Use Area and collecting eudc.tte, build a substitute-character mapping table and replace, and in the new system reject Private Use Area input in validation and do not create new gaijiSurvey: scan the Private Use AreaIdentify: build the substitute-character mapping tableReplace: replace according to the tableBlock: create no new gaijiCollect eudc.tte from each siteImage or note only where no mapping exists

Figure 9: Migrate gaiji in four stages, survey, identify, replace, and block, and create no new gaiji.

The direction is the same on the government side: the stated policy is to uniquely identify the gaiji that municipalities have created on their own (said to number about two million characters nationwide) against the Standard Characters for Administrative Affairs described below, and to stop using them.9 “Do not add gaiji; identify them against a standardized character set” is becoming the established migration pattern in both the public and private sectors.

6. The Government Character Platform — From Koseki Unified Characters to Standard Characters for Administrative Affairs

When designing a system that handles personal names, knowing the government’s character platform gives you material for deciding “how far to accept”.

First, Separate the Names and Roles of the Character Platforms

Name Steward Outline
Koseki Unified Characters Ministry of Justice About 56,000 characters organized for the computerization of family registers (koseki). Searchable on the Ministry of Justice site7
Juki-net Unified Characters Japan Agency for Local Authority Information Systems (J-LIS) About 21,000 characters used on the Basic Resident Register Network
Character Information Platform (MJ) Character Information Technology Promotion Council About 60,000 characters used in administrative work, organized and managed under MJ character-glyph names; the IPAmj Mincho font and the MJ character-information list are published. Developed as an IPA project and now transferred to the council8
Standard Characters for Administrative Affairs (MJ+) Digital Agency A character set that extends the Character Information Platform with, among others, family-register characters that cannot be identified against MJ. Standard-conforming systems use this character set for names and the like, with JIS X 0221:2020 as the character encoding9

Separate Exchange Within Government from Exchange with the General Outside World

For municipal core-business systems (standard-conforming systems), the standard specification is a two-tier arrangement: use the Standard Characters for Administrative Affairs for information exchange of names and the like, and exchange within the range of JIS X 0213:2012 with external systems that have no unified exchange rules, such as smartphones.9

This arrangement, “hold a wide character set internally and exchange with the outside within the range a general environment can display”, is itself a useful reference for private-sector systems.

The two-tier exchange of a standard-conforming systemA municipal standard-conforming system uses the Standard Characters for Administrative Affairs for information exchange of names and the like, and exchanges within the range of JIS X 0213:2012 with external systems such as smartphones that have no unified exchange rulesMunicipal standard-conforming systemInformation exchange of names and the likeExchange with external systemsStandard Characters for Administrative AffairsRange of JIS X 0213:2012Counterparts with no rules, such as smartphonesA wide character set is held internally

Figure 10: A two-tier arrangement: government exchange uses the Standard Characters for Administrative Affairs, and external exchange with no rules uses JIS X 0213:2012.

Decide the Range Your Own System Accepts

As practical guidance for a general business system, we recommend the following.

  • Decide the accepted character set and state it in both the specification and the input validation. For example, “the range of JIS X 0213:2012”, “no Private Use Area and no combining characters”, or “IVS not accepted (or accepted, with display guaranteed only in an IPAmj Mincho environment)”
  • Do not accept without limit. A design of “it is Unicode, so anything goes” will break somewhere along display, print, or integration
  • Decide in advance how out-of-range characters are handled. The rule for substituting an alternative representation (a new-style form or katakana) and the wording used to explain it to the person are part of the system specification
  • Where a downstream party such as a government agency or a financial institution has character-set rules, treat those as authoritative and conform to them
Designing and operating the accepted character setDecide the accepted character set and state it in both the specification and the input validation, accept in-range characters, and for out-of-range characters decide the operation in advance, including the rule for substituting an alternative representation and the wording used to explain it to the personYesNoDecide the accepted character setState it in the specificationState it in the input validationIn range?AcceptSubstitute an alternative representationThe wording explained to the person is part of the specification too

Figure 11: State the accepted character set in both the specification and the input validation, and decide the handling of out-of-range characters as well.

7. Choosing and Embedding Fonts — Aligning the Screen and the Printed Form

Here the order of thinking is: confirm that the font exists in the environments you use → use the same font on screen and on the form → confirm the license and embed it in the PDF.

7.1. The Character of the Standard Fonts

Font Availability Character and where to use it
MS Gothic / MS Mincho Standard on Windows Veterans designed for low-resolution screens. Default glyphs are JIS2004-based2. Still in service for keeping compatibility with legacy forms
Meiryo Vista onward A modern screen typeface that assumes ClearType. Appeared together with the Vista-generation JIS2004 transition1
Yu Gothic / Yu Mincho Windows 8.1 onward Has family members shipped on both Windows and macOS, which makes it easier to align the look of documents
BIZ UD Gothic / BIZ UD Mincho Windows 10 1809 onward Universal-design typefaces by Morisawa. The first candidate on projects that prioritize the readability of forms and screens14
Noto Sans JP Installed separately Provided as open source, and easy to bundle on servers or Linux environments and to serve on the web

Check the Environments in Use, Not Just the Font Name

What matters in the choice is not typeface preference but whether the font exists in every environment involved in display, print, and PDF generation.

The Japanese supplemental fonts on Windows 10/11 (BIZ UD and others) may not be installed depending on the configuration, and in a configuration that generates PDFs on the server side, whether the font is present on the server has a direct effect.

The environments to check when choosing a fontWhen choosing a font, what matters is not typeface preference but whether the font exists in every environment involved in display, print, and PDF generation, and the configuration of supplemental fonts and whether the font is present on the server have a direct effectCandidate fontPresent in every environment?Display environmentPrint environmentPDF-generation serverSupplemental fonts may be absent depending on the configurationFont presence on the server has a direct effect

Figure 12: Choose a font not by typeface preference but by whether it is present in every environment for display, print, and PDF generation.

7.2. The Basics of Form Design — Align, Then Embed

Use the Same Font on Screen and on the Form

Specify the same font on screen and on the form. If the fonts differ, the same data can look like a different glyph, and you get the complaint from the opening. A configuration like “Meiryo on screen, MS Mincho on the form” should at least be checked for glyph differences in the 168 JIS2004 characters.

Embed in the PDF, After Confirming the License

Embed the font in the PDF. If you do not, the viewer’s side renders with whatever font it has on hand, and not only the glyphs but the layout can change.

Whether embedding is allowed is decided by the license. An OpenType font declares its embedding permissions in the fsType field (Installable, Restricted, Preview & Print, Editable, no subsetting, and so on), and you must not embed a font whose embedding is not permitted.10 For commercial fonts, checking the contract is mandatory.

Make subset embedding the default. If you embed only the glyphs of the characters used, you do not have to carry a whole Japanese font (several MB to tens of MB).

Consider PDF/A for Long-Term Retention

If long-term retention is a requirement, use PDF/A. PDF/A (ISO 19005) is a standard that makes the resources needed for display self-contained within the file, and font embedding is mandatory.11 It is also the most reliable way to prevent “we opened it ten years later and the glyphs had changed”.

The decision flow for font embeddingBefore embedding a font in a PDF, check the embedding license in fsType, make subset embedding the default if it is permitted, and consider PDF/A, where embedding is mandatory, if long-term retention is a requirementPermittedNot permittedLong-term retention requirementEmbed the font in the PDFEmbedding permitted by fsType?Subset embedding is the defaultMust not embedOnly the glyphs of the characters usedConsider PDF/AFont embedding is mandatory

Figure 13: Embedding presupposes checking the fsType license; subset embedding and PDF/A are the baseline.

How to choose an implementation approach for printing and PDF output is covered in detail in “Printing and PDF Output in Windows Business Apps”.

8. Font Linking and Fallback — The “A Different Font Gets Mixed In” Phenomenon

A character for which the specified font has no glyph is not left blank; the default behavior of modern rendering stacks is to render it with a different font instead.

In GDI, the “font linking” defined in the registry (FontLink\SystemLink) does this; in DirectWrite, WPF, and browsers, “font fallback” does.15

The flow of font linking and fallbackIf the specified font has the glyph it is displayed as is, if not it is rendered with a linked or fallback font instead, and if no glyph exists anywhere it turns into □, but the data is usually still intactYesNoYesNoDisplay a characterDoes the specified font have the glyph?Displayed in the specified fontDoes a linked or fallback font have it?Rendered with a different font insteadThe cause of a mixed typeface feel□ (tofu) is displayedThe data is usually still intact

Figure 14: □ is the trace of a failed fallback; whether substitute rendering succeeds is the fork between “mixed” and “tofu”.

Read the Result of Substitute Rendering from the Symptom

Knowing this mechanism lets you explain the following familiar cases.

  • The typeface feel differs between alphanumerics and Japanese: a Latin font was specified first, so only the Japanese portion is drawn in the linked or fallback Japanese font
  • Only the kanji in a Japanese sentence take on Chinese-style glyphs: the fallback resolved to a Chinese font. This happens easily in web pages and apps that do not pass language information (a lang attribute or a locale) correctly
  • Tofu (□) appears: neither the specified font nor the fallback font has the glyph. In other words, □ is the “trace of a failed fallback”, and the data is usually still intact

Fallback Is No Substitute for Design

Fallback is a safety net; it is not a substitute for choosing the right font from the start.15

In a business app, the sound position is “the main display and print paths are complete with the designed fonts alone, and fallback is insurance against unexpected characters”. For how to think about font selection in a multilingual UI, see also “Localizing WinForms/WPF Apps”.

9. An Implementation Checklist for Business Apps

Finally, the table below summarizes the points to check at each layer, from input through integration. Do not stop at “it displays on screen”; check storage, print, and integration as well.

Layer Typical failure Design and implementation points
Input Environment-dependent characters, IVS-bearing characters, and Private Use Area characters come in from the IME Decide the accepted character set and validate against it. Treating out-of-range input as guidance (offering an alternative representation) rather than an error keeps counter work moving
Normalization Unintended conversions under NFKC, such as ㈱ to (株), fullwidth and halfwidth merged, and ① to 1. Even NFC replaces a CJK compatibility ideograph (for example 神 at U+FA19) with the unified ideograph U+795E Do not apply NFKC to names and addresses. Limit normalization to specific uses (such as generating search keys) and store the original as entered12
Storage Column lengths too short for surrogate pairs and IVS; truncation by code unit Store in UTF-8/UTF-16 and give column lengths headroom in code units. Extract substrings in grapheme units
Display □ because the font has no glyph; glyphs change through fallback Explicitly specify a font that can display the target character set, and check what the target OS ships as standard
Print and PDF Glyph differences between screen and form; substitute rendering on the viewer’s side Use the same font on screen and on the form, and subset-embed it in the PDF after checking the license10
Integration with other systems Shift_JIS (CP932) conversion garbles JIS X 0213 additional kanji, IVS, and gaiji into ? or 〓 State the character encoding and character set in the integration specification. Where CP932 integration remains, implement detection of unconvertible characters and substitution rules

Normalization in particular is the trap that is the very theme of this article: a process applied “with good intentions” that flattens the distinctions between variant characters and between fullwidth and halfwidth. Keep the original as is; process a copy is the principle.

Encoding trouble in CSV integration is covered in detail in “CSV Is Not "Just Text"”.

Keep the original as is and process a copyStore the entered string as the original exactly as entered, apply normalization to a copy limited to uses such as generating search keys, and note that applying NFKC to the original loses the distinctions between variant characters and between fullwidth and halfwidthEntered stringOriginal: stored as enteredCopy: normalized for a limited useGenerating search keys and the likeNFKC on the original flattens the distinctions

Figure 15: Limit normalization to a specific use and apply it to a copy; store the original as entered.

10. Summary

Investigation: Separate Data from Appearance

  • Split character trouble first into the “data layer (character encoding)” and the “appearance layer (fonts)”. � signals a data-layer failure, □ an appearance-layer failure.
  • JIS X 0213:2004 changed the exemplar glyphs of 168 characters, and Windows defaults to the JIS2004 glyphs from Vista onward. 葛, 辻, and 飴 looking different across environments is font history, not data corruption.

Design: Decide How Glyphs Are Specified and What Is Accepted

  • The standard means of fixing a glyph in the data is IVS, but without a supporting font and a supporting app it falls back to the default glyph. Do not forget the implementation impact that one character can be up to four UTF-16 code units.
  • Gaiji (EUDC) are assets specific to that PC and cannot travel with the data. The realistic answer is to inventory them at migration time, replace them through a mapping table to regular characters or IVS, and stop creating new ones.
  • A system that handles personal names decides the accepted character set and states it. The government is standardizing toward the Standard Characters for Administrative Affairs on the foundation of the Koseki Unified Characters and the Character Information Platform, and systems that exchange data with it need to follow that movement.

Implementation and Output: Protect the Original and Align the Display Paths

  • For forms and PDFs, the baseline is “use the same font as the screen, confirm the license, and embed”. Consider PDF/A for long-term retention.
  • NFKC normalization, slicing by code unit, and CP932 conversion are the three big points that quietly break variant characters and gaiji. Make storing the original and processing in grapheme units the rule.

The next time someone tells you “the character is different”, start by asking this: are the code points the same, or different? If they are the same, it is a font problem; if they differ, it is a data problem. This one move keeps you from entering the investigation through the wrong door.

The first question that decides the entry point of the investigationWhen told that a character is different, first compare whether the code points are the same or different, and start the investigation as a font problem if they are the same and as a data problem if they differSameDifferentTold that the character is differentAre the code points the same?Font problemData problem

Figure 16: If the code points are the same, start the investigation as a font problem; if they differ, as a data problem.

KomuraSoft LLC handles the design and investigation of character handling in business systems. We work from both the code layer and the font layer: isolating the cause of symptoms such as “the character differs on screen and on the form” or “a name turned into □ after the migration”, surveying gaiji and building substitute-character tables during migration from legacy systems, designing the accepted character set for systems that handle personal names, and reviewing the font-embedding configuration of forms and PDFs.

References

  1. Morisawa Inc., [JIS X 0213:2004 (JIS2004) Font Glossary](https://www.morisawa.co.jp/culture/dictionary/1927). On JIS X 0213:2004, following the Hyogai Kanji Jitaihyo, changing the exemplar glyphs of 168 kanji to the printing-standard forms (the so-called Kangxi Dictionary forms), and on JIS2004-compliant fonts shipping as standard in Windows Vista.

    ↩ ↩2 ↩3 ↩4

  2. Microsoft Learn, MS Gothic font family. On the MS Gothic family’s default glyphs being JIS2004-based, and on the JIS90 legacy glyphs being accessible through the OpenType ‘jp90’ feature. ↩ ↩2 ↩3

  3. Microsoft Learn, The Unicode standard. On a variation sequence consisting of a base character plus a variation selector (VS1 through VS256, U+FE00 through U+FE0F and U+E0100 through U+E01EF), on the example of using U+845B 葛 versus U+845B followed by U+E0100 (VS17) (Nishi-Kasai Station and Katsuragi City), and on a supporting font being required for display. ↩ ↩2 ↩3

  4. Microsoft Learn, cmap — Character to Glyph Index Mapping Table (OpenType spec). On OpenType fonts implementing Unicode Variation Sequences in cmap subtable format 14, on the distinction between default and non-default UVSs, and on usage examples in JIS2004-compliant fonts. ↩ ↩2

  5. Microsoft Learn, End-User-Defined and Private Use Area Characters. On gaiji (EUDC) and Private Use Area (PUA) characters being defined independently by each user or organization, and on the same code point being assigned differently from computer to computer and therefore able to collide. ↩ ↩2

  6. Microsoft Learn, Character Sets and Fonts. On the PUA (U+E000 through U+F8FF and other ranges) being used for EUDC purposes in Unicode, on creating glyphs in the Private Character Editor, and on EUDC fonts being installed hidden as .tte files and associated with fonts through the HKEY_CURRENT_USER\EUDC registry key. ↩ ↩2

  7. Ministry of Justice, Koseki Unified Character Information: Search Criteria. The official search site for the Koseki Unified Characters provided by the Ministry of Justice. On being able to search the glyphs, readings, and related information of the characters used in family registers. ↩ ↩2

  8. Character Information Technology Promotion Council, Character Information Platform Development Project. On the Character Information Platform (the MJ character glyphs, the MJ character-information list, and the IPAmj Mincho font), developed by IPA with the support of the Ministry of Economy, Trade and Industry and others and covering about 60,000 kanji used in administrative work, now being transferred to the council and published there. ↩ ↩2

  9. Digital Agency, Report of the Study Group on the Operation of Character Requirements in Local Government Information Systems (July 2024). On the gaiji used in municipalities being said to number about two million characters, on the “Standard Characters for Administrative Affairs” (commonly MJ+), an extension of the Character Information Platform, being the character set for names and the like in standard-conforming systems with JIS X 0221:2020 as the character encoding, on using the Standard Characters for Administrative Affairs for information exchange of names and the like and JIS X 0213:2012 for exchange with smartphones and the like, and on the policy of uniquely identifying conventional gaiji against the Standard Characters for Administrative Affairs and no longer using them. ↩ ↩2 ↩3 ↩4

  10. Microsoft Learn, OS/2 — OS/2 and Windows Metrics (OpenType spec). On a font’s fsType field defining the embedding license (Installable / Restricted License / Preview & Print / Editable, the no-subsetting bit, and so on), and on applications not being permitted to embed a font whose embedding is not licensed. ↩ ↩2 ↩3

  11. PDF Association, PDF/A Basics. On PDF/A (ISO 19005) for long-term retention requiring the elements needed to display the document to be contained in the file, with font embedding as the representative mandatory example. ↩ ↩2

  12. Microsoft Learn, Using Unicode Normalization to Represent Strings. On the four Unicode normalization forms NFC/NFD/NFKC/NFKD, and on the KC and KD forms unifying compatibility characters such as fullwidth and halfwidth characters and losing information, which makes them generally unsuitable as the canonical stored form of a string. ↩ ↩2

  13. Unicode Consortium, Ideographic Variation Database. The registry of IVSs based on UTS #37. On collections such as Adobe-Japan1 (2007), Hanyo-Denshi (2010), and Moji_Joho (2014) being registered, and on additional registrations to the Moji_Joho collection in the August 2026 edition as well. ↩

  14. Microsoft Learn, BIZ UDGothic font family. On BIZ UD Gothic, a universal-design typeface by Morisawa, being included as a Japanese supplemental font from Windows 10 version 1809 onward. ↩

Recent articles sharing the same tags. Deepen your understanding with closely related topics.

These topic pages place the article in a broader service and decision context.

This article connects naturally to the following service pages.

Frequently Asked Questions

Common questions about the topic of this article.

Why does the same 葛 character look different depending on the PC or the printed form?
It is most likely a difference in font glyphs, not mojibake. JIS X 0213:2004 (JIS2004) changed the exemplar glyphs of 168 kanji to the printing-standard forms, and Windows, too, made the JIS2004 glyphs the default in MS Gothic, MS Mincho, and other fonts from Vista onward. 葛, 辻, and 飴 are representative examples: the Unicode code point (the data) stays the same, and only the glyph the font holds (the appearance) has changed. Compare the data and it matches; a form image from the XP era and the screen of a new PC disagreeing in shape is the specified behavior. If you want the glyphs to match as well, use the same font on screen and on the form, or specify the glyph with an ideographic variation selector.
If we use ideographic variation selectors (IVS), does that solve every glyph problem in personal names?
It does not. IVS is a mechanism that places a selector from U+E0100 onward immediately after the base character to specify the glyph as data, and the specified glyph is displayed only when a supporting font such as IPAmj Mincho and a supporting app are both present. In an environment without support, the correct behavior is for the selector to be ignored and the base character's default glyph to be displayed; in some environments the selector may also show up as □. Furthermore, one IVS-bearing character can be up to four code units in UTF-16, which affects character counting, substring extraction, and the design of DB column lengths. If you adopt it, confirm the scope of support through display, print, and every downstream system before you use it.
Can a character registered as a gaiji (EUDC) be displayed on another PC or in a PDF?
As a rule, no. A gaiji is a mechanism in which the user registers a glyph in that PC's eudc.tte file at a code point in the Unicode Private Use Area (U+E000 onward), so on another PC the same code point is undefined or a different glyph. It is therefore the fate of gaiji to turn into □ or look like a different character once they reach email, a PDF, or another system. If you already hold data that contains gaiji, the realistic approach at migration is to find every use of the Private Use Area, build a mapping table to regular Unicode characters or ideographic variation selectors, and replace them. Creating new gaiji in a new system should be avoided.
How far should a business system go in accepting the characters of personal names?
The first step is to decide the character set you accept and state it explicitly as a specification. Family registers hold about 56,000 Koseki Unified Characters, and the government's standard-conforming systems are moving toward the Standard Characters for Administrative Affairs, an extension of the Character Information Platform, but a general business system is under no obligation to accept the same level without limit. A realistic design decides a range such as "up to the scope of JIS X 0213" or "no ideographic variation selectors and no Private Use Area", validates at input time, and handles out-of-range characters with an alert or an alternative representation. Only systems that exchange data with government systems or municipalities need to follow the developments in the Standard Characters for Administrative Affairs and the JIS X 0221-based exchange requirements.
How do we make a printed form or PDF show the same characters as the screen?
The basics are to specify the same font on screen and on the form, and to embed the font in the PDF. If the fonts differ, the same data can produce different glyphs, and if the viewer's PC lacks the font, a substitute font is used for rendering and the appearance breaks. Whether embedding is allowed is decided by the font's license (OpenType fsType), so check it yourself rather than leaving it to the report library. Subset embedding, which embeds only the characters used, also keeps the file size down. If long-term retention is a requirement, consider PDF/A, where font embedding is mandatory.

Author Profile

Profile page for the article author.

Go Komura

Representative of KomuraSoft LLC

Focused on Windows software development, technical consulting, and investigations into failures that are difficult to reproduce.

Back to the Blog