Wikiwand AI

CJK Compatibility Ideographs

Unicode character block From Wikipedia, the free encyclopedia

CJK Compatibility Ideographs is a Unicode block created to contain mostly Han characters that were encoded in multiple locations in other established character encodings, in addition to their CJK Unified Ideographs assignments, in order to retain round-trip compatibility between Unicode and those encodings. However, it also contains 12 unified ideographs sourced from Japanese character sets from IBM.

RangeU+F900..U+FAFF
(512 code points)
PlaneBMP
ScriptsHan
Assigned472 code points
Quick facts Range, Plane ...
CJK Compatibility Ideographs
RangeU+F900..U+FAFF
(512 code points)
PlaneBMP
ScriptsHan
Assigned472 code points
Unused40 reserved code points
Source standardsKS X 1001
Big5
IBM 32
JIS X 0213
ARIB STD-B24
KPS 10721-2000
Unicode Version History
1.0.1 (1992)302 (+302)
3.2 (2002)361 (+59)
4.1 (2005)467 (+106)
5.2 (2009)470 (+3)
6.1 (2012)472 (+2)
Unicode documentation
Code chart ∣ Web page
Note: [1][2]
Range was initially part of the Private Use Area in Unicode 1.0.0,[3] and removed from it in Unicode 1.0.1.
Close

The block has dozens of ideographic variation sequences registered in the Unicode Ideographic Variation Database (IVD).[4][5] These sequences specify the desired glyph variant for a given Unicode character.

Character sources

Sources for the original collection of CJK Compatibility Ideographs include:

  • South Korean KS X 1001 (U+F900–U+FA0B, 268 characters; see that page for the explanation)
  • Taiwanese Big5 (U+FA0C–U+FA0D, 2 characters)
  • "IBM 32": 32 Japanese characters from IBM (U+FA0E–U+FA2D; see below)

In ensuing versions of the standard, more characters have been added to the block from:

  • South Korean KS X 1001 (U+FA2E–U+FA2F, 2 characters)
  • Japanese JIS X 0213 (U+FA30–U+FA6A, 59 characters)
  • Japanese ARIB STD-B24 (U+FA6B–U+FA6D, 3 characters)
  • North Korean KPS 10721-2000 (U+FA70–U+FAD9, 106 characters)

The "IBM 32" characters

IBM Japanese double-byte EBCDIC includes several kanji which do not exist in, or do not round-trip from, JIS X 0208. These were included as gaiji in extensions to Shift JIS and EUC-JP from IBM (e.g. code page 942), NEC, the Open Software Foundation, and Microsoft (e.g. Windows code page 932). However, they were not used as a source for the original Unified Repertoire and Ordering (URO). Instead, 32 of the IBM extension kanji, those which had not been included in the URO from other sources, were included in the CJK Compatibility Ideographs block in the range U+FA0E–U+FA2D.

Of these 32 characters:

  • 19 are unifiable with characters in the URO, and are therefore compatibility ideographs in the strict sense.
  • 12 are kokuji characters which are actually unified ideographs (with the Unified_Ideograph property, and which do not change upon normalisation). In spite of their inclusion in the CJK Compatibility Ideographs block and their algorithmically generated character names beginning with "CJK COMPATIBILITY IDEOGRAPH", they are not duplicates of characters in the original CJK Unified Ideographs block in any respect;[6][7] 11 of these 12 are completely non-duplicate, while U+FA23 﨣 CJK COMPATIBILITY IDEOGRAPH-FA23 was later unintentionally duplicated in CJK Unified Ideographs Extension B as U+27EAF 𧺯 CJK UNIFIED IDEOGRAPH-27EAF. They are placed there because they do not have a URO encoding, yet IBM 32 is one of the encodings where duplicate encodings are of concern. All of them are rarely used or are variants of common kanji. They are as follows:
  • U+FA0E 﨎 CJK COMPATIBILITY IDEOGRAPH-FA0E
  • U+FA0F 﨏 CJK COMPATIBILITY IDEOGRAPH-FA0F
  • U+FA11 﨑 CJK COMPATIBILITY IDEOGRAPH-FA11
  • U+FA13 﨓 CJK COMPATIBILITY IDEOGRAPH-FA13
  • U+FA14 﨔 CJK COMPATIBILITY IDEOGRAPH-FA14
  • U+FA1F 﨟 CJK COMPATIBILITY IDEOGRAPH-FA1F
  • U+FA21 﨡 CJK COMPATIBILITY IDEOGRAPH-FA21
  • U+FA23 﨣 CJK COMPATIBILITY IDEOGRAPH-FA23
  • U+FA24 﨤 CJK COMPATIBILITY IDEOGRAPH-FA24
  • U+FA27 﨧 CJK COMPATIBILITY IDEOGRAPH-FA27
  • U+FA28 﨨 CJK COMPATIBILITY IDEOGRAPH-FA28
  • U+FA29 﨩 CJK COMPATIBILITY IDEOGRAPH-FA29
  • Uniquely, (U+FA20 蘒 CJK COMPATIBILITY IDEOGRAPH-FA20) is intended to be encoded as the kyūjitai form of a kokuji which received a separate encoding for a variant that is straightforwardly the (extended) shinjitai form U+8612 蘒 CJK UNIFIED IDEOGRAPH-8612. The URO only encoded the shinjitai form, and uses its stroke count to place it in this position. It is furthermore one variant of the many variants of the jinmeiyō kanji U+8429 萩 CJK UNIFIED IDEOGRAPH-8429 (i.e. Kummerowia). U+FA20 was assigned a normalisation to U+8612, even though the 龜 and 亀 components, while both forms of radical 213, are not usually considered unifiable.[8]

Block

CJK Compatibility Ideographs[1][2][3]
Official Unicode Consortium code chart (PDF)
 0123456789ABCDEF
U+F90x 豈更車賈滑串句龜 龜契金喇奈懶癩羅
U+F91x 蘿螺裸邏樂洛烙珞 落酪駱亂卵欄爛蘭
U+F92x 鸞嵐濫藍襤拉臘蠟 廊朗浪狼郎來冷勞
U+F93x 擄櫓爐盧老蘆虜路 露魯鷺碌祿綠菉錄
U+F94x 鹿論壟弄籠聾牢磊 賂雷壘屢樓淚漏累
U+F95x 縷陋勒肋凜凌稜綾 菱陵讀拏樂諾丹寧
U+F96x 怒率異北磻便復不 泌數索參塞省葉說
U+F97x 殺辰沈拾若掠略亮 兩凉梁糧良諒量勵
U+F98x 呂女廬旅濾礪閭驪 麗黎力曆歷轢年憐
U+F99x 戀撚漣煉璉秊練聯 輦蓮連鍊列劣咽烈
U+F9Ax 裂說廉念捻殮簾獵 令囹寧嶺怜玲瑩羚
U+F9Bx 聆鈴零靈領例禮醴 隸惡了僚寮尿料樂
U+F9Cx 燎療蓼遼龍暈阮劉 杻柳流溜琉留硫紐
U+F9Dx 類六戮陸倫崙淪輪 律慄栗率隆利吏履
U+F9Ex 易李梨泥理痢罹裏 裡里離匿溺吝燐璘
U+F9Fx 藺隣鱗麟林淋臨立 笠粒狀炙識什茶刺
U+FA0x 切度拓糖宅洞暴輻 行降見廓兀嗀﨎﨏
U+FA1x 塚﨑晴﨓﨔凞猪益 礼神祥福靖精羽﨟
U+FA2x 蘒﨡諸﨣﨤逸都﨧 﨨﨩飯飼館鶴郞隷
U+FA3x 侮僧免勉勤卑喝嘆 器塀墨層屮悔慨憎
U+FA4x 懲敏既暑梅海渚漢 煮爫琢碑社祉祈祐
U+FA5x 祖祝禍禎穀突節練 縉繁署者臭艹艹著
U+FA6x 褐視謁謹賓贈辶逸 難響頻恵𤋮舘
U+FA7x 並况全侀充冀勇勺 喝啕喙嗢塚墳奄奔
U+FA8x 婢嬨廒廙彩徭惘慎 愈憎慠懲戴揄搜摒
U+FA9x 敖晴朗望杖歹殺流 滛滋漢瀞煮瞧爵犯
U+FAAx 猪瑱甆画瘝瘟益盛 直睊着磌窱節类絛
U+FABx 練缾者荒華蝹襁覆 視調諸請謁諾諭謹
U+FACx 變贈輸遲醙鉶陼難 靖韛響頋頻鬒龜𢡊
U+FADx 𢡄𣏕㮝䀘䀹𥉉𥳐𧻓 齃龎
U+FAEx
U+FAFx
Notes
1.^As of Unicode version 17.0
2.^Yellow background: CJK unified ideographs (not compatibility ideographs)
3.^Grey areas indicate non-assigned code points

History

The following Unicode-related documents record the purpose and process of defining specific characters in the CJK Compatibility Ideographs block:

More information Version, Final code points ...
Close

See also

References

Related Articles

Timelines

Top Qs

Fact Checks