Skip to Content

Char

A char is one extended grapheme cluster, written in single quotes. One English letter is a char, one CJK character is a char, and an emoji or a character composed of a combining pair is a char too.

print('a', 'é'); print('\x41', '\u{1F600}'); print('a' + 1);
a é A 😀 98

Literals and escapes

char literals support the full string escape set (\n \r \t \0 \ \" \' \xHH \u{HEX}), and the content must be exactly one grapheme cluster: '\u{1F600}' is one char; '\r\n' is also one cluster (CR+LF composes into a single cluster). Zero clusters ('') or more than one ('ab') reports MS1008.

Code point arithmetic and conversion

When a char meets integer arithmetic, it takes part as its code point and the result is an integer: the code point of 'a' is 97, so 'a' + 1 yields the integer 98, no longer a char.

char and integer conversions go through the code point: int(c) takes the first code point of the cluster, char(n) builds the character for a code point (an illegal code point raises MS6000 at runtime):

print(int('A'), char(66)); val first: char = "Mush"[0]; print(first);
65 B M

The string’s unit of length is the char, and indexing always yields one complete glyph ("Mush"[0] gives 'M').

Notes

  • No normalization: equality compares code point sequences exactly, so precomposed and decomposed forms are not equal. Verified: '\u{65}\u{301}' == 'é' is False (letter plus combining acute accent vs precomposed é);
  • The pitfalls of multi-code-point glyphs are collected on the Unicode and Grapheme Clusters page.
Last updated on October 11, 2026