Skip to Content

Unicode and Grapheme Clusters

MushScript’s character unit is not the code point; it is the extended grapheme cluster. A string’s Length, indexing, slicing, and method surface all count clusters.

One “character” can span multiple code points

print("e\u{301}".Length, "é".Length, "e\u{301}" == "é"); print("甲".Length, "👨‍👩‍👧".Length);
1 1 False 1 1

e plus a combining acute accent is two code points, yet the character a human sees is one cluster, so Length is 1. The precomposed é is likewise one cluster. But the two are not the same value: string equality compares code point sequences exactly, and the language performs no implicit NFC normalization. Emoji work the same way: the family emoji is a long run of code points with a cluster count of 1.

Iterating and indexing by cluster

val s = "aé甲🚀"; for c in s { print("cluster", c); } print(s.CharAt(1), s.CharAt(3)); print(s.Substr(0, 2));
cluster a cluster é cluster 甲 cluster 🚀 é 🚀 aé

for-in yields one cluster per step, CharAt(index) fetches the character at a cluster number, and Substr’s start and length count clusters too. CJK characters, emoji, and combining characters all follow the same cluster rules — you never end up holding half a character.

Code points and ASCII helpers

print("hello".CodePointAt(0), "甲".CodePointAt(0)); print('a'.IsAsciiLetter(), '5'.IsAsciiDigit(), 'x'.ToAsciiUpper());
104 30002 True True X

When you need the raw code point, use CodePointAt(index): it returns the code point of the cluster’s first scalar. The ASCII classifications and case conversions hang directly off char: IsAsciiDigit / IsAsciiLetter / IsAsciiUpper / IsAsciiLower / ToAsciiUpper / ToAsciiLower.

Notes

  • Interpolation and formatting do not deal in clusters; whatever value goes into a hole is stringified by that value’s rules.
Last updated on October 11, 2026