Unicode and Grapheme Clusters
MushScript’s character unit is not the code point; it is the extended grapheme cluster. A string’s Length, indexing, slicing, and method surface all count clusters.
One “character” can span multiple code points
print("e\u{301}".Length, "é".Length, "e\u{301}" == "é");
print("甲".Length, "👨👩👧".Length);1 1 False
1 1e plus a combining acute accent is two code points, yet the character a human sees is one cluster, so Length is 1. The precomposed é is likewise one cluster. But the two are not the same value: string equality compares code point sequences exactly, and the language performs no implicit NFC normalization. Emoji work the same way: the family emoji is a long run of code points with a cluster count of 1.
Iterating and indexing by cluster
val s = "aé甲🚀";
for c in s
{
print("cluster", c);
}
print(s.CharAt(1), s.CharAt(3));
print(s.Substr(0, 2));cluster a
cluster é
cluster 甲
cluster 🚀
é 🚀
aéfor-in yields one cluster per step, CharAt(index) fetches the character at a cluster number, and Substr’s start and length count clusters too. CJK characters, emoji, and combining characters all follow the same cluster rules — you never end up holding half a character.
Code points and ASCII helpers
print("hello".CodePointAt(0), "甲".CodePointAt(0));
print('a'.IsAsciiLetter(), '5'.IsAsciiDigit(), 'x'.ToAsciiUpper());104 30002
True True XWhen you need the raw code point, use CodePointAt(index): it returns the code point of the cluster’s first scalar. The ASCII classifications and case conversions hang directly off char: IsAsciiDigit / IsAsciiLetter / IsAsciiUpper / IsAsciiLower / ToAsciiUpper / ToAsciiLower.
Notes
- Interpolation and formatting do not deal in clusters; whatever value goes into a hole is stringified by that value’s rules.
