Half the emoji list is skin-tone arithmetic
โHow many emoji are thereโ is a question with an unsatisfying answer, because the number depends entirely on what you agree to count. Pick Emoji indexes 3,963 records. That is a real number, taken from Unicodeโs own data, and it is also a misleading one if you read it as โ3,963 different picturesโ.
1,923 of those records are distinct base characters. The other 2,040 are skin-tone variants. Slightly more than half the list โ 51.5% โ is one character wearing a modifier.
Where all 2,040 of them live
They are not spread across the dataset. Every single skin-tone variant in this dataset sits in one group, people-and-body, and that group is almost entirely made of them: 2,430 records, of which only 390 are base characters. No emoji in any of the other eight groups โ not animals-and-nature, not objects, not activities โ takes a tone at all.
Of those 390 people-and-body base characters, 332 accept a skin tone and 58 do not. The 58 are worth looking at, because they are not a random remainder:
- 25 in the family subgroup, plus 5 more filed under person-symbol โ 30 family sequences in all, none of them tone-capable
- 12 body parts โ mechanical arm and leg, brain, anatomical heart, lungs, tooth, bone, tongue, mouth, biting lip, eye, eyes
- 11 in person-symbol โ speaking head, the two silhouette busts, people hugging, footprints, fingerprint, and the generic family glyphs
- 8 fantasy beings โ genie, zombie, troll, hairy creature
- 2 sports โ person fencing and skier, both drawn fully covered
Most of that list makes immediate sense: a bone has no skin tone, a silhouetteโs whole point is that it has none, a fencer is wearing a mask. But the largest single block, the families, is there for a different reason entirely โ and it is the same reason this post is about.
313 characters ร 5, and 19 characters ร 25
Here is the part that produces the number. Of the 332 tone-capable base characters, the split is not gradual. It is two groups, and nothing in between:
- 313 characters take exactly 5 variants each โ light, medium-light, medium, medium-dark and dark. One person in the picture, one modifier. 313 ร 5 = 1,565 records.
- 19 characters take exactly 25 variants each. 19 ร 25 = 475 records.
1,565 + 475 = 2,040, which is the total exactly. There is no third case.
The nineteen are every emoji in the dataset that depicts two people at once:
- handshake
- people, men and women with bunny ears (3)
- people, men and women wrestling (3)
- people holding hands, women holding hands, men holding hands, woman and man holding hands (4)
- kiss, plus its woman-man, man-man and woman-woman forms (4)
- couple with heart, plus the same three forms (4)
Each of the two people gets a tone independently, so the combinations are 5 ร 5 = 25, not 5. Every pairing exists as its own encoded sequence with its own codepoints: there is a light-skinned-and-dark-skinned handshake, and it is a different record from the dark-skinned-and-light-skinned one, because the sequence puts the hands in a fixed order.
And that is where the families come back in. Twenty-five records is the price of two people. Ten of those thirty family sequences depict four people. Tone them independently and one emoji becomes 5โด = 625 records; the three-person families would be 125 apiece. Thirty family sequences expanded that way would add several thousand entries to a list that currently holds 3,963 in total. The families are not tone-capable because the arithmetic that gives you 25 does not stop at 25.
The cost of adding one emoji
This arithmetic has a consequence that is easy to miss when you read a release announcement.
Emoji 18.0 added two one-person hand gestures, a leftwards thumb sign and a rightwards thumb sign. Two characters. They generated twelve records โ the two bases plus five tones each. That is already a 6ร multiplier on what the announcement said.
Had either of them been a two-person gesture, it would have generated twenty-six records on its own. One drawing, twenty-six entries in the list, twenty-five of which an artwork set has to draw separately if it wants full coverage.
This is not hypothetical accounting. It is most of why Blobmoji, a volunteer-maintained set, is 1,347 records short of the dataset: 1,059 of those missing records are skin-tone variants. The base characters are largely drawn. The combinatorial expansion on top of them is what nobody has time for.
Why every variant still gets its own page
Given all that, it would be reasonable to fold the variants into their baseโs page and be done with it. This site does not, for a plain reason: each variant is a separate character sequence, with its own codepoints, that a person can copy and paste and search for. Someone looking for the exact sequence they pasted out of a chat window should land on a page about that sequence, not on a page about a relative of it.
But a variant page and its base page would otherwise be near-identical, so the variants are deliberately thin. A tone page links straight back to its base, and it reuses the baseโs social-card image rather than generating 2,040 more images that differ only in hue โ 2,040 images that would have to be built, stored and served to say almost nothing.
The number to actually remember
When you see a headline count of how many emoji exist, the useful follow-up is not โis that number rightโ. It is โhow many of those are picturesโ.
For this dataset the answer is 1,923 pictures and 2,040 modifiers. Both numbers are true. Only one of them is what most people mean.