I needed letter frequencies for biasing dataset generation (coming soon!) but I was surprised at the paucity of sources online for Indonesian. I ended up just calculating this myself. Feel free to use this for whatever you'd like.

These were calculated from News (2024), Newscrawl (2016), Web (2018), Web-public (2017), Wikipedia (2021), Mixed (2013) of the Indonesian corpus of the Leipzig Corpora Collection [1]. I used the 1M version of each register, totalling 6,000,000 sentences, 95,262,491 word tokens, and 564,745,272 letters. For simplicity, tokenization here is just whitespace splitting.

LetterMean %Range (pp)LetterMean %Range (pp)
A19.130.60N10.100.32
E8.000.50I7.950.19
T5.150.16R5.100.56
U4.960.14K4.880.52
S4.640.27M4.400.16
D4.060.45G3.750.22
L3.410.22P3.150.42
B2.670.24H2.250.22
O2.030.18Y1.650.24
J1.010.21C0.610.17
W0.490.09F0.340.08
V0.170.07Z0.060.03
X0.0260.01Q0.0180.01
Range = max − min across the six registers.

Digraphs

For fun, I also looked at what the most common digraphs were. Unsurprisingly, ng (like yang, dengan, hingga) dominated at 24.7 occurrences per 1,000 letters, with ny (like hanya, banyak, sebelumnya) second at 7.0 occurrences per 1,000 letters.

References

[1] D. Goldhahn, T. Eckart & U. Quasthoff: Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages.
In: Proceedings of the 8th International Language Resources and Evaluation (LREC 2012), 2012.