Charsets & encoding in Java
UTF-8, StandardCharsets, platform default pitfalls.
A charset is a codebook
A charset maps characters to bytes. UTF-8 can encode every Unicode character: plain ASCII takes 1 byte, é takes 2, € takes 3 and most emoji take 4. Note: String.length() counts chars, not bytes.
Characters vs bytes
What does this print?
String s = "€5";
byte[] b = s.getBytes(StandardCharsets.UTF_8);
System.out.println(s.length());
System.out.println(b.length);2 22 44 4
Show the answer
2 characters, but 4 bytes: € needs 3 bytes in UTF-8 and 5 needs 1.
Mojibake
Decode bytes with a different charset than they were encoded with and you get garbled text — mojibake. UTF-8 writes é as two bytes; ISO-8859-1 reads each byte as its own character, producing é. Rule: decode with the same charset you encoded with.
Wrong codebook
What does this print?
byte[] b = "é".getBytes(
StandardCharsets.UTF_8);
String s = new String(b,
StandardCharsets.ISO_8859_1);
System.out.println(s.length());
System.out.println(s.equals("é"));1 false2 true2 false
Show the answer
2 and true — é's two UTF-8 bytes were decoded as two separate ISO-8859-1 characters: "é", classic mojibake.
The default charset
Since Java 18 (JEP 400) the default charset is UTF-8 on every platform. Before that it depended on the OS — often windows-1252 on Windows. Passing **StandardCharsets.UTF_8** explicitly keeps code correct on any JVM.
String text = Files.readString(path,
StandardCharsets.UTF_8);
byte[] raw = text.getBytes(
StandardCharsets.UTF_8);Reading Arabic text on Java 11 / Windows
var r = new FileReader("ar.txt");Pre-Java 18 this decodes with the OS default (e.g. windows-1252): garbage.
var r = new FileReader("ar.txt",
StandardCharsets.UTF_8);Correct on every JVM and every OS.
Is Base64 a charset?
UTF-8, ISO-8859-1, UTF-16, Base64 — which one is not a character encoding?
Think about it, then reveal the answer
Base64. It encodes binary bytes as ASCII text (for emails, JSON, URLs); it doesn't map characters to bytes. The other three are real charsets, all available in StandardCharsets.
Encoding bugs in the wild
Mojibake shows up in CSV exports opened in Excel, emails with broken accents, and names like "José" in databases. The professional fix is boring and effective: UTF-8 everywhere, named explicitly at every boundary — files, HTTP headers and database connections.
Key takeaways
- String.length() counts chars, not bytes
- UTF-8: ASCII is 1 byte; é is 2; most emoji are 4
- Decode with the same charset you encoded with
- Java 18+ default is UTF-8 (JEP 400); older JVMs used the OS
UTF-8 was famously sketched out by Ken Thompson and Rob Pike on a placemat in a New Jersey diner in 1992.
Practice questions
What does this print?
String s = "café";
System.out.println(s.length());
System.out.println(
s.getBytes(StandardCharsets.UTF_8).length);- 4 4
- 4 5
- 5 5
- 5 4
Check your answer
4 5. The string has 4 characters, but é needs 2 bytes in UTF-8, so the encoded form is 5 bytes.
Which one is NOT a character encoding?
- UTF-8
- ISO-8859-1
- UTF-16
- Base64
Check your answer
Base64. Base64 encodes binary bytes as ASCII text; it doesn't map characters to bytes. The others are charsets that StandardCharsets provides.