📁 I/O, Files & Networking · Intermediate

Charsets & encoding in Java

UTF-8, StandardCharsets, platform default pitfalls.

🧩 The mysteryA customer named "José" signs up, and your report prints "José". The database is fine, the code compiles… so who mangled his name?

A charset is a codebook

A charset maps characters to bytes. UTF-8 can encode every Unicode character: plain ASCII takes 1 byte, é takes 2, € takes 3 and most emoji take 4. Note: String.length() counts chars, not bytes.

🔮 Predict it

Characters vs bytes

What does this print?

String s = "€5";
byte[] b = s.getBytes(StandardCharsets.UTF_8);
System.out.println(s.length());
System.out.println(b.length);
  1. 2 2
  2. 2 4
  3. 4 4
Show the answer

2 characters, but 4 bytes: € needs 3 bytes in UTF-8 and 5 needs 1.

Mojibake

Decode bytes with a different charset than they were encoded with and you get garbled text — mojibake. UTF-8 writes é as two bytes; ISO-8859-1 reads each byte as its own character, producing é. Rule: decode with the same charset you encoded with.

🔮 Predict it

Wrong codebook

What does this print?

byte[] b = "é".getBytes(
    StandardCharsets.UTF_8);
String s = new String(b,
    StandardCharsets.ISO_8859_1);
System.out.println(s.length());
System.out.println(s.equals("é"));
  1. 1 false
  2. 2 true
  3. 2 false
Show the answer

2 and true — é's two UTF-8 bytes were decoded as two separate ISO-8859-1 characters: "é", classic mojibake.

The default charset

Since Java 18 (JEP 400) the default charset is UTF-8 on every platform. Before that it depended on the OS — often windows-1252 on Windows. Passing **StandardCharsets.UTF_8** explicitly keeps code correct on any JVM.

String text = Files.readString(path,
    StandardCharsets.UTF_8);
byte[] raw = text.getBytes(
    StandardCharsets.UTF_8);

Reading Arabic text on Java 11 / Windows

✗ Platform default
var r = new FileReader("ar.txt");

Pre-Java 18 this decodes with the OS default (e.g. windows-1252): garbage.

✓ Explicit charset
var r = new FileReader("ar.txt",
    StandardCharsets.UTF_8);

Correct on every JVM and every OS.

🤔 Think first

Is Base64 a charset?

UTF-8, ISO-8859-1, UTF-16, Base64 — which one is not a character encoding?

Think about it, then reveal the answer

Base64. It encodes binary bytes as ASCII text (for emails, JSON, URLs); it doesn't map characters to bytes. The other three are real charsets, all available in StandardCharsets.

💼 In the real world

Encoding bugs in the wild

Mojibake shows up in CSV exports opened in Excel, emails with broken accents, and names like "José" in databases. The professional fix is boring and effective: UTF-8 everywhere, named explicitly at every boundary — files, HTTP headers and database connections.

Key takeaways

  1. String.length() counts chars, not bytes
  2. UTF-8: ASCII is 1 byte; é is 2; most emoji are 4
  3. Decode with the same charset you encoded with
  4. Java 18+ default is UTF-8 (JEP 400); older JVMs used the OS
🤯 Did you know?

UTF-8 was famously sketched out by Ken Thompson and Rob Pike on a placemat in a New Jersey diner in 1992.

Practice questions

What does this print?

String s = "café";
System.out.println(s.length());
System.out.println(
    s.getBytes(StandardCharsets.UTF_8).length);
  1. 4 4
  2. 4 5
  3. 5 5
  4. 5 4
Check your answer

4 5. The string has 4 characters, but é needs 2 bytes in UTF-8, so the encoded form is 5 bytes.

Which one is NOT a character encoding?

  1. UTF-8
  2. ISO-8859-1
  3. UTF-16
  4. Base64
Check your answer

Base64. Base64 encodes binary bytes as ASCII text; it doesn't map characters to bytes. The others are charsets that StandardCharsets provides.

Next: the modern way to talk to the file system — Path and Files.