Character Encoding Errors

Last reviewed on May 11, 2026

Open a CSV exported from a Windows system in a macOS spreadsheet and the accented letters turn into garbage. Read a UTF-8 source file with a tool that assumes Windows-1252 and the comments fill up with question marks. Paste a copied paragraph from a PDF into a database and discover later that what looked like a hyphen was actually a different character entirely. These are all the same problem with different costumes: the file contains bytes encoded one way, and software is interpreting them another.

This guide is the broad version of the topic. For the narrow TXT case, the TXT encoding guide covers the day-to-day fixes; for CSV, the CSV import guide has format-specific detail. The aim here is to explain what's actually going on, so the fixes make sense.

What encoding actually is

A computer stores text as numbers. An encoding is the rulebook that maps numbers to characters in both directions. For decades, most encodings were small (one byte per character) and regional — Windows-1252 in Western Europe, Shift JIS in Japan, GB2312 in mainland China, ISO-8859-1 in Latin scripts, and many more. Each rulebook used the same byte values to mean different characters.

UTF-8 is the modern default. It's a variable-length encoding of Unicode, which is a unified character set that covers essentially every script in use. UTF-8 has the useful property that it's backward-compatible with plain ASCII for the first 128 characters and unambiguous for everything else. Almost every problem described on this page comes from a file or a tool that hasn't caught up with UTF-8 yet.

The three failure modes

Encoding problems all reduce to one of three patterns:

Detecting the actual encoding

Tools that "auto-detect" encoding are guessing. Their accuracy is high for big, clearly written texts and poor for short or unusual files. A practical approach in order:

  1. Check for a BOM. UTF-8 files sometimes start with EF BB BF, UTF-16 with FF FE or FE FF. If the first bytes match, the encoding is unambiguous. (See the file signatures guide for how to read those bytes.)
  2. Ask the producer. If you control the system that created the file, set its export to UTF-8 explicitly and re-export. Most "encoding problems" disappear at this step.
  3. Open in a tool that shows the detected encoding. Code editors such as VS Code, Sublime Text, and Notepad++ show the current encoding in the status bar and let you re-read the file with a different one.
  4. Use a command-line detector. file -i on macOS and Linux reports a guessed encoding; chardet (Python) and uchardet (C) do the same with confidence values.

A confident wrong guess is the most common failure of automated detection. If a detector says "UTF-8" but the output still looks wrong, the file probably contains a mix of encodings (typical of legacy CSVs that have been edited by several tools in turn). The mix can't be fixed by re-reading; the file has to be normalised.

Converting between encodings

Once the actual encoding is known, conversion is a one-step operation:

Always write the conversion to a new filename. If the conversion is wrong (because the assumed source encoding was wrong), the original file is still there to try again.

The look-alike problem

Some characters are visually identical but encoded differently. The Cyrillic letter а and the Latin letter a look the same in most fonts. ASCII apostrophe ' and the typographic apostrophe ’ render as the same shape in many editors. Hyphen, en dash, em dash, and minus sign all live at different code points. These confuse:

The fix is to normalise. Most languages and editors offer Unicode normalisation (NFC or NFKC) which collapses compatible variants into canonical forms. Where the data needs to remain ASCII (database keys, identifiers, filenames in cross-platform systems), an explicit ASCII-only filter — rejecting or transliterating anything else — prevents the problem at the boundary.

Preventing the problem at the source

Three habits remove most encoding pain before it starts:

Common mistakes

Where this leads

Most encoding problems are diagnosed once and fixed forever by changing the exporting tool's default. For legacy data that's already broken, the conversion steps above are usually enough. When the underlying file is also damaged in some other way — truncated, partial, or containing binary data interleaved with text — see the corrupted files guide for the broader diagnostic flow.