Character Encoding Errors
Last reviewed on May 11, 2026
Open a CSV exported from a Windows system in a macOS spreadsheet and the accented letters turn into garbage. Read a UTF-8 source file with a tool that assumes Windows-1252 and the comments fill up with question marks. Paste a copied paragraph from a PDF into a database and discover later that what looked like a hyphen was actually a different character entirely. These are all the same problem with different costumes: the file contains bytes encoded one way, and software is interpreting them another.
This guide is the broad version of the topic. For the narrow TXT case, the TXT encoding guide covers the day-to-day fixes; for CSV, the CSV import guide has format-specific detail. The aim here is to explain what's actually going on, so the fixes make sense.
What encoding actually is
A computer stores text as numbers. An encoding is the rulebook that maps numbers to characters in both directions. For decades, most encodings were small (one byte per character) and regional — Windows-1252 in Western Europe, Shift JIS in Japan, GB2312 in mainland China, ISO-8859-1 in Latin scripts, and many more. Each rulebook used the same byte values to mean different characters.
UTF-8 is the modern default. It's a variable-length encoding of Unicode, which is a unified character set that covers essentially every script in use. UTF-8 has the useful property that it's backward-compatible with plain ASCII for the first 128 characters and unambiguous for everything else. Almost every problem described on this page comes from a file or a tool that hasn't caught up with UTF-8 yet.
The three failure modes
Encoding problems all reduce to one of three patterns:
- Mojibake. A UTF-8 file is read as a single-byte encoding (or vice versa). One character in the source becomes two or three garbage characters in the display. The classic
éin place oféis UTF-8 read as Windows-1252. - Replacement characters. A tool reading a file in an encoding that has no slot for a character substitutes a placeholder — typically
?,�, or an empty box. The information is permanently lost at that point. - Silent miscoding. A character looks normal but isn't what you expect. Smart quotes, em dashes, non-breaking spaces, and visually identical Unicode look-alikes all cause this. Software comparing strings byte-for-byte fails to match them; database imports may succeed but produce subtly wrong values.
Detecting the actual encoding
Tools that "auto-detect" encoding are guessing. Their accuracy is high for big, clearly written texts and poor for short or unusual files. A practical approach in order:
- Check for a BOM. UTF-8 files sometimes start with
EF BB BF, UTF-16 withFF FEorFE FF. If the first bytes match, the encoding is unambiguous. (See the file signatures guide for how to read those bytes.) - Ask the producer. If you control the system that created the file, set its export to UTF-8 explicitly and re-export. Most "encoding problems" disappear at this step.
- Open in a tool that shows the detected encoding. Code editors such as VS Code, Sublime Text, and Notepad++ show the current encoding in the status bar and let you re-read the file with a different one.
- Use a command-line detector.
file -ion macOS and Linux reports a guessed encoding;chardet(Python) anduchardet(C) do the same with confidence values.
A confident wrong guess is the most common failure of automated detection. If a detector says "UTF-8" but the output still looks wrong, the file probably contains a mix of encodings (typical of legacy CSVs that have been edited by several tools in turn). The mix can't be fixed by re-reading; the file has to be normalised.
Converting between encodings
Once the actual encoding is known, conversion is a one-step operation:
- Code editors: open the file, use "Reopen with encoding" to read it correctly, then "Save with encoding" as UTF-8.
- Command line (macOS, Linux, WSL):
iconv -f WINDOWS-1252 -t UTF-8 input.txt > output.txt. - PowerShell:
Get-Content -Encoding default input.txt | Set-Content -Encoding UTF8 output.txtfor legacy files; adjust the source encoding as needed. - Spreadsheets: use the "Import text" or "Get external data" path rather than File → Open, which lets you choose the source encoding explicitly.
Always write the conversion to a new filename. If the conversion is wrong (because the assumed source encoding was wrong), the original file is still there to try again.
The look-alike problem
Some characters are visually identical but encoded differently. The Cyrillic letter а and the
Latin letter a look the same in most fonts. ASCII apostrophe ' and the typographic
apostrophe ’ render as the same shape in many editors. Hyphen, en dash, em dash, and minus sign
all live at different code points. These confuse:
- Database imports where unique keys are compared byte-for-byte.
- URL handling, where a look-alike domain can be used to mask a phishing target (the IDN homograph problem).
- Source code, where a curly quote pasted from a word processor breaks a string literal.
- Search and replace operations that look right and silently miss matches.
The fix is to normalise. Most languages and editors offer Unicode normalisation (NFC or NFKC) which collapses compatible variants into canonical forms. Where the data needs to remain ASCII (database keys, identifiers, filenames in cross-platform systems), an explicit ASCII-only filter — rejecting or transliterating anything else — prevents the problem at the boundary.
Preventing the problem at the source
Three habits remove most encoding pain before it starts:
- Export as UTF-8 by default. Most modern spreadsheets, databases, and code editors offer this; choose it once in the application's settings and the problem disappears for new files.
- Declare the encoding. HTML uses
<meta charset="utf-8">; XML uses an encoding declaration in the prolog; HTTP and email useContent-Typeheaders. When the format supports a declaration, use it. When software has to guess, it sometimes guesses wrong. - Test the round trip. When a workflow exports text and re-imports it elsewhere, run a sample file with accents, non-Latin characters, smart quotes, and an em dash through the full pipeline. If anything looks wrong at the other end, fix the pipeline once and the rest of the work follows.
Common mistakes
- "Fixing" mojibake by hand. Find-and-replace on
é,è, and friends works for the cases you noticed and leaves the cases you didn't. Convert the encoding properly instead. - Assuming UTF-8 because the file looks fine. If the file contains only ASCII characters, every common encoding renders it correctly; the encoding is decided as soon as a non-ASCII character appears.
- Trusting the file extension. A
.csvcan be encoded in anything; the extension says nothing about the bytes inside. - Re-saving as UTF-8 after mojibake has already happened. The damage was done when the file was first read with the wrong encoding; saving the broken interpretation just freezes the corruption.
- Mixing BOM and no-BOM UTF-8. Some tools require one and refuse the other; if a file works in one application and breaks in another, the BOM is usually the difference.
Where this leads
Most encoding problems are diagnosed once and fixed forever by changing the exporting tool's default. For legacy data that's already broken, the conversion steps above are usually enough. When the underlying file is also damaged in some other way — truncated, partial, or containing binary data interleaved with text — see the corrupted files guide for the broader diagnostic flow.