Turkish Character Repair
Read broken Turkish characters back, or strip them to ASCII for a slug.
An example
Not your input
- Broken
- Türkçe karakter düzeltme
- Best reading
- Türkçe karakter düzeltme
The text as pasted scores -15; this reading scores +9. More than one encoding produces it, which is why the text can be certain even when the code page that broke it is not.
There is no table of replacements here. The round trips are performed instead: the text is written back out in the encoding it was wrongly read as, those bytes are read as the encoding it could have been, and each result is scored for how much valid Turkish it yields. A lookup that swaps one broken character for one letter answers the commonest case and nothing else. It cannot see the same mistake made twice, and it has no opinion at all about the letters that went through a Western table. Worked out on this server with mbstring; nothing is fetched and nothing is stored.
What it cannot tell you: which reading is right. It guesses from the bytes it was given, and where the breakage threw information away there is nothing left to recover, because a replacement character is a byte that no longer exists. Two layers of damage are undone; text that was broken three times over comes back still broken, and the scores say so rather than pretending otherwise.
A short string often carries no evidence either way, and the score rewards Turkish letters, so text in another language scores weakly even when the reading is right. Icelandic is where it reads wrongly on purpose: þ, ð and ý are real letters there and the same bytes as ş, ğ and ı here, so this page will offer to turn Icelandic into Turkish. The text as you pasted it is always listed with its own score, which is how you refuse.