Turkish Character Repair

Read broken Turkish characters back, or strip them to ASCII for a slug.

An example

Not your input

Broken
Türkçe karakter düzeltme
Best reading
Türkçe karakter düzeltme

The text as pasted scores -15; this reading scores +9. More than one encoding produces it, which is why the text can be certain even when the code page that broke it is not.

Paste it exactly as it arrived, broken and all. Read on this server and nowhere else: nothing is fetched, nothing is stored, and the text is gone when the page finishes rendering.

Two different jobs. Repairing works out what the bytes could have been and gives you every reading with a score. Stripping is a table and guesses nothing: ş becomes s, ğ becomes g, İ becomes I, ı becomes i.

Lower case, and everything outside a to z and 0 to 9 becomes one hyphen. The lower case is Turkish, so ISPARTA gives isparta and İSTANBUL gives istanbul. Left off, the letters are folded and the case is untouched, which is usually what a file name wants. Applies to the second mode only.

There is no table of replacements here. The round trips are performed instead: the text is written back out in the encoding it was wrongly read as, those bytes are read as the encoding it could have been, and each result is scored for how much valid Turkish it yields. A lookup that swaps one broken character for one letter answers the commonest case and nothing else. It cannot see the same mistake made twice, and it has no opinion at all about the letters that went through a Western table. Worked out on this server with mbstring; nothing is fetched and nothing is stored.

What it cannot tell you: which reading is right. It guesses from the bytes it was given, and where the breakage threw information away there is nothing left to recover, because a replacement character is a byte that no longer exists. Two layers of damage are undone; text that was broken three times over comes back still broken, and the scores say so rather than pretending otherwise.

A short string often carries no evidence either way, and the score rewards Turkish letters, so text in another language scores weakly even when the reading is right. Icelandic is where it reads wrongly on purpose: þ, ð and ý are real letters there and the same bytes as ş, ğ and ı here, so this page will offer to turn Icelandic into Turkish. The text as you pasted it is always listed with its own score, which is how you refuse.