Skip to content
100% local

Language detector

Estimate which language a piece of text is written in, ranked by confidence.

Input
Output

Language detector

Paste any text and this tool estimates which language it is written in, ranking the most likely candidates by a confidence score. It compares the text against a profile of common function words and diagnostic letters for each supported language, so it works on anything from a single sentence to a long document, log excerpt or list of user comments in an unknown language.

The language group option narrows the search to a family — Slavic, Germanic, Romance or other — which helps when you already know roughly what you're dealing with and want a sharper answer. Minimum length sets how many characters the tool expects before it treats a result as reliable; shorter snippets still get a guess, just with a warning attached. Evaluate lets you score the whole text as one unit, or each line or paragraph on its own, useful for text that switches language partway through. Sharpen close languages gives extra weight to letters that are unique to one side of a commonly confused pair, such as Slovak's ľ and ô against Czech's ř and ě, or Croatian's đ against Slovenian. Scores can be shown as percentages or raw values, and the output can be the full ranked list or just the best-guess language code.

Detection is approximate by nature — very short text, code-switched sentences, or languages that share most of their common words will lower confidence, and the tool says so rather than guessing with false certainty. When it suspects a text mixes more than one language, it adds a note instead of silently picking one.

Everything runs locally in your browser; the text you paste is never uploaded anywhere. Copy the result, download it as a .txt file, or send it straight into another tool to continue working on the same text.

FAQ

How accurate is the detection?
It works well on a full sentence or more in one of the supported languages, especially once diacritics or other distinctive letters are present. Very short snippets, numbers-only text, or languages that share most common words score lower and are flagged as less reliable.
Why does it list several languages instead of just one?
The ranking shows how confident the tool is relative to the alternatives. A single dominant language usually scores well above the rest; a close race between two or three means the text is ambiguous or short.
What does "sharpen close languages" actually change?
It adds extra weight to letters that are effectively unique to one language in a confusable pair, such as Slovak versus Czech or Croatian versus Slovenian, which usually widens the gap between them without affecting other languages.
Can it detect languages not in the language group list?
No — it only ranks the languages it has a profile for. Text in an unsupported language will show low, roughly even scores across the board rather than a confident match.
Is my text uploaded anywhere?
No. Detection runs entirely in your browser using a small built-in word and letter profile — your text never leaves your device.