Advertisement

By spotfreeai.com Team
Productivity

Free AI Translation Tools 2026: Tested for Accuracy

By SpotFreeAI.com · 20 min read

Quick Answer

Free AI translation tools convert text between languages at no cost. In 2026 the strongest free picks are DeepL for European pairs, Google Translate for sheer language coverage, and ChatGPT or Claude when context matters. Accuracy drops sharply for low-resource languages, so check anything that carries real consequences.

The most accurate free AI translation tool in 2026 depends on your language pair, and anyone who gives you a single winner is selling something. DeepL leads on major European pairs. Google Translate covers more languages than anything else by a wide margin. ChatGPT and Claude often read more naturally on Chinese, Japanese and Korean because they weigh context across a whole passage instead of translating sentence by sentence. All of them are free enough for everyday use, and all of them get meaningfully worse the further you go from English and the big European languages.

That last part is the bit worth understanding, because it is where people get burned. Here is what the research actually shows, and how to work out which tool fits what you are translating.

What Are the Best Free AI Translation Tools in 2026?

Four options cover almost every situation, and they are good at genuinely different things.

DeepL is the specialist. It was built for translation and nothing else, and it shows in how natural the output reads on German, French, Spanish, Italian and the other major European languages. The free tier handles a decent chunk of text at a time and includes document upload.

The old knock on DeepL was coverage, and that criticism is out of date. Its supported languages documentation now lists over 180 language entries, with the company putting the Translator figure above 100. Worth knowing the distinction though, because it's where the marketing gets slippery: the flagship next-generation model covers 36 languages, and the wider list runs on older engines. So a language being on the list is not the same as that language getting DeepL's best work.

Google Translate is the generalist, and its advantage is still breadth. It lists around 249 languages, including many that no other free tool touches. Its 2024 expansion added 110 languages in one go, and the regional detail there is genuinely useful: Acehnese, Balinese, Madurese, Minang and Iban from Indonesia, plus Bikol, Hiligaynon, Kapampangan, Pangasinan and Waray from the Philippines. If you work with regional languages rather than just national ones, that list is the reason to keep Google open. It also does camera translation and offline packs, which nothing else matches for travel.

ChatGPT and Claude are not translation tools, which is exactly why they are useful. Because they process a whole passage at once, they handle tone, idiom and ambiguity better than a sentence-level engine. You can also just ask them to explain a choice, or to redo something more formally. If you are unsure which chatbot to reach for, our comparison of ChatGPT, Claude and Gemini breaks down the free tiers. For a full comparison of ChatGPT and Claude against other AI models, see the ChatGPT vs Claude breakdown on WhichAIBest.

Meta's NLLB models sit behind the scenes rather than in a polished consumer app, but they matter because they are the reason low-resource translation improved at all. More on that in a moment.

How Accurate Is Free AI Translation, Really?

Better than it was, and still uneven in ways the marketing never mentions.

The clearest recent benchmark comes from Meta's No Language Left Behind work. Costa-jussà and colleagues, publishing in Nature in 2024, built a single model covering 200 languages and evaluated it across more than 40,000 translation directions using the FLORES-200 benchmark, human raters, and a toxicity detector. Against the previous state of the art it delivered an average 44 percent improvement in translation quality as measured by BLEU.

A 44 percent jump sounds enormous, and it is. But read what it is measured against. The comparison is with earlier machine translation systems, not with a competent human. A large relative gain on a low starting point still leaves you well short of publishable. That distinction is the single most useful thing to hold onto when someone tells you AI translation is solved.

It is also worth knowing what BLEU actually does. It scores how closely output matches a reference translation, word for word. It rewards literal accuracy and it is blind to whether something reads naturally, or lands the right tone, or would embarrass you in front of a client. A tool can win on BLEU and still produce stiff, obviously-machine prose.

The benchmark underneath most of these claims is worth knowing by name, because it explains why the numbers are narrower than they sound. Goyal, Gao, Chaudhary and colleagues built FLORES-101 from just 3,001 sentences pulled from English Wikipedia, professionally translated into 101 languages under a controlled process. It is a genuinely good benchmark and it is still Wikipedia prose. Encyclopedia sentences are clean, declarative and context-light. Your customer complaint, your rental contract, your group chat full of slang and half-finished thoughts are none of those things. Scores earned on Wikipedia do not transfer cleanly to the messy text you actually need translated.

Do LLMs Beat Dedicated Translation Tools Now?

Sometimes, and there's finally a properly independent answer instead of vendor marketing.

Every year the Conference on Machine Translation runs a General Machine Translation shared task, which is the closest thing this field has to a neutral referee. For the 2024 round, documented in the ACL Anthology, organisers covered 11 language pairs with test sets spanning three to five domains each. Alongside the systems that researchers submitted, they pulled in translations from 8 different large language models and 4 commercial online translation providers, so the chatbots and the dedicated engines got graded on identical text.

The evaluation method is the part that matters. Instead of BLEU, they used professional human annotators working with a protocol called Error Span Annotations, where a person marks the exact stretch of text that's wrong and how badly. That's a direct fix for the BLEU problem described above. A machine counting word overlap can't tell you that a sentence is grammatically fine and means the opposite of the original. A trained annotator can.

Two results stand out. English to Spanish came out the easiest pair, with the best systems producing close to flawless translation. And the speech domain was the hardest, largely because errors from the speech recognition step get inherited by the translation. So if you're translating auto-generated captions, you're stacking two error rates on top of each other.

The organisers named the paper "The LLM Era Is Here but MT Is Not Solved Yet." That's a fair summary and a useful one to keep in your head. Chatbots have genuinely joined the top tier for some pairs. Nobody has swept the board.

The 2025 round went further, and the title alone tells you where the field's head is at. Kocmi and 28 co-authors called it "Time to Stop Evaluating on Easy Test Sets", presented at the Tenth Conference on Machine Translation in Suzhou in November 2025. They scaled up hard: 30 language pairs and 60 systems, of which 36 came from participating researchers and 24 were pulled in from large language models and popular online translation providers. Professional annotators again scored them with Error Span Annotation, with a stricter protocol called Multidimensional Quality Metrics used on two pairs.

Two changes matter more than the headline scores. They built a difficulty sampling technique to deliberately select harder source material instead of the clean, easy sentences benchmarks have leaned on for years. And they handed annotators whole documents with multiple paragraphs rather than isolated sentences, which is the only way to catch a system that translates each line competently while losing the thread across a page.

That's the same complaint about Wikipedia-flavoured benchmarks made a section ago, except now it's coming from the people who run the field's main evaluation. When the referees start saying the test was too easy, treat every accuracy claim built on the old test sets as an upper bound rather than an estimate.

What happens when a specialist goes up against a chatbot

Korean is the cleanest test case, because you'll read everywhere that Naver's Papago beats the big engines on it. Papago is free, it's built specifically for Korean, and that reputation is repeated in roughly every tool roundup online. So it should win.

Jeong-Hwa Lee at Dankook University checked. Writing in the International Journal of Contents in 2025, Lee took 590 matched Korean-to-English dialogue samples from the screen adaptation of Pachinko and ran them through both Papago and ChatGPT. ChatGPT got 343 of the 590 correct, or 58.1 percent. Papago got 223, or 37.8 percent. The error breakdown tells you why: Papago produced 301 lexical errors against ChatGPT's 190, with wrong word meanings the single biggest category for both.

Be fair about what that does and doesn't prove. Screen dialogue is tone-heavy, idiomatic, and full of implied context, which is exactly the terrain where a passage-level model has the advantage. Papago would likely close the gap on a menu or a train timetable. But it does puncture the idea that picking the specialist is automatically the safe move, and it's a reminder that the confident rankings in tool roundups are rarely tested against anything.

The other number worth carrying: a separate error analysis of 777 spoken Korean sentences put through Papago found 55.9 percent translated correctly and 44.1 percent incorrectly, with inaccurate meaning accounting for 83.4 percent of the errors. Roughly two in five sentences wrong, from a tool built for that exact language. That's the realistic baseline for free translation, not the marketing one.

Why vendor benchmarks aren't the same thing

Worth knowing how to read the claims you'll run into while shopping. DeepL's own March 2026 testing reports that it won 100 percent of language pair matchups against Google Translate. That may well be accurate on the pairs and text types DeepL chose. But the company ran the test, picked the samples, and published the result, which is a different kind of evidence from professional annotators scoring 12 systems blind on shared test sets.

Here's what happens when someone independent runs the same matchup. A team at the University of Geneva put DeepL and Google Translate head to head on French medical research abstracts and published it in PLOS ONE in 2024. Scored on nine ROUGE metrics, the differences in medians were not statistically significant. Not smaller than advertised. Not significant at all.

So one party reports a clean sweep and a peer-reviewed test on real professional text finds the engines roughly tied. Both can be technically true, because they measured different things on different material. That's the whole lesson. None of this makes vendor numbers worthless. Just discount them the way you'd discount any self-reported score, and weight independent evaluation higher when the two disagree.

What Changed When Google Put Gemini Into Translate?

Google rebuilt the engine underneath Translate, and it's the single biggest change to free translation in the period this guide covers. On 12 December 2025 the company announced that Translate now uses Gemini models to handle text, aimed squarely at the thing statistical engines have always been worst at: idioms, slang and local expressions. Google's own example is "stealing my thunder", which used to come out word by word and now gets read as a phrase.

Read the rollout details before you assume this applies to you, because most coverage skipped them. The text upgrade launched in the United States and India only, and it covers English paired with roughly 20 languages including Spanish, Hindi, Chinese, Japanese and German. If you're translating Thai, Vietnamese, Tagalog or Bahasa, you were not in that first wave. The language practice features went wider, reaching close to 20 more countries including Germany, Sweden and Taiwan, but practice tools are a different product from the translation engine.

Speech is where the bigger jump landed. The December announcement included a beta live translation mode on Android in the US, Mexico and India, running through your headphones across more than 70 languages while trying to hold on to tone, emphasis and cadence. Six months later Google went further and shipped Gemini 3.5 Live Translate on 9 June 2026, a streaming speech-to-speech model that detects over 70 languages and preserves the speaker's intonation, pacing and pitch instead of flattening everything into one synthetic voice. It stays a few seconds behind the person talking rather than waiting for them to finish a sentence. That one rolled out globally on both Android and iOS in the Translate app, and it's free.

So has the speech problem been solved? Nobody independent has checked yet, and that distinction is the whole point of the section above this one. Everything in the last two paragraphs comes from Google announcing Google's work. The WMT shared task that found speech the hardest domain of all was scoring systems from before any of this shipped, so it isn't evidence against the new models. It also isn't evidence for them. Until professional annotators put Gemini 3.5 Live Translate on a shared test set alongside everything else, treat the demos the way you'd treat DeepL's clean sweep against Google: plausible, self-reported, and measured on material the vendor picked.

What you can act on right now is narrower than the headlines suggest. Live speech translation in the Translate app got materially better and costs nothing, so it's worth trying again if you wrote it off a year ago. The text quality upgrade may or may not have reached your language pair yet. And the underlying failure mode has not changed at all: the model still has to recognise speech before it translates it, so a noisy room, a strong accent or an unfamiliar script still breaks step one and step two still translates the mistake with total confidence.

Why Do Southeast Asian Languages Get Worse Results?

Because these models learn from text, and there is far less text in these languages for them to learn from.

The scale of the gap is easy to underestimate. Lovenia and around sixty collaborators, presenting SEACrowd at EMNLP 2024, pointed out that Southeast Asia is home to more than 1,300 indigenous languages and roughly 671 million people, yet almost none of that shows up proportionally in AI training data. Their project pulled together corpora spanning close to 1,000 SEA languages and evaluated models across 36 indigenous languages and 13 tasks, precisely because that evaluation data did not exist before.

What that means in practice is a tier system nobody advertises. Indonesian, Vietnamese and Thai do reasonably well, since they have big populations and a decent web presence. Tagalog, Burmese, Khmer, Lao and the many regional languages beneath them do noticeably worse, and the output can read as stilted or plain wrong even when it looks confident.

There is a second problem stacked on the first. Benkirane and colleagues, also at EMNLP 2024, tested hallucination detection across 16 language directions and found that the methods for catching invented output work considerably better on high-resource languages than low-resource ones. So the languages most likely to produce a bad translation are also the languages where a bad translation is hardest to detect. If you are translating into Burmese, you get worse output and less warning that it went wrong.

The good news is that this is finally being measured properly. The Conference on Machine Translation, the main academic venue where translation systems get benchmarked, is running a dedicated Chinese to Southeast Asian multilingual task for WMT26, covering bidirectional translation between Chinese and seven regional languages: Thai, Vietnamese, Lao, Burmese, Khmer, Indonesian and Malay. It scores efficiency alongside quality, including model size and inference speed, which matters for anyone deploying this cheaply. A separate track evaluates video subtitle translation into Thai, Indonesian and Malay. None of that fixes the data shortage overnight, but sustained attention from the field is how these gaps eventually close, and Lao in particular has been almost entirely absent from large multilingual benchmarks until now.

Has anything actually improved for SEA languages?

Yes, and it changes the practical advice enough to be worth saying plainly. DeepL used to be irrelevant for this region because it simply didn't support the languages. It does now. Its documentation lists Indonesian, Thai, Vietnamese, Malay and Tagalog, with Indonesian, Thai and Vietnamese getting the fuller feature treatment including glossaries and translation memory.

The reason it happened tells you something about how these gaps close. DeepL added Vietnamese and Thai in June 2025 and was unusually direct about why, saying the launch was "a direct response to demand we're hearing from our customers, especially in manufacturing." Not linguistic justice. Supply chains. The company cited roughly 86 million Vietnamese speakers and 40 million Thai speakers as the business case.

That's worth sitting with if you work with a smaller regional language. Commercial demand is what moves these roadmaps, so languages attached to export manufacturing and international trade get support first. Khmer, Lao and Burmese have large speaker populations too, and they wait longer because fewer companies are billing anyone in them.

So the honest update: for Indonesian, Vietnamese, Thai, Malay and Tagalog you now have a real second opinion instead of just Google plus a chatbot. Run your text through both and compare. For anything below that tier, the old advice stands unchanged, because so does the data shortage underneath it.

If you are running a business across the region, our guide to AI tools for small businesses in Southeast Asia covers where this actually bites.

Can you even trust the accuracy scores for these languages?

Less than you'd like, and this is the layer underneath everything above. It should make you cautious about any accuracy claim you read for a lower-tier language, including the ones on this page.

Plenty of multilingual benchmarks get built by taking an English test set and machine-translating it into the target languages. Read that again and the problem is obvious. The technology being graded also helped write the exam. The team behind SEA-HELM said it plainly: automatic translation "often misses the cultural nuances inherent in the target language, and can result in translation errors and biases." They add that benchmarks assembled this way, with little input from the community, raise questions about their cultural authenticity and reliability.

SEA-HELM is the attempt to do it the hard way instead. Susanto and colleagues at AI Singapore and the National University of Singapore, with support from Stanford's Center for Research on Foundation Models, built it with native speakers involved at every stage of dataset planning and construction rather than translating an English original. For the linguistic diagnostics, the examples were handcrafted from scratch by linguists working with native speakers and reviewed round after round.

And here's the number that tells you how thin the ground is. After all that work, SEA-HELM covers five languages: Filipino, Indonesian, Tamil, Thai and Vietnamese. Five, in a region with more than 1,300. Lao, Khmer and Burmese aren't in it, and that isn't neglect so much as arithmetic, because building a proper human-authored test set in each language is slow expensive work that has only recently started.

So when you meet a confident accuracy percentage for a language outside the top tier, the useful question is where the test set came from. If the answer is machine translation, the score is partly measuring the thing it's meant to be grading.

Which Tool Should You Pick for Your Language Pair?

Work backwards from what you are translating and who reads it.

The habit worth building is running the same text through two tools. Where they agree, you are probably fine. Where they diverge sharply, that sentence is where something has gone wrong.

Does It Matter Which Direction You Translate?

Yes, more than almost anything else you can control for free. And most people get it backwards without realising.

Translating into a language you read well is a fundamentally different job from translating out into one you cannot check. Same tool, same text, wildly different odds of catching a problem. The research bears this out in a way that's worth knowing before you send anything.

Sun, Wang and Jia ran a controlled study on this and published it in PLOS ONE in 2025, working with 31 participants across Chinese and English in both directions. They compared machine translation post-editing against translating from scratch, and measured time, editing effort, and final quality.

Three results stand out:

That last point is the one to sit with. Machine help isn't uniformly useful. It's most useful exactly where you needed it least, and it thins out in the direction where you're already weakest.

The same asymmetry shows up at the model level, not just the human one. Translation into English tends to score higher than translation out of English across the WMT evaluations, and the reason isn't mysterious. These models see far more English during training, so generating fluent English is the easy half of the job. Producing natural Thai or Vietnamese is the hard half, and that's the half you're relying on when you send something out.

So what do you do with this?

If you're reading something in a language you don't speak, you're in the good direction. Output lands in a language you can judge, and if a sentence reads oddly you'll notice. Trust your instincts here. Weird phrasing usually means something went wrong upstream.

If you're sending something out, you're in the hard direction and you've lost your own quality check. This is where people get burned, because the output looks confident and you have no way to tell. Keep sentences short and plain. Cut idioms, jokes, and anything clever, since those are the first things to break. Then find someone who reads the target language, even briefly. Ten minutes from a colleague beats any amount of back-translation.

And be honest about which direction actually matters for you. A lot of people optimise their tool choice around reading foreign text when the risky half of their work is the outbound message nobody checks.

Can Free AI Tools Translate Documents, Images and Speech?

All three, free, though the quality drops as you move from typed text to a photo to live speech. Worth knowing which one you're actually asking for, because the failure modes are different.

Documents. Google Translate takes uploaded files and hands back a translated copy with the layout roughly intact. DeepL does the same on its free tier with a monthly cap on how many files you get. Both struggle with anything heavily formatted, so tables, multi-column layouts and slide decks tend to come back scrambled even when the words are fine. For a PDF that started life as a scan, you're really asking for image translation, and you should read the next paragraph instead.

Images and camera. Google Translate's camera mode is still the standout here and nothing else free comes close for travel. Point it at a menu, a road sign or a packet of something in a supermarket and you get an overlay in real time. The chatbots will also read an uploaded photo and translate it, and they're often better at a dense block of text like a letter or a form because they can reason about the whole thing. The catch with any image translation is that it's two systems in a row: text recognition, then translation. Bad handwriting, low light, or an unfamiliar script means the first step fails and the second step confidently translates whatever it thought it saw.

Speech and subtitles. This is the weakest of the three, and there's a measurement to back that up rather than just a hunch. In the WMT24 shared task described earlier, the speech domain came out the hardest of all the domains tested, specifically because errors from the speech recognition step get inherited by the translation. You're stacking two error rates. So live conversation mode is genuinely useful for ordering food or asking directions, and genuinely unreliable for a meeting where the details matter. That measurement predates Google's Gemini speech models, though, and the section on those covers what has actually been independently checked since.

Auto-generated subtitles deserve their own warning. If you turn on automatic captions and then automatic translation, you're watching output that has passed through speech recognition and machine translation, neither of which knew what the other got wrong. It's fine for gist. It is not a transcript.

Does Offline Translation Work as Well as Online?

No. Offline is a real step down in quality, and it is still the thing you want on your phone before you land somewhere without a signal.

Google Translate is the only one of the four that does this properly. You download language packs to the device, and once a pack is on there, typed translation and camera translation both keep working with no connection. The tell that these are cut-down models is in Google's own instructions: it offers the option to upgrade a downloaded language to a higher-quality pack, which only makes sense if the one you got by default is the smaller one.

The reason is physics rather than stinginess. A model that runs on your phone has to fit in phone memory and answer in phone time, and shrinking a translation model costs accuracy. This is a live enough problem that researchers design architectures specifically around it. Tan, Yang, Zhang, Liu, Sun and Liu built one such approach for on-device translation, published in IEEE/ACM Transactions on Audio, Speech, and Language Processing in 2022, reporting gains of up to 1.7 BLEU on WMT14 English to German and 1.8 BLEU on WMT20 Chinese to English while running up to 1.5 times faster than a comparable baseline. You can read it on arXiv. Nobody spends that effort squeezing quality back out of a small model unless the small model was costing them something.

Two practical things follow, and the second one matters more in this region.

Use it the way you would use a phrasebook. Signs, menus, a short exchange with a driver. For anything you would be embarrassed to get wrong, wait until you have a connection and run it through the full model instead.

Can You Use Free AI Translation for Professional Work?

For some professional jobs, yes, and the evidence is better than you'd expect. For anything with legal or clinical consequences, no.

The Geneva study is worth unpacking properly, because it's one of the few times anyone has tested free translation on genuine professional writing rather than benchmark sentences. The researchers took ten abstracts published in 2021 in two high-impact bilingual medical journals, CMAJ and Canadian Family Physician, where a professional English version already existed alongside the French. Then they ran the French through DeepL, Google Translate and CUBBITT and compared all three against that human reference, using nine ROUGE metrics plus fluency ratings from ten separate raters.

Two findings, and they point in slightly different directions. On the automated metrics, nothing separated the three engines to a statistically significant degree. On human fluency ratings, CUBBITT edged ahead of DeepL and Google Translate, and here's the part that surprises people: it scored above the original English text written by the authors. The researchers concluded that French-speaking academics could reasonably use any of the three when preparing an English version of their work.

Read that carefully before you generalise from it. A research abstract is close to the ideal case for machine translation. It's formal, structured, jargon-heavy in predictable ways, and free of idiom or implied meaning. The author is a domain expert who reads English well enough to catch a mangled sentence. That's a draft-then-check workflow with a competent reviewer at the end, which is very different from pasting a contract into a box and sending the output to a client.

So the honest rule of thumb. Free AI translation is genuinely good enough to produce a first draft of formal, technical prose when someone who understands the content reviews it afterwards. It is not good enough to be the last step on anything where a wrong word costs money or harms somebody. Consent forms, tenancy agreements, dosage instructions, safety notices, employment terms: get a human who can be held accountable. The cost gap between a free tool and a certified translator looks huge right up until the moment something goes wrong.

If translation is one piece of a wider freelance or client workflow, our guide to AI tools for freelancers in Southeast Asia covers where to draw the same line on other tasks.

Is It Safe to Paste Confidential Text Into a Free Translator?

Assume it isn't. Free translation is the single easiest way for company information to walk out of a building, and almost nobody thinks of it as a data transfer while they're doing it.

The UK's National Cyber Security Centre puts the principle plainly in its guidance on AI and cyber security. Queries you send to a public AI service are visible to the provider, may be used to develop the service, and sit with a company that could later be acquired by someone with different privacy standards or suffer a breach of its own. Their recommendation is the one to memorise: don't put confidential or sensitive information into queries to public services, and don't submit anything that would cause you problems if it became public.

That advice exists because the failure has already happened at scale. In September 2017 the Norwegian broadcaster NRK reported that text employees had pasted into the free site Translate.com was turning up in ordinary Google searches. The material found included notices of dismissal, plans for workforce reductions and outsourcing, contracts, and passwords. Journalists at the translation industry publication Slator went looking and surfaced more of the same from other organisations, including a bank's staff performance report and medical correspondence complete with names, emails and phone numbers. The Oslo Stock Exchange responded by blocking staff access to online translation sites outright.

The mechanism was mundane, which is what makes it instructive. Submitted text was stored in the cloud so human translators could work on it, and those pages were reachable by search engines. Nobody was hacked. The tool worked as designed and the design was the problem.

So a workable set of rules:

None of this means avoid free translation. It means separate the two jobs. Understanding a foreign-language email about a delivery is fine. Translating a redundancy letter, a patient note, or an unsigned contract is a different act, and the free tools are the wrong place for it.

How Do the Free Limits Compare?

Free tiers move constantly, so treat any specific number you read as a snapshot rather than a fact.

The broad shape as of September 2026: Google Translate has no meaningful cap for normal web and app use, which remains its quiet advantage. DeepL's free tier lets you paste text in the web interface but allows only a small number of document uploads a month, with a file size cap, and that allowance has been cut over time rather than raised. Third-party pricing trackers disagree on the current figure, so check DeepL's own file translation page before you plan around a number. The document allowance is what pushes heavy users toward paying. ChatGPT and Claude cap you by message volume on their free tiers rather than by word count, so long documents mean working in chunks.

There's a live example of why "free" needs checking rather than assuming. DeepL restructured its developer tiers during 2026, and its own changelog confirms API Developer and API Growth as current plans in entries dated June and July 2026. The old API Free plan, the one with 500,000 characters a month, has now closed to new signups, and this is confirmed by DeepL itself rather than just by pricing trackers. Its API plans help page states that API Free and API Pro can no longer be purchased, and that existing Free or Pro subscriptions keep working as before. So if you already had one, nothing breaks. If you didn't, that door is shut.

What replaced it matters more than the closure. The Developer plan gives you one million characters in total, a one-time allowance for building and evaluating, not a monthly refill. Read that twice if you were planning around the old 500,000 a month, because a one-off million runs out and then stops, where the old plan reset every month. Growth is the recurring paid tier if you need translation running continuously, and Enterprise handles large committed volumes.

Either way, the consumer web translator stayed free throughout. If you paste text into deepl.com, none of this touches you. It only matters if you were planning to wire translation into something.

One thing to watch when you go looking for exact numbers. Most of the conflicting figures floating around come from mixing up two different products. The consumer web and app versions are what you use when you paste text into a box, and their limits are per-translation rather than monthly. The developer APIs are a separate thing with separate monthly character allowances, and those are the numbers that usually get quoted in comparison posts. If a figure does not say which one it is describing, it is not much use to you. Check the provider's own pricing page for whichever product you actually plan to use.

If avoiding accounts entirely matters to you, Google Translate and DeepL both work without signing in, while the chatbots increasingly want an email. Our roundup of free AI tools that need no account tracks which ones still let you in cold.

How Do You Spot a Bad Translation in a Language You Cannot Read?

This is the practical skill, and you do not need to speak the language to get most of the way there.

  1. Translate it back. Run the output through a different tool, back into English. Meaning that survives the round trip is usually safe. Meaning that mutates is a flag.
  2. Check the numbers, names and dates. These are where machine translation quietly fails, and they are checkable without knowing a word of the target language.
  3. Watch the length. Output dramatically shorter than the source often means something got dropped. Dramatically longer can mean the model padded.
  4. Be suspicious of fluency. Hallucinated translation reads smoothly. Fluent output is not evidence of accurate output, and confidence is not a quality signal.
  5. Ask a person for anything public. If it goes on a website, a contract or a product label, a native speaker glancing over it costs very little next to the cost of getting it wrong.

Does AI Translation Invent Gender That Was Not in Your Text?

Yes, constantly. And it's the one accuracy failure that the back-translation trick above will never catch, which is why it sits here rather than earlier.

The mechanism is simple. English lets you write "the doctor called" or "my colleague said" without stating anyone's gender. Plenty of languages won't let you do that. Spanish, French, German, Greek, Arabic and Hebrew all force a grammatical choice. So the model has to pick one, and it picks based on what it saw during training.

How badly do the tools actually get this wrong?

Stanovsky, Smith and Zettlemoyer built the first proper test at ACL 2019. Their WinoMT benchmark is 3,888 English sentences, each containing an occupation plus a pronoun that makes the gender unambiguous, translated into ten languages that mark gender grammatically. The professions were chosen against US Bureau of Labor Statistics data so that half the sentences match the stereotype for that job and half deliberately cut against it.

The design is the clever bit. A system that simply guessed the stereotype every single time would score exactly 50 percent. They ran four commercial systems and two academic models through it and found all of them, in their words, significantly prone to gender-biased translation errors.

That was 2019, so the obvious question is whether it got fixed. It didn't.

Mastromichalakis and colleagues at the National Technical University of Athens published GAMBIT in 2025, a much bigger test: 9,805 English texts covering all 436 occupations in the ISCO-08 international job classification, written across five formats from news reports to casual conversation. They pushed it into Greek and French through six systems, including Google Translate, Meta's NLLB models and Claude 3.5 Sonnet.

The results are blunt. Of the 436 occupations, roughly 374 of them, about 86 percent, got assigned one gender in more than 80 percent of translations. For 293 occupations the systems were 100 percent consistent toward a single gender. Nurses, midwives and cleaners came back feminine. Nearly everything else came back masculine, including plenty of jobs that are genuinely mixed in real life.

The comparison against reality is the number worth carrying. The models produced masculine forms nearly twice as often as actual labour statistics would justify, and their output correlated more strongly with gender stereotypes (0.41 to 0.74) than with real employment data (0.21 to 0.65). So this isn't a model quietly reflecting the world as it is. It's a model exaggerating it.

And notice Claude 3.5 Sonnet was in that lineup. Reaching for a chatbot instead of a translation engine doesn't get you out of this one.

Why does back-translation miss it?

Because the evidence disappears on the return trip. Put "the doctor called" into Spanish and you may well get "el médico llamó". Translate that back into English and you get "the doctor called". Clean round trip, no flags raised.

But a gender you never wrote is now sitting in the Spanish version, and the check you just ran was structurally blind to it, because English doesn't carry the marking that would expose it. The length check passes. The numbers and dates check passes. It reads fluently, which the article already warned you is not a quality signal. Every safety net in the section above lets this one through.

So handle it directly instead:

Does AI Translation Get Politeness and Formality Right?

Often not, and this is the error most likely to embarrass you in Asia specifically. Your grammar can be perfect and your meaning intact while the register is completely wrong, so the message reads as either rude or bizarrely stiff. You won't see it. The person receiving it will see nothing else.

The problem is structural. English mostly carries politeness through word choice, so "are you sure?" works with your boss and your flatmate. Korean, Japanese, Vietnamese, Thai and Javanese bake it into the grammar instead. There's no neutral option to fall back on. The engine has to commit to a level of deference on every sentence, and your English source gave it nothing to base that on.

Researchers have a name for this. Nadejde and colleagues at Amazon built CoCoA-MT, a dataset and benchmark for exactly this problem, and published it in Findings of the Association for Computational Linguistics: NAACL 2022. They covered six target languages and showed that models fine-tuned on labelled contrastive data could hit 82 percent accuracy on the formality level they were asked for in domain, dropping to 73 percent on text outside that domain.

Read those numbers the right way round. That's a system specifically built and trained to control formality, and roughly one sentence in four still comes out at the wrong level once you move it to unfamiliar text. The free consumer tools you're pasting into aren't running anything like that.

The field took the problem seriously enough to make it a shared task. The IWSLT 2023 formality control track asked systems to produce both a formal and an informal version of the same source segment, scored with a matched-accuracy metric measuring how often the output actually hit the requested level. The two supervised language pairs were English to Korean and English to Vietnamese, with 400 annotated training segments and 600 test sentences each. Portuguese and Russian ran as zero-shot pairs with no training data at all.

That language choice tells you something. When researchers wanted the hardest, most consequential formality problems to work on, they picked Korean and Vietnamese. If you're translating into either, you're working in a pair the field itself flags as difficult.

Here's what makes this worse than the gender problem above. Back-translation won't catch it. Run a too-casual Korean sentence back into English and you get a perfectly polite English sentence, because English flattens the distinction that was lost. The check that catches other errors is blind to this one by design.

So handle it on the way in:

What If Your Text Mixes Two Languages?

Then you're outside what most translation tools were built for, and the tool you'd normally reach for is probably the wrong one.

Mixed-language text is completely ordinary across this region. Singlish moves between English, Malay, Hokkien and Tamil inside a single sentence. Taglish does the same with Tagalog and English. Plenty of WhatsApp threads in Kuala Lumpur or Jakarta switch language halfway through a message without anyone thinking about it. Researchers call this code-switching, and translation engines have a hard time with it because they were trained on the assumption that a sentence is in one language.

The clearest picture of how badly this can go comes from a 2023 study at the Workshop on Computational Approaches to Linguistic Code-Switching. Zheng-Xin Yong and a large team tested models across seven Southeast Asian languages: Indonesian, Malay, Chinese, Tagalog, Vietnamese, Tamil and Singlish (ACL Anthology).

The results were uneven in a way that matters if you're picking a tool. BLOOMZ and Flan-T5-XXL simply couldn't produce text with phrases from more than one language. ChatGPT did generate fluent, natural Singlish. But on the English and Tamil pair it mostly returned output that was grammatically wrong or meaningless. It also sometimes threw in a language nobody had asked for. The authors' recommendation was blunt: don't use these models for this without extensive human checking.

That fits a broader finding from EMNLP the same year. Ruochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata and Alham Fikri Aji tested multilingual LLMs across four code-switching tasks, including translation, and titled the paper "Multilingual Large Language Models Are Not (Yet) Code-Switchers". Their conclusion was that being multilingual doesn't automatically mean being good at mixed-language text, and that the big models kept losing to fine-tuned models a fraction of their size.

So which tool actually handles mixed input best?

Here's where it gets more useful, because the picture changed quickly.

A 2024 paper at LREC-COLING by Muhammad Huzaifah, Weihua Zheng, Nattapol Chanpaisit and Kui Wu benchmarked six LLMs across seven datasets specifically on translating code-switched input (ACL Anthology, pages 6381 to 6394). GPT-4 and GPT-3.5 both performed strongly against supervised translation models and commercial engines. GPT-4 in particular held up well across different code-switching conditions rather than only the easy ones.

Put those together and you get a genuinely practical rule. Asking a model to write mixed-language text is still unreliable. But feeding it mixed-language text and asking for a clean translation out is the one job where a general chatbot beats a dedicated engine like Google Translate or DeepL, because those engines try to identify one source language and then commit to it.

So if your input is a Singlish voice note transcript, a Taglish customer message, or a group chat that keeps switching, paste it into a chatbot and say what languages are in there. If you want help choosing between them, our comparison of ChatGPT, Claude and Gemini goes through the free tiers.

Two cautions worth carrying. Tamil pairs were the weak spot in the 2023 testing and nothing since has clearly fixed that, so treat Tamil output as needing a human read. And a model that quietly introduces a language you didn't ask for is a failure mode you won't catch unless you know what the output should look like, which is exactly the problem covered further up in the section on spotting bad translations.

What Else Do People Ask About Free AI Translation?

What is the most accurate free AI translation tool?

It depends entirely on your language pair, which is the honest answer nobody wants. DeepL tends to lead on major European pairs. Google Translate covers far more languages than anyone else. For Chinese, Japanese and Korean, general chatbots like ChatGPT and Claude often read more naturally because they handle context better than a sentence-by-sentence translator. Test your own pair rather than trusting a single leaderboard.

Is Google Translate still worth using in 2026?

Yes, for coverage. It lists around 249 languages, and its 2024 expansion added 110 at once including regional languages like Acehnese, Balinese, Iban, Hiligaynon and Waray that nothing else free touches. DeepL has closed much of the gap on national languages, so the case for Google is now less about raw count and more about the regional tail plus camera mode and offline packs.

Why is AI translation worse for Southeast Asian languages?

Training data. Southeast Asia has over 1,300 indigenous languages and roughly 671 million people, but very little of that shows up in the text these models learn from. The SEACrowd team at EMNLP 2024 documented exactly this shortage. Less data means weaker translation, and it hits smaller languages like Burmese, Khmer and Lao hardest. Indonesian, Vietnamese and Thai now get decent support because commercial demand pulled them forward first.

Can AI translation make things up?

Yes, and it is called hallucination. The model produces fluent output that has drifted from your source text. It is a bigger risk in low-resource languages, and it is harder to catch there too. Benkirane and colleagues found at EMNLP 2024 that hallucination detection works noticeably better for high-resource languages than for low-resource ones.

Should you use free AI translation for legal or medical documents?

No, not on its own. Free tools are fine for understanding the gist of something, for travel, and for informal messages. Anything with legal, medical or financial consequences needs a human translator who can be held accountable. Use AI to prepare a draft if you like, then pay someone qualified to check it.

Sources: Costa-jussa M.R., Cross J., Celebi O., Elbayad M., Heafield K., Heffernan K. et al., Scaling neural machine translation to 200 languages, Nature, vol 630, issue 8018, pages 841 to 846, 2024. Lovenia H. et al., SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages, EMNLP 2024. Benkirane K., Gongas L., Pelles S., Fuchs N., Darmon J., Stenetorp P., Adelani D.I., Sanchez E., Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models, Findings of ACL: EMNLP 2024. Susanto Y., Hulagadri A.V., Montalan J.R., Ngui J.G., Yong X.B., Leong W., Rengarajan H., Limkonchotiwat P., Mai Y., Tjhi W.C., SEA-HELM: Southeast Asian Holistic Evaluation of Language Models, arXiv 2502.14301, 2025, AI Singapore and National University of Singapore with support from Stanford CRFM, covering Filipino, Indonesian, Tamil, Thai and Vietnamese, test sets built with native speakers at each stage rather than machine-translated. Goyal N., Gao C., Chaudhary V., Chen P.J., Wenzek G., Ju D., Krishnan S., Ranzato M., Guzman F., Fan A., The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation, Transactions of the Association for Computational Linguistics, 2022. Findings of the WMT24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet, Proceedings of the Ninth Conference on Machine Translation, ACL Anthology 2024.wmt-1.1, 11 language pairs, 8 LLMs and 4 online providers, human evaluation via Error Span Annotations. Kocmi T. et al., Findings of the WMT25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets, Proceedings of the Tenth Conference on Machine Translation, Suzhou, November 2025, ACL Anthology 2025.wmt-1.22, 30 language pairs and 60 systems of which 24 were LLMs and online providers, difficulty sampling and document-level evaluation. WMT26 Chinese to Southeast Asian Multilingual Machine Translation shared task, Conference on Machine Translation. Lee J.H., A Comparative Study of ChatGPT and Papago Korean-to-English Translations of Dialogue from the Adaptation of Pachinko, International Journal of Contents, vol 21, issue 3, 2025, 590 matched sample pairs. National Cyber Security Centre (UK), AI and cyber security: what you need to know, guidance on queries to public AI services. NRK, Translate.com data exposure, reported 3 September 2017, subsequently documented by Slator and CSO Online. Stanovsky G., Smith N.A., Zettlemoyer L., Evaluating Gender Bias in Machine Translation, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, ACL 2019, WinoMT challenge set of 3,888 sentences across ten languages with professions drawn from US Bureau of Labor Statistics data. Menis Mastromichalakis O., Filandrianos G., Symeonaki M., Stamou G., Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms, 2025, GAMBIT dataset of 9,805 English texts covering all 436 ISCO-08 occupations, evaluated into Greek and French across six systems including Google Translate, NLLB and Claude 3.5 Sonnet. DeepL supported languages documentation, developers.deepl.com, over 180 language entries listed including Indonesian, Thai, Vietnamese, Malay and Tagalog. DeepL blog, launch of Vietnamese, Thai and Hebrew, June 2025, 36 languages on the next-generation model and stated speaker populations of approximately 86 million Vietnamese and 40 million Thai. DeepL API changelog and release notes, developers.deepl.com, entries dated 24 June 2026 and 22 July 2026 confirming API Developer and API Growth plans. Reported closure of the DeepL API Free plan to new signups is based on third-party pricing trackers and is not confirmed in DeepL's own changelog. Google Translate Help, What's new in Google Translate: More than 100 new languages, 110 languages added June 2024 including Acehnese, Balinese, Madurese, Minang, Iban, Bikol, Hiligaynon, Kapampangan, Pangasinan and Waray. Google blog, Gemini capabilities bring new translation upgrades to Google Translate, 12 December 2025, text translation live in the United States and India across English and approximately 20 languages, live translation beta on Android in the United States, Mexico and India across more than 70 languages. Google blog, Fluid, natural voice translation with Gemini 3.5 Live Translate, 9 June 2026, streaming speech-to-speech model detecting over 70 languages, preserving intonation, pacing and pitch, global rollout on Android and iOS in the Google Translate app. Neither Gemini translation release has yet been evaluated in an independent shared task. Yong Z.X. et al., Prompting Multilingual Large Language Models to Generate Code-Mixed Texts: The Case of South East Asian Languages, Proceedings of the 6th Workshop on Computational Approaches to Linguistic Code-Switching, CALCS 2023, ACL Anthology 2023.calcs-1.5, seven Southeast Asian languages covering Indonesian, Malay, Chinese, Tagalog, Vietnamese, Tamil and Singlish. Zhang R., Cahyawijaya S., Cruz J.C.B., Winata G.I., Aji A.F., Multilingual Large Language Models Are Not (Yet) Code-Switchers, EMNLP 2023, arXiv 2305.14235, four code-switching tasks including machine translation. Huzaifah M., Zheng W., Chanpaisit N., Wu K., Evaluating Code-Switching Translation with Large Language Models, LREC-COLING 2024, ACL Anthology 2024.lrec-main.565, pages 6381 to 6394, six LLMs benchmarked across seven datasets. Free-tier limits and language counts checked August 2026 and subject to change.