Accept No Imitations! How to Get the Most Out of AI Language Tools
By Amanda Crain
If we’re to use AI for text production, we must understand not only what AI systems can do but also what they cannot do.
The computer scientist and cryptanalyst Alan Turing tackled the question “Can machines think?” in 1950. Instead of redefining “thinking”—a concept that didn’t apply to machines the way people imagined—he reframed the problem, referring to it as “the imitation game” and asking: Can a machine perform a task normally associated with humans to the point where a human cannot tell the difference? Turing predicted there would one day be machines capable of passing what became known as the Turing Test, a test designed to gauge a machine’s ability to exhibit intelligent behavior equivalent to a human.
Turing’s expectations have been fulfilled in many ways. Today, we have systems that use complex algorithms to analyze vast amounts of data, to recognize patterns within it, and to reproduce and vary those patterns. In some situations, AI chatbots, for example, can pass for human, thus passing the Turing Test. Yet the popular expectation that AI could learn to match humans in all things is far from being realized. As Turing recognized, the technology doesn’t “think,” reason, ponder, or judge the way humans do. It operates on mathematical probabilities. Yet these probabilities are of only limited use when applied to the improbable and seemingly illogical nature of human language. Therefore, any AI-generated text can only be a starting point for a piece of writing intended to communicate a truthful and reliable message.
When AI programs are given relevant data and applied to suitable tasks, they’re both fast and immensely useful. Yet the subtleties of human language are hard to quantify. If we’re to use AI for text production, we must understand not only what AI systems can do but also what they cannot do.
When it comes to academic writing, AI is alluring. Instead of staring at a blank screen and having to come up with a structure and ideas, users can now start the writing process by employing prompts to generate a text on the desired subject. It saves a lot of typing. But users need to be aware that AI language processing works by trawling through internet texts to calculate the most likely next word or phrase for a given prompt; currently, its horizon is no broader than that. It plagiarizes, remixes, and regurgitates. It cannot produce anything original or insightful, except by accident. And what it produces may or may not be true. An AI text program can make anything up if, by its mathematical parameters, that thing is probable.
AI solves problems by making countless repetitive rules-driven decisions—something humans aren’t good at. For instance, AI can convert your Harvard-style bibliography into Oxford style in seconds, although you would be wise to check it. But AI has major limitations in processing human language, precisely because much of what we say isn’t probable or predictable. AI language processing systems face an additional drawback: the training data available to them is distorted, not representative.
The Example of AI Translation
Programs like ChatGPT draw on such enormous data pools that it can be impossible to establish from where and on what criteria they assemble text. But in more controlled AI language processing, such as translation, the picture becomes a little clearer. AI translation is, in effect, intensely prompted text generation. This turns every sentence into a case study.
The Benefit of Hindsight
I’ve been using translation systems almost daily for many years. The data sets used in translation memory systems a dozen years ago were tiny compared to the later machine translation and today’s “large language” and “large reasoning” models. But their manageable size did allow the user to gain insights into how automated systems go about the mathematical construction of human language.
Even in the current data deluge, it’s possible to identify language situations in which AI cannot yield anything but a translation that’s misleading or simply nonsense. That’s the very latest point at which a human needs to step in, with human understanding and a big red pen.
Million-to-One Chances Crop Up Nine Times Out of Ten
The patterns of human expression are so subtle and varied that any single, mildly creative or original sentence is, in itself, an extremely improbable collection of words. Think of the predictive text on your phone: the most likely answer isn’t always right, and the right answer is often not very likely.
Depending on the quality and type of text it’s given, AI-based programs like DeepL and Google Translate may be able to turn out a passable translation. They choose the statistically most likely words, lining them up like beads on a string. This reductive approach enables them to deal capably with texts constructed in a typical pattern: “Johannes Kepler was born in 1571 in Weil der Stadt”; “Insert tab A into slot B”; “Once upon a time, there was a king who had three daughters.”
Yet even the stable patterns of biographies, operating instructions, and fairytales are more varied and structurally complex than we realize. Once we depart from very basic patterns—which we do all the time—translation, like all writing, depends on meaning, not math. Without the human capacity for context, translations can go badly astray. Here are a few of many examples. Because of the ever-changing AI data pools, they’re no longer reproducible. But these types of errors will always occur because human language references outside contexts that AI cannot include in its calculations.
In the first example from a German text, a keen new professor picks up a screwdriver and works on some equipment:
- Professor S. hat sich gleich am ersten Tag im Labor die Handschuhe angezogen und mit mir am Strömungskanal herumgeschraubt.
- DeepL rendered this as: Professor S. put on her gloves on the very first day in the laboratory and screwed around with me on the flow channel.
The program reproduced a predictable pattern but couldn’t recognize that to “screw around” has quite another meaning from its innocent literal translation
in German.
Sometimes the right context requires extra definition. In the following example, DeepL gets lost in the woods:
- Jeder Forst ist zugleich auch ein Wald, aber nicht jeder Wald, und wäre er auch noch so groß, ein Forst.
- DeepL’s translation: Every forest is at the same time also a forest, but not every forest, no matter how large it is, is a forest.
Even sein is not always simply to be, as a look at Article 1 of the Basic Law for the Federal Republic of Germany, Germany’s constitution, shows:
- Die Würde des Menschen ist unantastbar.
- DeepL and Google Translate yield variations on: Human dignity is inviolable.
In English, this is a declarative and clearly untrue statement because human dignity is violated every day. Article 1 is a declaration of principle: Human dignity shall be inviolable.
The Calculation Game
If a word in the source language has four possible translations—one occurring 40% of the time and three each occurring with 20% frequency—AI translation will choose the statistically most frequent option. Put simply, it will choose incorrectly 60% of the time when that word occurs.
Star or starling? A bird becomes a celestial body:
- So wie ein großer Schwarm von Vögeln ganz andere Bewegungen am Himmel macht, im Vergleich zu der Bahn eines einzelnen Stars.
- DeepL renders this as: Just as a large flock of birds makes completely different movements in the sky compared to the path of a single star.
Bacteria take the waters:
- Bakterien in heißen Quellen der Tiefsee
- DeepL’s translation: bacteria in hot springs in the deep ocean
- Edited version: bacteria in hydrothermal vents on the ocean floor
And there’s a special opportunity for university academics to clean up when a lecture-free universitärer Verfügungstag becomes university disposal day.
When the Feature Is Also the Bug
AI can make “informed guesses.” This may solve one problem while creating an utterly different one that no human would expect. For example, researchers at the University of Tübingen reconstructed an ancestral natural antibiotic. With the German language’s beautiful capacity for compound words, this became Urantibiotikum. DeepL didn’t recognize this word, so corrected it to two words it did recognize—Uran and Antibiotikum—and translated them as uranium antibiotic.
A Game of Telephone
When translating between several languages, DeepL and Google Translate use English as a “pivot language” through which all other translations run. That means, for example, when going from German into French, your German text will first be processed into English and then that result will be processed into French. This can introduce or amplify errors.
The Recycled Data Economy
AI-generated texts are posted online and then gobbled up as new training data. This self-affirming cycle leads to the boosting of trends as AI suggestions are taken up and amplified. (For example, delve is currently a popular word.) In translation, common mistakes become institutionalized. (Tage can mean days, but is frequently used in the sense of Tagung, a meeting or conference. A translation into days leaves non-German speakers wondering what students do for the rest of the year if there are only five Study Days.)
First the Quick Fix, Then the Hard Yards
Just as we walk past the bike to the car when we’re tired, running late, or it looks like rain, we’re going to use AI for text generation simply because it’s quicker. In making this Faustian bargain, we switch from the creative process of formulating and expressing our own ideas to the critical process of working out what’s wrong and fixing it. We must examine every sentence carefully because there’s no guarantee that what an AI-generated text says, or even what we think it says, is true.
Alan Turing didn’t speak in terms of “intelligence.” Instead, he envisaged machines with the calculating power to mimic human behavior so well that humans could be fooled. In 2026, we take the Turing Test every time we log on the internet. The question to be asking now is not “How smart are the machines?” but “How smart are we?”
Tips for Using AI Translation Programs
- Always consider data privacy. Whatever you upload to a free translation website will be in the public domain, instantly.
- Beware of rubber-stamping a translation of a source text in your native language. You may understand every word of the translation, but that may be a false positive. After all, you know what it’s meant to say. An international audience may be left utterly confused. Translating back into the source language can help identify and fix errors.
- Always pre-edit the source text. Ideally, avoid sentences with more than one relative clause, make grammatically clear sentences, and scan for words that could easily mean something else and replace them or rephrase to avoid ambiguities.
- Input full sentences or short paragraphs so you can check your translation piece by piece.
- Remember that every text is written for a particular audience, in a particular social and cultural context, and that the act of translation will place it before a different audience, possibly unfamiliar with that context. The translator must always ask: Will the target audience get this?
- Be aware that at any moment, AI translation can introduce a mistake you’ve never dreamed of. There are no guarantees.
Amanda Crain has worked as a German-to-English translator at the University of Tübingen, Germany, for 15 years. Previously, she was an editor and producer in the Deutsche Welle TV English newsroom. She has a master’s degree in German studies and is a qualified legal translator in Germany. amanda.crain@uni-tuebingen.de


