For as long as there has been a localization process, linguistic errors have not been far behind. The introduction of machine translation (MT) has gone some way to even out the bumpy road of translation quality.
However, it comes at a cost: neural machine translation (NMT) engines may introduce linguistic errors that no human ever would. And with complicated, specialized content, this risk only increases.
The recent progress in AI, however, makes it possible to deploy algorithms that help clean language assets or help automate the process of language QA. Let’s look at a few such innovations.
Take translation memories (TM), for example. The sad truth of the matter is that TMs deteriorate over time, depreciating their value. This could be because multiple target equivalents of one source segment are saved to the TM (creating inconsistency). It could be because segments that are split in the source language cannot be split in the target language in the same way due to word order differences. Or it could be that the terminology or linguistic style used in the past is no longer applicable. In any case, a little spring cleaning every so often is necessary.
Why should I run an AI TM cleanup?
Submitting a TM for manual review is expensive, as this is a time-consuming and complex task. This means most TMs rarely undergo the necessary maintenance.
A semantic analysis feature (more about this later) allows our artificial intelligence (AI) model to process segments accurately in just a fraction of the time it would take for a human to do the same. The AI TM Cleanup tool allows for cost and time savings and improves linguistic quality in all future translation projects. You can have your TM back as good as new!
How does the cleanup work?
This tool relies on AI’s capabilities to automatically target potentially low-quality TM segments and flag them for review.
Step 1: Our AI model analyzes the health of the TM before defining a scope of content required for review alongside auxiliary services.
Step 2: A distribution report is generated from the analysis, and a quote is rendered based on the total number of issues.
Step 3: After the AI model has detected errors, misalignments, etc., our professional linguists make the necessary corrections, followed by QA (quality assurance) checks.
Step 4: A comparison report between the status and TM health pre- vs. post-cleanup is generated.
Our AI TM Cleanup tool’s main feature is its ability to perform semantic analysis. This feature allows the model to process, for example, 100,000 segments and then identify thousands of segments within the larger set that are of suspicious quality. This allows you to flag things that need to be reviewed by a human translator and what doesn’t.
What else can I use AI for?
Back translation: This is a standard element of some Quality Assurance (QA) workflows, especially in regulated industries. Traditionally, this involves a third-party linguist translating the target text back into the source language, manually comparing the back translation to the source. Sounds time-consuming, right? That’s where AI technology comes in.
What makes AI back translation better?
Aside from the aspect of time, traditional back translation may be complicated. For example, mistranslations that still make sense in the context make it unclear that it’s a mistake. The same also goes with textual omissions. If the translated text still makes sense even with part of the text left out, it’s hard for a linguist to identify the error.
By integrating AI into the back translation process, we can detect potential omissions and mistranslations, errors that are notorious for being difficult to spot.
How AI back translation works
AI enhances this traditional process by having two separate models “analyze” the differences between source and the back-translated target, assigning each model a binary score depending on the likelihood of a genuine error in the segment. Scores are added together to create a final score.
- Score 0: Neither AI model suspects any error within a segment.
- Score 1: One of the AI models suspects there is an error within a segment.
- Score 2: This occurs when both AI models suspect an error within a segment. This is likely an omission.
The back translation comparison report prioritizes segments with the highest score (2).
On top of using these models for error detection, we also use a third AI model to identify potentially erroneous segments in the report. This is done by training the model with historical data and creating artificial errors.
Beyond accelerating the back translation process, this tool aims to reduce a quality specialist’s exposure to false positives, increasing the accuracy and efficiency of the report review.
Get ready for a smoother process!
AI Backtranslation and AI TM Cleanup are two examples of innovative use of “small AI.” These tools have provided tangible improvements to the translation process and are only set to continue advancing over time.





