How to Successfully Deploy AI in Localization (So It Works)
The GenAI revolution is here, and the exponential progress of large language models (LLMs) in recent months may have led many to believe that human linguists will soon become obsolete. However, the opposite is likely to be true: The future will be shaped by the complementary nature of AI capabilities and human expertise. And quality is the reason why humans should be the ones in the driver’s seat.
Our landscape after the ChatGPT storm
AI is revolutionizing the localization industry. At the beginning of 2023, most practitioners were still skeptical, or at least cautious, about LLMs. But a few months later, after the release of GPT-4, it is hard to find anyone who has not yet joined the GenAI revolution. The innovation race is on, and new solutions are appearing on the market every week. Chillistore, a subsidiary of Argos Multilingual, is keeping up — the group stepped into the race in 2020 by forming its own Innovation Lab and commencing development work on new AI-powered solutions.
One of the widely discussed topics currently is how to manage quality in AI-assisted localization. It is becoming increasingly clear that standard quality management processes and tools will need to be thoroughly transformed, and a novel approach will need to be developed. This is particularly true when LLMs are not only used for content translation, but also for content generation.
Moving to multilingual language models
However, because LLM’s training data is primarily based on English text, there is a significant disparity in experience between English-speaking ChatGPT users and users of other languages. To address this issue, transitioning to truly multilingual language models has become a top priority. These models use training data that is distributed equally across language groups and leverage transfer learning techniques to establish connections between languages and apply what they have already learned from other languages.
Building multilingual language models means an increased demand for linguists who can evaluate the quality of language data produced by the model. They may also need to post-edit the data to improve its quality for further model retraining. Various studies have confirmed that linguists trained in evaluation should be involved to ensure that issues are correctly identified and resolved.
Quality management in the GenAI world
Although we are still on the cusp of the GenAI revolution, we can already see companies (both buyers and service providers) experimenting in various areas of quality management. Here are just a few examples of the topics we’d like to explore:
Quality screening
In addition to their use in content creation and translation, research suggests that LLMs might also aid in understanding target content quality. Nowadays, LLMs are being used to pioneer quality assessment. The accuracy of quality evaluation produced with the assistance of LLMs appears promising, as evidenced by a recent study conducted by one of the industry leaders on a few high-resource languages.
Pre-assessment of content quality
Although the gap between high- and medium-/low-resource languages still needs to be bridged to ensure comparable evaluation accuracy across languages, the conclusion that LLMs can be used for quality evaluation even when there is no human reference available opens up new possibilities for the use cases and implementation of language quality evaluations.
For example, with the use of LLMs, problematic parts of the content may be identified, and related quality management activities might be scoped to proactively mitigate risks. These activities may include additional revisions, enhanced or additional language quality evaluation (LQE), additional layers of automated quality checks, and cleanup of translation memories, among others.
Linguistic asset management
It is a well-known fact that translation memories (TMs) significantly reduce the cost and time required for translations. They are also essential for training and customizing machine translation (MT) engines. In the world of GenAI, high-quality linguistic assets, particularly TMs, will gain even more value due to their potential to help customize language models.
AI-powered translation memory (TM) clean-up
TMs are vulnerable databases. They are modified and processed by many users with little or no control. A typical TM would contain hundreds of thousands of segments translated into multiple languages. However, as it grows with time, its quality usually deteriorates. Bringing TM quality back to a healthy state is very costly and time-consuming. As a result, TM owners would rather adopt a compromise solution than invest in a genuinely thorough TM clean-up.
This is where large language models (LLMs) might become a real game-changer. Initial experiments with TM clean-ups using LLMs for semantic analysis and identification of potentially erroneous segments have shown significant time and cost savings. For instance, Argos’ AI-powered TM clean-up is 90% faster than if the same activity was performed in the traditional way. See the case study on page 12 for more details.
Content insights
In today’s world, understanding the full range of content attributes is crucial for determining the most effective content delivery strategies. LLMs allow for a wide range of perspectives when analyzing vast amounts of content. The following are just a few examples of AI-assisted content analysis.
Content profiling
With the enormous amount of content produced, it is crucial to map the source content and cluster it based on various criteria, such as complexity, specific knowledge and expertise requirements, or the implementation of particular technologies, processes, and workflows. LLMs enable gathering all types of information.
In addition to various “counts” (character count, word count, sentence count, duplicates count, etc.), data on sentiment, spelling quality, grammar check, and text complexity can be obtained. Moreover, a quality manager can obtain information on critical elements like product names, part numbers or referential content, or extract terminology and validate it against the existing glossary.
AI-assisted content analysis can also help collect invaluable information for risk analysis and address reputational, legal, or general quality risks, or risks impacting user experience and accessibility, diversity, and inclusion. Content profiling is key to specifying quality requirements relevant to content use cases, right sourcing, defining the production workflow, scoping various quality controls, testing, and more.
Identification of style and tone of voice
Insights obtained from AI-powered content analysis can be helpful in understanding the tone of voice or language register. Specifics found in the given content type can be reflected in detailed style guide rules and tone of voice guidelines. This guidance is important not only for linguists, but it also helps fine-tune LLMs to better understand the particular nuances necessary for content creation.
Competitive analysis and sentiment detection
AI-assisted content analysis may be extended to include the content of other brands and user-generated content published on social media and forums. This analysis could, for instance, focus on capturing language attributes that are typical for other brands’ content and their user personas. Alternatively, sentiment detection could help understand user’s feelings about brands and their suite of products.
AI-assisted checks of language inclusivity
LLMs have the potential to detect anomalies and screen large volumes of data to identify offensive or inappropriate content, as well as content that does not comply with required diversity and inclusion standards. By leveraging AI’s capabilities to identify more complex inclusivity issues beyond just “banned words and concepts,” we can efficiently make content more inclusive. Any market- and language-specific challenges for delivering inclusive content should be considered when setting up the regular content creation process.
AI and language quality evaluation (LQE)
AI is rapidly transforming localization technology and tools in all aspects of the content creation life cycle. The potential for AI-powered language quality evaluation (LQE) is both fascinating and tempting from a quality management perspective, as it offers exciting opportunities for exploration and experimentation.
Both options are possible: high-level content-quality information as well as detailed error reporting
Even though it is still early days, information on text cohesion, compliance with terminology, grammar, or various specific rules can be collected to provide a high-level overview of content quality.
Currently, there are numerous experiments with AI-powered analytical language quality evaluation. The system is trained with precisely defined error typologies and severities, usually with multidimensional quality metrics (MQM), and with all mandatory references (glossaries, style guides, branded and do-not-translate terminology, etc.). The outcome of the AI-driven LQE is a quality report that lists segments with detected errors, including specifications of error types and severities, and suggestions for correct translations.
New opportunities for extending content quality evaluation coverage
By incorporating AI into content quality evaluation, it is now possible to evaluate all content types and use cases, a task which was previously seen as prohibitively time-consuming and costly. The AI-powered LQE (“LQE assistant”) can work alongside humans to improve quality at scale. The LQE reports generated by AI can aid LQE experts or trained evaluators to be more efficient. For less visible or less important content, raw quality data from the reports can be collected and analyzed, and only outliers or suspicious information would be revised by humans.
LQE experts are more than quality evaluators
New horizons will also open up for LQE experts. They will have the opportunity to observe the “behavior” of the AI algorithm and learn how to provide actionable feedback to prompt engineers and data scientists, or to receive training on creating and improving prompts themselves.
Content creation in the GenAI world
Challenges and potential solutions
Using generative AI for content creation is currently one of the most intensively explored topics. The limitations of LLMs for content creation have already been well described, so let’s focus on the most frequently mentioned ones:
| Limitation | Details and Solutions |
|---|---|
| Factual inaccuracies and hallucination issues. | Content may contain wrong statements or incorrect information. The ethical challenge of publishing AI-generated content that does not correctly reflect reality underlines the irreplaceable role of humans in detecting and fixing such issues. |
| Challenges in achieving consistent quality and performance across languages. | User experiences vary considerably across cultures. The disparity between the amount of English and non-English data is the key root cause. Solution: the creation of multilingual language models. |
| Current language models have not yet been developed to a state that would enable them to cater to the diverse needs and preferences of different user communities. | Content created by GenAI often conveys values that are encoded in English, rather than offering values that are relevant to the culture and language of the reader. As a result, the user may not be emotionally engaged, even if the text is technically correct from a linguistic standpoint. |
| Creating unique content for a specific brand is difficult. | It is often hard to distinguish content across multiple brands. This is why having detailed brand guidelines for each individual language has become more important than ever before. |
| Data anonymization is necessary for openly available language models. | To avoid extra cost and time, development of alternative solutions is key. |
| There is a somewhat negative sentiment and mistrust towards AI-generated content. | Content often feels unnatural and lacks engagement, even though it is technically correct. Humans tend to be much more critical of machine-generated content than they are of human-generated content, despite the fact that both may have similar issues. There are often exaggerated expectations regarding the capabilities of GenAI, while a lot of research and experimentation is still needed. Only human experts, such as scientists, linguists, and engineers can enhance the quality of AI-generated output and help improve overall sentiment towards it. Additionally, regulatory frameworks for the use of AI can help mitigate risks and increase trust. |
Linguists and their role
Looking around, we see that limitations often raise curiosity and motivate us to seek solutions. GenAI is a tool, and as such, humans are the driving force behind it. To improve the quality and performance of GenAI, we need to redefine the traditional roles of linguists.
In addition to the ever-growing importance of expert knowledge (such as domain, market, language processing, socio-economic, and cultural knowledge), there will be a high demand for “post-AI editors.” These editors may work with a checklist like this to improve the output quality of GenAI:
- Carefully fact-check to eliminate any potential misinformation.
- Apply creative edits that add a human touch and make the text intellectually engaging.
- Correct linguistic issues, especially in content generated in rare languages.
- Adjust the style to align with the branded tone of voice.
- Handle branded terminology.
- Screen the output for potential bias.
- Suggest improvements to fixed parts of prompts to minimize recurring quality issues.
- Re-prompt and adjust non-fixed parts of prompts for their specific one-time use.
- Guide AI to deliver proper structure for SEO needs, including keywords and linking structure.
- Ensure the uniqueness of the output by embedding human-originated words and phrases in prompts.
- Perform enhanced editing on chunks of text to help (re)train, fine-tune, and create multilingual models.
Generative AI offers one powerful solution: prompts can be modified multiple times to tailor the output, bringing it as close as possible to the desired results. Linguists play a key role in the content generation process in the world of AI. If trained, they can enhance the quality of the content immediately by fine-tuning the prompts themselves. Instant collaboration between linguists and engineers is the future of AI content generation.
So what’s the future?
The rapid pace of this technological revolution presents unforeseen opportunities, and nobody wants to fall behind. As human beings, we naturally assume that content or products will resonate with our emotions and expectations. Today’s reality, however, proves that significant and thoughtful investments are still needed if we want to offer comparable experiences across languages, cultures, and communities.
It’s no wonder, then, that there’s still a reserved sentiment about AI-generated content.
The internet is already flooded with low-quality content created solely for profit, which unfortunately drives away users. To combat this issue, Google has been penalizing auto-generated spammy content designed to manipulate search rankings for years. Additionally, the exponential growth of disinformation is a concerning issue.
To address this, the European Commission urges tech giants operating in the EU to label content generated by AI. Furthermore, AI content detectors are emerging on the internet and continuously improving in accuracy. Although it is currently not difficult to trick such detectors, they may become much more reliable very soon, especially after authorities introduce legal regulations for AI technology.
From the user’s perspective, content generated by human experts helps establish trust based on the belief that there are dedicated individuals with relevant expertise behind the material. In contrast, AI-generated content, even if it is fluent and easy to read, lacks “the human touch.”
The language may be error-free but not emotionally engaging, and the style and/or word choice may not match the particular user community’s preferences, especially if the extralinguistic context is not reflected. Except perhaps for English, users can recognize that content was not created by humans, which in certain cases may leave them feeling deceived or even threatened.
…the future is actually bright for us, language professionals
One thing is becoming increasingly clear: successful deployment of language models relies on humans at the core. It’s not just about researchers, data scientists, or software engineers; it’s also about linguists.
In the GenAI era, a linguist’s value will grow with any type of expertise offered on top of “just being the translator.” Domain expertise (financial, legal, medical, socio-cultural), in-depth product and/or market knowledge, and other types of expertise will be critical.
Linguists will be urgently needed to help enhance language models, fill gaps in language datasets for low-resource languages, and document all types of language-specific rules, brand guidelines, and terminologies used to improve the performance of language models.
In other words, humans won’t just be in the loop; they will be in control when it comes to using AI for anything related to language.






