Enterprise Translation

September 08, 2026

|

4 min read

Ecommerce Translation Quality Tiers: What Needs Human Review

Not all ecommerce content deserves the same translation workflow. This guide maps product pages, checkout, reviews, support content, and campaign copy to five quality tiers, with the decision rules, cost trade-offs, and quality thresholds that tell you when to move a content type to AI-only.

LILT Team

LILT Team

Ecommerce Translation Quality Tiers: What Needs Human Review

Most ecommerce content does not need human translation review. Product attributes, filters, and long-tail catalog copy can run AI-only without anyone noticing. Checkout, returns policies, regulated claims, and your highest-traffic product pages need a human. The decision comes down to one question: what does an error on this page actually cost you?

Retailers who route everything through the same workflow pay too much for the content that does not matter and move too slowly on the content that does. This guide maps the quality tiers, matches them to specific ecommerce content types, and sets out the rules for deciding which tier a piece of content belongs in.

On this page

  • What are the quality tiers for ecommerce translation?
  • Which ecommerce content needs human review?
  • Five questions that route any content type
  • When is raw AI translation good enough for product content?
  • What does machine translation post-editing actually mean?
  • How do you move content to a cheaper tier over time?
  • How do you measure whether the routing is working?
  • How LILT routes ecommerce content
  • Frequently asked questions

What are the quality tiers for ecommerce translation?

Translation quality is not a single dial running from bad to good. It is a set of distinct workflows, each with a defined amount of human involvement and a defined set of checks. Two of them are anchored in international standards, which makes them worth naming precisely.

TierWhat it isHuman touch pointsWhat gets checkedStandardBest for

1. Raw AI

Machine output published as-is

0

Nothing

None

Internal use, gisting, dead-stock long tail

2. AI Review

AI agents check and correct AI output before publication

0

Terminology, do-not-translate lists, flagged errors

None yet

High-volume catalog attributes

3. Light post-editing

Linguist corrects accuracy, no stylistic polish

1

Accuracy, meaning, terminology

ISO 18587 (light)

Mid-tier product pages, support macros

4. Full human verification

Linguist brings output to publication quality

1 to 2

Accuracy, terminology, grammar, style, locale conventions

ISO 18587 (full), ISO 17100

Checkout, legal, flagship product pages

5. Transcreation

Message recreated for the market

2+

Brand intent, cultural fit, campaign performance

None

Campaigns, homepage, ad copy

Tier 1: Raw AI

Machine output goes live with no review of any kind. This is defensible for content where comprehension is the only requirement and nobody is making a purchase decision on it. It is indefensible anywhere a customer forms an impression of your brand.

Tier 2: AI Review

AI agents check the machine output before it publishes, enforcing glossaries and do-not-translate lists and flagging segments that look wrong. There is still no human in the loop, but the output is no longer unchecked. This tier did not exist a few years ago and there is no standard governing it yet, which is why evaluating it means asking a vendor what the agent actually checks rather than accepting the label.

Tier 3: Light post-editing

A linguist reads the machine output and fixes anything factually or linguistically wrong. They do not polish style. ISO 18587 defines this as producing a text that is comprehensible and accurate, without expecting it to read like it was written by a native marketer.

Tier 4: Full human verification

A linguist brings the output to publication quality: accuracy, terminology, grammar, style, and locale conventions such as units, currencies, and address formats. ISO 18587 calls this full post-editing and holds it to the same output standard as ISO 17100 human translation.

Tier 5: Transcreation

The linguist is given the intent rather than the sentence, and writes new copy that achieves the same effect in the target market. A slogan that depends on an English pun does not survive translation at any tier below this one.

You may see AI post-editing, or AIPE, used as a sixth category. It is an emerging term with no standard behind it, so treat it as a vendor label rather than a defined workflow.

Which ecommerce content needs human review?

The tiers only become useful when they are attached to specific content. This is the mapping most retail catalogs converge on. Treat it as a starting position to argue with, not a prescription, because the cost of an error varies by category and market.

Content typeCost of an errorRecommended tierWhy

Product attributes and variants

Low, filters break

2

Highly repetitive, very high translation memory reuse

Long-tail product descriptions

Low, minor confusion

2

Volume makes human review uneconomic

Flagship and high-traffic product pages

High, lost conversion

4

Small page count carrying a large share of revenue

Category and collection pages

Medium, navigation and search visibility

3

Keyword-sensitive, few pages

Checkout and payment flow

Very high, abandoned carts

4

Trust-critical, tight character limits

Returns, shipping, and terms

Very high, regulatory exposure

4 or 5

Local law varies, may need specialist review

Customer reviews and user content

Low

1 or 2

Volume, and shoppers expect unpolished authenticity

Support macros and help articles

Medium, ticket deflection

3

Reused across thousands of tickets

Campaign and homepage copy

High, brand perception

5

Needs recreation, not translation

Ad copy and paid search

High, wasted spend

5

Character limits plus cultural fit

SEO metadata and hreflang

Medium, discoverability

3

Requires native keyword research, not translation

Images containing text

Medium

3 or 4

Needs visual QA as well as linguistic review

Two rows are worth dwelling on, because they are where retailers most often get it wrong in both directions.

Customer reviews are routinely over-invested in. Shoppers reading a translated review are looking for signal about the product, not polished prose, and a lightly awkward translation reads as authentic rather than careless. Checkout is routinely under-invested in. It is a handful of strings, so it looks trivial, but a mistranslated shipping term or an overflowing button label costs you the sale after you already paid to acquire the visit.

Five questions that route any content type

When a content type is not on the list above, these five questions will place it.

  1. Does a customer make a purchase decision on this content? Product pages, comparison tables, and size guides do. Internal metadata does not.
  2. What does an error actually cost: confusion, a return, or a regulator? Confusion tolerates tier 2. Returns justify tier 3 or 4. Regulators require tier 4 or 5 with a specialist reviewer.
  3. Is this repeated across thousands of SKUs? High repetition means high translation memory reuse, which means the marginal cost of a higher tier falls sharply after the first pass. Repetition can justify better treatment, not worse.
  4. Does it sit in a character-limited UI component? Buttons, banners, filter labels, and mobile navigation can be linguistically perfect and still break the layout. These need in-context visual review regardless of tier.
  5. Is it discoverable in search? Content that people find through search engines or AI assistants needs native keyword research. Translating your English keywords produces text nobody searches for.

When is raw AI translation good enough for product content?

Raw AI, with no review at all, is genuinely fine in a narrower set of cases than most vendors will admit and a wider set than most brand teams will accept.

It works when the content is internal, when it exists purely so someone can understand roughly what a thing is, or when the alternative is no translation at all.

It stops working the moment content is customer-facing and load-bearing. Three failure modes recur in retail: product terminology that is correct in general usage but wrong in your category, regulatory and safety claims where a plausible-sounding error creates liability, and anything with a character limit, where output that is linguistically perfect still breaks the component it lives in.

The practical test: if you would be uncomfortable seeing this string in a screenshot on social media, it does not belong in tier 1.

What does machine translation post-editing actually mean?

Machine translation post-editing, or MTPE, is the workflow where a machine produces a first draft and a human linguist corrects it. It is the most widely used paid translation workflow in the industry, and ISO 18587 defines two levels of it.

Light post-editing aims at content that is accurate and comprehensible. The post-editor fixes mistranslations, omissions, and errors that change meaning. They leave stylistic awkwardness alone. The result reads as functional rather than polished.

Full post-editing aims at output indistinguishable from a human translation. The post-editor corrects accuracy, terminology, grammar, style, and locale conventions. ISO 18587 holds the result to the same output standard as ISO 17100, the human translation standard.

The distinction matters commercially, because vendors quote both as "post-editing" and the price gap between them is large. Ask which level a quote refers to and what the post-editor is instructed to leave alone.

MTPE is also not the same thing as human-in-the-loop translation. In classic MTPE, the model produces output, a human fixes it, and the correction disappears. In a human-in-the-loop system, the correction feeds back and the model stops making that error. Over a large catalog, that difference compounds into a very different cost curve.

How do you move content to a cheaper tier over time?

The point of tiering is not to lock content into a workflow permanently. It is to start conservative and earn your way down as the evidence supports it.

Three mechanisms drive content down the tiers.

Translation memory reuse. Once a string is approved, it should never be paid for again. In catalogs with heavy repetition across variants and seasons, reuse rates climb steeply after the first full pass, and the volume actually requiring processing drops with it.

Model specialization. A model trained on your catalog, your glossaries, and your approved past translations makes category-specific errors less often than a general-purpose one. Fewer errors in the raw output means less correction needed, which means a lower tier produces acceptable results.

Correction feedback. In systems where reviewer corrections retrain the model instantly, every edit a linguist makes on tier 4 content reduces the correction burden on similar content later.

For localization teams, this changes the job rather than removing it. Instead of processing queues, the work becomes setting thresholds, deciding which categories have earned a tier change, and governing quality across markets. That is a more defensible remit than being the bottleneck every content request has to pass through.

What quality thresholds should trigger a tier change?

Set the threshold before you start. A workable pattern: track corrections per thousand words for a content category over a defined period, and when the rate stays below an agreed level across two consecutive review cycles with no severity-one errors, move the category down one tier. Keep a sample of the demoted content in human review as an audit, so a quality regression surfaces before customers find it.

Never move more than one tier at a time, and never move regulated content on metrics alone.

How do you measure whether the routing is working?

Subjective spot checks are how most retailers evaluate translation quality, and they are how most retailers end up arguing about a handful of unrepresentative examples. Measure instead.

Linguistic measures. The MQM error typology is the industry framework here, classifying errors by category, accuracy, fluency, terminology, style, locale conventions, and design, and by severity. Categorized errors tell you whether a problem is a model issue, a glossary gap, or a brief that was never written.

Operational measures. Corrections per segment, edit distance between raw output and published text, terminology compliance rate, translation memory reuse rate, rework rate, and on-time delivery.

Coverage measures. The one most retailers cannot answer and every executive eventually asks: what percentage of the active catalog is translated and current in each market? Without it, you cannot tell whether a market is underperforming because of demand or because half the catalog is missing.

Agree these metrics during onboarding or a proof of concept, with thresholds attached. Metrics without thresholds become reporting rather than decisions.

How LILT routes ecommerce content

LILT is a multilingual agentic AI platform rather than a translation vendor or a standalone translation management system, and the tiering above is built into how content moves through it rather than being a service you order.

Content is triggered from the system of record. More than 100 native integrations connect LILT to Shopify, Magento, Salesforce Commerce Cloud, and the PIM, DAM, and CMS systems where product content originates, so new and changed items are detected, routed, and returned to the correct locale without manual submission.

Routing happens automatically by content type, audience, and risk. AI Review agents handle tier 2, correcting errors before a person sees them and prioritizing error-prone content for escalation. Expert human verifiers handle tiers 3 through 5, with domain and in-market expertise for regulated and campaign content. Every correction retrains the model in real time, so the share of content requiring human review falls over time rather than staying fixed.

ASICS increased translation velocity by 60% and reduced costs by 70% working this way. The INKEY List uses LILT to translate its Shopify store and customer collateral.

If you want to see how your own catalog would route across these tiers, book a demo and bring your content inventory rather than a sample file.

Frequently Asked Questions

Is AI translation good enough for product descriptions?

For most of a large catalog, yes. Long-tail product descriptions and structured attributes are repetitive and low-risk, and AI output with automated review is usually indistinguishable to shoppers. Your highest-traffic product pages are the exception. They carry a disproportionate share of revenue, there are few of them, and human verification on that small set is inexpensive relative to what a conversion drop costs.

What is the difference between MTPE and human translation?

In machine translation post-editing, a machine produces the first draft and a human corrects it. In human translation under ISO 17100, a human translates from scratch and a second linguist revises. Full post-editing under ISO 18587 targets the same output quality as human translation, so the difference is process and cost rather than the standard the finished text is held to.

Does machine translation hurt multilingual SEO?

Unreviewed machine translation can hurt visibility, but not for the reason most people assume. The problem is rarely a penalty. It is that translated text uses the words your English copy used rather than the words local shoppers actually search for. Metadata, category names, and headings need native keyword research, which is a research task rather than a translation task.

Should you translate customer reviews?

Usually yes, at a low tier. Translated reviews measurably increase buyer confidence in markets where shoppers cannot read the original, and the volume makes anything above tier 2 hard to justify. Shoppers also read reviews expecting informal, imperfect language, so light awkwardness costs less here than almost anywhere else on the site.

What is ISO 18587?

ISO 18587:2017 is the international standard for post-editing machine translation output. It defines the requirements for the post-editing process, distinguishes light from full post-editing, and sets out the competences a post-editor must have. It is the companion to ISO 17100, which covers human translation. Asking a vendor which of the two a given workflow conforms to is a fast way to find out what you are actually buying.

How do you know when to stop human-reviewing a content type?

Set a measurable threshold in advance, such as corrections per thousand words staying below an agreed level across two consecutive review cycles with no severity-one errors. When a category clears it, move it down one tier and keep auditing a sample. Never demote regulated content on metrics alone, and never move more than one tier at a time.

Contact Us

Learn more about how LILT can simplify your translations with AI.

Book a Meeting

Share this post

Copy link iconCheckmark