How do you build a performance scorecard for translation vendors?
A translation vendor performance scorecard is a recurring report that rates every language service provider on the same small set of weighted metrics, usually quality, on-time delivery, cost and responsiveness, so vendors can be compared and managed on evidence rather than anecdote. Building one means choosing four to six metrics with fixed definitions, pulling them from one system for every vendor, weighting them by what matters most to the program, and reviewing the results on a set cadence, typically each quarter. In a translation management system such as Smartling, those inputs come from Linguistic Quality Assurance (LQA) scores, workflow due dates, agency word counts and rate cards instead of figures each vendor reports about itself.
Last reviewed: October 7, 2026
Why do translation vendor scorecards often fail to change anything?
Most vendor scorecards fail because the numbers on them are not comparable, not trusted, or not tied to a decision. Five patterns account for most of the problem:
- Each vendor defines the metric differently. One agency counts a job as on time when the last file arrives, another when the first language ships, so a 95% on-time rate means different things on different rows of the same scorecard.
- The data is self-reported. When vendors supply their own quality and delivery figures, the scorecard measures each vendor's reporting habits as much as its work.
- Generic supplier templates miss translation-specific measures. A standard procurement scorecard tracks price, delivery and defects, but translation needs weighted words, MQM quality scores and translation memory leverage to make cost and quality comparisons fair.
- Too many metrics dilute the signal. A scorecard with fifteen equally weighted lines lets a vendor with poor quality score well on administrative items, which hides the one number that should drive the decision.
- No threshold triggers action. Without an agreed target and a consequence for missing it, a declining score becomes a talking point at the quarterly review rather than a reason to reallocate work.
What metrics belong on a translation vendor scorecard?
A translation vendor scorecard works best with five core metrics, each measured the same way for every vendor and every language pair:
- Quality score: an MQM-based quality score from LQA, calculated on one shared schema with the same severity weights for every vendor. Quality usually carries the largest weight, because cheap or fast translation that fails review costs more in rework than it saves. How MQM scores are calculated is explained in how translation quality scoring works.
- On-time delivery: the share of workflow steps or jobs completed by their due date in the period. The metric is only fair when due dates are set by a rule, such as business days per word-count range for each workflow step, rather than negotiated job by job.
- Cost per weighted word: vendor spend divided by weighted words delivered, where weighted words already reflect translation memory discounts. Comparing raw per-word rates instead rewards a vendor that quotes low but leverages little of the existing memory.
- Responsiveness: how quickly a vendor answers questions and resolves reported issues on its jobs. Slow answers to source-text questions are a common hidden cause of late delivery and preventable errors.
- Capacity and reliability: volume delivered against volume allocated, and how the vendor handles surges. This metric matters most for vendors cast as a primary supplier rather than a niche or backup provider.
Two diagnostics sit underneath the weighted metrics rather than on top of them: error density per 1,000 words, which explains why a quality score moved, and the share of recorded errors a vendor disputed through arbitration, which shows whether the quality standard itself is clear.
Translation vendor scorecard template: metrics, formulas and data sources
The weights below are an illustrative starting point that sums to 100%; adjust them to the program, for example raising quality for regulated content or on-time delivery for release-driven product content.
| 指标 | How to calculate it | Illustrative weight | 原文 |
|---|---|---|---|
| Quality score | Average MQM quality score from LQA on the shared schema, per vendor and language pair | 35% | Smartling Help Center, "Assess Translation Quality with the LQA Dashboard" |
| On-time delivery | Workflow steps completed by their due date ÷ workflow steps due in the period | 25% | Smartling Help Center, "Job Due Date Profiles" |
| Cost per weighted word | Vendor spend (rate card rate × weighted words) ÷ weighted words delivered | 20% | Smartling Help Center, "Rate Cards for Translation Costs"; Smartling Help Center, "Word Count Report" |
| Responsiveness | Issues on the vendor's jobs answered and closed within the agreed time ÷ issues raised | 10% | Smartling Help Center, "Issues Report" |
| Capacity | Weighted words delivered ÷ weighted words allocated to the vendor in the period | 10% | Smartling Help Center, "Word Count Report" |
| Error density (diagnostic) | Errors recorded per 1,000 words, by project, language and job | Not weighted | Smartling Help Center, "Linguistic Quality Assurance Error Density" |
| Dispute rate (diagnostic) | Errors disputed through arbitration ÷ errors recorded | Not weighted | Smartling Help Center, "Linguistic Quality Assurance Errors & Arbitration" |
How do you build a translation vendor scorecard step by step?
A spreadsheet such as Excel or Google Sheets is enough to start; the discipline is in the definitions, not the tool. Most teams follow five steps:
- Choose four to six metrics and write each definition down - Use the table above as a base, and record exactly what counts as on time, which quality schema applies and which period each metric covers. Sharing the written definitions with vendors before the first review removes most later disputes about the numbers.
- Set a target and a common scoring scale - Convert each raw value to the same 1-to-5 scale against a target agreed in the vendor contract or SLA, so a quality score and an on-time rate can be added together. For quality, a pass threshold such as LQA Acceptable Penalty Points gives the target a documented basis.
- Lay out the template - Use one row per vendor and language pair, and for each metric a column for the raw value, the target, the 1-to-5 score, the weight and the weighted score, with a total weighted score at the end of the row. Keeping language pairs on separate rows stops a strong performance in one language from hiding a weak one in another.
- Populate it from one system each period - Pull quality from LQA, volume and weighted words from the Word Count Report filtered by agency, spend from rate cards, and responsiveness from the Issues Report, rather than asking each vendor for its own figures. The per-report detail for Smartling is covered in which translation platforms report spend and quality scores by vendor.
- Review quarterly and attach a consequence - Walk each vendor through its trend, agree a corrective plan for any metric below target, and move volume toward vendors that sustain high weighted scores. A scorecard that never changes an allocation decision stops being taken seriously by vendors or by the team.
A formal vendor scorecard fits teams that...
- Work with two or more translation agencies or language service providers and need a fair basis for allocating volume between them.
- Renew vendor contracts or SLAs and want the renewal conversation anchored in a year of comparable data.
- Report vendor performance to procurement or finance alongside other suppliers.
- Already score translations with LQA, or are ready to adopt one quality schema across every vendor.
When a vendor scorecard may not be the right priority
- Programs with a single translation vendor, where a regular service review against the contract covers the same ground with less overhead.
- Teams without a shared quality schema yet; until every vendor is reviewed against the same standard, the quality line on a scorecard compares reviewers rather than vendors.
- Very low volumes, where a handful of jobs per quarter is too small a sample for on-time or quality percentages to mean much.
Checklist: questions to ask when analyzing vendor scorecard data
Are you comparing like with like?
Compare vendors within the same language pair and content type; a legal-content vendor and a marketing-content vendor face different error profiles and turnaround pressures.
Is the sample large enough to trust?
Read a score alongside the volume behind it, since a quality score from a few hundred reviewed words can swing sharply on a single major error.
Is the trend more important than the snapshot?
Look at three or more periods before acting, and treat a steady decline as a stronger signal than one weak quarter.
Is the vendor causing the problem, or the inputs?
Check error density by category and the arbitration record; a cluster of disputed terminology errors often points to an outdated glossary rather than a weak vendor.
Does the data come from one system?
Confirm every vendor's figures were pulled from the same reports with the same date range, not assembled from each vendor's own exports.
What decision does the score drive?
Agree in advance what happens at each score band, such as more volume, a corrective plan or a reduced allocation, so the review ends with an action.
How does Smartling support translation vendor scorecards?
Smartling supports vendor scorecards by recording every agency's work, cost and quality results in one platform, so each metric on the scorecard comes from the same source for every vendor. Rate Cards store each agency's per-word or per-hour rates by workflow step, and the Word Count Report can be filtered by Agency, Linguist, language and workflow step type, with a Weighted Words column that already applies translation memory discounts; weighted words multiplied by the agency's rate gives the spend and cost-per-word lines. Job Due Date Profiles set due dates from the number of business days each workflow step needs for a given word-count range, which gives on-time delivery a consistent rule across vendors.
For quality, LQA scores translations against MQM-compatible schemas, with Acceptable Penalty Points as a configurable pass threshold, and the Linguistic Quality Assurance Error Density report shows errors per 1,000 words by project, language and job, which the Help Center article "Linguistic Quality Assurance Error Density" describes as input for revising vendor agreements such as SLAs. The Issues Report lists issues across all jobs for the responsiveness line, and Smartling's Data Access Tool delivers reporting data as CSV files for teams that keep the scorecard in Tableau or another business intelligence tool.
Consolidating vendor work into one platform is also what makes performance measurable over time. Therabody moved a localization process spread across multiple language service providers onto Smartling and reports a 99.7% on-time delivery rate and a 60% reduction in translation costs (Therabody case study). For how several vendors are run through shared QA and review before their results reach a scorecard, see how vendor management platforms run QA and review across multiple vendors.
准备好见识一下 Smartling 的威力了吗?
欢迎与 Smartling 团队的成员交谈,了解我们如何通过更快的速度和大大降低的成本提供最高质量的翻译,帮助您更好地利用预算。