The writer's problem
A publisher weighing machine translation for a novel needs a benchmark more structured than a vendor's demo. The Ninth Conference on Machine Translation ran a dedicated literary-translation shared task, and its organizers, Tencent AI Lab and China Literature Ltd., published a findings paper describing what was actually compared and how.
What the documents show
The findings paper's own abstract states the task received 10 submissions from 5 academia and industry teams, evaluated with both automatic metrics, including BLEU, chrF, COMET and a document-level d-BLEU score, and human judgment, with the official system ranking based on the human evaluation rather than the automatic scores. Systems compared included Google and GPT-4 as baselines alongside participant systems such as Cloudsheep, HW-TSC, NLP2CT-UM and SJTU-LoveFiction. The task's own organizer page announced, ahead of the event, that Chinese-German and Chinese-Russian data would be added alongside the existing Chinese-English pair, describing this as new for the year. The findings paper's own limitations section states plainly that this year's task in fact focused only on the Chinese-English direction, with the additional language pairs deferred to a future edition. The organizer page's advance announcement and the findings paper's final account of what was actually run do not fully match on scope, a gap worth noting rather than assuming away.
The editorial choice
A publisher should read a shared-task ranking as a comparison under one specific evaluation protocol and one text domain, web fiction translated from Chinese, not as a general verdict on any system's literary-translation quality in other language pairs or genres. Editorially, citing this benchmark to justify a translation workflow decision should specify the language direction and note that the official ranking reflects human judges' assessment of these particular submissions, not an automated score alone.
What stays with the author
Neither document evaluates a translated novel as a finished literary object read start to finish, and neither substitutes for a human translator's judgment on register, idiom or voice across a full book. The shared task measures system performance on a defined slice of text; deciding whether a translated passage reads as literature remains outside its scope.
- Does this benchmark's Chinese-English web-fiction domain resemble the text actually being translated?
- Would the ranking change under a different evaluation protocol or a different genre?
- What does the gap between the announced scope and the evaluated scope suggest about reading any shared-task page as final before its findings paper appears?
The two documents together describe a real, structured comparison, bounded to one direction and one domain, not a general ranking of machine translation for literary work.
Follow the source.
The findings paper's own abstract and limitations section state 10 submissions from 5 teams were evaluated by automatic metrics and human judgment, with official rankings based on human judgment, and that the task ultimately covered only the Chinese-English direction.
Source date: 15 Nov 2024 · Retrieved: 16 Sept 2026
The organizers' own task page, current as retrieved, announced Chinese-German and Chinese-Russian datasets as new additions for 2024 alongside the existing Chinese-English pair and describes the A/B testing evaluation method introduced that year.
Source date: Not established · Retrieved: 16 Sept 2026
Site publication is not established by an event date. Original record ID: 0030-bf-096. This local design review does not change its editorial status.