Meta AI found that language-model judges preferred AI-written paper sections to the expert originals in 63.5% to 84.6% of comparisons. When the researchers gave an AI judge a standard grading rubric instead, it scored the model-written section higher every time. A lab that trains a writer to please such a grader risks teaching it to produce more of the writing people call AI slop.
Meta AI researcher Jason Weston made a bigger claim when he posted the work on X on September 29: “Claim: we’ve solved the AI slop problem (!)” The evidence behind “solved” is 38 human judgments, 18 of them the authors rating sections of their own papers.
In a September 2026 research blog, Meta presents a way to make graders favor expert writing and reports substantial gains for a model trained against them.
The judge preferred the replacement
The researchers took a published computer-science paper, removed its abstract, introduction, related-work section, or conclusion, and asked a model to fill the gap. That gave them two sections answering the same assignment: the author’s original and the model’s replacement. In pairwise tests, the judge saw the candidates in both orders to control for a preference for whichever appeared first.
The model could not simply win by writing more. In early tests without a length instruction, replacements ran two or three times as long as the originals. For the main paper-writing task, Meta asked writers to stay within roughly 15% of the original section’s word count.
Even then, the judges usually chose the AI text. The 63.5%-84.6% range depended on which model wrote and which judged. Meta tested Anthropic’s Opus 4.8 and OpenAI’s GPT-5.6 in both roles. Meta also had GPT-5.6 generate task-specific grading criteria; under that standard rubric approach, AI-written sections scored higher in 100% of comparisons.
The blog describes the training consequence directly:
“If AI slop is judged to be better than human writing by the LLM grader, then training will encourage more slop.”, Meta AI research blog
A polished replacement can cover every part of a paper and still be a worse section. Meta says its initial rubrics tended to penalize human authors for deliberate omissions and reward model text for breadth and surface fluency. The grader was scoring what was easy to recognize, not necessarily what an expert had done well.

Meta calls its repair Reinforcement Learning from eXpert-Aligned Rubrics, or RL-XAR. It collects expert human passages, generates model alternatives, and revises the instructions used to create grading rubrics until the expert passages score higher. It then trains the writer against those rubrics and repeats the process to find weaknesses in the improved writer.
A preliminary check used 52 paper-section examples from eight training papers and five held-out validation papers. Across seven rubric-revision rounds, the validation score gap, human minus model, went from −4.2 to +2.76, crossing into human-favoring territory at round four. That shows the grading instructions could be changed to recognize distinctions the initial ones missed.
A better writer, measured by Meta’s graders
Meta then trained Qwen3.5-27B to write against the learned rubrics. A separate, larger Qwen model judged rewards during training; Meta used GPT-5.6 to judge its reported rubric evaluations. Its paper-writing dataset contained 2,243 training sections drawn from 561 papers, with another 360 sections from 90 papers for validation.
After two training rounds, the strongest paper-writing checkpoint scored 9.60 against a human-normalized 10 on the weakest of three learned rubrics. Meta tested across multiple rubrics because a model trained to satisfy one set of criteria might do poorly on another.
In blind tests using the researchers’ own papers, they preferred the trained Qwen writer to the untrained Qwen baseline 16 times to 2. Meta says the baseline often tried to turn an introduction into a miniature account of the entire paper, while the trained model stayed closer to the section’s job.

Story writing showed a similar improvement over baseline. On Meta’s learned rubric, the trained model rose from 2.8 to 8.2 against a human-normalized 10; the next-best frontier writer in that test scored 6.8. Blind readers preferred the trained model to its untrained baseline 19 times to 1.
Where the result stops
Wikipedia writing was much less responsive. Meta reports that training moved Qwen from 2.7 to 4.0, a 1.3-point gain calculated from those scores, while human writing was normalized to 10. The trained model remained below every frontier writer Meta reported for that task: GPT-5.6 scored 7.9, and two others scored 7.2. Meta points in part to the limits of the Qwen judge used during training.
The researchers also acknowledge that their learned rubrics may be “still potentially biased towards our model”. Using a different model to judge the final scores addresses one concern, the training and evaluation judges are not the same, but the rubrics themselves remain part of Meta’s method. No one outside Meta has tested RL-XAR yet.
Meta has shown that a writing grader can reward the wrong thing, and that changing the grader can improve a writer in the tasks it tested. Its full technical report is still forthcoming. For now, the Wikipedia result is the clearest boundary on the claim: the same training recipe that brought Qwen close to human scores on Meta’s paper rubric left it at 4.0 out of 10 on another kind of writing.
Further Reading
- Towards RL for Superhuman Text: Unslopping AI, Meta AI’s account of its judge tests, training method, and writing evaluations.
