A new arXiv preprint proposes that an image-generating model can get better by criticising its own work. It is a submission, not a verified result: no reviewer has accepted or validated it.
The method, called UniEvo-VL, splits one model into two roles. The student sees only the plain request. The teacher sees the same request plus a written critique of an earlier attempt, a hint the student would not normally have. Training then nudges the student toward what the teacher would have done. All of this is tied to test-time compute, meaning it happens while the model is being used, not during its original training.
Built on the open-source Qwen-image-2512, the authors report GenEval moving from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. Those are the authors’ own measurements. They also say a stronger outside critic, GPT5.6-Luna, points to a higher ceiling, while text-rendering results were mixed.
Teams eyeing self-critique loops for image quality should treat text-heavy output as unproven until someone else reproduces the work.
Based on the UniEvo-VL preprint posted to arXiv, submitted 30 September 2026.