Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Arena Combines Rewards to Post-Train Image Models

Arena Combines Rewards to Post-Train Image Models

Arena AI·Saturday, October 3, 2026
  • •Arena combines human preference votes with prompt-specific rubric rewards to post-train text-to-image models.
  • •FLUX.2-dev gained 69 Arena points; Ideogram 4 scored 1224 and passed listed open-source models.
  • •Ensembled reward configurations reached a 66.0% win rate on 1K held-out Arena prompts.
  • •Arena combines human preference votes with prompt-specific rubric rewards to post-train text-to-image models.
  • •FLUX.2-dev gained 69 Arena points; Ideogram 4 scored 1224 and passed listed open-source models.
  • •Ensembled reward configurations reached a 66.0% win rate on 1K held-out Arena prompts.
  • •Arena combines human preference votes with prompt-specific rubric rewards to post-train text-to-image models.
  • •FLUX.2-dev gained 69 Arena points; Ideogram 4 scored 1224 and passed listed open-source models.
  • •Ensembled reward configurations reached a 66.0% win rate on 1K held-out Arena prompts.
  • •Arena combines human preference votes with prompt-specific rubric rewards to post-train text-to-image models.
  • •FLUX.2-dev gained 69 Arena points; Ideogram 4 scored 1224 and passed listed open-source models.
  • •Ensembled reward configurations reached a 66.0% win rate on 1K held-out Arena prompts.

Arena introduced a post-training method for text-to-image models that combines human preference scores with rubric-based rewards. The method uses about 5 million pairwise votes from Text-to-Image Arena, collected across more than 100 models, and rubrics generated with language and vision-language models. The rubrics check whether images follow prompts, honor constraints and avoid reward-hacking behaviors. Arena reports that FLUX.2-dev gained 69 Arena points on its live T2I leaderboard, while Ideogram 4 reached 1224 and surpassed every publicly listed open-source model by Sep 04, 2026.

The preference component is a Bradley–Terry reward model: given a prompt and image, it returns a score reflecting broad human judgments such as visual quality, composition and aesthetics. Arena says the model outperformed other state-of-the-art models on the MMRB2 benchmark, and that post-training performance improved with more reward-model training data. For prompt faithfulness, a language model turns each prompt into a tree-structured checklist of yes-or-no questions about objects, attributes, spatial relationships and style. A vision-language model checks the generated image against the questions; the share of satisfied criteria becomes the faithfulness reward.

Separate rubric rewards check conflicts with user intent, such as added objects, changed styles or violations of negative instructions, and recurring exploits such as garbled text or photorealism when a non-photographic style was requested. The constraint reward is applied when prompts call for strict adherence. When an anti-hacking detector flags an image, it gates the preference reward so a positive preference score is canceled; without a flag, that score is preserved. Arena says this avoids adding a separate objective that a model could game.

Arena trained on 10K real user prompts and evaluated on a held-out set of 1K Arena prompts. Each post-trained checkpoint was compared with its frozen base model using the MMRBv2 pairwise evaluation protocol, with Gemini-3.5-Flash as judge; each image pair was shown in both orders to reduce position bias. Adding faithfulness and intent-gated constraint rewards raised the win rate to 64.2%; ensembling policies trained with complementary reward configurations reached 66.0% with Arena RM. Arena says the recipe also worked with the open-source Pickscore reward model.

Arena introduced a post-training method for text-to-image models that combines human preference scores with rubric-based rewards. The method uses about 5 million pairwise votes from Text-to-Image Arena, collected across more than 100 models, and rubrics generated with language and vision-language models. The rubrics check whether images follow prompts, honor constraints and avoid reward-hacking behaviors. Arena reports that FLUX.2-dev gained 69 Arena points on its live T2I leaderboard, while Ideogram 4 reached 1224 and surpassed every publicly listed open-source model by Sep 04, 2026.

The preference component is a Bradley–Terry reward model: given a prompt and image, it returns a score reflecting broad human judgments such as visual quality, composition and aesthetics. Arena says the model outperformed other state-of-the-art models on the MMRB2 benchmark, and that post-training performance improved with more reward-model training data. For prompt faithfulness, a language model turns each prompt into a tree-structured checklist of yes-or-no questions about objects, attributes, spatial relationships and style. A vision-language model checks the generated image against the questions; the share of satisfied criteria becomes the faithfulness reward.

Separate rubric rewards check conflicts with user intent, such as added objects, changed styles or violations of negative instructions, and recurring exploits such as garbled text or photorealism when a non-photographic style was requested. The constraint reward is applied when prompts call for strict adherence. When an anti-hacking detector flags an image, it gates the preference reward so a positive preference score is canceled; without a flag, that score is preserved. Arena says this avoids adding a separate objective that a model could game.

Arena trained on 10K real user prompts and evaluated on a held-out set of 1K Arena prompts. Each post-trained checkpoint was compared with its frozen base model using the MMRBv2 pairwise evaluation protocol, with Gemini-3.5-Flash as judge; each image pair was shown in both orders to reduce position bias. Adding faithfulness and intent-gated constraint rewards raised the win rate to 64.2%; ensembling policies trained with complementary reward configurations reached 66.0% with Arena RM. Arena says the recipe also worked with the open-source Pickscore reward model.

Read original (English)·Oct 2, 2026
#text to image#arena#flux.2 dev#ideogram 4#preference learning#rubric rewards#reward model#prompt faithfulness