Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Researchers Introduce SCAPO Training Method

Researchers Introduce SCAPO Training Method

HuggingFace·Friday, October 9, 2026
  • •Westlake University researchers introduced SCAPO, a token-level method for training language models with verifiable rewards
  • •SCAPO raised AIME 2024–2026 accuracy over GRPO by 5.63 points on Qwen3-4B-Base and 4.17 on Qwen3-1.7B-Base
  • •The method led among compared approaches on most tested math benchmarks and all tested out-of-distribution benchmarks
  • •Westlake University researchers introduced SCAPO, a token-level method for training language models with verifiable rewards
  • •SCAPO raised AIME 2024–2026 accuracy over GRPO by 5.63 points on Qwen3-4B-Base and 4.17 on Qwen3-1.7B-Base
  • •The method led among compared approaches on most tested math benchmarks and all tested out-of-distribution benchmarks
  • •Westlake University researchers introduced SCAPO, a token-level method for training language models with verifiable rewards
  • •SCAPO raised AIME 2024–2026 accuracy over GRPO by 5.63 points on Qwen3-4B-Base and 4.17 on Qwen3-1.7B-Base
  • •The method led among compared approaches on most tested math benchmarks and all tested out-of-distribution benchmarks
  • •Westlake University researchers introduced SCAPO, a token-level method for training language models with verifiable rewards
  • •SCAPO raised AIME 2024–2026 accuracy over GRPO by 5.63 points on Qwen3-4B-Base and 4.17 on Qwen3-1.7B-Base
  • •The method led among compared approaches on most tested math benchmarks and all tested out-of-distribution benchmarks

Westlake University researchers introduced Semifactual Credit-Augmented Policy Optimization (SCAPO), a method for training large language models with verifiable rewards. The paper was published on September 30 and submitted to Hugging Face on October 8. It examines why models can change their answers when prompts contain features unrelated to the task, even when the underlying problem and correct answer stay the same.

The researchers report that token-level sensitivity varies substantially. During decoding, suppressing high-drift token candidates improved reasoning accuracy without updating model weights. They also identify a limitation in Group Relative Policy Optimization (GRPO): it assigns the same outcome-based advantage to every response token, which may reinforce irrelevant as well as useful reasoning.

SCAPO adds token-level credit assignment based on semifactually stable responses. It measures changes in token probabilities when a fixed response is evaluated under prompt interventions that preserve the problem and answer. Normalized stability scores reduce the advantage given to relatively unstable tokens during early training; stability by itself receives no additional credit.

On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improved AIME 2024–2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model sizes, it achieved the best results among compared methods on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks. The authors report that the code is available on GitHub.

Westlake University researchers introduced Semifactual Credit-Augmented Policy Optimization (SCAPO), a method for training large language models with verifiable rewards. The paper was published on September 30 and submitted to Hugging Face on October 8. It examines why models can change their answers when prompts contain features unrelated to the task, even when the underlying problem and correct answer stay the same.

The researchers report that token-level sensitivity varies substantially. During decoding, suppressing high-drift token candidates improved reasoning accuracy without updating model weights. They also identify a limitation in Group Relative Policy Optimization (GRPO): it assigns the same outcome-based advantage to every response token, which may reinforce irrelevant as well as useful reasoning.

SCAPO adds token-level credit assignment based on semifactually stable responses. It measures changes in token probabilities when a fixed response is evaluated under prompt interventions that preserve the problem and answer. Normalized stability scores reduce the advantage given to relatively unstable tokens during early training; stability by itself receives no additional credit.

On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improved AIME 2024–2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model sizes, it achieved the best results among compared methods on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks. The authors report that the code is available on GitHub.

Read original (English)·Oct 9, 2026
#scapo#policy optimization#reinforcement learning#verifiable rewards#qwen3#aime#token level credit assignment#grpo