Tongyi Lab Debuts FIPO Algorithm, 32B Model Inference Outperforms o1-mini

Tongyi Lab's Intelligent Computing team today unveiled a new post-training algorithm for large language models — FIPO (Future-KL Influenced Policy Optimization). This algorithm introduces a novel "Future-KL" mechanism that effectively resolves the common technical challenge of "inference length stagnation" in pure reinforcement learning (Pure RL) training.
During training for long-text reasoning and complex logical alignment, traditional reinforcement learning often fails to precisely identify key decision points in lengthy sequences. The FIPO algorithm, developed by the Tongyi team, applies differentiated reward allocation to key tokens, encouraging the model to adopt a more forward-looking approach during chain-of-thought (CoT) generation.
Experimental results indicate that, in a pure reinforcement learning setup with a 32B-scale model, the model using the FIPO algorithm has already outperformed similarly sized models like DeepSeek-Zero-MATH and OpenAI's o1-mini, representing significant advancement in the logical reasoning and mathematical computation abilities of domestic large models.
Currently, the competition among large models is shifting from pre-training scale to deep alignment at inference time. The release of FIPO not only offers a new way to assess the quality of "thinking processes" in logical reasoning models, but also signals that the open-source community and leading domestic labs are gradually forging an independent technological path in their pursuit of world-class reasoning models.
Related article
How to fix Core Web Vitals for better SEO rankings
How to Build a Free Website: Free Domain, Hosting & AI Website BuilderTable of ContentsIntroductionAI Website Builder 1: Hookous AIStep 1: Sign UpStep 2: Choose Business Type/NicheStep 3: Select ServicesStep 4: Define Website GoalsStep 5: Enter Busin
Apple, Google Partner With Anthropic to Address 27-Year-Old Vulnerability via Glass Wing Protection
As artificial intelligence advances rapidly in code generation and logical reasoning, the cybersecurity landscape faces unprecedented challenges. Recently, the prominent AI startup Anthropic officially launched a cross-industry collaboration called *
OpenAI Chief Scientist Addresses AI Reasoning Transparency Debate: Complexity Steady, No Sudden Jump
On September 2, Jakub Pachocki, OpenAI’s Chief Scientist, addressed public concerns on X regarding the AI model Astra, clarifying claims that it operates without oversight and lacks transparent reasoning.Why the Controversy Erupted: Deep Recurrence O
Related Special Topic Recommendations
Comments (0)
0/500

Tongyi Lab's Intelligent Computing team today unveiled a new post-training algorithm for large language models — FIPO (Future-KL Influenced Policy Optimization). This algorithm introduces a novel "Future-KL" mechanism that effectively resolves the common technical challenge of "inference length stagnation" in pure reinforcement learning (Pure RL) training.
During training for long-text reasoning and complex logical alignment, traditional reinforcement learning often fails to precisely identify key decision points in lengthy sequences. The FIPO algorithm, developed by the Tongyi team, applies differentiated reward allocation to key tokens, encouraging the model to adopt a more forward-looking approach during chain-of-thought (CoT) generation.
Experimental results indicate that, in a pure reinforcement learning setup with a 32B-scale model, the model using the FIPO algorithm has already outperformed similarly sized models like DeepSeek-Zero-MATH and OpenAI's o1-mini, representing significant advancement in the logical reasoning and mathematical computation abilities of domestic large models.
Currently, the competition among large models is shifting from pre-training scale to deep alignment at inference time. The release of FIPO not only offers a new way to assess the quality of "thinking processes" in logical reasoning models, but also signals that the open-source community and leading domestic labs are gradually forging an independent technological path in their pursuit of world-class reasoning models.
How to fix Core Web Vitals for better SEO rankings
How to Build a Free Website: Free Domain, Hosting & AI Website BuilderTable of ContentsIntroductionAI Website Builder 1: Hookous AIStep 1: Sign UpStep 2: Choose Business Type/NicheStep 3: Select ServicesStep 4: Define Website GoalsStep 5: Enter Busin
Apple, Google Partner With Anthropic to Address 27-Year-Old Vulnerability via Glass Wing Protection
As artificial intelligence advances rapidly in code generation and logical reasoning, the cybersecurity landscape faces unprecedented challenges. Recently, the prominent AI startup Anthropic officially launched a cross-industry collaboration called *
OpenAI Chief Scientist Addresses AI Reasoning Transparency Debate: Complexity Steady, No Sudden Jump
On September 2, Jakub Pachocki, OpenAI’s Chief Scientist, addressed public concerns on X regarding the AI model Astra, clarifying claims that it operates without oversight and lacks transparent reasoning.Why the Controversy Erupted: Deep Recurrence O





Home






