News
Gemini 4 Pro Leaked Online, Crushing Astra and Fable - Wall Street News
新智元
2026.09.18
·Web·by JunhaRyu#AI Agent#Gemini 4 Pro#LLM#RSI
Key Points
- 1Google has reportedly launched a new, powerful AI model under the alias "gemini-3.8-flash" on the AI Arena, which is widely suspected to be the next-generation Gemini 4 Pro.
- 2Benchmark results indicate that this model outperforms competitors like OpenAI’s Astra and Anthropic’s Fable 5.1 in complex tasks, including coding, reasoning, and advanced SVG or 3D generation.
- 3The accelerated development of Gemini 4 Pro is attributed to Google’s successful implementation of a Recursive Self-Improvement (RSI) loop, marking a significant shift in their AI training strategy.
The provided text reports on the emergence of a high-performance, unannounced AI model in the LMSYS Chatbot Arena, identified under the moniker "gemini-3.8-flash." Industry speculation and technical assessments strongly suggest this is the next-generation flagship, Gemini 4 Pro, from Google DeepMind.
Core Developments and Performance
Gemini 4 Pro represents a significant leap in performance, reportedly outperforming current industry leaders OpenAI’s "Astra" and Anthropic’s "Fable 5.1." Performance metrics cited from the Arena include:- DeepSWE v1.1: Achieved an 88% success rate in autonomous coding agent tasks, surpassing competitors by nearly 2%.
- GDPval-AA v2: Secured a landmark 2064 Elo rating in complex, real-world knowledge tasks.
- Terminal-bench 2.1: Demonstrated 95.3% proficiency in terminal-based coding environments.
- OSWorld-2.0: Achieved 86.8% accuracy in complex computer-use and operating system navigation.
Technical Methodology: Recursive Self-Improvement (RSI)
The paper posits that the accelerated development cycle and superior intelligence of Gemini 4 Pro are driven by a successful implementation of Recursive Self-Improvement (RSI).- RSI Mechanism: Unlike traditional supervised fine-tuning, RSI establishes a closed-loop system where the AI acts as an autonomous agent to evaluate, critique, and improve its own underlying logic and search strategies. This effectively creates an "AI-training-AI" architecture that transcends traditional exponential scaling, potentially allowing for a super-linear performance trajectory.
- Supporting Evidence: The disclosure of Google’s "Dream-RSI" research suggests a paradigm shift where agents optimize their internal search algorithms through experiential learning. This recursive feedback loop allowed Google to expedite the pre-training phase of Gemini 4, compensating for the cancellation of the previously aborted Gemini 3.5 Pro project.
Capabilities and Technical Utility
The model demonstrates marked advancements in multi-modal generation and complex reasoning:- Creative and Generative UI/UX: The model exhibits high-fidelity generation in SVG/3D environments, including real-time rendering of complex objects like H145 helicopters and interactive, physics-based 3D games.
- Architectural Efficiency: With a 10 million input token window and 256,000 output token limit, the model supports permanent cross-session memory and native web-connectivity capabilities without requiring external API dependencies.
- Cost Efficiency: Despite its superior performance, the model maintains a competitive pricing structure at 11.25 per million output tokens, positioning it as a highly cost-effective solution within the current "Big Three" LLM landscape.