Gemini 4 Pro Leaked Online, Crushing Astra and Fable - Wall Street News
News

Gemini 4 Pro Leaked Online, Crushing Astra and Fable - Wall Street News

新智元
2026.09.18
·Web·by JunhaRyu
#AI Agent#Gemini 4 Pro#LLM#RSI

Key Points

  • 1Google has reportedly launched a new, powerful AI model under the alias "gemini-3.8-flash" on the AI Arena, which is widely suspected to be the next-generation Gemini 4 Pro.
  • 2Benchmark results indicate that this model outperforms competitors like OpenAI’s Astra and Anthropic’s Fable 5.1 in complex tasks, including coding, reasoning, and advanced SVG or 3D generation.
  • 3The accelerated development of Gemini 4 Pro is attributed to Google’s successful implementation of a Recursive Self-Improvement (RSI) loop, marking a significant shift in their AI training strategy.

The provided text reports on the emergence of a high-performance, unannounced AI model in the LMSYS Chatbot Arena, identified under the moniker "gemini-3.8-flash." Industry speculation and technical assessments strongly suggest this is the next-generation flagship, Gemini 4 Pro, from Google DeepMind.

Core Developments and Performance

Gemini 4 Pro represents a significant leap in performance, reportedly outperforming current industry leaders OpenAI’s "Astra" and Anthropic’s "Fable 5.1." Performance metrics cited from the Arena include:
  • DeepSWE v1.1: Achieved an 88% success rate in autonomous coding agent tasks, surpassing competitors by nearly 2%.
  • GDPval-AA v2: Secured a landmark 2064 Elo rating in complex, real-world knowledge tasks.
  • Terminal-bench 2.1: Demonstrated 95.3% proficiency in terminal-based coding environments.
  • OSWorld-2.0: Achieved 86.8% accuracy in complex computer-use and operating system navigation.

Technical Methodology: Recursive Self-Improvement (RSI)

The paper posits that the accelerated development cycle and superior intelligence of Gemini 4 Pro are driven by a successful implementation of Recursive Self-Improvement (RSI).

  • RSI Mechanism: Unlike traditional supervised fine-tuning, RSI establishes a closed-loop system where the AI acts as an autonomous agent to evaluate, critique, and improve its own underlying logic and search strategies. This effectively creates an "AI-training-AI" architecture that transcends traditional exponential scaling, potentially allowing for a super-linear performance trajectory.
  • Supporting Evidence: The disclosure of Google’s "Dream-RSI" research suggests a paradigm shift where agents optimize their internal search algorithms through experiential learning. This recursive feedback loop allowed Google to expedite the pre-training phase of Gemini 4, compensating for the cancellation of the previously aborted Gemini 3.5 Pro project.

Capabilities and Technical Utility

The model demonstrates marked advancements in multi-modal generation and complex reasoning:
  • Creative and Generative UI/UX: The model exhibits high-fidelity generation in SVG/3D environments, including real-time rendering of complex objects like H145 helicopters and interactive, physics-based 3D games.
  • Architectural Efficiency: With a 10 million input token window and 256,000 output token limit, the model supports permanent cross-session memory and native web-connectivity capabilities without requiring external API dependencies.
  • Cost Efficiency: Despite its superior performance, the model maintains a competitive pricing structure at 2.25permillioninputtokensand2.25 per million input tokens and11.25 per million output tokens, positioning it as a highly cost-effective solution within the current "Big Three" LLM landscape.

Strategic Context

Google’s stealth release of Gemini 4 Pro signals a strategic pivot to reclaim technological dominance. By integrating RSI, Google intends to shift the competitive landscape from static chat interfaces to advanced, long-horizon autonomous agents capable of complex research and software engineering. The model serves as the critical centerpiece in maintaining Google's competitive edge against the emerging flagship models from OpenAI and Anthropic.