Google Releases Gemini 3.5 Flash With Agentic Gains
- •Gemini 3.5 Flash achieves a 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash.
- •The model reaches speeds over 280 output tokens/s, marking a 70% improvement in velocity over the previous version.
- •Costs have risen significantly, with a 3x price increase to $1.50/$9.00 per 1M input/output tokens.
Google DeepMind has introduced Gemini 3.5 Flash, the latest addition to its model family, which achieves a score of 55 on the Artificial Analysis Intelligence Index. This represents a 9-point improvement over the previous Gemini 3 Flash, largely driven by enhanced performance in agentic tasks (autonomous operations using tools) and a significant reduction in model hallucinations. On the AA-Omniscience benchmark, the hallucination rate dropped by 31 points to reach 61%.
The model demonstrates significant advancements in agentic capabilities, particularly in GDPval-AA, where it achieved an Elo rating of 1656. This performance places it ahead of Gemini 3 Flash (1204) and Gemini 3.1 Pro (1314), trailing only GPT-5.4 (1674). Despite these intelligence gains, the model maintains high efficiency with speeds exceeding 280 output tokens/s, approximately 70% faster than its predecessor.
Operational costs have increased, with the model being 5.5x more expensive to run on the Intelligence Index than Gemini 3 Flash and 75% costlier than Gemini 3.1 Pro. The pricing is set at $1.50 per 1M input tokens and $9.00 per 1M output tokens, representing a 3x increase in unit price compared to the 3 Flash version. Users can benefit from a 90% discount on cached input tokens. The model maintains a 1M context window and supports multimodal inputs including image, video, and speech. In the MMMU-Pro multimodal evaluation, it achieved a record-high score of 84%, surpassing Gemini 3.1 Pro at 82%.