Alibaba Releases Qwen3.7-Plus Multimodal Agent Model
- •Alibaba released Qwen3.7-Plus, a multimodal agent model unifying vision and language capabilities.
- •The model achieves 70.3 on Terminal Bench 2.0 and 90.3 on the GPQA Diamond benchmark.
- •Qwen3.7-Plus automated a full 11-hour software development cycle, generating 10,000+ lines of code.
Alibaba introduced Qwen3.7-Plus, a multimodal agent model designed to unify vision and language capabilities into a single, versatile foundation. This model builds upon the Qwen3.7 text backbone to support end-to-end task execution, including GUI (graphical user interface) navigation, CLI (command line interface) operations, and complex software engineering workflows. Qwen3.7-Plus is now available via Alibaba Cloud Model Studio, offering developers an agent capable of perceiving real-world scenes, writing code from visual references, and automating multi-step productivity tasks.
In performance evaluations, Qwen3.7-Plus demonstrated strong results across various text and coding benchmarks. On Terminal Bench 2.0, the model achieved a score of 70.3, and it recorded a 62.1 on QwenWorldBench. Its reasoning capabilities were validated through scores of 90.3 on GPQA Diamond and 92.9 on the HMTT 2026 February benchmark. The model also showed robust performance in general visual understanding, scoring 86.9 on RealWorldQA and 77.0 on CountQA, while maintaining 88.0 on VideoMME.
The model functions as a multimodal interactive hybrid agent that integrates a "see, think, write, act, and verify" loop. In practical tests, the system operated continuously for over 11 hours to automate the research and development cycle of a mobile application. During this process, it generated more than 10,000+ lines of code and triggered over 1,000+ agent calls. It supports API integration compatible with OpenAI’s specifications, allowing developers to use features like the preserve_thinking parameter for enhanced agentic workflows in complex environments.