Alibaba Cloud has unveiled its latest multimodal AI model, "Qwen3.8-Omni-Flash." Designed with Google's Gemini Flash in mind, this new model achieves a dramatic reduction in inference costs while maintaining high benchmark performance.
"Qwen3.8-Omni-Flash" combines rapid response times with advanced multimodal processing capabilities. Optimized specifically for cost-efficiency, it aims to provide an environment where enterprises and developers can build high-performance AI applications at a lower cost.
In benchmarks, Qwen3.8-Omni-Flash delivers performance comparable to existing major competing models. The reduction in inference costs has been achieved through optimizations in the model architecture, which is expected to become a crucial factor in enhancing competitive pricing within the AI market.
Through this model, Alibaba Cloud intends to support enterprises focusing on cost-efficiency in their AI adoption. Moving forward, the company plans to expand API availability and strengthen its ecosystem to accelerate deployment across diverse use cases.