×
LLMs

Qwen3.8-Max: A New Bar for Coding and Cowork

Today, we are officially releasing Qwen 3.8-Max, the most capable model in the Qwen family to date.

Alibaba Unveils Qwen3.8-Max: Its Largest and Most Capable Flagship Model to Date

Qwen3.8-Max exhibits advanced capabilities in coding, real-life work, research, and long-horizon tasks.

Deep Dive: How Kimi's AI Agent Runs on Alibaba Cloud

This article explains how Kimi leverages Alibaba Cloud's ACK and ACS to build a secure, instantly elastic infrastructure capable of supporting hundreds of thousands of concurrent AI Agent sandboxes.

Quest 1.0: Refactoring the Agent with the Agent

Tokens Produce Deliverables, Not Just Code

Qwen2.5: A Party of Foundation Models!

This article introduces the latest addition to the Qwen family, Qwen2.5, along with specialized models for coding and mathematics.

The Hermetic AI Sandbox: Deploying Sovereign Qwen Models in Fully Air-Gapped VPCs

This article provides a step-by-step architectural guide to deploying Qwen LLMs in a fully air-gapped, zero-internet Alibaba Cloud VPC to ensure strict data sovereignty and regulatory compliance.

Qwen Enables Rapidly Expanding AI Hardware Ecosystem

Qwen, Alibaba’s inhouse LLM and multimodal AI model family, is increasingly becoming a foundational component for next-generation AI hardware.

The Second Half of the Enterprise Agent Era: How to Make Agents Smarter the More They Are Used?

This article introduces AgentLoop, Alibaba Cloud's one-stop platform that enables enterprise AI agents to continuously self-evolve through full-stack observability and automated evaluation.

DeepSeek V4-Flash dalam Skala Besar: Panduan Deployment Berbasis Benchmark

Memilih cara men-deploy large language model di lingkungan produksi adalah salah satu keputusan paling konsekuensial — sekaligus paling membingungkan — yang dapat diambil sebuah tim AI.

Alibaba Cloud AI Gateway FinOps Features Officially Launched | Making Every Token's Consumption "Visible and Controllable"

This article introduces Alibaba Cloud AI Gateway's new FinOps features for LLM cost governance and token quota management.

Say Goodbye to Goldfish Memory: PolarDB Mem0 Gives AI Agents long-term Memory

This article introduces PolarDB Mem0, a managed long-term memory service that gives AI agents persistent, structured recall through integrated vector and graph capabilities.

Code Harness or Natural-Language Harnesses?

This article introduces Natural-Language Agent Harnesses (NLAH), replacing traditional code-based agent control with executable natural language strategies.

I Tested 19 LLM API Workloads on Real Calls and Cut Costs 79% — Here's the Data

518 real API calls. $33.99 → $7.06 in a single run. The same parameter change projects $15,667/year saved on a healthcare workload — here's the exact code, the math, and every scenario I measured.

AgentScope Java 2.0: Building a Distributed, Enterprise-Grade Foundation for AI Agents

This article introduces AgentScope Java 2.0, an open-source framework for building distributed, enterprise-grade AI agents with production-ready features.

Alibaba Launches Qwen3.7-Plus, AI Swine Diagnosis Assistant and Model Studio CLI

This article introduces the launch of Qwen3.7-Plus multimodal model, an AI swine diagnosis assistant with Muyuan Group, and Model Studio's open-source CLI for AI agents.

Qwen-VLA: From Understanding the World to Acting in It

This article introduces Qwen-VLA, a general-purpose Vision-Language-Action model that extends multimodal perception and reasoning into continuous action generation for embodied intelligence.

Beyond 'Demo-Grade' Architecture: Building a Highly Available Production Foundation for Dify with SAE × SLS

This article introduces Alibaba Cloud SAE, a serverless platform that simplifies application modernization and accelerates AI deployment with zero node management.

DeepSeek V4-Flash at Scale: A Benchmark-Driven Deployment Guide

Choosing how to deploy a large language model in production is one of the most consequential — and confusing — decisions an AI team can make.

LoongCollector + ACS Agent Sandbox: Build a Production-grade AI Agent Runtime Platform

This article introduces a production-grade AI Agent runtime platform combining ACS Agent Sandbox for security and LoongCollector for observability.

Alibaba Cloud Tair KVCache Simulation Analysis: High-Precision Computational and Caching Simulation Design and Implementation

This article introduces Tair-KVCache-HiSim, a high-fidelity CPU-based simulator for optimizing multi-tier KV Cache configurations in LLM inference.