Skip to content
Dashboard

Ling 3.0 Flash

Ling 3.0 Flash is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.

ReasoningTool UseImplicit Caching
import { streamText } from 'ai'
const result = streamText({
model: 'inclusionai/ling-3.0-flash',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingMore models by Inclusionai

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Release Date
256K
1.2s
385tps
Free
Free
novita logo
08/06/2026