Optimized for coding and agentic workflows at $0.75/M input / $3.75/M output (introductory through Dec 2026), it's the most capable Flash yet and a direct cost win for any pipeline running 3.6 Flash today.
Why it matters: Half-price swap-in with improved coding and agent performance — a straightforward upgrade path for any team running Flash-tier models in production.
Powered by Cerebras wafer-scale hardware, the new API tier runs 14× faster than standard, targeting coding, financial research, and support use cases; in limited enterprise preview with pricing TBD.
Why it matters: Real-time inference speed unlocks UX patterns (live streaming, sub-second response) that previously required sacrificing model quality.
Claude sessions can now discover and message each other via @ mentions; subagent forks inherit the full conversation and prompt cache by default, reducing setup overhead for multi-agent pipelines.
Why it matters: Native coordination between Claude sessions reduces complexity for any multi-agent workflow — less glue code, faster iteration on agentic builds.
SynthID-Text watermarks (no quality, speed, or cost impact) roll out to future Claude models for EU AI Act compliance; a detection API for verifying AI-generated content is coming soon.
Why it matters: The upcoming detection API gives teams a programmatic way to verify AI-generated content at scale — useful for brand integrity and compliance checks in content automation.
28B parameters, 262k native context (extensible to 1M), full multimodal (text, image, video), commercially deployable on the day of release.
Why it matters: Strong self-hosted candidate for teams where data residency, commercial licensing, or cost control rules out API-only options.