Late August 2026 model routing: DeepSeek peak pricing, Qwen3.8 open weights, Grok on Bedrock, and Sol cost lanes
Between August 13 and August 26, 2026 the cheap and mid lanes moved again: DeepSeek GA plus peak/off-peak rates, Alibaba Qwen3.8 open weights, Grok 4.6 on Amazon Bedrock, OpenAI Ultrafast Sol, and Sol promotional pricing. Here is the updated routing board I use so Chinese open models, Bedrock governance, and Sol speed tiers do not collide.
In this post (7 sections)
Introduction
Early August was Grok Bot and Grok 4.6 in Cursor. Mid-to-late August is the routing board catching up: Chinese open weights you can self-host, DeepSeek rates that punish naive always-on queues, Grok inside AWS governance, and Sol becoming both faster and cheaper depending on which surface you use. This post updates August 2026 model routing: Grok, Qwen, DeepSeek with what shipped after August 13.
DeepSeek V4-Pro: GA plus schedule-aware pricing
DeepSeek moved V4-Pro to general availability with unchanged model name `deepseek-v4-pro`, native OpenAI Responses API compatibility, and low/high/max thinking effort. Peak and off-peak billing starts August 16. That turns queue design into a cost control, not an ops afterthought.
- Batchable worker jobs: prefer off-peak windows when latency SLOs allow.
- Interactive agents: budget peak rates explicitly in the spend dashboard.
- Do not assume the July flat rates still apply in your unit economics.
Source: DeepSeek API updates.
Qwen3.8 open weights: self-host without API lock-in
Alibaba's Qwen3.8-27B under Apache 2.0 plus flagship open weights on Hugging Face and ModelScope matters if your buyer constraint is data residency or offline worker farms. Treat 27B as a worker candidate. Keep larger frontier models for planner turns until your evals say otherwise. Pair with stop paying frontier prices.
Source: Alibaba Cloud Qwen3.8 announcement.
Grok 4.6 on Amazon Bedrock
August 19 put Grok 4.6 on Bedrock at $2 / $6 per million tokens (cached input $0.50) with 500K context and configurable reasoning effort. If your control plane is already AWS, this is how Grok enters the same IAM, logging, and private networking story as Claude on Bedrock. Enable in a sandbox account first. Compare against Cursor and Copilot Grok lanes you already tested after August 12–14.
Source: xAI Grok 4.6 on Bedrock.
OpenAI Sol: Ultrafast and promotional pricing
Ultrafast for GPT-5.6 Sol (preview, August 13) targets latency-critical loops. Later August promotional pricing cut Sol toward $4 / $20 per million tokens for a limited window. ChatGPT also received August Sol/Luna snapshots that can diverge from Codex and Work versions. Treat chat, API, and Codex as separate lanes until you pin snapshots: unified agent spend dashboard.
| Lane | Candidate | Watch-out |
|---|---|---|
| Cheap batch workers | DeepSeek V4 off-peak / Qwen3.8-27B | Peak rates and self-host ops |
| AWS-governed long agents | Grok 4.6 on Bedrock | Account enablement and eval parity |
| Latency-critical Sol | Ultrafast Sol | Preview capacity and quality |
| Frontier planners | Fable 5.1 / Sol / Opus 5 | Do not demote planners to Flash alone |
Action checklist
- 01DeepSeek schedulesAdd peak/off-peak columns to your DeepSeek cost model.
- 02Qwen worker benchBench Qwen3.8-27B on two worker tasks you already run on Flash or Luna.
- 03Bedrock Grok sandboxEnable Grok 4.6 in a Bedrock sandbox and compare to Cursor Grok quality.
- 04Pin Sol snapshotsPin Sol snapshots per product surface before celebrating promo pricing.
- 05Publish the boardUpdate the routing board doc your on-call actually reads.
Conclusion
Late August did not invent a new architecture. It changed the prices, hosts, and schedules under the architecture you already have. Rebuild the board with DeepSeek timing, Qwen self-host options, Bedrock Grok, and Sol speed/price tiers, then promote only what beats your harness.
Sources: api-docs.deepseek.com — updates ; alibabacloud.com ; x.ai — grok 4 6 amazon bedrock ; OpenAI ; OpenAI
Agentic AI patterns, delivered Thursdays
What I am shipping, watching, and pruning out of client stacks each week. One email. No fluff.