From the field.
What Jigar learns building and training, shared as posts. Specifics over slogans.
Agent Plugins 1.0: package skills and MCP once, stop maintaining five client manifests
Agent Plugins 1.0.0 landed as an open standard on August 6, then hit GA in VS Code, Copilot CLI, and the Copilot app on August 12. One plugin.json can carry skills/ and optional mcp.json across launch clients, with vendor extras under reverse-domain namespaces. Here is how I package it, the mistakes I expect, and the honest gap that Claude Code is not on the launch client list.
Claude's early-August 2026 enterprise stack: Cowork Chrome, Compliance API, Auto mode, and self-hosted runners as one playbook
Between August 5 and August 12, 2026 Anthropic shipped five enterprise-facing controls that only make sense together: Cowork in Chrome, Compliance API for Cowork and Claude Code transcripts, Auto mode as the Pro/Max/Team default, self-hosted runners in your network, and inference hooks for DLP. Here is how I sequence the rollout for Team versus Enterprise.
OpenAI Daybreak Blue and Red: cyber models, agent sandboxes, and why GPT-5.6-Cyber stays out of production coding agents
On August 10, 2026 OpenAI expanded Daybreak: Blue for defensive work with frontier models, Red for purpose-trained cyber models including GPT-5.6-Cyber under identity checks and legal attestation. Pair that with Codex safer auto-review defaults and the July 31 sandbox escape checklist. Do not put Cyber in your production coding loop.
Claude cyber-eval incidents: the sandbox escape checklist I run before any offensive agent test
On July 30, 2026 Anthropic disclosed three incidents where Claude models reached live production systems during third-party cybersecurity evaluations. The root cause was a harness misconfiguration that left internet access open while prompts claimed there was none. Here is the defense-in-depth checklist I now require before any capture-the-flag or offensive agent eval.
Cursor's agent swarm economics: why frontier planners and cheap workers beat solo GPT bills
Cursor's July 2026 swarm write-up rebuilt SQLite in Rust from the docs and measured planner vs worker spend. Frontier-only GPT-5.5 hit about $10.5k; Opus 4.8 planning with Composer 2.5 workers landed near $1.3k at similar quality. Here is how I translate that into routing rules for production graphs.
MCP 2026-07-28 is live: day-one field notes after the cutover
July 28 shipped. MCP's fifth spec release makes the protocol stateless at the core, hardens OAuth, graduates Tasks and Apps into a versioned extensions framework, and updates Tier 1 SDKs. Here is what I verify in the first 48 hours that the pre-GA checklist could not prove until clients actually moved.
MCP GA is four days out: the final cutover checklist I run before July 28 goes live
The MCP 2026-07-28 final specification publishes July 28. You already did inventory, header security, and handle migration. These four days are cutover discipline: freeze pins, delete sticky sessions only after round-robin passes, confirm cacheScope with two identities, and staff an on-call window for clients still sending Mcp-Session-Id.
Graph engineering with Claude Code: the 14-step roadmap I use when linear agents stall
A straight-line agent is a degenerate graph. Claude Code dynamic workflows move orchestration into JavaScript so subagents fan out, verify, and converge without stuffing every intermediate result into one context window. Here is the 14-step roadmap I map onto production work, with topology diagrams, contracts, and the patterns I already ship.
Gemini 3.6 Flash shipped, 3.5 Pro did not: the routing checklist I run after July 21
On July 21, 2026 Google released Gemini 3.6 Flash ($1.50/$7.50 per million), 3.5 Flash-Lite ($0.30/$2.50, ~350 tok/s), and a limited-access 3.5 Flash Cyber via CodeMender. Gemini 3.5 Pro stays in partner testing. Here is how I re-benchmark cost per completed task against Sonnet 5 and Fable without betting the roadmap on Pro.