Topic Pillar

Enterprise AI Automation.Agents that replace operational work, not just write emails.

Enterprise automation is where agentic AI pays back fastest: operational workflows, support and service pipelines, ERP integrations, content-at-scale. The wins are concrete — hours saved per ticket, headcount that scales sub-linearly, decisions logged for audit. The patterns matter more than the model.

135 cluster pages· 45 posts· 7 notes· 82 updates· 1 events

Where automation pays back fastest

Operational triage (routing, processing, escalation), customer support agents that hit ticket systems directly, and content-plus-SEO pipelines (research → outline → write → review). These are workflows with clear inputs, measurable outputs, and decisions that can be logged.

Why IT services teams ship faster than expected

IT services teams have the right combination: existing ops workflows to automate, domain context to write good tool descriptions, and developer headcount to maintain the systems. The blocker is rarely capability — it is usually scoping the first agent narrowly enough to actually ship.

What we deliver on consulting engagements

Architecture design, production-grade implementation using Claude API and MCP, full observability (Langfuse), structured outputs (Pydantic), retry semantics, and a real handoff so your team can maintain and extend it. Working agents, not slides about agents.

45 blog posts

Deep dives on Enterprise AI Automation

Production

GitHub Copilot’s September 4 wave: GPT-6 Astra GA, Agent Merge preview, and how to govern the new model lanes

On September 4, 2026 GitHub made GPT-6 Astra generally available in Copilot and published a weekly release that expands Fable 5.1 and Gemini 3.8 Flash access while putting Agent Merge into public preview. This guide covers what changed and how enterprises should enable the new agentic coding lanes.

Sep 5, 202614 min
Read the post
Production

GPT-6 Astra for enterprise agents: Critical cyber designation, routing decisions, and the September 2026 rollout checklist

OpenAI released GPT-6 Astra on September 3, 2026 as its most capable broadly deployed model for coding, computer use, research, and long-horizon agent work. It is also OpenAI’s first model to reach Critical cybersecurity capability under the Preparedness Framework. This guide explains what changed, who should enable it, and how to route it safely against Sol and Fable 5.1.

Sep 4, 202616 min
Read the post
Production

Muse Spark 1.3 for agentic coding: what Meta changed, how to route it, and when it belongs in your stack

Meta released Muse Spark 1.3 on September 2, 2026 with stronger long-horizon agent collaboration, cleaner coding style, and lower tool-call overhead. The model is available through Muse Code and the Meta Model API. This guide explains the practical routing implications for platform and engineering teams.

Sep 3, 202614 min
Read the post
Production

Claude Fable 5.1 migration checklist: cache economics, breaking tool_choice rules, and when Mythos 5.1 actually matters

On September 1, 2026 Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1 for long-running coding and knowledge work. Same list prices as Fable 5 except cache reads drop 75% to $0.25 per million tokens. Forced tool use breaks. Thinking blocks bind tighter. Here is the migration sequence I run before flipping planner lanes.

Sep 2, 202617 min
Read the post
Production

Gemini 3.8 Flash and Fairwind Cyber: the September 2026 worker-lane and defender-lane checklist

Google shipped Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2, 2026, the third Flash release in six weeks. Standard 3.8 Flash stays at the $0.75 / $3.75 introductory rate through year-end. Flash Cyber and the Fairwind Program gate vulnerability discovery and patching for trusted defenders. Here is how I update routing boards without treating Cyber as a casual coding model.

Sep 2, 202615 min
Read the post
Production

Late August 2026 model routing: DeepSeek peak pricing, Qwen3.8 open weights, Grok on Bedrock, and Sol cost lanes

Between August 13 and August 26, 2026 the cheap and mid lanes moved again: DeepSeek GA plus peak/off-peak rates, Alibaba Qwen3.8 open weights, Grok 4.6 on Amazon Bedrock, OpenAI Ultrafast Sol, and Sol promotional pricing. Here is the updated routing board I use so Chinese open models, Bedrock governance, and Sol speed tiers do not collide.

Aug 26, 202615 min
Read the post
Production

Grok Bot as always-on teammates on a shared cloud computer: the August 2026 playbook

On August 11, 2026, xAI opened Grok Bot early beta: always-on AI teammates with a persistent cloud computer, available on SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium (desktop and iOS). The product pitch is finishing work in the real apps. The production detail that matters is sharper: every Bot you create shares one user-scoped cloud computer (files, browser sessions, app logins). Here is how I hand off real jobs, set request boundaries and Auto Review, and refuse to treat separate Bots as a security boundary.

Aug 13, 202618 min
Read the post
MCP

Agent Plugins 1.0: package skills and MCP once, stop maintaining five client manifests

Agent Plugins 1.0.0 landed as an open standard on August 6, then hit GA in VS Code, Copilot CLI, and the Copilot app on August 12. One plugin.json can carry skills/ and optional mcp.json across launch clients, with vendor extras under reverse-domain namespaces. Here is how I package it, the mistakes I expect, and the honest gap that Claude Code is not on the launch client list.

Aug 12, 202615 min
Read the post
Production

Claude's early-August 2026 enterprise stack: Cowork Chrome, Compliance API, Auto mode, and self-hosted runners as one playbook

Between August 5 and August 12, 2026 Anthropic shipped five enterprise-facing controls that only make sense together: Cowork in Chrome, Compliance API for Cowork and Claude Code transcripts, Auto mode as the Pro/Max/Team default, self-hosted runners in your network, and inference hooks for DLP. Here is how I sequence the rollout for Team versus Enterprise.

Aug 12, 202619 min
Read the post
Production

OpenAI Daybreak Blue and Red: cyber models, agent sandboxes, and why GPT-5.6-Cyber stays out of production coding agents

On August 10, 2026 OpenAI expanded Daybreak: Blue for defensive work with frontier models, Red for purpose-trained cyber models including GPT-5.6-Cyber under identity checks and legal attestation. Pair that with Codex safer auto-review defaults and the July 31 sandbox escape checklist. Do not put Cyber in your production coding loop.

Aug 10, 202617 min
Read the post
Production

Claude cyber-eval incidents: the sandbox escape checklist I run before any offensive agent test

On July 30, 2026 Anthropic disclosed three incidents where Claude models reached live production systems during third-party cybersecurity evaluations. The root cause was a harness misconfiguration that left internet access open while prompts claimed there was none. Here is the defense-in-depth checklist I now require before any capture-the-flag or offensive agent eval.

Jul 31, 202616 min
Read the post
Multi-Agent

Cursor's agent swarm economics: why frontier planners and cheap workers beat solo GPT bills

Cursor's July 2026 swarm write-up rebuilt SQLite in Rust from the docs and measured planner vs worker spend. Frontier-only GPT-5.5 hit about $10.5k; Opus 4.8 planning with Composer 2.5 workers landed near $1.3k at similar quality. Here is how I translate that into routing rules for production graphs.

Jul 29, 202615 min
Read the post
Architecture

Graph engineering with Claude Code: the 14-step roadmap I use when linear agents stall

A straight-line agent is a degenerate graph. Claude Code dynamic workflows move orchestration into JavaScript so subagents fan out, verify, and converge without stuffing every intermediate result into one context window. Here is the 14-step roadmap I map onto production work, with topology diagrams, contracts, and the patterns I already ship.

Jul 22, 202622 min
Read the post
Production

Gemini 3.6 Flash shipped, 3.5 Pro did not: the routing checklist I run after July 21

On July 21, 2026 Google released Gemini 3.6 Flash ($1.50/$7.50 per million), 3.5 Flash-Lite ($0.30/$2.50, ~350 tok/s), and a limited-access 3.5 Flash Cyber via CodeMender. Gemini 3.5 Pro stays in partner testing. Here is how I re-benchmark cost per completed task against Sonnet 5 and Fable without betting the roadmap on Pro.

Jul 21, 202614 min
Read the post
Production

Claude Code Week 29: live MCP artifacts change the blast radius of a shared dashboard

Claude Code Week 29 (July 13–17, v2.1.207–v2.1.212) lets published artifacts call MCP connectors on each view through the viewer's own connections, with first-call approval. Public sharing links, editor roles, screen reader mode, and auto-background for long MCP calls shipped in the same window. Treat live artifacts like production integrations, not pretty exports.

Jul 17, 202612 min
Read the post
MCP

MCP just got a rival enterprise protocol: the portability playbook I run before teams pick a side

In mid-July 2026, reporting put Google, Microsoft, Salesforce, Snowflake, and ServiceNow behind a shared enterprise agent backend protocol framed as a counter to Anthropic's MCP. Protocol wars are procurement stories dressed as plumbing. Here is how I keep tool contracts, auth, and gateways portable while MCP 2026-07-28 still ships on July 28.

Jul 16, 202613 min
Read the post
Production

Claude Code Week 28: the desktop browser and /doctor checklist I run before enabling it org-wide

Claude Code Week 28 (July 6–10, releases v2.1.202 through v2.1.206) shipped a sandboxed in-app browser on Desktop and upgraded /doctor from a read-only report into a repair tool. Auto mode also blocks transcript tampering. Here is the governance checklist I run before teams turn browsing and auto-fix on for every engineer.

Jul 14, 202612 min
Read the post
Production

Claude Cowork is not Claude Code for civilians: the knowledge-worker playbook after the mobile launch

Anthropic shipped Claude Cowork on mobile and web July 8, starting with Max subscribers. Usage data from 1.2 million sessions shows more than 90% of Cowork work is non-technical: memos, RFP reviews, inbox triage, decks. Tasks run in the cloud, sync across devices, and continue when you close the app. Here is how I govern Cowork without treating it like a coding agent.

Jul 8, 202614 min
Read the post
Production

Claude Sonnet 5 is the new default: the migration checklist I run before swapping claude-sonnet-4-6 in production

Anthropic shipped Sonnet 5 on June 30 as default on Free and Pro chat and as claude-sonnet-5 on the Platform API at $2/$10 per million through August 31. Three breaking changes matter for agent builders: adaptive thinking on by default, manual extended thinking removed, and non-default sampling parameters return 400. The tokenizer adds roughly 30% tokens for the same text. Here is the swap checklist.

Jul 5, 202614 min
Read the post
Production

Fable 5 is back globally: the redeployment routing checklist I run before July 7

Anthropic restored Claude Fable 5 worldwide on July 1 after US export controls lifted June 30. Pro through Enterprise plans get up to 50% of weekly usage limits included through July 7, then usage credits. A new safety classifier blocks the Amazon-reported jailbreak in 99%+ of cases. Here is how I re-promote Fable without repeating June's config chaos.

Jul 4, 202614 min
Read the post
Architecture

Codex Record and Replay turns one demo into a Computer Use skill: how I inspect generated skills before trusting them unattended

Codex app 26.616 adds Record and Replay on macOS: perform a workflow once, Codex packages it into a skill you replay with different inputs. Thread handoff and automation run history ship alongside. Computer Use must be enabled. Here is the review checklist I run before any recorded skill runs unattended.

Jul 3, 202613 min
Read the post
Production

One agent spend dashboard for Cursor, Claude Code, and Copilot: what Copilot's ai_credits_used field unlocks

GitHub added ai_credits_used to the Copilot usage metrics API on June 19. It is a per-user total, not yet split by feature or model, but it closes a gap I have been papering over with spreadsheets. Here is how I unify Copilot credits with Claude API keys and Cursor team usage into one attributable spend view.

Jul 1, 202613 min
Read the post
Production

Fable 5 got suspended worldwide in three days: the frontier model adoption checklist I now run on every engagement

Anthropic launched Claude Fable 5 on June 9, promised included subscription access through June 22, then suspended Fable and Mythos 5 globally on June 12 after a US export-control directive. Most teams never finished piloting. Here is the governance checklist I use when capability, retention, pricing, and access can flip faster than your rollout calendar.

Jun 25, 202614 min
Read the post
Production

Claude Code Artifacts turn terminal output into live review pages: what Team and Enterprise buyers should pilot first

Artifacts in Claude Code beta publish self-contained HTML to claude.ai that republishes to the same URL as the session progresses, with version history and org-only sharing. Strict CSP, no external fetch, no backend. Requires Team or Enterprise and claude.ai login. Here is the workflow I use for PR walkthroughs and incident timelines without screenshot threads in Slack.

Jun 22, 202613 min
Read the post
MCP

MCP Enterprise-Managed Authorization is stable: how IdP-provisioned connector access replaces per-server OAuth hell

EMA makes the organization IdP the decision-maker for which MCP servers a user can reach. Admins enable connectors once; clients exchange an Identity Assertion JWT for scoped tokens without redirecting every employee through OAuth per server. Anthropic ships it across Claude, Claude Code, and Cowork; VS Code supports it; Okta is the first IdP. Here is the pilot I run before July 28 stateless transport work lands.

Jun 19, 202614 min
Read the post
Architecture

Cursor cloud subagents in 2026: /in-cloud, /babysit, and /automate without losing your local guardrails

Cursor 3.7 lets you spin subagents in cloud VMs with /in-cloud, iterate on a PR until merge-ready with /babysit, and hand off between local and cloud sessions. Cursor 3.8 adds /automate and five GitHub review triggers. Here is the workflow I use so parallel cloud work does not bypass Auto-review, environment snapshots, or pre-push /review.

Jun 18, 202613 min
Read the post
Production

Agentjacking is real: poisoned Sentry errors can hijack Cursor, Claude Code, and Codex without touching your repo

Tenet Threat Labs injected a fake stack trace through a public Sentry DSN and watched 100+ coding agents execute attacker commands during normal triage. No git write access required. The agent treats the error as ground truth. Here is how I harden observability MCP feeds, scope triage prompts, and block auto-exec on untrusted telemetry.

Jun 17, 202613 min
Read the post
Production

The June 15 Claude billing change: Agent SDK credits, model retirement, and the checklist I run before anything breaks

Two Anthropic changes land on the same day: programmatic Claude usage moves to a separate monthly credit pool, and claude-opus-4-20250514 plus claude-sonnet-4-20250514 stop answering on the API. Interactive Claude Code is fine. Cron jobs and CI agents are not. Here is how I audit auth paths, claim credits, and grep for retiring model IDs before the first failed run.

Jun 15, 202614 min
Read the post
Production

Governing agent autonomy in 2026: Auto-review, permissions.json, and pre-push review

Cursor Auto-review uses permissions.json autoRun.allow_instructions and block_instructions so low-stakes shell/MCP calls flow and high-stakes actions pause. Here is how I wire that pattern into local agents, SDK headless runs, and CI — without mistaking a classifier for a hard security boundary.

Jun 11, 202614 min
Read the post
Architecture

Agentic RAG vs vanilla RAG: why a Sufficient Context Agent beats retrieve-then-pray

Google Research shipped Agentic RAG on Gemini Enterprise with a Sufficient Context Agent that refuses to answer when retrieval is incomplete. On factuality benchmarks they report up to 34% higher accuracy versus standard RAG. Here is when one-shot RAG is still enough, when you need iterative retrieval, and how I wire the pattern without blowing latency budgets.

Jun 6, 202614 min
Read the post
Production

Agentic transformation is an operating-model problem, not a model problem

Microsoft published a 6-step playbook for rolling agents out across an enterprise, and the line that matters is "you do not need a bigger model, you need a better operating model." That matches what I see in consulting: the pilots that die do not die on model quality, they die on ownership, evals, and governance. Here is how I read the playbook for IT services teams, and the operating-model gaps that actually stall agent rollouts.

Jun 4, 202611 min
Read the post
Production

Your agent's supply chain is the attack surface now

A poisoned VS Code extension spent eighteen minutes on the marketplace and walked off with Claude Code credentials and MCP configs. The model was never the target. Your agent's supply chain is: the extensions, skills, MCP servers, tool definitions, and keys it is allowed to touch. Here is how I harden all four layers, and the checklist I run on every deployment.

May 27, 202612 min
Read the post
Multi-Agent

Inside Recruiting Atelier: a runnable reference for the primitives of an agentic system

A working open studio that vets duplicates, plans the run, screens, scores, shortlists, and notifies. The whole pipeline lives in roughly ninety lines of supervisor code and a tool registry you can read in one sitting. Here is what is inside, why every piece is there, and what you can copy into your own stack.

May 24, 202614 min
Read the post
Production

How an agentic studio screens, scores and shortlists candidates for your hiring team

Open Recruiting Atelier and you do not see a generic AI dashboard. You see five named specialists doing the work a screening team would do: catching duplicates, checking the brief, scoring on four dimensions, ranking, drafting the dispatch. Drop one CV or fifty. Click any candidate to see exactly why they landed where they did. This is what AI for recruitment looks like when it respects your judgment instead of replacing it.

May 24, 202610 min
Read the post
Architecture

Code agents vs skill agents: when to give an agent the keyboard and when to give it the toolbox

Two ways to let an agent act in the world. Code agents write fresh code into a sandbox. Skill agents pick from a curated menu. The choice should be made in the kickoff, not the postmortem. Here is the framing I use with clients, the four axes where they diverge, and the hybrid pattern most production systems become.

May 22, 202611 min
Read the post
Tool Design

Tool registry design for agentic AI: how the wrong registry kills accuracy before the prompt is read

I reviewed a system last month with 47 tools in its registry and a 22 percent wrong-tool-selection rate. The team was about to migrate from Sonnet to Opus to fix it. The prompt was fine. The registry was the bug. This is the audit pattern I run on every client codebase before we change anything else, the seven failure modes I see in production, and the numbers from the cleanup.

May 22, 202612 min
Read the post
Architecture

AI agent vs agentic AI: what the distinction actually means when you ship one

Vendors blur the line because "agentic" sells. The two terms describe different architectures, with different cost shapes, different observability needs, and different scoping conversations. Here is the framing I use with clients and the three-question test for which one your project actually needs.

May 22, 202612 min
Read the post
Production

MCP governance just became a product: what Databricks Unity AI Gateway changes for enterprise agents

Every enterprise MCP deployment I have audited in the last six months has been hand-rolling tool-access policy, payload logging, and per-team cost limits on top of a gateway someone wrote in two days. Databricks just shipped that as a product. Here is what it actually changes, where the gaps still are, and the migration I would run for a Databricks shop.

May 20, 202612 min
Read the post
Production

The cheapest LLM call is the one you do not make. GitHub's 19-62% token cut, decoded

GitHub published an instrumented analysis of their agentic CI workflows and reported 19-62% token-cost reductions. The savings are the headline. The technique (pre-agentic data fetching and tool-registry hygiene) is the story most teams will miss.

May 11, 20269 min
Read the post
MCP

MCP 1.0 is here. What changes for the servers you already wrote

The protocol stabilised. Most working servers will keep working. Three places the new spec actually requires changes (auth profile, server registry, streaming-response semantics) with diffs from a real migration.

May 1, 20268 min
Read the post
Multi-Agent

Why I am replacing supervisor patterns with handoffs

Supervisors looked clean on paper and shipped slow in production. Handoffs read messier in the code but recover better when an agent loses the plot. Two real systems and where supervisors still earn their keep.

Apr 26, 20268 min
Read the post
Production

Prompt caching is not optional anymore. Measuring a 47% cost drop

A walkthrough from a client engagement: identifying stable prefixes, restructuring the system prompt for cacheability, and the telemetry that proved caching was actually working.

Apr 19, 20267 min
Read the post
Production

The agent observability stack we ship to every client

Traces, spans, evals, cost-per-completed-task, and the one dashboard panel that catches 80% of regressions. Vendor-agnostic; covers Langfuse, Honeycomb, and rolling your own.

Mar 28, 20268 min
Read the post
Multi-Agent

Haiku 4.5 made our router 5x cheaper. The trade-off matters

Replacing Sonnet with Haiku in the dispatcher role cut our orchestration cost dramatically. It also cost us in two specific places I did not predict.

Feb 22, 20267 min
Read the post
Production

Eval datasets: stop testing your agents on the happy path

If your eval set is the demos you showed the client, you are testing the wrong thing. How we build evals from production failures and the minimum viable suite to ship.

Jan 19, 20268 min
Read the post
82 ship-news updates

Latest in Enterprise AI Automation

Tools

GPT-6 Astra becomes generally available in GitHub Copilot for Pro+, Max, Business, and Enterprise

September 4, 2026 · via GitHub
Tools

GitHub Copilot adds Agent Merge public preview and expands Fable 5.1 and Gemini 3.8 Flash model choice

September 4, 2026 · via GitHub
Tools

Google ships Gemini 3.8 Flash for agentic coding at the same introductory price as 3.7 Flash

September 2, 2026 · via Google
Architecture

Google launches Fairwind with Gemini 3.8 Flash Cyber for trusted defender vulnerability and patching agents

September 2, 2026 · via Google
Tools

Meta releases Muse Spark 1.3 for longer-horizon agentic workflows and more efficient coding agents

September 2, 2026 · via Meta
Claude

Anthropic announces Enterprise Frontier Safeguards to pair zero data retention with frontier misuse controls

September 1, 2026 · via Anthropic
Tools

DeepSeek launches experimental V4-Flash-Vision-Exp for multimodal agent workflows on the public API

August 21, 2026 · via DeepSeek
Open Source

NVIDIA Megatron Core 0.19.0 adds DeepSeek V4 HybridModel paths, MLA, DSA, and MXFP8 training support

August 19, 2026 · via NVIDIA
Research

Anthropic reports Claude designed protein binders against 14 of 15 targets with 22 to 35% hit rates

August 18, 2026 · via Anthropic
Frequently asked

Enterprise AI Automation — the questions teams actually ask

Go deeper on this topic

New breakdowns on this and related agentic AI topics, plus what I am shipping for clients — one email on Thursdays.