Insights

NOTES FROM THE FORGE.

Field notes from shipping AI into real businesses — what works, what doesn't, and what's next.

STRATEGYFeb 20266 min read

Why most AI pilots never reach production

90% of pilots die between the demo and the deploy. The causes are predictable — and avoidable. Here's the checklist we run before any build.

Read article

AGENTSJan 20268 min read

The agent stack: what actually works in 2026

Planning loops, tool use, memory, evaluation. A pragmatic tour of the architecture patterns that survive production traffic.

Read article

ROIDec 20255 min read

A practical framework for AI ROI

How to price the value of an hour saved, an error caught, or a deal accelerated — before you spend a rupee on building.

Read article

CULTURENov 20257 min read

Human + AI: designing teams around intelligence

The companies winning at AI aren't replacing teams — they're redesigning work around what humans do best. Lessons from 40+ deployments.

Read article

Every failed pilot shares the same autopsy: it was built to impress, not to work. The demo had curated inputs, a friendly audience and no stakes. Production has none of those comforts.

Before we build anything, we run four checks: Is there a measurable number? If the system can't move a tracked metric, it's a toy. Is the data actually accessible? Most "AI problems" are plumbing problems in disguise. Who owns the failure mode? Someone must be accountable when the system is wrong. What's the rollback plan? Confidence comes from knowing you can turn it off.

Pilots that pass these four checks reach production at a startlingly higher rate — because they were never really pilots. They were first versions of real systems.

The most reliable agents in production are boringly simple: a strong model, a tight tool surface, and an evaluation harness that catches regressions before users do.

Planning: explicit step plans beat emergent reasoning for business tasks — auditability matters more than magic. Tools: fewer, well-described tools outperform sprawling toolboxes; every tool is a chance to be wrong. Memory: structured state in your database beats fuzzy long-term memory for operational work. Evaluation: if you can't measure the agent's task success rate, you don't have an agent — you have a liability.

The teams winning with agents treat them like employees: clear job descriptions, limited permissions, and performance reviews every week.

AI ROI is usually calculated with too much optimism and too little arithmetic. Our framework is deliberately dull. Value = hours saved × loaded hourly cost + error reduction × cost per error + revenue acceleration × margin. Nothing mystical.

Then we apply the haircut: automation rarely removes 100% of a task. A system that saves "an hour a day" usually saves 40 minutes. Great systems save the whole hour and improve quality — but you earn that in versions two and three, not version one.

The honest math still favors building. It just favors building the right thing, in the right order, with expectations set like an engineer's — not a marketer's.

Across 40+ deployments, the pattern is unambiguous: companies that design workflows for human + AI collaboration outperform companies that use AI as a headcount tool — on both output and retention.

The winning operating model has three layers. AI does the volume: the repetitive 70% where consistency beats creativity. Humans do the judgment: exceptions, relationships and decisions with real stakes. The system learns from every human correction — so the more it runs, the less it needs supervising.

The result isn't a smaller team. It's a team whose ceiling just moved.

YOUR NEXT COMPETITIVE ADVANTAGE MIGHT BE AI.

Let's turn an idea into an intelligent system that actually works.

Start a Project