Deep-dive guides on AI agents, agent orchestration, MCP, and developer tooling.
7 posts found
Codex, Copilot coding agent, Cursor Background Agent, Jules, Devin, and cloud Claude Code sessions all follow the same five-stage pipeline — the differences that matter are in sandbox scope, repo access, and where the real trust boundary sits.
Tools and resources get all the attention in MCP, but sampling is the primitive that inverts the relationship — letting a server without its own model borrow the client's LLM through a two-gate approval flow.
When an LLM agent calls the wrong tool or sends malformed arguments, the postmortem usually blames the model — but the actual defect is almost always in the JSON Schema the tool was registered with.
In agent systems, instructions live across four surfaces — system prompt, tool schemas, tool results, and few-shot text — not one. Most prompt debugging still only looks at the first.
Long-running agents rarely fail because they run out of context window — they fail because nobody designed what happens to attention quality once the transcript outgrows what the model can usefully weigh. Here is how tiered compaction, tool-output pruning, and sub-agent isolation actually work.
The three dominant multi-agent orchestration topologies each fail in a different, predictable way once you move past the demo — here is how to pick one based on where your task actually breaks, not which pattern sounds more sophisticated.
Vector similarity search answers what text is topically related — but long-running agents need to know what is true right now. Conflating the two is why agents keep resurrecting overturned decisions.