
Modern coding agents such as Claude Code, Cursor, and Codex are commonly described as 'AI pair programmers,' but the description hides the real engineering underneath. Under the surface, every such tool is a wrapper around the same machinery: an iterative loop that calls a language model, parses its output, executes a tool, and feeds the result back in. This article is a technical teardown of that machinery for developers who want to understand internals rather than just pick a product. The scope is deliberately narrow: the agent loop, context assembly, tool wiring, and the harness differences between Claude Code, Codex, and Cursor. It does not cover benchmarks, pricing, sandboxing, or multi-agent orchestration, which are addressed in sibling articles in this series.

A coding agent cannot answer a real request in a single inference call. Given a goal like "add JWT auth to the API," the model has to read existing files, decide what to change, write code, run tests, observe errors, and try again. None of that fits in one completion. So every coding agent — Claude Code, Cursor, Codex, Aider, Cline, VS Code Copilot — is built around the same machinery: the model proposes something, the harness executes it against the real environment, the result is fed back, and the loop runs again until a stop condition fires. Autocomplete can only suggest; an agent edits files and runs your shell because of this cycle.
The vocabulary for this loop comes from the 2022 paper ReAct: Synergizing Reasoning and Acting in Language Models by Shunyu Yao et al. (arXiv:2210.03629, a Princeton University and Google Research collaboration, later published at ICLR 2023). ReAct formalized the Thought-Action-Observation cycle that virtually every modern harness implements today:
The key contribution is interleaving reasoning and action rather than asking the model to plan everything up front.
In practice, one pass through the loop runs six concrete steps:
tool_calls.Anthropic describes Claude Code's runtime as a "dumb loop": the harness is intentionally simple and all reasoning lives in the model. Mechanically, the orchestration layer is often just a while statement; the complexity sits in everything around it (context compaction, permissions, hooks, subagent isolation). The harness does not need to be clever, only reliable at handing control back and forth between the model and its tools.
VS Code's agent documentation uses a parallel three-stage framing (VS Code agent concepts): Understand (read files, search the codebase, look up docs to see what needs to change), Act (modify code, run commands, call services), Validate (run tests, check compiler output, iterate if results fail). The stages are not separate sub-loops; they describe what the same ReAct cycle looks like at a higher level of abstraction.
This is the mental model to carry into the rest of the article: every harness, regardless of brand, is a wrapper around this same six-step loop.

The user prompt is never what the model actually receives. Every coding agent runs an assembly step before each model call that concatenates the user's request with relevant files, prior conversation turns, instruction documents, environment metadata, and tool definitions. This assembled payload is what allows grounded, repository-specific answers instead of generic completions (VS Code Agents docs).
Codex stores conversation history as ResponseItems and rebuilds each prompt from two parts:
AGENTS.md / user instruction files, and environment context.Rather than indexing the whole repo, Codex uses targeted tool reads (e.g., read_file with offset/limit) to pull in only what's needed. To keep the prompt within budget, the harness applies several techniques:
TokenCount events surfaced to the UI.When usage crosses a configurable threshold — default roughly 90% of the model context window, clamped to a minimum — Codex triggers auto-compaction. The compacted history is rewritten into three components: the initial context, recent user messages under a budget cap, and a handoff summary bridging the gap (OpenAI community thread on context management).
Claude Code's harness is a single TypeScript process that builds its prompt around a stable prefix: system instructions, tool schemas, and CLAUDE.md content. Prompt caching breakpoints are placed at the boundaries of that prefix so the model API can reuse already-processed tokens at substantially lower read cost across turns (Anthropic prompt caching docs). Claude Code compacts rarely and structurally — it trims tool outputs without a model call and only invokes a summarizer near the window limit, while preserving CLAUDE.md and project skills across the reset (Louis Bouchard, Context Engineering 2026).
Cursor markets a 200K context window, but multiple forum threads report usable context of roughly 70K–120K tokens after internal truncation. For routine tasks this is usually enough; for large multi-file refactors, Claude Code's reliable full-window delivery (200K on individual plans, 500K on Team/Enterprise) is reportedly the more predictable choice (Builder.io comparison, Cosmic JS comparison).
The practical takeaway: prompt caching, targeted retrieval, and threshold-triggered compaction are not optional polish — they are the engineering that turns a generic LLM into a useful coding agent.

A tool call is a JSON-structured request emitted by the model rather than free-form text. The model is fine-tuned to produce an object containing the tool name and its arguments; the harness parses that object, validates it against a JSON Schema, dispatches it to a registered function, runs the function in a sandbox, and returns the result as an "observation" string the model can read on the next turn. OpenAI shipped this pattern as function calling on June 13, 2023, and Anthropic's equivalent reached general availability on May 30, 2024; the mechanics are identical across providers (nango.dev).
In execution terms, a tool call maps cleanly onto a JSON-RPC-style request: a unique id, a method name (the tool), and a params object (the arguments). The harness is the dispatcher that authenticates, sandboxes, and routes that request to a real function — usually a shell process, a file-API call, or an HTTP request to an external service.
The taxonomy in arXiv 2604.03515v2 shows convergence across coding agents on four capability categories — read, search, edit, and execute — with a fifth validate category appearing only in Moatless Tools. In production harnesses these expand into roughly six families (lilianweng.github.io, agent-cookbook.com):
read, write, edit, multi_edit, apply_patch, glob, grep, lsbash, PowerShelllsp, git_status, git_diff, git_commitweb_search, web_fetch, browser toolsspawn_agent, resume_agent, wait_agent, list_agents, close_agent, interrupt_agentThe same paper distinguishes five strategies a scaffold uses to decide which tools the model sees:
built_tools model reconstructs the visible tool list each turn.For Claude Code, the external-tool layer is dominated by MCP, an Anthropic-developed client-server protocol with three primitives: Tools (model-controlled executable functions), Resources (readable data sources), and Prompts (reusable templates) (mcp.so, modelcontextprotocol.io). MCP servers are registered through a JSON configuration entry specifying command, args, env, and optional cwd or disabled flags, and the protocol is supported across Claude, ChatGPT, VS Code, and Cursor.
Two further Claude Code primitives complete the picture: slash commands (user-invoked actions such as /commit or /review) and hooks (event-driven shell scripts that fire on tool lifecycle events, typically PreToolUse and PostToolUse). Together with MCP, they form the three surfaces a Claude Code user can use to extend what the harness can do.

A subagent is an independent agent spawned by a main agent to handle a focused subtask, such as researching a topic or analyzing a function, and report the results back. The defining property is not what it can do but what it cannot see: it runs in its own context window and returns only a summary to the caller. Everything that happened inside, every file read, every search query, every intermediate token, stays sealed in that subagent's window. As the VS Code documentation puts it, "without subagents, every file read, search result, and intermediate step during research accumulates in the main agent's context window, potentially crowding out important information" (VS Code Agents docs). This makes subagents primarily a context-optimization technique rather than a capability upgrade. The work would, in principle, be possible in the main window; the subagent exists to keep the main conversation legible.
VS Code's reference behavior is explicit about the boundaries:
That synchronous contract is a deliberate design constraint: it prevents the main agent from hallucinating a placeholder while a research task is still in flight, at the cost of stalling the outer loop.
Claude Code keeps the same in-process mental model but optimizes it for terminal-driven, repo-wide work. Each subagent has its own independent context window, and Claude Code documents three sub-agent modes, including Fork (which duplicates parent context), Teammate (an independent window that communicates via a file-based mailbox), and Worktree (an independent Git branch). Because a subagent can be pinned to a different model — including cheaper ones for read-only exploration — one session can plan on a frontier model and delegate execution to faster variants (Anatomy of an AI Agent Harness).
Codex takes a different position entirely. Rather than running subagents in the same process, it spins up short-lived cloud VMs, clones the repo into each, and uses Git worktrees to queue separate tasks in parallel without merge conflicts. The output arrives later as a pull request. This is closer to a background-worker model than a pair-programmer model, and it sidesteps the context-window problem by isolating work at the filesystem level rather than the context level (daily.dev comparison).
In every harness, subagents are an architectural choice that trades latency for context cleanliness. You pay wall-clock time — sometimes tens of thousands of tokens of exploration inside the subagent, distilled down to a thousand-token summary — in exchange for a main window that stays focused on the task at hand.

Claude Code's harness, CLI, and model orchestration all live in one TypeScript process with one main loop calling Claude, executing tools, and feeding results back (Addy Osmani — Agent Harness Engineering; nimbalyst — OpenCode vs Claude Code). The system prompt is closed and only publicly known via Piebald-AI's reverse-engineering extraction: roughly 2,900 tokens assembled from around 110 conditional strings per session, tuned for Anthropic's own models, and not user-editable (Firecrawl — Claude Code vs OpenCode). The harness relies on a stable system-prompt prefix so Anthropic's prompt caching keeps input-token costs predictable, and exposes three primitives — slash commands, MCP, and hooks — plus per-subagent context windows that isolate exploration from the main loop (penchan — Coding Tools Comparison; alexop.dev — Four Types of Memory). The bet is speed inside one model family; the cost is rigidity — you cannot swap the model or fork the harness (nimbalyst — OpenCode vs Claude Code).
OpenAI's harness is built for asynchronous, parallel execution rather than live pair programming. Tasks run in short-lived cloud VMs, each cloned from the repo, and Codex uses Git worktrees to queue independent jobs in parallel without merge conflicts, then returns a pull request (daily.dev — AI Coding Agents Comparison; alekseialeinikov.com — Coding Agents 2026). Context is assembled from normalized ResponseItems plus targeted read_file calls — there is no full-repo embedding. Tool outputs are middle-truncated under a token/byte policy, function output items are budgeted, and auto-compaction fires when usage crosses roughly 90% of the model's context window, clamped and overridable (community.openai.com — Context Management). The codex-rs architecture is layered into prompt/context generators, loop and multi-agent graphs, and OS sandboxes (kenhuangus.substack.com — Harness Engineering).
Cursor is a VS Code fork with agents wired into the editor: Composer for multi-file plan-and-edit, Agent mode running a read-edit-run-react loop, MCP server support, and a model picker that can route per request to Claude, GPT, Gemini, or other frontier providers (futureagi.com — Best AI Coding Agents 2026; cosmicjs.com — Claude Code vs Copilot vs Cursor). Because the harness varies by chosen model, effective context and tool wiring also vary, and there are reports of advertised context being truncated in practice (daily.dev — AI Coding Agents Comparison). Cursor mitigates this with paging through large outputs and session tie-in to the running IDE (LinkedIn — Hierarchical Memory Management).
The three harnesses share a common skeleton — assemble context, call model, parse tool call, execute, feed back — and differ in what they optimize. The clearest contrast comes from comparing Claude Code with OpenCode: Claude Code's single-process loop avoids network hops and runs fast against its own models, while OpenCode's client-server split (Go TUI over HTTP to a Bun/JavaScript server) buys model swap, alternative clients, and remote attachment at the cost of a localhost hop per tool call (nimbalyst — OpenCode vs Claude Code; datacamp.com — What Is OpenCode). Cursor sits between them — IDE-native with model flexibility — and Codex leans the other way: server-managed VMs and background-worker semantics. None is a best design; they are different bets on the speed/flexibility/control axis.

The case that a coding agent needs a loop is short: any tool that solves real engineering work must iterate, because real engineering work cannot be solved in a single forward pass. SWE-bench Verified — the 500-instance, human-validated subset of SWE-bench curated by OpenAI and Princeton in 2024 — makes the point through its task shape. Each instance hands the model two inputs that cannot be concatenated into one prompt:
A patch counts as solved only when the project's own fail-to-pass test suite goes green. There is no shortcut: the agent must read the bug report, locate the relevant code in a codebase it has never seen, reproduce the failure, edit the right files, and re-run the tests until they pass. That sequence is the explore-edit-verify cycle, and it is what an agent loop exists to drive — one model call, one tool execution, and one observation fed back into the next prompt. No single-prompt interface, no matter how large the context window, can compress those steps out of existence.
A useful counterpoint to the assumption that loops must be elaborate is mini-SWE-agent. Released in July 2025, it is roughly 100 lines of Python and reaches 65% on SWE-bench Verified. The harness is deliberately bare — a single bash tool, a short system prompt, and the same iterate-until-done structure described above. The score matters less than the implication: the loop is the load-bearing component, not the surrounding scaffolding. If ~100 lines can extract most of the available signal, most of the engineering effort in commercial harnesses is going into convenience features (search, edit primitives, retries, parallel sub-agents) rather than into the loop itself.
This is also why SWE-bench-style benchmarks have displaced older generation-style tests such as HumanEval as the primary measure of coding-agent capability. HumanEval asks a model to complete an isolated Python function, fits trivially in a single forward pass, and has been saturated by frontier models. SWE-bench Verified demands the explore-edit-verify cycle by construction, which is why top models cluster above 80% on it as of mid-2026 and why the headline coding number in any new model launch now comes from this benchmark family rather than from single-function completion tests.

The harness is the "dumb loop" — most of the intelligence lives in the model, and the harness just manages turn-taking. But deciding when the loop ends is one of the few places where harness engineering actually matters. An agent that never stops is not autonomous; it is a bug with a token bill.
Across production harnesses, the same four conditions reappear, in slightly different shapes:
maxSteps config; when it fires, the agent is forced into a text-only summarization. Cline models its inner loop recursively with a reflection limit. Without this cap, a model that is confused will keep calling tools until something else stops it.Stop semantics are an active design surface. Anthropic's June 2026 loop taxonomy lists four loop types, not just stop checks: turn-based loops, goal-based loops triggered by /goal, time-based loops via /loop and /schedule, and proactive loops that fire without an explicit user prompt. Each loop type implies different defaults for when it should end, and each commits the harness to a particular failure mode if the goal is unreachable.
There is also a recurring design tradeoff: soft limits with human-in-the-loop escalation preserve long-running workflows but require interrupt infrastructure; hard limits guarantee budget control but may abort legitimate work. LangGraph models this as a configurable maximum iteration count functioning as a workflow circuit breaker, sitting above infrastructure-level rate limiting that cannot see internal reasoning.
The loop body is largely commoditized. What changes between Claude Code, Cursor, and Codex is less the prompt-to-tool plumbing and more which stop conditions are enforced, which can be overridden, and what the harness does when they fire. Choosing a coding agent is, in practice, choosing which stop semantics you are willing to trust with your budget.