An agent harness is often described as "the thing around the model," but the engineering substance lives in a single iteration of a while loop. Every Claude Code session, every OpenDev turn, and every LangGraph run reduces to the same mechanical cycle executed by the same set of components. This article walks through that cycle one step at a time, using Claude Code and OpenDev as concrete reference points. The goal is to expose the unit of execution that sits beneath every agent framework — the discrete, observable sequence a harness repeats from a user message to a final answer. Readers should leave with a precise mental model of prompt assembly, model call, tool dispatch, state mutation, and telemetry capture as distinct, addressable operations.
An agent loop can only run if the agent already exists. Before the first user message reaches the model, the harness must produce a fully assembled agent object — one whose system prompt is compiled, tool schemas are rendered, subagents are registered, and permission context is loaded. OpenDev and Claude Code both pay this one-time setup cost up front so that the subsequent while not done loop can dispatch its first iteration with no assembly latency and no nondeterministic prompt state (arxiv.org).
BaseAgentOpenDev enforces the discipline through its type foundation. Every agent inherits from a BaseAgent abstract class that accepts (config, tool_registry, mode_manager) and declares four abstract methods: build_system_prompt(), build_tool_schemas(), call_llm(), and run_sync(). The design choice that matters is eager construction: BaseAgent.__init__() invokes both build_system_prompt() and build_tool_schemas() before the constructor returns. When __init__ completes, the agent is already ready to serve requests. A concrete refresh_tools() method re-runs both builders if the tool registry changes.
System-prompt assembly is itself a four-step pipeline (PromptComposer): filter each section against an environment snapshot, sort survivors by priority, load markdown files and resolve ${VAR} placeholders, and join the result with core role text and a dynamic environment block. The default action-mode agent produces five functional tiers — core identity, tool definitions, safety and rules, environment context, and task framing — all materialised before the loop starts.
Claude Code implements the same discipline, but distributed across its session bootstrap. Reverse engineering shows its prompt is composed of ~110 conditional strings, each evaluated against the session environment and assembled into a single ~2,900-token prompt that is not user-editable. In parallel, the bootstrap loads MCP server tool schemas into the toolset, indexes CLAUDE.md files (which are re-read on every turn), and reads permission grants from project configuration before exposing any tool to the model. Claude Code gates roughly 40 discrete tool capabilities through a three-stage model — trust establishment at project load, per-call permission check, and explicit user confirmation for high-risk operations — all of which must be configured before the first tool call (firecrawl.dev).
Treat scaffolding as a phase distinct from execution. Lazy assembly inside the cycle body makes prompt state nondeterministic across turns and breaks cache-key stability for prompt caching, since any re-evaluation can shift the prefix the provider hashes. Stable prefixes are exactly what enables Claude Code's prompt caching (ENABLE_PROMPT_CACHING_1H) to amortise input cost over a session, so any code path that mutates prompt composition per turn directly undermines that optimisation.
Every agent loop begins by reconstructing the prompt the model will see. Assembly is hierarchical — each layer slots into a defined position rather than being string-concatenated at the end. The widely cited stack is system prompt, tool definitions, memory files, conversation history, and current user message.
OpenAI Codex implements this as a strict priority ordering. From highest to lowest, Codex layers: server-controlled system message, tool definitions, developer instructions, user instructions, and conversation history (Agent Harness Components; Anatomy of an AI Agent Harness). Lower-priority layers cannot override higher-priority ones. Claude Code takes a different shape: reverse engineering shows roughly 110 conditional strings merged per session, producing a core prompt of about 2,900 tokens (Claude Code vs OpenCode comparison).
CLAUDE.md is reloaded every turn. Its semantic memory — the CLAUDE.md hierarchy with user-level (~/.claude/CLAUDE.md) and project-level (./CLAUDE.md) files — is not loaded once per session. It is injected into the prompt on every turn, so the size of project context files multiplies directly across the loop (Four Types of Memory in AI Coding Agents; Claude Code token optimization). A 200-line project CLAUDE.md therefore pays for itself again on every iteration of the cycle.
MCP tool schemas and registered skills are unconditional. Every connected MCP server injects its tool definitions into the prompt whether or not the agent will invoke them. Anthropic documents a typical five-server MCP setup at roughly 55,000 tokens of tool definitions, and broader measurements report 45–50k tokens with four to five servers and as much as 72% of the context window consumed before the user sends a single message (MCP vs Tool Calls for AI Agents; LangChain vs MCP). Registered skills occupy the same fate: they sit in context even when never invoked. The contrast with on-demand skills — which load via SKILL.md only when a tool call references them — illustrates that the harness, not the protocol, decides what is unconditional overhead.
Because the prompt is the unit the model actually reasons over, a harness that hides assembly is a harness that cannot be debugged. The precondition for diagnosing drift is the ability to print the exact bytes the model sees at turn N, and to diff them against turn N − 1. Without that diff, a developer cannot tell whether a behavior change came from new conversation content, a reloaded memory file, a newly connected MCP server, or a system-prompt mutation — and the loop becomes a black box whose internals can only be inferred from outputs.
Once the harness has assembled the prompt, the entire conversation state — system instructions, tool schemas, history, scratchpad — leaves the harness boundary as one HTTP request and returns as one structured response. There is no second chance to reshape the input mid-flight; whatever reaches the model is what the model sees.
The spectrum of model capability. Claude Code ships a closed prompt tuned for a single vendor and treats the Anthropic Messages API as the canonical surface; tool calls must match its chat template exactly, otherwise production-blocking serialization bugs appear. OpenDev, by contrast, binds five distinct LLM roles — Action, Thinking, Vision, Compact, and Embedding — to user-configurable providers and falls back through a capability registry when a role is unavailable. This contrast illustrates the two ends of the spectrum: a closed, hand-tuned single-model binding on one side, and an open, role-routed compound system on the other. Either way, the call itself is the same shape: prompt in, structured message out.
Native tool calling. Modern harnesses do not parse free text for Action: … or json fences. The model emits tool_calls objects whose names and argument schemas are constrained at decode time, often by compiling the JSON schema into a state machine that masks invalid tokens. This eliminates the brittle regex-parsing layer that older agent frameworks were notorious for. OpenAI's structured outputs, LangChain's Pydantic output parsers, and Anthropic's tool-use blocks all push validation into the schema layer rather than into post-hoc string matching, raising compliance with declared shapes to effectively100%.
Three-branch classification. After the response lands, the harness classifies it into one of three branches that determine what happens next:
tool_calls — the loop continues; results will be fed into the next inference.Where the harness earns its keep. Anthropic frames the runtime as a "dumb loop": all intelligence lives in the model, and the harness only manages turns. The classification step is exactly where that turn management pays off. A misclassification — treating a tool call as prose, or missing a handoff signal — silently corrupts every subsequent turn, because the wrong branch's state-update logic runs and the loop carries the error forward. The harness does not need to be intelligent, but it must be reliable at this one handoff.
Once the model returns a tool_call, the harness does not execute it immediately. Instead, it walks an ordered approval chain in which the first match wins, so the sequence of checks is itself the policy. A representative ladder runs in this order:
/permissions command or a prior "always allow" prompt).read, edit, execute, mcp, or other) checked against the standing session-level policy.A "YOLO" or auto-approve mode is typically inserted between the deny gate and the session-grant check, so anything that survives the deny list is approved without prompting unless the user has set stricter rules.
Claude Code gates approximately 40 discrete tool capabilities independently and applies them across three lifecycle stages, as described in analyses of Anthropic's harness architecture:
CLAUDE.md rules, and hook configuration are read once when the session starts.A load-bearing design choice is that permission enforcement is architecturally separate from model reasoning: the model decides what to attempt, but the tool system decides what is allowed. A hallucinated or compromised tool_call still hits a hard boundary, because no amount of persuasive text in the model's output can rewrite the policy chain.
OpenDev reaches a similar guarantee through a different mechanism. Rather than muting tools at runtime, it spawns a Planner subagent whose tool schema simply omits write tools. The LLM never sees definitions it cannot use, so write attempts during planning are eliminated at the schema level rather than blocked by a permission check. The main agent calls spawn_subagent(type="Planner"), the subagent runs in an isolated context with read-only tools, and only its returned plan enters the parent loop. This pattern generalizes: any subagent in OpenDev (Code-Explorer, PR-Reviewer, Security-Reviewer) is constructed eagerly with a tool schema tailored to its role.
Once a tool call is approved, the harness dispatches it inside a sandbox. Best practice is for tools to call an abstraction like sandbox.exec() rather than child_process.exec() directly, so the backend can be swapped (local, in-memory copy-on-write, or remote VM) without changing tool code. Each sandbox is created, executed against, and torn down independently, with strict limits on CPU, memory, filesystem scope, and network access.
After execution, the harness captures raw output and runs any configured hooks. PreToolUse hooks handle validation, authorization, rate limiting, and audit logging before side effects occur; PostToolUse hooks handle output filtering, quality checks, and follow-up triggers. In Claude Code, the hooks system is the primary extensibility point for organizational policy: teams can run security scans, linters, or secrets checks at specific lifecycle stages without modifying the core agent.
Every step in this path — approval verdict, sandbox ID, raw input, raw output, hook return codes, latency — is emitted to the telemetry stream. Treating telemetry as a structured side product of the loop (rather than a debug afterthought) is what makes the cycle observable, replayable, and auditable downstream.
The closing move of every cycle iteration is a translation step. The raw return value of a tool — a file's bytes, a JSON blob from an API, a stdout stream, a stack trace — is not a legal input to a chat model. The harness converts each tool's output into a message-shaped object with a uniform structure, typically a tool_result content block paired with the originating tool_use id. Because successes and errors share the same envelope, the model sees a single, homogeneous observation stream regardless of what actually happened inside the sandbox. The Anthropic Messages API and most other chat-completion interfaces require this pairing, and a missing or malformed pair is one of the most common causes of API rejections mid-loop.
When a tool throws — a 404, a permission denial, a timeout, an unparseable argument — the harness does not swallow the exception and continue. It catches the failure at the harness layer and repackages it as an error result addressed to the same tool_use id. This is what allows the loop to recover from transient failures: the model reads the error message in the next inference just as it would read any other observation, and can issue a corrected call on the following turn. Silent error handling collapses this feedback channel and turns recoverable mistakes into permanent stalls.
Once packaged, the new messages are appended to the conversation history. Three classes of state survive the cycle boundary and feed the next prompt assembly: the canonical message list, any scratchpad or working-memory notes the harness maintains, and the artifacts that tools wrote to disk or external stores. In Claude Code, for example, the most recently accessed files are also re-injected so the model retains a working set beyond the literal transcript.
As the message list grows, the harness must keep the next request under the model's context budget. In the Extended ReAct Loop, OpenDev treats this as a discrete sub-phase that runs before the next inference rather than as a panic-mode emergency at the limit. Its Adaptive Context Compaction is staged: at roughly 70% utilization a warning is logged; at 80% older tool-result messages are replaced in place with compact reference pointers such as [output offloaded to scratch file]; later stages apply progressively heavier summarization. Anthropic's API exposes a similar capability through the compact-2026-01-12 beta, which inserts a typed compaction block once a configured token threshold is crossed and discards everything preceding it on the next request.
Compaction is lossy by construction, and silent compaction is the single most common source of context rot in long-running agents. Production harnesses therefore surface compaction events in telemetry: what was summarized, what was dropped, what tokens remain. A developer reading the trace should be able to answer, at any point in the loop, what the model currently sees and what has been elided. Without that visibility, debugging a degraded run becomes archaeology.
The harness's while loop is mechanically simple, but its termination predicate is the most safety-critical decision in the entire cycle. Anthropic's "dumb loop" framing captures the truth here: the harness does not ask the model whether it is done. Termination is driven by explicit, harness-enforced conditions, never by model judgment.
OpenDev's Extended ReAct loop runs four phases per turn — staged context compaction, an optional thinking phase, an optional self-critique phase, and the standard ReAct step — and only after those four phases does the harness evaluate whether to exit. The exit predicate is a fixed disjunction of five cases:
maxSteps config forces a text-only response once the iteration cap is hit, then prompts the model to summarize and recommend remaining work (opencode.ai/docs/agents).A sixth, often-implicit dimension is wall-clock and monetary limits. Production harnesses should enforce step, time, token, and cost budgets on every loop and terminate gracefully with a structured failure rather than hanging (maa1.medium.com).
The Codex-style Item/Turn/Thread protocol exposes each of these transitions as a typed JSON-RPC message, which is why deterministic stop conditions are a non-negotiable design rule. Because every cycle is observable on the wire, an ambiguous termination predicate becomes detectable on every client surface (github.com/ai-boost/awesome-harness-engineering).
The failure mode of getting it wrong is severe. Without explicit termination, a misbehaving model — looping on the same tool call, fishing for confirmation, or refusing to commit to a final answer — can run indefinitely, drain token budgets, and corrupt session state. The classic pattern to guard against is the "same tool call twice" rule. Issuing an identical call with identical parameters twice in a row is a strong stuck-indicator: the model is not gathering new information, it is spinning. At that point the harness should escalate to the user rather than continue (pristren.com). Treat iteration caps as a critical rule, not a tuning knob: an agent that never stops is not autonomous — it is a bug with a token bill.
The harness's while loop is invisible from the outside — every iteration collapses into a single assistant message — but observability tooling treats each pass as a structured tree of spans. The OTel GenAI Semantic Conventions define a standard gen_ai.* attribute vocabulary that lets one backend consume traces from any compliant harness, and a typical cycle resolves into a small, fixed set of span operations nested under a parent turn span:
invoke_agent — parents the iteration. Carries the agent name, model, and configuration. For a local harness like LangGraph or the Claude Agent SDK it is INTERNAL; for remote agents (OpenAI Assistants API, AWS Bedrock Agents) it becomes CLIENT and measures the round trip.chat (the generation span) — the actual model call. Attributes include gen_ai.provider.name, gen_ai.request.model, gen_ai.operation.name, and token counts.execute_tool — one per tool dispatch, with tool name, arguments, and result status.invoke_workflow when multiple agents coordinate within a single cycle.Two histogram metrics are effectively mandatory for any production deployment: gen_ai.client.operation.duration (latency in seconds) and gen_ai.client.token.usage (token consumption). Together with the span tree, they turn the loop into a queryable object rather than an opaque transcript.
The frameworks mentioned throughout this article ship OTel instrumentation by default. LangGraph, the OpenAI Agents SDK, the Claude Agent SDK, Pydantic AI, smolagents, Strands Agents, and LiveKit Agents all emit traces through the same surface, so a back end such as Langfuse can ingest them over the OTLP /api/public/otel endpoint without per-framework adapters. The Claude Agent SDK, for instance, exposes traces via OpenTelemetry so that every prompt, model response, and tool call lands in the trace store; Microsoft Agent Framework behaves the same way through configure_otel_providers().
The OTel conventions explicitly require that message content attributes be opt-in — instrumentations SHOULD NOT capture prompt and completion text by default, because those payloads carry whatever the user typed. Operators who need full content capture must enable it deliberately and route it to a compliant destination.
This is what makes the loop the right unit of analysis. Questions like why did this task take 40 steps or which tool call introduced the regression are not answerable from the final assistant message; they are answered by reading the trace tree, where every iteration's prompt assembly, model call, tool dispatch, state mutation, and stop decision is a separately addressable span. Without per-cycle spans, production debugging of an agent reduces to guessing, because the harness has already discarded the intermediate evidence.