← All posts

Building an agentic IDE: lessons from openrouter_agent

openrouter_agent started as a question: how hard is it to build an IDE where the model can actually work with the codebase - read files, edit them, run the tests - instead of just chatting about code in a side panel? The honest answer: the happy path took a weekend, and everything after it took months of small, instructive failures. This post is about those failures: the tool loop, context assembly, and the four problems that taught me the most about how agents really work.

The whole agent is a while loop

Strip away the marketing and an agent is three things: a model, a set of tools it can call, and a loop that feeds tool results back into the conversation until the model decides the task is done. That's it. No framework magic, no hidden intelligence - just a disciplined cycle of decide, act, observe.

while (steps++ < MAX_STEPS) {
  const res = await model.chat(messages, tools);
  if (!res.tool_calls?.length) break;   // model is done

  for (const call of res.tool_calls) {
    const result = await runTool(call); // read_file, edit_file, run_tests...
    messages.push({ role: "tool", content: cap(result, 2000) });
  }
}

The tool loop: user task, model decision, tool call, result, context growth, verification

Every iteration follows the same shape: the model sees the conversation, decides on the next action, the IDE executes it, and the result lands back in the context. The loop ends when the model stops calling tools - or when the step budget runs out, which matters more than you'd think (more on that below).

The first version of this loop was thirty lines of code, and it worked - in the sense that it could sometimes fix a failing test end to end. It was also dumb as a hammer: no verification, no context management, no guardrails. Everything interesting in this project happened while fixing that.

Context assembly is 90% of the quality

The model is only as good as what it sees. Early on, the agent saw the user's task and whatever it had read so far - and it flailed: editing functions it had never read, guessing at imports, breaking things two files away from the change it was making.

Context assembly: what the model sees before deciding the next action, and the rules that made it work

The fix was not a better prompt. It was treating context as an engineered artifact with explicit rules:

  • Full files, not hunks. A diff snippet hides the surrounding code the change has to fit into. Reading the whole file (trimmed of blank-line runs and long import blocks) made edits dramatically more accurate.
  • Trim aggressively, keep order stable. When the context overflows, cut the oldest tool results first - but never reorder. Models are surprisingly sensitive to a context that reshuffles under them mid-task.
  • Errors verbatim, never summarized. A stack trace summarized by the agent is a stack trace the model can't act on. The raw text goes in, capped in length, with a pointer for reading more.
  • Never drop the task itself. The original request stays pinned at the top no matter how long the session gets. Agents that "forget" what they're building mid-loop will happily wander off polishing something else.

Four failures that taught me the most

1. Tool results flooding the context

The first time the agent ran a test suite, the output was five thousand lines of npm noise. It went straight into the context, pushed out everything useful, and the next model call was garbage-in-garbage-out. The fix: every tool result passes through a cap - head and tail kept, middle truncated, with a note like "output truncated, use read_file(path, offset) for more". The agent gets enough to decide, and a tool to fetch the rest on demand.

2. Models are not interchangeable

openrouter_agent is built on the OpenRouter API, which is both its biggest feature and its biggest headache: a hundred plus models behind one interface, and "one interface" is a lie of convenience. Models differ in how strictly they follow tool schemas - some emit clean tool_calls, some wrap JSON in prose, some invent parameter names. The fix was a normalization layer plus a per-model quirks table, and a small scenario suite (twenty tasks like "fix this failing test") that runs against every model before it joins the default list. A model that passes a chat benchmark but fails the scenario suite does not ship.

3. Edits that don't apply

The edit tool is search-and-replace over file content, and models fail at it in a specific way: they reproduce the target snippet from memory with a whitespace difference, the match fails, and the model retries the same broken edit - sometimes five times in a row. Two fixes helped. First, a fuzzy fallback that tolerates indentation drift. Second, and more important: when an edit fails, the error message tells the model exactly what to do next - "match failed; call read_file on the target range before editing again". Failure messages are prompts. Write them like prompts.

4. Loops that never end

Left unbounded, an agent will happily burn forty steps retrying the same failing edit, or "improve" code nobody asked it to touch. Two guardrails fixed it: a hard step budget with an honest final message ("I could not complete X; here's where I got stuck"), and a repetition detector - when the same tool is called with the same arguments twice in a row, a system hint tells the model to change approach. Neither fix is clever. Both are essential.

The verification loop is what makes it usable

The single biggest quality jump in the whole project came from one change: giving the agent a run_tests tool and a rule - after every edit, run the tests. Before that, the agent produced plausible code and had no idea whether it worked. After that, it could catch its own mistakes, and the success rate on real tasks roughly doubled. An agent that can verify itself stops guessing; an agent that can't is just a very confident typo generator.

What I'd do differently

Three things, in order of regret. First, an eval set of real tasks from day one - for weeks I tuned by vibes, and every "improvement" was a coin flip. Second, streaming in the UI earlier: the agent was capable long before it felt capable, because users stared at a silent screen for thirty seconds. Third, a visible cost meter during development - agents burn tokens at a rate that surprises everyone, and watching the number live changes how aggressively you trim context.

The surprise ending: building the agent taught me more about prompting than any prompt engineering article. The prompt is maybe 10% of the system. The other 90% is what the model sees, what it can do, and what happens when it fails.

Building an agent or an AI-powered tool? Get in touch - or try openrouter_agent on GitHub.

← All posts

↑