Mubit
All field notes

Why Request-Path Authority Matters More Than Model Choice

Coding agents are not just model wrappers. Their internal structure decides what they can build, what they cost, and what they can expose.

Model choice isn't the security question

Coding agents are runtime systems around models. Their internal structure determines what they can build, what they cost, and what they can expose. Last week, thereallo reverse-engineered the Claude Code binary and found a prompt mutation path that most users would not have seen during normal operation. The finding was not in the model or the API. It was in the local client, before the request left the developer's machine. The client generated a normal-looking system prompt line:

Today's date is 2026-06-30.

Under certain condit2ions, the client could have changed the byte representation of that string. The apostrophe in Today's could be swapped between visually similar Unicode characters, and the date separator could change from - to /. The rendered prompt still looked ordinary in common interfaces, while the serialized request body contained environment-derived state. According to the writeup, those mutations encoded information related to proxy hostname classification and timezone behavior.

The security issue is the authority this revealed. The client was able to mutate prompt bytes before egress without exposing that mutation through the normal review surface. Prompt construction is therefore not a minor implementation detail. It is part of the trust boundary.

The model is only one part of the system

A coding agent is a runtime around a model. Most modern agents contain some version of this pipeline:

user task
  ↓
context collector
  ↓
prompt builder
  ↓
model call
  ↓
tool broker
  ↓
shell / editor / filesystem executor
  ↓
patch applicator
  ↓
test and lint feedback
  ↓
memory or routing feedback

This pipeline matters because the model only operates on the problem the scaffold gives it. A model that sees only the current file has limited repository awareness. A model connected to repository search, symbol lookup, patch tools, shell access, test output, and a repair loop can inspect the codebase, make changes, observe failures, and iterate. The capability comes from the model and the surrounding runtime together, not from the mode alone.

The design choices that change agent behavior

The visible UI hides most of the real design surface. Coding agents differ in how they collect context, structure the loop, apply edits, execute commands, verify work, retain memory, and route tasks across models.

LayerDesign choiceEffect on resultEffect on trust
ContextOpen files, repo map, embeddings, AST search, grep, MCPDetermines whether the model sees the right codeDetermines what code or metadata leaves the machine
LoopSingle-shot, ReAct, planner/editor, subagentsDetermines exploration and self-correctionDetermines how many actions happen before review
EditsFull-file rewrite, diff, search/replace, structured tool callDetermines patch reliabilityDetermines whether changes are auditable
ExecutionNo shell, local shell, sandbox, cloud workspaceDetermines whether the agent can verify workDetermines blast radius
VerificationLint, tests, build, fail-to-pass checksDetermines whether generation can become repairDetermines whether bad code stops early
MemoryConversation, project rules, persistent tracesDetermines continuityDetermines retention risk
RoutingFixed model, manual choice, learned routerDetermines cost and model-task fitDetermines who sees or mutates prompts

These layers produce different failure modes. Weak context selection can make the model patch the wrong abstraction. Weak edit application can produce patches that fail to apply cleanly. Limited execution support prevents the agent from learning from tests. Overbroad execution support gives the agent more authority than the task requires. Proxy-style routing can improve model selection, while also creating another component with request-path authority.

Context selection decides the problem

Repository tasks often fail before code is written because the agent did not see the files that matter. A bug fix might depend on a type definition, a test fixture, a hidden convention, a config file, and a caller several directories away. In practice, the agent has to build a smaller working set because full-repository context is usually too large and noisy for a model call, while a single open file rarely contains enough information to make a safe change. Agents build that working set in different ways. Common approaches include repo maps that compress files, symbols, classes, functions, and relationships; embeddings or semantic search for natural-language retrieval; AST-aware search over methods and classes; grep and file viewers for precise inspection; iterative exploration through tool calls; and MCP servers for external context such as GitHub issues, Linear tickets, database schemas, logs, internal docs, and deployment systems.

This choice defines the problem the model is asked to solve. A large context window increases capacity, not relevance. The agent still has to decide what is source code, what is generated, what is vendored, what is stale, what is test-only, and what should be ignored. The practical result is simple: context quality determines whether the model is reasoning over the real task or over a convincing fragment of it.

Edit format affects patch quality

After the model decides what to change, the scaffold still has to turn that decision into file bytes. This layer often gets less attention than model choice, even though patch format has a direct effect on reliability, cost, and reviewability.

Edit formatCommon failure modeWhy it matters
Full-file rewriteNoisy diffs, accidental deletion, high token costSimple to generate, expensive to review
Search-and-replace blockFails when the target text is not unique or has driftedCompact, fragile without validation
Unified diffMalformed hunks or patches that do not applyReviewable, dependent on model formatting discipline
Structured editor toolPoor results when the schema or validation is weakSafer when the operation boundary is explicit
AST-aware editLimited by parser coverage and language supportUseful for refactors that need structural guarantees

A strong model paired with a poor edit format can still produce unusable patches. A smaller model paired with strict patch validation and useful feedback can recover from intermediate mistakes. The edit layer is also a trust boundary because the user should be able to inspect the proposed change before it touches the workspace, reproduce the bytes before and after, and verify that the agent did not transform files outside the intended patch.

Shell access changes the risk model

Shell access turns a coding agent into a closed-loop repair system:

edit
  ↓
run tests
  ↓
observe failure
  ↓
inspect stack trace
  ↓
edit again

This loop is one reason agents can outperform a single model call. Tests turn generation into search, and command output gives the model feedback from the real system rather than from its own assumptions.

The same permission set used for verification can also reach project scripts, environment variables, package managers, network calls, git state, and files outside the intended scope. The useful review question is therefore not whether the agent has tools. The useful question is what the tool broker can do without review.

A serious coding-agent architecture needs explicit answers for filesystem scope, network access, shell approval, secret access, package installation, git mutation, workspace isolation, subagent permissions, and MCP server permissions. A restrictive sandbox reduces exposure and may limit verification. A permissive shell improves autonomy and increases risk. A cloud workspace isolates the local machine and moves execution into vendor infrastructure. Every design makes a tradeoff.

Subagents and external tools increase scope

Subagents are mainly useful because they separate work into different contexts. A main agent can delegate repository search, test triage, security review, or migration planning to separate workers with their own context windows. This keeps the main thread cleaner and allows parallel exploration, although it also creates more permission and coordination problems. Subagents can summarize away critical details, trust stale context, collide on edits, or inherit more tools than they need.

MCP expands the boundary further by letting a coding agent access issue trackers, databases, logs, docs, browser automation, design systems, and internal tools. That improves results because real engineering work often lives outside the repository. It also turns the coding agent into an integration runtime rather than a code editor with a chat interface. The more context an agent can aggregate, the more important prompt construction becomes. If the scaffold can combine repo files, tickets, logs, tool outputs, and environment state into a model request, then the scaffold is part of the trusted computing base. A Unicode marker is a low-bandwidth signal. Hidden context selection can carry much more information.

Request-path authority

A component has request-path authority when it can see the full model-bound request, modify prompt bytes before egress, inject system instructions, add or remove context, change tool schemas, route traffic, record responses, or branch on local environment. Many coding agents need some of this authority because request construction, tool mediation, and provider dispatch are core agent functions.

The authority needs to be visible. A local client that builds prompts, a proxy that forwards model calls, an MCP host that aggregates resources, and a cloud workspace that runs tests are all part of the trusted computing base. These components compile developer intent into model requests and machine actions. A prompt preview is useful only when it matches the transmitted request body. A rendered preview is not an audit log. The audit object is the exact serialized request body, plus the tools, files, and outputs that shaped it.

Requirements for an inspectable agent scaffold

A capable coding agent should still have an inspectable scaffold. The basic requirements are explicit inputs, deterministic prompt serialization, documented context selection, canonical Unicode handling, visible tool schemas, auditable permissions, sandboxed execution, reviewable patches, version-pinned behavior, and byte-level egress observability.

For the same task, repository state, config, and tool outputs, the system should be able to answer these questions:

What did the agent see?
What did it send?
What did it run?
What did it change?
What did it remember?
Why did it choose that model?

If the answer depends on invisible branches over hostname, timezone, locale, proxy configuration, account bucket, or experiment flag, the scaffold is not inspectable enough.

Where Minima fits

Model routing is becoming part of the coding-agent scaffold because coding tasks have different cost and quality profiles. A documentation edit, a single-file refactor, a failing-test repair, and a multi-file migration should not automatically go to the same model. The routing problem is easy to state:

Design choiceWhat it prevents
Recommends, never proxiesMinima cannot rewrite the live provider-bound prompt
User stack runs the modelProvider keys and transport stay under user control
Feedback is outcome metadataLearning does not require storing full prompts or responses
Routing learns from outcomesModel choice improves from real usage instead of static assumptions
Harness dispatch stays localThe scaffold can route per task without giving Minima custody of the request

The answer depends on the scaffold. Short documentation edits, mechanical changes, multi-file diffs, test-driven repairs, and architecture-heavy tasks require different model capabilities. Generic model leaderboards are not enough because production outcomes matter. That is what Minima is built for. Minima is a recommendation engine for LLM routing. It recommends a model before the call, then learns from the result afterward. It does not need to sit inside the live model request path. A typical flow is:

POST /v1/recommend
  ↓
your stack runs the selected model
  ↓
POST /v1/feedback

Minima learns from task-model-outcome history: what ran, whether it worked, how much it cost, how long it took, and how many tokens it used. The architectural point is that routing intelligence does not require prompt custody.

A proxy router can observe and rewrite the full request, which can be useful for centralized policy, logging, redaction, caching, and provider normalization. Model selection alone does not require full request-path authority. Minima's design keeps the live model call in your stack, so the model request, provider keys, transport, logs, and policy controls remain under your control. Minima advises before the call and learns after the call. It does not need to mutate the prompt because it does not need to own the prompt.

What Minima's design prevents

Minima can improve routing without becoming the model-call proxy. That design removes specific failure classes from the Minima boundary.
Coding agents are runtime systems that turn developer intent into repository context, prompt bytes, tool calls, shell commands, patches, tests, and feedback loops. That structure determines what the agent can build, expose, mutate, or hide. The practical questions are architectural:

Components that must sit in the request path need byte-level observability, deterministic serialization, explicit inputs, sandboxing, auditable permissions, version pinning, and reviewable edits. Components that do not need request-path authority should stay out of the path.

That is the Minima position. A coding agent may need a powerful scaffold. A model router does not need custody of the prompt. The agent scaffold should be capable, and the request path should be inspectable. Security patches repair individual implementations. Architecture removes unnecessary capabilities. We would rather build the routing layer that cannot mark your requests than the one that promises it will not.

Conclusion

Coding agents are runtime systems that turn developer intent into repository context, prompt bytes, tool calls, shell commands, patches, tests, and feedback loops. That structure determines what the agent can build, expose, mutate, or hide. The practical questions are architectural:

What does it see?
What does it send?
What does it run?
What does it change?
What does it remember?
What can it not do by design?

Components that must sit in the request path need byte-level observability, deterministic serialization, explicit inputs, sandboxing, auditable permissions, version pinning, and reviewable edits. Components that do not need request-path authority should stay out of the path.

That is the Minima position. A coding agent may need a powerful scaffold. A model router does not need custody of the prompt. The agent scaffold should be capable, and the request path should be inspectable. Security patches repair individual implementations. Architecture removes unnecessary capabilities. We would rather build the routing layer that cannot mark your requests than the one that promises it will not.

Give the next run
something to build on.