Skip to main content

Architecture

AI-Native Repository Architecture

How this portfolio's repository lets AI coding agents make useful changes safely: deterministic context routing before an edit, narrow agent roles during it, and the same validators, tests and human review after it.

Published October 5, 2026

This portfolio is changed every day by AI coding agents working alongside me. This case study is about the system around those agents: how the repository tells an agent what it needs to know before it edits anything, which work an agent may do on its own, and what checks every change has to pass before it reaches production. The premise is the same one behind the Photography Assistant: the model is a capable, probabilistic participant inside a deterministic system. AI increases what can be implemented in a day. The architecture of the repository decides how safely that capability can be used.

The problem

A coding agent can write code that compiles, passes its tests and still does the wrong thing for this repository. Generation quality is not the main risk. The risks are the things a model cannot see from the file it is editing:

  • Architecture. Which domain may import which, where database access is allowed, which layer owns a decision.
  • Conventions that look optional. Every user-facing string exists in English and German; every database row is validated with a schema before use; admin writes go through an authenticated procedure.
  • Constraints that fail late. Some mistakes pass the type checker, the linter, the unit tests and the production build, and only break in the deployed container.
  • Context that is plausible but wrong. A confident summary of a subsystem, from a model or from an outdated document, is easy to act on and hard to notice.

So the engineering problem is control in two directions: what context reaches the agent before it changes something, and what is checked deterministically after it does.

What I built

I designed and built the repository's agent-facing architecture: the instruction hierarchy, a deterministic routing layer for context, a validated knowledge layer, the architecture and repository-hygiene validators, the local hooks and CI gates that run them, and a small set of specialised agent roles with enforced limits. I also ran the experiments that decided which roles exist and which model each one uses.

It serves one developer (me) and several AI agents on a production codebase of about 4,500 tracked files. Some decisions below are shaped by that scale, and I say so where they are.

Architecture at a glance

  1. Request from me
  2. Primary agent classifies the taskrouting manifest → knowledge concepts · skills · CodeGraph entry points · validation commands
    • Bounded discoveryread-only explorer
    • Bounded implementationdecided package → bounded implementer (contract, write boundary, ownership check)
  3. Primary agent verifies + integrates
  4. Pre-commit hooklint · hygiene · architecture · knowledge · generated index
  5. Pre-push hookstaleness · knowledge health · hygiene · architecture · static analysis · types · unit tests
  6. CI (pull request, then main)static analysis · four validators · tests · build → deploy
  7. I review, commit and ship
  • Instruction hierarchy. One file (CLAUDE.md) holds the behavioural rules for every agent; the instruction files for other tools point to it instead of restating it. It says how to behave and links to where knowledge lives; it deliberately does not try to describe the system.
  • Routing manifest. A lookup table of twelve task categories, each listing the knowledge concepts, workflow skills, code-graph starting points and validation commands that category needs.
  • Knowledge layer. About forty short concept files on architecture, decisions, security, testing and frontend rules, each with machine-checked metadata and source citations.
  • CodeGraph. A local index of every symbol and call edge, which agents query instead of reading files one at a time.
  • Validators. Deterministic checks for architecture boundaries, repository hygiene, knowledge integrity and routing integrity, run by hooks and again in CI.
  • Agent roles. A primary agent that owns the task, plus narrow subagents for discovery and for already-decided implementation, each limited by tool configuration and guards.

The key architectural decision

Authority comes from configuration and checks, never from the model.

The agent that owns a task decides what the task means, which code owns the change and what counts as done. Work it hands off arrives at a subagent already bounded: a discovery question with a defined scope, or an implementation package with a written contract, an explicit list of paths it may write and acceptance commands that run. What a subagent may do is set by its tool list and by a guard script that runs before each of its tool calls, not by its instructions or by how capable the model behind it is. Everything a subagent returns is treated as evidence that the primary agent re-reads before acting on.

This is why the cheapest model in the system can do real work safely: its authority is the smallest. In the measured history of my own requests, none of 177 arrived already bounded enough to hand off directly; turning an ambiguous request into bounded work is the part that stays with the strongest reasoning in the loop.

Context before generation

Before a non-trivial edit, an agent classifies the task against the routing manifest by matching the category's trigger phrases, then opens only what that category lists: one to three concept files, any procedural skill, and the code-graph entry points. It states what it opened before editing. If nothing matches, it says so and falls back to a navigation page. A silent skip is the one outcome the instructions forbid.

This is deliberately not semantic retrieval. The manifest, the knowledge validator and the routing validator use no embeddings, no ranking and no model calls; they are static tables and linters. I chose that because the failure I wanted to prevent was an agent loading plausible context. A lookup table can be wrong, but it is wrong the same way every time, it can be reviewed in a diff, and a validator can check that every path it names exists and is the right kind of thing.

The code graph is different: it is a third-party tool with its own full-text ranking for search and exact traversal for callers, callees and impact. Agents use it to read the actual source of the symbols around a change in one call, instead of chaining searches and file reads.

The knowledge layer is validated rather than trusted. Each concept declares the source files and regions it describes. One check verifies metadata, links and navigation coverage on every commit; another compares each concept against git history before every push and reports which concepts may have drifted from their sources.

Guardrails

Each guard runs at a defined boundary, and the boundaries overlap so that skipping one doesn't remove the check.

On every commit, a hook runs the linter when code is staged, the staged-file hygiene check, the staged-file architecture check, knowledge validation and a check that the generated knowledge index is up to date.

On every push, a second hook runs the knowledge staleness and health reports, the full hygiene and architecture validators, static analysis over the whole repository, the TypeScript compiler and the unit tests. Both hooks stop at the first failure.

On every pull request and again on the path to production, CI runs static analysis, the same four repository validators as separate named steps, the tests and the production build. Deployment depends on that job, so a validator error stops a deploy.

The architecture validator is the core of it. Domains and their allowed imports are declared in one manifest, and it reports, as errors: forbidden cross-domain imports, cycles between domains, database access from any file that is not a declared owner, production code importing tests or diagnostic scripts, hard-coded storage paths, and a bundler directive misused on local imports. Declared owners and exceptions carry a written reason; an owner list is an inventory of what exists, not a pre-approval of what might. Warnings, such as documentation that no longer matches the manifest, print without blocking.

The hygiene validator protects the evaluation work the product depends on: canonical benchmarks and human judgment files cannot be deleted in a routine change, archived material cannot become an active dependency, and an evaluation script cannot reimplement the production retrieval code it is meant to measure.

The agent and tool boundary

Four specialised roles exist, each with a pinned model and a restricted tool set:

  • Repository explorer (Sonnet). Read-only tools plus a shell guarded against writes. It gathers evidence and labels every claim as observed, inferred or unknown.
  • Bounded implementer (Haiku). It can write source files. A guard refuses repository authority (commits, pushes, installs, deploys) and any edit to the files that define the rules judging its work. An ownership check afterwards compares what actually changed with the paths the contract allowed.
  • Execution manager (Haiku). It advances a plan the primary agent already wrote, step by step, through a deterministic execution ledger, and can only dispatch the two roles above. The ledger, not the manager, writes the final handoff.
  • Visual QA (Haiku). Experimental. It reviews screenshots and state captured by a deterministic browser script, and can't drive a browser itself.

Subagents run on an explicitly chosen model; the general default is Sonnet, not the primary agent's model. Architecture, product semantics, security and trust boundaries, acceptance criteria and anything ambiguous are never delegated. UI work is not delegated by default either, because there is no deterministic check for rendered state that could accept it.

The guards are pattern guards, not sandboxes. They make the intended use the easy path and block known mistakes; they don't make misuse impossible. That's why the real gates are the validators, the tests and my review, which run no matter who made the change.

Separately from the coding agents, the product exposes its cover-letter and CV pipelines as tools over the Model Context Protocol. Those tools are thin transport adapters over the same deterministic modules the web application uses, so an AI client can call the pipeline but can't change how evidence is selected.

What failed

Writing a rule down did not enforce it. Problem: the rule that only the query layer touches the database was stated four times in the agent instructions, and nothing checked it. Result: a repeated instruction is still only an instruction. Consequence: it became an architecture error with five declared owners, each with a reason. The instructions now also say which part of the rule is machine-checked and which remains a review convention, instead of stating every clause with equal weight.

Green builds that would break production. Problem: a bundler directive on local imports passed the type checker, the linter, the tests and the production build, and would only fail inside the standalone production image. A second defect hard-coded a storage path that should depend on the environment. Experiment: I turned both into validator rules and replayed them against the historical commits. Result: five findings and one finding at the defective commits, zero after the real fixes. The existing import graph could not have caught the first one: it ignores dynamically built import paths by design, and the misuse depends on exactly that. Consequence: a dedicated check. Two things an agent previously had to notice became things it no longer needs to reason about.

Local gates were the only gates. Problem: for a while the four validators ran only in local git hooks, and hooks can be skipped. Consequence: each validator became its own named step in pull-request CI and in the deploy pipeline. Some checks still have no CI equivalent, such as knowledge staleness, and the instructions document that exception instead of claiming skipping the hooks is harmless.

A staleness check that cried wolf. Problem: the knowledge staleness check flagged concepts whenever a file they cited changed. Experiment: every flag was audited by hand before the check could be added to CI. Result: 15 of 17 flags were false or conservative. The one hard failure was a false positive, while the one genuinely stale concept was only soft-flagged. Consequence: concepts now declare the regions of a file they depend on. Against that target the repaired detector had no false positives or false negatives. "Stale" now blocks only when a cited source is missing or unresolvable; anything else is advisory.

A routed path that produced nothing. Problem: the routing manifest pointed agents at the database schema as a code-graph entry point. The file existed, so validation passed, but the code graph extracts no symbols from SQL. On inspection, five of 21 entry points produced nothing. Consequence: a separate list for files to read directly, and a validator error when a code-graph entry isn't a symbol source. Whether a file is in the index is not the test; whether it produces symbols is.

The cheapest model as default explorer. Experiment: five audited discovery runs on Haiku. Result: material defects, including false claims and missed configuration facts, in three of five. Consequence: discovery runs on Sonnet. The larger cost turned out to be elsewhere: re-pricing 101 historical subagent runs showed most of the spend came from subagents silently inheriting the most expensive model, which is why the default is now pinned.

AI reviewers did not earn the role. Experiment: independent reviewers were given the historical defects that deterministic checks had missed, plus clean control changes. Result: a Sonnet reviewer caught two of seven against a threshold of three set before the first run, and its one unique catch did not reproduce. A cheaper hosted model caught none unaided and never returned an empty review on a clean change. Consequence: no AI review stage. Defect classes that recur become deterministic rules.

The contract was the limit, not the model. Experiment: eight frozen implementation packages, each run on Haiku and on Sonnet. Result: Haiku passed seven of eight with no corrections, Sonnet six. Both models produced the same wrong output on one package, because the contract described the rule imprecisely. On two others, the acceptance command could never pass. Consequence: Haiku owns bounded implementation, and contracts plus their acceptance commands are reviewed as carefully as code.

Comments that drifted from the code. Problem: in three months, comment density in the source tree grew more than tenfold. Result: an audit found most comments valuable, but every confirmed stale claim restated a fact owned elsewhere: a count, an "only caller", a data claim. One had removed a safety justification. Consequence: a rule that comments state the why and the invariant beside the code they protect, never facts owned by another file, plus a hygiene warning when a comment names a deleted file.

Trade-offs

  • More guardrails, more maintenance. Every rule needs fixtures, a reason for each exception and upkeep when the architecture changes. The bar I set is that an error rule must encode a real invariant that is currently at zero violations. Anything weaker stays a warning or a review convention.
  • More context is not better context. Routing deliberately loads one to three concept files, not the whole knowledge layer. That keeps an agent's window for reasoning, at the risk that a relevant concept is never routed.
  • Slower local pipelines. The pre-push hook runs the type checker and thousands of unit tests. That's minutes I accept so that a broken change rarely leaves my machine.
  • Delegation costs time. Handing work to a subagent saves the primary agent's context but is measurably slower when done step by step. It pays off for large discovery or parallel investigations, not for a single known file.
  • Central knowledge can go stale. Explicit knowledge only stays useful while it is true. The staleness check, source citations and the rule against restating facts owned elsewhere are the cost of keeping it honest.
  • Guards are not sandboxes. Pattern guards are simple and auditable, but they are not a security boundary. The real gates are the ones every change passes regardless of author.

What remains human

Agents don't decide whether something should exist. I decide what to build and why; whether an abstraction is justified or a duplicate is acceptable; how an interface should feel; what an ambiguous requirement means; and whether an experiment ships. Experiments here end with a verdict written against criteria fixed before the first run, and several of the most useful verdicts in this repository are rejections. I review every change and make every commit; an AI drafts the commit message from the staged diff, and I confirm it.

Result

What the architecture provides, and what I can demonstrate:

  • Architectural violations that matter are machine-detectable, with the same rule run locally and in CI.
  • An agent starts a task by retrieving the repository's own context for that kind of task, not generic advice or its own guess.
  • Validation runs before a change leaves my machine, and again before it can be deployed.
  • Architectural knowledge is written down, cited and checked against source, rather than living only in my head.
  • AI-written changes pass exactly the same gates as mine.

I don't claim a productivity multiple. What I can say is narrower: mistakes that were made here, by me or by an agent, became checks, and each check now applies to every change that follows.

Further reading

The engineering journal tells parts of this story in narrative form: From Portfolio Website to AI-Operable Engineering Platform, Agent Automation Without Losing Engineering Control, AI Doesn't Need More Context — It Needs Better Context and Designing Deterministic AI Workflows Around Claude. The Photography Assistant case study applies the same principle to a production LLM feature.