Architecture
Keeping a React Frontend Coherent Under AI-Assisted Development
How this portfolio's frontend stays coherent while it changes quickly with AI coding agents: an external design system as the foundation, a small set of shared decisions, local ownership for everything else, tests that pin invariants, browser measurement, and human judgement about the product.
Published October 7, 2026
This portfolio did not start as a 60,000-line React application. Over time it took on a photography application with a gallery, a focus viewer, a map and an interactive constellation of photographs; two conversational interfaces; admin tools; long-form journal and architecture pages; local speech input; and several rendering surfaces that exist only in the browser. Building another component stopped being the hard part. Keeping all of it feeling like one product while it kept changing became the hard part. This case study is about the frontend architecture I use to do that, and about where it stops working.
The coherence problem
The frontend is a Next.js 16 App Router application on React 19, with the React Compiler enabled since the first commit. It has 361 TSX files, about 61,000 lines of them, and 276 client modules. Photography alone accounts for about 38,600 lines of component code. Every user-facing string exists in English and German, and nearly every surface has a different shape on a phone.
Those numbers are not the point. They describe the conditions under which a frontend quietly fragments: a second button that is almost the first, a font size restated in a new file, a modal that traps focus slightly differently, a hover affordance that never reaches a touch screen. Each is reasonable on its own. Together they make a product that looks assembled rather than designed.
Speed makes this worse. When a change takes an hour instead of a day, more changes happen, and each is a chance to re-decide something already decided. Much of this frontend was built with AI coding agents in the workflow, so that pressure is real here. The question is not how to organise React components. It is how to keep a frontend coherent when implementation can move faster than its decisions can be re-derived.
My answer has four parts:
- Architecture: where code lives and which way dependencies point.
- Conventions: the product's decisions, written down once.
- Deterministic evidence: types, lint, tests that pin invariants, and measurements in a real browser.
- Human judgement: taste, product direction, and whether something is worth its complexity.
The layered frontend
Dependencies point one way: from routes, through product features, to a small shared layer, to an external component library at the bottom.
- Routes and layoutsserver components · locale · cached content
- Server-rendered pagesjournal · architecture · trust pages
- Interactive feature rootshome · photography · conversations · admin
- Shared primitivescontrols · disclosure · motion · conversation shell
- Style policysurface constants · type roles · breakpoints · conversation chrome
- Tokens and motion vocabularyroot custom properties · shared keyframes
- KERNexternal components and tokens
- Server by default. The locale layout loads messages, site content and the page background on the server. Content reads use Next's
"use cache"with explicit lifetimes and cache tags that admin edits invalidate. - Client at the feature root. The client boundary sits where interaction begins: the home page, the photography application, the two conversation surfaces, the admin. Below it a feature owns its state. There is no global store.
- A shim for the component library. KERN ships as one bundle that creates React context when it loads, which crashes a server component that imports it. Server-rendered pages reach it through one small client re-export module, and that is what lets the journal and architecture pages stay on the server.
- Two data paths. Server components read through the cached layer. Interactive features use tRPC with TanStack Query; the two AI conversation endpoints use a streaming link so answers can arrive in phases.
- Heavy runtimes behind lazy islands. A Three.js avatar, a MapLibre map with its own worker, a Canvas2D constellation and the homepage focus viewer load on demand and client-only. WebGPU launcher icons fall back to SVG. Local speech recognition runs in a worker.
- Components own their waiting. There is no route-level loading or error file. Loading and failure states belong to the component that knows what it is waiting for.
A photography component may use a shared primitive; a shared primitive never knows about photography. For the photography interface those import directions are declared and checked mechanically. Elsewhere they are a convention.
Not another design system
I did not build a second design system for this portfolio. The formal foundation is KERN, an open-source UX standard for German public administration, used here through its React kit: 211 component files import it for typography, buttons, alerts, inputs, dialogs and cards. Above it is something deliberately smaller, a convention layer for the decisions this product makes that KERN doesn't: a few product primitives, typography roles, a motion vocabulary, one owner for breakpoints, the geometry of floating surfaces, and the chrome the two conversation interfaces share.
There is no catalogue, no token pipeline and no versioned package, and most of the layer's rules are written down rather than enforced. That distinction matters for everything that follows.
Some primitives are used widely. Counting files that import each from outside its own folder:
ChipButton, 28: action chips and link chips, typed as a discriminated union so a chip is a button or a link, never an ambiguous mixture.IconActionButton, 16: its accessible name is a required prop.EditorialAnnotation, 13: byline and metadata rows.EyebrowLabel, 13: the small uppercase kicker above a heading.MotionReveal, 11: a mount entrance with no state of its own.DisclosureGroup, 10: native<details>, used instead of KERN's accordion outside admin forms.
Others have few consumers, and that is not automatically a failure. The conversation shell has two importers and the shared question field five, because there are exactly two conversation products and they share both. Section, a headed glass card, is down to two after the homepage was recomposed. The shared dropdown has one consumer, the language switcher, which appears in both desktop and mobile navigation. A low count says how many places need a decision, not whether the decision deserved a name.
The rule that governs the layer is short: promote a shared decision, not a repeated number. Recurring 999px pill radii, 8px gaps, blur strengths and local z-index values were each examined and left alone. They are the same numbers expressing different decisions, and a token for each would couple surfaces with no reason to change together. A z-index registry was judged unnecessary for the same reason.
These primitives were mostly extracted from existing features rather than designed in advance, and several were renamed or reshaped later.
Local ownership before abstraction
A decision moves into shared code only when two surfaces make the same decision, and then only as far up as both need. I think of it as the nearest common owner.
Shared: the same decision, made twice
transcript chrome visitor turn, bubble, 12/32px rhythm
conversation shell floating panel / mobile sheet
question field one composer for both products
GPU identity layer one launcher artwork, two products
breakpoints one owner for device tiers
typography roles eyebrow, group title, pill label
Local: similar on the surface, decided separately
Curator type scale 14px body · 12px metadata
Engineering scale 15px answer · 13px metadata
role label also the focus-view eyebrow
domain shells one per conversation product
motion primitives different lifecycles
999px · 8px · zIndex same numbers, different reasons
The two conversation surfaces show it best. The photography Curator answers in caption-sized prose: 14px body, 12px metadata. Engineering Ask answers questions about how I work, and there the answer is the content, so it sits one step larger at 15px and 13px. One merged scale would have been wrong for one of them.
Where they genuinely agreed, the decisions moved: five value-identical transcript decisions now have one owner beside the conversation shell. The role label was evaluated for the same move and rejected, because in the Curator it is also the focus view's eyebrow.
The same reasoning stopped two larger abstractions. A universal AI chat panel was never built; the conversation shell already was the shared layer, and each product keeps its own shell for its own state. The motion primitives were not merged either. An audit found the problem was not how many there were but that their lifecycles were implemented inconsistently, so each lifecycle was fixed instead.
The largest feature component, the photography assistant, peaked at 3,055 lines and is 2,303 today. Part of that drop was a later pass removing stale comments, so the line count is context, not the result. The result was a render boundary: typing in the Curator used to re-render the gallery behind it, 384 gallery-card renders over four keystrokes. With the question draft and the conversation core moved into their own component and hook, it is zero. The seams followed what had to re-render, not what looked large. Extracting the map and constellation hosts was considered and declined, because each would have needed nine or ten props tunnelled through it.
Styling as architecture
The styling model is unusual, and I don't think it is universally right. It is a set of ownership decisions that fit this codebase:
- Inline styles by default: about 1,600 inline style objects across 248 files, with spacing from one
spacing()helper on an 8px base. - A CSS module only where inline styles cannot express the CSS: pseudo-elements,
:has(),@starting-style, keyframes, hover and focus-visible states, and media queries a component's own layout needs. There are 18, each owned by one component and opening with a comment that says why it exists. - Global CSS for foundations: root tokens, including five for motion duration and easing; resets; the native dialog entrance; and a shared motion vocabulary.
- KERN underneath, with its own styles and tokens.
The global stylesheet once reached 903 lines. One migration moved rules that belonged to individual components into the first eleven modules and took it to 286. It is 542 now, with 18 modules. That is not a story of CSS getting smaller; it is a story of rules getting owners. What stays global is foundational, and much of its length is rationale written beside each rule.
The most useful moment in that history was an audit that changed no styles at all. Its conclusion was that the styling architecture was healthier than its documentation. The code was coherent; the written guidance contradicted it:
- One workflow forbade CSS modules that the main instructions allowed.
- The instructions pointed at a spacing helper in a file that did not exist.
- The motion guidance called the global reduced-motion rule the only mechanism, while twenty-two animations already carried their own reduced-motion branches. Followed literally, it would have let an infinite animation ship with no reduced-motion handling at all.
So the first change was to the rules, with no production change. Reduced motion became a procedure with one input: is the animation timed by a motion token? If so, the global collapse covers it. If not, it needs its own media query. If JavaScript drives it, it reads one shared hook.
Only then did two narrow consolidations follow: the transcript chrome above, and one owner for device breakpoints that three TypeScript files had each re-declared. For the first, 940 computed style fields measured in a browser were identical before and after. For the second, 50 of 50 fresh-load probes resolved the expected device tier, and no value or behaviour changed. Turning down the other consolidation candidates was part of the work.
Performance boundaries
I keep two kinds of performance statement apart: what was measured, and what the architecture intends.
Measured:
- WebGPU launcher. The dimensional launcher icon adds 1.2 KB gzip to the initial path; the other 44 KB loads only when the GPU artwork runs. In idle CPU profiles it used less main-thread time than the CSS-animated SVG it can replace.
- A loop that never stopped. A
requestAnimationFrameloop in the shared launcher identity ran about 577–613 callbacks every ten seconds in every state, including while its launcher was not mounted. Gating it on state the effect already depended on brought that to zero, with no visible change. - Journal layout shift. Navigating from the index to an article on a throttled connection shifted the layout by up to 0.1991 CLS at phone width. The fix was not a better skeleton: the placeholder was removed, so a chapter arrives in its first server response. CLS is now 0 at both measured widths.
- Header fold. A first design that folded the mobile header on scroll measured CLS 0.046, because the island changed width while centred. The later design moves nothing by layout and measures 0.
- Typing in the Curator. 384 gallery-card renders over four keystrokes became zero, as described above.
Structural simplification, fewer moving parts rather than measured speed: homepage reveal observers went from 17 to 9, and four state variables and one timer that existed only to animate were removed.
Architectural intent, not measured: server rendering and cached reads; heavy runtimes behind lazy, client-only boundaries; photographs generated at four widths in WebP and AVIF, with per-photo sizes in the focus viewer; and the React Compiler, which I have never measured against its absence. I believe these choices are right. I don't have numbers showing they made the site faster, so I don't claim that.
Accessibility changes the architecture
The accessibility decisions that mattered most were about which element or primitive a feature is built on.
- Native elements over custom widgets. The mobile conversation sheet is a native
<dialog>opened withshowModal(), so focus trapping, Escape and an inert background come from the browser. Disclosure uses<details>. A custom navigation drawer was replaced by KERN's dialog. - Accessible names in the type contract. The icon-only action button and the spatial loading indicator cannot be rendered without a label; it is a required prop.
- Hidden means inert. Visually hidden chrome is also made
inert. Opacity is never the accessibility state. - Content visible without JavaScript. With JavaScript disabled, the homepage used to be blank below the hero: all seven scenes at opacity zero, permanently. Entrances now start from visible server HTML and arm themselves only for content that is genuinely off-screen.
- A composer that keeps its identity. On mobile the conversation's text field was destroyed and recreated whenever layout density changed, losing focus mid-typing. It is now one persistent element.
Two defects found by measuring the rendered page show why this has to be checked in a browser. The gallery's 72 cards each nested interactive elements, a serious axe violation; after the fix there were none. And on a common desktop size, the floating conversation launcher covered the footer's AI transparency link: none of five hit-test points on the link was reachable by pointer, and after the fix all five were. The fix derives the launcher's clearance from the module that already owns the insets and stacking of floating surfaces, instead of adding another local offset.
This is not a claim of WCAG conformance; no conformance audit has been done. The automated checks described here, including axe, run locally in the development workflow, not as a CI gate.
Responsive design changes ownership
Sometimes a phone needs smaller dimensions. Sometimes it needs a different owner for the interaction.
- Navigation. Desktop has a pill navigation bar. Below 1280px the header becomes a floating island whose menu opens as a bottom sheet with its own dismiss gesture. Building it surfaced three existing defects: scroll lock stuck after a language switch, Escape in the language list closing the whole menu, and classic scrollbars shifting the page sideways.
- Conversation. The conversation shell is a floating panel on desktop and a native modal sheet on mobile. On mobile, the fixed chrome took 249px of the sheet while the keyboard cost 49px, so a density model now decides how much chrome each stage of a conversation shows.
- Hover. Hover text on gallery cards was removed. It showed a title that in production was always just the photo's ID, plus details repeated in the focus view. The card actions stayed.
- Safe areas. Controls anchored to a screen edge add that edge's safe-area inset in CSS, with no JavaScript detection. Extending the page under the notch with
viewport-fit=coverwas tried and reverted. - Images. Photo
sizesare built from the same TypeScript breakpoint owner that CSS mirrors, and a test checks that the mirrors agree.
Responsive logic has three legitimate owners: CSS media queries for layout, one hook for render decisions that need JavaScript, and HTML sizes strings. Not everything goes through the hook, deliberately.
All of this was checked with touch emulation in Chrome. Several mobile behaviours, including the iOS keyboard and safe-area handling, have not yet been verified on a physical device.
Where the agent fits
Everything so far is the frontend. AI coding agents work inside it, and the useful question is how an agent can make a change without rediscovering every decision above. Mostly the answer is context:
- One instruction file with the styling rules (layer order, inline-first, the CSS-module exception, decisions-not-numbers, progressive and reduced motion), the KERN rules, and a list of existing components to reuse before creating new ones.
- Frontend knowledge files on architecture, motion, accessibility, AI interaction surfaces and the KERN server/client boundary, plus a short written frontend constitution.
- Workflow skills for adding a component, adding motion, adding localised text, validating a UI change in a browser, backing a performance claim with a measurement, and running React Doctor on the changed files.
- A code graph for finding existing symbols and their callers before writing new ones.
- Module comments that say why a file exists.
The repository-wide side of this, including context routing, the code graph and the validators, is described in AI-Native Repository Architecture. This article covers only what it means for frontend work.
- Taska request about a surface or behaviour
- Contextinstructions · frontend knowledge · skills
- Discover what existscode graph · existing primitives · module comments
- Bounded changein the owning layer
- Teststypes · lint · unit and source-pinned tests
- Browser evidencerendered-state checks · measurement
- Human reviewtaste · product direction · complexity
- Keep, modify or delete
I directed, reviewed and committed the work. AI coding agents were used inside that workflow for repository analysis, implementation support, measurement and report writing. Git history does not record which lines an agent wrote, and I don't treat any of this as autonomous frontend development.
What the agent cannot decide
An agent that can measure more things does not make every question objective. The frontend reports in this repository record decisions that no test, measurement or agent could settle.
- The WebGPU launcher. A dimensional launcher icon passed every platform gate: it rendered in the real launcher, fell back to SVG on every failure and cost less at idle. It was still a worse icon, because at a device pixel ratio of 1, on ordinary desktop monitors, it muddied into grey-blue dots and lost the identity of the existing glyph. The verdict was to keep the SVG. A later darker material was rendered in three palettes and one was chosen on review, raising the icon's interior contrast from 1.44:1 to 4.44:1. The other two were removed from the code, not left switchable.
- The camera artwork. Enlarging the Curator's camera artwork from 71.5% to 79.2% of its frame was built, measured and reverted on review.
- The header fold. The scroll-folding header was built and measured, then removed at my direction because the contact actions belonged on the island at all times. The next day I reversed that decision. The redesign moved nothing by layout, and CLS went from 0.046 to 0.
- Gallery hover. Removing the hover text was easy to justify with data. Removing the card actions too was on the table, and I rejected it.
- Constellation depth and atmosphere. Two constellation experiments worked technically. Information-bearing depth passed 16 new tests and all 185 constellation tests, and in a real browser its effect was imperceptible at production values. A volumetric atmosphere did not earn its complexity. Both were removed.
- The iOS keyboard. The persistent composer proved the text field no longer remounts. Whether the keyboard appears and stays open on an iPhone needs a physical phone, and that is still open.
Agent Deterministic Human
────────────────────── ───────────────────── ──────────────────────
repository analysis types visual taste
implementation support lint product direction
measurement tests, source pins worth its complexity?
consistency checks import boundaries physical devices
report drafts rendered-state checks what ships
The agent produces evidence. Deterministic software turns some of that evidence into checks on every change. I decide what the product is.
Testing without a large DOM test stack
This is the part of the architecture I would describe most carefully, because it is a trade-off, not a recommendation.
The repository has 866 test files. About 94 test frontend code by reading component source and asserting its structure; others test pure policy modules; one renders server markup and checks the HTML. There is no jsdom, no Testing Library, no browser end-to-end suite in CI and no persistent visual-regression baseline.
What makes this workable is where decisions live. Much of the frontend's behaviour sits in small React-free modules (conversation density, header compaction, panel layout, focus-view image sizing, breakpoints), so it can be tested as ordinary functions. Source-pinned tests guard architectural shape:
MotionRevealholds no React state, effect or timer.- Every close of the focus view goes through one guarded path.
- Safe-area handling stays in CSS.
- Breakpoints have one owner, and every CSS mirror agrees with it.
These run on every push and in CI. They do not cover real interaction well: keyboard order, focus returning to the right place, gestures, anything that depends on how a browser lays out and paints.
Rendered-state verification covers part of that gap. A local command drives a real browser across declared surfaces and viewports and checks six invariants: the intended state was reached; no interactive control is covered by another; axe shows no regression against a baseline; no raw translation keys are visible; nothing overflows horizontally; no image is broken. It found both accessibility defects above. It depends on a browser tool that a fresh checkout doesn't have, so it runs in the development workflow, not in CI.
I also evaluated a multimodal model as a visual judge. It scored best of the approaches tried (F1 84.2, against 74.7 for pixel diffing) and still missed two of its five predeclared gates. It can help triage. It cannot pass or fail a change.
CI: every pull request, and again before deploy
TypeScript, strict
Biome, recommended lint rules
Unit, policy-module and source-pinned tests
Import boundaries for the photography interface
─────────────────────────────────────────── CI boundary
Local: run in the development workflow
React Doctor on the changed files
Rendered-state verification, including axe
Human visual and product review
Failure as evidence
Fast iteration is only useful if rejecting an experiment is cheap. Five rejections that shaped this frontend:
- A WebGPU orb would make a better launcher icon. Evidence: every platform gate passed, but at a device pixel ratio of 1 the icon lost its identity. Decision: keep the SVG for now.
- A universal AI chat panel would remove duplication. Evidence: the conversation shell already held the shared decisions. Decision: no further abstraction.
- Information-bearing depth would make the constellation easier to read. Evidence: 185 tests green, effect imperceptible in a browser. Decision: remove it, module and tests included.
- A React 19.3 fragment ref could replace a focus wrapper. Evidence: a compiling type probe showed the fragment instance has no equivalent of
.contains(), which the blur handling needs. Decision: keep the wrapper. - Repeated values should become tokens. Evidence: the same values expressed different decisions. Decision: keep them local and write the rule down.
Rejection usually meant deletion: the depth module and its tests, the two unchosen palettes, an older scroll-reveal component. Very little survives as a dormant flag.
The development loop
A substantial frontend change here tends to go through the same loop: an audit of the area, inspection of what already owns the problem, a bounded change in that owner, unit and source-pinned tests for its invariant, measurement in a real browser, product review by me, and then keep, modify or delete.
That is different from prompting for a piece of UI and shipping what comes back. The styling work is the clearest example: an audit, a correction to the written rules with no production change, then two narrow consolidations, each measured in a browser to change nothing visible. Not every change followed the loop that cleanly. Some audits landed in the same commit as their fix, the original ticket prompts are not stored, and a human decision is on record only where a report says so.
Limitations
The current boundaries:
- No formal internal design system. A convention layer over KERN, not a component library of its own.
- Many conventions are prose. Nothing mechanically enforces KERN-only components, inline-first styling, where
"use client"belongs, reuse of an existing primitive, or parity between the English and German message files. - No frontend routing category. The repository's context routing has no category for frontend tasks, so the frontend knowledge has to be opened deliberately.
- The component inventory can drift. The list of existing components is prose that nothing checks against the code.
- Browser verification is local. Rendered-state checks and React Doctor are part of the workflow, not CI.
- No interaction test suite and no persistent visual baseline.
- Incomplete device testing. Several mobile behaviours have only been checked in emulation.
- Authorship is not recorded. Git cannot show which lines an AI agent wrote.
- Documentation drifts too. Written guidance has fallen behind the code more than once, and it will again.
Conclusion
The coherence problem does not go away when an agent can write the code. It gets sharper, because more decisions are touched each day.
This architecture is not meant to let an agent make every frontend decision. It is meant to reduce how many decisions have to be invented during implementation. Shared primitives carry the decisions the product has already made more than once. Local ownership keeps a decision with its surface until a second surface makes the same one. Tests pin the invariants that should not drift. Browser measurement checks what the rendered page actually does. And a human stays responsible for whether the product is any good.
AI raised how much of this frontend could change in a day. Whether those changes added up to one product or to entropy was decided elsewhere: by who owned each decision, what was measured, and who said no.
Further reading
The repository-wide system behind this, including context routing, validators and agent roles, is in AI-Native Repository Architecture. The Photography Assistant and Photography Search case studies cover the features whose interfaces appear throughout, and Voice Interaction Architecture covers the local speech input.