Skip to main content

Architecture

Photography Assistant Architecture

How a grounded LLM assistant over my photo archive works: retrieval and an eight-photo evidence pack bound what the model sees; validators gate what it may say.

Published October 5, 2026

The Photography Assistant is a production LLM system that answers natural-language questions about my photo archive — "quiet streets at night", "what did I shoot with the 23mm?", "why do these belong together?" — with a short answer and the photographs that support it. The architectural premise is simple to state and most of the work to keep: the language model writes the answer, but it never decides what is true. Retrieval decides what the model may see, and deterministic validators decide what it may say.

The problem

A conversational interface over a personal archive looks like the textbook case for "question → LLM → answer". It isn't, for four reasons.

  • The facts live in metadata the model has never seen. Which lens, which city, which evening, which mood — none of it is world knowledge. An answer is only correct if it is grounded in this archive.
  • Plausible is the failure mode. An ungrounded model will happily describe a photo that doesn't exist, or a relationship between two photos that isn't there. In a portfolio, a confident wrong answer is worse than no answer.
  • Meaning and metadata disagree. Semantic search finds photos that sound right; metadata finds photos that are provably right. Neither is enough on its own.
  • Inference must stay bounded. Every model call adds latency, cost and a new way to fail. The system has to work when a provider is slow, rate-limited or wrong.

What I built

The Assistant runs in production behind the Photography page, where visitors meet it as the Curator: type or speak a question, get a short answer, and see the photographs it cites. I designed and built the whole path — retrieval, evidence construction, the generation contract, the validators, the streaming answer UI and the evaluation tooling around it.

It is a production system on a small corpus: a few hundred published photographs, one owner, public visitors. That shapes several decisions below, and I say so where it does.

Architecture at a glance

Visitor question
        │
        ▼
Rate limit · escalation check
        │
        ▼
Context resolver
  (deterministic; planner LLM
   only when it cannot decide)
        │
        ▼
Candidate pool
  lexical + semantic + filters
        │
        ▼
Rank · rerank top 16
        │
        ▼
Admission (corroboration)
        │
        ▼
Evidence Pack: ≤ 8 photos
  + computed relationships
        │
        ▼
LLM → typed JSON
        │
        ▼
Gates: contract · claims ·
  citations · evaluations ·
  scope
  (at most one retry)
        │
        ▼
Answer + cited photos
  • Rate limit and escalation check. Ten questions per minute per IP address; a pattern check refuses requests that try to reach beyond the gallery ("dump all evidence") before any retrieval runs.
  • Context resolver. Follow-ups ("only the ones at night", "more like this one") are resolved by deterministic rules against a small, server-authored context the browser echoes back. A planner model is consulted only when the rules cannot decide, at most once per turn, and its output can only select one of the plans the rules already define.
  • Candidate pool and ranking. Deterministic and semantic retrieval run together; a reranker reorders the head of the list. Nothing here is generated.
  • Admission. A candidate has to earn its place in the evidence. This is where most precision is won or lost.
  • Evidence Pack. At most eight photos, with their stored facts and the relationships computed between them. This is everything the model will ever see.
  • Generation and gates. The model returns a typed object, not free text. Every structured assertion in it is checked against the pack before anything reaches the visitor.

The key architectural decision

The language model is a component inside a deterministic system, not the system itself. Concretely, it is given exactly two jobs that rules did not do well — choosing which of the admitted photos actually answer the question, and writing a short, specific answer about them.

It does not control:

  • which photos exist or are published;
  • what is retrieved, or how it is ranked;
  • which photos are admitted as evidence;
  • which relationships exist between photos (they are computed before the model runs);
  • what the response contains structurally (a schema enforced at the provider and again in the application);
  • whether its own claims are accepted.

Two consequences follow. First, the model has no tool, no database access and no path to anything outside the Evidence Pack — that, not a prompt instruction, is the real security boundary. Second, a provider change is a configuration change: production generation runs through a fallback chain of hosted models (GLM 5.3 Flash on a pinned endpoint, then GLM 5.3, then Claude Haiku 4.5), each tier with its own deadline, and every tier passes through the identical validation chain.

Retrieval before reasoning

Retrieval is application code, shared with the Photography search but configured for the Assistant's job. Apart from the embedding and reranking models it calls, every step in it is deterministic.

  • Lexical and filter retrieval matches the question against captions, tags, moods and locations, and turns recognisable constraints into hard filters: EXIF predicates ("shot at f/1.4", "wider than 23mm"), places, countries and time ranges. Technical questions are filtered and ranked from EXIF data by rules; the model only narrates the result.
  • Semantic retrieval embeds the question and searches photo descriptions in a vector index, which finds photos that match in meaning rather than wording — and also finds photos that merely sound related.
  • Facet eligibility. When a question combines several requirements ("rainy street at night"), candidates must satisfy all of them. If nothing does, one requirement may be relaxed, and the answer says so explicitly rather than quietly returning a near miss.
  • Reranking. For descriptive questions the top 16 candidates are reordered by a reranking model (Voyage rerank-2.5). It fails open: if the reranker is slow or down, the deterministic order is used. Technical and browsing questions skip it entirely.
  • Admission. The top eight candidates then have to earn their place; nothing is backfilled. When at least one of them shares wording with the question, a frozen rule decides: a photo is admitted if its own caption or tags corroborate the question, or if its capped semantic score plus its rerank score clears a threshold (1.01) chosen by cross-validation in an offline experiment and not tunable at runtime. When none of them does — typically a mood or descriptive question — a photo found only by semantic similarity must be corroborated by its caption, tags, mood or visual concepts, with one narrowly scoped exception described under trade-offs.

The admitted photos, at most eight and in ranked order, become the evidence.

The Evidence Pack

The Evidence Pack is the contract between retrieval and generation. For each of at most eight photos it carries the stored caption, tags, mood and location; why the photo matched and how strongly; the camera, lens and capture settings, plus derived categories ("wide aperture", "evening"); and the relationships to other photos in the pack.

Those relationships are computed, not inferred. Seven kinds — same location, shared mood, shared tag, same lens, same day, similar colour, same orientation — are derived from stored data before the model runs, with at most eight related photos per photo. When the answer says two photos share a lens, that edge already existed in the pack.

The pack is deliberately small and deliberately text-only. Eight photos is enough for a conversational answer and small enough that every one of them can be checked; the model sees no pixels and no image URLs, and is instructed not to describe what an image looks like beyond what the stored evidence says. Bounding the pack bounds everything downstream: prompt size, cost, latency and — most importantly — the set of things the model could possibly cite.

Structured generation

The model's output is a typed object with four fields: the answer prose, the photo IDs it cites, a list of relationship claims (each a relationship kind plus two to eight photo IDs, at most twenty), and optional per-photo evaluations (a short observation backed by one to five named pieces of stored evidence).

The schema is sent to the provider as native structured output, so malformed responses are largely prevented at the source, and it is validated again in the application, because provider enforcement is not a guarantee. When I moved from prompt-described JSON to native structured output, structural failures in the evaluation runs went from six to zero.

A valid schema is not a correct answer. It guarantees that the answer has the right shape — that every claim names a known relationship kind and every evaluation names its evidence — which is what makes the next step possible. Whether those claims are true is the validators' job.

Citation and claim validation

Every response passes the same gates, in a fixed order:

  1. Contract — the response matches the schema.
  2. Claims — every relationship claim must correspond to an edge that exists in the pack. One unsupported claim rejects the response.
  3. Citations — every cited photo must be in the Evidence Pack. A response that cites anything else is rejected outright, without a retry; a fabricated photo is the one failure that is never repaired.
  4. Evaluations — every evidence reference must point at stored data, and at least one must be about the photo's content.
  5. Leak redaction — internal identifiers are stripped from the prose.
  6. Result scope — counting words in the prose ("all three", "both") must agree with the number of photos actually cited, in English and German.

A failure in the contract, claims, evaluation or scope gate allows exactly one regeneration, with an instruction that names what was wrong. A second failure becomes a refusal that shows no photos — the visitor gets an honest "I couldn't answer that" rather than a partly verified answer. The one exception is a second counting mismatch: the evidence is valid, so the answer is delivered and the mismatch logged rather than discarding a correct result set over wording. While the model is still writing, its prose streams to the page as clearly provisional text; only the validated response is ever the source of results.

What this guarantees: every photo shown was retrieved and admitted; every relationship stated as a claim exists in the data; every evaluation cites real stored evidence. What it cannot guarantee: that every sentence of free prose is true. The prose is constrained by the prompt and checked for counting and leaks, but it is not fact-checked sentence by sentence, and the evidence itself is only as good as the stored captions and tags. That boundary is deliberate and stated, not hidden.

What remains deterministic

  • Rate limiting, input validation and the escalation check.
  • Follow-up resolution, apart from the bounded planner call on ambiguous turns.
  • All retrieval: filters, EXIF predicates, location and time scope, facet eligibility and relaxation.
  • Ranking, diversity limits, the admission rule and its threshold.
  • Evidence Pack construction and every relationship in it.
  • Schema validation, all six gates, the retry policy and the refusal text.
  • The conversation context, which is authored by the server and re-validated on every turn.
  • Result ordering, summaries and photo metadata shown in the UI.

Besides the answer writer and the embedding model behind semantic search, three steps use models, each with a narrow mandate and a deterministic fallback: the reranker (falls back to the deterministic order), the planner (falls back to the deterministic plan) and a semantic-evidence judge described under trade-offs (falls back to the deterministic admission result). None of them can get anything past the gates.

What failed

Replacing the model's photo selection with rules. Problem: the model chooses which pack photos to cite; that is its least auditable job. Hypothesis: a transparent rule over retrieval signals could do it as well. Experiment: 106 benchmark questions, human-judged. Result: the model (Claude Haiku at the time) reached 84.4% precision, the best rule 80.9%, citing the whole pack 79.5%; the model's decisive drops — such as citing nothing from a pack of eight irrelevant photos — correlated with no signal available before generation, and a quality-preserving rule would have saved only about 220 prompt tokens. Consequence: selection stayed with the model; the case where it had been forced to explain away irrelevant photos ("shot at f/1.4") was fixed upstream with deterministic EXIF predicates instead.

Retrying without making things worse. Problem: validation was rejecting a meaningful share of answers, each one a refusal. Experiment: an A/B/C over 17 deliberately failure-heavy questions, five repeats each. Result: a better prompt alone barely moved first-attempt failures (28% → 27%); a single retry told exactly which check failed brought final failures to 14%, at an average of 1.26 model calls per request. A later audit found that the retry for rejected claims was quietly changing which photos were cited — destroying correct citations while fixing the claims. Freezing the cited set during that retry preserved it in 26 of 28 runs versus 14 of 28, and delivered 104 human-relevant photos versus 82. Consequence: one bounded, informed retry; claim repairs may not change the evidence. Mean precision had hidden the harm entirely — the lesson was to measure what a visitor actually receives.

Compressing the evidence to save tokens. Problem: the prompt is the largest cost. Experiment: a compact tabular encoding (TOON) for the evidence, the output, or both, in interleaved runs on frozen packs. Result: the lossless version saved 7.4% of tokens — about 46 ms on a request of roughly 4.5 seconds — while encoding the output compactly dropped first-pass reliability to 66.7%. A 16% saving in an early prototype turned out to come from silently dropping evidence. Consequence: the evidence format stayed, and the rule became "fewer tokens without a measurable latency gain is a rejection".

Choosing the model. Cheaper hosted models first failed on latency (2.9× slower at the median) and availability (a third of requests erroring), for a saving of a few dollars a year — not a reason to switch. A later tournament found a model (GLM 5.3) that beat Claude Haiku on blinded usefulness and first-pass reliability, but appeared to emit no relationship claims in German at all. A follow-up showed that was an artefact of the test set: every German question was a listing question, for which the contract says to emit none. Consequence: a provider-neutral generation layer and a fallback chain, with the switch made as configuration — and a standing suspicion of any evaluation whose fixtures can't see the feature being measured.

Evaluation and reliability

These are three different mechanisms, and none substitutes for another.

Offline evaluation decides whether a change ships. Retrieval changes are measured against a frozen benchmark of 106 questions over 191 photos with 2,460 human relevance judgments. Generation changes are measured with interleaved runs on frozen evidence packs against predeclared gates and budgets. One of the most useful findings was about the benchmark itself: many candidates tie on score, and the way ties happened to be broken flattered production by about four NDCG points. Production still beat all 200 random tie orders — but rankings are now reported over many tie orders, not one.

Production safeguards run on every request: the gates, the single bounded retry, fail-open retrieval components, fail-safe refusals, per-tier deadlines and rate limits.

Runtime observability records what happened: an event per gate, the candidate funnel from retrieval to evidence, planner, rerank and admission decisions, and a log of public questions and their outcomes. When tracing is enabled, it records metadata only — never prompts or answers.

Trade-offs

  • More code than direct retrieval-augmented generation. Six gates, a retry policy, an admission rule and a context resolver are a lot of machinery for a photo archive. Each exists because a measured failure needed it.
  • Honest refusals cost answers. A response that fails twice is refused even if most of it was right. I chose an occasional "I couldn't answer that" over an occasional confident fabrication.
  • Validators and prompt must evolve together. The prompt carries 24 numbered rules, and several gates exist to check them. Changing one without the other is a regression.
  • Eight photos is a ceiling. The bound keeps answers checkable and cheap, and it also limits how broad an answer can be.
  • Recall for precision, explicitly. For questions where no candidate shares the question's wording, deterministic corroboration was very precise and missed much genuinely relevant evidence. I added a narrowly scoped judge model that may only promote a rejected semantic-only candidate: in English it raised recall on that residual from 33% to 95% while precision fell from 100% to 91%. In German it improved recall similarly but missed its own precision bar; I shipped it anyway, as a product decision for an exploratory surface, because its authority is bounded: it can only turn a rejection into an admission, never the reverse; it cannot reorder anything; its verdict is never itself evidence the answer can cite; and any failure returns the deterministic result.
  • Provider dependence. In a September measurement the primary model tier was rate-limited on roughly half of all turns; the fallback chain carried them, at the cost of multi-second answers and more moving parts.

What I would change at larger scale

Some of the current simplicity is deliberate for one owner and a few hundred photos; some of it would not survive more.

  • Larger corpora: the vector index lives in the application's SQLite database, and the eight-photo pack and rerank depth were tuned on this corpus. Both would need re-measuring; I would expect to move to a dedicated vector store well before the pipeline shape needed to change.
  • More traffic: rate limits are held in process memory, which is correct for a single instance and wrong for several. They would move to a shared store, and model calls would need a budget and queueing policy rather than just deadlines.
  • Multiple users or tenants: the evidence boundary already scopes what the model can see per request; tenancy would make that scope part of retrieval and authorisation, not just of publication state.
  • Stricter SLOs: offline evaluation is run deliberately rather than on every change. With an availability target, a small, cheap regression subset would run in CI, and latency would get an explicit budget per stage.

The conversation design would not need to change: the server keeps no session, and each turn carries only a small, validated context.

What this changed in how I build AI systems

I started this expecting model quality to be the main lever. It mattered — the current model was preferred in blinded comparison — but it worked inside limits set elsewhere: no model could cite a photo that retrieval had not admitted, and several of the largest measured changes in what visitors receive came from admission rules, evidence boundaries and the shape of a single retry. Two of the most consequential findings were errors in the measurements themselves. What made the system trustworthy was giving the model the jobs rules couldn't do as well — choosing and explaining — and making everything around those jobs checkable.

Further reading

The engineering journal covers parts of this path in narrative form: Every Match Needs an Explanation, Pipelines Beat Prompts and Designing Search for Memories.