Architecture
Photography Search Architecture
How my photo archive's search keeps metadata, meaning, appearance and concepts as separate signals, and lets a human-judged benchmark decide what ships.
Published October 5, 2026
Photography Search is the retrieval system behind my photo archive: it finds the photographs that answer a question, the ones that resemble a photo you are looking at, and the ones that share a photographic technique. Its central decision is that these are different questions. Exact metadata, descriptive meaning, visual appearance and photographic concepts each get their own signal with a bounded job, and a frozen, human-judged benchmark decides when two of them may be combined. Most proposed combinations failed that test. The ones that shipped are deliberately narrow.
The problem
"Find similar photos" sounds like one problem. In a photography archive it is at least four:
- Literal match. "Shot with the 23 mm", "Lisbon", "black and white", "2023". The answer is a fact in the metadata, and a photo either has it or it doesn't.
- Descriptive meaning. "Quiet streets at night", "Straßen bei Regen". No single field holds the answer. It has to be inferred from captions, tags and mood, and it may be phrased in English or German.
- Visual appearance. "More that look like this one": colour, light and shape, independent of what the caption says.
- Photographic technique. Leading lines, backlighting. These are compositional ideas that neither keyword nor caption similarity reliably captures.
These signals disagree. Semantic similarity happily ranks a photo that sounds related above one that provably matches. Image similarity finds a similar palette in an unrelated scene. Keyword matching misses a relevant photo whose caption uses a different word. Treating the scores as interchangeable, or adding them together because they all look like numbers, produces a ranking that is wrong in ways that are hard to see.
What I built
I designed and built the retrieval layer behind the Photography page. Visitors reach it in three places:
- Questions to the Curator. Natural-language questions run through the shared retrieval and ranking described below. The Curator then adds its own evidence admission and a language model that writes the answer; that part is covered in the Photography Assistant case study.
- Similar photos in Focus view. A default "similar" set plus four explicit lenses.
- Concept exploration. "Follow this technique" for a small set of photographic concepts.
The text-ranking function is also exposed as a standalone search procedure. That procedure is what the offline benchmark measures, and it is the only path that uses visual rank fusion (described below). It is a production system on a small corpus: a few hundred published photographs, one owner, anonymous visitors.
Architecture at a glance
Text retrieval: one path, two channels.
Question (EN / DE)
│
▼
Query interpretation
EXIF · places · time ·
vocabulary · colour · mood
│
┌────┴─────┐
▼ ▼
Lexical Semantic
+ filters (text vectors)
└────┬─────┘
▼
Eligibility
every requested facet
│
▼
Ranking
lexical + capped semantic
│
▼
Diversity caps
│
▼
Results + one true reason
Similar photos and concept exploration are separate paths with their own signals:
Similar photos (Focus view)
General ─ meaning vector
Subject · Atmosphere ·
Composition ─ facet vectors
Appearance ─ image vector
Concept exploration
stored concept assertion
→ middle similarity band
→ up to six photos
Nothing in the second diagram is fused with anything else. That is a measured decision, not an omission.
Why hybrid retrieval
Neither channel is enough on its own, so each is given the job it does reliably.
Lexical retrieval and filters own exact constraints. Query interpretation turns recognisable phrases into structured filters before any ranking happens:
- EXIF predicates ("f/1.4", "wider than 23 mm", "slow shutter");
- places, countries, regions and distance ("near Lisbon");
- dates, seasons and relative time;
- orientation, perceptual colour families and curated mood families;
- tags, matched through a controlled vocabulary.
Lexical scoring is a sum of fixed points per matched field: an exact tag or location match is worth 5, a mood match 4, a colour match 3. It runs over every published photo, so a photo with real metadata evidence can never be missed for lack of a vector match.
Semantic retrieval owns descriptive meaning. Whatever part of the question is not consumed by filters is embedded and matched against photo descriptions. When the remaining words carry no content, for example in a question that is entirely structured ("portrait photos from 2023 shot at f/1.4"), no semantic retrieval runs at all. The filters are authoritative, and similarity would only add photos that sound related.
The vocabulary bridges the two, with guard rails. A query word that is not a known tag, mood, place or colour can be promoted to its nearest known term by embedding similarity. A promoted word becomes exact evidence worth full lexical points, so a wrong promotion outranks everything semantic. Each category therefore has its own calibrated threshold (tags 0.55, moods 0.62, places 0.65, colours 0.70). Function words and the request language itself are not eligible. Before that rule existed, the German "und" could resolve to the tag "dock".
Eligibility comes before ranking. When a question combines several requirements ("rainy street at night"), a candidate must satisfy every requirement that is a hard property of the photo before it is ranked at all. Softer, experiential terms do not exclude candidates. On the Curator's path, if nothing satisfies every requirement, one may be relaxed and the answer says so. The standalone search returns an honest empty result instead.
Semantic retrieval
The document. Each published photo has one text document, built from its stored metadata: title, caption and alt text in both English and German, sorted tags, mood, colours and location. It is embedded with OpenAI's text-embedding-3-small (1,536 dimensions) and stored in a sqlite-vec index inside the application's SQLite database.
Versioning. Every vector carries a document version and a hash of the text it was built from. A vector is used only while both match the photo's current metadata, so editing a caption cannot leave a stale vector answering queries.
The query side. The question, or the part of it the filters did not consume, is embedded with the same model and matched by exact k-nearest-neighbour search. The recall window is three times the requested result count, bounded between 20 and 150.
Failure. If the embedding provider fails, the semantic channel returns nothing and lexical results are still ranked and returned.
What semantic similarity can establish: that a photo's description is close in meaning to the question. What it cannot: that the photo has a property. Two photos about "street at night" and "street at dawn" are close in embedding space, and only one of them answers the question. That limit is why semantic similarity contributes to ranking but never decides eligibility.
The bilingual document is load-bearing. German questions are answered almost entirely through the semantic channel: of the relevant photos found in a paired English/German study, 84.7% of the German ones arrived through semantic retrieval, against 27.0% of the English ones, because English words resolve to tags far more often. Removing the German half of the document lowered German NDCG@10 from 0.7392 to 0.6872 in a separate experiment. A multilingual alternative embedding the question and photo text in Cohere's space scored lower again (0.5527), so the single bilingual document stayed.
Ranking and fusion
Every eligible candidate gets one score: finalScore = lexicalScore + min(2 × semanticScore, 2).
The cap is the point of the formula. Semantic similarity can contribute at most 2 points, and a single exact tag match is worth 5. So no amount of semantic similarity can outrank exact evidence: a photo tagged "street" always ranks above an untagged photo that merely sounds like a street.
An earlier version did the opposite. It ranked primarily by similarity with small boosts for filters, and a photo tagged with the queried word could lose to an unrelated photo with high similarity. A weighted average would make that unlikely; the cap makes it impossible.
In practice the semantic term does two jobs:
- Ordering semantic-only candidates among themselves.
- Breaking near-ties between lexical candidates. An experiment shrank it to a pure tie-breaker for lexically matched photos, and the ranking came out byte-identical to production on all 99 scored benchmark questions.
Exact ties are common on a curated corpus, because many photos carry the same tags: half of all eligible candidates share both their score and their evidence tier. They are broken first by evidence strength: structured exact match, then curated equivalence, then weak colour evidence, then semantic-only, then no evidence. Only after that does a fixed photo order apply. As the evaluation section explains, that last step mattered more than it looks.
Visual rank fusion is the one place appearance enters text ranking:
- The question. Embedded with Cohere's multimodal model (embed-v4.0, 1,536 dimensions) into the same space as the photos' image vectors.
- The fusion. The five nearest images are combined with the production order by reciprocal-rank fusion, k = 60. Ranks are fused because cosine similarity in an image space and an additive evidence score are not on comparable scales, and every raw-score blend tested failed.
- What it may change. It reorders and never rescores. It may add an image-only candidate only when the question carries no structured constraint.
- When it runs. Only on the standalone search path, only for questions with visual descriptive content, and only when production's own top two results are clearly separated (more on that below).
The Curator's retrieval does not use visual fusion. It reuses the ranking with a stricter per-tag lexical cap and then applies its own evidence admission.
Diversity
Pure relevance ranking can still produce a poor gallery. A question that matches thirty photos from one evening in one city returns that evening thirty times. After ranking, two caps apply:
- at most five results per location;
- at most six per mood.
The top-ranked result is always kept, and photos with no location or mood are never capped. When the question explicitly asks for a place, the location cap is switched off, because demoting Venice photos in an answer to "Venice" would be diversity working against the question.
The cost is real: a capped photo can be more relevant than one that replaces it. The cap trades a little precision at the tail for a result page that shows the range of the archive. The reported match count is taken before diversity, so the cap never hides how many photos actually matched.
Visual similarity
Similar photos is a different intent from text search: there is no question, only a photo the visitor is looking at.
The default ("General") uses the same meaning vector as semantic retrieval: the six nearest photos by description, excluding the source, published photos only. It deliberately does not use image similarity. Two photos with the same palette and no shared subject are a poor default answer to "more like this".
Four explicit lenses sit beside it, each a single vector space with no blending:
- Subject, Atmosphere and Composition: vectors of short descriptions of each aspect, written by a vision model at upload and embedded separately.
- Appearance: the photo's image vector itself (Cohere embed-v4.0). Operationally, appearance means what the pixels look like (colour, light, tonality, shape), not what the photo is about. The vectors are computed at upload, so this lens makes no provider call at request time.
Appearance is good at "same light, same tonality, same kind of frame". It fails exactly where a caption would succeed: it cannot tell a harbour at dusk from a car park at dusk if they share the light.
I tested fusing the lenses into one "unified" result. Over several rounds it looked promising on calibration seeds. On 13 untouched seeds with 540 new judgments, the locked fusion configuration scored below General on NDCG@6 (−0.031) and P@6 (−0.026), and was worse on 6 of the 13 seeds. General stayed the default, and the lenses stayed separate and explicit.
Visual Concepts
Visual concepts let a visitor explore photographic structure: "show me more leading lines". The design goal was to make that possible without turning it into opaque AI categorisation.
- A small, fixed vocabulary. Four concepts are defined in code: leading lines, backlighting, atmospheric perspective and subject isolation. A model cannot invent a fifth.
- Assertions with evidence. At upload, the vision model may assert at most four concepts per photo, each with a confidence and a one-sentence piece of evidence about that photo. Exploration shows that stored sentence, never a generated claim about how photos relate.
- A narrower set for exploration. Only leading lines and backlighting are offered as exploration paths: they are the two concepts the selection policy below was evaluated on.
The interesting decision is which concept photos to show. Showing all photos with the concept returned nearly the same set from every starting photo: cross-source overlap (Jaccard) was 0.46–0.52. Ranking them by similarity to the current photo duplicated what Similar photos was already showing.
The shipped policy keeps only the middle third of the similarity distribution: related to the current photo, but not its nearest neighbours. Within that band the order is fixed, so similarity cannot creep back in through the ordering. It was selected by a predeclared rule, won on blind-graded pivot quality (2.50 of 3 against 2.44 for the next eligible strategy), and held on a holdout. The samples were small (9 calibration and 6 holdout sources), which is one reason the feature stays narrow.
A follow-up that tried to turn concepts into recommendation trails was rejected. Concepts remain provenance (why a photo belongs), not a navigation engine.
Why this result matched
Every result can show why it matched, and that explanation is a projection of the evidence retrieval recorded, not a story written afterwards.
Each candidate carries a record of exactly which signals matched: which tags, which place, which camera property, whether the match came from semantic similarity. The UI selects one reason, the highest-priority one that is true (tag before place before camera, and so on, with semantic similarity last), or shows nothing if no specific signal applies.
Two earlier behaviours were removed because they claimed more than retrieval knew:
- A generic "strong match" label. It was shown whenever several weak signals coincided.
- A visible match percentage. It was relative to the top result, not a probability, and read as confidence it did not have. It is still computed internally, but no longer shown to visitors.
The rule: an explanation may never be stronger than the evidence behind it.
Evaluation
Search changes are decided offline, against a frozen benchmark, because a live corpus changes under you and live impressions confuse "looks better" with "is better".
The benchmark:
- 106 questions over a frozen snapshot of 191 photos, in 21 query families. They range from literal to natural-language, structured, relational and mixed phrasing, include negative questions, and cover English and German.
- 2,460 human relevance judgments on a 0–3 scale, graded blind with every model signal hidden.
- Candidates pooled from six independent retrieval channels so the benchmark is not tuned to production's own output.
- When 17 questions came back with every candidate relevant, they were re-pooled at twice the depth. Five remained saturated, and that is recorded as a limitation rather than patched over.
The metrics, and why there are several:
- NDCG: graded ordering quality.
- MRR: how soon the first relevant photo appears.
- P@5 / P@10: precision of the visible page.
- R@10 / R@20: recall within it.
They disagree often enough that no change ships on one of them. Results are also broken down by query family and language, because English and German retrieve through different channels (above) and a pooled average can hide a regression in either.
Measuring the measurement. The most consequential evaluation finding was about the benchmark itself:
- Half of all eligible candidates (50.3%) sat in exact ties, sharing both score and evidence tier.
- Ties were broken by photo order, and every judged photo happened to precede every unjudged one, which counts as irrelevant.
- Production's tie-break therefore systematically pushed judged photos up, inflating every metric.
Re-scored over 200 random tie orders:
- NDCG fell from 0.634 to a tie-neutral mean of 0.593 (95% interval 0.584–0.604).
- MRR fell from 0.671 to 0.633.
- P@5 fell from 0.533 to 0.475.
Production still beat all 200 tie orders on every metric, so earlier ranking verdicts held. But absolute baselines were invalidated, and rankings are now reported as a tie-neutral mean with an interval.
What failed
A better ranking that was an artefact.
- Problem: German and descriptive queries seemed to suffer because semantic evidence was weighted too lightly against lexical points.
- Hypothesis: rebalancing the two terms improves ranking.
- Experiment: seven formulations on 99 scored benchmark questions, among them raising the cap, rank-based scaling, role-aware weights and removing the semantic term for lexical candidates.
- Result: nothing moved except removing the semantic term, which gained 4.16 NDCG points. That gain was produced entirely by the tie artefact above. Under paired tie-neutral comparison the same change was worse (−1.70 points, interval −3.07 to −0.36).
- Architectural consequence: the scoring formula stayed. Tie-neutral reporting became the default protocol for any ranking experiment.
Fusion that hurt where the ranking was uncertain.
- Problem: visual rank fusion had shipped on an earlier benchmark (+1.8 NDCG points, +10.6 MRR points), gated by a check for visual intent.
- Experiment: an audit on 117 frozen judged queries.
- Result: the gate admitted 42 of the 49 queries fusion harmed. Net, fusion was 1.9 points per query worse than no fusion at all, while an undeployable oracle showed 7 points of headroom. No runtime feature (pool size, facet count, visual similarity) separated helpful from harmful activations.
- A follow-up on 136 judged queries found where the harm lived: almost entirely in questions where production's own top results were tied or nearly tied.
- Architectural consequence: fusion now runs only when production's top two results are separated by at least 0.05.
- Against no fusion that is +1.58 NDCG@10 points (0.6879 against 0.6721), and +0.88 P@5. Activation fell from 89.7% to 31.6%, and helped/harmed queries moved from 41/43 to 23/8.
- The rule beat no fusion on 182 of 200 held-out splits, and any threshold from 0.02 to 0.70 behaved similarly, so it is a mechanism boundary, not a fitted constant.
- Retiring fusion was considered and rejected: the image vectors stay for Similar photos and concepts anyway, so retirement would have saved nothing and lost the gain.
More recall that nothing could reach.
- Problem: maybe relevant photos never reach the results because candidate generation misses them.
- Experiment 1: a BM25 caption index as an extra discovery channel. It found real additional candidates (+1.17 points of candidate recall), but none of them reached the top 5, 10 or 20 under any admission policy, including a ground-truth oracle (NDCG +0.00).
- Experiment 2: a deeper semantic recall window (150 instead of 60). It raised candidate recall from 93.2% to 97.1% and lowered NDCG by 0.35 points; no newly found relevant photo reached the top 20.
- Architectural consequence: candidate generation is not the bottleneck; scoring and evidence are. No new retrieval channel was added. The question "what would a perfect oracle gain?" now comes before building one.
Trade-offs
- More than one retrieval path. Lexical, semantic, image and concept signals each need their own storage, versioning and failure behaviour. That is more moving parts than vector search alone, accepted because each one answers a question the others answer badly.
- A cap instead of a learned blend. The additive cap is simple and guarantees exact evidence wins, but it is a hand-set structure, not a model fitted to relevance. A learned ranker might do better on average. On a few hundred photos there is not enough judged data to train one without overfitting, and it would give up the guarantee.
- Diversity demotes relevant photos. The caps trade tail precision for a page that represents the archive.
- Visual similarity is not relevance. Appearance answers a different question, so it lives behind an explicit lens and a narrow fusion gate, never in the default.
- Hosted embedding providers. Semantic retrieval and fusion depend on external models. Both fail open to the lexical path, so the cost of an outage is quality, not availability.
- Offline metrics are only as good as the judgments. All of them come from one human grader. Five questions are saturated, and the German slice is small (11 questions). Several findings above came from correcting the benchmark, not the ranker.
- Simple infrastructure because the corpus is small. Exact nearest-neighbour search, in-memory lexical scoring and a vector index inside the application database are the right choices for a few hundred photos. They are not choices I would keep at a different scale.
Scaling
Today: a personal archive of a few hundred photographs, a single-owner upload workflow that enriches and embeds each photo once, and anonymous public reads. Here is what would change:
- 10,000 photos. The in-memory lexical scan becomes measurable per query. Lexical matching would move to an inverted index, as an index structure for tags and filters, not as BM25 ranking, which was measured to hurt. Exact vector search is still viable. The benchmark would need re-pooling, because judgments for 191 photos do not cover 10,000.
- 1,000,000 photos. Exact k-nearest-neighbour search and full-scan image similarity are no longer viable, so approximate indexes (HNSW or similar) in a dedicated vector store would replace them. Approximation changes recall, so the recall window, the diversity caps and the fusion gap threshold would all need re-measuring. Human judgment would have to be sampled rather than pooled exhaustively.
- Multiple users or tenants. Tenant scope becomes part of retrieval and authorisation, not just publication state. Vocabularies, mood families and concept registries are currently global and curated by one owner; they would need per-tenant ownership or a shared governance model.
- Higher query volume. Each question may make several provider calls (vocabulary promotion, semantic embedding, concept resolution, fusion). Embeddings are cached today; at volume they would need batching, budgets and per-client limits, and fusion's extra call would be the first candidate to drop.
- Incremental indexing. Already in place in small form: vectors are versioned by document version and content hash, so re-enrichment replaces exactly what changed. At scale this becomes a queue with backfill and reindex jobs.
What would remain: separate signals with bounded jobs, exact evidence that similarity cannot outrank, eligibility before ranking, diversity as a deliberate final step, and a frozen human-judged benchmark with tie-neutral reporting deciding what ships.
Engineering lessons
- Different notions of "similar" should stay separate until a measurement says otherwise. Each fusion I tested looked reasonable. Only one survived, and only in a narrowed form.
- Bound a signal's job structurally. The semantic cap is a guarantee, not a tuned weight, and an experiment confirmed it already behaves as a tie-breaker where exact evidence exists.
- Measure the measurement. A tie-break order flattered production by four NDCG points and briefly made a worse ranking look better.
- An average can hide where harm lives. Fusion's net loss was concentrated in tied orderings; finding that boundary turned a net-negative feature into a net-positive one.
- Ask what a perfect oracle would gain before adding recall. Twice, more candidates turned out to be unreachable by the ranking that would have to use them.
Further reading
The Photography Assistant case study covers how the Curator turns this retrieval into grounded answers. The engineering journal tells parts of the search story in narrative form: Designing Search for Memories and Every Match Needs an Explanation.