Architecture
Photo Visual Constellation
Every published photograph on one map, placed beside the photographs it resembles. The layout is computed offline, React holds only what the visitor selects, and the camera, animation and pixels run on a single canvas outside React.
Published October 7, 2026
Photo Visual Constellation is the "Visual Connections" view of my photo archive: every published photograph placed on one map, close to the photographs it resembles. You can pan across it, zoom from a field of stars down to individual photographs, and follow one photo to its nearest visual neighbours. Its central decision is a strict division of labour. The layout is computed offline and never changes while you look at it. React owns what you have chosen. Everything that changes every frame, including the camera, the animations and the pixels, lives outside React, on a single canvas.
The problem
The archive holds 381 published photographs. I wanted a view that answers a question the grid and the map cannot: what does my photography look like as a whole, and which pictures belong together visually?
That turns into four requirements:
- The whole corpus must fit on one screen, so that clusters, gaps and outliers are visible at a glance.
- Any region must become readable. Zooming in has to turn anonymous points into actual photographs, at their real aspect ratio.
- Movement has to keep up with the hand. Dragging, scrolling and pinching cannot wait for anything.
- The map must stay the same between visits. If the arrangement shifted every time I added a photo, nobody could learn it.
Only the last requirement is about data. The first three are about rendering: hundreds of images, continuously re-sized, cross-faded and re-positioned, at the rate a pointer moves. So the architecture is mostly an answer to one question: what has to happen per frame, and what does not?
What I built
The view is one lens of the Photography page, next to the grid and the map. When it opens:
- At first, every photograph is a small star tinted with its dominant colour, over a soft colour field that marks the dense regions.
- Dragging pans; the scroll wheel, a pinch or the zoom buttons zoom; a reset button frames everything again.
- Zooming in turns stars near the centre of attention into circular thumbnails, and further in, into photographs at their true shape.
- Hovering a photograph lifts it and, after a short pause, lights up its closest visual neighbours.
- Clicking selects it. Lines connect it to its eight nearest neighbours, the rest of the map quietens, and a side panel shows the photo with actions such as opening it, starting a "visual journey" through neighbouring images, or seeing where it was taken.
- Clicking the selected photograph again opens it in the full-screen view.
The neighbours come from an image-similarity search, the same one behind "similar photos" elsewhere on the page. The map itself has no edges. Position is the relationship, and the page tells screen-reader users so: position is illustrative, not a measurement.
Architecture at a glance
- Offline layoutUMAP-2D over image embeddings · aligned to the previous run
- Stored pointsphotoId, x, y · 381 points · 23 KB
- React surfacelayout fetched on open · selection, neighbours, inspector
- Derived geometry refonce per layout · spacing, density, territories
- Camera refoffset + scale · pointer, wheel, pinch, buttons
- One frame looprequestAnimationFrame · stops when nothing moves
- One Canvas2D surfacestars → circles → photographs
- The offline layout reduces each photograph's image embedding to two coordinates. It runs on demand from the admin, never during a page request.
- The stored points are the only thing the browser receives about the layout: three numbers per photograph.
- The React surface fetches those points when the lens opens, joins them to the photo metadata the page already has, and owns everything a visitor chooses: the selection, its neighbours, the side panel and the text read to assistive technology.
- Derived geometry is computed once per layout. It covers point spacing, density and colour territories.
- The camera is a plain object in a ref, changed directly by input.
- One frame loop advances whatever is animating and redraws, then stops.
- The canvas draws every photograph at whatever representation the current zoom calls for.
Where the positions come from
The positions are computed offline:
- A layout run takes the image embedding of every published photograph (1,536 dimensions) and projects it to 2D with UMAP. The run is seeded, so the same inputs always produce the same map.
- It then rotates, scales and, if needed, mirrors the new result to best match the previous layout before storing it. This is a closed-form 2D alignment.
- Each run is a versioned row with its parameters, a hash of the photo set, and how well the alignment fit. The newest completed run is the one served.
The alignment exists for returning visitors. UMAP is free to produce a rotated or mirrored version of the same structure, and without alignment a rebuild after a new upload could flip the whole map. The last production rebuild moved existing photographs by about 0.3% of the layout's radius.
Keeping this offline was not a performance optimisation first. The projection needs the vectors, and those never leave the server. Running the projection per request would also cost far more than serving its result: 23 KB of coordinates for 381 photographs.
The browser never normalises the coordinates either. UMAP units are arbitrary, so every constant on the client is relative: zoom limits are multiples of the scale that fits the whole corpus, and sprite sizes follow the typical distance between neighbouring points. A layout that comes out twice as large renders identically.
One canvas, not a node per photograph
Each photograph is not an element in the page. The stage is a single <canvas> drawn with the browser's 2D API, plus a few decorative layers, one native button around the canvas and one live region for announcements. That count is the same at every zoom level and with any selection. Adding photographs adds pixels, not elements.
Three things pushed me there:
- Representation changes continuously. A photograph is a star, then a growing circle, then a rounded frame that morphs to its true aspect ratio, with cross-fades between all three, and the change depends on both zoom and position. As DOM nodes, that would mean hundreds of elements changing size, shape and opacity at frame rate.
- The underlay is a raster. The colour territories are a density field. On a canvas they are one small bitmap, painted through the same camera as the photographs, so they cannot drift out of alignment with them.
- No library was needed. There is no Three.js, WebGL, D3 or physics engine here. Canvas2D is enough at this size.
I did build WebGL experiments for atmosphere effects. They were removed for how they looked, not for performance.
The cost is real. Pixels have no semantics: no focus, no accessible name, no hit area. Everything an element would give me for free, I either rebuilt (hit testing, one keyboard-operable control for the whole stage), moved elsewhere (labels, the side panel, a selection announcement), or don't have (per-photograph semantics and spatial keyboard navigation; see below).
React owns intent, refs own frames
The split I enforced is about frequency:
- React state, changed when the visitor decides something: the selected photo, its neighbours, the side panel, projected populations.
- React state, changed only when the photo under the pointer changes: the hovered photo.
- A ref, changed on every pointer, wheel and touch event: the camera (offset and scale).
- Refs, changed every frame while animating: emphasis ramps, camera easing, glow position.
- Refs, computed once per layout: spacing, density field, colour territories.
- Refs, filled once per photograph: loaded images, sprite caches.
The canvas component has no useState at all. Props arrive through ordinary rendering and are copied into refs after each commit, so the draw loop always reads the latest committed values. A render React throws away never reaches the screen. Commands go the other way through an imperative handle: fit the corpus, frame a neighbourhood, zoom in one step. The parent asks the canvas to move the camera; it never holds the camera itself.
That has one visible consequence. The zoom buttons cannot show a disabled state at the zoom limit, because that would mean copying the camera into React just to grey out a button. They clamp silently instead.
Measured on a production build against the real 381 photographs:
- Panning, wheel zoom, button zoom, resizing and sitting idle cause zero React renders of the constellation.
- Hover re-renders only when the hovered photograph changes. That was 23 renders for a sweep of 61 pointer moves across the whole map.
React Compiler is on in this repository. It compiles the React side of the feature (the stage owner, the side panel, the controls). It declines to compile the canvas component, because of counters mutated inside closures. That costs nothing here, since the canvas does no reactive work worth memoising.
Coordinates and the camera
There are four coordinate spaces, and only one transform between the first two:
- UMAP spacearbitrary units · aligned between rebuilds
- Camerafit: scale = min((W−2p)/spanX, (H−2p)/spanY) · { offsetX, offsetY, scale }, scale ∈ [0.5, 32] × fit
- CSS pixelssx = x · scale + offsetX · the stage's canvas box
- Device pixelsctx.scale(dpr), dpr ≤ 2 · the canvas backing store
- The camera is three numbers. Panning adds the pointer delta to the offset.
- Wheel zoom multiplies the scale by
exp(−deltaY × 0.0015)and keeps the point under the cursor fixed. A pinch does the same around the midpoint of the two fingers. - The buttons step by exactly one wheel notch, about 1.16×, and ease over 400 ms.
- Zoom is limited to between half and 32 times the scale that fits the whole corpus.
Two rules keep the camera predictable:
- Until the visitor touches it, the camera keeps refitting whenever the stage resizes. The first measurement of a freshly mounted element is often not its final size. Once the visitor pans or zooms, the camera is theirs, and nothing moves it without an explicit action.
- Selecting a photograph moves the camera the minimum distance needed to keep it comfortably on screen, and never changes the zoom. A visitor who clicked a photo has not asked to be taken anywhere.
Device pixel ratio is capped at 2. On a phone with a ratio of 3 that saves more than half the backing-store pixels, for a difference I could not see on small sprites.
One frame loop
All motion goes through one requestAnimationFrame chain. Each tick:
- advances the camera easing (400 ms), the hover and selection emphasis (160 ms), and the glow that follows the selected neighbourhood;
- lets a few newly loaded photographs appear;
- redraws once;
- schedules itself again only if something is still moving.
There is never a second chain. Two independent loops would let two subsystems each draw in the same browser frame.
The exception is the star field. At corpus scale the stars twinkle slowly, on periods of several seconds, so the loop stays alive while stars are visible, throttled to about 30 draws per second. Sampling a multi-second motion every 33 ms is indistinguishable from 60 fps and halves the idle cost. As soon as photographs take over the screen, the loop stops by itself. With reduced motion requested, the stars are static and the idle page draws nothing: 0 draws in five seconds, against 148 without the preference.
Not every draw goes through the loop. Dragging and wheel zoom redraw directly inside the event handler, so the picture never lags a frame behind the pointer.
The loop also has a lifecycle, not just a current frame. Images keep arriving after the view closes, and each one used to ask for a redraw. When the view is closed, the loop is disposed and refuses every later request, so a late image can no longer start it again.
Stars, circles, photographs
What each photograph looks like is a pure function, recomputed for every photograph on every frame. It takes the current zoom relative to the corpus fit, and the photograph's distance from the centre of attention: the selected photo if there is one, otherwise the middle of the stage.
- Below about 1.1× fit, everything is a star.
- Between about 1.1× and 2.2×, photographs near the centre of attention fade in as circles while their stars fade out.
- From about 3.5× to 7×, circles morph into rounded frames at the photo's real aspect ratio. The aspect comes from stored dimensions, so the shape is right before the image has even downloaded.
Two refinements keep the transition gradual:
- Each photograph's threshold is shifted slightly by how much room it has around it, so the map does not turn photographic in one step.
- The selected photograph and its neighbours get a floor, so they resolve first.
On a mouse, the area around the pointer is also brought forward; on touch this does not exist.
The same function decides the hit area. A photograph currently drawn as an image gets an image-sized target; a star keeps a 14-pixel one. Because click handling calls exactly what drawing calls, what you click is always what you see.
Computing this for all 381 photographs costs about 49 µs per frame, measured outside the browser. That is why there is no spatial index or virtualization: a full pass over the corpus every frame is cheaper than maintaining either.
Loading the photographs
Images are the real cost of this view, and they were also the source of its worst measured problem.
What is fetched. The canvas uses each photo's existing 800-pixel thumbnail derivative, in its WebP version. There is no canvas-specific image to generate or store. It is served with a one-year immutable cache header and drawn into the canvas, center-cropped to the current shape. There are no <img> elements in the stage.
When. Nothing is fetched at the opening view; the whole corpus as stars needs no images (0 requests, measured). Fetching starts slightly before a photograph becomes visible as one. It is limited to the visible area plus a margin, and the selection and its neighbours are requested at high priority.
How they appear. An early version fetched every thumbnail the moment the view opened. With 221 photographs at the time, the arriving images produced one 723 ms main-thread task, during which dragging froze. Two changes fixed it:
- nothing loads until the first zoom;
- newly loaded photographs are revealed at most six per frame, through the same frame loop.
Long tasks at load went from 881 ms to zero. I also tried moving decoding off the main thread with img.decode(). It measured worse (four long tasks, up to 386 ms of drag lag), so I reverted it.
What it still costs. Fetching runs ahead of drawing. The first zoom step from the full view requests 63 thumbnails; after one deep zoom 261 of 381 have been requested, about 16 MB, while 160–180 are drawn as photographs at a time. The real cost is not those references, though. My own image cache holds under 20 MB of encoded data. What grows is the browser's decoded-image cache: in Chromium's memory accounting it rose from 76 MB to about 440 MB across one zoom-in and a few pans, and the browser later reclaimed it to about 210 MB on its own. The size is driven by the thumbnails being larger than they need to be: they are 800 pixels, and sprites are never drawn above 112 CSS pixels. So instead of evicting images in JavaScript, which would not touch that cache, the change worth making is a smaller canvas-specific thumbnail. I've deferred it until the archive or a real-device profile calls for it.
Responsive and mobile
The stage is a square, capped at 788 pixels wide on desktop, with the side panel beside it. On tablets the panel moves below. On phones the square gives way to 70% of the viewport height, at least 480 pixels, because a square on a portrait phone wastes the height.
Nothing else is special-cased. The camera, zoom limits and sprite sizes are all relative to the fitted scale and the stage, so a phone and a desktop reach the same representation at the same relative zoom.
Touch needed real changes:
- Gestures. One finger pans, two fingers pinch, and a tap selects. A pinch keeps the point under your fingers fixed, so the same rule covers zooming and moving with two fingers. There is no hover, and the pointer-follow effect simply does not exist for touch, so nothing stays highlighted after a tap.
- Scroll interception. Browsers deliver React's wheel and touch-move listeners as passive, which means they cannot stop the page from scrolling or zooming behind the canvas. So those listeners are attached natively to the canvas and marked non-passive, together with
touch-action: none. Through a measured pinch the page did not move. - The cost. A drag that starts on the stage can never scroll the page, which is why the stage leaves 30% of the screen free.
On an emulated phone with the CPU slowed four times, the view opens with one long frame of about 300–360 ms. A CPU profile splits it between the browser's own rendering work and React mounting the canvas: the one-off geometry, the star sprites and the first draws. There is no single part worth optimising yet, so I've left it.
Accessibility
Spatial proximity cannot be made non-visual. What can be made accessible is the content the map leads to, and most of the route already existed. Once a photograph is selected, the side panel offers its actions:
- the full-screen view, which steps through the photograph's visual neighbourhood with Previous and Next;
- a visual journey, which walks from neighbour to neighbour and moves the map along.
All of these are ordinary buttons. What was missing was a way to make the first selection without a pointer.
So the stage is a single native button. Its name says what it does, "Select the photograph at the centre of the map", and its description says what the map shows:
Whole portfolio arranged by visual resemblance. Photographs are positioned by how visually similar they look to each other. Position is illustrative, not an exact measurement.
The keyboard path works like this:
- Focusing the stage with the keyboard highlights the photograph nearest the centre, the same point the zoom buttons zoom around.
- Enter or Space selects it. A polite live region announces the selection, and the button's name changes to "Open the selected photograph", so pressing it again opens the photo, just like a second click.
- Tab continues into the side panel, then the zoom controls.
- Escape closes the full-screen view and returns focus to the map.
What I did not build:
- No arrow-key navigation through 2D space, and no hidden list of all 381 photographs. That list would just repeat the gallery.
- The map's own idea, which photographs sit together, stays visual. The relationships are reachable through the neighbourhood views instead.
I validated this against the browser's accessibility tree, not yet with VoiceOver or NVDA.
Measurements
Production build, the real 381-photo database, headless Chromium at 1440×900 and device pixel ratio 2. Frame rates are not reported because a headless browser has no real display cadence.
- Open the view: 1 React render, 2 long frames (>50 ms; longest 100 ms). No thumbnails requested.
- Idle: 0 renders, 0 long frames. About 30 draws per second for the stars; none with reduced motion.
- Pan (40 pointer moves): 0 renders, 0 long frames.
- Wheel zoom (12 steps): 0 renders, 0 long frames.
- Hover across the map (61 moves): 23 renders, 1 long frame (67 ms). Renders happen only when the hovered photo changes.
- Select a photograph: 3 renders, 0 long frames. The neighbours and the side panel load.
- Resize the window: 0 renders, 0 long frames. Handled by a ResizeObserver.
Frame-loop callbacks, including the redraw, stayed at or below 1.7 ms at the 95th percentile on desktop. The layout response is 23 KB. The canvas code, loaded only when the view first opens, is 15 KB compressed.
What the audit found
Writing this article meant measuring the feature. That turned up real defects, and I fixed them before publishing:
- A frame loop that outlived the view. Leaving the map while thumbnails were still downloading let a late image restart the loop after the canvas was gone. Twenty seconds later it was still running on every frame. The loop is now disposed with the view; I measured again in a production build, and nothing keeps running.
- A cancelled pointer counted as a click. The browser's
pointercancelwas handled likepointerup, so an interrupted touch could select or open a photograph. Cancellation now only cleans up. - A pinch that also panned. The last finger's own pointer events moved the map while the pinch zoomed it. A photograph under a spreading pinch slid about 100 pixels sideways; it now stays under the fingers.
- An unstable value and a few stale comments. The journey breadcrumb was rebuilt on every render, and three comments described behaviour the code no longer had.
Each fix has a regression test, apart from the pinch, which I measured in the browser. What remains is described above: the decoded-image cost, the long frame on slow phones, and the limits of the keyboard path.
Trade-offs
- Canvas over elements. The DOM stays the same size however many photographs there are, and the zoom transitions are continuous. In exchange, the photographs have no built-in semantics, focus or hit areas. The stage gets one button, and the relationships are reached through the side panel instead.
- Canvas2D over WebGL. No shader pipeline, at the price of CPU rasterisation. At this size it was never the bottleneck I measured.
- The camera outside React. No renders during movement, but React cannot see the camera, so a zoom button cannot know it is at the limit.
- Offline layout. The map is stable and the client is cheap, but a new photograph only appears after a rebuild.
- No virtualization. One simple pass per frame, which works because 381 is small.
- Reusing an existing thumbnail. Nothing new to generate, store or cache, at the price of downloading more pixels than the canvas ever draws.
What I would change at larger scale
The per-frame work is not where this breaks. The once-per-layout work is. Several of the derived-geometry steps compare every photograph with every other one. On synthetic, larger versions of the real layout:
- 381 photographs (today): geometry ≈20 ms once, representation math ≈0.05 ms per frame.
- 1,000 photographs: ≈60 ms once, ≈0.1 ms per frame.
- 4,000 photographs: ≈650 ms once, ≈0.4 ms per frame.
Somewhere past a thousand photographs I would:
- compute that geometry offline next to the layout, or in a worker;
- request a canvas-sized thumbnail (about 256 pixels) instead of the 800-pixel one;
- evict images that have been far from the viewport for a while;
- add a spatial index for hit testing only if profiling showed it mattered.
Independently of scale, I would test the keyboard path with real screen readers, and profile memory and the opening frame on a real mid-range phone rather than an emulated one.
Further reading
The Photography Search case study explains the image-similarity search that supplies each photograph's neighbours. The map itself is on the Photography page, under the "Visual Connections" lens.