04 · Road to 60

The road to sixty

F-Zero X runs its logic at 60 Hz and the port never lowered that. Every frame the CPU walks the game's display list, translates it, and feeds the PICA. The road to 60 was the road from a 35–58 ms frame build to under 16.7 ms, measured first in the emulator and then, from 08-21, on the console. This page is that road, lever by lever.

51.359.6
median fps, hardware, round 1 → round 5 of the LOCKED-60 campaign
39.251.4
tenth percentile (the crowd floor)
9%53%
of one-second beats at the 60 cap
12–20
fps in the first emulator race, 08-14
35–55
fps on the first hardware run, 08-21 (no stereo)
2
levers built to completion and shelved on hardware evidence

Where the frame went: the CPU-bound verdict

The measurement pass of 08-20 settled the strategy. GPU drawing time was flat at 0.4 ms per frame and the loop never once waited on the PICA; CPU build time was 26 to 58 ms. Fill rate, model decimation and texture compression were struck from the plan as irrelevant to frame time (the whole race used 8 KB of texture bytes and 957 vertices per frame). The bottleneck was the software display-list interpreter and the bridge that feeds it. Every lever from then on attacked CPU time, and every profile was reported as a set of named buckets:

BucketWhat it is
brthe bridge pre-pass: converting the game's raw display list into the runtime's command form, resolving addresses
dspinterpreter dispatch: walking the converted list and executing every command
vtxvertex transform: projecting N64 vertices to clip space on the CPU
triper-triangle work: tile extents, shader lookup, clip parameters, packed VBO emission
imptexture import: looking up or uploading textures for a draw
drwissuing the draw to citro3d, including the stereo target switch

The emulator era: 08-14 → 08-21

Before the console arrived every number was an Azahar proxy, known to be pessimistic on menus and unrepresentative in absolute terms. The shape of the work still held.

DateLeverMeasured (emulator)
08-14First race12–20 fps race, 15–20 fps menus, debug logging on
08-20Frame pacer double-throttle fix, HLE audio producer to core 2, malloc histogrammeasurement pass: CPU 26–58 ms vs GPU 0.4 ms
08-21Loadblock discovery. Per-opcode profiler named G_LOADBLOCK at 22.34 ms per frame: each 4 KB load did 512 shared_ptr copies (a thousand atomic RMWs), a 4 KB memcpy and a live getenvF3 22.34 → 2.21 ms; dispatch 27.6 → 8.9 ms; menu wall 35–38 → 16.7–19.9 ms. User: "blazing fast"
08-21Flat dispatch table replacing libultraship's layered handler tablesneutral on time, enabled per-opcode profiling
08-21 nightS7 interpreter memo (raw dispatch, geo-diag gate, tri-state memo), async frame-mirror copy, vblank-stall skip, upload thrash 155 → 0.5 per frame, resolutions 93 → 0menu 15 → 19.9 fps; race build 26–58 → 24–46 ms; title wall 34 → 21–27 ms
08-21Texrect state fast path (batch same-state runs)E4 texrect bucket 6–9.7 ms attacked

Hardware, before the campaign: 08-21 → 08-28

DateBuildHardware result (user + log)
08-21First .cia and .3dsx, rspolish35–55 fps; median 48–50, p95 58, max 59 over 8 races; crowd dips 25–35. "wow it runs amazing"
08-27Fleet drop: traffic grind (bridge region cache, binary-search asset lookup, pipesync no-op), sky wedge, CCMUX 11, stereo armed, shadow"performance is much better! 3D works great, doesnt seem to affect performance much at all"; one hard crash (filelog OOM, fixed same day)
08-28Touch menu, full-bleed display, rival detail option, triloop packed VBO, leak fix (upstream LUS bug)median 51, p95 60, max 60; MINIMAL rival detail = +8–9 fps in crowds, floor ~40 (was 25–35)
08-28GPU vertex transform moonshotimperceptible on hardware; venue floor loses texture mapping (24-bit uniform precision). Shelved.

The LOCKED-60 campaign: 09-01 → 09-03

Five hardware rounds in three days, each a staged branch with a test plan, each measured from the console's own once-per-second fps beats in the log. Stereo was accidentally off in rounds 1 and 2 and on from round 3, which makes the later gains larger than they look.

medianp10
0204060cap 60Round 1|Mainline before the campai · median fps: 51.351.3Round 1|Mainline before the campai · p10 fps: 39.239.2Round 1Mainline before the campaiRound 2|Bridge memos (brfast) · median fps: 57.257.2Round 2|Bridge memos (brfast) · p10 fps: 46.646.6Round 2Bridge memos (brfast)Round 3|HUD atlas · triangle memo · median fps: 56.056.0Round 3|HUD atlas · triangle memo · p10 fps: 48.048.0Round 3HUD atlas · triangle memo Round 4|Render thread on core 2 (p · median fps: 57.657.6Round 4|Render thread on core 2 (p · p10 fps: 44.644.6Round 4Render thread on core 2 (pRound 5|Ahead mode · bridge on mai · median fps: 59.659.6Round 5|Ahead mode · bridge on mai · p10 fps: 51.451.4Round 5Ahead mode · bridge on maifps (1-second beats)
Median and floor, round by round. Round 1 is mainline before the campaign (754 beats over 63 minutes). Round 5 is a full Grand Prix with the render thread in ahead mode, the bridge on the main core and automatic rival detail.
02550Round 1|Mainline before the campai · % of beats at 60: 99Round 1Mainline before the campaiRound 2|Bridge memos (brfast) · % of beats at 60: 2929Round 2Bridge memos (brfast)Round 3|HUD atlas · triangle memo · % of beats at 60: 3232Round 3HUD atlas · triangle memo Round 4|Render thread on core 2 (p · % of beats at 60: 1414Round 4Render thread on core 2 (pRound 5|Ahead mode · bridge on mai · % of beats at 60: 5353Round 5Ahead mode · bridge on mai% of beats ≥ 59.5 fps
Time spent at the cap. The fraction of one-second windows pinned at 60 went from one in eleven to more than half.
RoundWhat it carriedConditionsBeatsMedianp10At cap
Round 1Mainline before the campaignstereo off · trace off754 beats · 63 min51.339.29%
Round 2Bridge memos (brfast)stereo off · trace off78 beats · 6.5 min57.246.629%
Round 3HUD atlas · triangle memo · TMEM · audiostereo on · trace off237 beats · 20 min56.048.032%
Round 4Render thread on core 2 (pipe)stereo on · trace off58 beats · 5 min57.644.614%
Round 5Ahead mode · bridge on main · auto rival detailstereo on · trace off · full GP229 beats · 19 min59.651.453%
Round 1, beforeRound 5, after
0153046610–30 · Round 1: 550–30 · Round 5: 110–3030–35 · Round 1: 0030–35 · Round 5: 0030–3535–40 · Round 1: 6635–40 · Round 5: 0035–4040–45 · Round 1: 101040–45 · Round 5: 1140–4545–50 · Round 1: 191945–50 · Round 5: 3345–5050–55 · Round 1: 303050–55 · Round 5: 111150–5555–58 · Round 1: 151555–58 · Round 5: 111155–5858–59.5 · Round 1: 7758–59.5 · Round 5: 181858–59.5≥59.5 · Round 1: 99≥59.5 · Round 5: 5353≥59.5% of beats
The distribution moved, not just the median. Round 1 has a long tail below 45 fps (crowd starts, transitions). Round 5 stacks its mass in the top bin.

Where the milliseconds came from

Hardware profiles with the trace on (which itself inflates the numbers by a few milliseconds) on the same course, crowd windows (30 machines in view) and steady windows. Buckets in milliseconds per frame.

bridgedispatchvertextriangleimportdraw
0.05.711.317.022.6Campaign levers off|crowd · bridge: 3.0Campaign levers off|crowd · dispatch: 8.3Campaign levers off|crowd · vertex: 1.1Campaign levers off|crowd · triangle: 5.6Campaign levers off|crowd · import: 0.8Campaign levers off|crowd · draw: 1.520.2Campaign levers offcrowdCampaign levers on|crowd · bridge: 3.1Campaign levers on|crowd · dispatch: 7.8Campaign levers on|crowd · vertex: 1.2Campaign levers on|crowd · triangle: 4.6Campaign levers on|crowd · import: 0.8Campaign levers on|crowd · draw: 1.218.8Campaign levers oncrowdRender thread (pipe)|crowd · bridge: 3.1Render thread (pipe)|crowd · dispatch: 4.9Render thread (pipe)|crowd · vertex: 1.0Render thread (pipe)|crowd · triangle: 4.2Render thread (pipe)|crowd · import: 0.7Render thread (pipe)|crowd · draw: 1.215.1Render thread (pipe)crowdCampaign levers off|steady · bridge: 2.2Campaign levers off|steady · dispatch: 7.6Campaign levers off|steady · vertex: 0.7Campaign levers off|steady · triangle: 4.0Campaign levers off|steady · import: 0.8Campaign levers off|steady · draw: 1.216.4Campaign levers offsteadyCampaign levers on|steady · bridge: 2.2Campaign levers on|steady · dispatch: 7.3Campaign levers on|steady · vertex: 0.7Campaign levers on|steady · triangle: 3.2Campaign levers on|steady · import: 0.7Campaign levers on|steady · draw: 0.814.9Campaign levers onsteadyRender thread (pipe)|steady · bridge: 2.7Render thread (pipe)|steady · dispatch: 3.9Render thread (pipe)|steady · vertex: 0.7Render thread (pipe)|steady · triangle: 3.1Render thread (pipe)|steady · import: 0.7Render thread (pipe)|steady · draw: 0.912.0Render thread (pipe)steadyms per frame (trace on)
Per-bucket frame time on hardware. The campaign levers shaved the triangle and dispatch buckets; the render thread took dispatch from 7.8 to 4.9 ms in crowds by overlapping it with game logic on core 2. The 16.7 ms budget line is the whole bar; anything under it is a 60 fps frame.
ProfileWindowbrdspvtxtriimpdrwTotal
Campaign levers off control, same coursecrowd2.968.261.125.580.781.4820.2
Campaign levers off control, same coursesteady2.27.630.663.980.761.1816.4
Campaign levers on atlas · tri memo · TMEM · brfastcrowd3.077.841.184.640.781.2418.8
Campaign levers on atlas · tri memo · TMEM · brfaststeady2.177.350.653.220.720.7914.9
Render thread (pipe) core 2 renders, main waitscrowd3.124.891.014.210.681.1615.1
Render thread (pipe) core 2 renders, main waitssteady2.723.870.73.120.660.9512.0

Every lever, and its verdict

LeverIdeaMeasuredVerdict
Loadblock span storeOne record per 4 KB load instead of 512 shared_ptr copies; same-content skip; getenv latchF3 22.3 → 2.2 ms, dispatch 27.6 → 8.9 ms (emu)mainline
Flat dispatch + [prof]/[profop]Plain table dispatch; per-bucket and per-opcode timersneutral; enabled everything aftermainline
S7 interpreter memosRaw-pointer dispatch, gated geo diagnostics, tri-state memobuild 34 → 23.5 ms (emu); one 2× regression bisected out (dlcache)mainline
Async frame mirror, vblank-stall skipMenus stop waiting on a copy and on a late vblankmenu 15 → 19.9 fps; wVbl 10–16 ms → 0 (emu)mainline
Texrect fast pathBatch same-state rectangle runsE4 bucket attacked; superseded by the HUD atlasmainline
Traffic grindBridge region cache (no svcQueryMemory per probe), binary-search asset lookup, pipesync no-opbr 19.7 → 10.4 ms, wall 55.7 → 46.7 ms (emu)mainline
Prim/env value-change flushCorrectness first: flush on colour changefixes machine tint; small costmainline
Rival detail (NATIVE/REDUCED/MINIMAL)Bias the game's own LOD tiers for non-player machines beyond the 5 nearest+8–9 fps in crowds on hardwaremainline
Triloop packed VBOSingle-write PICA-layout emission from the interpreterkills the double vertex copymainline
Leak fix (uninitialised ColorCombinerKey::shader_id)Upstream LUS bug: combiner map missed per drawheap flat; hidden per-draw flush tax gonemainline
Double-height stereo targetOne 400×480 target, per-eye viewport, fewer PICA flushesuser observed 2D ≈ 3D on hardware; citro3d source says flush cost is GPU-side and the GPU idlesshelved
GPU vertex transform (moonshot)Unlit vertex transform on the PICA vertex shaderworks (220 k draws, zero fallback, −62% vtx emu); imperceptible on HW; 24-bit uniform precision breaks ±32000-unit floorsshelved
Bridge output cacheCache converted display lists across frames, epoch-validatedcensus: 87% of walked commands are host-built per frame; ceiling ~1.5 msNO-GO
brfast (bridge memos)Per-list facts memo, address-resolve memo, placeholder/raw-copy/range-class memosbr 11.5 → 4.9 ms (emu); HW br 6.43 → 3.73 in crowds; median 56.0 → 57.2, p10 42.0 → 47.2mainline
HUD texrect atlas (trectbatch)Pack HUD rectangles into atlas pages; batches stay open across viewscrowd draws 136 → 26 (E4 −2.6 ms emu)mainline
trifastFast packed triangle loop; found the S7 memo had never actually hit (flag bug)tri −15..−19% (emu); HW tri 5.58 → 4.64 ms crowdmainline
tmemfastTMEM load bookkeepingF3 −22..−35% (emu)mainline
audioprime / audioprime2Boot audio underrun: title font sample load blocked on a 10.7 MB inflatestore audio_table uncompressed in the o2r: preload 4.3 → 1.97 smainline
dspfast (batch-break fold)Fold prim/env changes to avoid batch splitscensus: texture switches split anyway; cannot paydefault off
Machine texture atlas (atlas3d)Atlas machine-part textures to cut switchesonly ~21 imports per frame split a batch; 0.1–0.5 ms best case; two thirds of machine textures wrap or mirrordefault off
Render thread, core 2 (mode 1, pipe)Interpreter and draws on core 2, forked at osSpTaskStartGoHW: stable, HOME clean; median 49 → 56, p10 42 → 48 (trace on); not yet overlapping (waitMain 10–15 ms)mainline
Render thread mode 2 (ahead)DP-done acknowledged when the game parks; 2-deep backpressure; true one-frame pipelineHW round 5: median 59.6, p10 51.4, 53% at capmainline
Bridge on main (bridgemain)Bridge pre-pass runs on core 0 while core 2 rendersrender-thread br 4 → 0 ms; brMain 3.6 ms; balanced coresmainline
Auto rival detail (dynlod)Raise the LOD tier when render > 15 ms, lower when < 12 ms; user setting is the floorshipped in round 5; thresholds untuned on HWmainline
Render-thread-owned texture cacheFixes the round-5 data abort: clears become requests drained on the render threadreceipt texcacheMainMut=0; awaiting HWmainline

What is left

Round 5 is the last measured state: 53% of beats at the cap and a median a hair under 60. The floor is still the crowd start. The remaining known levers, in the order the research ranked them: double command buffers on the render thread (CPU build and GPU are still serialised on core 2), a texture import-lookup memo (imp 0.7 ms, 83 imports per frame), tuning the auto-LOD thresholds against hardware, and, if ever wanted, GPU vertex transform phase 2 on lit machine geometry with a large-coordinate CPU fallback.