- The budget
- Device measurements
- First pass: NEON
- Second pass: control rate, matrix, PGO
- Third pass: output stage, 16-bit tables, frame cache
- Open question: memory traffic
- CPU guard
- Wavetable import
- Considered and left out
MPC renders 128-frame blocks: 2902 µs per block, per plugin instance. A plugin passes at p99 ≤ 15% of the block (warns up to 35%). PolyForce's ceiling, 8 voices × 8 unison with both oscillators, has to stay within 15%.
Every optimisation is checked two ways:
- On the device with
make bench-device(see Building). - Between device runs, as ARM instructions per 128-frame block: the same
-O3 -mcpubuild, counted underqemu-arm. A proxy for the device, not a substitute.
The output of every optimised path is compared with the plain scalar engine: within 2e-6 relative RMS on x86, 1.8e-4 on ARM (reciprocal estimates and fused multiply-adds), over 132 test scenes.
make bench-device, p99 in % of the 2902 µs block. Patch: both oscillators, filter 1 LP24 + drive
→ filter 2 LP12, held chords; bench pinned to core 1 with MPC running.
| Date | Build | 8 v × 1 | 8 v × 4 | 8 v × 8 | 16 v × 16 | 256-frame import |
|---|---|---|---|---|---|---|
| 2026-10-04 | Phase 0 (16-voice engine) | 8.1 | 11.5 | 15.5 | 45.1 | 582 ms, 22 MB |
| 2026-10-04 | v0.0.2 (8 × 8 cap) | 8.2 | 11.4 | 15.2 | — | (same code) |
| 2026-10-05 | v0.0.3 release build (CI: GCC 11, glibc 2.31, profile-guided; all three passes) | 5.3 | 7.7 | 11.1 | — | — |
By voice count:
| Voices | v0.0.2 × 1 | × 4 | × 8 | v0.0.3 × 1 | × 4 | × 8 |
|---|---|---|---|---|---|---|
| 1 | 1.7 | 1.9 | 2.3 | |||
| 2 | 2.4 | 3.0 | 3.7 | |||
| 4 | 4.3 | 5.8 | 7.8 | 2.9 | 4.2 | 5.9 |
| 8 | 8.2 | 11.4 | 15.2 | 5.3 | 7.7 | 11.1 |
v0.0.3 with a busy matrix (8 slots, both LFOs): 6.2 at 8 × 1, 11.2 at 8 × 8; the worst single block
11.9%. Every case passes. The release .so loads (VSTPluginMain, Classic built) in 42 ms. The
three passes below came between the two builds: at 8 × 8 the block costs 27% less than v0.0.2's.
Other Phase 0 numbers: VSTPluginMain (then building the four built-in tables) took 163 ms on the
Force; it now builds only Classic and the classic oscillator shapes, the other built-ins on first
use (6–73 ms each on x86).
An x86 sweep of 379 commercial sample tables: all load and play, 334 ms average import, 468 ms worst.
At v0.0.2 the filters were about 55% of a voice's cost, so they came first.
The engine renders each control chunk in passes over the sounding voices: control (matrix, mod
envelope, pitch), sources (oscillators, subs, noise into per-voice buses), filter 1, filter 2,
output. The buses hold the voices side by side per sample, so the state-variable filters run
four voices per NEON vector (dsp/simd.h: GCC vector types, NEON on the Force, SSE on x86, so
the tests run the same code). Drive is a pass of its own.
Also in this pass:
- the wavetable oscillators and subs read four samples per step with 64-bit pair loads and an unzip
- pitch, cutoff and pan use short polynomials instead of libm
- the envelopes run without a per-sample switch
- the plugin only rebuilds the patch and rescans the parameters for MPC when something changed
| ARM instructions per block | Before | After |
|---|---|---|
| 1 voice | 91.5 k | 56.8 k (−38%) |
| 8 voices × 1 | 439 k | 232 k (−47%) |
| 8 voices × 8 unison, busy matrix | 999 k | 567 k (−43%) |
-
32-sample control rate with glides. The matrix, the envelopes' control outputs, filter coefficients and wave position are computed every 32 samples instead of 16, and every value glides across the chunk it applies to instead of stepping:
- oscillator level × pan, morph, sub, noise and the voice gain: per sample
- the state-variable filters' g, k and drive: every 16 samples, coefficients rebuilt per step so the filter stays stable
- comb delay and feedback: per sample; vowel formants: every 16 samples
A 40 Hz LFO on level, volume or pan moves the output no faster than the waveform itself (a check in
test/m5_test.cppthat the old engine fails). Time constants stay referenced to 16 samples, so envelopes, smoothing and drift keep their timing. -
Cheaper matrix and LFOs: slots resolved once per patch, shared sources once per chunk, a polynomial LFO sine.
-
Profile-guided build (
make arm-pluginwhenqemu-armis installed): an instrumented copy plays 132 patches under qemu (tools/pgo_train.cpp, about 10 s), then the.sois compiled with that profile. Functions the trainer never ran are optimised as usual; inside the ones it ran, paths it never took (other LFO shapes, glide, mono) count as cold, so the trainer covers the common patches broadly.PGO=0builds without;make test-arm-pgoruns the suite against the shipped objects.
| ARM instructions per block | NEON pass | + control rate, matrix | + PGO |
|---|---|---|---|
| 1 voice | 56.8 k | 52.1 k (−8%) | 47.7 k (−16%) |
| 8 voices × 1 | 232 k | 193 k (−17%) | 178 k (−23%) |
| 8 voices × 8 unison, busy matrix | 567 k | 518 k (−9%) | 505 k (−11%) |
- Output stage across voices. The last pass (amp envelope, steal fade, gliding gains, the sum
into the output) runs four voices per vector like the filters, and four samples' sums are
reduced into one vector. Up to four voices the envelope's closed form
c1 + c2 · q^(i+1)is worked out inside the sum; with more voices, or in a chunk with a steal fade or a stage change, it goes through a table (eight lanes of everything would not fit NEON's registers). - 16-bit wavetables. Samples are
int16with one float scale per frame (the largest value of the frame over all its levels maps to 32767): a 256-frame table takes 4.5 MB instead of 9.0 MB, so the shared 96 MB table cache holds twice as many, and the built-ins shrink the same way. The rounding sits about 96 dB under each frame's own peak, so a quiet frame keeps its resolution. The importer quantises frame by frame (no full-size float copy while loading). While the position moves, the oscillator reads the 16-bit frames directly: one 32-bit load per pair of neighbouring samples, unzip, widen, convert, about as many instructions per sample as the float tables took. - Frame cache for positions that hold still. A position that is the same as in the previous chunk plays a float copy of its frame at that mip level, two frames premixed at the morph: one table read per sample instead of two, nothing to convert. 24 slots per instance (about 200 KB), shared by its voices, least recently used goes first, each voice remembering its last slot; at most two full-size frames are (re)filled per 32-sample chunk, past that the oscillator reads the table directly. Sub oscillators use it too.
- Comb filters no longer clear their delay lines (64 KB per voice) when a note starts: a read further back than the note has written counts as silence.
- The plugin copies the parameters for the engine only when one was written since the last block.
ARM instructions per block, counted with the same patch before and after (both oscillators, F1 LP24 with drive into F2 LP12, notes held; "moving": envelope 2 sweeps the position, "still": it stays put):
| ARM instructions per block | Before | After | PGO before | PGO after |
|---|---|---|---|---|
| 1 voice, moving | 56.3 k | 49.3 k (−12%) | 52.8 k | 45.6 k (−14%) |
| 1 voice, still | 57.0 k | 47.2 k (−17%) | 52.8 k | 43.2 k (−18%) |
| 8 voices × 1, moving | 198.5 k | 183.3 k (−8%) | 184.8 k | 169.6 k (−8%) |
| 8 voices × 1, still | 198.7 k | 164.7 k (−17%) | 183.2 k | 150.0 k (−18%) |
| 8 × 8 unison, moving | 496.2 k | 477.3 k (−4%) | 483.2 k | 476.6 k (−1%) |
| 8 × 8 unison, still | 497.3 k | 349.2 k (−30%) | 481.2 k | 332.7 k (−31%) |
| 8 × 8 unison, busy matrix | 528.0 k | 440.8 k (−17%) | 509.3 k | 434.1 k (−15%) |
(In the busy patch one oscillator's modulated position rests at the end of its table, so it plays from the cache.) qemu counts no memory stalls: what half the table memory does for cache misses only the device bench shows.
The 16-bit samples are the one deliberate change in the sound: over the 132 test scenes the output
is within about 1e-5 relative RMS of the float tables, 1.2e-4 in the quietest (LP24 into HP24 at
3 kHz, little left but the rounding). The frame cache and the direct reads agree within 2.3e-6;
the output stage matches the old one within 4.7e-7 (summation order). test/m3_test.cpp plays
still, gliding and over-full cases with the frame cache on and off (Synth::setFrameCache).
At 8 × 8 the oscillators are still about 70% of the block when the position moves. A 256-frame
table is 4.5 MB now; the Force's L2 cache is about 1 MB. A position that holds still reads an 8 KB
float copy whatever the table's size; a moving one reads the table itself. Whether that waits on
memory is what the stage bench answers on the device: make bench-device WAVETABLES=<folder> runs a profiling build (polyforce_stages.so) that
reports each pass's time per block, then plays the same voices with the positions swept on the
built-in Classic table and on a 256-frame table from that folder.
- If the sources pass still grows clearly with the large table, prefetching the frames a moving position is heading for is the next thing to try.
- If not, memory is not what limits the oscillators.
The profiling build is a plain (not profile-guided) build, so its times read a little higher than
the shipped .so's.
The plugin adds up the engine's own CPU time (not MPC's callbacks) against the real-time budget,
over windows of at least one 128-frame block, so process() sub-blocks and small host blocks are
judged like MPC's.
| Window load | Action |
|---|---|
| Two windows in a row over 40% | Fade out the quietest voice that is only ringing out |
| One window over 65% | Fade out two at once |
| A single window between 40% and 65% (a patch rebuild, a burst of note-ons) | Nothing |
Fades take 3 ms, like a steal. Held and pedal-sustained notes are never touched. The status line
shows GUARD n while it acts; PF_CPU_GUARD=0 turns it off (the test suite does).
A 256-frame table first took 582 ms and 22 MB on the Force. Two changes (Milestone 1):
- Shorter high mip levels: level k stores
clamp(8 · (1024 >> k), 256, 2048)samples instead of 2048 — 2.4× less memory, the audible levels 0–2 unchanged. - Two frames per FFT: two real frames share one complex FFT, halving the FFT count.
Since the third pass the samples are 16-bit with a scale per frame, quantised as each frame pair
is built. Now about 120 ms and 4.5 MB on x86 (-O2); make bench-device times it on the device.
Loading always happens off the audio thread.
- Float copies for moving positions: the frame cache only pays off for a position that stays put; a moving one would need a new copy every chunk, and float tables would double the memory.
- Voices on a second core: MPC already runs instances on its audio workers in parallel, and handing voices to another thread inside a 2.9 ms block risks dropouts that can't be tested without the device.
- More than 8 voices or 8 unison: 16 × 16 measured 45% of a block before the NEON pass; 8 × 8 stays the ceiling until the device bench of the current build says otherwise.