A clean-room, MIT-licensed Rust engine for PEEC (partial-element equivalent circuit) extraction of resistance and inductance of 3-D conductor geometries: spirals, busbars, bond wires, on-chip interconnect. FastHenry's problem class, none of its code, and built for today's hardware (portable SIMD, all cores).
FastHenry (MIT RLE, 1994) is still the reference tool for frequency-dependent
R/L extraction, and it is unmaintained, single-threaded C under a license that
permits only internal, noncommercial use and forbids redistribution. Every
open-source EDA flow that needs partial inductances either shells out to it in
a legal grey zone or does without. fasterhenry is the replacement: the
published method, implemented from the papers, released under MIT, fast enough
to sit inside a design loop.
What exists today: a filament model with uniform, surface-graded and
skin-depth-graded subdivision; partial self/mutual inductance kernels;
ground planes with holes and graded contact regions; coupling truncation;
mesh assembly with a dense complex solve over a frequency sweep; and a
matrix-free precorrected-FFT operator with a GMRES solve for large problems,
which the CLI switches to automatically above 10 000 filaments.
The physics is cross-checked against independent references — PyPEEC and
the Greenhouse closed forms (docs/validation.md) and
a head-to-head with FastHenry itself (docs/benchmarks.md);
the precorrected-FFT path is tested against the dense solve. The CLI reads FastHenry
.inp decks and writes JSON, a MAT v4 Zc.mat, or a SPICE subcircuit.
Ruehli's PEEC formulation; the FastHenry approach of Kamon, Tsuk and White
(FASTHENRY: a multipole-accelerated 3-D inductance extraction program, IEEE
Trans. MTT 42(9), 1994); Grover/Rosa closed forms for parallel filament
partial inductances; numerical quadrature for arbitrary orientation. See
docs/ as the design lands.
This project must never contain code derived from MIT's FastHenry or FastCap,
in any branch or mirror (ediloren/FastHenry2, wrcad/xictools, …). Their
notice grants "internal, noncommercial" use only and prohibits distribution
of copies or derivatives — a port could never be released. Contributors work
from the papers and from this repository's own code. Reading the FastHenry
.inp deck format is fine (a file format is not code); reading FastHenry
source while contributing here is not. See CONTRIBUTING.md.
Rust (stable), nalgebra + simba/wide for portable SIMD, rayon for
parallel matrix fill, num-complex, and rustfft (pure Rust, MIT OR
Apache-2.0) for the precorrected-FFT operator. No BLAS/LAPACK, no C
dependencies: one static binary on arm64 and x86-64.
The committed, regenerated numbers live in
docs/benchmarks.md:
a cargo bench (criterion) sweep of assembly- and solve-stage throughput
against filament count (fasterhenry/benches/assembly.rs,
fasterhenry/benches/solve.rs; kernel-level throughput is
fasterhenry/benches/kernels.rs), regenerated on demand by
.github/workflows/bench.yml on a pinned ubuntu-latest runner —
parsed straight from criterion's own JSON output, not hand-typed off
whichever machine ran it last. The perf_smoke CI tests back the same
posture with hard budgets, on every run: a 2 000-filament dense sweep
within 10 s on that same pinned runner class, and a 4 760-filament
matrix-free sweep within 120 s and under half the dense path's working
set (fasterhenry/tests/perf_smoke.rs).
The dense path is parallel across all cores and SIMD-batched in the
kernels, so it is dramatically faster than a single-threaded 1994-era
solver on modern hardware — for problems that fit the dense regime,
because its working set is 8n² bytes of partial inductances before the
factorization is counted. Beyond that regime the matrix-free path takes
over: a precorrected-FFT operator (#42) under GMRES (#43), linear in
memory and near-linear in time.
Measured on one machine (AWS 8 vCPU, one thread, 2026-09-25) over
assembly plus a one-frequency sweep of a square-meander fixture, dense
against pFFT + GMRES — full table, method and caveats in
docs/benchmarks.md. The wall
times predate #58 and #60 (which only shorten them); the working-set column,
and the memory-based threshold below, are unaffected:
| filaments | dense | pFFT + GMRES | dense working set |
|---|---|---|---|
| 2 400 | 4.2 s | 4.7 s | 0.15 GB |
| 9 800 | 156 s | 19.4 s | 2.5 GB |
| 29 928 | not attempted | 61 s | 23 GB |
| 99 224 | not attempted | 207 s | 256 GB |
GMRES converges in 5 iterations at every size, and Z(ω) is reproducible
run to run and thread-count to thread-count.
The wall-clock crossover is near 3 000 filaments, but the default
switches at fasterhenry::DENSE_PATH_MAX_FILAMENTS = 10 000: the dense
path is exact where the matrix-free one approximates the far field
(< 1e-4 on Z), so the handover is placed where dense stops being
affordable on any geometry rather than where it stops being fastest.
--solver dense / --solver iterative (library: SolverChoice) force
either path at any size, and the CLI says on stderr when a run leaves the
dense path.
Measured head-to-head against the original FastHenry (operator-run,
one machine, self-authored decks; method, hardware and caveats in
docs/benchmarks.md):
fasterhenry's dense path is faster on wall clock at every size up to
~20 000 filaments (3× at small sizes, ~1.2× at 20 k, where FastHenry's
multipole stays 14× ahead per thread — that dense-path edge is
parallelism; the algorithmic answer is the pFFT + GMRES path above).
On shared segment fixtures the two engines
agree to better than 0.1 % on the extracted impedance — the
cross-validation behind the "replacement" claim.
| crate | what |
|---|---|
fasterhenry |
the library: geometry, filaments, kernels, assembly, solve |
fasterhenry-cli |
fasterhenry binary: .inp/JSON in, JSON out |
This repository is developed with Loom orchestration. To drive an approved issue through Curator → Builder → Judge → Doctor → Merge:
cd fasterhenry
/loom:sweep <issue>Merges into main are gated on CI: branch protection requires every job in
.github/workflows/ci.yml (clean-room, cargo-deny, cargo-about, the three
Rust legs, and Package) to pass. If a PR's checks went green before main
moved, re-run them rather than merging through; merge-pr.sh refuses
required checks older than the base tip. Adding or renaming a CI job means
updating the required-checks list on main to match.
MIT — see LICENSE.
Dependencies are restricted to permissive licenses (MIT, Apache-2.0, Zlib,
Unlicense, Unicode-3.0) — nothing copyleft, and nothing that imposes a
source-disclosure obligation on the distributed fasterhenry binary, which
links every transitive dependency. Dependencies must also come from crates.io —
no git revisions and no private registries, so every input to a release build is
an immutable published version. Both policies are enforced in CI by cargo deny check licenses sources; the allowlist, the allowed registry and the reason each
entry is there are in deny.toml. Run it locally the same way:
cargo deny check licenses sourcesThe Apache-2.0-only dependencies (nalgebra, simba, approx,
nalgebra-macros) add an attribution term the MIT notice above does not
discharge: Apache-2.0 §4(a) requires that a recipient of a distributed binary
receive a copy of the Apache-2.0 license text. That is discharged by
THIRD_PARTY_LICENSES.md — a generated bundle
carrying the full text of every license in the normal + build dependency
closure, which is what a binary actually links. It is committed, not
hand-maintained: CI regenerates it and fails the PR if the committed copy is
stale, if the two allowlists (deny.toml, about.toml) have drifted apart, or
if a workflow ships a built artifact without shipping the bundle with it. Any
binary distribution of fasterhenry — a release asset, a container image, a
distro package — must include that file. Regenerate it with:
./tools/third-party-licenses.sh # rewrite the bundle
./tools/third-party-licenses.sh --check # what CI runsNone of the four Apache-2.0-only crates carries a NOTICE file, so §4(d) is not
engaged. Installing from source (cargo install fasterhenry-cli, the only thing
release.yml publishes today) is not a binary distribution by this project —
cargo fetches each dependency, with its own license file, from crates.io — so
§4(a) attaches to the bundle above only once a prebuilt artifact is shipped.