Contents
Cargo makes a crate wait until each of its dependencies has been fully checked, function bodies included, before it starts. But a dependent compiles against a dependency’s interface, the signatures, types and trait impls in its .rmeta file. It does not need to know that the bodies type-check. A patch series published this week, headstart, removes that wait for rustc and cargo, and the way it does so says a lot about what a Rust dependency edge actually carries.
This is an experiment, not a feature. The author’s own assessment says it is ready for a design discussion and a draft PR, and not ready to merge. Every number below is the author’s measurement, not an independent one.
The mechanism
Today rustc writes a crate’s metadata after every function body has been type-checked and borrow-checked. Headstart adds a new query, analysis_interfaces, that runs only the item-level checks (well-formedness, coherence). The driver then writes a second file, libfoo-hash.early-rmeta, and announces it with an early-metadata artifact notification. Only afterwards does analysis check the bodies. Cargo, behind a -Zheadstart flag, starts dependents on that notification.
For cargo check that is the whole story, since check never generates code. For cargo build it is not, because code generation can inline and instantiate generic functions from the dependency, which needs their MIR. So a dependent does all of its analysis against early metadata, then pauses before code generation until the dependency’s full metadata exists, and hands its job slot back while it waits.
The early file is not simply “the interface”. Some of what dependents read comes from bodies, so the encoder forces those first:
- the hidden type behind an
impl Traitin a signature, which comes from type-checking the defining function; - the MIR of consts,
const fns and array lengths; - statics, evaluated before the write;
async fnandasyncblock state machines, because dependents read them to decide whether a future isSend.
Everything else is left out, including optimized MIR for ordinary functions, exported symbols and cross-crate inlinability.
Why the swap is the dangerous part
A dependent that compiled against early metadata must, before code generation, replace it with the full file. The design doc is specific about what that involves: a new crate-metadata entry is created while the early one stays alive, the type cache entries for that crate are dropped because positions in the blob mean something different, and source files are imported again from the full blob.
The invariant that matters is that nothing before the swap may ask an early-metadata crate a question whose answer changes afterwards. The readiness doc gives a concrete hole. The MIR inliner, at optimization level 1 or higher, asks each dependency whether it has MIR for a callee. A crate still loaded from early metadata answers no, and the query system caches that answer past the swap. A dependent then fails to find MIR it needs. The same document says the invariant is currently guarded by a hand-kept list of queries, checked only under -Zearly-metadata-verify, and that an upstream version should fail loudly instead.
The failure model is simpler. Both files carry the same crate hash, computed from the early metadata bytes plus a hash of the HIR, so any body change changes it, which keeps incremental compilation sound. If a dependency fails after writing early metadata, waiters get its lock without the file and stop with an error. Cargo holds the diagnostics of any unit that started early until all its dependencies finish cleanly, and drops them otherwise. The author’s claim is that a failing build prints the same diagnostics and exit status as today, and the repo ships scripts/check-errors.sh to test it. The cost is thrown-away work downstream.
What was tried before
Nicholas Nethercote tried moving metadata writing before type-checking in 2019 (rust-lang/rust#64112). With a single metadata file, the early write still had to carry the MIR needed for code generation, which forced most of the analysis anyway. He saw at most 1.07x and dropped it, citing a small win and slightly delayed error messages, per the headstart design doc. Splitting the file into an analysis-only early file and a full file for code generation is the difference, and it also covers cargo check, which cargo did not pipeline at all.
What the numbers say, and what they depend on
On 16 cores, the author reports clean builds of 13 real projects up to 54% faster for cargo check and up to 42% for cargo build, none slower. rust-analyzer is the headline case (54% check, 42% build). On 4 cores the same project is 24% faster for check and 13 to 15% for build, and wide workspaces (typst, lldap, bevy) come out even. The README ties the gain directly to cores the build would otherwise leave idle.
That explains the shape of the results. A workspace whose crates form a long dependency chain leaves most cores idle while each crate waits for the previous one’s bodies. The README shows this for a clean cargo build of codex-rs on 16 cores, reported 37% faster. A wide graph already has plenty of runnable work, so there is nothing to fill. On a saturated machine the author also measures 3 to 7% more rustc CPU time with headstart on, from the early write and extra compilations competing for cache and memory.
The open items in the readiness doc are the ones that decide whether this ships:
- Only cargo can drive it. The loader assumes the rlib sits next to the
.rmetaand that the dependency compiles on the same machine, which Bazel and Buck2 sandboxed or remote rmeta actions do not guarantee. - A resumed compilation does not wait for a jobserver token, because waiting starved the critical path. Inside
make -jthat breaks the jobserver contract. - Cargo does not cap paused compilers, so peak memory can rise. The only peak-memory figures (from minus 4% to plus 39%) come from the first, check-only version.
- Nothing has been run on Windows, or on macOS since the locks replaced polling.
What to do with it
You should not run a patched toolchain in CI on the strength of this. The practical takeaways are smaller. If your build is slow on a many-core machine and cargo build --timings shows a staircase of workspace crates with idle cores underneath, that is the shape this targets, and the staircase is the thing to measure first. If your workspace is wide, or your CI runners have four cores, expect little. And if you maintain a build tool that parses --json=artifacts, note that the series adds three notification kinds, early-metadata, wait-metadata and resume.
Sources
- PowderworksCode/headstart README, checked October 5, 2026.
- Headstart design notes: early metadata contents, swap, locks, risks and prior art.
- Headstart readiness assessment: open issues and 4-core numbers.
- rust-lang/rust#64112: the 2019 attempt, as summarized in the design notes.