Premature de-optimization is the root of all evil
TL;DR — Popular languages make three choices for you: memory has no location, the machine has no shape, and the source is lowered to a low-level IR almost at once. Each choice throws away something the programmer knew. Together they make programs slow before anyone writes them, and a set of tools that few engineers master exists to win the speed back. Vx puts memory and topology in the program and keeps them there until code generation.
Knuth warned against optimizing before you know where the time goes. The more common mistake today is choosing, before you start, a language that cannot say where the time goes.
Three choices made before the first line
Memory is hidden. In C, C++, Rust, Go, Java and Python, memory is one flat
array of bytes. A float * does not say whether the floats are in the local socket's
DRAM, the other socket's DRAM, a GPU's HBM across PCIe, or a CXL pool across the rack. Those
differ in latency by orders of magnitude, and the type system treats them as one thing. The
programmer cannot write down “this is small and hot, keep it close”, and the compiler
cannot ask “does this even fit?”
Topology is hidden. A server has sockets, NUMA domains, cores that share some caches and not others, GPUs behind particular PCIe switches, and NVLink or InfiniBand between boxes. None of it appears in the language. A thread is a thread, and where it runs is decided by the OS scheduler, a runtime, or an environment variable in a launch script. Code cannot say “run next to that data”, so it runs wherever it lands and the data travels to meet it.
Lowering happens early. Clang turns C++ into LLVM IR in one walk over the syntax tree. Rust keeps MIR for borrow checking and then hands the same optimizer LLVM IR. Java and Python become bytecode. From then on the optimizer sees loads, stores, pointer arithmetic and branches, and spends much of its effort rebuilding what the source already said: loop analysis finds the loops again, scalar evolution finds the induction variables again, alias analysis guesses which pointers overlap. LLVM even has a pass, InferAddressSpaces, whose job is to work out which GPU memory a generic pointer points into, after the language dropped that fact from the type. When a guess fails, the optimizer has to assume the worst, and assuming the worst is slow.
All of them model a tape
Put the three together and every mainstream language targets the same machine: Turing's tape. One line of cells, one head, every cell the same distance away and the line never ending. The hardware is a tree of memories of very different sizes and speeds, joined by links of very different bandwidths, with compute scattered across the leaves. An earlier post walks through that tree for one H100 node.
A program written for the tape and run on the tree is slow from the start. That is the de-optimization in the title, and no programmer chose it. The language chose it for every program at once.
The tools that win it back
The information the language will not carry is recovered afterwards, by hand:
- Measuring:
perf,perf c2cfor false sharing, VTune, Nsight Systems and Nsight Compute, likwid, cachegrind. - Placing:
numactl,taskset,hwloc,libnuma,CUDA_VISIBLE_DEVICES, and first-touch page placement that depends on which thread happened to write a page first. - Arranging: hand-padded structs,
alignas(64), arrays of structs rewritten as structs of arrays, huge pages switched on by an administrator.
These tools are good, and few engineers use them well. The usual cycle is to ship, profile,
find half the time going to remote memory or PCIe copies, add numactl --membind to a
launch script, and hope nobody changes the machine. The knowledge ends up in scripts, wiki pages
and people's heads. The compiler never sees it, cannot check it, and the next refactor can break
it without a warning.
C++ allocators go part of the way
C++ is the one mainstream language that tries. std::allocator, and later
std::pmr memory resources, let a container ask a chosen allocator for memory. You can
write one that calls numa_alloc_onnode, or one that hands out pinned host memory. For
NUMA placement that helps. Beyond it, very little:
- The allocator decides where memory comes from, and returns a plain
T *. From then on nothing in the type says where the memory is, so nothing stops a thread on socket 1 from hammering pages on socket 0. std::pmrreaches its resource through a virtual call, so the compiler learns even less than before.- Nothing says where code runs, or what moving data between two places costs.
- Nothing checks capacity. A working set that is too large shows up as a failed
cudaMalloc, or as the OOM killer a few hours into a job.
Libraries go further, and stop at the same place. SYCL and Kokkos type the memory region inside a
device. Neither can make where a function runs part of the function's type, because C++ has
nowhere to put it. Kokkos shows the limit plainly: when
SpaceAccessibility<...>::accessible is a compile-time false, meaning
the compiler has already proved the access illegal, the library can only emit a run-time
Kokkos::abort. A language can do better, so Vx is one.
What Vx does with each of the three
Memory has a location, and it is part of the type
A Vx tensor's type names the memory it lives in, and moving it is an explicit
transfer that produces a value of a new type:
fn main() -> i32 {
let a = Tensor<f32, [4, 4]>::fill(2.0);
let b = Tensor<f32, [4, 4]>::fill(3.0);
let mut c = Tensor<f32, [4, 4], Memory::NPU_HBM>::uninit();
let a_npu = transfer(a, Memory::NPU_HBM);
let b_npu = transfer(b, Memory::NPU_HBM);
spawn on(Topology::NPU[0]) {
c = a_npu @ b_npu;
}
let c_host = transfer(c, Memory::CPU_DRAM);
print(c_host[0][0]);
return 0;
}
The allocator's answer is kept for the value's whole life, where C++ drops it at the first
T *. So every later use can be checked, and mistakes the tape model compiles are
refused on a laptop:
| Mistake | C++ / CUDA | Vx |
|---|---|---|
| Host reads device memory | segfault or cudaErrorIllegalAddress | E6003, naming the missing transfer and its cost |
| Working set larger than the memory | failed cudaMalloc, or OOM hours in | E6009 / E6010, against the declared capacity |
| Copy between memories with no link | run-time error or a slow fallback | E6002, no path exists |
The program prints 24. Replace transfer(a, Memory::NPU_HBM) with plain
a, and the compiler refuses it:
Error[E6003] at 10:9: 'a_npu' lives in CPU_DRAM but NPU[0] sees only [NPU_HBM];
insert an explicit transfer to NPU_HBM (cost 50 on the declared path)
The machine is an input, and code says where it runs
Vx reads the machine from a declaration file. This is the H100 description that ships in
fleet/, without its citations:
Memory HBM {
capacity: 80 GiB, bandwidth: 3.35 TB/s, managed: explicit, scope: device
}
Memory L2 {
within: Memory::HBM, capacity: 50 MiB, bandwidth: 12 TB/s, managed: cached, scope: device
}
Memory SMEM {
within: Memory::L2, capacity: 228 KiB, bandwidth: 128 B/cyc, clock: 1.98 GHz,
granule: 1 KiB, managed: explicit, scope: sm, replicas: 132
}
Topology Device {
arch: nvptx64,
memory: Memory::HBM,
visible: [Memory::HBM, Memory::L2, Memory::SMEM],
transfer Memory::CPU_DRAM -> Memory::HBM : 63 GB/s,
transfer Memory::HBM -> Memory::CPU_DRAM : 63 GB/s,
transfer Memory::HBM -> Memory::L2,
transfer Memory::L2 -> Memory::SMEM copy_engine
}
That is the tree the tape leaves out, written as data: how big each memory is, which sits
inside which, who can see it, how many copies exist, and which links join them at what speed.
spawn on(Topology::NPU[0]) names where a block runs in the same vocabulary. With both
declared, the compiler does ahead of time what the tools above do afterwards:
- A block can read only the memories its topology lists as
visible. - A transfer that crosses several hops is routed by the lowest predicted cost, from the declared bandwidths.
--diagnostics-jsonreports the bytes each transfer and eachspawnregion reads and writes, per memory and per buffer, derived from the generated code.vxc --machine fleet/h100-sxm.vxchecks a program against an H100 from a MacBook. Swapping the file checks it against an A100, a B200, an MI300X, or a Cortex-M7 whose tightly-coupled memory plays the part of shared memory.
NUMA, the case C++ allocators are built for, is the same relation as two GPUs on one node: a
memory some cores reach directly and others reach across a priced link. A Memory space
declares node: N, the runtime binds its pages there with mbind, and the
model predicted a remote access at 2.26× a local one against 2.35× measured.
The NUMA that wasn't there has the numbers, and
the machine that lied about them.
The facts survive until code generation
Declaring placement is worth little if the code generator never sees it. The Vx front end
lowers to a flat HIR and then to its own MLIR dialect, where placements and spawn
regions are still visible. Only after that does code go down to LLVM IR
for CPUs, NVVM and PTX for NVIDIA GPUs, or CoreML for the Apple Neural Engine. Because the
structure is still there, the backend reads facts where a C++ backend would have to infer them:
- A tile placed in
Memory::SMEMbecomes.sharedstorage in the PTX, and the launch geometry comes straight from the loop bounds. See Tiles, without a tile type. - A placed region that is one
linalg.matmulis still a matrix product when it reaches the runtime, so the runtime hands it to cuBLAS instead of running generated loops. - Whether the host may read a memory (
managed:in the machine file) is passed as an argument to the runtime call that allocates and copies, so the allocator and the type checker agree about which memory the host can touch. - A space's
node: Nreaches the runtime as a number it binds pages to.
No pass has to guess where a value lives, because nothing ever threw that away.
What is not built yet
This post argues that hiding costs is expensive, so here is what Vx does not do yet:
- The cost model has no fixed per-transfer cost. It prices a copy as bytes over bandwidth. On an A100 that is 98% low for a 4 KiB copy, the size of decode-step traffic. Adding the constant makes the cheapest route depend on the size of the transfer.
- The programmer still chooses placements. The compiler checks and prices them. Choosing them automatically needs the program as a graph of which region produces the bytes another consumes. The per-region byte counts exist. The edges between them do not.
- The host backend runs a kernel on one thread. NUMA placement per node measured 1.34× over interleaving with threads placed too; through today's single-threaded dispatch it is worth about 1.5× local against remote, and no more.
- Raw pointers carry no memory space.
*mut Tmeans host memory, as it does in C, and a device tensor cannot yet be passed to a C library through anexterndeclaration (#742). - Loop-invariant code motion and vectorization still come from LLVM, after lowering.
Each of these is work on the compiler. None needs the language to say more than it already says. A tape-model language can't be patched into seeing the tree, because the facts have to be in the source from the start; once they are, improving what the compiler does with them is ordinary work.
The point
Premature optimization spends an afternoon on speed before you know where the time goes. Premature de-optimization picks a language that can never say where the time goes, and so gives up speed in every program written in it, then pays for a set of tools to win part of it back.
On current hardware the time goes to moving data. Training runs sit around 40% of peak FLOPs, and decode on large models is in single digits, because the arithmetic units are waiting on memory. A language that cannot talk about memory and topology cannot talk about most of what its programs cost. Vx is at 0.0.2 and the list above is long, but it starts from the machine the program will actually run on.
Machine files covers every field and what the compiler derives from it, and the heterogeneous model covers placement and transfer. For the pieces of this argument in more depth, see Nonlinear Turing tape, float * of C/C++/Rust fails to distinguish host DRAM from GPU HBM and What a pointer forgets.