Premature de-optimization is the root of all evil

TL;DR — Popular languages make three choices for you: memory has no location, the machine has no shape, and the source is lowered to a low-level IR almost at once. Each choice throws away something the programmer knew. Together they make programs slow before anyone writes them, and a set of tools that few engineers master exists to win the speed back. Vx puts memory and topology in the program and keeps them there until code generation.

Knuth warned against optimizing before you know where the time goes. The more common mistake today is choosing, before you start, a language that cannot say where the time goes.

Three choices made before the first line

Memory is hidden. In C, C++, Rust, Go, Java and Python, memory is one flat array of bytes. A float * does not say whether the floats are in the local socket's DRAM, the other socket's DRAM, a GPU's HBM across PCIe, or a CXL pool across the rack. Those differ in latency by orders of magnitude, and the type system treats them as one thing. The programmer cannot write down “this is small and hot, keep it close”, and the compiler cannot ask “does this even fit?”

Topology is hidden. A server has sockets, NUMA domains, cores that share some caches and not others, GPUs behind particular PCIe switches, and NVLink or InfiniBand between boxes. None of it appears in the language. A thread is a thread, and where it runs is decided by the OS scheduler, a runtime, or an environment variable in a launch script. Code cannot say “run next to that data”, so it runs wherever it lands and the data travels to meet it.

Lowering happens early. Clang turns C++ into LLVM IR in one walk over the syntax tree. Rust keeps MIR for borrow checking and then hands the same optimizer LLVM IR. Java and Python become bytecode. From then on the optimizer sees loads, stores, pointer arithmetic and branches, and spends much of its effort rebuilding what the source already said: loop analysis finds the loops again, scalar evolution finds the induction variables again, alias analysis guesses which pointers overlap. LLVM even has a pass, InferAddressSpaces, whose job is to work out which GPU memory a generic pointer points into, after the language dropped that fact from the type. When a guess fails, the optimizer has to assume the worst, and assuming the worst is slow.

All of them model a tape

Put the three together and every mainstream language targets the same machine: Turing's tape. One line of cells, one head, every cell the same distance away and the line never ending. The hardware is a tree of memories of very different sizes and speeds, joined by links of very different bandwidths, with compute scattered across the leaves. An earlier post walks through that tree for one H100 node.

A program written for the tape and run on the tree is slow from the start. That is the de-optimization in the title, and no programmer chose it. The language chose it for every program at once.

The tools that win it back

The information the language will not carry is recovered afterwards, by hand:

These tools are good, and few engineers use them well. The usual cycle is to ship, profile, find half the time going to remote memory or PCIe copies, add numactl --membind to a launch script, and hope nobody changes the machine. The knowledge ends up in scripts, wiki pages and people's heads. The compiler never sees it, cannot check it, and the next refactor can break it without a warning.

C++ allocators go part of the way

C++ is the one mainstream language that tries. std::allocator, and later std::pmr memory resources, let a container ask a chosen allocator for memory. You can write one that calls numa_alloc_onnode, or one that hands out pinned host memory. For NUMA placement that helps. Beyond it, very little:

Libraries go further, and stop at the same place. SYCL and Kokkos type the memory region inside a device. Neither can make where a function runs part of the function's type, because C++ has nowhere to put it. Kokkos shows the limit plainly: when SpaceAccessibility<...>::accessible is a compile-time false, meaning the compiler has already proved the access illegal, the library can only emit a run-time Kokkos::abort. A language can do better, so Vx is one.

What Vx does with each of the three

Memory has a location, and it is part of the type

A Vx tensor's type names the memory it lives in, and moving it is an explicit transfer that produces a value of a new type:

fn main() -> i32 {
  let a = Tensor<f32, [4, 4]>::fill(2.0);
  let b = Tensor<f32, [4, 4]>::fill(3.0);
  let mut c = Tensor<f32, [4, 4], Memory::NPU_HBM>::uninit();

  let a_npu = transfer(a, Memory::NPU_HBM);
  let b_npu = transfer(b, Memory::NPU_HBM);

  spawn on(Topology::NPU[0]) {
    c = a_npu @ b_npu;
  }

  let c_host = transfer(c, Memory::CPU_DRAM);
  print(c_host[0][0]);
  return 0;
}

The allocator's answer is kept for the value's whole life, where C++ drops it at the first T *. So every later use can be checked, and mistakes the tape model compiles are refused on a laptop:

MistakeC++ / CUDAVx
Host reads device memorysegfault or cudaErrorIllegalAddressE6003, naming the missing transfer and its cost
Working set larger than the memoryfailed cudaMalloc, or OOM hours inE6009 / E6010, against the declared capacity
Copy between memories with no linkrun-time error or a slow fallbackE6002, no path exists

The program prints 24. Replace transfer(a, Memory::NPU_HBM) with plain a, and the compiler refuses it:

Error[E6003] at 10:9: 'a_npu' lives in CPU_DRAM but NPU[0] sees only [NPU_HBM];
  insert an explicit transfer to NPU_HBM (cost 50 on the declared path)

The machine is an input, and code says where it runs

Vx reads the machine from a declaration file. This is the H100 description that ships in fleet/, without its citations:

Memory HBM {
  capacity: 80 GiB, bandwidth: 3.35 TB/s, managed: explicit, scope: device
}
Memory L2 {
  within: Memory::HBM, capacity: 50 MiB, bandwidth: 12 TB/s, managed: cached, scope: device
}
Memory SMEM {
  within: Memory::L2, capacity: 228 KiB, bandwidth: 128 B/cyc, clock: 1.98 GHz,
  granule: 1 KiB, managed: explicit, scope: sm, replicas: 132
}
Topology Device {
  arch: nvptx64,
  memory: Memory::HBM,
  visible: [Memory::HBM, Memory::L2, Memory::SMEM],
  transfer Memory::CPU_DRAM -> Memory::HBM : 63 GB/s,
  transfer Memory::HBM -> Memory::CPU_DRAM : 63 GB/s,
  transfer Memory::HBM -> Memory::L2,
  transfer Memory::L2 -> Memory::SMEM copy_engine
}

That is the tree the tape leaves out, written as data: how big each memory is, which sits inside which, who can see it, how many copies exist, and which links join them at what speed. spawn on(Topology::NPU[0]) names where a block runs in the same vocabulary. With both declared, the compiler does ahead of time what the tools above do afterwards:

NUMA, the case C++ allocators are built for, is the same relation as two GPUs on one node: a memory some cores reach directly and others reach across a priced link. A Memory space declares node: N, the runtime binds its pages there with mbind, and the model predicted a remote access at 2.26× a local one against 2.35× measured. The NUMA that wasn't there has the numbers, and the machine that lied about them.

The facts survive until code generation

Declaring placement is worth little if the code generator never sees it. The Vx front end lowers to a flat HIR and then to its own MLIR dialect, where placements and spawn regions are still visible. Only after that does code go down to LLVM IR for CPUs, NVVM and PTX for NVIDIA GPUs, or CoreML for the Apple Neural Engine. Because the structure is still there, the backend reads facts where a C++ backend would have to infer them:

No pass has to guess where a value lives, because nothing ever threw that away.

What is not built yet

This post argues that hiding costs is expensive, so here is what Vx does not do yet:

Each of these is work on the compiler. None needs the language to say more than it already says. A tape-model language can't be patched into seeing the tree, because the facts have to be in the source from the start; once they are, improving what the compiler does with them is ordinary work.

The point

Premature optimization spends an afternoon on speed before you know where the time goes. Premature de-optimization picks a language that can never say where the time goes, and so gives up speed in every program written in it, then pays for a set of tools to win part of it back.

On current hardware the time goes to moving data. Training runs sit around 40% of peak FLOPs, and decode on large models is in single digits, because the arithmetic units are waiting on memory. A language that cannot talk about memory and topology cannot talk about most of what its programs cost. Vx is at 0.0.2 and the list above is long, but it starts from the machine the program will actually run on.

Machine files covers every field and what the compiler derives from it, and the heterogeneous model covers placement and transfer. For the pieces of this argument in more depth, see Nonlinear Turing tape, float * of C/C++/Rust fails to distinguish host DRAM from GPU HBM and What a pointer forgets.