How Vx frees memory
TL;DR — Vx frees memory in four ways, and the type checker decides almost all of it. A tensor is freed right after its last use. A value whose type implementsDropis dropped at the end of its block, in reverse order, as in Rust. A tensor in GPU memory is freed through the plugin that allocated it. Raw memory frommallocis freed by hand. A week ago the compiler instead rebuilt ownership from its intermediate code, and most bugs in that approach were fixed by skipping code and leaking.
A compiler that already tracks who owns each value, to stop you using one after it moved, knows everything it needs to free it. Until this month Vx did not use that knowledge.
Vx used to free heap memory in three unrelated ways: a compiler pass over its intermediate
code (MLIR) that freed the tensors the compiler allocated, a second pass for the buffers that transfer makes, and a free()
you called by hand on Vec, Box and String. None of them used
the ownership the checker tracks. The first rebuilt it from the MLIR, which has lost it, so a
tensor in a struct field was leaked on purpose, and a program that placed data in GPU memory was
skipped and leaked whole.
Now the checker writes each drop into the program, and both of Vx's code generators turn it
into a call that frees the memory. This post
walks through the four ways memory is freed today, with the drop points the compiler prints. Every
example was compiled and run on both of Vx's code generators, from commit e602443b on
main.
| Way | What it frees | When |
|---|---|---|
| Drops the checker writes | Every tensor a function owns, including one in GPU memory | After its last use |
The Drop trait | Structs, enum payloads, Vec, Box, String, files and sockets | At the end of its block |
| Transfer frees | A transfer result nothing names | After its last use |
| By hand | malloc'd and mmap'd memory | When you call free or munmap |
Tensors: freed after their last use
1 fn total(t : Tensor<f32, [4]>) -> f32 {
2 return t[0] + t[1];
3 }
4
5 fn main() -> i32 {
6 let a = Tensor<f32, [4]>::fill(1.0);
7 let b = Tensor<f32, [4]>::fill(2.0);
8 print(a[0]);
9 let s = total(b);
10 let c = Tensor<f32, [4]>::fill(3.0);
11 if s > 1.0 {
12 let d = c;
13 print(d[0]);
14 }
15 print(s);
16 return 0;
17 }
With VX_PRINT_DROPS=1 the compiler prints where it put each drop:
drop in total: t, before the return on line 2
drop in main: d, after the statement on line 13
drop in main: c if it was not moved, after the statement on line 11
drop in main: a, after the statement on line 8
Each one follows a rule you could check by hand:
ais read for the last time on line 8, so it is freed right after that line, whilemainkeeps running. The checker finds that line with a liveness analysis: for each block, it works out the last statement that uses each value.bis passed by value tototal, so it belongs tototalnow, andtotalfrees it before it returns.mainnever frees it.cis moved intod, but only when theifis taken. Whether it still needs freeing is only known at run time, so the compiler keeps a hiddenboolthat the move sets, and the drop after theiffreesconly ifcwas not moved.downs whatcdid, and is freed after its last use on line 13.
The program prints 134 on both code generators. A view of a tensor, such as a row
q[i] or a field read, owns nothing and is never freed. It borrows its owner instead,
so the owner is not freed before the view's last use.
Values with a Drop impl: freed at the end of their block
Freeing a tensor changes nothing a program can see, so the earlier the better. A
Drop impl is different: it can print, unlock, or close a file, and when it runs is
part of what the program means. So a value whose drop runs a Drop impl waits for the
end of its block, and values are dropped in reverse order of declaration, as in Rust.
import std::vec;
struct Noisy { id : i32 }
impl Drop for Noisy {
fn drop(self : &mut Noisy) -> void {
print(self.id);
print!(" ");
return;
}
}
struct Pair { left : Noisy, right : Noisy }
enum Slot { Full(Noisy), Empty }
fn main() -> i32 {
let p = Pair { left : Noisy { id : 1 }, right : Noisy { id : 2 } };
let mut v : Vec<Noisy> = Vec::new();
v.push(Noisy { id : 3 });
v.push(Noisy { id : 4 });
let s = Slot::Full(Noisy { id : 5 });
print!("end ");
return 0;
}
end 5 3 4 1 2
That is the line the same program prints in Rust, on both of Vx's code generators. Nothing is
dropped before end, though p, v and s are never
used again after they are made. Then s, declared last, goes first and drops the
payload of its variant (5). v drops its elements first to last (3 4) and then frees
its buffer. p drops its fields in the order they are declared (1 2).
Vec, Box, String, File,
TcpStream, UdpSocket and TcpListener all implement
Drop, so the free() they used to need is gone. A struct with no
Drop impl that only holds tensors has nothing visible to do when it is dropped, so it
is freed after its last use, like a tensor. core::mem::drop(x) drops a value early, and
core::mem::forget(x) never drops it.
Tensors in GPU memory
A tensor you transfer to another memory is freed by its drop too, through the
allocator it came from: memory from cudaMalloc must go back through
cudaFree. Here Device is a small GPU-like machine, declared in the
program with its own memory, HBM:
1 Topology Device {
2 memory: Memory::HBM,
3 visible: [Memory::HBM],
4 transfer Memory::CPU_DRAM -> Memory::HBM : 10 GB/s,
5 transfer Memory::HBM -> Memory::CPU_DRAM : 10 GB/s
6 }
7
8 fn main() -> i32 {
9 let host = Tensor<f32>([1024]);
10 let mut dev = transfer(host, Memory::HBM);
11 spawn on(Topology::Device) {
12 for i in 0..1024 {
13 dev[i] = 2.0;
14 }
15 }
16 let back = transfer(dev, Memory::CPU_DRAM);
17 print(back[0]);
18 return 0;
19 }
drop in main: host, after the statement on line 10
drop in main: dev, after the statement on line 16
drop in main: back, after the statement on line 17
The same rule as before: each tensor goes after its last use. What differs is the call each
drop becomes. In the generated code (MLIR's LLVM dialect), host's drop is a call to
free, while the drops of dev and back are calls to
vx_plugin_free with the memory's topology:
llvm.call @free(%18)
llvm.call @vx_plugin_free(%20, %13)
llvm.call @vx_plugin_free(%32, %4)
vx_plugin_free asks the runtime for that topology. It calls cudaFree
for GPU memory and free for host memory, or, when the buffer lives on a remote worker,
sends that worker a FREE message. The program prints 2.
One kind of buffer has no owner to drop: a transfer whose result is never named, as in
print(transfer(dev, Memory::CPU_DRAM)[0]). An older compiler pass, written
before drops existed, frees a transfer's result after its last use, and it still handles exactly
that case.
By hand
Memory you take from C stays yours to give back. std::alloc exposes
malloc, realloc and free, and std::mmap exposes
mmap and munmap, all unsafe, and nothing frees them for
you:
import std::alloc;
fn main() -> i32 {
unsafe {
let p = malloc(16);
*p = 7i8;
print(*p);
free(p);
}
return 0;
}
Memory that needs no free
- A kernel's stack. A small fixed-size tensor made at the start of a
spawnregion becomes a stack slot in the kernel, and a GPU shared-memory tile lives only as long as the kernel. Both are gone when the kernel returns. Any other tensor a region makes is freed by its drop, inside the kernel. - The runtime's staging pool. When the runtime copies a tensor to a GPU for one
kernel launch, that copy belongs to the runtime, and goes back to a pool when the launch ends.
The pool keeps the memory for the next launch, and calls
cudaFreeonly when keeping it would take the pool past its limit.
How we check it
Vx's differential fuzzer writes random programs together with a Rust copy of each, and checks that both print the same thing. It can also preload a small library that counts the blocks a compiled program allocates and frees, and report a program that leaves any behind. The commit that added that counter reports 3,000 programs from the generator that moves tensors through calls, branches and loops, each printing what its Rust copy prints and each freeing everything it allocates on the default code generator. The commit that made placed tensors drop checked its new test, a GPU tensor returned from one function and given to another, the same way on both code generators.
What is left
Writing this post turned up one bug, and the drop work lists three more:
- A
matchon a variant whose name another enum also uses picks the wrong variant on the legacy code generator, at random. A drop of an enum is written as amatch, so aDropimpl on the payload ran in only 2 of 8 compiles of the same program (#1342). - An enum with a number and a struct at the same payload position is never dropped (#1343).
String::as_c_strleaks a copy on every call (#723).- The legacy code generator leaks a placed tensor returned inside
Verified<..>(#1335).
The design is in
docs/implementation_plans/drop_semantics.md, and the work is tracked in
#1041. Vx is Apache 2.0 with the LLVM
exception — install it and run any program with
VX_PRINT_DROPS=1 to see where its values are freed.