How Vx frees memory

TL;DR — Vx frees memory in four ways, and the type checker decides almost all of it. A tensor is freed right after its last use. A value whose type implements Drop is dropped at the end of its block, in reverse order, as in Rust. A tensor in GPU memory is freed through the plugin that allocated it. Raw memory from malloc is freed by hand. A week ago the compiler instead rebuilt ownership from its intermediate code, and most bugs in that approach were fixed by skipping code and leaking.

A compiler that already tracks who owns each value, to stop you using one after it moved, knows everything it needs to free it. Until this month Vx did not use that knowledge.

Vx used to free heap memory in three unrelated ways: a compiler pass over its intermediate code (MLIR) that freed the tensors the compiler allocated, a second pass for the buffers that transfer makes, and a free() you called by hand on Vec, Box and String. None of them used the ownership the checker tracks. The first rebuilt it from the MLIR, which has lost it, so a tensor in a struct field was leaked on purpose, and a program that placed data in GPU memory was skipped and leaked whole.

Now the checker writes each drop into the program, and both of Vx's code generators turn it into a call that frees the memory. This post walks through the four ways memory is freed today, with the drop points the compiler prints. Every example was compiled and run on both of Vx's code generators, from commit e602443b on main.

WayWhat it freesWhen
Drops the checker writesEvery tensor a function owns, including one in GPU memoryAfter its last use
The Drop traitStructs, enum payloads, Vec, Box, String, files and socketsAt the end of its block
Transfer freesA transfer result nothing namesAfter its last use
By handmalloc'd and mmap'd memoryWhen you call free or munmap

Tensors: freed after their last use

 1  fn total(t : Tensor<f32, [4]>) -> f32 {
 2    return t[0] + t[1];
 3  }
 4
 5  fn main() -> i32 {
 6    let a = Tensor<f32, [4]>::fill(1.0);
 7    let b = Tensor<f32, [4]>::fill(2.0);
 8    print(a[0]);
 9    let s = total(b);
10    let c = Tensor<f32, [4]>::fill(3.0);
11    if s > 1.0 {
12      let d = c;
13      print(d[0]);
14    }
15    print(s);
16    return 0;
17  }

With VX_PRINT_DROPS=1 the compiler prints where it put each drop:

drop in total: t, before the return on line 2
drop in main: d, after the statement on line 13
drop in main: c if it was not moved, after the statement on line 11
drop in main: a, after the statement on line 8

Each one follows a rule you could check by hand:

The program prints 134 on both code generators. A view of a tensor, such as a row q[i] or a field read, owns nothing and is never freed. It borrows its owner instead, so the owner is not freed before the view's last use.

Values with a Drop impl: freed at the end of their block

Freeing a tensor changes nothing a program can see, so the earlier the better. A Drop impl is different: it can print, unlock, or close a file, and when it runs is part of what the program means. So a value whose drop runs a Drop impl waits for the end of its block, and values are dropped in reverse order of declaration, as in Rust.

import std::vec;

struct Noisy { id : i32 }

impl Drop for Noisy {
  fn drop(self : &mut Noisy) -> void {
    print(self.id);
    print!(" ");
    return;
  }
}

struct Pair { left : Noisy, right : Noisy }

enum Slot { Full(Noisy), Empty }

fn main() -> i32 {
  let p = Pair { left : Noisy { id : 1 }, right : Noisy { id : 2 } };
  let mut v : Vec<Noisy> = Vec::new();
  v.push(Noisy { id : 3 });
  v.push(Noisy { id : 4 });
  let s = Slot::Full(Noisy { id : 5 });
  print!("end ");
  return 0;
}
end 5 3 4 1 2

That is the line the same program prints in Rust, on both of Vx's code generators. Nothing is dropped before end, though p, v and s are never used again after they are made. Then s, declared last, goes first and drops the payload of its variant (5). v drops its elements first to last (3 4) and then frees its buffer. p drops its fields in the order they are declared (1 2).

Vec, Box, String, File, TcpStream, UdpSocket and TcpListener all implement Drop, so the free() they used to need is gone. A struct with no Drop impl that only holds tensors has nothing visible to do when it is dropped, so it is freed after its last use, like a tensor. core::mem::drop(x) drops a value early, and core::mem::forget(x) never drops it.

Tensors in GPU memory

A tensor you transfer to another memory is freed by its drop too, through the allocator it came from: memory from cudaMalloc must go back through cudaFree. Here Device is a small GPU-like machine, declared in the program with its own memory, HBM:

 1  Topology Device {
 2    memory: Memory::HBM,
 3    visible: [Memory::HBM],
 4    transfer Memory::CPU_DRAM -> Memory::HBM : 10 GB/s,
 5    transfer Memory::HBM -> Memory::CPU_DRAM : 10 GB/s
 6  }
 7
 8  fn main() -> i32 {
 9    let host = Tensor<f32>([1024]);
10    let mut dev = transfer(host, Memory::HBM);
11    spawn on(Topology::Device) {
12      for i in 0..1024 {
13        dev[i] = 2.0;
14      }
15    }
16    let back = transfer(dev, Memory::CPU_DRAM);
17    print(back[0]);
18    return 0;
19  }
drop in main: host, after the statement on line 10
drop in main: dev, after the statement on line 16
drop in main: back, after the statement on line 17

The same rule as before: each tensor goes after its last use. What differs is the call each drop becomes. In the generated code (MLIR's LLVM dialect), host's drop is a call to free, while the drops of dev and back are calls to vx_plugin_free with the memory's topology:

llvm.call @free(%18)
llvm.call @vx_plugin_free(%20, %13)
llvm.call @vx_plugin_free(%32, %4)

vx_plugin_free asks the runtime for that topology. It calls cudaFree for GPU memory and free for host memory, or, when the buffer lives on a remote worker, sends that worker a FREE message. The program prints 2.

One kind of buffer has no owner to drop: a transfer whose result is never named, as in print(transfer(dev, Memory::CPU_DRAM)[0]). An older compiler pass, written before drops existed, frees a transfer's result after its last use, and it still handles exactly that case.

By hand

Memory you take from C stays yours to give back. std::alloc exposes malloc, realloc and free, and std::mmap exposes mmap and munmap, all unsafe, and nothing frees them for you:

import std::alloc;

fn main() -> i32 {
  unsafe {
    let p = malloc(16);
    *p = 7i8;
    print(*p);
    free(p);
  }
  return 0;
}

Memory that needs no free

How we check it

Vx's differential fuzzer writes random programs together with a Rust copy of each, and checks that both print the same thing. It can also preload a small library that counts the blocks a compiled program allocates and frees, and report a program that leaves any behind. The commit that added that counter reports 3,000 programs from the generator that moves tensors through calls, branches and loops, each printing what its Rust copy prints and each freeing everything it allocates on the default code generator. The commit that made placed tensors drop checked its new test, a GPU tensor returned from one function and given to another, the same way on both code generators.

What is left

Writing this post turned up one bug, and the drop work lists three more:

The design is in docs/implementation_plans/drop_semantics.md, and the work is tracked in #1041. Vx is Apache 2.0 with the LLVM exception — install it and run any program with VX_PRINT_DROPS=1 to see where its values are freed.