v5.01

SimonGiebenhain/MonoNPHMv5.01Jul 14, 2026by Zaneham

AI Summary

Major update rebranding from BarraCUDA to Booth, adding significant CPU compilation support for kernels previously requiring GPUs.

Key Highlights

  • Renamed compiler from BarraCUDA to Booth
  • Added CPU-side support for scalar arithmetic on x86-64 and RV64
  • Tensix emits native machine code for RISC-V cores
  • Inlined __device__ calls on GPU and vector backends
  • Improved diagnostics resembling Clang's output

Breaking Changes

  • Renamed compiler binary and version jump

New Features

  • CPU compilation support for CUDA, HIP, and Triton kernels
  • Native machine code generation for RISC-V cores

Full Release Notes

So long, and thanks for all the fish.

BarraCUDA was a good pun and a bad description, so it swam off: the compiler is Booth now, the binary is `kath`, and the version jumps to mark the line. This is the first release under the new name.

Since 0.5 the compiler grew a spine on the CPU side. x86-64 and RV64 both gained most of their scalar arithmetic, so a CUDA, HIP or Triton kernel now compiles and runs on a machine with no GPU in it at all. Tensix learned to emit native machine code for the baby RISC-V cores instead of leaning on a C++ handoff. `__device__` calls are inlined away before isel, so device helpers with control flow finally work on the GPU and vector backends. Diagnostics were rebuilt to read like Clang's, carets and colour and all, and the frontend picked up a pile of coverage.

Thanks to the people who sent patches this cycle: @GauthamMK-0 for `__hip_bfloat16` and the `warpSize` builtin, @nataliakokoromyti for lowering `tl.where` to select and keeping the PTX header ASCII, and @kstppd for fixing the Makefile on non-x86 machines. Much appreciated.