Burn 0.22.0: Faster Builds, Easier Extensions, and Smarter Autotuning
Faster builds, easier custom operations, adaptive memory management, and smarter autotuning…
Blog
Explore what's new, our releases, technical work, and practical guides from across the Tracel ecosystem.
Faster builds, easier custom operations, adaptive memory management, and smarter autotuning…
A dedicated repository, broader model coverage, and a new ONNX compliance gate…
Distributed training, lower overhead, improved kernels, and early reinforcement-learning support…
Specialized asynchronous channels and memory arenas accelerate Burn communication…
CubeCL unifies optimized CPU and GPU kernels, with Blackwell support…
A year of performance work and tighter integration across Burn and CubeCL…
Faster builds, easier custom operations, adaptive memory management, and smarter autotuning…
A dedicated repository, broader model coverage, and a new ONNX compliance gate…
Distributed training, lower overhead, improved kernels, and early reinforcement-learning support…
Specialized asynchronous channels and memory arenas accelerate Burn communication…
CubeCL unifies optimized CPU and GPU kernels, with Blackwell support…
A year of performance work and tighter integration across Burn and CubeCL…
How asynchronous GPU behavior can hide the real performance bottleneck…
Quantization, multi-GPU training, LLVM CPU execution, and a renewed training loop…
New funding supports Tracel’s mission to broaden access to AI compute…
An alpha inference engine for LLMs, VLMs, training, and fine-tuning…
Portable matrix-multiplication kernels designed to rival vendor-specific libraries…
Performance, reliability, and optimization improvements across the framework…
A Metal backend, expanded tensor fusion, and CubeCL-powered runtimes…
Why memory bandwidth makes efficient model quantization essential for inference…
Compiler-focused techniques cut matrix-multiplication benchmark build time dramatically…
Faster tensor operations, experimental runtimes, and new quantization foundations…
A roadmap for performance, portability, and flexibility across the compute stack…
Major tensor-operation improvements and experimental multi-backend capabilities…
A performance program for AI beyond hardware and software constraints…
Build a ResNet workflow in Burn and import ImageNet-trained weights…
An introduction to Burn datasets, batchers, and the data-loading workflow…
CubeCL joins Burn for cross-platform GPU programming, alongside broader improvements…
New tensor operations, autodiff advances, backend refactoring, and framework features…
Burn’s eager tensor streams enable fused kernels without static graphs…
How Burn selects high-performing GPU kernels for each workload and device…
A new company from Burn’s creator, built around high-performance and portable AI infrastructure…
Burn-Compute brings asynchronous execution, memory management, and autotuning to backends…
A Rust GPU backend built for portable deep-learning deployments…
Rust ownership patterns make tensor allocation and reuse substantially more efficient…
Why Rust’s safety and concurrency abstractions fit modern deep-learning systems…
No posts match these filters.
Loading the demo request form… Contact us
Tell us how you work with compute. We'll take it from there.
Select the hardware you primarily use. A mixed setup is welcome.
Select at least one.
Let's put your compute to work
A conversation starts here
We'll be in touch shortly.