Related reading
Nvidia announces native GPU programming in Rust
NVIDIA released CUDA‑Rust, letting you write GPU kernels directly in Rust and compile to PTX. Two programming models are supported: the traditional SIMT model via the `cuda-oxide` backend (nightly Rust, custom codegen) and the newer Tile model via `cutile‑rs` (stable Rust, JIT‑compiled Tile IR). Both provide Rust‑typed safety guarantees (e.g., `DisjointSlice`, tensor partitioning) and simple Carg…
Hacker News front pagenvidia.com11 minHN961402Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
The paper introduces mutation analysis as a quantitative adequacy metric for GPU‑kernel benchmark oracles. By injecting 10,303 deterministic faults into 188 verified CUDA kernels (7,384 with a known kill witness), they show the official KernelBench checker misses 16.9% of faults—especially 78.6% of precision‑related faults. Their analysis quantifies the impact of existing patches (e.g., KernelBen…
Hugging Face Daily Papersarxiv.org1 minpaperThe work by Valve's Timur Kristóf on improving old AMD GPUs on Linux
Timur Kristóf from Valve has significantly improved Linux driver support for old AMD GCN 1.0/1.1 GPUs, enabling them to use the modern AMDGPU kernel driver and RADV Vulkan driver. This work addresses display and power management issues, and adds soft reset support, leading to better performance and functionality for aging hardware.
Hacker News front pagephoronix.com1 minHN482102Building a Linux GPU Driver for the M4 Mac Mini in One Month
Built a clean‑room OpenGL ES 3.0 Linux driver for Apple‑silicon M4/A18 Pro GPUs in ~4 weeks, covering reverse‑engineered firmware ABI, a Rust kernel driver, a custom IR/shader compiler, and user‑space Metal translation; achieved 200 fps Minecraft and WebGL demos, with heavy LLM assistance for debugging and code generation.
Hacker News front pagecodyho.dev15 minHN416281GTR: Gated Token Recurrence for Efficient Dense Prediction
GTR replaces quadratic softmax attention with gated linear attention and spatial recurrence, delivering a fast, softmax‑free vision backbone. It reaches 58.9 COCO box AP with ~1.9 ms latency on RTX 4090 and runs efficiently on edge GPUs via TensorRT.
Hugging Face Daily Papersarxiv.org1 minpaper

