Related reading
Nvidia announces native GPU programming in Rust
NVIDIA released CUDA‑Rust, letting you write GPU kernels directly in Rust and compile to PTX. Two programming models are supported: the traditional SIMT model via the `cuda-oxide` backend (nightly Rust, custom codegen) and the newer Tile model via `cutile‑rs` (stable Rust, JIT‑compiled Tile IR). Both provide Rust‑typed safety guarantees (e.g., `DisjointSlice`, tensor partitioning) and simple Carg…
Hacker News front pagenvidia.com11 minHN961402The work by Valve's Timur Kristóf on improving old AMD GPUs on Linux
Timur Kristóf from Valve has significantly improved Linux driver support for old AMD GCN 1.0/1.1 GPUs, enabling them to use the modern AMDGPU kernel driver and RADV Vulkan driver. This work addresses display and power management issues, and adds soft reset support, leading to better performance and functionality for aging hardware.
Hacker News front pagephoronix.com1 minHN482102Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
Flash-dLLM is a training-free framework that accelerates Diffusion LLM inference by addressing GPU memory I/O bottlenecks with an I/O-aware KV-cache kernel. It also introduces a KV-cache-driven draft-and-verify decoding strategy, achieving significant speedups (up to 11x) over prior methods.
Hugging Face Daily Papersarxiv.org1 minpaperOpenJev
OpenJev is a browser‑only demo that lets you load small LLM checkpoints (e.g., MiniCPM‑5 2B, Qwen3 0.6B) onto your GPU and compare two inference paths: reading raw logits for a set of options versus prompting the model to emit a JSON with option probabilities token‑by‑token. The page reports model sizes, download times, balanced accuracy on a few benchmarks, and wall‑clock timings measured with `…
Hacker News front pageopenjev.com2 minreleaseHN709288

