Related reading
HelixWorld: A Real-time Interactive Audio-Visual World Model
HelixWorld is a real-time interactive audio-visual world model that generates synchronized visual scenes and camera-grounded spatial stereo sound. It achieves drift-free joint audio-visual rollouts at 24 FPS on a single GPU, matching visual fidelity of silent models while significantly improving spatial-acoustic immersion.
Hugging Face Daily Papersarxiv.org1 minpaperThe Rise of Audio AR
Hacker News front pagedbreunig.comHN3423OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.
Hugging Face Daily Papersarxiv.org1 minpaperStepAudio 3 Realtime Technical Report
StepAudio 3 Realtime is an audio‑language foundation model that runs a continuous listen‑converse‑think‑act loop. It introduces Deep Perception for rich acoustic cue extraction, Seamless Duplex for handling pauses/back‑channels, and a Think‑While‑Speaking mechanism that lets the model reason in parallel with speech output. On benchmarks it scores 73.0 macro avg on StepAudioChat, 90.6 on MMSU, 98.…
Hugging Face Daily Papersarxiv.org2 minpaperThe Effect of CRTs on Pixel Art
A nostalgic‑but‑technical look at how CRT displays interact with pixel art, covering scanlines, color bleed, signal quality, and how 8‑ and 16‑bit developers used dithering and anti‑aliasing to compensate for low resolution and limited palettes.
Hacker News front pagedatagubbe.se14 minHN313127Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
The authors present VGBench, a 1,018‑item diagnostic suite that tests whether audio LLMs correctly mute actions when the acoustic context shifts (e.g., speaker switches). Training with VoxGate dramatically improves gating, raising mute rates from <15% to >90% while preserving tool‑call accuracy.
Hugging Face Daily Papersarxiv.org1 minpaper
