proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersLei Ke, Jiahao Pan, Zeyue Tian1 min readpaperadvanced

HelixWorld: A Real-time Interactive Audio-Visual World Model

Summary

HelixWorld is a real-time interactive audio-visual world model that generates synchronized visual scenes and camera-grounded spatial stereo sound. It achieves drift-free joint audio-visual rollouts at 24 FPS on a single GPU, matching visual fidelity of silent models while significantly improving spatial-acoustic immersion.

  • HelixWorld is a real-time interactive world model generating synchronized visuals and spatial stereo sound.
  • It achieves 24 FPS for joint audio-visual rollouts on a single GPU, enabling low-latency interaction.
  • A high-fidelity spatial audio-visual dataset with true stereo acoustics was curated for training.
  • A teacher model is distilled into a few-step streaming student for causal interaction.

This work is significant for researchers and developers building immersive virtual environments, as it addresses the critical gap of integrating realistic, synchronized spatial audio into interactive world models.

8/10

Related reading

  1. Helix: The internal tool powering our Shopify app's native migration

    Helix is Shopify’s internal LLM‑driven framework for migrating React Native screens to native iOS/Android. It breaks a screen into tiny, reviewable checkpoints and forces each through four strict gates—behavior tests via a CLI, Gemini‑powered visual diff, two adversarial code reviewers, and a final engineer sign‑off—while remembering feedback to become more autonomous over time.

    Shopifyshopify.engineering8 minHN1
  2. OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

    OSWorld-Science introduces a benchmark suite of 146 scientific software tasks for evaluating visual language model agents across domains like molecular design and image analysis. The authors provide a harness for systematic comparison and show that current state‑of‑the‑art VLMs still struggle with many scientific workflows.

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. ROWBench: Do Video Models Render What the Program Specifies?

    PROWBench is a new benchmark designed to evaluate the visual fidelity of programmable world models to fine-grained, program-specified events and interactions. It comprises 170 programmatically constructed episodes and 600 proxy videos, using VLM-based metrics to check generated videos against observable consequences of program execution.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper