proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJunjie Chen, Fei Wang, Kun Li1 min readpaperadvanced

EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold

Summary

EvolvingAvatar is a novel 3D head generation model that adapts to ongoing conversations using test-time training and a self-supervised dyadic context prediction objective. It improves conversational motion statistics by learning from user audiovisual input during interaction, reducing expression mismatch by up to 11.1%.

  • EvolvingAvatar uses test-time training to adapt 3D head generation to live conversational dynamics.
  • A dyadic context prediction objective enables self-supervised learning from audiovisual context without explicit motion labels.
  • It employs persistent fast weights for long-term adaptation and transient jaw adaptation for immediate responses.
  • The model's adaptation improves generation, especially on out-of-distribution data, reducing expression mismatch.

This work is significant for researchers and developers building highly realistic and interactive virtual avatars, as it addresses the critical challenge of dynamic, context-aware conversational motion.

8/10

Related reading

  1. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

    OmniVChat defines native audio‑visual dialogue where a model consumes raw audio and video streams and replies in text. The authors build OmniVChat‑Studio, a multi‑agent simulator that generates single‑ and multi‑turn audio‑visual conversations, and use it to create OmniVChat‑Bench, a benchmark covering five dialogue abilities. They also propose OmniVChat‑RL, a reinforcement‑learning reward that b…

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. SteerDuplex: Steerable Duplex Speech Dialogue Models

    The paper presents SteerDuplex, a full‑duplex speech dialogue model that can be steered along tone, persona, and speed via instruction following, and introduces the SteerBench benchmark to evaluate such steerability. Supervised training yields a 44.5 % pass‑rate lift, and reinforcement‑learning fine‑tuning improves interruption handling and reduces pause barge‑ins, though reward hacking remains a…

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

    The paper presents AV‑GRPO, a diffusion‑based reinforcement‑learning framework that treats audio and video generation as separate but coordinated tasks by anchoring rollouts to each modality and freezing the opposite tower during optimization. Evaluated on the new 5DAV dataset and benchmarks, it achieves higher fidelity, better text‑modality alignment, and tighter audio‑video sync than the previo…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Article: Architecting Secure and Scalable Facial Verification Systems

    A real‑world post‑mortem of a high‑volume face verification service that moved from a naïve synchronous API to an async, layered pipeline (edge validation, preprocessing, decoupled detection/verification, decision engine) to achieve 8.5k rpm, p99 < 1.8 s, 30 % cost savings, and strict privacy controls.

    InfoQinfoq.com15 min
  6. StepAudio 3 Gen Technical Report

    StepAudio 3 Gen is a general‑purpose audio generation model that replaces diffusion with a discrete autoregressive generator over residual vector quantization tokens. Using a 16‑layer RVQ tokenizer and progressive pretraining, it reaches state‑of‑the‑art zero‑shot TTS and voice‑design performance while handling speech, vocals, sound effects, and music.

    Hugging Face Daily Papersarxiv.org2 minpaperHN2