proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZehao Jin, Junran Wang, Ruixuan Deng1 min readpaperadvanced

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

Summary

PersonaDose trains a description‑conditioned FLAS controller to steer LLM persona traits and calibrates flow time to hit a requested intensity, without needing intensity‑labeled data. Across Llama‑3.1‑8B, Qwen3‑8B, and Gemma‑3‑4B it boosts trait expression by up to 33 points and achieves mean targeting errors of 4.7‑6.2 points over reachable targets.

  • PersonaDose separates the learned behavioral range of a controller from the accuracy of intensity requests within that range.
  • Calibration of flow time yields mean targeting errors of 4.7‑6.2 points across 14‑22 reachable intensity targets per model.
  • On the Persona Vectors coherence floor of 75, trait expression improves by 33.2 (Llama‑3.1‑8B), 18.3 (Qwen3‑8B), and 17.8 (Gemma‑3‑4B) points versus contrastive activation addition.
  • The method requires no paired training data linking responses to target intensities, relying only on trait descriptions.

LLM engineers who need fine‑grained, controllable persona behavior can use PersonaDose to set trait intensity reliably without expensive labeled data.

7/10

Related reading

  1. Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations

    Deep Persona is a three‑layer architecture that models LLM personas with hierarchical expression, beliefs, and motivations, and enforces behavior through scripted determinism and bounded agency. The authors also introduce a reference‑free evaluation using psychological instruments, finding that structured personas improve alignment with human dialogue but LLMs still lag in emotion and joint atten…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs

    Personalized LLMs often lose user-specific characteristics when explicit style instructions are applied, a problem called personalization collapse. PsPLUG is a lightweight plug-in that addresses this by learning a user-specific residual, allowing dynamic control over personalization strength.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Dynamically Scaled Activation Steering

    Dynamically Scaled Activation Steering (DSAS) is a method‑agnostic framework that learns per‑token, per‑layer scaling factors to turn existing activation‑steering interventions on only when a model is likely to produce undesired output (e.g., toxic text). The scaling can be optimized jointly with any steering function, improves the toxicity‑utility trade‑off on language models, transfers to text‑…

    Apple Machine Learning Researchapple.com1 minpaper
  4. Mitigating the Length-Scaling Tax with Online Distillation

    The authors define the length‑scaling tax (LST) as excess response length without accuracy gain and propose Length Self‑Distillation (LSD), an online EMA‑based teacher that requires no external model. Experiments show LSD matches or exceeds RL performance while cutting LST from 19% to -3.7% on single‑turn and from 31.4% to 13.7% on multi‑turn tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation

    The paper shows that when specialist LLMs are trained only on QA pairs (no explicit reasoning supervision), their optimization implicitly selects a latent distribution of reasoning trajectories. By treating the distilled student as an agnostic probe—since it inherits only the sampled trajectories—the authors empirically demonstrate a strong correlation (across 27 specialist‑student pairs) between…

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness

    SkillSpec introduces a Hoare‑style framework that turns heterogeneous agent skill artifacts into a unified graph and reasons about correctness via intent‑masked specifications. In a study of 515 real‑world skills it flagged 763 confirmed defects with 61.2% precision, especially exposing intent‑implementation mismatches.

    Hugging Face Daily Papersarxiv.org1 minpaper