proomt

Search

Search posts, papers, and topics

probing

RSS
  1. 1

    A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

    The paper presents Probe of Internal Recognition (PIR), a reference‑free technique that reads a language model’s internal activations to detect which answer it recognizes, achieving 70‑87% balanced accuracy across eight LLMs. PIR reliably distinguishes deliberate concealment from lack of knowledge, enabling audits of sandbagging and unlearning.

    Hugging Face Daily Papersarxiv.org1 minpaper