Hugging Face Daily PapersSimon Schug, Brenden M. Lake1 min readpaperintermediate
Thought without systematicity? Evaluating reasoning models on rule induction tasks
Summary
The paper builds compositional rule‑induction task families and generates isomorphic variants to probe systematicity in reasoning models. Experiments reveal that models frequently fail on equivalent variants despite solving the original task, indicating a lack of systematic reasoning.
- Models can solve a rule‑induction task yet often fail on structurally equivalent variants created via recombination or substitution.
- The authors construct task families with compositional structure and define isomorphisms to test systematicity.
- Empirical results show a systematicity gap: correct performance on one task does not guarantee performance on its isomorphic counterpart.
- Systematicity should be incorporated into benchmark suites for reasoning models to assess robust generalization.
Anyone evaluating or developing LLM reasoning abilities needs to know that current models may not generalize systematically across equivalent tasks.
5/10
