Hugging Face Daily PapersHao Liang, Qihan Lin, Meiyi Qiang1 min readpaperadvanced
OmniEdu: Open Foundation Models for Learning and Teaching
Summary
OmniEdu is an open family of LLMs (4B‑27B) fine‑tuned on a curated, capability‑balanced educational corpus (≈70k examples, 16M tokens) covering subject competence, curriculum grounding, diagnostic reasoning, and pedagogical scaffolding. Across model scales it improves on K‑12 benchmarks (K12‑Bench EM 63.12 %/F1 76.69 %, MathFish 85.89 %, EDUMATH 86.95 %, MathTutorBench Scaffold 78.74 %) and achie…
- A systematic instruction‑tuning pipeline can produce a compact (≈70k) yet diverse educational dataset without massive annotation effort.
- Balancing four capabilities—subject knowledge, curriculum alignment, diagnostic reasoning, and pedagogical action—yields consistent performance gains across model sizes.
- Even a 4B model benefits from the curated data, but the 27B OmniEdu reaches state‑of‑the‑art scores on multiple K‑12 math and tutoring benchmarks.
- Evaluation includes both problem‑solving metrics (Exact Match, F1) and teaching‑oriented metrics (LongTutor teaching average), highlighting the need for multi‑facet assessment in educational LLMs.
Educational LLMs have traditionally been either problem‑solvers or tutors, with data mixed by source rather than by pedagogical function. OmniEdu demonstrates that a capability‑balanced dataset can be assembled at modest scale and still deliver measurable improvements on both academic and teaching…
8/10





