Hugging Face Daily PapersTristan Kirscher, Vivian Metzger, Philippe Meyer1 min readpaperadvanced
Assessing nnU-Net Generalization across Brain Tumor Populations in BraTS-GoAT 2026
Summary
Paper evaluates a standard 3‑D nnU‑Net on the new BraTS‑GoAT benchmark, training on 1,351 cases with five‑fold cross‑validation and test‑time mirroring. It reports Dice scores of 0.78/0.83/0.89 (ET/TC/WT) and shows a ~0.07 drop on heterogeneous validation, with limited benefit from ensembling or mirroring and failure linked to small, fragmented tumors.
- 5‑fold nnU‑Net trained on 1,351 BraTS‑GoAT cases achieved DSC 0.78 (ET), 0.83 (TC), 0.89 (WT) on pooled validation.
- Dice drops ~0.07 when moving from source out‑of‑fold to heterogeneous validation, indicating limited generalization.
- Test‑time mirroring gives marginal gains; ensembling across folds shows no clear benefit.
- A residual‑encoder variant matches nnU‑Net performance (0.8282 mean Dice) without extra complexity.
Engineers building medical image segmentation pipelines need to understand nnU‑Net's generalization limits on diverse brain tumor datasets.
6/10


