proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameKaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun201542 min readpaperadvanced

Deep Residual Learning for Image Recognition

Summary

The paper proposes reformulating deep layers as residual functions with identity shortcut connections, making it easy to train networks far deeper than before. Using this design, a 152‑layer ResNet achieved 3.57% top‑5 error on ImageNet, winning ILSVRC 2015.

  • Residual blocks learn F(x) = H(x)‑x, adding the input back via an identity shortcut, which mitigates vanishing gradients.
  • Identity shortcuts add no parameters or compute cost, yet enable training of networks with >100 layers.
  • A 152‑layer ResNet outperformed much shallower VGG nets while using fewer FLOPs.
  • Residual learning generalizes to other tasks (detection, segmentation) and datasets (CIFAR‑10, COCO).

Anyone building or researching deep convolutional models should understand residual connections to scale depth without optimization headaches.

9/10

Related reading

  1. ImageNet Classification with Deep Convolutional Neural Networks

    AlexNet introduced a deep convolutional network with ReLU activations, dropout regularization, and multi‑GPU training, achieving 15.3% top‑5 error on ImageNet 2012, far surpassing prior results. The paper demonstrated that large‑scale CNNs are feasible and set the foundation for modern deep vision.

    Hall of Famenips.cc24 minpaper
  2. Generative Adversarial Nets

    This paper introduces Generative Adversarial Nets (GANs), a novel framework for training generative models. It pits a generator (G) against a discriminator (D) in a minimax game, where G tries to produce data that D cannot distinguish from real data, and D tries to correctly classify real vs. generated samples.

    Hall of Famearxiv.org20 minpaper
  3. Adam: A Method for Stochastic Optimization

    This paper introduces Adam, a first-order gradient-based optimization algorithm that adaptively estimates first and second moments of gradients. It computes individual learning rates for different parameters, making it efficient and robust for large-scale, high-dimensional machine learning problems with noisy or sparse gradients.

    Hall of Famearxiv.org32 minpaper
  4. VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

    VisionHOPE introduces a visual backbone that updates its own parameters on‑the‑fly using five coupled memories, with a stability‑matched step‑size scheme that keeps updates non‑expansive. The approach attains competitive ImageNet, COCO and ADE20K performance, showing self‑modifying learning systems can serve as practical general‑purpose vision models.

    Hugging Face Daily Papersarxiv.org1 minpaper