proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameAlex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton201224 min readpaperadvanced

ImageNet Classification with Deep Convolutional Neural Networks

Summary

AlexNet introduced a deep convolutional network with ReLU activations, dropout regularization, and multi‑GPU training, achieving 15.3% top‑5 error on ImageNet 2012, far surpassing prior results. The paper demonstrated that large‑scale CNNs are feasible and set the foundation for modern deep vision.

  • ReLU non‑linearity trains several times faster than tanh, enabling deep networks on large data.
  • Dropout applied to fully‑connected layers dramatically reduces overfitting.
  • Model parallelism across two GPUs with limited inter‑GPU communication improves accuracy and training speed.
  • Architecture of 5 conv + 3 FC layers (~60 M parameters) achieved 15.3% top‑5 error, a record at the time.

Deep‑learning engineers and researchers building large‑scale vision models should know the techniques that made modern CNNs practical.

9/10

Related reading

  1. Deep Residual Learning for Image Recognition

    The paper proposes reformulating deep layers as residual functions with identity shortcut connections, making it easy to train networks far deeper than before. Using this design, a 152‑layer ResNet achieved 3.57% top‑5 error on ImageNet, winning ILSVRC 2015.

    Hall of Famearxiv.org42 minpaper
  2. Generative Adversarial Nets

    This paper introduces Generative Adversarial Nets (GANs), a novel framework for training generative models. It pits a generator (G) against a discriminator (D) in a minimax game, where G tries to produce data that D cannot distinguish from real data, and D tries to correctly classify real vs. generated samples.

    Hall of Famearxiv.org20 minpaper
  3. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

    VisionHOPE introduces a visual backbone that updates its own parameters on‑the‑fly using five coupled memories, with a stability‑matched step‑size scheme that keeps updates non‑expansive. The approach attains competitive ImageNet, COCO and ADE20K performance, showing self‑modifying learning systems can serve as practical general‑purpose vision models.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

    This paper evaluates GPT-6 Astra and five other frontier general-purpose AI systems across 34 computer vision capabilities and 55 benchmarks. It finds these systems excel at semantic interpretation and reasoning, but struggle with metric geometric accuracy, faithful reconstruction, and fine-grained specialized knowledge.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. Scaling Telco Autonomy: Leveraging GNNs with Distributed GraphFlow

    Google Cloud’s blog introduces Distributed GraphFlow (DGF), an open‑source Python library for building and scaling Graph Neural Networks (GNNs) on a Spanner‑backed digital twin of telecom networks. The post outlines the three‑layer architecture (digital twin on Spanner Graph, ML layer with DGF, AI agents) and highlights DGF’s high‑level API (5‑line example) and low‑level primitives, but provides…

    Google Cloud Bloggoogle.com3 min