proomt

Search

Search posts, papers, and topics

All posts

Hacker News front page19 min readintermediate

Writing Efficient C++ Code

Summary

Writing efficient C++ code requires understanding hardware memory hierarchy and cache behavior. Data-Oriented Design, which prioritizes contiguous data layout, is key to maximizing performance by reducing cache misses and enabling easier parallelization.

  • Data-Oriented Design (DOD) prioritizes data layout in memory to optimize for cache efficiency and parallel processing.
  • Scattered objects and pointer indirection common in OOP lead to frequent cache misses and hinder safe parallelization.
  • Contiguous data structures like std::vector improve cache hits by ensuring frequently accessed data resides in the same cache line.
  • Memory access latency is hundreds of CPU cycles; optimizing cache usage is more critical than minimizing instruction count.

Engineers working on performance-critical C++ applications, especially in domains like game development or real-time processing, will find practical guidance on optimizing for modern hardware.

7/10

Related reading

  1. What Every Programmer Should Know About Memory

    The paper explains how modern CPU caches, memory controllers, and NUMA architectures affect program performance and what developers can do to write cache‑friendly code. It provides concrete advice and tooling to reduce cache misses, avoid false sharing, and place memory near the accessing cores.

    Hall of Famefreebsd.org397 minpaperHN20346
  2. Article: Your Next DSL Author Is a Language Model

    Typed Domain Grounding (TDG) embeds a DSL inside a mainstream language the LLM already knows (e.g., Kotlin) and uses the host compiler as an oracle. The author describes five building blocks—embedding, choosing a host language with high training‑data frequency, compiler‑driven type safety, a generate‑compile‑repair loop, and an on‑demand teaching tool—and shows measured results from kUML, a Kotli…

    InfoQinfoq.com18 min
  3. Gallery of Processor Cache Effects

    This article illustrates various processor cache effects using C# code examples and measured performance. It demonstrates how cache lines, L1/L2 cache sizes, instruction-level parallelism, and cache associativity impact program execution times.

    Hall of Fameigoro.com12 minHN3
  4. Textbook review: Is Parallel Programming Hard, And, If So, What Can You Do About It?

    A detailed, personal review of Paul McKenney’s free online textbook on parallel programming. The author, coming from a TLA⁺/distributed‑systems background, finds the early chapters excellent for building intuition about CPU caches, memory ordering, and false‑sharing, but notes gaps (e.g., shallow coverage of C++11 atomics and MESI). The review is concrete, cites specific chapters, and offers prac…

    Lobstersahelwer.ca8 minHN12757lobste.rs48
  5. Smashing the Stack for Fun and Profit

    The article explains how stack‑based buffer overflows work on x86 Linux, showing stack layout, how overwriting the saved return address can hijack control flow, and demonstrates a simple C exploit. It remains a foundational guide for understanding classic memory‑corruption attacks.

    Hall of Fameberkeley.edu31 minHN21
  6. System Design Interviews for Data Roles: What to Actually Practice

    The piece shows that data‑role system design interviews evaluate how you turn vague requirements into a defensible architecture, not which tools you name, and it gives a concrete prep framework: clarify scope, quantify load, pick batch vs streaming with trade‑offs, and address failure handling. Candidates should rehearse explaining these decisions aloud with numbers rather than just drawing diagr…

    SitePointsitepoint.com6 min