proomt

Search

Search posts, papers, and topics

All posts

The Go BlogJunyang Shao and David Chase11 min readintermediate

Arch-specific SIMD in Go

Summary

Go 1.27 introduces experimental `arm64` and `wasm` support to the `archsimd` API, allowing Go developers to use architecture-specific SIMD instructions without writing assembly. The API prioritizes sensible naming and compiler optimizations over direct hardware mirroring, making low-level SIMD more accessible.

  • The `archsimd` package provides a Go API for architecture-specific SIMD intrinsics, now supporting `amd64`, `arm64` (NEON), and `wasm` (128-bit SIMD).
  • API design uses sensible names (e.g., `ShiftAllLeft`) and relies on compiler optimizations for operations like masked additions, rather than exposing raw intrinsics.
  • Vectors are represented by distinct struct types (e.g., `Float32x4`), and zero-cost type reinterpretations are handled via `ToBits()` and `ReshapeToUint<W>s()`.
  • Opaque `Mask` types abstract hardware-specific mask implementations, with the compiler optimizing operations like `x.Add(y).Masked(m)` into single instructions.

Engineers working on performance-critical Go applications can now leverage architecture-specific SIMD instructions for significant speedups in numerical and bit manipulation tasks without resorting to assembly.

7/10

Related reading

  1. Size-Specialized Memory Allocation

    Go 1.27 adds a set of span‑class‑specific malloc functions for allocations ≤ 80 bytes. By generating a tiny, constant‑size allocator per span class the runtime can inline size‑dependent work (e.g. zero‑clear) and skip span‑class lookup, yielding 20‑30 % faster small allocations and ~1 % overall speed‑up for allocation‑heavy programs. The implementation is generated automatically via an AST inline…

    The Go Bloggo.dev7 minHN315
  2. Fearless SIMD v1.0 is here

    Fearless SIMD v1.0 is a Rust library that enables safe, portable, and performant SIMD programming by eliminating `unsafe` blocks for most operations. It provides abstractions for autovectorization, multiversioning, and safe access to intrinsics, ensuring performance without compromising memory safety.

    Lobsterslinebender.org4 minHN29147lobste.rs66
  3. When to choose x86-64 vs aarch64

    When provisioning PlanetScale Postgres clusters you must pick either x86‑64 or aarch64, because the two architectures differ in core counting, clock speed, SIMD width, and on‑disk file compatibility. ARM gives more cores per vCPU and lower power, which helps many small queries, while x86‑64 offers higher single‑core performance and wider vector instructions, which helps heavy queries and certain…

    PlanetScaleplanetscale.com6 min