The Go BlogJunyang Shao and David Chase11 min readintermediate
Arch-specific SIMD in Go
Summary
Go 1.27 introduces experimental `arm64` and `wasm` support to the `archsimd` API, allowing Go developers to use architecture-specific SIMD instructions without writing assembly. The API prioritizes sensible naming and compiler optimizations over direct hardware mirroring, making low-level SIMD more accessible.
- The `archsimd` package provides a Go API for architecture-specific SIMD intrinsics, now supporting `amd64`, `arm64` (NEON), and `wasm` (128-bit SIMD).
- API design uses sensible names (e.g., `ShiftAllLeft`) and relies on compiler optimizations for operations like masked additions, rather than exposing raw intrinsics.
- Vectors are represented by distinct struct types (e.g., `Float32x4`), and zero-cost type reinterpretations are handled via `ToBits()` and `ReshapeToUint<W>s()`.
- Opaque `Mask` types abstract hardware-specific mask implementations, with the compiler optimizing operations like `x.Add(y).Masked(m)` into single instructions.
Engineers working on performance-critical Go applications can now leverage architecture-specific SIMD instructions for significant speedups in numerical and bit manipulation tasks without resorting to assembly.
7/10

