proomt

Search

Search posts, papers, and topics

All posts

CodeName OneShai Almog11 min readintermediate

Honey, I Shrunk Java

Summary

ParparVM reduces Java object headers to four bytes and rewrites the compiler to emit native C, achieving 22‑57 % lower elapsed time and up to 30 % lower peak memory versus JDK 25 across multiple platforms. The redesign also shrinks object sizes and enables faster collection loops, though allocation‑heavy workloads remain slower.

  • ParparVM's 4‑byte object header uses a 16‑bit class index plus two 8‑bit fields, cutting per‑object metadata from 8‑12 B to 4 B.
  • Self‑translation benchmarks show 22‑57 % lower elapsed time and ~10‑30 % lower peak memory versus JDK 25 on 16 hardware configurations.
  • Sequential array access runs ~2.8× faster, but allocation‑heavy workloads still take ~3× longer than HotSpot.
  • Typical objects (e.g., VarOp, ArrayList) shrink from 32 B to 24 B, and overall RSS drops ~13 MB in a one‑core container.

Java engineers building low‑footprint or embedded services, and VM designers, should care about the memory and speed gains from a four‑byte object header.

7/10

Related reading

  1. We Didn't Want to Build Another Java Server

    Codename One introduced an experimental native Java backend that compiles Java controllers to a tiny native executable via ParparVM. Benchmarks show sub‑millisecond startup, 10‑40 MiB memory, and up to 20 % higher request throughput than a Go fasthttp server.

    CodeName Onecodenameone.com15 min
  2. What Go Taught Us About Java Garbage Collection

    ParparVM’s GC was tuned by lowering the allocation‑trigger floor, adding configurable thresholds, parallel marking, mutator assistance, and proper weak/soft reference handling. These changes cut RSS from 98 MB to 38 MB, reduced worst‑case GC pauses from seconds to sub‑second, and improved cache hit rates with a recency‑based eviction policy.

    CodeName Onecodenameone.com7 min
  3. Faster Maps: Chasing Swiss Speed

    ParparVM’s HashMap suffered catastrophic miss latency due to linear probing on dense integer keys. By adopting CPython‑style perturbed probing (Swiss‑table style) and extending tagged immediate values to more primitives, miss latency dropped from 32 s to ~45 ms, allocation pressure fell dramatically, and overall performance stayed roughly flat despite a modest hit‑time slowdown.

    CodeName Onecodenameone.com8 min
  4. 1 points

    Saving another 100TB of RAM with math (and Rust)

    Cloudflare reduced the memory footprint of its Pingora Backend Router by re‑examining the consistent‑hashing implementation in the pingora‑ketama library. By increasing the number of virtual hash points per server from the default 1 to the standard 160 (and applying weighted hashing based on disk capacity), they cut the per‑node overhead enough to reclaim >100 TB of RAM across the fleet. The post…

    Hacker News front pagecloudflare.com13 minHN478120lobste.rs33
  5. Size-Specialized Memory Allocation

    Go 1.27 adds a set of span‑class‑specific malloc functions for allocations ≤ 80 bytes. By generating a tiny, constant‑size allocator per span class the runtime can inline size‑dependent work (e.g. zero‑clear) and skip span‑class lookup, yielding 20‑30 % faster small allocations and ~1 % overall speed‑up for allocation‑heavy programs. The implementation is generated automatically via an AST inline…

    The Go Bloggo.dev7 minHN315
  6. Lies, Damn Lies and Benchmarks

    Codename One engineers dissect why benchmark numbers can be misleading, then share concrete work on GC tuning, proper weak/soft references, and a new probing sequence for their open‑addressed HashMap that cuts miss‑probe counts from >16 k to ~1.5 per lookup.

    CodeName Onecodenameone.com20 min