Coming Soon

Every byte for inference.

Kelvin is a minimal, immutable Linux distribution that turns Apple Silicon hardware into a private AI inference appliance. Plug in, run models. Nothing else.

macOS holds 2-5 GB of non-evictable wired memory hostage and caps GPU memory at 75%. Kelvin reclaims all of it. The result: fit models and context lengths that simply don't fit under macOS, on hardware you already own. Need more? Buy a second Mac — Kelvin clusters them automatically over Thunderbolt, zero config, and runs models too big for any single machine.

Built on Linux, not from scratch. For inference, the OS isn't on the performance-critical path — the silicon is. We keep the mature GPU compiler, networking stack, and memory subsystem. We hack the memory layout where it matters (huge pages, contiguous reservations, TLB-friendly streaming) and spend our effort on what actually moves the needle: the engine and clustering.