Offline local inference on macOS 0 ▲ Universal Bits 2 hours ago · Tech · hide · 0 comments Local inference lets me run agentic workflows on sensitive data (e.g., financial or health records) while keeping both the computation and the data offline. That’s my number one reason for it: confidentiality. Ensuring that (rogue) agents don’t wander around is a hot topic. My setup pairs: Sandboxed inference on macOS, which uses seatbelt to harden the runtime and prevent remote connections. An offline Apple container that bundles pi and a few useful tools in a NixOS image. Zebra stripes represent sandboxing or isolation. Threat modeling llama-cpp runs under seatbelt with read-only access to the model files and no internet access. Models and chat templates are downloaded beforehand. The sandbox adds a level of defense in case of vulnerabilities in the inference engine (e.g., in Jinja template parsing). pi runs in an Apple container with DNS and external networking disabled. It communicates with llama-cpp over a Unix socket mounted within the container. A small relay exposes the Unix… No comments yet. Log in to reply on the Fediverse. Comments will appear here.