11–16× Faster LLMs in macOS VMs on Apple Silicon
macOS VMs on Apple Silicon can run LLM inference far faster when the guest reports conservative Metal capabilities. A process-scoped capability shim steers llama.cpp onto newer Metal kernels, yielding reported 11–16× speedups while keeping the same Virtualization.framework GPU path.