$ inference --device=local --cloud=none
Little Brother runs open LLMs entirely on your Mac and iPhone. Private by physics, not by policy.
// 100% on-device
Little Brother runs curated MLX models directly on Apple Silicon. Your prompts, your documents, and every token generated stay on your device — there is no server to trust because there is no server at all.
// 32 curated models
Qwen, Llama, Phi, Gemma, Nemotron, SmolLM3 and more — each with on-device memory estimates so you know what runs well on your Mac or iPhone. Thinking models included for when you need deeper reasoning.
// local RAG
Attach PDFs and ask questions. Little Brother indexes them locally with on-device embeddings and picks the right retrieval strategy — whole document, chunked, or summarized — without your files ever leaving the machine.
// live web search
When a question needs fresh information, models can call a live web search tool and answer with clickable citations — the search happens on demand, and only the query leaves your device.
// cross-model memory
Switch from a 1B model to an 8B model mid-conversation and keep the thread. Conversation memory is shared across every model in the catalog, so your context travels with you.
// sovereign_vault_protocol
Download Little Brother and run your first model in minutes. No account required. Nothing leaves your device.