$ inference --device=local --cloud=none

Your models.
Your machine.
Your data.

Little Brother runs open LLMs entirely on your Mac and iPhone. Private by physics, not by policy.

Little Brother app icon
Little Brother chat screen

// 100% on-device

No cloud. No account. No leaks.

Little Brother runs curated MLX models directly on Apple Silicon. Your prompts, your documents, and every token generated stay on your device — there is no server to trust because there is no server at all.

Little Brother model catalog

// 32 curated models

A model catalog that fits your hardware.

Qwen, Llama, Phi, Gemma, Nemotron, SmolLM3 and more — each with on-device memory estimates so you know what runs well on your Mac or iPhone. Thinking models included for when you need deeper reasoning.

Little Brother knowledge vault

// local RAG

Chat with your documents.

Attach PDFs and ask questions. Little Brother indexes them locally with on-device embeddings and picks the right retrieval strategy — whole document, chunked, or summarized — without your files ever leaving the machine.

Little Brother chat with cited sources

// live web search

Grounded answers, cited sources.

When a question needs fresh information, models can call a live web search tool and answer with clickable citations — the search happens on demand, and only the query leaves your device.

Little Brother conversation history

// cross-model memory

Conversations that survive model swaps.

Switch from a 1B model to an 8B model mid-conversation and keep the thread. Conversation memory is shared across every model in the catalog, so your context travels with you.

// sovereign_vault_protocol

Enter the Sovereign Vault

Download Little Brother and run your first model in minutes. No account required. Nothing leaves your device.