Icon for Atomic Chat

Atomic Chat

Open Source Free Trial

Open-source local AI chat app for running open-weight models on desktop and mobile

Atomic Chat is an open-source app for running open-weight LLMs locally, on macOS, Windows, Linux, iOS, and Android. It runs inference on-device with no account or sign-up, and the offline mode keeps requests loopback-only so nothing leaves the machine. Models like Llama, Qwen, Gemma, Mistral, and Phi load from Hugging Face, and it exposes an OpenAI-compatible endpoint at localhost:1337/v1 so existing tooling can point at it. Apache-2.0 licensed.

Pricing: Free

Hosting Self-hosted
HQ 🇪🇪 Estonia
License APACHE-2.0
Screenshot of Atomic Chat webpage

Atomic Chat sits in the same space as Ollama, Jan, GPT4All, and LM Studio: a local runtime that downloads open-weight models and serves them on-device, with a chat UI on top. What sets it apart is coverage down to mobile (iOS and Android in addition to the three desktop OSes) and an OpenAI-compatible server at localhost:1337/v1 that lets agents and existing OpenAI clients target it directly. It can connect to MCP servers for tools, file access, and web search.

The headline feature is TurboQuant, a 3-bit quantization path. Atomic Chat's site claims up to 8x faster inference, 6x less memory, and zero accuracy loss; the project's own repository frames the gains more narrowly (for example a smaller KV-cache footprint and per-model throughput boosts on llama.cpp), so treat the marketing figures as vendor claims rather than measured results and benchmark on your own models before relying on them. Pricing is free with no sign-up, and the source is Apache-2.0, so the claims are inspectable. Built by SPACESHIPINTELLIGENCE OU, based in Tallinn, Estonia.

Work on Atomic Chat? Feature it at the top of Frameworks & Stacks.

Is your product missing?

Add it here →