Skip to content

Llooma

A (less) minimal LLM chat app. Run it in your browser alone, or host it for your team.

A fork of Hollama by fmaclen

The sidebar, a conversation, the Library, and the settings
Your conversations, your Library, your servers. Nothing of it leaves the machine unless you send it.

Two ways to run it

Local keeps everything in your browser, with your own provider keys. Server signs users in, stores their data per account, and keeps the API keys where a browser can never read them.

Choose a mode →

Bring your own models

Ollama, OpenAI, Claude, Infomaniak and anything OpenAI-compatible. Several at once, each with its own label and colour, so you always know which endpoint is about to be billed.

See the providers →

Documents, read in the tab

A PDF, a spreadsheet, a Word file. Parsed in your browser and never uploaded, in server mode too. Optional OCR for scans, and a vision fallback when a page is a picture of words.

How it works →

Conversations that keep fitting

/compact replaces a long history with a structured summary so it still fits in the window. Nothing is deleted, and one click puts it all back.

Read about compaction →

Six themes, each with a light and a dark ramp

Section titled “Six themes, each with a light and a dark ramp”

Classic, Dracula, Catppuccin, Gruvbox, Nord and Solarized. Pick a palette without committing to a mode, or follow the system without committing to a palette.

The six themes, alternating light and dark ramps

Llooma is a small amount of glue around a lot of other people’s work.

Not sponsored

Running Ollama on your own machine means no third party at all, and that’s what I’d recommend first. It needs a decent GPU though, and no phone has one. If that’s not realistic, someone else runs the model, and it’s worth caring who.

Infomaniak is the one I like: Swiss datacentres they own and run, open-weight models, prompts they don’t log or train on, and the GPU heat goes into warming homes nearby. Any other OpenAI-compatible provider works just as well here.