Llama icon

Llama

A cosy home for your LLMs

Free LLM Server Open Source

Overview

Llama (formerly LlamaBarn) is a lightweight native macOS menu bar app from the llama.cpp team that makes running local large language models effortless. Built in Swift and weighing only 4 MB, it offers a curated model catalog filtered to what your Mac can actually run, one-click installs from Hugging Face, and an OpenAI-compatible API at localhost:8080/v1 for use from chat apps, editors, and scripts. Models load on demand and unload when idle, and it reuses the Hugging Face cache so models are shared with other llama.cpp tools.

Architecture: Apple Silicon

Key Features

  • Native macOS menu bar app built in Swift - only 4 MB
  • Curated model catalog filtered to what your Mac can run
  • One-click model installation from Hugging Face
  • OpenAI-compatible API at localhost:8080/v1
  • Built-in web UI for chatting with models directly
  • Models load on demand and unload when idle
  • Reuses the Hugging Face cache, so models are shared with other llama.cpp tools
  • Install via Homebrew (brew install --cask llama-app) or direct download
  • MIT licensed and fully open source

Tags

chattext generation