Your local AI,
finally under control

Llama Control is a desktop control panel for llama.cpp. Manage GGUF models, chat with your AI, benchmark performance, auto-tune server flags, and browse HuggingFace, all from one beautiful interface.

Up and running in 4 steps

No command line. No config files. No hassle. From download to your first chat in under 5 minutes.

01

Download & Install

Download the latest installer from GitHub Releases. About 110 MB, a single download, no account required.

02

Point to Your Models

Open Settings and point Llama Control to your GGUF model folder. It auto-detects LM Studio and common paths. The app scans all .gguf files and displays their metadata instantly.

03

Install llama.cpp

Use the built-in llama.cpp installer in Settings to download the right prebuilt for your GPU (CUDA, Vulkan, or CPU). One click, no manual extraction needed.

04

Run & Chat

Select a model, optionally use Auto-Tune for optimized flags, and hit start. Chat with your local AI instantly, no terminal, no config files, no command line.

Six powerful tools, one interface

From model management to AI-powered auto-tuning, Llama Control gives you full command over your local AI experience.

Your GGUF models, beautifully organized

Browse local GGUF files with rich metadata parsing. Edit names, descriptions, and tags. Start or stop the server per-model with custom launch flags, and organize your collection with favorites and custom folders.

  • One-click llama-server start with per-model flags
  • GGUF metadata parser, architecture, quantization, parameters
  • Server profiles for quick flag switching
  • Real-time GPU VRAM, CPU, and RAM usage stats
  • Duplicate, export, and delete models

Local Models

4 models
Llama-3.1-8B-Instruct
Q4_K_M - 4.9 GB - 6.1 GB VRAM
Running
Mistral-7B-Instruct
Q5_K_M - 5.6 GB - 6.8 GB VRAM
Qwen2.5-72B-Instruct
Q3_K_S - 35.2 GB - 38.4 GB VRAM
Phi-3.5-mini-instruct
Q4_K_M - 2.2 GB - 3.0 GB VRAM
Chat
Explain how KV cache quantization works in llama.cpp
KV cache quantization reduces memory usage by storing key and value tensors in lower precision (e.g., Q4, Q8).
Type a message...

Chat with your local AI

Streamed token responses from your running llama-server. Manage multiple chat sessions, edit system prompts, and switch to terminal mode with xterm.js for advanced workflows.

  • Streamed real-time responses from local model
  • Multiple sessions with full history
  • Customizable system prompts
  • Terminal mode with multi-session xterm.js
  • Reasoning trace display for CoT models

Browse & download from HuggingFace

Search GGUF models directly from HuggingFace. Get hardware-aware VRAM-fit hints, check disk space before downloading, and stream files straight into your models folder with resumable downloads.

  • Search by query, category, or popularity
  • VRAM-fit hints based on your GPU
  • Disk space check before downloads
  • Resumable multi-GB downloads
  • GGUF vendor detection badges
llama-3.1 8b instruct
Llama-3.1-8B-Instruct-Q4_K_M
Q4_K_M~6.1 GB VRAMdown 12.4K
Llama-3.1-8B-Instruct-Q5_K_M
Q5_K_M~6.8 GB VRAMdown 8.7K
Llama-3.1-8B-Instruct-Q3_K_S
Q3_K_S~4.1 GB VRAMdown 3.2K
  • Llama
  • Mistral
  • Qwen
  • Gemma
  • DeepSeek
  • Nemotron

Benchmark Results

Speed: 47.3 tokens/s
Math92%
Logic88%
Code76%
Knowledge95%
Format84%
Traps71%
AI Judge Score91.2/100

Measure. Compare. Optimize.

Run llama-bench speed tests to measure tokens/second. Evaluate quality with 18 deterministic tasks covering math, logic, code, and knowledge. Grade everything with AI judges.

  • llama-bench speed benchmark (tokens/s)
  • 18 quality tasks: math, logic, code, format, knowledge, traps
  • Perplexity test (wikitext-2)
  • AI quality judging via Claude/OpenAI/Gemini
  • Multi-model comparison slots

Let AI tune your server

Forget manual flag tweaking. The AI research phase suggests optimal launch flags, then automatically benchmarks each candidate through a trial loop. KV cache quantization, chunk sizes, fit-to-VRAM, all auto-optimized.

  • AI research phase suggests optimal flag candidates
  • Automatic benchmarking trial loop
  • KV cache quantization ladder testing
  • Context and chunk size tuning
  • Fit-to-VRAM auto-calibration

Auto-Tune Pipeline

01
AI Research
Analyzes your model and GPU to suggest optimal flag candidates
02
Trial Loop
Benchmarks each candidate through automated llama-bench runs
03
Selection
Chooses the winner based on speed + quality metrics
Best config foundq8_0 | --ngl 35 | --batch 2048

Usage Analytics

247
Total Chats
1.2s
Avg Response
12.4K
Messages
2.1M
Tokens Gen
Weekly Activity

Know your AI habits

Track your AI usage patterns with a comprehensive dashboard. Monitor total chats, messages, response times, and character counts. Identify your most active days and top conversations.

  • Total chats and message counts
  • User vs. assistant message breakdown
  • Average and median response times
  • Characters generated vs. sent ratio
  • Most active day tracking

Common questions

Everything you need to know about Llama Control.

Ready to take control?

Download Llama Control for free and start managing your local AI experience today. Windows 10+, no account needed.

v0.1.60, Windows 10/11 (64-bit), ~110 MB installer