For llama.cpp
llama.cpp on Edge, the leanest LLM you can run
Built for CPUs, so it runs on Edge today. GPU instances are coming soon.
Why llama.cpp on Edge
The smallest AI footprint that still does the job
Made for CPUs
Hand-tuned CPU kernels make quantised models practical on a normal VM. Ollama started out as a wrapper around it.GGUF, the common format
Most open-weight models ship GGUF quantisations on Hugging Face. Q4_K_M is the usual balance of size and quality.OpenAI-compatible server
llama-server exposes /v1/chat/completions with an --api-key flag. Swap the base URL in your SDK and go.Your binary, your flags
Build from source with CMake, pin a commit, and tune threads, context size and parallel slots. Nothing hidden.A small footprint
A 3B model at Q4 is about 2 GB. A Performance VM serves it with room to spare, and fits an 8B Q4 too.Weights in your bucket
Cache GGUF files in Edge Storage. New VMs copy them in with zero egress and no Hugging Face rate limits.
Reference architecture
How llama.cpp maps to Edge
Scripts and agents
OpenAI SDK, curl, LangChain
Internal tools
triage, tagging, summaries
llm.acme.com
A record → VM, anycast
llama-server :8080
s-4vcpu-8gb · Caddy for TLS
llm-weights
GGUF files, one per quant
- Compute
Runs llama-server on a Performance VM (4 vCPU, 8 GiB), with Caddy for TLS.
- Storage
Holds your GGUF files so any VM can copy them in quickly, egress-free.
- DNS
Anycast DNS points llm.acme.com straight at the VM.
- GPU Compute
Coming soon. GPU instances for CUDA builds and larger models; join the waitlist.
- CDN
Optional: add a deployment later for DDoS protection in front of a public endpoint.
Deploy
llama-server live in five steps
- 01
Build llama-server on first boot
A plain CPU build. CMake picks up the instruction sets the VM supports.
llamacpp-build.sh#!/bin/bash set -e apt-get update apt-get install -y build-essential cmake git libcurl4-openssl-dev caddy curl -fsSL https://edge.network/install.sh | sh git clone https://github.com/ggml-org/llama.cpp /opt/llama.cpp cd /opt/llama.cpp cmake -B build -DCMAKE_BUILD_TYPE=Release cmake --build build -j"$(nproc)" --target llama-server ln -s /opt/llama.cpp/build/bin/llama-server /usr/local/bin/ - 02
Create the VM and point DNS at it
A Performance VM handles 3B models comfortably and fits an 8B Q4 in memory.
shell$ edge compute scripts create --name llamacpp-build \ --file llamacpp-build.sh $ edge compute create --name llm --size s-4vcpu-8gb \ --image ubuntu-24 --region london --script llamacpp-build $ edge dns records add zone-abc123 --type A --name llm --data <vm-ip> - 03
Cache the model in Edge Storage
Download once from Hugging Face, keep it in your bucket, and copy it onto any VM that needs it.
shell$ hf download bartowski/Llama-3.2-3B-Instruct-GGUF \ --include "*Q4_K_M.gguf" --local-dir . $ edge storage cp Llama-3.2-3B-Instruct-Q4_K_M.gguf llm-weights/ # on the VM $ edge storage cp \ llm-weights/Llama-3.2-3B-Instruct-Q4_K_M.gguf /opt/models/ - 04
Run it as a service
llama-server listens on localhost and requires the API key on every request.
/etc/systemd/system/llama.service[Unit] Description=llama.cpp server After=network-online.target [Service] EnvironmentFile=/etc/llama.env # LLAMA_API_KEY=... ExecStart=/usr/local/bin/llama-server \ -m /opt/models/Llama-3.2-3B-Instruct-Q4_K_M.gguf \ --host 127.0.0.1 --port 8080 -c 8192 \ --api-key ${LLAMA_API_KEY} Restart=always [Install] WantedBy=multi-user.target - 05
Add TLS with Caddy
Caddy fetches a certificate for llm.acme.com and streams tokens as they're generated.
/etc/caddy/Caddyfilellm.acme.com { reverse_proxy 127.0.0.1:8080 { flush_interval -1 } }
Prefer to hand it off? Give the job to your AI agent or have our engineers do it.
What it costs
The cheapest path to your own LLM endpoint
- A fixed monthly price, however many tokens it generates
- No GPU needed for 1B–8B models at Q4
- Weights copied from your bucket with zero egress
- Billed hourly from a prepaid balance, with hard caps
Estimated monthly bill on Edge
Internal tools · Llama 3.2 3B Q4 on 4 vCPU
- Compute$38.47Performance VM · 4 vCPU · 8 GiB · 160 GB NVMe
- Storage$0.1515 GB of GGUF quants, first 5 GB free
- DNS$0.00Zone and A record for llm.acme.com
- Egress$0.00
FAQ
llama.cpp on Edge, answered
llama.cpp, Ollama or vLLM?
How fast is it on CPU?
Which quantisation should I use?
Can it serve several users at once?
When will GPU instances be available?
Drop-in services
Add these without touching the stack
- Edge ShieldBot protection
POST /siteverify → score: 94Score requests from public forms before they reach llama-server, so bots don't queue up behind your users.
Free for most sites. No puzzles, no widget to style.Explore - Edge AssistAI answers
<script src="…/assist.js">If you only need answers grounded in your site's content, Assist does it with one script tag and no model to run.
Free tier: 250 answered questions a monthExplore
Run LLMs the lean way
Start free on a CPU VM. Bring your GGUF, keep your prompts.
Free tiers hard-cap. Nothing bills until you add a card.