Skip to content

For llama.cpp

llama.cpp on Edge, the leanest LLM you can run

Build llama-server, load a Q4 GGUF from your bucket and serve an OpenAI-compatible API from a 4 vCPU VM. No GPU, no framework stack, no per-token bill.

Built for CPUs, so it runs on Edge today. GPU instances are coming soon.

Why llama.cpp on Edge

The smallest AI footprint that still does the job

llama.cpp is written to make quantised models practical on ordinary CPUs, which makes it the natural fit for Edge VMs today.
  • Made for CPUs

    Hand-tuned CPU kernels make quantised models practical on a normal VM. Ollama started out as a wrapper around it.
  • GGUF, the common format

    Most open-weight models ship GGUF quantisations on Hugging Face. Q4_K_M is the usual balance of size and quality.
  • OpenAI-compatible server

    llama-server exposes /v1/chat/completions with an --api-key flag. Swap the base URL in your SDK and go.
  • Your binary, your flags

    Build from source with CMake, pin a commit, and tune threads, context size and parallel slots. Nothing hidden.
  • A small footprint

    A 3B model at Q4 is about 2 GB. A Performance VM serves it with room to spare, and fits an 8B Q4 too.
  • Weights in your bucket

    Cache GGUF files in Edge Storage. New VMs copy them in with zero egress and no Hugging Face rate limits.

Reference architecture

How llama.cpp maps to Edge

The smallest serious LLM setup: one VM, one binary, one GGUF file. Caddy on the VM handles TLS and DNS points straight at it, with no CDN needed.
  • Compute

    Runs llama-server on a Performance VM (4 vCPU, 8 GiB), with Caddy for TLS.

  • Storage

    Holds your GGUF files so any VM can copy them in quickly, egress-free.

  • DNS

    Anycast DNS points llm.acme.com straight at the VM.

  • GPU Compute

    Coming soon. GPU instances for CUDA builds and larger models; join the waitlist.

  • CDN

    Optional: add a deployment later for DDoS protection in front of a public endpoint.

Deploy

llama-server live in five steps

A CPU build, one GGUF file, a systemd unit and Caddy. Everything else is optional.
  1. 01

    Build llama-server on first boot

    A plain CPU build. CMake picks up the instruction sets the VM supports.

    llamacpp-build.sh
    #!/bin/bash
    set -e
    apt-get update
    apt-get install -y build-essential cmake git libcurl4-openssl-dev caddy
    curl -fsSL https://edge.network/install.sh | sh
    git clone https://github.com/ggml-org/llama.cpp /opt/llama.cpp
    cd /opt/llama.cpp
    cmake -B build -DCMAKE_BUILD_TYPE=Release
    cmake --build build -j"$(nproc)" --target llama-server
    ln -s /opt/llama.cpp/build/bin/llama-server /usr/local/bin/
  2. 02

    Create the VM and point DNS at it

    A Performance VM handles 3B models comfortably and fits an 8B Q4 in memory.

    shell
    $ edge compute scripts create --name llamacpp-build \
        --file llamacpp-build.sh
    $ edge compute create --name llm --size s-4vcpu-8gb \
        --image ubuntu-24 --region london --script llamacpp-build
    $ edge dns records add zone-abc123 --type A --name llm --data <vm-ip>
  3. 03

    Cache the model in Edge Storage

    Download once from Hugging Face, keep it in your bucket, and copy it onto any VM that needs it.

    shell
    $ hf download bartowski/Llama-3.2-3B-Instruct-GGUF \
        --include "*Q4_K_M.gguf" --local-dir .
    $ edge storage cp Llama-3.2-3B-Instruct-Q4_K_M.gguf llm-weights/
    
    # on the VM
    $ edge storage cp \
        llm-weights/Llama-3.2-3B-Instruct-Q4_K_M.gguf /opt/models/
  4. 04

    Run it as a service

    llama-server listens on localhost and requires the API key on every request.

    /etc/systemd/system/llama.service
    [Unit]
    Description=llama.cpp server
    After=network-online.target
    
    [Service]
    EnvironmentFile=/etc/llama.env   # LLAMA_API_KEY=...
    ExecStart=/usr/local/bin/llama-server \
      -m /opt/models/Llama-3.2-3B-Instruct-Q4_K_M.gguf \
      --host 127.0.0.1 --port 8080 -c 8192 \
      --api-key ${LLAMA_API_KEY}
    Restart=always
    
    [Install]
    WantedBy=multi-user.target
  5. 05

    Add TLS with Caddy

    Caddy fetches a certificate for llm.acme.com and streams tokens as they're generated.

    /etc/caddy/Caddyfile
    llm.acme.com {
      reverse_proxy 127.0.0.1:8080 {
        flush_interval -1
      }
    }

Prefer to hand it off? Give the job to your AI agent or have our engineers do it.

What it costs

The cheapest path to your own LLM endpoint

One Performance VM, a few gigabytes of GGUF files and free DNS. Nothing on the bill counts tokens.
  • A fixed monthly price, however many tokens it generates
  • No GPU needed for 1B–8B models at Q4
  • Weights copied from your bucket with zero egress
  • Billed hourly from a prepaid balance, with hard caps
See compute pricing

Estimated monthly bill on Edge

Internal tools · Llama 3.2 3B Q4 on 4 vCPU

USD
  • Compute$38.47Performance VM · 4 vCPU · 8 GiB · 160 GB NVMe
  • Storage$0.1515 GB of GGUF quants, first 5 GB free
  • DNS$0.00Zone and A record for llm.acme.com
  • Egress$0.00
Total$38.62

FAQ

llama.cpp on Edge, answered

Something else? Ask an engineer.
llama.cpp, Ollama or vLLM?
Choose llama.cpp for full control of the binary and the smallest footprint. Ollama wraps the same ideas in a friendlier model manager. vLLM is built for high-throughput GPU serving; on CPU alone it's usually slower.
How fast is it on CPU?
It depends on model size and memory bandwidth. On a 4 vCPU VM, expect single-digit to around ten tokens per second for a 3B Q4 model, and slower for 8B. Measure with llama-bench before you commit.
Which quantisation should I use?
Q4_K_M is the usual starting point. Q5_K_M or Q6_K keep more fidelity at a larger size and lower speed; Q3 variants are smaller still. Test on your own prompts.
Can it serve several users at once?
Yes. llama-server has parallel slots (-np) and batches requests across them, splitting the context between slots. On CPU the total throughput is still limited, so it suits a team rather than a public product.
When will GPU instances be available?
GPU instances are coming soon. Join the waitlist at /compute/gpus. When they launch, rebuild with -DGGML_CUDA=ON and add -ngl to offload layers to the GPU.

Run LLMs the lean way

Start free on a CPU VM. Bring your GGUF, keep your prompts.

Free tiers hard-cap. Nothing bills until you add a card.