Skip to content

For Ollama

Ollama on Edge, open models on your own VM

Pull a quantised Llama, Qwen or Gemma model onto a CPU VM and serve it through Ollama's OpenAI-compatible API. Prompts stay on your infrastructure, weights stay in your bucket, and there's no per-token bill.

Runs on CPU VMs today. GPU instances are coming soon — join the waitlist.

Why Ollama on Edge

An OpenAI-style API, with your data and your costs under control

Ollama makes open models easy to run. On an Edge CPU VM it's a practical fit for small quantised models, internal tools and anything that shouldn't leave your network.
  • Quantised models on CPU

    1B–8B models at Q4 run on an 8 vCPU VM. A good fit for internal tools, agents, classification and team-scale RAG.
  • One binary, many models

    ollama pull llama3.2, qwen2.5, gemma3 or phi4-mini, then choose the model per request. Embedding models too.
  • OpenAI-compatible API

    Point the OpenAI SDK at https://ai.acme.com/v1 and your code runs unchanged: chat, completions and embeddings.
  • Weights cached in storage

    Sync the models directory to Edge Storage once. A new VM pulls from your bucket with zero egress, not from the registry.
  • Prompts stay private

    Customer prompts and responses never leave your VM. Useful for regulated data and anything under NDA.
  • No per-token bill

    Hosted APIs charge per million tokens. A VM is a fixed hourly price, however much it generates.

Reference architecture

How Ollama maps to Edge

One CPU VM runs Ollama behind Nginx, which checks an API key. Model blobs are cached in a bucket, so a replacement VM is serving again without re-downloading from the registry.
  • Compute

    Runs Ollama and Nginx on an 8 vCPU, 16 GiB VM: enough for quantised models up to about 8B.

  • GPU Compute

    Coming soon. GPU instances for larger models and higher throughput; join the waitlist.

  • Storage

    Caches model blobs so new or rebuilt VMs pull weights from your bucket, egress-free.

  • CDN

    Terminates TLS and adds DDoS protection in front of the API; completions bypass the cache.

  • DNS

    Anycast DNS for ai.acme.com.

  • Shield

    Scores the public chat form in front of your model, so bots don't burn CPU time.

Deploy

Ollama behind an API key in five steps

Ollama has no authentication of its own, so it listens on localhost and Nginx checks a key in front of it.
  1. 01

    Write a bootstrap script

    Installs Ollama and the Edge CLI, keeps Ollama on localhost and stores models where the sync step expects them.

    ollama-setup.sh
    #!/bin/bash
    set -e
    curl -fsSL https://ollama.com/install.sh | sh
    curl -fsSL https://edge.network/install.sh | sh
    mkdir -p /var/lib/ollama/models
    chown -R ollama:ollama /var/lib/ollama
    mkdir -p /etc/systemd/system/ollama.service.d
    cat > /etc/systemd/system/ollama.service.d/edge.conf <<'EOF'
    [Service]
    Environment="OLLAMA_HOST=127.0.0.1:11434"
    Environment="OLLAMA_MODELS=/var/lib/ollama/models"
    Environment="OLLAMA_KEEP_ALIVE=24h"
    EOF
    systemctl daemon-reload && systemctl restart ollama
    apt-get install -y nginx
  2. 02

    Create the VM

    8 vCPU and 16 GiB leave room for an 8B Q4 model plus an embedding model. --disk 160 gives space for several models.

    shell
    $ edge compute scripts create --name ollama-setup \
        --file ollama-setup.sh
    $ edge compute create --name llm --size s-8vcpu-16gb --disk 160 \
        --image ubuntu-24 --region london --script ollama-setup
  3. 03

    Pull models and cache them

    Pull once, then sync to your bucket. A replacement VM syncs back before Ollama starts.

    shell
    $ ollama pull llama3.2:3b
    $ ollama pull nomic-embed-text
    $ edge storage sync /var/lib/ollama/models llm-weights/ollama/
    
    # on a new VM, before starting Ollama
    $ edge storage sync llm-weights/ollama/ /var/lib/ollama/models
  4. 04

    Check the key in Nginx

    OpenAI SDKs send the key as a Bearer token. Buffering is off so tokens stream as they're generated.

    /etc/nginx/sites-enabled/ollama
    server {
      listen 443 ssl;
      server_name ai.acme.com;
      ssl_certificate     /etc/ssl/origin.pem;   # origin certificate
      ssl_certificate_key /etc/ssl/origin.key;
    
      location /v1/ {
        if ($http_authorization != "Bearer sk-acme-change-me") {
          return 401;
        }
        proxy_pass http://127.0.0.1:11434;
        proxy_buffering off;
        proxy_read_timeout 300s;
      }
    }
  5. 05

    Put the endpoint behind the CDN

    Add a bypassCache rule for /v1/** in the deployment's configuration, then call it like any OpenAI endpoint.

    shell
    $ edge cdn create --name ai
    $ edge cdn domains add cdn-a1b2c3 \
        --domain ai.acme.com --origin https://<vm-ip>
    
    $ curl https://ai.acme.com/v1/chat/completions \
        -H "Authorization: Bearer $OLLAMA_KEY" \
        -H "Content-Type: application/json" \
        -d '{"model": "llama3.2:3b",
             "messages": [{"role": "user", "content": "Hi"}]}'

Prefer to hand it off? Give the job to your AI agent or have our engineers do it.

What it costs

A fixed price for your own model endpoint

No GPU instance is priced here because none is available yet. This is an 8 vCPU VM, priced from the published per-resource rates, plus the bucket that holds your weights.
  • A fixed hourly price, whatever you generate
  • Weights pulled from your bucket with zero egress
  • Billed hourly from a prepaid balance, with hard caps
  • Move to a GPU instance when they launch
See compute pricing

Estimated monthly bill on Edge

Team assistant · llama3.2:3b on 8 vCPU · weights cached

USD
  • Compute · vCPU$23.368 vCPU × $0.004/hr × 730 hrs
  • Compute · memory$29.9016 GiB × $0.00256/GiB-hr × 730 hrs
  • Compute · disk$11.84160 GiB NVMe × $0.074/GiB-mo
  • Storage$0.3025 GB of model weights, first 5 GB free
  • CDN$0.00API traffic, inside the 500k free requests
  • DNS$0.00Zone and records for ai.acme.com
  • Egress$0.00
Total$65.40

FAQ

Ollama on Edge, answered

Something else? Ask an engineer.
Can Ollama really run without a GPU?
Yes. Ollama runs on CPU out of the box. Small quantised models (1B–3B at Q4) are usable for interactive work; 7B–8B models run, but slower. Expect single-digit to low-double-digit tokens per second, not hosted-API speeds.
Which model should I start with?
llama3.2:3b or qwen2.5:3b for chat and tool use, a 7B–8B model at Q4 if you need better answers and can accept slower output, and nomic-embed-text for embeddings.
When will GPU instances be available?
GPU instances are coming soon. Join the waitlist at /compute/gpus to hear first. If you need GPU capacity now, contact us: we may be able to arrange early access or an interim option.
How does this compare to OpenAI or other hosted APIs?
Your prompts stay on your infrastructure and the price doesn't scale with tokens. The trade-off: small open models on CPU are slower and less capable than frontier hosted models, so pick them for tasks they do well.
How do I secure the API?
Ollama has no built-in authentication, so never expose port 11434. Keep OLLAMA_HOST on 127.0.0.1, check a key in Nginx as in step four, and let the CDN handle public TLS.
How do I scale beyond one VM?
Ollama holds no state between requests, so add VMs from the same bootstrap script, sync weights from your bucket, and balance across them with Nginx or Caddy.

Run open models on your terms

Start on a CPU VM today, and join the GPU waitlist for bigger models.

Free tiers hard-cap. Nothing bills until you add a card.