For Ollama
Ollama on Edge, open models on your own VM
Runs on CPU VMs today. GPU instances are coming soon — join the waitlist.
Why Ollama on Edge
An OpenAI-style API, with your data and your costs under control
Quantised models on CPU
1B–8B models at Q4 run on an 8 vCPU VM. A good fit for internal tools, agents, classification and team-scale RAG.One binary, many models
ollama pull llama3.2, qwen2.5, gemma3 or phi4-mini, then choose the model per request. Embedding models too.OpenAI-compatible API
Point the OpenAI SDK at https://ai.acme.com/v1 and your code runs unchanged: chat, completions and embeddings.Weights cached in storage
Sync the models directory to Edge Storage once. A new VM pulls from your bucket with zero egress, not from the registry.Prompts stay private
Customer prompts and responses never leave your VM. Useful for regulated data and anything under NDA.No per-token bill
Hosted APIs charge per million tokens. A VM is a fixed hourly price, however much it generates.
Reference architecture
How Ollama maps to Edge
Your app and agents
OpenAI SDK, baseURL /v1
n8n, Open WebUI
same endpoint, own keys
ai.acme.com
anycast DNS
TLS in front of /v1
completions bypass the cache
ollama serve :11434
s-8vcpu-16gb · Nginx key check
llm-weights
ollama/ blobs and manifests
- Compute
Runs Ollama and Nginx on an 8 vCPU, 16 GiB VM: enough for quantised models up to about 8B.
- GPU Compute
Coming soon. GPU instances for larger models and higher throughput; join the waitlist.
- Storage
Caches model blobs so new or rebuilt VMs pull weights from your bucket, egress-free.
- CDN
Terminates TLS and adds DDoS protection in front of the API; completions bypass the cache.
- DNS
Anycast DNS for ai.acme.com.
- Shield
Scores the public chat form in front of your model, so bots don't burn CPU time.
Deploy
Ollama behind an API key in five steps
- 01
Write a bootstrap script
Installs Ollama and the Edge CLI, keeps Ollama on localhost and stores models where the sync step expects them.
ollama-setup.sh#!/bin/bash set -e curl -fsSL https://ollama.com/install.sh | sh curl -fsSL https://edge.network/install.sh | sh mkdir -p /var/lib/ollama/models chown -R ollama:ollama /var/lib/ollama mkdir -p /etc/systemd/system/ollama.service.d cat > /etc/systemd/system/ollama.service.d/edge.conf <<'EOF' [Service] Environment="OLLAMA_HOST=127.0.0.1:11434" Environment="OLLAMA_MODELS=/var/lib/ollama/models" Environment="OLLAMA_KEEP_ALIVE=24h" EOF systemctl daemon-reload && systemctl restart ollama apt-get install -y nginx - 02
Create the VM
8 vCPU and 16 GiB leave room for an 8B Q4 model plus an embedding model. --disk 160 gives space for several models.
shell$ edge compute scripts create --name ollama-setup \ --file ollama-setup.sh $ edge compute create --name llm --size s-8vcpu-16gb --disk 160 \ --image ubuntu-24 --region london --script ollama-setup - 03
Pull models and cache them
Pull once, then sync to your bucket. A replacement VM syncs back before Ollama starts.
shell$ ollama pull llama3.2:3b $ ollama pull nomic-embed-text $ edge storage sync /var/lib/ollama/models llm-weights/ollama/ # on a new VM, before starting Ollama $ edge storage sync llm-weights/ollama/ /var/lib/ollama/models - 04
Check the key in Nginx
OpenAI SDKs send the key as a Bearer token. Buffering is off so tokens stream as they're generated.
/etc/nginx/sites-enabled/ollamaserver { listen 443 ssl; server_name ai.acme.com; ssl_certificate /etc/ssl/origin.pem; # origin certificate ssl_certificate_key /etc/ssl/origin.key; location /v1/ { if ($http_authorization != "Bearer sk-acme-change-me") { return 401; } proxy_pass http://127.0.0.1:11434; proxy_buffering off; proxy_read_timeout 300s; } } - 05
Put the endpoint behind the CDN
Add a bypassCache rule for /v1/** in the deployment's configuration, then call it like any OpenAI endpoint.
shell$ edge cdn create --name ai $ edge cdn domains add cdn-a1b2c3 \ --domain ai.acme.com --origin https://<vm-ip> $ curl https://ai.acme.com/v1/chat/completions \ -H "Authorization: Bearer $OLLAMA_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "llama3.2:3b", "messages": [{"role": "user", "content": "Hi"}]}'
Prefer to hand it off? Give the job to your AI agent or have our engineers do it.
What it costs
A fixed price for your own model endpoint
- A fixed hourly price, whatever you generate
- Weights pulled from your bucket with zero egress
- Billed hourly from a prepaid balance, with hard caps
- Move to a GPU instance when they launch
Estimated monthly bill on Edge
Team assistant · llama3.2:3b on 8 vCPU · weights cached
- Compute · vCPU$23.368 vCPU × $0.004/hr × 730 hrs
- Compute · memory$29.9016 GiB × $0.00256/GiB-hr × 730 hrs
- Compute · disk$11.84160 GiB NVMe × $0.074/GiB-mo
- Storage$0.3025 GB of model weights, first 5 GB free
- CDN$0.00API traffic, inside the 500k free requests
- DNS$0.00Zone and records for ai.acme.com
- Egress$0.00
FAQ
Ollama on Edge, answered
Can Ollama really run without a GPU?
Which model should I start with?
When will GPU instances be available?
How does this compare to OpenAI or other hosted APIs?
How do I secure the API?
How do I scale beyond one VM?
Drop-in services
Add these without touching the stack
- Edge ShieldBot protection
POST /siteverify → score: 94Score each message to a public chatbot before it reaches Ollama, so scripted abuse doesn't eat your CPU time.
Free for most sites. No puzzles, no widget to style.Explore - Edge AssistAI answers
<script src="…/assist.js">Need answers from your own docs without running a model? Assist is one script tag, with 250 free answers a month.
Free tier: 250 answered questions a monthExplore
Run open models on your terms
Start on a CPU VM today, and join the GPU waitlist for bigger models.
Free tiers hard-cap. Nothing bills until you add a card.