For vLLM
vLLM on Edge: build the serving stack now
vLLM's CPU backend works today. GPU instances are coming soon — join the waitlist.
Why vLLM on Edge
The production serving stack, ready before the GPUs are
The real vllm serve
Same CLI, same OpenAI-compatible server. What you build on CPU carries over to a GPU machine with few changes.CPU backend for small models
Practical for models of roughly 0.5B–3B parameters: staging, evals, embeddings and light internal traffic.OpenAI-compatible server
Chat, completions and embeddings endpoints. Existing SDKs work with a new base URL and the --api-key you set.Weights in object storage
Keep safetensors in an Edge Storage bucket and sync at boot. No repeated Hugging Face downloads, and zero egress.Continuous batching
vLLM batches concurrent requests as they arrive. It's the reason teams choose it, and it pays off most on GPUs.Private by default
Prompts and outputs stay on your VM. The API key and the CDN's TLS are the only way in.
Reference architecture
How vLLM maps to Edge
RAG service
OpenAI SDK · new base URL
Eval suite in CI
nightly runs against /v1
inference.acme.com
TLS; /v1/* bypasses cache
vllm serve :8000
CPU backend · s-8vcpu-16gb
llm-weights
safetensors, one prefix per model
- Compute
Runs vLLM's CPU backend on 8 vCPU and 16 GiB for small models, staging and evals.
- GPU Compute
Coming soon. GPU instances for production vLLM throughput; join the waitlist.
- Storage
Keeps model weights in a bucket so a new VM starts from your copy, egress-free.
- CDN
TLS and DDoS protection in front of the API; inference paths bypass the cache.
- DNS
Anycast DNS for inference.acme.com.
Deploy
vLLM on CPU in five steps
- 01
Write a bootstrap script
Builds vLLM with the CPU backend into a virtual environment, and installs the Edge CLI for weight syncs.
vllm-cpu.sh#!/bin/bash set -e apt-get update apt-get install -y build-essential libnuma-dev python3-dev git curl -fsSL https://edge.network/install.sh | sh curl -LsSf https://astral.sh/uv/install.sh | sh . $HOME/.local/bin/env uv venv /opt/vllm --python 3.12 && . /opt/vllm/bin/activate git clone https://github.com/vllm-project/vllm.git /opt/vllm-src cd /opt/vllm-src uv pip install -r requirements/cpu-build.txt --torch-backend cpu uv pip install -r requirements/cpu.txt --torch-backend cpu VLLM_TARGET_DEVICE=cpu uv pip install . --no-build-isolation - 02
Create the VM and mirror the model
Upload weights to your bucket once. For multi-gigabyte files the docs recommend an S3 tool with multipart uploads.
shell$ edge compute scripts create --name vllm-cpu --file vllm-cpu.sh $ edge compute create --name inference --size s-8vcpu-16gb \ --disk 160 --image ubuntu-24 --region london --script vllm-cpu $ hf download Qwen/Qwen2.5-1.5B-Instruct --local-dir qwen2.5-1.5b $ aws s3 sync qwen2.5-1.5b s3://llm-weights/qwen2.5-1.5b \ --endpoint-url https://storage.edge.network - 03
Run it as a service
Weights sync from the bucket before each start. VLLM_CPU_KVCACHE_SPACE sets the KV cache size in GiB.
/etc/systemd/system/vllm.service[Unit] Description=vLLM OpenAI-compatible server After=network-online.target [Service] EnvironmentFile=/etc/vllm.env # EDGE_API_KEY, VLLM_API_KEY Environment=VLLM_CPU_KVCACHE_SPACE=4 ExecStartPre=/usr/local/bin/edge storage sync \ llm-weights/qwen2.5-1.5b/ /opt/models/qwen2.5-1.5b/ ExecStart=/opt/vllm/bin/vllm serve /opt/models/qwen2.5-1.5b \ --served-model-name qwen2.5-1.5b \ --host 127.0.0.1 --port 8000 \ --api-key ${VLLM_API_KEY} --max-model-len 8192 Restart=always [Install] WantedBy=multi-user.target - 04
Put the endpoint behind the CDN
Proxy :443 on the VM to port 8000, and add a bypassCache rule for /v1/** in the deployment's configuration.
shell$ edge cdn create --name inference $ edge cdn domains add cdn-a1b2c3 \ --domain inference.acme.com \ --origin https://<vm-ip> - 05
Call it from your app
Any OpenAI SDK works. The same client code will point at a GPU-backed server later.
client.tsimport OpenAI from 'openai' const client = new OpenAI({ baseURL: 'https://inference.acme.com/v1', apiKey: process.env.VLLM_API_KEY, }) const res = await client.chat.completions.create({ model: 'qwen2.5-1.5b', messages: [{ role: 'user', content: 'Summarise this ticket…' }], })
Prefer to hand it off? Give the job to your AI agent or have our engineers do it.
What it costs
Staging for your inference stack, at a fixed price
- The same OpenAI-compatible API you'll run on GPUs
- Weights served from your bucket with zero egress
- Billed hourly from a prepaid balance, with hard caps
- No per-token charges at any volume
Estimated monthly bill on Edge
Staging and evals · Qwen2.5-1.5B on 8 vCPU
- Compute · vCPU$23.368 vCPU × $0.004/hr × 730 hrs
- Compute · memory$29.9016 GiB × $0.00256/GiB-hr × 730 hrs
- Compute · disk$11.84160 GiB NVMe × $0.074/GiB-mo
- Storage$0.1515 GB of weights, first 5 GB free
- CDN$0.00API traffic, inside the 500k free requests
- Egress$0.00
FAQ
vLLM on Edge, answered
Can vLLM run without a GPU?
vLLM, Ollama or llama.cpp on CPU?
When will GPU instances be available?
Which models fit on 16 GiB?
How do I scale it?
Drop-in services
Add these without touching the stack
- Edge ShieldBot protection
POST /siteverify → score: 94Score requests to any public form that feeds the model, so automated abuse never reaches the inference queue.
Free for most sites. No puzzles, no widget to style.Explore - Edge AssistAI answers
<script src="…/assist.js">For answers grounded in your own docs, Assist is a hosted alternative: one script tag, 250 free answers a month.
Free tier: 250 answered questions a monthExplore
Build your serving stack now
Start on a CPU VM, keep the OpenAI-compatible API, and join the GPU waitlist for production scale.
Free tiers hard-cap. Nothing bills until you add a card.