Skip to content

For vLLM

vLLM on Edge: build the serving stack now

Run vllm serve on an Edge CPU VM for small models, staging and evals, with weights cached in your bucket and the OpenAI-compatible API behind TLS. Take the same setup to a GPU instance when they launch.

vLLM's CPU backend works today. GPU instances are coming soon — join the waitlist.

Why vLLM on Edge

The production serving stack, ready before the GPUs are

vLLM is built for GPUs, and Edge GPU instances are on the way. Its CPU backend lets you stand up the real server, API and weight pipeline today, then scale the hardware later.
  • The real vllm serve

    Same CLI, same OpenAI-compatible server. What you build on CPU carries over to a GPU machine with few changes.
  • CPU backend for small models

    Practical for models of roughly 0.5B–3B parameters: staging, evals, embeddings and light internal traffic.
  • OpenAI-compatible server

    Chat, completions and embeddings endpoints. Existing SDKs work with a new base URL and the --api-key you set.
  • Weights in object storage

    Keep safetensors in an Edge Storage bucket and sync at boot. No repeated Hugging Face downloads, and zero egress.
  • Continuous batching

    vLLM batches concurrent requests as they arrive. It's the reason teams choose it, and it pays off most on GPUs.
  • Private by default

    Prompts and outputs stay on your VM. The API key and the CDN's TLS are the only way in.

Reference architecture

How vLLM maps to Edge

One VM runs vllm serve with the CPU backend. Weights sync from a bucket at boot, and the CDN puts TLS in front of an endpoint that already speaks the OpenAI API.
  • Compute

    Runs vLLM's CPU backend on 8 vCPU and 16 GiB for small models, staging and evals.

  • GPU Compute

    Coming soon. GPU instances for production vLLM throughput; join the waitlist.

  • Storage

    Keeps model weights in a bucket so a new VM starts from your copy, egress-free.

  • CDN

    TLS and DDoS protection in front of the API; inference paths bypass the cache.

  • DNS

    Anycast DNS for inference.acme.com.

Deploy

vLLM on CPU in five steps

Follows vLLM's build-from-source steps for x86 CPUs. Check docs.vllm.ai for the current release before you pin a version.
  1. 01

    Write a bootstrap script

    Builds vLLM with the CPU backend into a virtual environment, and installs the Edge CLI for weight syncs.

    vllm-cpu.sh
    #!/bin/bash
    set -e
    apt-get update
    apt-get install -y build-essential libnuma-dev python3-dev git
    curl -fsSL https://edge.network/install.sh | sh
    curl -LsSf https://astral.sh/uv/install.sh | sh
    . $HOME/.local/bin/env
    uv venv /opt/vllm --python 3.12 && . /opt/vllm/bin/activate
    git clone https://github.com/vllm-project/vllm.git /opt/vllm-src
    cd /opt/vllm-src
    uv pip install -r requirements/cpu-build.txt --torch-backend cpu
    uv pip install -r requirements/cpu.txt --torch-backend cpu
    VLLM_TARGET_DEVICE=cpu uv pip install . --no-build-isolation
  2. 02

    Create the VM and mirror the model

    Upload weights to your bucket once. For multi-gigabyte files the docs recommend an S3 tool with multipart uploads.

    shell
    $ edge compute scripts create --name vllm-cpu --file vllm-cpu.sh
    $ edge compute create --name inference --size s-8vcpu-16gb \
        --disk 160 --image ubuntu-24 --region london --script vllm-cpu
    
    $ hf download Qwen/Qwen2.5-1.5B-Instruct --local-dir qwen2.5-1.5b
    $ aws s3 sync qwen2.5-1.5b s3://llm-weights/qwen2.5-1.5b \
        --endpoint-url https://storage.edge.network
  3. 03

    Run it as a service

    Weights sync from the bucket before each start. VLLM_CPU_KVCACHE_SPACE sets the KV cache size in GiB.

    /etc/systemd/system/vllm.service
    [Unit]
    Description=vLLM OpenAI-compatible server
    After=network-online.target
    
    [Service]
    EnvironmentFile=/etc/vllm.env   # EDGE_API_KEY, VLLM_API_KEY
    Environment=VLLM_CPU_KVCACHE_SPACE=4
    ExecStartPre=/usr/local/bin/edge storage sync \
      llm-weights/qwen2.5-1.5b/ /opt/models/qwen2.5-1.5b/
    ExecStart=/opt/vllm/bin/vllm serve /opt/models/qwen2.5-1.5b \
      --served-model-name qwen2.5-1.5b \
      --host 127.0.0.1 --port 8000 \
      --api-key ${VLLM_API_KEY} --max-model-len 8192
    Restart=always
    
    [Install]
    WantedBy=multi-user.target
  4. 04

    Put the endpoint behind the CDN

    Proxy :443 on the VM to port 8000, and add a bypassCache rule for /v1/** in the deployment's configuration.

    shell
    $ edge cdn create --name inference
    $ edge cdn domains add cdn-a1b2c3 \
        --domain inference.acme.com \
        --origin https://<vm-ip>
  5. 05

    Call it from your app

    Any OpenAI SDK works. The same client code will point at a GPU-backed server later.

    client.ts
    import OpenAI from 'openai'
    
    const client = new OpenAI({
      baseURL: 'https://inference.acme.com/v1',
      apiKey: process.env.VLLM_API_KEY,
    })
    
    const res = await client.chat.completions.create({
      model: 'qwen2.5-1.5b',
      messages: [{ role: 'user', content: 'Summarise this ticket…' }],
    })

Prefer to hand it off? Give the job to your AI agent or have our engineers do it.

What it costs

Staging for your inference stack, at a fixed price

No GPU instance is priced here because none is available yet. This is an 8 vCPU VM from the published per-resource rates, plus the weights bucket.
  • The same OpenAI-compatible API you'll run on GPUs
  • Weights served from your bucket with zero egress
  • Billed hourly from a prepaid balance, with hard caps
  • No per-token charges at any volume
See compute pricing

Estimated monthly bill on Edge

Staging and evals · Qwen2.5-1.5B on 8 vCPU

USD
  • Compute · vCPU$23.368 vCPU × $0.004/hr × 730 hrs
  • Compute · memory$29.9016 GiB × $0.00256/GiB-hr × 730 hrs
  • Compute · disk$11.84160 GiB NVMe × $0.074/GiB-mo
  • Storage$0.1515 GB of weights, first 5 GB free
  • CDN$0.00API traffic, inside the 500k free requests
  • Egress$0.00
Total$65.25

FAQ

vLLM on Edge, answered

Something else? Ask an engineer.
Can vLLM run without a GPU?
Yes. vLLM has a CPU backend for x86, built from source as in step one. It runs noticeably faster on CPUs with AVX-512; check lscpu on your VM. It's best suited to small models.
vLLM, Ollama or llama.cpp on CPU?
On CPU alone, llama.cpp or Ollama with a Q4 GGUF model is usually faster and lighter. Choose vLLM on CPU when you're building towards GPU serving and want the same server, flags and API from day one.
When will GPU instances be available?
GPU instances are coming soon. Join the waitlist at /compute/gpus to hear first. If you need GPU capacity now, contact us: we may be able to arrange early access or an interim option.
Which models fit on 16 GiB?
Models of roughly 0.5B–3B parameters in bf16 fit with room for the KV cache, which VLLM_CPU_KVCACHE_SPACE controls. A 7B model in bf16 needs around 15 GB for weights alone, so it won't fit comfortably.
How do I scale it?
Add VMs from the same bootstrap script and balance across them. For real throughput and larger models, vLLM needs GPUs: join the waitlist and we'll let you know when instances launch.

Build your serving stack now

Start on a CPU VM, keep the OpenAI-compatible API, and join the GPU waitlist for production scale.

Free tiers hard-cap. Nothing bills until you add a card.