Skip to content

AI & machine learning

Run open models on infrastructure you control

Serve quantised open models from your own VMs, keep weights and datasets in S3-compatible storage, and add cited AI answers to your site. Prompts stay on your machines, and there's no per-token bill.

GPU instances are coming soon. Everything else on this page runs today.

The problem

AI features shouldn't come with a meter running

Most teams start on a hosted model API. It's quick until the invoice, the privacy review or the data transfer bill arrives.
  1. 01

    Every prompt and document goes to someone else's servers

    Run the model on your own VM, on a private network. Prompts and completions never pass through a third-party API.

  2. 02

    Per-token pricing punishes the features people actually use

    A VM bills by the hour, whatever it serves. Busy or quiet, the cost is the same, and you can stop it when it's idle.

  3. 03

    Moving weights, datasets and outputs costs egress every time

    Storage and CDN egress is $0. Pull a 5 GB model to every VM or serve generated images to millions without a transfer line.

  4. 04

    Not every AI feature needs a model of your own

    Edge Assist adds cited answers from your own content with one script tag, at a penny a question after 250 free each month.

How it works

The pieces of a private AI stack, available today

Standard tools on standard infrastructure: llama.cpp or Ollama on a VM, weights in a bucket, results through the CDN.

Compute

Small open models run well on a CPU VM

Quantised 3B–8B models in GGUF format serve comfortably from RAM on an ordinary VM. Pick the exact vCPU and memory you need, expose an OpenAI-compatible endpoint, and point your existing SDK at it.
  • llama.cpp and Ollama, with an OpenAI-compatible /v1 API
  • 1–32 vCPU and 1–64 GB RAM per VM, billed hourly
  • Resize live as your model or traffic changes
  • Larger models are GPU territory: join the waitlist

Storage and CDN

Weights, datasets and outputs move for free

Keep model files and training data in an S3-compatible bucket and pull them to any VM with the tools you already use. Generated images and cached responses go out through the CDN. None of it adds an egress line.
  • Works with the AWS CLI and any S3 SDK
  • Presigned URLs for private datasets and uploads
  • Outputs served from 988 edge locations
  • Egress $0 on Storage, Compute and CDN

Private networking

Prompts stay on your machines

Put the model on a private network next to your app and close its port to everything else. Customer data goes from your app to your model and back, without leaving your VMs.
  • Private VM-to-VM networks, with free internal bandwidth
  • Security groups with default-deny inbound rules
  • Open weights you choose, version and audit

Open-source runtimes

Start from a runtime you already know

Step-by-step guides for the common open-model runtimes. llama.cpp and Ollama run on CPU VMs today; vLLM and ComfyUI are GPU-first.
Machine learning at the edge

Coming soon

GPU instances are on the way

Larger models, fine-tuning and image generation need a GPU, and Edge doesn't offer them yet. GPU instances from NVIDIA, AMD and Intel are in the works, with hourly billing and no long-term commitment. Join the waitlist for early access, or talk to us if you need capacity sooner.

What it costs

A private 8B model, priced to the cent

This is the whole bill for a VM serving Llama 3.1 8B around the clock, from the published rate card. No per-token line, no egress line, and it stops accruing the moment you stop the VM.
  • Prepaid balance, so there's no surprise bill
  • Resize for a launch, scale back afterwards
  • Same price in every region
See the rate card

Estimated monthly bill

llm-box: 8 vCPU · 16 GB RAM · 80 GB NVMe, running 730 hours

USD
  • vCPU$23.368 × $0.004 per hour
  • RAM$29.9016 GiB × $0.00256 per GiB-hour
  • NVMe disk$5.9280 GiB × $0.074 per GiB-month
  • Model weights$0.004.9 GB in Storage, inside the 5 GB free tier
  • Per-token fees$0.00None. The VM is the bill.
  • Egress$0.00
Total$59.18

FAQ

Straight answers

Planning something bigger? Our engineers can help.
Can I rent a GPU on Edge today?
Not yet. GPU instances are coming soon and you can join the waitlist on the GPU page. Today, quantised models up to around 8B parameters run on CPU VMs, which suits low-volume inference, embeddings and development.
Which models can I run on a CPU VM?
Anything in GGUF format that fits in RAM. For production-quality responses, 4-bit quantisations of 3B–8B models are the practical range on CPU. Models of 13B and above really want a GPU.
Does my data leave my infrastructure?
Not when you self-host. The model runs on your VM, reached over a private network from your app, and nothing is sent to a third-party model provider. Edge Assist is a hosted service that answers from the content you give it.
How do I get model weights onto a VM quickly?
Keep them in an Edge Storage bucket and copy them with the AWS CLI or any S3 SDK. Transfers out of Storage are free, so every new VM can pull its own copy.

Put a private model to work this week

Start with a small VM and a bucket. Keep your data, drop the per-token bill, and join the GPU waitlist for what comes next.

Free to sign up. Nothing on Edge silently starts billing.