AI & machine learning
Run open models on infrastructure you control
GPU instances are coming soon. Everything else on this page runs today.
The problem
AI features shouldn't come with a meter running
- 01
Every prompt and document goes to someone else's servers
Run the model on your own VM, on a private network. Prompts and completions never pass through a third-party API.
- 02
Per-token pricing punishes the features people actually use
A VM bills by the hour, whatever it serves. Busy or quiet, the cost is the same, and you can stop it when it's idle.
- 03
Moving weights, datasets and outputs costs egress every time
Storage and CDN egress is $0. Pull a 5 GB model to every VM or serve generated images to millions without a transfer line.
- 04
Not every AI feature needs a model of your own
Edge Assist adds cited answers from your own content with one script tag, at a penny a question after 250 free each month.
How it works
The pieces of a private AI stack, available today
Compute
Small open models run well on a CPU VM
- llama.cpp and Ollama, with an OpenAI-compatible /v1 API
- 1–32 vCPU and 1–64 GB RAM per VM, billed hourly
- Resize live as your model or traffic changes
- Larger models are GPU territory: join the waitlist
Storage and CDN
Weights, datasets and outputs move for free
- Works with the AWS CLI and any S3 SDK
- Presigned URLs for private datasets and uploads
- Outputs served from 988 edge locations
- Egress $0 on Storage, Compute and CDN
Private networking
Prompts stay on your machines
- Private VM-to-VM networks, with free internal bandwidth
- Security groups with default-deny inbound rules
- Open weights you choose, version and audit
No model required
Ship AI features without hosting a model
Edge Assist
AI answers on your siteA search-and-ask box grounded in your own pages, with citations. It declines questions your content can't answer rather than making things up.
250 answered questions free each month, then $0.01 eachAgent API
Infrastructure for AI agentsGive your coding agent an access code. It discovers what it can do, dry-runs every change, stays inside your budget cap and reports back in plain language.
Compute, CDN, Storage and DNS through one API
Open-source runtimes
Start from a runtime you already know
- llama.cppQuantised LLM inference, CPU or GPU
- OllamaRun open LLMs on your own infrastructure
- vLLMProduction-grade LLM serving
- ComfyUINode-based Stable Diffusion workflows
Coming soon
GPU instances are on the way
Larger models, fine-tuning and image generation need a GPU, and Edge doesn't offer them yet. GPU instances from NVIDIA, AMD and Intel are in the works, with hourly billing and no long-term commitment. Join the waitlist for early access, or talk to us if you need capacity sooner.
Recommended products
What an AI workload uses on Edge
Compute
CPU VMs for quantised models, embeddings and the app that calls them.
Storage
Model weights, datasets and outputs in S3-compatible buckets, zero egress.
CDN
Generated images, files and cacheable responses served close to users.
Assist
Cited AI answers on your site from one script tag.
GPU Compute
Coming soon. Join the waitlist for early access.
What it costs
A private 8B model, priced to the cent
- Prepaid balance, so there's no surprise bill
- Resize for a launch, scale back afterwards
- Same price in every region
Estimated monthly bill
llm-box: 8 vCPU · 16 GB RAM · 80 GB NVMe, running 730 hours
- vCPU$23.368 × $0.004 per hour
- RAM$29.9016 GiB × $0.00256 per GiB-hour
- NVMe disk$5.9280 GiB × $0.074 per GiB-month
- Model weights$0.004.9 GB in Storage, inside the 5 GB free tier
- Per-token fees$0.00None. The VM is the bill.
- Egress$0.00
FAQ
Straight answers
Can I rent a GPU on Edge today?
Which models can I run on a CPU VM?
Does my data leave my infrastructure?
How do I get model weights onto a VM quickly?
Put a private model to work this week
Start with a small VM and a bucket. Keep your data, drop the per-token bill, and join the GPU waitlist for what comes next.
Free to sign up. Nothing on Edge silently starts billing.