December 4, 2025
Edge Inference
Inference where your users already are
Edge Inference runs smaller models at network edge locations, cutting round-trip latency for classification, moderation and retrieval tasks that do not require frontier-scale models.
What it does
- Runs models at edge locations near users
- Cuts latency for classification and moderation
- Bills per inference with no idle capacity
Full review
Read the full Cloudflare review with scored criteria and pricing at /software/cloudflare.