Cloudflare Workers AI runs 50+ open-source models (Llama 3.1/3.2, Mistral, Whisper, embeddings, vision) on serverless GPUs at the edge for $0.011 per 1,000 Neurons.
Cloudflare Workers AI runs 50+ open-source models (Llama 3.1/3.2, Mistral, Whisper, embeddings, vision) on serverless GPUs at the edge for $0.011 per 1,000 Neurons.
Cloudflare Workers AI gives you serverless GPU inference on Cloudflare's global edge network with 50+ open-source models and a pay-per-use pricing model denominated in 'Neurons' ($0.011 per 1,000 Neurons). It went Generally Available with model coverage spanning Llama 3.1 8B fp8-fast ($0.045 / M input tokens, $0.384 / M output), Llama 3.2 1B, 3B, and 11B Vision Instruct, Whisper for transcription, BGE and m2-bert for embeddings, plus Markdown Conversion (beta), Function Calling (beta), JSON Mode, Asynchronous Batch API (beta), and LoRA fine-tune adapters. Both Workers Free and Workers Paid plans give you 10,000 Neurons/day free; above that, Workers Paid bills at $0.011 per 1,000 Neurons. Workers AI has first-party bindings inside Cloudflare Workers + Pages, REST API access, OpenAI-compatible endpoints, and integrations with <a href="/tools/openrouter">OpenRouter</a>'s Vercel AI SDK and HuggingFace Chat UI. Compare with <a href="/tools/groq">Groq</a> for fast Llama inference, <a href="/tools/together-ai">Together AI</a> for broader open-source coverage, <a href="/tools/replicate">Replicate</a> for OSS model variety with per-second pricing, and <a href="/tools/hugging-face">Hugging Face</a> Inference Endpoints for dedicated deployments.
Was this helpful?
Cloudflare Workers AI transforms AI model deployment through global edge distribution and serverless architecture. The comprehensive model catalog of 50+ open-source models, transparent neuron-based pricing, and zero infrastructure management make it ideal for production teams already invested in the Cloudflare ecosystem. Where it excels is the seamless integration with Workers, Vectorize, R2, and AI Gateway — building a complete RAG or agent pipeline without leaving the platform is genuinely frictionless. The free tier is generous enough for prototyping, and pay-as-you-go pricing keeps costs predictable at scale. The main trade-off is the absence of frontier closed-source models and somewhat uneven feature support across the catalog. Teams needing GPT-4-class reasoning or Claude-level long-context performance will still need to proxy those through AI Gateway. For workloads that fit within the open-model catalog, Workers AI delivers a compelling combination of low latency, low cost, and operational simplicity.
Deploy AI models across 300+ edge locations worldwide, leveraging Cloudflare's anycast network to route requests to the nearest available GPU for optimized performance. Latency varies by model size and GPU availability at each location — smaller models like Mistral 7B and Gemma 2B typically achieve median latencies well under 100ms from nearby locations, while larger models may route to a more limited set of GPU-equipped data centers. The system automatically balances proximity, capacity, and load to deliver the best available response time for each request.
Use Case:
Building AI-powered applications serving global audiences where response time directly impacts user experience, such as real-time chat assistants or interactive content generation.
Access 50+ curated open-source models including Meta's Llama 3.3 and Llama 4 Scout, Mistral 7B for efficient text generation, Google's Gemma for lightweight inference, and Stable Diffusion XL for image generation. The catalog spans text generation, embeddings (BGE), speech-to-text (Whisper), translation, and image models — all optimized for edge deployment and accessible through a unified API.
Use Case:
Multi-modal AI applications requiring text generation, image creation, speech processing, and embedding generation without managing multiple AI service providers.
Pay-per-use pricing at $0.011 per 1,000 neurons with 10,000 neurons free daily. Neurons represent normalized compute units across different model types, providing predictable billing without idle costs or minimum commitments. Each model in the catalog publishes its neuron cost per request, enabling developers to estimate expenses before deploying. For example, a typical Llama 3.1 8B text generation request costs approximately 50 neurons (~$0.00055), while image generation models consume more neurons per request due to higher compute requirements.
Use Case:
Startups and variable-workload applications where traditional GPU instance pricing creates financial uncertainty or forces over-provisioning for peak capacity.
Zero infrastructure management with automatic scaling, batching optimization, and resource allocation. Models warm automatically based on usage patterns to minimize cold start latency. The platform handles GPU provisioning, model loading, request queuing, and scaling entirely behind the scenes, allowing developers to treat AI inference as a simple API call without any operational burden.
Use Case:
Production applications requiring elastic scaling during traffic spikes without pre-provisioning capacity or managing GPU clusters and model deployment pipelines.
Native integration with AI Gateway for observability and control, Vectorize for vector storage, Workers for edge computing, and R2/D1 for data storage, creating complete AI application stacks. AI Gateway provides unified caching, rate limiting, retry logic, fallback routing, and real-time analytics across both Workers AI and external providers. Vectorize enables semantic search and RAG pipelines, while D1 and R2 handle structured and object storage respectively.
Use Case:
Building end-to-end AI applications with semantic search, RAG capabilities, and edge processing without assembling multiple disparate cloud services.
Support for function calling, structured JSON outputs, reasoning tasks, vision processing, and multi-turn conversations with extended context windows for document processing applications. LoRA adapter loading enables fine-tuned model variants without redeploying base models, and batch processing handles high-volume offline workloads efficiently. The Agents SDK and Workflows product enable stateful, multi-step agent pipelines combining inference with durable execution.
Use Case:
Agentic AI workflows requiring tool usage, complex reasoning, document analysis, and multi-modal understanding for sophisticated automation and decision-making systems.
Free
$5/month Workers Paid base, then $0.011 per 1,000 Neurons
Custom quote
Ready to get started with Cloudflare Workers AI?
View Pricing Options →Cloudflare Workers AI works with these platforms and services:
We believe in transparent reviews. Here's what Cloudflare Workers AI doesn't handle well:
Weekly insights on the latest AI tools, features, and trends delivered to your inbox.
Through late 2025 and into 2026, Cloudflare expanded Workers AI with broader Llama 3.3 and Llama 4 Scout family availability, additional reasoning-tuned open models from DeepSeek and Qwen, and deeper Agents SDK and Workflows integration for building stateful multi-step agent pipelines. The AI Gateway received major updates including unified analytics across Workers AI and third-party providers, improved caching for repeated prompts, and fallback routing between multiple model backends. GPU capacity was expanded to additional edge locations, improving global coverage and reducing queueing for popular models during peak demand. LoRA adapter support was broadened to cover more base model architectures, and batch processing capabilities were enhanced for high-volume offline inference workloads. Cloudflare also introduced improved observability tooling with per-request cost tracking and latency breakdowns in the dashboard.
AI Model Hosting & Inference
Run, fine-tune, and deploy thousands of community AI models with a single HTTP API — covering image, video, audio, language, and embedding models, billed per-second of GPU time.
AI Model Hosting & Inference
AI-native cloud for inference, fine-tuning, and dedicated GPU clusters, offering 200+ open-source and frontier-class models behind an OpenAI-compatible API plus reserved H100/H200/B200 capacity.
No reviews yet. Be the first to share your experience!
Get started with Cloudflare Workers AI and see if it's the right fit for your needs.
Get Started →Take our 60-second quiz to get personalized tool recommendations
Find Your Perfect AI Stack →Explore 20 ready-to-deploy AI agent templates for sales, support, dev, research, and operations.
Browse Agent Templates →