Run, fine-tune, and deploy thousands of community AI models with a single HTTP API — covering image, video, audio, language, and embedding models, billed per-second of GPU time.
Run, fine-tune, and deploy thousands of community AI models with a single HTTP API — covering image, video, audio, language, and embedding models, billed per-second of GPU time.
Replicate is the AI equivalent of an app store and serverless runtime in one. Any model published on Replicate (and many of the most-used image, video, and audio models are first-published here — FLUX, Stable Diffusion variants, Bria, Whisper, MusicGen, Stable Video Diffusion, Llama, plus thousands of community fine-tunes) can be called as a versioned HTTP endpoint without provisioning GPUs. Replicate handles autoscaling, queuing, cold-start optimization, and webhook delivery for long-running predictions. Developers can also push their own models using Cog, Replicate's open-source containerization tool, which turns a Python file + cog.yaml into a portable, GPU-ready prediction service that runs the same locally and in Replicate's cloud. Pricing is per-second of GPU time on the underlying hardware (with cheap, fast endpoints for popular models priced per-output — e.g., per image — to abstract the GPU detail away). For teams that need stability, Replicate offers Deployments (private, autoscaling endpoints with your own minimum-warm pool and scale rules) and Enterprise (dedicated capacity, SLAs, security review).
Was this helpful?
Feature information is available on the official website.
View Features →Per-second GPU billing (T4/A40/A100/L40S/H100 tiers) or per-output for popular fast models (FLUX, Whisper, etc.)
Per-second GPU billing on private autoscaling endpoints
Custom
Ready to get started with Replicate?
View Pricing Options →Weekly insights on the latest AI tools, features, and trends delivered to your inbox.
AI Model Hosting & Inference
AI-native cloud for inference, fine-tuning, and dedicated GPU clusters, offering 200+ open-source and frontier-class models behind an OpenAI-compatible API plus reserved H100/H200/B200 capacity.
AI Model Hosting & Inference
Production inference platform for open-weight LLMs, multimodal models, and custom fine-tunes — known for very fast serving (FireAttention/FireOptimizer), reliable function calling, and JSON mode at low per-token prices.
Deployment & Hosting
Baseten helps engineering teams deploy, autoscale, and monitor custom or open-source AI models behind production-ready inference APIs.
AI Cloud Infrastructure
GPU cloud with on-demand Pods, serverless inference, and multi-node clusters across 31 global regions — per-second billing on H100, H200, B200, and RTX GPUs.
No reviews yet. Be the first to share your experience!
Get started with Replicate and see if it's the right fit for your needs.
Get Started →Take our 60-second quiz to get personalized tool recommendations
Find Your Perfect AI Stack →Explore 20 ready-to-deploy AI agent templates for sales, support, dev, research, and operations.
Browse Agent Templates →