Open-source CLI and library for testing, evaluating, and red-teaming LLM prompts, models, and RAG pipelines — runs locally on your machine or in CI.
Developer-focused open-source CLI and library for local or CI-based LLM evaluation, red-teaming, and RAG regression testing.
Promptfoo is best for engineering and security teams that need open-source, repeatable LLM evaluation, red-teaming, and regression testing in local development or CI, with a free Community tier and quote-based Enterprise or On-Premise options for teams that need shared dashboards, SSO, SLAs, managed cloud, or customer-controlled deployment. Public Promptfoo documentation describes the core product as an open-source CLI and library for evaluating and red-teaming LLM apps, and the website expands that positioning across evaluations, red teaming, guardrails, model security, MCP Proxy, and code scanning. That source context matters because Promptfoo is not just a hosted evaluation dashboard: it can run locally, fit into CI workflows, compare prompts and providers, test RAG behavior, apply assertions, and help teams catch regressions before release. The Community plan is listed as Free Forever at $0/month and includes all LLM evaluation features, all model providers and integrations, custom integration with your own app, local or self-hosted operation, vulnerability scanning, community support, and red teaming up to 10k probes per month. Enterprise and On-Premise are both public Custom plans rather than self-serve paid tiers. Promptfoo does not publish exact Enterprise or On-Premise monthly prices, annual prices, billing periods, conversion thresholds from Community, paid seat limits, standard usage caps, data-retention terms, or minimum contract lengths in the provided pricing content. The clearest public conversion path is therefore usage and procurement driven: teams can start on Community, then contact sales when they need Enterprise capabilities such as team sharing, continuous monitoring, centralized security and compliance dashboards, configurable attack profiles, SSO, granular permissions, Promptfoo API access, managed cloud deployment, professional services, priority support, or SLA guarantees. On-Premise is the higher-control custom option for deployment on customer infrastructure, complete data isolation, a dedicated runner, an assigned deployment engineer, and enterprise security and deployment support. For buyers, the practical pricing takeaway is simple: Promptfoo is transparent and generous at the free tier, but paid planning requires a sales conversation because final Enterprise and On-Premise cost, billing frequency, limits, and contractual obligations are not publicly specified. The tool is strongest when a team wants evaluation definitions and security tests to live close to engineering workflows, especially for prompt changes, model comparisons, RAG pipelines, agent behavior, MCP-related security boundaries, and release gates in CI/CD. It is less ideal as a purely nontechnical prompt management workspace or as a replacement for full production observability, because its public materials emphasize repeatable evaluation, red-team testing, vulnerability scanning, guardrails, MCP Proxy security, and development-time checks more than trace-first monitoring.
Was this helpful?
Promptfoo evaluates prompts, models, and RAG pipelines so teams can compare behavior across changes. This is useful for regression testing, factuality checks, hallucination reduction, and validating whether a model or retrieval change improves real application outputs.
The Red Teaming product is designed to proactively identify and fix vulnerabilities in AI applications. Teams can use it to test jailbreak resistance, adversarial prompts, unsafe completions, and other security risks before users encounter them.
Promptfoo’s Guardrails are positioned as real-time protection against jailbreaks and adversarial attacks. This makes the platform relevant not only for offline evaluation but also for teams considering runtime safety controls around LLM applications.
The MCP Proxy is described as a secure proxy for Model Context Protocol communications. This is important for agentic systems that use MCP connections and need a security boundary around model-to-tool or model-to-context interactions.
Promptfoo’s Code Scanning product finds LLM vulnerabilities in IDE and CI/CD workflows. That lets engineering teams catch AI-specific security issues earlier in the software development process instead of relying only on manual review or production monitoring.
$0/month
Custom
Custom
Ready to get started with Promptfoo?
View Pricing Options →We believe in transparent reviews. Here's what Promptfoo doesn't handle well:
Weekly insights on the latest AI tools, features, and trends delivered to your inbox.
The scraped website content states “Promptfoo is now part of OpenAI” and shows © 2026 Promptfoo, Inc. The provided content does not include a dated release note or detailed 2025-2026 changelog beyond that update.
LLM Observability
Braintrust is an evals-first LLM observability platform combining production tracing, prompt playgrounds, autoevals, and Topics-based pattern discovery for teams shipping AI in production.
AI Observability
LangSmith is LangChain's commercial observability, evaluation and prompt management platform for LLM apps and agents in production.
LLM evaluation and governance
an LLM development platform for prompt management, evaluations, logging, and trustworthy AI product iteration; the homepage announces the team joining Anthropic.
Testing & Quality
Open-source LLM evaluation framework with 50+ research-backed metrics including hallucination detection, tool use correctness, and conversational quality. Pytest-style testing for AI agents with CI/CD integration.
No reviews yet. Be the first to share your experience!
Get started with Promptfoo and see if it's the right fit for your needs.
Get Started →Take our 60-second quiz to get personalized tool recommendations
Find Your Perfect AI Stack →Explore 20 ready-to-deploy AI agent templates for sales, support, dev, research, and operations.
Browse Agent Templates →