Arize Phoenix provides open-source tracing, evaluations, and MCP tooling for ai observability workflows. It includes MCP compatibility.
Arize Phoenix provides open-source tracing, evaluations, and MCP tooling for ai observability workflows. It includes MCP compatibility.
Arize Phoenix is an open-source observability and evaluation platform for LLM applications and agents. Developers use it to capture traces, inspect retrieval and tool calls, organize datasets, run evaluations, and compare behavior over time. Phoenix is the open-source project; Arize AX is the hosted commercial platform shown on the vendor pricing page. Teams can self-host Phoenix, use AX SaaS, or negotiate Enterprise deployment.
Tracing exposes model calls, retrieval, tools, latency, tokens, errors, and intermediate agent steps. Evaluations score properties such as relevance or correctness across datasets. Experiments enable repeatable comparisons when models, prompts, or retrieval settings change. Phoenix uses OpenTelemetry-oriented instrumentation and offers MCP tooling relevant to agent workflows.
The September 2026 Arize page lists AX Free at $0: 10 Signal issues, 25,000 traced spans, 1 GB ingestion monthly, 15-day retention, SaaS deployment, unlimited users, and unlimited evaluations. AX Pro is $50 monthly with 25 Signal issues, 50,000 spans, 10 GB ingestion, 30-day retention, unlimited users, and unlimited evaluations. Enterprise is custom-priced with unlimited issues, custom spans, ingestion and retention, and SaaS or self-hosted deployment. Open source avoids a license fee but still incurs storage and operations.
Benefits include self-hosting, detailed trace timelines, and repeatable regression tests. Limits include instrumentation effort, telemetry growth, and evaluator calibration. Fifty thousand spans can disappear quickly if one run creates dozens. Thirty-day Pro retention may be short for audits. LLM judges can be biased and need human labels. Traces can contain prompts, retrieved documents, personal data, or secrets unless redacted before collection.
Run a two-week pilot on one bounded production-like workflow. Use at least 50 representative tasks and preserve a manual baseline. Record completion rate, factual or technical correctness, p50 and p95 latency, total usage cost, setup hours, correction time, and failure categories. Test permissions with an account that should not see the target data. Confirm export, deletion, outage handling, rate limits, support response, retention, subprocessors, regional processing, and whether customer content trains models. For every write-capable workflow, begin in a sandbox, require human approval, and document rollback before granting production access.
Do not choose Arize Phoenix from a feature checklist alone. Calculate annual cost at realistic volume with a 25% usage buffer and include implementation, monitoring, reviewer labor, and external model charges. Compare the same test set with langfuse, langsmith, helicone, braintrust, ai agent observability how to monitor debug and trace agents in production. A good purchase produces measurable net time savings after review and governance work; an impressive demo that creates more corrections does not.
This tool is most useful when its specific integrations and workflow match an existing bottleneck. It is a weaker fit when the team cannot define an owner, representative evaluation set, permission boundary, or fallback process. Recheck vendor pricing and limits at purchase because cloud plans can change after this research date.
Was this helpful?
Leading open-source LLM observability platform offering comprehensive tracing, evaluation, and experimentation without vendor lock-in. Ideal for teams with DevOps capacity who need deep analytical insights into LLM application behavior, RAG pipeline quality, and multi-agent workflow debugging. Phoenix stands out for its OpenTelemetry foundation, which ensures trace portability and avoids ecosystem lock-in, and its robust evaluation framework that supports both automated LLM-as-a-judge scoring and human annotation workflows. The self-hosted model with zero licensing costs makes it particularly attractive for regulated industries and cost-conscious teams, though the operational overhead of managing infrastructure and the steeper learning curve compared to polished SaaS alternatives like LangSmith should be weighed against these benefits. With over 18,000 GitHub stars and strong backing from Arize AI, the project demonstrates sustained momentum and community adoption.
$0 license fee
$0
$50/month
Custom
Ready to get started with Arize Phoenix?
View Pricing Options →Arize Phoenix works with these platforms and services:
We believe in transparent reviews. Here's what Arize Phoenix doesn't handle well:
Weekly insights on the latest AI tools, features, and trends delivered to your inbox.
Through late 2025 and into 2026, Phoenix has expanded agent-focused tracing with deeper support for LangGraph, CrewAI, and AutoGen, including visualizations for multi-agent coordination and tool-call sequence inspection. The evaluation framework has been enhanced with new built-in evaluators for code generation quality, multi-turn conversation coherence, and structured output validation. Session and thread-based tracing now provides better visibility into conversational AI applications, grouping related interactions and tracking context evolution across turns. The prompt playground has been upgraded with multi-model comparison capabilities, allowing teams to test prompts against several providers simultaneously and feed results directly into experiments. Guardrails integration enables teams to define and monitor safety boundaries alongside performance metrics. The annotation workflow has been streamlined with bulk labeling tools, inter-annotator agreement metrics, and API-driven integration with external labeling platforms. Infrastructure improvements include faster trace ingestion, improved query performance for large datasets, and better support for high-cardinality span attributes in production environments.
AI Observability
LangSmith is LangChain's commercial observability, evaluation and prompt management platform for LLM apps and agents in production.
AI observability
An open-source observability and evaluation platform for language-model applications.
LLM Observability
Open-source LLM observability, gateway, and cost analytics platform — proxy your OpenAI, Anthropic, or Bedrock calls through Helicone and get traces, caching, retries, rate limiting, and cost tracking in one line of code.
MLOps
End-to-end MLOps and AI developer platform — Models (experiment tracking, sweeps, model registry) plus Weave (LLM/agent observability and evals) — used by frontier labs and enterprise ML teams.
No reviews yet. Be the first to share your experience!
Get started with Arize Phoenix and see if it's the right fit for your needs.
Get Started →Take our 60-second quiz to get personalized tool recommendations
Find Your Perfect AI Stack →Explore 20 ready-to-deploy AI agent templates for sales, support, dev, research, and operations.
Browse Agent Templates →