Mistral's open-weight agentic coding model family — Devstral Small and Devstral Medium — purpose-built for OpenHands, SWE-agent, and IDE coding agents.
Mistral's open-weight agentic coding model family — Devstral Small and Devstral Medium — purpose-built for OpenHands, SWE-agent, and IDE coding agents.
Mistral Devstral is a focused option for agentic coding and repository-level software work. Its clearest differentiator is a coding-specific model designed with the OpenHands team, an advertised 128K context window, and an Apache 2.0 open-weight Small variant. The staging record identifies these concrete capabilities: Purpose-built for agentic coding harnesses, Co-developed with OpenHands team, Devstral Small open-weight (Apache 2.0), 128K context window, Strong SWE-bench Verified performance, Cheap and self-hostable. Those details make the product worth evaluating for a defined workflow, but they are not a substitute for a hands-on pilot with the files, prompts, permissions, latency requirements, and failure cases your team actually has.
Pricing needs special care. The existing record lists Devstral Small (weights): Free (Apache 2.0) (Self-host from Hugging Face); API Devstral Small: $0.10/$0.30 per M tokens (Input/output on La Plateforme); API Devstral Medium: ~$0.40/$2 per M tokens (Frontier tier). During this scheduled run, both the vendor page and the expected pricing route were requested directly with curl, but neither returned usable HTML. Every figure and plan label therefore needs manual verification before it appears in a budget or purchasing decision. Do not assume that an open-source component makes hosting, support, storage, model inference, or implementation free. Ask for a written breakdown of subscription or usage charges, minimum commitments, overages, support, data egress, implementation, renewal terms, and cancellation rights. Model a 12-month total cost that includes engineering, security review, monitoring, and human quality control.
A sensible comparison set is Claude, Aider, Cursor Agent, Continue. These products are adjacent rather than perfectly interchangeable. Compare the part of the workflow each one owns, deployment options, regional availability, authentication, audit logs, data retention, export formats, rate limits, and the work required to recover from a failed call. Mistral Devstral should win only when its specific workflow produces a measurable result, not because a demonstration looks polished. For a fair test, run at least 20 representative cases, including several intentionally difficult ones, and preserve the inputs and expected outputs so competing products see the same workload.
The main advantages are specific: Open-weight Small variant can be self-hosted under Apache 2.0; Designed for agents that inspect files, edit code, and run tests; 128K context supports work across larger repositories. The tradeoffs are equally important: Current API token prices and benchmark results could not be rechecked in this run; Self-hosting a 24B-parameter model requires meaningful GPU memory and operations work; Coding benchmarks do not guarantee reliability on a team’s private repositories. Validate every generated answer, image, code change, or automated action at the boundary where an error becomes expensive. For production use, define who approves high-impact actions, how credentials are scoped, where logs are stored, and how a person can stop or reverse a workflow. Sensitive-data users should verify encryption, retention, deletion, subprocessors, data residency, training policy, and incident-response commitments in writing.
Practical use cases include Powering an OpenHands-style issue-to-patch workflow; Running a private coding assistant on controlled infrastructure; Automating multi-file refactors with tests in the loop; Evaluating an open-weight alternative to proprietary coding models. Pick one as the pilot rather than attempting a broad rollout. Capture baseline completion time, review time, correction rate, failure rate, unit cost, and user adoption before introducing the product. Then repeat those measurements with Mistral Devstral, counting human review and retries instead of treating them as free. A useful pilot has an owner, acceptance thresholds, a rollback path, and a fixed end date. Buy or standardize only if the measured gain exceeds license and operating costs without weakening accuracy, security, or accountability. That disciplined test is more informative than vendor benchmarks and protects the team while current pricing and product details await manual confirmation.
Was this helpful?
Feature information is available on the official website.
View Features →Free (Apache 2.0)
$0.10/$0.30 per M tokens
~$0.40/$2 per M tokens
Ready to get started with Mistral Devstral?
View Pricing Options →Weekly insights on the latest AI tools, features, and trends delivered to your inbox.
No reviews yet. Be the first to share your experience!
Get started with Mistral Devstral and see if it's the right fit for your needs.
Get Started →Take our 60-second quiz to get personalized tool recommendations
Find Your Perfect AI Stack →Explore 20 ready-to-deploy AI agent templates for sales, support, dev, research, and operations.
Browse Agent Templates →