Coding-agent native
Your coding agent is the interface for evals. It uses repository context to create scenarios, define checks, run evals, and investigate failures.
Halios brings evals to your coding agent. Add the Halios skill to your preferred coding agent, then build eval suites, investigate failures, run experiments, and improve your agent - just by prompting.
Halios makes a few deliberate choices about how agent evals should be built, run, and shipped faster.
Your coding agent is the interface for evals. It uses repository context to create scenarios, define checks, run evals, and investigate failures.
Use the same scenarios and checks while developing, in CI, and against production traces.
Describe the situation you want to test. Halios runs fresh multi-turn trials against the current agent instead of scoring a fixed conversation.
Keep scenarios and checks in Git alongside the code they evaluate. Review, branch, and change them together.
Send standard OpenTelemetry traces from your existing instrumentation or let your coding agent add it. No proprietary tracing SDK required.
Run evals, investigate failures, make changes, and verify the fix with fresh trials—all from the coding-agent workflow.
From setup to continuous improvement: prompt your coding agent to build scenarios, simulate multi-turn trials, catch and debug failures.
Create evaluation suite for your agent with realistic scenarios, edge cases, and checks.
Start with a generous free tier. Pay only as your evaluation workload grows.
Free allowances reset monthly. Additional usage requires a payment method.
Enterprise plans include on-prem hosting, SSO, RBAC, volume pricing, and priority support.
It installs the open-source Halios skill and CLI, authenticates with Halios, inspects your agent, and helps your coding agent create scenarios, define checks, run evaluations, surface failures, and more. You see and approve all the actions.
Fast-moving teams building AI agents. Developers, engineering leads, and product managers responsible for agent outcomes can use the same coding-agent-led workflow to build evals, investigate failures, and fix and improve their agents.
The Halios skill and CLI are open source and available on GitHub. The hosted evaluation runtime is proprietary.
A saved transcript records what your agent did once. Re-evaluating it only scores the same execution again. Halios starts from a scenario and runs a new interaction each time, testing how the current agent actually behaves.
Halios helps you understand where and why an agent is failing, but infrastructure observability is not the problem we are trying to solve. If your primary need is real time visibility into infra telemetry, dedicated observability tools are a better fit. Halios focuses on evaluating agent behavior, investigating failures, and verifying improvements.
Halios is built around a coding-agent-led workflow for agent evaluation. Your coding agent can create and maintain eval suites, run scenario-based simulations, investigate failures, and verify fixes using Halios. Instead of stitching together disconnected libraries and workflows built around static datasets, Halios gives you a simple, opinionated loop for continuously evaluating and improving agents.
Halios receives the agent traces needed for evaluation, which may include conversations, prompts, tool definitions, tool calls, and execution context. Halios does not require or receive your source code.
You control what leaves your environment. Sensitive fields and PII can be filtered or redacted through OpenTelemetry before traces are sent to Halios. Hosted data is securely stored, isolated between customers, and retained only for the configured retention period. Private and on-prem deployments are available for stricter requirements.
You own your data. Halios does not sell customer data or use your traces, prompts, or evaluation results to train models. Data is deleted according to your configured retention period.
Yes. Enterprise deployments can run privately or on-prem. You can also bring your own model keys or endpoints instead of using Halios-managed models.
Install the skill in your coding agent and start creating test scenarios, automated checks, and CI quality gates today.