Evals for any agent in minutes.

Halios brings evals to your coding agent. Add the Halios skill to your preferred coding agent, then build eval suites, investigate failures, run experiments, and improve your agent - just by prompting.

refund-agent / Claude Code
HALIOS RUNTIME
YOU

Simple, but opinionated.

Halios makes a few deliberate choices about how agent evals should be built, run, and shipped faster.

Coding-agent native

Your coding agent is the interface for evals. It uses repository context to create scenarios, define checks, run evals, and investigate failures.

Local → CI → Production

Use the same scenarios and checks while developing, in CI, and against production traces.

Scenarios, not static datasets

Describe the situation you want to test. Halios runs fresh multi-turn trials against the current agent instead of scoring a fixed conversation.

Repo-native evals

Keep scenarios and checks in Git alongside the code they evaluate. Review, branch, and change them together.

OpenTelemetry first

Send standard OpenTelemetry traces from your existing instrumentation or let your coding agent add it. No proprietary tracing SDK required.

Continual improvement

Run evals, investigate failures, make changes, and verify the fix with fresh trials—all from the coding-agent workflow.

What can you do with Halios?

From setup to continuous improvement: prompt your coding agent to build scenarios, simulate multi-turn trials, catch and debug failures.

Create evaluation suite for your agent with realistic scenarios, edge cases, and checks.

refund-agent / Codex
YOUprompt

Set up Halios evals for this refund agent. Start with the critical refund path.

Loaded Halios skill
Registered eval hooks and scenario schema in .halios/
Inspected refund_agent.py
Read 4 files · found 4 tools and refund policies
Created evaluation suite
Generated 4 scenarios · 4 checks for refund behavior
Generated .halios/scenarios.yml and .halios/eval.yml
Eval suite ready for baseline runs
app.halios.ai/refund-agent / Scenarios
Halios Runtime4 scenarios · 4 checks
Halios Scenarios product interface

Start building at no cost.

Start with a generous free tier. Pay only as your evaluation workload grows.

ResourceMonthly allowanceAdditional usage
Trace Ingestion2 GB/month included$2 per GB
Evaluation Checks10K/month included$0.50 per 1K checks
Managed Evaluation Usage1M tokens/month included$2 per 1M tokens
Trace Retention14 days included30 days + $0.25/GB/month
No base fee. No per-seat fee. Bring your own model key anytime.

Free allowances reset monthly. Additional usage requires a payment method.

Need enterprise deployment or support?

Enterprise plans include on-prem hosting, SSO, RBAC, volume pricing, and priority support.

Frequently Asked Questions

Getting started

It installs the open-source Halios skill and CLI, authenticates with Halios, inspects your agent, and helps your coding agent create scenarios, define checks, run evaluations, surface failures, and more. You see and approve all the actions.

Fast-moving teams building AI agents. Developers, engineering leads, and product managers responsible for agent outcomes can use the same coding-agent-led workflow to build evals, investigate failures, and fix and improve their agents.

The Halios skill and CLI are open source and available on GitHub. The hosted evaluation runtime is proprietary.

How Halios works

A saved transcript records what your agent did once. Re-evaluating it only scores the same execution again. Halios starts from a scenario and runs a new interaction each time, testing how the current agent actually behaves.

Halios helps you understand where and why an agent is failing, but infrastructure observability is not the problem we are trying to solve. If your primary need is real time visibility into infra telemetry, dedicated observability tools are a better fit. Halios focuses on evaluating agent behavior, investigating failures, and verifying improvements.

Halios is built around a coding-agent-led workflow for agent evaluation. Your coding agent can create and maintain eval suites, run scenario-based simulations, investigate failures, and verify fixes using Halios. Instead of stitching together disconnected libraries and workflows built around static datasets, Halios gives you a simple, opinionated loop for continuously evaluating and improving agents.

Data, security & deployment

Halios receives the agent traces needed for evaluation, which may include conversations, prompts, tool definitions, tool calls, and execution context. Halios does not require or receive your source code.

You control what leaves your environment. Sensitive fields and PII can be filtered or redacted through OpenTelemetry before traces are sent to Halios. Hosted data is securely stored, isolated between customers, and retained only for the configured retention period. Private and on-prem deployments are available for stricter requirements.

You own your data. Halios does not sell customer data or use your traces, prompts, or evaluation results to train models. Data is deleted according to your configured retention period.

Yes. Enterprise deployments can run privately or on-prem. You can also bring your own model keys or endpoints instead of using Halios-managed models.

Build Resilient Agents With Halios

Install the skill in your coding agent and start creating test scenarios, automated checks, and CI quality gates today.