Why we built Halios
We built Halios because agent evals kept turning into a pile of scripts, dashboards, and one-off test harnesses.
Testing an agent means more than checking an input and an output. You need to re-run conversations, tools, edge cases, and failures, then understand what changed when you try to fix something.
We wanted that work to stay close to the code. So instead of building another harness, we decided to embed ourselves in coding workflow developers already use. Halios gives coding agents the tools to build and run evals, investigate failures, and verify changes without requiring developers to adopt a separate evaluation workflow.
It's the fastest way to get started with agent evaluation. Give it a try!
Our approach
We believe teams should keep concise, reviewable evaluation specifications near their application code while Halios manages the execution pipeline and evidence. Scenarios can follow branches and pull requests, coding agents can understand them in context, and multiple people can collaborate through familiar Git workflows. Halios handles fresh multi-turn simulations, trace evidence, evaluation execution, result history, and the operational work required to use the same evaluation model from local development through CI and production.
Halios is framework- and model-provider-independent and uses OpenTelemetry-compatible traces so teams can keep their instrumentation portable. The Halios Agent Skill and CLI are open source; the hosted evaluation runtime is operated by Halios.
Company
Halios is developed and operated by Anomalytica, Inc., a San Francisco, California company building developer infrastructure for reliable AI agents. We work with teams shipping chatbots, RAG systems, tool-using agents, voice agents, coding agents, and custom agent workflows. For product, support, privacy, security, partnership, or enterprise deployment questions, visit the Halios contact page.