EvalCore
Snapshot testing for LLM apps and agents
Eval-core/evalcore · Open source
An AI app can keep running after a prompt edit, model swap, or dependency update while its answers quietly get worse. EvalCore catches those changes in CI with repeatable evals and offline replay. It ships as one binary and works from YAML configs and JSONL cases, with baselines, trials, model comparisons, and agent traces. I built it with Abhishek Manyam, and it is open source under Apache-2.0.
- Stack
- Rust
- Started
- Jul 2026
- Last push
- Jul 2026
