EvalCore

Snapshot testing for LLM apps and agents

Eval-core/evalcore · Open source

An AI app can keep running after a prompt edit, model swap, or dependency update while its answers quietly get worse. EvalCore catches those changes in CI with repeatable evals and offline replay. It ships as one binary and works from YAML configs and JSONL cases, with baselines, trials, model comparisons, and agent traces. I built it with Abhishek Manyam, and it is open source under Apache-2.0.

Stack
Rust
Started
Jul 2026
Last push
Jul 2026

Recent commits