#evals (11 Repositories)
Ranked open-source repositories tagged with #evals, scored by pull request acceptance likelihood and maintainer engagement velocity.
55.8%
57.0h
11 repositories tagged #evals
minghinmatthewlam/openbench
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
AgentEvalHQ/AgentEval
AgentEval is the comprehensive .NET toolkit for AI agent evaluation—tool usage validation, RAG quality metrics, stochastic evaluation, and model comparison—built first for Microsoft Agent Framework (MAF) and Microsoft.Extensions.AI. What RAGAS, PromptFoo and DeepEval do for Python, AgentEval does for .NET
spences10/my-pi
Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry
Jwuthri/Tracely-ai
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
NiceEval/NiceEval
build eval for your agent in 10 mins
benchflow-ai/benchflow
Research infra for creating RL environments, post-training, and evals.
hud-evals/hud-python
RL environments + evals for AI agents. Define once, train anything.
harbor-framework/harbor
Framework for evaluating and improving agents
tikalk/adlc-team-skills
🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams
LilMGenius/paperthin
Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.
Jwuthri/Tracely
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.