Back to Topics Directory
Topic Hub

#evals (11 Repositories)

Ranked open-source repositories tagged with #evals, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

55.8%

Avg Review Latency

57.0h

Filter by language

11 repositories tagged #evals

B TierPython 118

minghinmatthewlam/openbench

Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost

91.4%
Merge Rate
14h
First Review
100%
1st-Timers
0
Maintainers
B TierC# 136

AgentEvalHQ/AgentEval

AgentEval is the comprehensive .NET toolkit for AI agent evaluation—tool usage validation, RAG quality metrics, stochastic evaluation, and model comparison—built first for Microsoft Agent Framework (MAF) and Microsoft.Extensions.AI. What RAGAS, PromptFoo and DeepEval do for Python, AgentEval does for .NET

92.6%
Merge Rate
16d
First Review
100%
1st-Timers
1
Maintainers
B TierTypeScript 110

spences10/my-pi

Composable Pi coding agent with MCP, LSP, agent chains, prompt presets, and local eval telemetry

90.4%
Merge Rate
<1h
First Review
100%
1st-Timers
1
Maintainers
B TierPython 1.0k

Jwuthri/Tracely-ai

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

84.2%
Merge Rate
3d
First Review
75%
1st-Timers
0
Maintainers
B TierTypeScript 105

NiceEval/NiceEval

build eval for your agent in 10 mins

82.7%
Merge Rate
4d
First Review
100%
1st-Timers
0
Maintainers
A TierPython 327

benchflow-ai/benchflow

Research infra for creating RL environments, post-training, and evals.

67.1%
Merge Rate
1d
First Review
50%
1st-Timers
16
Maintainers
A TierPython 294

hud-evals/hud-python

RL environments + evals for AI agents. Define once, train anything.

69.8%
Merge Rate
1d
First Review
50%
1st-Timers
9
Maintainers
B TierPython 4.8k

harbor-framework/harbor

Framework for evaluating and improving agents

35.1%
Merge Rate
8h
First Review
28%
1st-Timers
65
Maintainers
D TierShell 130

tikalk/adlc-team-skills

🐙 ADLC Team Skills — Agentic SDLC for Engineering Teams

0.0%
Merge Rate
-
First Review
0%
1st-Timers
2
Maintainers
D TierShell 121

LilMGenius/paperthin

Low-level agentic design patterns. Turning old engineering wisdom into reflexes your agent reaches for on its own—on any agent.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 755

Jwuthri/Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Evals Open Source Repositories & C-Rank™ | GetMerged