#agent-evaluation (7 Repositories)
Ranked open-source repositories tagged with #agent-evaluation, scored by pull request acceptance likelihood and maintainer engagement velocity.
32.5%
4.8h
7 repositories tagged #agent-evaluation
AMD-AGI/AgentKernelArena
AgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, and GEAK—can be evaluated side-by-side on the same GPU kernel tasks, using objective and reproducible metrics.
coze-dev/coze-loop
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.
NVIDIA/SkillEvaluator
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
yzhao062/awesome-auditable-ai
Auditing AI agents: a curated list of papers, tools, datasets, benchmarks, and standards covering reliability, monitoring, failure attribution, and decision records.
samarailly51-pixel/claimpilot-harness
Crash-test insurance claim AI agents before production.
alphadl/AdaRubrics
AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories
chirpz-ai/pandaprobe
open source agent engineering platform: traces, evals, and metrics to debug and improve your AI agents. Integrates with LangGraph, CrewAI, Claude Agent SDK, and more.