Back to Topics Directory
Topic Hub

#evaluation (30 Repositories)

Ranked open-source repositories tagged with #evaluation, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

38.5%

Avg Review Latency

34.2h

Filter by language

30 repositories tagged #evaluation

B TierTypeScript 282

CodeSoul-co/Hypha

Harness-oriented agent system framework for production-grade LLM agent applications

95.2%
Merge Rate
3d
First Review
100%
1st-Timers
1
Maintainers
A TierPython 3.3k

modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

84.5%
Merge Rate
10h
First Review
62%
1st-Timers
40
Maintainers
A TierPython 27.7k 6 GFIs

mlflow/mlflow

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

66.1%
Merge Rate
16h
First Review
32%
1st-Timers
104
Maintainers
A TierC# 1.2k 1 GFIs

ncalc/ncalc

NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.

97.8%
Merge Rate
9d
First Review
100%
1st-Timers
2
Maintainers
A TierTypeScript 3.5k

langwatch/langwatch

The platform for LLM evaluations and AI agent testing

78.0%
Merge Rate
1d
First Review
63%
1st-Timers
15
Maintainers
A TierPython 2.0k

Cloud-CV/EvalAI

:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI

85.4%
Merge Rate
3d
First Review
40%
1st-Timers
4
Maintainers
B TierPython 881 1 GFIs

nolabs-ai/deepfabric

Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline

54.9%
Merge Rate
<1h
First Review
100%
1st-Timers
1
Maintainers
A TierPython 571

TIGER-AI-Lab/ClawBench

Open-source benchmark for browser AI agents on daily tasks.

84.9%
Merge Rate
11d
First Review
80%
1st-Timers
5
Maintainers
A TierGo 5.7k

coze-dev/coze-loop

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.

73.7%
Merge Rate
1d
First Review
86%
1st-Timers
3
Maintainers
A TierTypeScript 24.4k 1 GFIs

promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

72.9%
Merge Rate
3d
First Review
54%
1st-Timers
38
Maintainers
A TierGo 1.7k

trpc-group/trpc-agent-go

A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

73.8%
Merge Rate
2d
First Review
32%
1st-Timers
34
Maintainers
A TierPython 1.1k

NVIDIA-NeMo/Gym

Evaluate and improve models and agents using environments

58.1%
Merge Rate
1d
First Review
59%
1st-Timers
66
Maintainers
A TierPython 4.4k

EvolvingLMMs-Lab/lmms-eval

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

72.3%
Merge Rate
4d
First Review
75%
1st-Timers
18
Maintainers
B TierPython 7.4k

open-compass/opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

72.2%
Merge Rate
3d
First Review
53%
1st-Timers
12
Maintainers
B TierPython 115

THU-BPM/Omni-SafetyBench

[ACM MM 2026 Oral] Code for paper "Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models".

66.7%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
B TierPython 1.0k

Jwuthri/Tracely-ai

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

20.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 181

NVIDIA/SkillEvaluator

Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.

0.0%
Merge Rate
5h
First Review
0%
1st-Timers
5
Maintainers
D TierPython 180

strands-agents/evals

A comprehensive evaluation framework for AI agents and LLM applications.

0.0%
Merge Rate
1d
First Review
0%
1st-Timers
2
Maintainers
D TierHAHaskell 3.5k

sdiehl/write-you-a-haskell

Building a modern functional compiler from first principles. (http://dev.stephendiehl.com/fun/)

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 129

NOAA-OWP/inundation-mapping

Flood inundation mapping and evaluation software configured to work with U.S. National Water Model.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 15.5k

vibrantlabsai/ragas

Supercharge Your LLM Application Evaluations 🚀

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 105

FeiZhuNiU-INFJA/LIFT

An Evaluation Framework for Self-Evolving-Agent (Loaded Impact on Holdout Final Task)

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 755

Jwuthri/Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 822

suyoumo/ClawProBench

ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJava 491

alipay/ant-application-security-testing-benchmark

xAST评价体系,让安全工具不再“黑盒”. The xAST evaluation benchmark makes security tools no longer a "black box".

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 475

RecList/reclist

Behavioral "black-box" testing for recommender systems

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 499

caserec/CaseRecommender

Case Recommender: A Flexible and Extensible Python Framework for Recommender Systems

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 139

MirroS-Lab/HarnessEval-W

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 213

meituan-longcat/WBench

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC# 475

zzzprojects/Eval-Expression.NET

C# Eval Expression | Evaluate, Compile, and Execute C# code and expression at runtime.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers