#evaluation (30 Repositories)
Ranked open-source repositories tagged with #evaluation, scored by pull request acceptance likelihood and maintainer engagement velocity.
38.5%
34.2h
30 repositories tagged #evaluation
CodeSoul-co/Hypha
Harness-oriented agent system framework for production-grade LLM agent applications
modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
ncalc/ncalc
NCalc is a fast and lightweight expression evaluator library for .NET, designed for flexibility and high performance. It supports a wide range of mathematical and logical operations.
langwatch/langwatch
The platform for LLM evaluations and AI agent testing
Cloud-CV/EvalAI
:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI
nolabs-ai/deepfabric
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
TIGER-AI-Lab/ClawBench
Open-source benchmark for browser AI agents on daily tasks.
coze-dev/coze-loop
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
trpc-group/trpc-agent-go
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
NVIDIA-NeMo/Gym
Evaluate and improve models and agents using environments
EvolvingLMMs-Lab/lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
open-compass/opencompass
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
THU-BPM/Omni-SafetyBench
[ACM MM 2026 Oral] Code for paper "Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models".
Jwuthri/Tracely-ai
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
NVIDIA/SkillEvaluator
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
strands-agents/evals
A comprehensive evaluation framework for AI agents and LLM applications.
sdiehl/write-you-a-haskell
Building a modern functional compiler from first principles. (http://dev.stephendiehl.com/fun/)
NOAA-OWP/inundation-mapping
Flood inundation mapping and evaluation software configured to work with U.S. National Water Model.
vibrantlabsai/ragas
Supercharge Your LLM Application Evaluations 🚀
FeiZhuNiU-INFJA/LIFT
An Evaluation Framework for Self-Evolving-Agent (Loaded Impact on Holdout Final Task)
Jwuthri/Tracely
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
suyoumo/ClawProBench
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
alipay/ant-application-security-testing-benchmark
xAST评价体系,让安全工具不再“黑盒”. The xAST evaluation benchmark makes security tools no longer a "black box".
RecList/reclist
Behavioral "black-box" testing for recommender systems
caserec/CaseRecommender
Case Recommender: A Flexible and Extensible Python Framework for Recommender Systems
MirroS-Lab/HarnessEval-W
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
meituan-longcat/WBench
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
zzzprojects/Eval-Expression.NET
C# Eval Expression | Evaluate, Compile, and Execute C# code and expression at runtime.