Back to Topics Directory
Topic Hub

#llm-evaluation (19 Repositories)

Ranked open-source repositories tagged with #llm-evaluation, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

35.6%

Avg Review Latency

22.4h

Filter by language

19 repositories tagged #llm-evaluation

B TierJavaScript 139

JasonColapietro/suede-creator-skills

Open-source Agent Skills for Claude Code and Codex: ship-DAG orchestration, A-F code review, AI evals, CI gates, design, copy, SEO/AEO/GEO, Instagram growth, app shipping, creator rights, and consumer recovery.

96.2%
Merge Rate
2d
First Review
100%
1st-Timers
0
Maintainers
B TierPython 1.0k

Jwuthri/Tracely-ai

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

84.2%
Merge Rate
3d
First Review
75%
1st-Timers
0
Maintainers
A TierJava 1.0k

juanjuandog/FinSight-AI

AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.

66.0%
Merge Rate
1h
First Review
100%
1st-Timers
3
Maintainers
A TierPython 21.6k

comet-ml/opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

78.5%
Merge Rate
2d
First Review
73%
1st-Timers
44
Maintainers
B TierPython 1.8k 116 GFIs

future-agi/future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

53.8%
Merge Rate
2d
First Review
18%
1st-Timers
74
Maintainers
B TierPython 5.8k 1 GFIs

Giskard-AI/giskard-oss

🐢 Open-Source Evaluation & Testing library for LLM Agents

62.0%
Merge Rate
2d
First Review
27%
1st-Timers
28
Maintainers
B TierPython 4.4k

EvolvingLMMs-Lab/lmms-eval

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

71.9%
Merge Rate
4d
First Review
67%
1st-Timers
18
Maintainers
B TierRust 1.2k

microsoft/prompty

Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.

25.0%
Merge Rate
3h
First Review
0%
1st-Timers
1
Maintainers
B TierPython 311

OpenBMB/UltraEval-Audio

Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation

100.0%
Merge Rate
15h
First Review
0%
1st-Timers
1
Maintainers
B TierPython 18.0k 2 GFIs

confident-ai/deepeval

The LLM Evaluation Framework

39.0%
Merge Rate
2d
First Review
31%
1st-Timers
70
Maintainers
D TierPython 159

weavebench/WeaveBench

[EMNLP 2026] WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 343

palico-ai/palico-ai

Build, Improve Performance, and Productionize your AI Application

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierShell 377

ukanwat/aaabench

A long-horizon benchmark harness: give a coding agent a real game engine, professional conditions and time, and ask it to build an open-world game. Harness only, no results.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 354

alphadl/AdaRubrics

AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 197

EverMind-AI/SkillCorpus

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 2.0k

msoedov/agentic_security

Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMulti-language 646

ValueByte-AI/Awesome-LLM-in-Social-Science

Awesome papers involving LLMs in Social Science.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 755

Jwuthri/Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMulti-language 6.4k

jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Llm-evaluation Open Source Repositories & C-Rank™ | GetMerged