#llm-evaluation (19 Repositories)
Ranked open-source repositories tagged with #llm-evaluation, scored by pull request acceptance likelihood and maintainer engagement velocity.
35.6%
22.4h
19 repositories tagged #llm-evaluation
JasonColapietro/suede-creator-skills
Open-source Agent Skills for Claude Code and Codex: ship-DAG orchestration, A-F code review, AI evals, CI gates, design, copy, SEO/AEO/GEO, Instagram growth, app shipping, creator rights, and consumer recovery.
Jwuthri/Tracely-ai
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
juanjuandog/FinSight-AI
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.
comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
future-agi/future-agi
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Giskard-AI/giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
EvolvingLMMs-Lab/lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
microsoft/prompty
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandability, and portability for developers.
OpenBMB/UltraEval-Audio
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation
confident-ai/deepeval
The LLM Evaluation Framework
weavebench/WeaveBench
[EMNLP 2026] WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
palico-ai/palico-ai
Build, Improve Performance, and Productionize your AI Application
ukanwat/aaabench
A long-horizon benchmark harness: give a coding agent a real game engine, professional conditions and time, and ask it to build an open-world game. Harness only, no results.
alphadl/AdaRubrics
AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories
EverMind-AI/SkillCorpus
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
msoedov/agentic_security
Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪
ValueByte-AI/Awesome-LLM-in-Social-Science
Awesome papers involving LLMs in Social Science.
Jwuthri/Tracely
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
jeinlee1991/chinese-llm-benchmark
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。