#evaluation-framework (7 Repositories)
Ranked open-source repositories tagged with #evaluation-framework, scored by pull request acceptance likelihood and maintainer engagement velocity.
43.0%
26.4h
7 repositories tagged #evaluation-framework
EuroEval/EuroEval
The robust European language model benchmark.
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
future-agi/future-agi
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
EleutherAI/lm-evaluation-harness
A framework for few-shot evaluation of language models.
confident-ai/deepeval
The LLM Evaluation Framework
aiverify-foundation/moonshot
Moonshot - A simple and modular tool to evaluate and red-team any LLM application.
MaurizioFD/RecSys2019_DeepLearning_Evaluation
This is the repository of our article published in RecSys 2019 "Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches" and of several follow-up studies.