agent-evaluation
Comprehensive framework for evaluating AI agents through multi-modal grading, benchmarks, and production monitoring.
Install
npx skills add https://github.com/supercent-io/skills-template --skill agent-evaluationStats
| Total installs | 10,066 |
| Weekly installs | 10.1K |
| GitHub stars | 88 |
| First seen | Jan 24, 2026 |
| Source | @supercent-io/skills-template |
Summary
- Covers three grader types (code-based, model-based, human) with trade-offs and best practices for each agent category
- Provides an 8-step roadmap from initial task creation through production monitoring, including environment isolation, outcome-focused grading, and saturation detection
- Includes benchmarks for major agent types: SWE-bench for coding, WebArena for computer use, τ2-Bench for conversational agents
- Offers CI/CD integration patterns, A/B testing templates, and production sampling strategies for real-time quality monitoring
Tags
Related skills
| Skill | Installs | vs agent-evaluation |
|---|---|---|
| create-readme | 8,622 | -1,444 |
| ralph | 10,688 | +622 |
| genkit | 10,389 | +323 |
| prompt-repetition | 10,488 | +422 |
| firecrawl-crawl | 8,324 | -1,742 |
FAQ
- How many installs does agent-evaluation have?
- agent-evaluation has 10,066 total installs and 10.1K installs this week.
- Where is agent-evaluation hosted?
- agent-evaluation is published by @supercent-io/skills-template at https://github.com/supercent-io/skills-template.
- How many GitHub stars does agent-evaluation have?
- agent-evaluation has 88 GitHub stars.
- When was agent-evaluation first indexed?
- OrangeBot.AI first indexed agent-evaluation on Jan 24, 2026.
- How do I install agent-evaluation?
- Run: npx skills add https://github.com/supercent-io/skills-template --skill agent-evaluation