coder_eval
Playwright for coding agents. Test that your skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, A/B experiments, CI gates.
- Type: Skill
- Imported from GitHub
- Popularity: 148 GitHub stars
- License: Apache-2.0
- Source: https://github.com/UiPath/coder_eval
- Repository: https://github.com/uipath/coder_eval
- Tags: agent-evaluation, agent-skills, agent-testing, anthropic, claude, claude-code, claude-code-plugins-marketplace, claude-code-skills, claude-skills, codex, coding-agents, evaluation-framework
- Updated: 2026-10-01
More skills
- skills — Public repository for Agent Skills
- ponytail — Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- agent-skills — Production-grade engineering skills for AI coding agents.
- open-design — 🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becom…
- claude-mem — Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and inject…
- Understand-Anything — Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about…
- archify — Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion a…
- awesome-claude-skills — A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows