#evals (5)
2026-09-30 posts What I changed in my Claude Code setup after the Opus 5.5 and Sonnet 5.5 posts 2026-09-12 wiki Plugin evals: the ablation arm is the finding 2026-09-05 wiki Commerce agents blueprint — Anthropic's own answer to skills vs subagents 2026-05-22 wiki DeepEval — pytest-native LLM evaluation framework 2026-04-09 wiki Agent toolkit landscape — from prompt libraries to autonomous project managers