agent-design-review
Review an LLM agent design and find where it will be unreliable, expensive, or unsafe. Use when asked to review an agent architecture, critique a multi-step/tool-using agent, debug an agent that loops or goes off-task, or harden an agent before launch. Produces a structured review — task fit, control flow, tools, memory/context, failure handling, cost, and safety — with prioritised findings and fixes.
git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills /tmp/agent-design-review && cp -r /tmp/agent-design-review/exports/openclaw/agent-design-review ~/.claude/skills/agent-design-reviewSKILL.md
# Agent Design Review Skill
Most agents don't fail because the model is weak — they fail because the *design* lets them loop, call the
wrong tool, lose the thread across steps, or burn tokens with no stopping rule. This skill reviews an agent's
architecture against the decisions that actually determine reliability, and ranks the fixes — so "it works in
the demo but not in prod" becomes a specific list of changes. (Writing a new agent spec? Use
[`agent-spec`](../agent-spec/SKILL.md).)
## Working from a brief
Given a sketch ("a research agent that searches, reads, and writes a report"), **deliver the full review
anyway** — infer the likely control flow and tools, label the inference, and flag what to confirm. Never
withhold the review for missing detail.
## Required Inputs
Ask for these only if they aren't already provided (else infer and label):
- **What the agent does** — its goal, and what a successful run produces.
- **Control flow** — single prompt, plan-then-execute, ReAct loop, or multi-agent; and the stopping condition.
- **Tools & actions** — what it can call, and which actions have side effects (write, send, pay).
- **Memory & context** — what state carries across steps, and how context is kept in budget.
- **Constraints** — latency, cost per run, and the trust boundary (untrusted input? real-world actions?).
## Output Format
### Agent Review: [agent]
**1. Summary** — will this be reliable in production? The top 3 risks and the single change that helps most.
**2. Findings by dimension** — for each, what's sound and what's fragile:
| Dimension | Finding | Severity | Fix |
|---|---|---|---|
| Control flow | no max-steps / no progress check → loops | High | step budget + "am I making progress?" check + halt |
| Tool use | overlapping tools confuse selection | Med | fewer, sharply-described tools; allowlist |
| Context | full history re-sent each step → cost + drift | High | summarise/scope memory per step |
| Failure handling | one tool error aborts the run | Med | retry/backoff + graceful degradation |
| Safety | acts without confirmation on writes | High | human/confirm gate on side-effecting actions |
**3. Reliability checklist** — termination guarantee (it always stops), error recovery, idempotency of
side-effecting actions, and determinism where it matters.
**4. Cost & latency** — where tokens/steps are spent and how to cut them (cheaper model for sub-steps, caching, fewer round-trips) without losing quality. Pair with [`llm-cost-latency-budget`](../llm-cost-latency-budget/SKILL.md).
**5. Safety** — untrusted input/tool output handled as data not instructions, least-privilege tools, and
confirmation gates on high-impact actions. Pair with [`llm-guardrails-spec`](../llm-guardrails-spec/SKILL.md).
**6. Prioritised fix plan** — ordered by impact-to-effort.
## Quality Checks
- [ ] The agent has a guaranteed stopping condition (step/budget cap + progress check) — no unbounded loops
- [ ] Side-effecting actions are idempotent or gated by a confirmation
- [ ] Tools are few and sharply described so selection is unambiguous; access is least-privilege
- [ ] Context strategy keeps the window in budget across steps (no naive full-history resend)
- [ ] Tool errors are recovered, not fatal — retry/backoff and graceful degradation
- [ ] Findings are severity-ranked and the fix plan is ordered by impact
## Anti-Patterns
- [ ] Do not approve an agent with no termination guarantee — "it usually stops" is an outage waiting to happen
- [ ] Do not let it take irreversible actions without a confirmation gate
- [ ] Do not give it many overlapping tools — selection accuracy drops as the toolset grows
- [ ] Do not resend the whole history every step — cost and drift both climb
- [ ] Do not treat tool/retrieved output as trusted instructions — it's the injection surface
## Based On
LLM agent design practice — bounded control flow, least-privilege tool use, context management, error recovery, and safety gating.Conduct a structured ethical review of an AI or ML feature, model, or product. Use when preparing to deploy an AI system, assessing algorithmic risk, auditing a model for bias, or producing a responsible AI impact assessment. Produces a structured ethics review covering fairness, transparency, privacy, safety, accountability, and societal impact with a risk tier score, pre-deployment checklist, and prioritised mitigations.
Structure AI and ML product decisions with the rigour of any product decision. Use when building AI-powered features, evaluating LLM integrations, designing AI products, or assessing AI readiness. Produces a complete AI product canvas covering problem definition, model approach, data requirements, evaluation framework, UX design, responsible AI checklist, and launch monitoring plan.
Transform feature briefs into structured design briefs that give designers the context they need before opening Figma. Use when asked to write a design brief, create a design handoff, brief a designer on a new feature, or translate a PRD into design requirements. Produces a brief with user goal, emotional context, success criteria, constraints, edge cases, and out-of-scope boundaries.
Design statistically rigorous A/B tests and interpret experiment results. Use when asked to design an experiment, run an A/B test, calculate sample size, interpret test results, or assess whether an experiment was successful. Produces a complete experiment design with hypothesis, sample size, run time, success criteria, and risk flags — or a results interpretation with ship/iterate/kill recommendation.
Synthesises user signals from multiple research sources into a unified, weighted insight brief. Use when you have data from interviews, support tickets, NPS verbatims, app reviews, or sales calls and need to reconcile contradictions, surface the underlying need behind requests, or answer 'what are users really telling us'. Produces ranked insights with confidence ratings, source weighting rationale, divergent signal analysis by user segment, and a research gap identification section.
Structure a product data analysis, metric deep-dive, funnel analysis, or cohort study. Use when asked to analyse product metrics, investigate a drop in conversion, explain a data change to stakeholders, or find the root cause of a metric movement. Produces a structured analysis with question, root cause, confidence level, and recommended action.
Interpret product metrics against goals and surface actionable signals. Use when asked to analyse product health, review key metrics, investigate a performance issue, produce a health report, or assess product-market fit signals. Produces a structured health report with RAG status, trend analysis, root cause hypotheses, and prioritised actions.
Structure a retention analysis, churn investigation, or engagement deep-dive for any product team. Use when asked to analyse user retention, investigate churn, measure DAU/MAU, or build a retention improvement plan. Produces a retention snapshot with root cause hypotheses, aha-moment correlation, and prioritised interventions.