AI Coaching Tool Evaluation Checklist
Use it when...
Use before choosing or piloting any AI coaching tool, for yourself, a team or an organization.
What this is
Chapter 2 sorts AI coaching tools into three layers and gives two evaluation instruments: five key questions and a ten-item Tool Evaluation Checklist to answer before adopting any AI coaching tool. The book's point is that tools without a framework are shiny objects, and that the choice for a leader is an infrastructure decision, not a software purchase.
Use when / Do not use when
Use before choosing or piloting any AI coaching tool, for yourself, a team or an organization. Do not use it to rank vendors. The book gives no scoring or pass mark, and its vendor names may be dated (this page names categories only).
The Instrument
The three layers (categories only)
| Layer | What it is | Strengths (book) | Limitations (book) | Best for | Not good for |
|---|---|---|---|---|---|
| 1. Foundation models | General-purpose AI systems | Flexible, accessible, free or cheap; good for brainstorming and getting unstuck | No memory between sessions; no personalization from your history; no accountability structure; no built-in development frameworks | Ad hoc thinking partner; writing help on difficult emails or presentations; research synthesis; reflection prompts | Sustained development over time; behavior change that needs tracking; skill building that needs structured practice |
| 2. Professional platforms | Integrated workflow tools with AI built in | Live where you already work; context-aware from your documents and projects; low friction | Not coaching-specific; no development frameworks built in | Productivity within existing workflows; documentation and analysis; quick assistance on current tasks | Deliberate practice; feedback loops that track growth; structured development planning |
| 3. Coaching-specific AI | Purpose-built for development | Coaching frameworks for skill building; memory of the development journey; progress tracking; accountability; escalation protocols to a human | Cannot do transformational work or replace human judgment on complex identity or ethical questions | Skill practice with feedback; goal tracking with accountability; habit formation with reinforcement; regular reflection with pattern recognition | Transformation work |
Book's one-line summary of Layers 1 and 2: Layer 2 tools "make you more efficient. They don't make you more capable." Layer 3 subcategories (book): Role-specific (built for a function such as sales, customer success, engineering leadership, product management); Skill-specific (particular capabilities such as executive presence, difficult conversations, strategic thinking); General development (career planning, habit formation, self-improvement across domains).
Where to start (book)
- Individual: start with Layer 1, build the habit of AI as a thinking partner, then explore Layer 3 matched to a development focus.
- Manager: Layer 3 extends reach (available outside working hours; repeated role-play of difficult conversations).
- Leader: an infrastructure decision that shapes how the organization develops capability.
Five key questions (verbatim)
- Does it personalize based on your context? Or does it give everyone the same generic advice?
- Does it track progress over time? Or does every conversation start from zero?
- Does it integrate with your existing workflows? Or is it one more thing you have to remember to use?
- Does it know its limits? Does it tell you when something needs human judgment?
- Does it escalate appropriately? Is there a path to a human coach when needed?
Red flags (verbatim)
- Tools that claim to do everything (they can't)
- Tools with no human escalation path (dangerous for edge cases)
- Tools that feel like chatbots with coaching language bolted on (most of them)
- Tools that don't explain how they work or what data they use
The book adds that the best tools are transparent about capabilities and limitations and position themselves as partners to human coaching, not replacements.
The Tool Evaluation Checklist (verbatim): "Before adopting any AI coaching tool, ask:"
- What specific problem does this solve that we can't solve now?
- How does it integrate with systems we already use?
- What data does it need, and how is that data protected?
- What training will users need to get value from it?
- How will we measure whether it's working?
- What happens when someone needs help the AI can't provide?
- Who owns this tool: HR, L&D, IT, managers?
- What's the path from pilot to scale?
- How does this tool handle bias and fairness concerns?
- What's the exit strategy if this doesn't work?
Categorize tools by use case (verbatim labels)
- Daily reflection: Foundation models, journaling apps with AI
- Skill practice: Role-specific coaching platforms, simulation tools
- Goal tracking: Platforms with accountability features
- Performance analysis: Tools that integrate with performance data
- Career development: Long-term planning and guidance tools
How to run it (Performance Lab suggested practice)
- Name the use case from the list above before looking at any tool.
- Identify which layer the candidate belongs to.
- Answer the five key questions, then the ten checklist items, in writing, with evidence (vendor documentation, trial, internal owner).
- Check each red flag explicitly.
- Record unanswered items as open; do not fill them with assumptions.
- Where item 6 or Key Question 5 has no answer, consider "Escalation Decision Tree and Non-Delegable List" before piloting.
Record sheet (Performance Lab suggested practice)
Tool category / layer: ______ Use case: ______ Date: ______
| # | Checklist item | Answer | Evidence | Open? |
|---|---|---|---|---|
| 1-10 | (as above) | |||
| Key questions 1-5 answered: ______ Red flags found: ______ Decision (pilot / hold / reject) and who decides: ______ |
From Performance Amplified by Chad T. Dyar, Ch.2.