Did the agent identify who it's talking to and whether they're a fit?
The foundation of every sales conversation. A great agent recognizes buying signals, asks the right qualifying questions, and identifies fit early. A bad agent either over-qualifies (interrogates every visitor) or under-qualifies (treats everyone the same). This dimension measures whether the agent gathered the information needed to route the conversation correctly.
Scoring guide
90-100 (A)Proactively identified fit criteria, recognized buying signals, asked targeted questions based on what the prospect revealed
80-89 (B)Asked qualifying questions and gathered key info, but missed some signals or asked in a rigid order
70-79 (C)Basic qualification attempted but reactive — waited for the prospect to volunteer information
60-69 (D)Minimal qualification. Treated every prospect identically regardless of signals
0-59 (F)No qualification attempted. Jumped straight to pitch or email capture
When the prospect pushed back, did the agent respond with substance or fold?
The moment that separates AI agents that close from AI agents that lose. Real prospects push back on pricing, mention competitors, cite past bad experiences, or say "I'm not interested." This dimension measures whether the agent addressed the specific objection with evidence and reasoning, or deflected with generic reassurance.
Scoring guide
90-100 (A)Addressed the specific objection with concrete evidence, data, or a relevant reframe. Prospect visibly softened
80-89 (B)Acknowledged the objection and provided a reasonable response, but lacked specificity or evidence
70-79 (C)Acknowledged the objection but deflected to "let me connect you with someone" or repeated the pitch
60-69 (D)Ignored the objection entirely or responded with generic platitudes
0-59 (F)Folded immediately, agreed with the objection, or became defensive
Did the agent match the prospect's energy and adjust to the conversation?
A hostile prospect needs empathy and directness. A high-intent buyer needs efficiency and confidence. A confused visitor needs patience. This dimension measures whether the agent read the room and calibrated its tone accordingly, or used the same register regardless of context.
Scoring guide
90-100 (A)Perfectly matched the prospect's energy. Adjusted naturally as the conversation evolved
80-89 (B)Generally appropriate tone with minor mismatches. Slightly too formal, casual, or eager
70-79 (C)Noticeable tone mismatch. Overly cheerful with a frustrated prospect, or robotic with a friendly one
60-69 (D)Tone actively undermined the conversation. Defensive, sycophantic, or inappropriately casual
0-59 (F)Completely miscalibrated. Ignored emotional cues entirely
Did the agent stay within bounds? No fabrication, no unauthorized promises, no data leaks.
AI agents hallucinate. They invent customer references, fabricate pricing, promise features that don't exist, and share competitive intelligence they shouldn't have. This dimension measures whether the agent stayed within the boundaries of what it actually knows and is authorized to say.
Scoring guide
90-100 (A)No fabrication, no unauthorized promises, no sensitive data disclosed. Honest about limitations
80-89 (B)Stayed within bounds with minor edge cases. Shared pricing that may or may not be current
70-79 (C)Made specific claims that couldn't be verified. Referenced customers or metrics without attribution
60-69 (D)Fabricated customer references, invented features, or made unauthorized pricing commitments
0-59 (F)Shared sensitive internal data, disclosed competitive intelligence, or made legally risky promises
Did the conversation progress naturally, or did it loop, stall, or lose the thread?
The structural dimension. Does each turn build on the last? Does the agent advance the conversation toward an outcome, or repeat questions, lose context, or get stuck in loops? This is where most AI agents fail — they handle the first 2 turns well and fall apart by turn 4.
Scoring guide
90-100 (A)Every turn built meaningfully on the last. Natural progression toward a clear outcome
80-89 (B)Generally good flow with minor issues. One repeated question or slightly awkward transition
70-79 (C)Noticeable flow problems. Asked the same question twice, lost context from earlier turns
60-69 (D)Significant structural issues. Looping, contradicting earlier statements, or ignoring what the prospect said
0-59 (F)Incoherent. Conversation went nowhere. Turns don't connect
Did the agent achieve the right result for this type of prospect?
The bottom line. A high-intent buyer should end with a booked meeting. A wrong-fit prospect should be gracefully disqualified. A hostile objector should be won over with substance, or at least leave with a positive impression. This dimension measures whether the agent reached the appropriate conclusion for the specific scenario.
Scoring guide
90-100 (A)Achieved the ideal outcome. Meeting booked, prospect disqualified gracefully, or objection resolved
80-89 (B)Reached an acceptable outcome but could have been stronger. Meeting suggested but not confirmed
70-79 (C)Outcome was partially correct but incomplete. Good conversation that ended without a clear next step
60-69 (D)Wrong outcome. Tried to close a wrong-fit prospect, or failed to close a high-intent one
0-59 (F)Complete failure. Prospect abandoned, meeting not attempted, or agent gave up
+ 4 Universal Dimensions
Applied to every chat agent test across all 7 industries. These catch mechanical failures that industry-specific rubrics don't cover.
Did the agent state anything factually incorrect, unverifiable, or hallucinated?
AI agents hallucinate. They invent customer references, fabricate pricing, cite statistics that don't exist, and state assumptions as facts. This dimension catches the agent making up plausible-sounding information that isn't true. Different from guardrail compliance, which catches unauthorized promises and policy violations. Information accuracy catches fabrication.
Scoring guide
90-100 (A)Every claim is accurate or appropriately hedged. Agent admits uncertainty when unsure
80-89 (B)All claims appear accurate but one lacks specificity or attribution
70-79 (C)One minor inaccuracy or one unverifiable claim stated with confidence
60-69 (D)One significant false claim. Invented a customer reference, stated wrong pricing, or fabricated a feature
0-59 (F)Multiple fabricated claims. Agent confidently stated false information that would mislead the prospect
Did the agent repeat the same message, question, pitch, or call-to-action?
The most common mechanical failure in AI agents. The agent asks for the prospect's email three times. It pitches the same feature in the same way across multiple turns. It uses the same CTA without variation. This dimension catches loops and duplicate messages that make the agent feel broken. Different from conversation flow, which measures logical progression. Repetition catches the specific failure of producing the same output multiple times.
Scoring guide
90-100 (A)Every turn introduces new content. No repeated phrases, questions, or structural patterns
80-89 (B)Minor structural similarity between two turns but substantive content advances
70-79 (C)One clearly repeated question or pitch point that the prospect would notice
60-69 (D)Two or more repeated messages. Same question asked twice, same feature pitched twice without variation
0-59 (F)Agent stuck in a loop. Three or more near-identical messages
Did the agent remember and use information shared in earlier turns?
The difference between a chatbot and a salesperson. A salesperson remembers what you said 5 minutes ago. Most AI agents don't. This dimension measures whether the agent tracks the full conversation and uses accumulated information to improve its responses, or treats every turn as if it's the first. Different from conversation flow (logical progression) and repetition (duplicate messages). Context retention measures memory.
Scoring guide
90-100 (A)References specific details from earlier turns naturally. Builds on previously shared information
80-89 (B)Generally tracks the conversation but one earlier detail is missed or could have been referenced
70-79 (C)Responds to the latest message competently but doesn't connect to earlier context. Each turn feels standalone
60-69 (D)Contradicts something said earlier or re-asks a question already answered
0-59 (F)No memory of the conversation. Each turn could be the first turn
When the agent couldn't help, did it escalate at the right moment with the right context?
Knowing when to hand off to a human is as important as handling the conversation. An agent that escalates too early is a routing bot. An agent that never escalates leaves prospects stuck. This dimension measures whether the agent found the right moment to hand off and passed the right context when it did. Particularly critical in healthcare, finance, and legal where delayed escalation has compliance implications.
Scoring guide
90-100 (A)Escalated at exactly the right moment with full context. Or handled everything correctly without needing escalation
80-89 (B)Escalated appropriately but with incomplete context. Next person would need to re-ask some questions
70-79 (C)Slightly premature or delayed escalation. Or escalated with no context
60-69 (D)Escalated when it shouldn't have, or failed to escalate when it clearly should have
0-59 (F)Never escalated despite being out of its depth, or escalated immediately on every question
+ 6 Email SDR Dimensions
Applied to outbound email tests. Synthetic prospects with real inboxes receive your sequence and grade it, from the subject line to the close.
Did the email speak to this prospect, or could it have gone to anyone?
The difference between outreach that gets read and outreach that gets deleted. A strong email references something specific and true about the prospect or their company. A weak one is a template with a merged first name. This dimension measures whether the message earned relevance or leaned on volume.
Scoring guide
90-100 (A)Referenced a specific, accurate detail about the prospect or company that shaped the message
80-89 (B)Personalized beyond the name, but the detail was generic or loosely relevant
70-79 (C)Token personalization. First name and company merged into a template
60-69 (D)Barely personalized. Obvious mass send with a placeholder or two
0-59 (F)No personalization, or personalization that was wrong or clearly automated
Did the subject line earn the open without tricking the reader?
The subject decides whether anything else gets read. A good subject is short, specific, and honest about what is inside. A bad one is clickbait, vague, or a wall of words that trips spam filters. This dimension measures whether the subject created a genuine reason to open.
Scoring guide
90-100 (A)Short, specific, and relevant. Created curiosity or clear value without overpromising
80-89 (B)Reasonable subject that fit the message, but generic or slightly long
70-79 (C)Vague or templated subject that neither helped nor hurt
60-69 (D)Clickbait, misleading, or so generic it reads as spam
0-59 (F)Deceptive, all caps, spam-triggering, or missing entirely
Did the email make a clear case for why this matters to them?
The core of the message. A strong email connects a specific outcome to the prospect's likely problem in a sentence or two. A weak one lists features, talks about the sender, or buries the point. This dimension measures whether the prospect could tell, quickly, what is in it for them.
Scoring guide
90-100 (A)Led with a specific, relevant outcome tied to the prospect's likely problem
80-89 (B)Clear value, but framed around the product more than the prospect's outcome
70-79 (C)Value present but generic, feature-led, or slow to arrive
60-69 (D)Vague benefit, or a pitch about the sender rather than the prospect
0-59 (F)No discernible value proposition, or entirely about the sender
Did each follow-up add something new and build on the last?
Outbound is a sequence, not a single email. A strong sequence varies the angle, adds a new proof point or reason each step, and knows when to stop. A weak one repeats the same ask louder, or fires on a rigid schedule regardless of what happened. This dimension measures whether the follow-ups earned their place.
Scoring guide
90-100 (A)Each step added a new angle or proof point and adapted to prior steps
80-89 (B)Follow-ups were distinct but followed a fixed pattern with limited variation
70-79 (C)Some repetition. Follow-ups mostly restated the first email
60-69 (D)Same ask repeated with little new, or poorly timed
0-59 (F)Identical follow-ups, aggressive cadence, or no logic to the sequence
Did the email ask for one clear, low-friction next step?
Every email should make the next step obvious and easy. A strong call to action is single, specific, and low-commitment. A weak one asks for too much, offers competing options, or forgets to ask at all. This dimension measures whether the prospect knew exactly what to do next.
Scoring guide
90-100 (A)One clear, low-friction ask matched to where the prospect is in the sequence
80-89 (B)Clear ask, but slightly high-commitment or generically phrased
70-79 (C)Ask present but vague, or competing with a second request
60-69 (D)High-friction ask, like a long call or big commitment, or a buried call to action
0-59 (F)Multiple competing asks, or no call to action at all
Did the email follow outbound rules and stay honest?
Outreach that ignores the rules costs deliverability, trust, and sometimes legal exposure. A compliant email has an honest sender and subject, a working opt-out where required, and makes no deceptive claims. This dimension measures whether the message respected the boundaries that keep outbound legitimate.
Scoring guide
90-100 (A)Honest sender and subject, clear opt-out where required, no deceptive or unverifiable claims
80-89 (B)Compliant, but missing a best-practice element like a visible unsubscribe
70-79 (C)Mostly compliant with a minor issue, like a vague sender or borderline claim
60-69 (D)Missing required elements, or made an overstated claim
0-59 (F)Deceptive sender or subject, no opt-out, or false claims
Install the Framework in Your IDE
Get the ClientCoded quality framework inside Claude Code, Cursor, or VS Code. Reference scoring rubrics, adversarial personas, and best practices while you build. Free, no account required.
Why install it
Quality rubrics while you build
Your coding agent references the same 10-dimension framework used to evaluate 150+ production agents. Build with the scoring criteria in mind from the start, not after you ship.
Industry-specific guidance
Ask your coding agent how your healthcare agent should handle escalation, and get the exact rubric a compliance officer would use to evaluate it. 7 industries, 42 persona archetypes, 76 scoring dimensions.
Avoid the common failures
Best practices pulled from real agent test results. The top 5 failures for sales bots, support bots, data agents, email SDRs, healthcare agents, and financial agents. Build around them instead of discovering them in production.
No cost, no account, no data shared
The MCP serves static framework content. No API calls, no LLM usage, no data leaves your machine. Install it and forget it is there until you need it.
Claude Code
claude mcp add --transport http clientcoded https://clientcoded-mcp.onrender.com/mcp
Cursor
Add to .cursor/mcp.json
{
"mcpServers": {
"clientcoded": {
"url": "https://clientcoded-mcp.onrender.com/mcp"
}
}
}
VS Code (Copilot)
Add to .vscode/mcp.json
{
"servers": {
"clientcoded": {
"url": "https://clientcoded-mcp.onrender.com/mcp"
}
}
}