If you've ever accepted an AI's first answer and moved on without a second look, Claude's Reflect dashboard can now put a number on exactly how much that habit is costing you.
Reflect is Claude's built-in scoring layer. It watches how you interact across conversations and rates your habits on four dimensions, each one named to start with the same letter: Delegation, Description, Discernment, and Diligence.
The scores aren't about whether Claude is performing well. They're about whether you are using it well. A low score on any dimension points to a specific, fixable habit, not a tool problem.
The core insight
Most people don't get poor results from AI because the model is bad. They get poor results because the habits around the model are underdeveloped. Reflect makes those habits visible for the first time.
The four scores, and the fix for each
Score 1
Delegation
Are you handing off tasks that actually fit, or sending things that are too large and vague to succeed? A low score here means the ask itself is the problem before Claude even reads it.
The fix
Break the ask into one clear deliverable before you send it. One output, one goal per prompt.
Score 2
Description
A low score here means your prompts are missing context, format, or constraints. Claude is filling in the blanks with assumptions, and those assumptions usually aren't what you wanted.
The fix
Name the role, the format, and the one thing it must NOT do. Three additions, noticeably better output.
Score 3
Discernment
This one stings the most. A low score means you're taking first outputs at face value and moving on. No checking, no pushback, no verification before the output goes somewhere that matters.
The fix
Before you use anything important, ask Claude to check its own answer for errors first. One extra message, much lower risk.
Score 4
Diligence
Are you revisiting and improving past work, or treating every conversation as one-and-done? A low score here means useful outputs are sitting idle when they could be refined and reused.
The fix
Schedule one re-check pass on anything that matters. Even a single follow-up prompt compounds the quality of earlier work.
What each fix looks like in practice
The four fixes are simple to state and easy to forget. Here is what applying each one actually looks like when you sit down to write a prompt:
Delegation fix in action. Instead of "help me with my marketing strategy," write "write three subject line options for a re-engagement email to lapsed subscribers who haven't opened in 90 days." One deliverable, one scope.
Description fix in action. Add a role, a format, and a constraint to any prompt that previously had none. "You are a plain-language editor. Rewrite the paragraph below as three bullet points. Do not add any information that isn't already in the source text."
Discernment fix in action. After getting an important output, send a second message: "Before I use this, check it for factual errors, unsupported claims, or anything that could be misread." Then evaluate what comes back before you act.
Diligence fix in action. Pick one output from this week that you used once and closed. Reopen it. Ask Claude what could be stronger, what's missing, or how it would improve the draft now that more context exists.
A prompt template that addresses all four at once
If you want a single structure that builds all four habits into each prompt by default, this is a starting point worth saving:
Four-dimension prompt template
Role: You are a [specific role relevant to the task].
Task: [One clear deliverable. One goal. Nothing compound.]
Format: [Bullet list / numbered steps / short paragraph / table, etc.]
Constraint: Do not [the one thing most likely to go wrong or drift].
After you respond: Check your answer for errors, missing context,
or anything that could be misread before I use it.
The last line is the Discernment fix baked directly into the request. You don't have to remember to send a second message because the self-check is already part of the first one.
A quick-reference table for low scores
If you pull up Reflect and see a low number, here is the plain-language translation of what it means and what to do first:
Low Delegation
Your tasks are too broad or too vague. Start by narrowing each ask to a single, concrete output before sending it.
Low Description
Your prompts lack context. Add a role, a format requirement, and one clear constraint to every prompt that matters.
Low Discernment
You're accepting first outputs without scrutiny. Build a self-check step into your prompts or send a follow-up before acting on anything important.
Low Diligence
You're treating AI outputs as finished rather than as drafts. Block time to revisit and improve work that will be reused or shared.
The honest bit
A high score on Delegation doesn't mean your prompts are good overall. It means your task scoping is reasonable. The other three dimensions are still independent and can be low at the same time.
The Discernment fix adds a step. Asking Claude to self-check means one extra message per important output. That is the correct trade-off, not a workaround to skip.
Diligence scores can mislead if you're working on genuinely final outputs. Not everything needs a follow-up pass. The dimension is most useful when applied to recurring, reusable, or high-stakes work.
The Description fix alone won't save a poorly scoped task. All four dimensions interact. Fixing one while ignoring the others produces incremental improvement, not reliable results.
Reflect reflects your habits, not Claude's capability. If your scores are low, that's information about how you're working, not a verdict on the model's performance.
Want one concrete fix per dimension, built for your actual workflow?
If you'd rather have a tailored set of prompt templates and habit adjustments for your specific use cases instead of starting from scratch, book a call and we'll work through it together.