The 11PromptPro Prompt Quality Framework
The open methodology behind the prompt quality score.
The Prompt Quality Framework is the open methodology behind the quality score that 11PromptPro computes for every generated prompt. It defines what we measure, how the score is calculated, and — just as importantly — what the score does not mean. We publish it so the number you see in the app is a verifiable measurement, not a black box.
The interview adapts — the yardstick doesn't.
11PromptPro asks you 4–11 adaptive questions — only the ones your specific task actually needs. A simple task gets a short interview; a complex one gets a longer one. Questions that would not improve your prompt are skipped entirely.
The measurement, on the other hand, never changes its ruler. Every prompt — whether it came from a 4-question or an 11-question interview — is graded against the same 7 dimensions, on the same 0–100 scale, every time. Adaptive on the way in, fixed on the way out.
The 7 dimensions
Each dimension is scored independently. Together they cover what makes a prompt work: a clear goal, the missing background, the right perspective, a defined output, explicit boundaries, internal craftsmanship, and enough detail to act without guessing.
1. Clarity
What it measures: whether the goal is stated unambiguously — one objective, zero guesswork about what the result should be.
Why it changes AI output: when a goal is vague, the model has to guess what you actually want — and every guess is a lottery ticket. A precise goal collapses the space of possible answers down to the ones you meant.
Before: “Help me with my marketing.”
After: “Write three subject-line variants for a spring-sale email to existing customers.”
2. Context
What it measures: whether the prompt supplies the background the model cannot know — your situation, your audience, the relevant facts.
Why it changes AI output: the model has never met your customers, read your reports, or seen your product. Without context it fills the gaps with generic averages; with context it can reason about your case instead of any case.
Before: “Write a product description.”
After: “Write a product description for a €40 insulated bottle aimed at bike commuters.”
3. Role
What it measures: whether the prompt places the model in a fitting perspective — a defined expertise or point of view the task calls for.
Why it changes AI output: a well-chosen role activates a consistent vocabulary, register, and standard of judgment for the whole answer. (For reasoning-optimized prompts this dimension is deliberately excluded — step-by-step reasoners perform worse with persona theatrics.)
Before: “Write a letter.”
After: “You are an HR specialist writing a formal rejection letter.”
4. Output format
What it measures: whether the prompt defines the structure of the answer — length, sections, shape.
Why it changes AI output: an undefined format means you spend your time re-cutting the answer instead of using it. Naming the format moves that work into the first draft.
Before: “Explain compound interest.”
After: “Explain compound interest in max. 150 words, as a bulleted list, with one worked example.”
5. Constraints
What it measures: whether the prompt sets explicit boundaries — brevity, tone, language, things to avoid.
Why it changes AI output: constraints are the difference between “an answer” and “a usable answer”. Without them, the model's defaults win — and its defaults are aimed at the average user, not at you.
Before: “Write a job ad.”
After: “Write a job ad — under 300 words, no buzzwords, no salary promises.”
6. Quality
What it measures: the craftsmanship of the prompt itself — internal coherence, no contradictions, no filler.
Why it changes AI output: a contradictory prompt (“short but very detailed, formal but casual”) forces the model to choose which instruction to ignore. A coherent prompt never puts it in that position.
Before: “Keep it short but cover everything in depth, and be formal yet playful.”
After: “Keep it under 200 words; precise, professional tone, one light remark allowed in the intro.”
7. Specificity (Effectiveness)
What it measures: the operational test — would a capable model execute this prompt without needing to ask a single clarifying question?
Why it changes AI output: a chat model cannot ask you back; every missing detail becomes a silent assumption. Specific prompts turn assumptions into instructions.
Before: “Make me a workout plan.”
After: “Make a 3-days-per-week, dumbbell-only workout plan for a beginner with 30 minutes per session.”
How the score is calculated
Per dimension: each of the 7 dimensions is graded from 0 to 100 using a rubric that describes score ranges (e.g. 0–25 vague, 51–75 fairly clear, 76–100 unambiguous) rather than single anchor points, so the grader interpolates instead of snapping to round numbers.
Total: the headline score is a weighted average of the seven dimension scores. The role dimension counts at half weight — the interview derives role fit from context rather than asking for a persona directly, so users cannot influence it explicitly. For reasoning-optimized prompts it is excluded entirely.
Completeness cap: unanswered interview questions limit the ceiling. The maximum possible total is 100 − 60 × (1 − c)², where c is the share of answered questions (counting relevant-but-unasked dimensions as missing). Answer everything and the cap is 100; answer nothing and it drops to 40. Skipped questions don't make a prompt bad — they make parts of it unverifiable, and the cap keeps the score honest about that.
What the score is NOT: it is not a guarantee that the AI's answer will be good, not a benchmark against other tools, and not a measure of the underlying model. It is a relative craftsmanship signal: it tells you how well a prompt is built, and where it can be improved — before a single answer is generated.
FAQ
Is a higher score always better?
Directionally, yes — a prompt scoring 85 is almost certainly better built than one scoring 40. But differences of a few points are within grader noise. Use the score to spot weak dimensions, not to chase a perfect 100.
Does the score measure the AI's answer?
No. It measures the prompt, before any answer exists. A high-scoring prompt can still receive a mediocre answer; a low-scoring one can get lucky. The score grades the instructions, not the outcome.
Why can my score be capped?
If you skip interview questions, the framework cannot verify the dimensions those questions cover. Rather than guessing, the completeness cap limits the maximum total — so an unanswered half of the interview can never masquerade as a perfect prompt.
Is the framework model-specific?
No. The 7 dimensions are model-agnostic — a clear, contextual, specific prompt helps every current model family. Only the weighting of the role dimension adapts: it is excluded for reasoning-optimized prompts, where personas are counterproductive by design.
