A fixed seven-axis rubric for scoring any prompt out of 100, and what each axis is worth.
/ Prompt Quality Score
Most prompt advice is a list of tips with no way to tell whether your prompt follows them. A rubric fixes that: the same prompt scores the same every time, so you can change one thing, re-score, and see whether it helped.
This is the rubric FixMyPrompt runs. Seven axes, fixed weights, adding up to 100. You can apply it by hand, or paste a prompt and have it scored.
| Axis | Points |
|---|---|
| Goal clarity | 20 |
| Audience | 15 |
| Format and structure | 15 |
| Constraints | 15 |
| Context | 15 |
| Tone and voice | 10 |
| Examples | 10 |
| Total | 100 |
Is the desired outcome stated, and is it measurable?
Scores low
Write me something about our new product.
Scores high
Write a 120-word product announcement for our new scheduling feature, aimed at existing customers, ending with a link to the changelog.
The heaviest axis, because everything else is downstream of it. A model that does not know what success looks like optimises for sounding reasonable, which is exactly the generic output people complain about.
Is the intended reader specified?
Scores low
Explain how our pricing works.
Scores high
Explain how our pricing works to a small business owner who has never used a tool like this and is worried about hidden fees.
Audience determines vocabulary, assumed knowledge, and what to leave out. Unstated, the model writes for a generic everyone, which lands with no one.
Is the shape of the output defined?
Scores low
Summarise this meeting.
Scores high
Summarise this meeting as three sections: decisions made, owners with deadlines, and open questions. Bullet points, no preamble.
Without a stated shape the model picks one, and it usually picks an essay. Most rework is reformatting rather than rewriting.
Are hard rules and exclusions stated?
Scores low
Write a follow-up email.
Scores high
Write a follow-up email under 80 words. No exclamation marks, do not mention discounts, do not ask for a meeting twice.
What the model must not do carries as much weight as what it should. Exclusions are the fastest way to kill the habits you keep editing out by hand.
Is there enough background to do the job?
Scores low
Draft a reply to this complaint.
Scores high
Draft a reply to this complaint. The customer has been with us two years, the delay was our fault, and we can offer a partial refund but not a full one.
Missing context is where hallucination comes from. The model fills gaps with plausible invention because that is the only thing it can do.
Is the style explicit, or anchored to an example?
Scores low
Make it sound good.
Scores high
Warm and direct, the way you would explain it to a colleague. No corporate hedging, no marketing superlatives.
Tone is the most common silent failure. The output is technically correct and still unusable because it does not sound like you.
Does the prompt show what good looks like?
Scores low
Write in our house style.
Scores high
Write in our house style. Here is a paragraph we published last month that gets it right: [paste].
One concrete example moves output quality further than a paragraph describing the same thing. Weighted lowest only because most prompts can score well without it.
Take a prompt most people would consider fine:
Write a blog post about our new reporting feature.
Walking the axes in order:
| Axis | Score | Why |
|---|---|---|
| Goal clarity | 8 / 20 | A blog post about a feature is a topic. Nothing says what the reader should do or understand afterwards. |
| Audience | 0 / 15 | Not stated. Existing users and prospects need completely different posts. |
| Format | 3 / 15 | “Blog post” implies prose and nothing else. No length, no sections, no headline treatment. |
| Constraints | 0 / 15 | Nothing excluded, nothing required. |
| Context | 2 / 15 | The feature is named and never described. The model will invent what it does. |
| Tone | 0 / 10 | Unspecified, so you get house-neutral AI voice. |
| Examples | 0 / 10 | None. |
| Total | 13 / 100 | Reads fine, scores terribly. |
That gap is the whole point of scoring. The prompt looks reasonable, so when the post comes back bland the natural conclusion is that the model is weak. It is not. Six of seven axes were left blank, and the model filled them in with averages.
The same request, scoring in the 80s:
Write a 600-word post announcing our new reporting feature, for existing customers on the starter plan who have never opened the reports tab. Open with the problem it solves (they currently export to a spreadsheet to see week-over-week numbers), then three short sections: what it does, how to turn it on, and what it does not do yet. Plain and direct, the way you would explain it to a colleague. No superlatives, no “game-changing”, no exclamation marks. End with a link to the docs.
Nothing there is clever. It is the same request with the blanks filled in, which is all a high score ever means.
Under 40 usually means the prompt is a topic rather than an instruction. The 40s and 50s are the common case: a clear goal with no format, constraints, or audience, which is why the output reads generic. Above 80 the prompt is doing its job, and further gains come from examples rather than more instruction.
Score a prompt against this rubric