Prompt Quality Score

A fixed seven-axis rubric for scoring any prompt out of 100, and what each axis is worth.

  • Home
  • / Prompt Quality Score

Most prompt advice is a list of tips with no way to tell whether your prompt follows them. A rubric fixes that: the same prompt scores the same every time, so you can change one thing, re-score, and see whether it helped.

This is the rubric FixMyPrompt runs. Seven axes, fixed weights, adding up to 100. You can apply it by hand, or paste a prompt and have it scored.

AxisPoints
Goal clarity20
Audience15
Format and structure15
Constraints15
Context15
Tone and voice10
Examples10
Total100

Goal clarity 20 points

Is the desired outcome stated, and is it measurable?

Scores low

Write me something about our new product.

Scores high

Write a 120-word product announcement for our new scheduling feature, aimed at existing customers, ending with a link to the changelog.

The heaviest axis, because everything else is downstream of it. A model that does not know what success looks like optimises for sounding reasonable, which is exactly the generic output people complain about.

Audience 15 points

Is the intended reader specified?

Scores low

Explain how our pricing works.

Scores high

Explain how our pricing works to a small business owner who has never used a tool like this and is worried about hidden fees.

Audience determines vocabulary, assumed knowledge, and what to leave out. Unstated, the model writes for a generic everyone, which lands with no one.

Format and structure 15 points

Is the shape of the output defined?

Scores low

Summarise this meeting.

Scores high

Summarise this meeting as three sections: decisions made, owners with deadlines, and open questions. Bullet points, no preamble.

Without a stated shape the model picks one, and it usually picks an essay. Most rework is reformatting rather than rewriting.

Constraints 15 points

Are hard rules and exclusions stated?

Scores low

Write a follow-up email.

Scores high

Write a follow-up email under 80 words. No exclamation marks, do not mention discounts, do not ask for a meeting twice.

What the model must not do carries as much weight as what it should. Exclusions are the fastest way to kill the habits you keep editing out by hand.

Context 15 points

Is there enough background to do the job?

Scores low

Draft a reply to this complaint.

Scores high

Draft a reply to this complaint. The customer has been with us two years, the delay was our fault, and we can offer a partial refund but not a full one.

Missing context is where hallucination comes from. The model fills gaps with plausible invention because that is the only thing it can do.

Tone and voice 10 points

Is the style explicit, or anchored to an example?

Scores low

Make it sound good.

Scores high

Warm and direct, the way you would explain it to a colleague. No corporate hedging, no marketing superlatives.

Tone is the most common silent failure. The output is technically correct and still unusable because it does not sound like you.

Examples 10 points

Does the prompt show what good looks like?

Scores low

Write in our house style.

Scores high

Write in our house style. Here is a paragraph we published last month that gets it right: [paste].

One concrete example moves output quality further than a paragraph describing the same thing. Weighted lowest only because most prompts can score well without it.

Scoring one prompt end to end

Take a prompt most people would consider fine:

Write a blog post about our new reporting feature.

Walking the axes in order:

AxisScoreWhy
Goal clarity8 / 20A blog post about a feature is a topic. Nothing says what the reader should do or understand afterwards.
Audience0 / 15Not stated. Existing users and prospects need completely different posts.
Format3 / 15“Blog post” implies prose and nothing else. No length, no sections, no headline treatment.
Constraints0 / 15Nothing excluded, nothing required.
Context2 / 15The feature is named and never described. The model will invent what it does.
Tone0 / 10Unspecified, so you get house-neutral AI voice.
Examples0 / 10None.
Total13 / 100Reads fine, scores terribly.

That gap is the whole point of scoring. The prompt looks reasonable, so when the post comes back bland the natural conclusion is that the model is weak. It is not. Six of seven axes were left blank, and the model filled them in with averages.

The same request, scoring in the 80s:

Write a 600-word post announcing our new reporting feature, for existing customers on the starter plan who have never opened the reports tab. Open with the problem it solves (they currently export to a spreadsheet to see week-over-week numbers), then three short sections: what it does, how to turn it on, and what it does not do yet. Plain and direct, the way you would explain it to a colleague. No superlatives, no “game-changing”, no exclamation marks. End with a link to the docs.

Nothing there is clever. It is the same request with the blanks filled in, which is all a high score ever means.

Reading your score

Under 40 usually means the prompt is a topic rather than an instruction. The 40s and 50s are the common case: a clear goal with no format, constraints, or audience, which is why the output reads generic. Above 80 the prompt is doing its job, and further gains come from examples rather than more instruction.

Score a prompt against this rubric