The 80-point headline is worse than the 45-point one
Type "email marketing" into the Headline Generator, pick List, and you get twelve headlines back, each with a coloured score badge. The top one — "10 Proven Ultimate Secrets to Boost Your Email Marketing" — scores 80/100, green, "excellent." A real headline like "Why your cold emails land in spam and the 3 fixes that work" scores 45/100 — orange, "needs improvement."
Read those two aloud. The 80 is generic marketing-speak; the 45 is the one a real person clicks. The score isn't measuring quality — it's counting checklist hits, and the checklist rewards familiar power words while punishing the length specificity requires. Stop mistaking it for a verdict.
What the score actually is
The 0-100 number is five things added up, capped at 100. No model, no comparison, no training data — just arithmetic:
| Component | Max points | Rule |
|---|---|---|
| Length | 25 | 40-60 chars = 100%; 30-39 = 75%; 61-70 = 65%; 71-80 = 40%; <30 = 50% |
| Power words | 20 | 8 pts each, capped at 20 (so 3+ words max it) |
| Emotional words | 15 | 5 pts each, capped at 15 |
| Contains a number | 10 | Any digit anywhere — 1 scores the same as 101 |
| Ends with "?" | 5 | Question bonus |
| Word count | 15 | 6-12 words = 15; 4-5 or 13-16 = 10 |
| Non-neutral sentiment | 5 | Positive or negative word majority |
That's the whole engine. The 80-pointer: 25 (56 chars, in band) + 20 (three power words: proven, ultimate, boost) + 5 (one emotional: ultimate) + 10 (has "10") + 15 (9 words) + 5 (positive) = 80. The 45-pointer: 25 (59 chars) + 0 power + 0 emotional + 10 (has "3") + 10 (13 words) + 0 = 45.
The difference isn't quality. It's that spam, cold, fixes, land aren't on the tool's hardcoded word lists, so they score nothing — even though they're the words doing the actual selling.
Specificity loses, familiarity wins
The trap is the word lists. They're fixed: ~50 power words (ultimate, proven, breakthrough, supercharge) and ~50 emotional words (shocking, secret, nightmare, truth). If your headline's strongest word isn't on either list — compliance, ROI, audit, refund, churn — it scores zero. The checklist underrates specific niche language and overrates the generic vocabulary that makes twelve headlines sound interchangeable.
So treat the score as a coverage meter — "did I include a number, a power verb, land in the length band?" — not a quality meter. When score and hook disagree, trust the hook.
The six formulas, and the one job you have to do yourself
The tool offers six shapes (How-To, List, Question, Controversial, Guide, Comparison), each with twelve templates. Pick the one that matches your traffic source and throw the rest away:
- How-To / List for search intent ("how do I…", "ways to…"). Literal wins.
- Question for the feed — curiosity, nobody searches it.
- Controversial sparingly, only when the body delivers the contrarian take. "Why SEO Is Dead" earns the click once; the next nine times it bounces.
- Comparison for decision queries ("X vs Y"), only with a real second option.
- Guide for pillar/evergreen content, not a 400-word news post.
The tool fills {topic} verbatim. Type "email marketing" lowercase and you get "How to Boost email marketing Like a Pro" — broken capitalization it won't fix. Title-case your input, or edit the output.
The 40-60 char band doesn't fit all three platforms
The length score peaks at 40-60, roughly where Google truncates a search title. But the same tool checks three platforms — Google (60), Facebook (40), email (50) — so a 58-char "perfect" headline goes green on Google and yellow "Close/Over" on Facebook, where it's cut. The single band is a search-engine answer to a question the other two answer differently. For a Facebook ad or email subject, aim for the platform you're publishing on, not the green band.
Gotchas
- It's a template engine, not a model. Despite the "AI" name and 600ms spinner, generation is
Math.random()over fixed filler lists — no LLM. Same topic + formula regenerates different headlines each click (random, not seeded), honestly useful, but the "AI" label oversells it. - The year filler is stale.
{year}comes from a hardcoded['2025', '2026']list. A "Top 10 … Strategies for {year}" headline can render "...for 2025" — already in the rear-view. Delete the year when it isn't part of the value. - "Reading level" is average letter count, not reading level. ≤4 letters/word = "simple," ≤6 = "moderate." "SEO" reads simple; "utilise" reads complex. A length heuristic with a misleading label, not a Flesch measure.
- The A/B "Winner" is the same formula on both sides. A/B mode compares two headlines by the identical checklist — which hits more items, not which gets more clicks. Real A/B needs impressions and a click goal; this is a tiebreaker, not an experiment.
Summary
- The 0-100 score is a deterministic checklist — length (25) + power words (20) + emotional words (15) + number (10) + question (5) + word count (15) + sentiment (5) — not a quality verdict. A specific 45 beats a generic 80.
- The word lists are fixed and domain-blind, so the score underrates niche-specific strong words (audit, refund, churn) and overrates familiar power words.
- Pick one of the six formulas to match your traffic source (how-to/list for search, question/controversial for the feed) and delete the rest.
- The 40-60 char "optimal" is a search-engine number; Facebook caps at 40, email at 50 — target the platform you're publishing on.
- Run options through the Headline Generator, treat the score as coverage not verdict, then polish with the grammar checker. For other surfaces: YouTube titles, email subject lines, slogans.