📗 This is the user manual for the Suite for ChatGPT Ads, the tool in the Ninja Scripts panel. Here you can search it, listen to it, highlight it and save it to favourites. It is updated with every release of the tool.
ChatGPT Ads has no Quality Score and there is no conversation report, so when an ad group underperforms the platform doesn't say why. This screen replaces that with what can be measured: the three texts OpenAI matches against the conversation — the group's hints, the ad and the landing page — and how much they talk about the same thing.
The four components
| What it measures | Weight | |
|---|---|---|
| A | Hints ↔ ads: for each hint, the ad that resembles it most. Those below the threshold are orphans | 35 % |
| B | Ads ↔ page: each ad against its own landing page | 30 % |
| C | Hints ↔ page: hints the page doesn't cover | 20 % |
| D | Cohesion among the hints, and whether the group mixes two intents | 15 % |
That yields an index from 0 to 100: ≥ 70 high · ≥ 45 needs work · below that, low. If a piece is missing (no ads, unreadable page), its weight is shared out and the screen says so; with fewer than two pieces, «no data».
What to do with each finding
- Orphan hints → no ad in the group talks about that. Either you write an ad for that intent, or that hint is surplus (🧭 Hints).
- An ad far from its page → the click lands on something that doesn't promise the same thing. Change the page or the ad.
- Hints the page doesn't cover → you're asking for conversations your page doesn't answer.
- Two intents in one group → the most common case and the most profitable to fix: split it into two groups (🧩 Ad groups), each with its hints and its ad.
The table also cross-references the index with each group's real 30-day CTR, and below it says whether the two correlate… but only when there are at least 500 impressions on each side. With two groups and four clicks there is no correlation worth the name, and claiming one would be the opposite of what this index promises.
The two methods, and why which one matters
Similarity between two texts is measured in one of two ways:
| Method | When | How good |
|---|---|---|
| OpenAI embeddings | If you've saved your OpenAI key in «My accounts» | It measures MEANING: it recognises synonyms and the same intent in other words |
| Approximate lexical | If there's no key | It compares words (without accents or plurals). Fine for orientation, but it falls short |
⚠️ The two scales are not comparable, and the panel acts accordingly. Measured on the same texts on the same day, the lexical method comes out 25 to 55 points below embeddings — enough to flip the verdict in all ten groups that were compared.
So with the approximate method the panel shows the index but does NOT judge with it: it doesn't open the «low relevance» recommendation, and the 🩺 Audit sends the relevance area to «can't be judged yet» instead of docking you points. Approximate is worth showing; accusing with it would be unfair.
The OpenAI key costs pennies (embeddings for a whole account are thousandths of a euro) and goes in My accounts.
When it's recalculated
Once a day, on the first read, and only for the groups whose hints, ads or page have changed — or whose last calculation is more than 7 days old. There's a Recalculate button to ask for it right away: it works in the background and the page refreshes itself.
And the honest caveat
It's a diagnosis, not a truth. It measures whether your texts talk about the same thing, which is what the platform matches against the conversation; it doesn't measure whether your offer is any good. A group with an index of 95 may sell little, and one at 40 may do well if its ad is irresistible. It's for knowing where to start looking.