The Experiment
We ran 30 AI productivity brands through an "AI Visibility" scanner that tests how well ChatGPT, Claude, and Gemini know, recommend, and accurately describe them. Think of it as SEO - but optimized for AI brains instead of Google crawlers.
Each brand was scored across four weighted probes by three frontier models:
- Brand Recognition -Does the AI know who you are?
- Recommendation Likelihood -Would it recommend you to users?
- Information Accuracy -Does it get your story right?
- Citation Willingness -Would it cite you as an authority?
The Full Ranking
| Rank | Brand | Score | Label |
|---|---|---|---|
| 1 | ChatGPT | 93.0 | AI Powerhouse |
| 2 | Claude | 90.3 | AI Powerhouse |
| 3 | Grammarly | 90.0 | AI Powerhouse |
| 4 | Notion | 89.4 | AI Powerhouse |
| 5 | Gemini | 88.2 | AI Powerhouse |
| 6 | Jasper | 86.2 | AI Powerhouse |
| 7 | Otter.ai | 86.1 | AI Powerhouse |
| 8 | QuillBot | 85.8 | AI Powerhouse |
| 9 | Fireflies | 85.1 | AI Powerhouse |
| 10 | Wordtune | 84.8 | AI Powerhouse |
| 11 | Perplexity | 84.5 | AI Powerhouse |
| 12 | Character.AI | 84.1 | AI Powerhouse |
| 13 | Rytr | 83.5 | AI Powerhouse |
| 14 | Tome | 83.1 | AI Powerhouse |
| 15 | Writesonic | 82.9 | AI Powerhouse |
| 16 | Sudowrite | 82.9 | AI Powerhouse |
| 17 | Copy.ai | 82.6 | AI Powerhouse |
| 18 | Beautiful.ai | 82.2 | AI Powerhouse |
| 19 | NovelAI | 81.7 | AI Powerhouse |
| 20 | ChatPDF | 80.8 | AI Powerhouse |
| 21 | Elicit | 80.7 | AI Powerhouse |
| 22 | Poe | 78.8 | Well Known |
| 23 | Consensus | 76.9 | Well Known |
| 24 | Scribe | 76.6 | Well Known |
| 25 | You.com | 76.5 | Well Known |
| 26 | Mem.ai | 75.5 | Well Known |
| 27 | Pi | 73.3 | Well Known |
| 28 | Humata | 73.0 | Well Known |
| 29 | Gamma | 71.9 | Well Known |
| 30 | Fathom | 71.8 | Well Known |
Surprise #1: Gemini Is Way Too Nice
Google's Gemini scored literally every single brand above 89. The lowest score Gemini gave anyone was Humata at 89.4. Meanwhile Claude handed out 46s and 50s.
This isn't a bug, it reveals something about how different models "know" brands. Gemini seems to have broader but shallower recall. It recognizes almost every brand name but doesn't always know what they actually do. Claude is the opposite: if it doesn't know you well, it admits it.
The biggest model disagreements were staggering. Gamma scored 91.6 from Gemini but only 46.4 from Claude , a 45-point spread. Pi, Mem.ai, and You.com all had spreads above 40 points.
Surprise #2: AI-Native Brands Aren't Winning
You'd think companies built on AI would dominate AI visibility. They don't.
Perplexity; literally an AI search engine, scored 84.5. Not bad, but not dominant. You.com scored 76.5, placing it in the bottom third. Pi (founded by DeepMind and LinkedIn legends) scored 73.3. Fathom (AI meeting notes) came in dead last at 71.8.
Meanwhile Grammarly (a grammar checker from 2009) scored 90.0 and Notion (a notes app) scored 89.4. The lesson: being AI-native doesn't make you AI-visible. Content history and brand ubiquity matter far more.
Surprise #3: Claude Is the Honest Critic
Claude consistently scored brands 10-20 points lower than Gemini. But Claude's scores also correlated more with how actual users think about these brands.
For example, Copy.ai got a 95.0 from Gemini but only 68.1 from Claude. Claude's explanation was candid: "I have solid foundational knowledge, but my information lacks recent specifics about their current product offerings." That's fair, Copy.ai pivoted from "AI writer" to "GTM platform" and most people missed it.
Claude's lower scores seem to reflect actual knowledge confidence, not random harshness. If you're optimizing for AI visibility, Claude might be the model you need to impress most.
Category Breakdown: Beat Your Neighbors, Not the World
The overall average was 82.1, but category averages tell a different story:
| Category | Average | Range |
|---|---|---|
| Chat / LLM | 84.6 | 73.3 – 93.0 |
| AI Writing | 83.5 | 72.0 – 86.2 |
| Productivity | 83.5 | 75.5 – 90.0 |
| Meeting / Transcription | 81.0 | 71.8 – 86.1 |
| Presentation | 79.1 | 71.9 – 83.1 |
| Research / Document AI | 78.7 | 73.0 – 84.5 |
If you're a meeting transcription tool scoring 80, you're above average for your category even if you're below the overall average. The goal isn't universal perfection, it's category dominance.
What This Means for Your Brand
You don't need a perfect score. You just need to beat your neighbors.
Different models know different things. Gemini knows everyone. Claude knows the winners. If Claude doesn't know you, you're not in the conversation when Anthropic-powered tools answer user questions.
The brands that score highest aren't necessarily the best products , they're the ones with the clearest, most consistently described value propositions across the web. AI models learn from the same content humans read. If your website, PR, and third-party mentions tell a coherent story, the models absorb it.
Check Your Own AI Visibility
We built a tool that runs the same analysis on any brand. It takes about 60 seconds and tests your visibility across ChatGPT, Claude, and Gemini simultaneously.
