Method
How we measure AI visibility
Published in enough detail that you can run it yourself without hiring us. If a method can only be trusted when it is hidden, it is not a method.
Vandro Labs is a Generative Engine Optimization agency for astrology and spiritual wellness apps — we work on getting apps named inside ChatGPT, Perplexity and Gemini answers. This is the method we run for paying clients. We publish it instead of hiding it. What we do · Free audit of your app
1. Clean sessions, one per engine
A logged-in assistant with memory enabled measures your own history, not the market. Every audit run is done in a clean session.
- ChatGPT — temporary chat, so nothing is read from or written to memory.
- Perplexity — incognito window, with the “use search history for personalisation” toggle off. That toggle is the personalisation lever most people miss.
- Gemini — incognito, or a Google profile kept separate from daily work.
We keep normal, logged-in accounts for daily work and never audit from them. Mixing the two is the most common way an agency ends up reporting a client's visibility as better than it is.
2. Ask as the customer's customer
Wording changes the answer more than the engine does. We write questions the way an end user types them, and we keep the wording neutral.
Three rules we apply, each of which came from getting it wrong first:
- Neutral phrasing. No adjectives that pre-select a winner, no naming the app being tested. A question containing the answer is a demo, not a measurement.
- Split near-identical wordings. “Best astrology software or app” and “best astrology app” are different questions and return different apps. We run both and log them separately.
- Reproduce the founder's own phrasing. When a founder tells us how they describe their product, we run that exact wording too. It is the fastest way to find out whether the market's vocabulary matches theirs — usually it does not.
3. Three levels of repetition
Because answers vary between runs of an identical prompt, a single run is a screening signal, never a measurement.
| Level | What it is | What it can conclude |
|---|---|---|
| Level 1 — screening | One run per app, many apps | “Worth a closer look” — nothing more |
| Level 2 — evidence | The same question, three or more runs, across at least two different days | A recommendation rate you can put in an email |
| Level 3 — tracking | The full question set, re-run on a fixed schedule | A trend line: did the work move it? |
A client retainer runs at level 3. The free audit we send to a prospect runs at level 2 — and we label it as such, including when the result is that nothing conclusive can be said yet.
4. Visit every cited source, and classify the page not the domain
A citation is only useful once you know what kind of page it is. We open every cited URL by hand and classify the individual page, because one domain can host several different kinds of page.
The same site may publish a genuine editorial review on one URL and a self-serving “top 10” that ranks its own product first on another. Classifying by domain averages those together and produces a tidy, wrong dataset. Our categories:
- App ranking itself on its own comparison page
- App Store or Play Store listing
- Another app blogging about the category
- Directory, aggregator or ranking site — including ones that present themselves as independent
- Genuine editorial publication
- Reddit or forum thread
- Dev-agency lead-generation blog
- YouTube reviewer
We also record whether the page shows signs of being deliberately optimised for AI citation, and whether it discloses its own conflict of interest. In Edition 1, most of the highest-performing pages did not disclose anything.
5. What we will and will not claim
“Not yet seen” is not “invisible”. If an app has not appeared in our runs, we say it has not appeared in our runs. Proving absence would require a far larger question set than anyone has run, and pretending otherwise is how audits become sales theatre.
Third-party revenue estimates are orders of magnitude, not figures. Where we use market data to decide who to study, we label it as an estimate and never present it as a company's accounts.
We separate what we verified from what we inferred. Internally every claim carries one of two labels: verified, or inference. Anything you receive from us that is an inference is marked as one.
6. Corrections we have already had to make
Publishing the method means publishing the times it produced a wrong belief. Three from Edition 1:
- “Biased comparison pages get filtered out.” False. We found multiple pages ranking themselves first, with no disclosure, cited by all three engines — including one that was the single most-cited source for the highest-intent question we tested.
- “Review count is a good proxy for importance.” False. It predicts neither citation nor revenue, and we had been using it to prioritise which apps to study.
- “Identified expert comments on Reddit are pointless.” False. The rule that gets people banned is pretending to be someone you are not; a comment that states who you are and where you work, and is genuinely useful, is allowed — and the comment itself becomes the citable source.
Run it on your own app
Everything above is enough to do a level-2 audit yourself in an afternoon. If you would rather have it done and dated for you, we will run your app through the same question set on all three engines and send the results back by email — free, and no call required.