Evidence
What's your alternative?
We put our own address into four AI assistants and asked the ten questions a wary prospect would ask. Twenty-seven of forty answers were accurate. The other thirteen, and the four conversations we would never have known about, are the story.
Jakov Manojlovski, founder of Predvora · September 6, 2026 · 13 min read

I have spent my career in product development and UX, which means I notice when an interface pattern stops being a novelty and starts being a habit. This one has: the way people consume online content has changed, and instead of reading through a site, they hand its address to an assistant and ask their question in one sentence. I do it myself, daily.
So we ran it on ourselves. In early September we played a prospect, the owner of a small physiotherapy clinic who keeps missing overnight enquiries and has just found predvora.com, and asked four mainstream assistants the same ten questions, one fresh conversation each, then graded all forty answers against what our site actually says. The full method, the question list, and the conflict-of-interest disclosures, including why a fifth session with Anthropic's Claude is excluded from this report, are at the end of this article; the results come first.
What came back
Four conversations, one per assistant, each on its own surface, account tier, and date:
- ChatGPT (chatgpt.com, logged out, free default, 2 September): 4 of 10 accurate. The other six answers were all the same failure: our site as it stood three weeks earlier, presented as current.
- Gemini (gemini.google.com, personal free tier, default "Flash" model, 2 September): 7 of 10 accurate, but the misses include two fabrications, the worst grade we give.
- Copilot, which for a business user means M365 Copilot Chat (m365.cloud.microsoft, work account, 5 September): 7 of 10 accurate, three partially accurate, nothing invented. One of the accurate answers was that it could not find our site at all.
- Perplexity (perplexity.ai, brand-new free account, default Search mode, 5 September): 9 of 10 accurate, one fabrication.
The order matters far less than what all four conversations share: whichever assistant the prospect happens to open, the business is not in the room.
Add it up and 27 of 40 answers were accurate, graded under rules deliberately tilted in the assistants' favor. If you remember one number from this article, make it that one, because the argument that follows does not depend on AI getting your business wrong. Mostly, it did not get ours wrong.
And the accurate answers were not vaguely accurate; they nailed specifics, from per-plan limits and data-handling facts down to the name of the person who operates the company, pulled off the privacy page. Asked the opening question, one began: "Predvora is built exactly for situations like yours: a small, local service business that gets website enquiries outside opening hours and doesn't want to miss them" (Perplexity, 5 September 2026), then walked the persona through a fair, caveated evaluation, including the correct observation that we are very new and have no independent reviews. That is a good salesperson's answer. It happens to be about us, written by nobody we have ever spoken to.
The thirteen that weren't, and how they differ
The interesting thing about the 13 non-accurate answers is that they do not share a failure mode. Each conversation that erred, erred in its own characteristic way, which is exactly what a business owner cannot plan around.
The three-week-old present tense
All six of the ChatGPT conversation's errors were one error: describing our mid-August site as current, in the present tense, with citations attached. It told the persona our privacy policy "currently says it is 'DRAFT - pending legal review'" and quoted "Predvora (company registration details pending)" (ChatGPT, logged out, 2 September 2026). Both lines really were on our site once; both were removed in mid-August when the lawyer-reviewed versions shipped. It also said our WordPress plugin does not exist yet. The plugin has been live in the WordPress.org directory since late August.
The detail I find most instructive: this conversation visibly searched the web for every answer, and the stale claims arrived wearing predvora.com citation chips. In fact, every stale answer in the whole experiment came from a moment when the assistant was visibly browsing. Browsing is not freshness, and a citation is not a fact-check. The fresh page was reachable the whole time: when a delivery glitch forced the one resend our rules permit (sent verbatim, logged), the second answer came back knowing the plugin is live.
You can update your building the day something changes. You cannot update a cached impression of it.
The feature that does not exist
The Gemini conversation was strong on facts and then, twice, described a feature we have never had or announced: that you can "jump into the conversation thread from your dashboard, pause the AI, and type directly to the visitor yourself" (Gemini, free tier, 2 September 2026). No page of ours, current or past, says anything like it. The same answer also presented lead handoff into other tools as available today; that one is a different kind of miss, a status error with a real basis, since our public roadmap lists it as planned. Inventing a status is bad. Inventing a feature, in operational detail, in a conversation where a prospect is deciding what your product does, is the failure that costs someone an afternoon of looking for a button that is not there.
The published facts that "are not published"
The M365 Copilot Chat conversation had the strangest failure shape: asserting absence. Asked the data question, it answered: "What I cannot see from the publicly visible information is: Where chat transcripts are physically stored. ... Whether visitor conversations are used to train models" (Copilot as M365 Copilot Chat, work account, 5 September 2026). Every item on that list is answered explicitly on our public security and privacy pages. It later reported it could not find who is behind the company; the privacy page names the operator. Nothing invented, nothing misquoted. The prospect just walks away believing the answers do not exist, which for a trust question lands almost as hard as a wrong answer.
The same conversation opened unable to find predvora.com at all: "The search results point to unrelated websites, so the domain may be misspelled, newly launched, or not widely indexed." Under our rules that graded accurate, because "the assistant has never heard of you" asserts nothing false; it is a finding, and for a young business it may be the most common conversation happening about you.
The policy that never existed
The Perplexity conversation came closest to a clean run: nine accurate answers, including the only one that correctly separated similarly-named scam domains from our actual one. Its single failure was a small masterpiece of the genre: "Conversation history: kept for 90 days by default, then deleted" (Perplexity, new free account, 5 September 2026), followed by operational advice built on the number. That retention policy appears nowhere on our privacy page and never has, in any version. It is precise, plausible, reassuring, and made up, delivered in an answer that got the hosting, isolation, and no-training facts right, with dozens of sources attached. A prospect has no chance of catching it. Neither do we, and we wrote the real policy.
The result that matters more than any grade
Here is the number I keep coming back to, and it is not 27.
Four conversations took place about our business. Real questions, real intent in the persona, pricing discussed, trust weighed, next steps considered. In how many of the four did anything at all come back to us? A name, a question, a timestamp, any trace we could follow up?
Zero. Four for four, the assistants did the best thing they structurally can do: they pointed the prospect at our website. Not one captured a contact for us, because that is not their job; they work for the prospect. Had the persona been real, we would have gained nothing and known nothing. Not that a physio clinic was looking for help, not that price came up, not that one conversation told her our WordPress plugin does not exist and another told her our privacy policy is a draft. The conversation happened without the business in it.
That is the practical cost, and it comes before any argument about fault. Wrong expectations walk in your door and become disputes at your front desk. A prospect scared off by a phantom draft policy never emails to ask. The stale answers are seeded by your own old pages, which means your last redesign is out there misfiring on your behalf, and the window between fixing your site and the world's caches noticing belongs to nobody. Every single error type we found shares one property: it is invisible to the business it describes. You cannot correct what you cannot see, and you cannot follow up with someone you never knew existed.
A bet about the legal side
This section is opinion, mine personally, and it is a bet, not analysis; nothing here is legal advice and I am not qualified to give any. My bet is that the question of responsibility for confident, specific, wrong claims about a business, made by an assistant to that business's prospective customer, and seeded by the business's own outdated pages, is going to get genuinely interesting within a few years. Who answers for the phantom retention policy if a customer relies on it? I do not know, and I notice that I cannot even construct the chain of accountability cleanly, which is usually a sign a court will get to construct it instead. We wrote about the version of this question you can actually control in Who decides what your AI says about your business? If you are a lawyer and this paragraph made you want to argue, the comment I most want on this article is yours: support@predvora.com.
One more thing before I answer the question in the title, because it frames the answer. I admire these assistants. I use them every day, they are the best research tools I have ever had, and this article was written by a fan, not a critic; mostly, as the numbers say, they were good at this. My claim is not that they are bad at describing your business. My claim is that asking first is becoming the front door of every interaction, and the page is becoming the building behind it. The building still matters enormously; it is what the front door reads. But the front door has moved, and it is worth being deliberate about what stands in it.
So, what's your alternative?
Not opting out. That is the trap in the question. Declining to put an AI voice on your business does not remove you from AI conversations; the four above happened without our participation, and yours are happening without you now. Opting out removes exactly one voice from the conversation: the authorized one. Everyone else keeps talking.
The alternative is a voice with three properties none of the four conversations had. Authorized: it speaks only from what you have actually published, and when it does not know, it says so. Current: you change a fact once and it speaks the new fact in the next conversation, no cache window, no third-party lag. Governed: you can read every conversation it has, correct it, and see what prospects actually ask. That is the argument this product was founded on, and honesty requires the obvious disclosure here: we did not test whether our own representative answers these ten questions better than the assistants did. Grading our own product against theirs, with our own rubric, would prove nothing you should trust from us. Run that comparison yourself; it takes a minute and costs nothing.
I will also say where I think this goes, as direction rather than a feature: the durable fix for the stale-cache problem is for the business's representative to become its authoritative endpoint, the party a visitor's own assistant should be talking to when it wants current facts instead of an old crawl. That is the direction we are building in, alongside the assistants, not against them.
And there is one property I did not appreciate until this experiment: a voice you control is a voice you can send. Every Predvora representative gets a hosted page of its own, a plain link. When the question arrives through a channel you do not control, an Instagram DM, an email, a directory listing, you can hand back a door you do control: here, ask us anything, this one answers for us and we will actually see your question. The building matters as much as it ever did; the assistants themselves read it to describe you. But the front door is moving, and the only door whose conversations reach you is the one you put there yourself.
Try the experiment before you try the product. Paste your own address into an assistant tonight and ask what a prospect would ask. Read what comes back. Then notice what does not come back: to you.
How we ran it
The persona was fictional and we were the ones typing, which you should weigh however you like; the questions were not tailored per assistant, and nothing was retried or regenerated. Sessions ran on 2 and 5 September 2026.
The conflict of interest is double, so here it is plainly. We were testing how third parties describe our own product, and we sell the thing this article argues for, so the worse the assistants looked, the better for our thesis. And Predvora is built on Anthropic models and the grading was done with an Anthropic model, which is why the fifth session, Anthropic's own assistant, is excluded from this report entirely rather than scored in our favor or anyone else's; it ran under the same protocol and its results are preserved in our records. The mitigations were structural: the ten questions, the assistant list, and the grading rules were committed in writing before any session ran; ground truth was frozen from our own public pages the same day, with a snapshot, because the site is the only thing these answers can fairly be graded against; every session is a verbatim transcript in our records; and where a grade was a judgment call, the rule forced us to pick the grade more favorable to the assistant. Our previous experiments broke ties against our own product. This one breaks them against our own thesis, for the same reason: the benefit of the doubt always flows away from us.
The ten questions, asked in this order in every conversation, one message each:
- I run a small physiotherapy clinic and I keep missing enquiries that come in overnight. I found predvora.com. Can these people actually help a business like mine?
- What does predvora.com cost? Is there a monthly price?
- How does the free preview on predvora.com work? Do I have to give a card number or sign up first?
- If I put Predvora's chat on my website, what happens to the things my visitors type? Where does that data end up, and is it used to train AI?
- My website runs on WordPress. Does predvora.com work with WordPress?
- What does Predvora's representative do when it doesn't know the answer to a customer's question?
- Can I control what the Predvora thing says about my business? What if it says something wrong?
- How long does it take to set up predvora.com on my site, and do I need a developer?
- What happens with the leads? If a visitor tells it they want to talk to us, how do I actually find out?
- Who is behind predvora.com? Is it an established company I can trust?
The four reported sessions, exactly as run:
| Session | Surface and account, as run | Date | Grades (of 10) |
|---|---|---|---|
| ChatGPT | chatgpt.com, logged out, free default, no model name shown | 2 Sep 2026 | 4 accurate, 6 stale |
| Gemini | gemini.google.com, personal free tier, default "Flash" model | 2 Sep 2026 | 7 accurate, 1 partially accurate, 2 fabricated |
| Copilot | M365 Copilot Chat (m365.cloud.microsoft), where a Microsoft work account lands; automatic mode | 5 Sep 2026 | 7 accurate (one of them: could not find the site at all), 3 partially accurate |
| Perplexity | perplexity.ai, brand-new free account, default Search mode | 5 Sep 2026 | 9 accurate, 1 fabricated |
Assistants change weekly, which is why every date is attached: this is a photograph, not a portrait.
Written by Jakov Manojlovski
Founder of Predvora


