Ask ChatGPT for the best accounting firm for startups and it will give you a confident shortlist in about ten seconds. Some founder, somewhere, is running that exact query right now instead of scrolling through Google results. The question that should keep you up at night is simple: when the query is about your category, does your company appear in the answer?
This is the first installment of a monthly series on AI search visibility. We are writing it because we have something most commentary on this topic lacks: our own server logs. StartupCFO tracks every AI crawler that touches our marketing site, and the numbers surprised us enough that we decided to publish them. This post covers what actually happens when someone asks an AI assistant a question in your category, how to tell whether AI systems are reading your site, and what to do about it.
What Actually Happens When a Founder Asks ChatGPT a Question
When someone types "best accounting firm for startups" into ChatGPT, one of three things happens, and often all three in sequence.
First, the model answers from what it already knows. Every large language model carries a compressed memory of its training data. If your company was written about, reviewed, compared, and discussed across the web before the model's training cutoff, the model may simply know you exist and mention you without touching the internet at all. This is the slowest lever to move, because it depends on content that was published and crawled months or years before the model shipped. But it is also the most durable: a brand baked into the model's weights shows up even when the user has browsing turned off.
Second, the assistant runs a live search. For queries that need current information, ChatGPT search retrieves results from a search index. OpenAI builds part of this index with its own crawler, OAI-SearchBot, and has said its search feature also draws on third-party search providers, with Microsoft Bing the most commonly cited. The practical implication gets missed constantly: if your site is poorly indexed in Bing, you have a handicap in ChatGPT search that no amount of Google optimization fixes. Perplexity works similarly with its own index, built by PerplexityBot.
Third, the assistant fetches specific pages in real time. When the search step surfaces a promising URL, the assistant sends a live fetcher to pull the actual page and read it before composing the answer. For ChatGPT this fetcher identifies itself as ChatGPT-User. For Perplexity it is Perplexity-User. These hits are different in kind from crawler traffic: each one corresponds to a real human, asking a real question, at the exact moment your page was pulled in to help answer it.
The answer the user sees is a synthesis of all three, usually with citations. Whether your brand appears depends on whether you exist in the training data, whether you rank in the underlying index, and whether your pages, once fetched, actually contain a quotable, attributable answer.
The Crawl-to-Click Gap: Our Own Numbers
Here is what this looks like in practice, from our own marketing site's last 30 days.
AI crawlers hit the site 57,630 times. Of those, 8,497 were ChatGPT-User, the fetcher that fires when a real person's ChatGPT query pulls a page in live. That is not background indexing. That is thousands of moments where an actual human asked ChatGPT something, and ChatGPT decided a StartupCFO page was worth reading before it answered.
In the same window, roughly 40 humans clicked through to the site from AI tools.
Sit with that ratio for a second. About 1,400 AI content consumptions for every human visit. AI assistants are reading our content at industrial scale, using it to answer questions, and sending almost nobody. The visitor gets their answer inside the chat window and never needs to click.
If you are still measuring your marketing purely in sessions and clicks, this entire channel is invisible to you. Your analytics dashboard shows 40 visits and shrugs. Your server logs show 57,630 reads. The gap between those two numbers is where an increasing share of your buyers are forming opinions about your category, and possibly about you.
The strategic conclusion is uncomfortable but clear: for AI channels, the unit of success is not the click. It is the mention. When ChatGPT answers "best accounting firm for startups," the win is your name in the answer with a citation next to it. The click, if it comes at all, is a bonus.
How to Tell If AI Reads Your Site
You do not need a vendor tool to answer this question. You need your server logs, or your CDN's analytics if you are behind Cloudflare, Vercel, or similar. Filter requests by user agent and look for these strings. Each one tells you something different.
GPTBot is OpenAI's training crawler. It collects content that may be used to train future models. GPTBot traffic means your content is a candidate for the model's long-term background knowledge. It respects robots.txt.
OAI-SearchBot is OpenAI's search indexer, comparable to a classic search engine crawler. It builds the index that ChatGPT search retrieves from. OAI-SearchBot traffic means you are being indexed for AI search surfacing, which is a prerequisite for being cited in live answers.
ChatGPT-User is the live fetcher. It is not a crawler in the traditional sense; it fires on demand when a user's query needs your page right now. This is the metric to watch. Every ChatGPT-User hit maps to a real query from a real person where your page made the retrieval cut. If this number is growing, your AI visibility is growing, whether or not referral clicks follow.
PerplexityBot is Perplexity's index crawler, doing for Perplexity what OAI-SearchBot does for ChatGPT. Perplexity-User is its live, on-demand counterpart, fetching pages when a user's question needs them.
Other agents worth knowing: Anthropic's ClaudeBot crawls for Claude, and Google-Extended is the token that controls whether Google can use your content for its AI models, separate from normal Googlebot indexing.
One nuance that matters when you read your logs: the training crawlers and index crawlers generally respect robots.txt, but the live fetchers often do not, because they are acting on behalf of a specific user request rather than crawling autonomously. Perplexity has been explicit that Perplexity-User generally ignores robots.txt for exactly this reason. So the three categories are not just different signals; they are different levers. You can block training while allowing search indexing, or allow everything, and the user agents let you make that choice per crawler.
The one-line health check: pull 30 days of logs, count hits per AI user agent, and note which pages they fetch most. If ChatGPT-User and Perplexity-User are hitting specific articles repeatedly, those pages are actively feeding AI answers. If you see only GPTBot and nothing else, you are training data but not yet a source. If you see nothing at all, check whether your robots.txt, WAF, or bot protection is silently blocking them, which is more common than you would think, because many bot mitigation products block AI fetchers by default.
The Playbook: How to Show Up in the Answer
None of this is magic, and most of it overlaps with plain good publishing. But the emphasis shifts. Here is what the evidence and our own logs support.
Write definition-first pages
Assistants assemble answers from passages, not pages. A page that opens with a crisp, self-contained definition of its topic, then expands, gives the model an extractable answer in the first hundred words. Look at how encyclopedia entries are structured and do that. Every important concept in your category deserves a page that a model could quote verbatim and have the quote stand alone. This is why we maintain a glossary of startup finance terms as standalone definition pages, and why posts like our startup accounting 101 guide open by defining the thing before discussing it.
Give assistants something quotable
Analyses of AI answer visibility consistently find that pages with structured lists, concrete statistics, and quotable claims get cited more than pages of unbroken prose. One widely cited study of 10,000 queries put the visibility lift for structured, statistic-rich content at 30 to 40 percent. Whatever the exact number, the mechanism is intuitive: a model composing an answer reaches for specific, attributable facts. "AI crawlers hit our site 57,630 times in 30 days while sending about 40 visitors" is quotable. A vague paragraph about AI traffic trends is not. Publish your own data when you have it. First-party numbers are the single hardest thing for competitors to replicate.
Name your brand inside answerable content
If your best explainer never mentions your company, the assistant can absorb your explanation and cite you namelessly or not at all. This does not mean stuffing your name into every paragraph. It means that when your content makes a claim only you can make, attribute it: your methodology, your data, your recommendation as a firm. Models associate brands with topics through co-occurrence in text. If "StartupCFO" appears repeatedly alongside substantive content about startup accounting and fractional CFO services, that association is exactly what surfaces when someone asks about the category.
Use FAQ content and structured data
Schema.org markup, particularly FAQPage, Article, and Organization schema, gives machines an unambiguous statement of what a page contains and who stands behind it. FAQ blocks also happen to mirror the format of the queries assistants receive: a question, then a direct answer in two to four sentences. Every substantial article on this site carries FAQ pairs for that reason. The honest caveat: structured data is more clearly proven for traditional search features than for LLM citation specifically, but it costs little, and the same question-and-answer discipline improves the prose for human readers anyway.
Publish an llms.txt file
llms.txt is an emerging convention: a markdown file at your site root listing your most important pages with brief descriptions, essentially a curated sitemap written for language models. We publish one for StartupCFO (you can find it at our site root, at the llms.txt path). Honesty requires a caveat here too: no major AI provider has committed to consuming llms.txt, and its real-world impact is debated. We treat it the way early adopters treated XML sitemaps in 2005: cheap, harmless, plausibly useful, and a signal that your site is deliberately machine-readable. Adopt it in an afternoon and move on to the things with stronger evidence.
Be present where AI looks for comparisons
When an assistant answers a "best X for Y" query, it leans heavily on roundups, comparison pages, and listicles, because those pages are pre-digested answers to exactly that query shape. Two moves follow. First, get into the third-party roundups and directories that cover your category, since those are the sources assistants cite for recommendation queries. Second, publish your own honest comparison content. Our guide to the best fractional CFO firms in 2026 exists partly because comparison pages are what retrieval systems reach for when a founder asks the comparison question. If the only comparison content in your category is written by competitors, you already know whose framing the model will repeat.
Keep your content technically reachable
Assistants read raw HTML. Content that only exists after client-side JavaScript renders, or behind a login, or inside a PDF viewer, is invisible to most fetchers. Server-render your marketing pages, keep answers in the HTML, and make sure your robots.txt explicitly allows the AI agents you want. Then verify in your logs that they actually get 200 responses, not 403s from an overzealous firewall.
What Not to Do
Do not block AI crawlers if you sell to people who ask AI for advice. There is a legitimate debate about training data and compensation for publishers whose content is their product. But if you are a B2B startup whose content exists to make buyers aware of you, blocking GPTBot and OAI-SearchBot is unilateral disarmament. You keep your content out of the models and out of the answers, while your competitors' content fills the space. The 1,400 to 1 ratio means AI answers are where the audience is; opting out of the answer does not make the question go away.
Do not keyword-stuff for machines. The temptation is to treat AEO like 2009 SEO: repeat the target phrase, spin up thin pages per keyword, and pad content with question headers that get two-sentence non-answers. Language models are, almost by definition, good at recognizing low-quality text. Content that reads as stuffed to a human reads as stuffed to a model, and the assistants' selection mechanisms increasingly favor sources that look authoritative and specific. Write the genuinely best page on the topic. Everything in the playbook above is formatting and distribution for substance, not a substitute for it.
Do not confuse AI-generated volume with visibility. Publishing 200 thin AI-written posts does not make you citable. Assistants citing sources tend to converge on a small set of pages that answer cleanly and carry evident expertise. One page with original data beats fifty pages of paraphrased consensus.
How We Will Track This: The Series Promise
Most writing about AI search visibility is speculation because the writers cannot see the data. We can, at least for one site, and we would rather publish real numbers from a sample size of one than hypotheticals from a sample size of zero.
So here is the commitment. Every month, we will publish an update from our own logs with a consistent set of metrics:
- Total AI crawler hits, broken out by agent (GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, and others as they appear)
- Live-fetch volume, the ChatGPT-User and Perplexity-User hits that map to real user queries
- Human referrals from AI tools, the actual click-throughs
- The crawl-to-click ratio, currently around 1,400 to 1
- Which pages AI fetches most, since that tells us which content is actually feeding answers
- What we changed, so that when the numbers move we can at least form a hypothesis about why
This month's baseline, once more, for the record: 57,630 AI crawler hits in 30 days, 8,497 of them ChatGPT-User live fetches, roughly 40 human visitors referred by AI tools.
We expect some months to be boring and some conclusions to be wrong in retrospect. That is fine. A public, dated record of real numbers is worth more than confident guesses, and if the ratio collapses or inverts, you will read about it here first.
The Takeaway
AI assistants are already answering your buyers' questions, with or without you. The channel is measurable today, from logs you already have, and the tactics that improve your standing are mostly things a good content operation should be doing anyway: definitions first, real data, honest comparisons, machine-readable pages, and a brand name attached to substance.
We spend our days as a fractional CFO firm helping founders see around corners in their financials, and we run our own marketing the same way: measure first, then act. Our AI finance platform gets built with the same instinct. If you want help with the finance side of that discipline, from your first clean chart of accounts to board-ready reporting, book a free consultation. And if you just want the AI visibility numbers, come back next month. The logs will be waiting.