How do I get my website into AI search results?
Make sure the page can be read with JavaScript switched off, put the answer to a real question in the first two or three sentences underneath that question as a heading, name the business in the third person beside the claim, mark the page up with structured data, and check you're not blocking the crawlers. That list is the whole technical job. Most local websites fail the first item before anything else even gets considered.
How an answer engine decides which business to name is a separate question, covered in the post on getting recommended by ChatGPT. This one is the set of things you do to the site.
How do I find out what a crawler can actually see?
Turn JavaScript off in the browser and reload your own website. That's the crude version of the test and it's accurate enough to act on.
GPTBot, ClaudeBot and PerplexityBot largely don't execute JavaScript. Anything that only exists after the page has loaded — text fetched from an API, an FAQ accordion that requests its answers when clicked, a review widget, a whole section that fades in on scroll — isn't there as far as they're concerned. The page looks perfect to you and reads as close to empty to them.
The other version of the same check is to view the page source and search it for a sentence you can see on the screen. If the sentence isn't in the source, it's not in the page.
Fixing it means having the server send the content rather than the browser assemble it. On a site built that way from the start it costs nothing. On a site that wasn't, it's the biggest item on this list, and it's still the one worth doing first.
Am I accidentally blocking them?
Read your robots.txt, which sits at your domain followed by /robots.txt and takes ten seconds to check. Some platforms and security products block AI crawlers by default, and some agencies added those rules deliberately a while ago without mentioning it to anybody.
Look for rules naming GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest. Blocking them is a defensible choice for a publisher whose business is selling its own words. For a local service business that wants to be recommended, it's self-harm.
While the file is open, confirm it points at your sitemap and isn't quietly blocking anything else you care about.
What should the page itself look like?
Plainer than the marketing instinct wants it.
- A question as the heading, in the words a person would say out loud.
- The answer complete in the first two or three sentences. A model lifting two sentences won't go hunting for the rest of the point.
- The business named in the third person beside the claim. "Metcalf Media builds", not "we build". A claim attached to nobody is a claim a model can't attribute, so it uses a page where it can.
- One subject per page. A page about four things is a weak answer to all four.
- Real detail. Figures you stand behind, steps that actually happen, the honest exceptions.
The third-person habit is the one owners resist hardest. It feels stilted to write, it takes an afternoon to apply across an existing FAQ, and it's the cheapest item on this entire list.
What's llms.txt, and do I need one?
A plain text file at the root of a site that tells an AI crawler what the business is, what it does, where it works and what may not be claimed about it. It's a young convention rather than a standard everybody honors, and it costs about an hour.
The valuable half is the constraints. A file stating plainly that the business publishes no pricing, offers no guarantee and serves a named list of towns gives a model something to defer to instead of filling the gap with an assumption. Metcalf Media writes one for every site it builds for exactly that reason.
Does structured data help?
It helps, and it's not the magic part. Schema markup states the facts of a page in a form that doesn't depend on reading the prose correctly: what kind of business this is, where it works, what a service page is about, which questions and answers sit on it.
Worth marking up on a local site: the business itself, the service pages, the questions and answers, the breadcrumbs, and articles. Worth never marking up is a rating you generated yourself. Metcalf Media publishes only real review data in structured form, because a self-published rating is a manual action waiting to happen and isn't eligible for stars anyway.
What would a builder put into all this?
A home builder is a useful case, because the questions are long and the sale runs a season. Somebody typing into a model at ten at night isn't asking who builds houses. They're asking what a custom build actually costs per square foot, how long it takes, what happens if their current house hasn't sold yet, whether anything can still be changed after framing.
Those are page-length answers. They're the questions a builder answers in every first meeting, and almost nobody has put them on a website in a form anything can quote. A builder whose pages answer them — question as the heading, answer up front, company named beside the claim, all of it in the HTML — becomes the available source in a field of brochures.
Place goes on the page too. Somebody searching from Sapulpa is asking about building in Sapulpa, and a model can only attach a company to a town when a page says so and says something real about it.
What do I do first?
In this order, because the first two are free:
- Load the site with JavaScript off and write down what disappears.
- Read robots.txt and remove any block you didn't intend to be there.
- Rewrite the openings of the pages that answer questions. Third person, answer first.
- Add structured data and an llms.txt file.
- Keep adding answers, because coverage is what decides how often a site gets quoted at all.
AI search visibility is mostly that list, done properly, on a site that was already worth finding. If you want somebody to run the checks and tell you which of them a site currently fails, that's a call.
Related questions.
Does my website need to be rebuilt to be readable by AI crawlers?
Only if the content genuinely doesn't exist until the browser builds it, which is the case on a good number of modern templates. Metcalf Media runs the JavaScript-off check first and will say plainly when a site needs a rebuild rather than a rewrite, because the rewrite is wasted effort on a page a crawler can't read.
Should I block AI crawlers from my site to protect my content?
Not if the business wants to be recommended, because a blocked site can't be cited by the thing it blocked. Metcalf Media leaves the crawlers open on a local service site, since the words on those pages exist to win a phone call rather than to be sold as content.
Will adding schema markup get my business into AI answers on its own?
No. Schema clarifies the facts on a page rather than creating a reason to quote it, so it helps a page already answering something plainly and does nothing for one that's not. Metcalf Media adds it to every page and treats the writing as the part that earns the citation.
How often do AI crawlers come back to a website?
Irregularly, and far less predictably than a search engine crawler does, so a change can take weeks to show up in anything. Metcalf Media publishes changes and keeps adding answers rather than watching for a re-crawl, because coverage over time is the lever and a single edit isn't.
Find out where your marketing is losing booked jobs.
We’ll go through your lead sources, your follow-up, your creative and your tracking, and you’ll leave knowing what I’d do in your position — including if the answer is that you shouldn’t do this yet.

