← Previous · All Episodes
Your Next Customer May Never Visit Your Website Episode 26

Your Next Customer May Never Visit Your Website

· 28:12

|

On the show this week, the most influential visitor to your website this year may never see your homepage. It won't notice your hero image. It will never watch your explainer video. It's an AI agent, and it's shopping and researching on behalf of a human who trusts it, and it's already reading your website thousands of times for every customer that it sends back to you.

So we're gonna look at the numbers behind the web's broken bargain, courtesy of Cloudflare. We'll also trace a 30-year effort to teach machines to read the web and meet the engineer now teaching websites to talk back. And we'll ask what it means to design an experience for a customer with no eyes.

That's all coming up on Marginally Better

Welcome to Marginally Better, a show about business, innovation, and the American economy. I'm Joe Taylor Jr. For 30 years, the web ran on a handshake. Search engines crawled your site, and in exchange, they sent you people, traffic, customers. So what happens when that handshake deal falls apart? And what are we building to replace it?

So let's start with Cloudflare. Now, if you don't know Cloudflare, they're a company that is kind of like a middleman for a big chunk of the web. They provide services that speed up traffic or secure it, depending on the relationship that a company has with the kinds of customers that visit it. It's especially helpful for companies that tend to get attacked.

They're good at stopping denial-of-service attacks and things like that. And there's a couple companies like this out in the world. Akamai is another one, Fastly. But Cloudflare itself sits in front of a huge slice of the world's websites, and it gives it a unique view of who's knocking on all those virtual doors.

And they have enough customers spread out around the world that they've got a pretty significant statistically reliable Data set. So in July 2025, by Cloudflare's count, OpenAI's crawlers were reading roughly 1,100 pages for every one human visitor that they referred back. Anthropic's ratio got clocked as high as 38,000 to one at one point, and other tools like Perplexity could go even higher.

Cloudflare found that 75% of the AI crawling that it could identify, was for training models, collecting content, versus 17% for AI search and just about 3% for agents that were doing something that a user actually asked for. Cloudflare notes that its methodology can undercount referrals from apps, but the overall pie chart is pretty hard to argue with.

The machines are reading way more than they're referring. So Cloudflare decided to change some of its rules. In July 2025, it began blocking AI crawlers by default for any new domains on its platform, and it floated a pay-per-crawl marketplace. Now, a year later, they've produced a progress report, and they admitted that a simple block or allow switch isn't quite enough.

So they've rolled out some categories: search crawlers, user-directed agents, training crawlers, and each one of those is manageable separately, even on free accounts. And starting this September, new ad-supported sites on Cloudflare are going to block training crawlers and agents by default while still welcoming search.

So that means deciding which robots get into your website is now a business decision and not an IT setting. Why does all this matter? Well, Pew Research Center gave us a number. Researchers examined nearly 69,000 real Google searches from 900 US adults. And when an AI-generated summary appears at the top of the results, people clicked through to an actual website just 8% of the time.

That's roughly half the 15% rate on pages without a summary, and that 15% rate is even down from a few years ago. And a quarter of the time, a page with an AI summary ended the person's browsing session altogether. That means they got their answer from that AI overview, and they left And your site may have supplied that answer, but you wouldn't know it, and you never got to meet that potential customer.

Now, if agents are where the customers are, Shopify wants its merchants there too. So last March, Shopify announced agentic storefronts. Those let millions of its merchants sell directly to ChatGPT users with connections to Microsoft's Copilot and Google's AI mode in Gemini managed from a single place. The Google connection runs on something called the Universal Commerce Protocol, the UCP.

That's a shared standard co-developed with Shopify, Etsy, Wayfair, Target, and Walmart, and endorsed by more than 20 other commerce and payment companies. Now, what that is is a common language that lets an AI agent discover products, check out, apply discounts, and manage orders directly with a merchant's systems.

The storefront is becoming something an agent can walk through on your behalf, no screen required. All of this creates a measurement problem. Adobe Analytics reported that traffic referred to US retail sites from generative AI tools grew six hundred ninety-three percent during the 2025 holiday season compared with the year before.

And between October 2024 and May 2026, AI referrals to retail sites grew over one thousand three hundred percent, with travel sites topping two thousand two hundred percent. That's growing fast from a small base, and you can see where the interest lies among AI users. They're shopping, they're looking for recommendations, sure, but they're trying to offload more and more of things like complex travel planning.

But those referrals are only a little bit of what we can see so far. Adobe's answer to all of this is a product that they're calling Brand Visibility, and it's built on nearly three hundred million AI search prompts designed to show companies how often and how favorably they appear inside ChatGPT

Inside ChatGPT, Copilot, Google's AI mode, and Perplexity. Analytics firms like Similarweb now track AI visibility as its own discipline, separate from referral traffic. Your brand can now win or lose a customer inside an AI conversation that you'll never see before your website's analytics register a thing.

There's one thread running through all of this, and it means the customer journey now starts, and sometimes it ends, somewhere you cannot see or control. And so we're asking this week, if an AI agent reads your website today, what would it come away believing about you and your brand?

When we come back, from the first RSS spec in 1999 to websites that hold conversations, we'll examine the 30-year project of an engineer named RV Guha. That's next on Marginally Better

Today's episode raises a question that you need to answer in your business right now. Your website may look clear to the people who visit it, but is the information underneath the interface organized well enough for search engines, AI assistants, and shopping agents to get it right? Our website reality check can show you where the experience breaks down before a visitor or an agent reaches the next step.

It works in your favorite podcast player. It's normally $27, but with the promo code BOOK, B-O-O-K, we're gonna take $20 off. It's just seven bucks with homework that you can follow once a day for the next 10 days. You can start it immediately. It's at websiterealitycheck.com, and you can also follow the link in our show notes

It's Marginally Better. I'm Joe Taylor Jr. In 1999, if you loved a website, you had one option, just keep coming back to it. Check it all the time, bookmark it, refresh it, visit it on Tuesday, come back on Thursday, see if anything changed. The web was a collection of individual doors that you just had to keep knocking on, until an engineer named RV Guha Came along.

And he thought that machines should do that knocking on doors for you. He had spent a lot of the early '90s at Apple, and he was building something called the Meta Content Framework. There was a team there working on this. It was a very early system for describing what a website contains in a form that software could process.

He ended up moving to Netscape. If you're an old head like me, you remember how excited you were to get in a box a copy of Netscape Navigator that you could load into your system. And so he, along with a bunch of senior engineers, moved over to Netscape, continued that work, and it fed into something that was called RSS.

Now, Guha wrote the first RSS specification. It was called RDF Site Summary. That was published in March 1999. And, a bunch of his Netscape colleagues contributed to it. His colleague Dan Libby simplified it into a version that really started to take root all over the web. Now, credit for RSS, which eventually became known as Really Simple Syndication, among other things, credit for this is famously shared and argued about.

Dave Winer was a notorious blogging pioneer, still is. his Scripting News format in the '90s was syndicating updates long before RSS itself had a name, but he kind of took over the format's development at UserLand after Netscape moved on from its interest there. And Dave Winer shipped the versions that most of the web ended up adopting.

He pushed RSS into every corner of blogging and eventually into podcasting. the feed that this show travels on is based on his branch of that family tree. But however you decide to divide that credit, the idea still stuck. a website could start announcing its own updates, and your computer could subscribe the way that you might subscribe to a magazine.

And it sounds quaint now, but it really did rewire how information moved in the '90s. Now, Guha kept pulling his end of the thread for the next few decades. He ended up creating RDF. That's a framework for describing data so that machines can work with it. And then as a Google fellow, he helped set up schema.org.

That's the labeling system that tells a search engine that this number is a price or this block is a recipe. If you've ever seen a search result that shows things like star ratings or show times before you click anything, that's a lot of the plumbing that Guha helped lay.

Each project was a variation on that same idea. The web was built for human eyes, but machines need a translation layer. And that brings us to the present, because Guha, who left Google after about two decades, ended up joining Microsoft as a technical fellow last year. And now he says he's working on what might be the last iteration of that idea.

It's called NLWeb. It was launched at Microsoft as an open project in May of twenty twenty-five, and instead of helping machines read a website, NLWeb lets people and machines talk to a website. So what's that actually mean in everyday terms? Today, your website speaks to visitors through choices that a designer made in advance: menus, filters, navigation labels, search box.

Guha has described NLWeb as freeing users from interacting only through the choices that a designer anticipated. So if you ask a site what you actually want to know in plain language, and it answers from its own content and data. Now there's a second layer. Every NL website can also function as what's called an MCP server.

This is a standard plug that lets external AI agents query the site's information directly if the publisher allows it. One implementation for two audiences. You've got humans who want a conversation and agents running errands for humans. Microsoft says the goal is to help websites interact, transact, and be discovered by AI agents,

Which sounds abstract until you remember the story we talked about in the first block of the show. Shopify is already routing merchants into ChatGPT. Google's Commerce Protocol lets an agent walk a checkout flow on a merchant's servers. The rails are already being laid while we're talking about it But a web where software can converse and transact raises two hard questions, and two more companies are racing to answer them.

The first question comes from Matthew Prince, Cloudflare's CEO, and it's about consent. Back on July 1st, 2025, he was calling it Content Independence Day, and that's the day that Prince flipped Cloudflare's defaults so that new sites block AI crawlers unless the owner says otherwise.

And that's when he proposed that access to content should be something publishers can price. Now, a year later came the more interesting move, the admission that block the bots was too blunt. So a crawler training a model on your content, the search engine hoping to cite you, and the agent trying to buy from you on a customer's behalf are really three different visitors with three different value propositions.

So Cloudflare's new controls treat them that way. So the door to your website now has something it never had before. You could think of it as a guest list. Now, the second question comes from Stripe, and it's about trust. If an AI agent shows up at your checkout with a credit card number, how do you know the human actually authorized this purchase?

Stripe, which worked with OpenAI on a standard called the Agentic Commerce Protocol, that's ACP, they point out that fraud systems built to detect suspicious humans can misfire on legitimate AI agents because automated behavior lacks normal human messiness. All the little signals that commerce platforms look for that indicate that you're a human and not a bot, well, if you've hired a bot to go do your work for you, how's it gonna know if that's real or not?

So the reverse risk is even worse. Imagine an attacker manipulating your agent into buying things that you never wanted. So Stripe's answer to this is a mechanism called shared payment tokens. So instead of handing the agent a card number, the customer's payment method gets wrapped in a token that only works for one seller for a limited time up to a maximum amount.

The agent can complete the errand it was sent on and nothing else. It's a leash in the best sense of that metaphor. Stripe says retailers from Etsy to Revolve to Urban Outfitters, they're already onboarding to its ACP platform. Now, put those three threads together. You've got Guha, who's out here giving websites a voice.

Prince is putting everybody on a guest list, and Stripe is giving their customers a leash for that robot errand runner. None of them is actively trying to kill the human website. Every one of these systems assumes there's still a person at the center with a goal delegating, and there's precedent for this.

Twenty-five years ago, RSS didn't replace websites. It just changed how you got attention for them, and the publishers who adapted early built audiences that the holdouts never quite caught. The pattern is repeating with higher stakes. Some sites say they're going to fight the agents. Some will end up just having to surrender to them.

But the ones that thrive are the ones who are going to actively decide, visitor by visitor, what traffic you want to welcome, what each visitor is welcome to do. Your website is becoming two things at once. It's a place people visit and a service that machines call. Running a business online now means designing for both of those situations and knowing which one your next customer will use first.

Still ahead, how to design a website for the customer that can't see it, and a test to run on your own site this week right after this

One more note before our final block. Everything that we covered today, it's gonna look DIFFERENT again in six months, and that is one of the problems that our team built Experience Helpdesk for. Because when a question lands on your desk, like, "Should we block this crawler?" Or, "Why is an AI assistant misreading our return policy?"

Send it to our team by voice, video, or text, and we're gonna answer within a business day. Our members also get a little UX challenge that we send out on Mondays, plus a library of more than 30 frameworks and checklists that we've been using with our clients for the past decade,

So, go to experiencehelpdesk.com, get all the details, or follow the link in our show notes

It's Marginally Better. I'm Joe Taylor Jr. Let's end somewhere a little strange. I want you to picture your website's newest customer. It's got no eyes, so your beautiful homepage carousel is completely wasted on it. It has no patience for your cookie banner. It does not appreciate that beautiful brand font that you paid for, and it is never, ever, ever gonna sign up for your newsletter.

And it might be the most valuable visitor that you get all week because a human sent it with a task and with a budget. Now, designing for this customer is turning a lot of the old advice inside out, and honestly, a lot of it turns out to be advice that a lot of us should have been following anyway. So start with the information under the interface.

A human can squint at an ambiguous product page and still usually figure it out from context. An agent can't Complete product attributes, accurate metadata, current prices, and structured data. Shopify's guidance for merchants entering agentic commerce reads like a librarian's wish list, and that is the point.

For all of my fellow information architecture nerds out there, this is your moment. Companies that treated content architecture as an afterthought are about to find out just how much of their sales depend on it. Here's my favorite implication of what's happening in this whole shift.

An AI agent cannot follow a return policy that's vague or that contradicts itself. Now, a human customer can read some of your legalese, they can read your vague shipping page, and they're gonna sigh and just call your support line anyway. An agent is gonna just stop. They will fail the task, or they will come back with guidance to not recommend you to their human.

Google's Commerce Protocol has merchants publish their capabilities in a machine-readable profile, including what you support, how you handle payment, what happens after the sale. Fuzzy policy has always been a customer experience problem, and now it's a distribution problem too. there's something almost poetic about that because after decades of burying the fine print and after all the time that folks like me have been saying to simplify your policies and procedures, businesses are now getting a very specific financial incentive to make policies so clear that robots can act on them.

And it's time for you to decide what each visitor to your site may do. Cloudflare's categories give you the vocabulary. It's okay for you to say that you want to welcome the search crawler that cites you. you can negotiate with a trainer that absorbs you, and this is very divisive. But I've talked to clients, some of whom are very eager to get their content into training libraries because that content includes advocacy for their points of view or for their brand.

I have other clients that do not want crawlers to train AI on their content. They consider their content, separate, not to be trained on, and it's up to you to decide what you want to do. But now is the time for you to turn these features on and let the platforms know what they should be doing. And finally, it's okay for you to turn off access to a scraper that's not gonna give you anything back.

There is still some debate as to whether some of these platforms are necessarily gonna honor the rules that you set, but at least now you've got the opportunity to welcome different guests in different ways through the same door, but with different rules.

Now, here's one more thing to think about this week. Measure the influence that you can't see. If an AI assistant compared you to three competitors last night and ended up recommending someone else, your analytics probably recorded nothing. Tools for tracking AI visibility are still young. Adobe's is brand new, and there are a bunch of early-stage startups stalking around this space.

But the question that they're all asking is one that you can start asking yourself now for free. Just sit down, ask a few AI assistants about your own category, and when you do this, do this in a separate browser from what you normally use. Do not use it logged in. Try to simulate as best as possible a total blank slate with no personalization, no customization.

And read what these AI agents are saying about you and your brand. This is the new secret shopping. None of this replaces designing for humans. AI adoption is still fairly minimal among the general public, but this is the first early warning of what's going to be a longer-term trend. The human at the end of this chain is still the point.

The agent is just running their errand, and that's exactly why this matters. Because when somebody's trusted assistant is asking your website a question, the answer that it finds for your customer is the customer experience, even if that customer never sets foot in your store and never reaches your own website.

So I want you this week to give your website a no-eyes test. Strip away all of the visuals in your head and ask, "Is the underlying information clear, current, and complete enough that a machine acting for your best customer would get it right?" And if the answer is no, you've found your next project. And here's one more thing.

Every time you have heard someone like me talk about how important it is to make sure that your content is accessible for folks using assistive technology like screen readers, this is the moment where if you've been making that investment all along, you get a great head start here.

Because the same accessibility standards that help folks access the web are gonna help those helpful agents as well. So that's our show for today. Whether you're selling products, publishing content, or booking appointments, I want you to remember that the web's newest visitors aren't looking at your design, just your substance.

So make sure that holds up. Thanks for listening to Marginally Better. If you like what you've heard, help us out by leaving a quick review on Apple Podcasts. It'll help us spread the word about the show to people just like you who care deeply about great customer experiences. If you want to get behind-the-scenes notes from me and the rest of the team, go to marginallybettershow.com or follow the link in our show notes.

Marginally Better is a Calufrax Radio production with research by Connie Evans. I'm Joe Taylor Jr.

View episode details


Subscribe

Listen to Marginally Better using one of many popular podcasting apps or directories.

Apple Podcasts Spotify Overcast Pocket Casts
← Previous · All Episodes