← Previous · All Episodes
What If Software Only Got Paid When You Succeeded? Episode 25

What If Software Only Got Paid When You Succeeded?

· 20:14

|

On the show this week, for decades, software companies got paid the same way whether you succeeded or failed. You bought the licenses, paid the bill, and if nobody at your company ever figured out how to use the thing, well, that was your problem. But a new generation of companies is making a much riskier promise.

If the software doesn't deliver the results, you don't pay. So we'll look at the support platforms betting their revenue on something they call an outcome, and the surprisingly slippery question of what counts as one. We'll also meet the former Salesforce co-CEO trying to rewrite the business model that he helped build, and we'll visit a 1960s jet engine factory that figured out all of this before software even existed.

That's all coming up on Marginally Better

Welcome to Marginally Better, a show about business, innovation, and the American economy. I'm Joe Taylor Jr. Software pricing used to be simple. You paid for seats, that's the number of people allowed to log in, use the tool, and then you would pay for usage, how much of the product you consumed. Now, this week's stories are about companies trying a third model.

You pay when the software actually produces a result that you wanted. We've got four quick stories, and then we'll spend some time with the man leading the charge. But first, Fin, the AI customer service product from Intercom, now charges ninety-nine cents per outcome with a fifty outcome monthly minimum.

The unit Fin bills for is the result itself, and Fin publishes exactly what it bills for, a resolution, a completed procedure that hands off to a human, a disqualified sales lead. Each of those costs ninety-nine cents, unless you get a qualified sales lead, and that costs nine dollars and ninety-nine cents.

Now, according to Fin's own pricing documentation, a resolution counts when the customer confirms the AI solved their problem or simply stops asking for help after Fin's final answer. That second part is doing a lot of work because the customer's silence after the last message helps determine whether the vendor gets paid.

If Fin asked a clarifying question and the customer went quiet, that doesn't count. If Fin delivered an answer and the customer went quiet, it might, but the wording of one message can decide whether money changes hands. We'll come back to why that matters. Now, Zendesk took a different route to the same destination.

That company has been billing for what it calls automated resolutions. Those are customer requests its AI agents resolved Without a human having to step in. But Zendesk doesn't count every conversation that ends quietly. According to its documentation, a large language model reviews each candidate conversation after the fact and judges whether the customer's request was actually resolved.

Zendesk now separates a contained resolution, meaning the AI handled it and the customer didn't come back, from a verified resolution. That's one that passed the review. Customers get billed only for the verified kind. So one AI talks to your customer, and a second AI decides whether the first one did a good enough job to charge for.

That architecture exists because Zendesk concluded that a conversation ending is not the same thing as a problem being solved. Not everyone is betting on outcomes, though. Salesforce introduced Flex Credits for its Agentforce AI platform in May twenty twenty-five. That is a consumption model where a standard agent action costs twenty credits.

That's about ten cents at rates published as of the day of this recording. Now, an action under these rules is a function that the agent performs. That means looking up a record, updating a contact, kicking off a workflow. An action proves that an agent did something. It doesn't prove that your problem went away.

An agent can flawlessly update a record, open a case, and trigger a workflow while your actual issue still sits there unresolved. Salesforce still offers per conversation and per user pricing, which makes it a useful reminder that the industry still really hasn't agreed on what the billable unit of AI customer service work should be.

Now, the cleanest version of this model that our team could find might come from a smaller company. Chargeflow handles chargeback disputes for merchants. Those are cases where a customer disputes a credit card charge and the bank claws back that money. Chargeflow assembles the evidence, submits the dispute, and charges a twenty-five percent success fee only on funds actually recovered.

If they get no recovery, they get no fee. Now, what makes this one work is that the outcome is decided outside Chargeflow's own walls. The issuing bank rules on the dispute. The recovered dollars land in the merchant account, and nobody has to interpret a customer's silence. That distinction turns out to be the whole ballgame.

Four companies, four definitions of success. If you sell anything, software or not, it's worth asking which of your own metrics would survive this test. Would you still get paid if your customers only paid for results? When we come back, the co-CEO who walked away from Salesforce to bet against the business model he helped build, and the sixty-year-old jet engine deal that says his gamble might pay off.

Stay with us.

Support for Marginally Better comes from Experience Helpdesk, a service from our team at Johns and Taylor. If you run a small or a mid-sized company, your website probably looks fine, but you still suspect it's leaving money on the table. Experience Helpdesk members send me their toughest customer experience questions by voice note, video, or text and get an answer within one business day, usually as a short screen recording walking through their actual sites.

No meetings required. Learn more at experiencehelpdesk.com.

It's Marginally Better. I'm Joe Taylor Jr. Bret Taylor has spent his career inside the machine that built modern software economics. He worked at Google early on. He co-created Google Maps. He co-founded FriendFeed, which Facebook bought, and became Facebook's chief technology officer for a while. He co-founded Quip, which Salesforce bought, and then he rose to co-CEO of Salesforce itself, the company that more than any other taught the world to rent software by the seat, by the month, forever.

So I think it means something that Taylor's next act is a bet against the idea of the seat license. In twenty twenty-three, Taylor and Clay Bavor founded Sierra. That's a company that builds AI agents to do customer-facing work. That's answering questions, processing claims, retaining subscribers. And Sierra prices its product in a way that would have sounded reckless in a Salesforce boardroom not that long ago.

Customers pay when the software achieves specific agreed-upon results. Now, Taylor has argued that autonomous agents demand a different economic model because the product is no longer helping an employee do a process. The product itself is the process, and in his framing, the completed job, not the software seat, becomes the unit that matters.

Sierra says this changes the company from the inside out. When a promised result doesn't happen, that miss shows up in Sierra's own revenue, so sales, product, support, and the engineers building the agents all feel it. The company describes customer success as baked into the profit and loss statement.

Under seat pricing, a vendor can cash your check for years while your team quietly gives up on the product. Under outcome pricing, your failure is their failure. Well, that's the pitch anyway, and this is the part that I'm interested in. Sierra itself admits the hard part. The company acknowledges that outcome pricing is far messier than seats or usage because both sides have to agree on what a successful outcome is, how to measure it, and who should get the credit.

So if Sierra's agent does most of the work but hands the last step to a human, did the agent create the value? If the customer got the right answer but hated the experience, was the outcome still a success? Those used to be philosophical questions for a product team offsite. Outcome pricing moves them into the contract.

Now, if this whole idea sounds like a bit you'd see on HBO's Silicon Valley show, I wanna take you somewhere else. To a British factory floor 60 years ago. Now, in the 1960s, Rolls-Royce, the jet engine maker, uh, they do that in addition to making cars, they developed a service model it trademarked as Power by the Hour.

The modern version is called TotalCare, and here's how it works. An airline doesn't pay separately for every repair, every inspection, every overhaul. It pays Rolls-Royce a fixed rate for every hour the engine spends flying. That arrangement shows up in real contracts. A TotalCare agreement filed with the US Securities and Exchange Commission spells out that Hawaiian Airlines would pay Rolls-Royce based on the flight hours that each covered engine accumulated.

And that kind of arrangement flips the incentives because if the engine sits broken in a hangar, Rolls-Royce earns nothing on it. The company only makes money when that engine is on a wing in the air doing its job. Rolls-Royce says so plainly in its own materials. TotalCare is charged per flying hour, so the company is rewarded only for engines that perform.

The manufacturer takes on the maintenance risk, it monitors engine health, it plans the shop visits because every hour of downtime comes out of its own pocket. Airlines don't have to buy a pile of maintenance services. They buy predictability from a partner whose interests pointed in the same direction as theirs.

So that idea has been proven. Sixty years of those jet engines in the air say that aligning payment with results can work. So why do we think that it took software this long for this idea to catch on? Because it's easier to measure what an engine flying hour actually is. It's recorded by instruments.

It's objective. Nobody has to interpret whether the engine really flew. Compare that with some of the customer service resolutions that we've been talking about so far on this show. Someone, or now some model, has to decide whether the problem was solved, whether the software deserves credit, and whether a customer's silence really means satisfaction or just acceptance.

And the moment a company's revenue depends on that judgment call, that judgment call comes under pressure. There's a name for this problem, and researchers call it a Goodhart effect after the economist Charles Goodhart. And the plain language version goes like this: when a measure becomes a target, it stops being a good measure.

A metric really can improve a system at first, but if you optimize it hard enough, especially when you attach money to it, the connection between the metric and the real goal starts to fray. Now, we've talked about Goodhart on the show before, and this isn't a hypothetical worry. A twenty twenty-three study from the National Bureau of Economic Research looked at Medicare's pay for performance program for dialysis facilities.

That's tied to quality scores. The researchers found incentives produced real effort to improve care, but they also found gaming. Patients whose health would drag down a facility's score were fourteen to seventy-one percent more likely to find themselves moved to another facility So you've got the same incentive, but both behaviors showing up at the same time.

Now, nobody has yet shown Sierra, Finn, or Zendesk cooking their books. I'm not suggesting they are. The dialysis study matters for a totally different reason. It tells us that paying for performance doesn't choose between honest improvement and metric-protecting behavior because it funds both. It can't tell the difference.

That means the measurement system itself has to be built to resist the pressure with transparency, with review, and with the customer's ability to push back. Rolls-Royce settled a big question decades ago. Paying for results can work, but the unsettled question is whether success survives becoming a line item, especially when the company selling you the product is also the company doing the accounting.

If you're evaluating a vendor who promises outcome pricing, borrow the skepticism of a good auditor. Ask who defines the outcome, who measures it, and what happens when you eventually disagree, and the answers will tell you more than the price. Coming up, what a ghosted chatbot, a closed ticket, and your favorite business metric all have in common.

That's next on Marginally Better.

Before we wrap up the show for this week, if this episode has you wondering what done really means on your own website, I would urge you to start with some evidence. Our Website Reality Check audiobook shows you where your site is losing people and what to fix first. It's just $27 one time. You know what?

Let's bring that down to $7. Go ahead and take $20 off with the promo code BOOK, B-O-O-K, and that brings it down to just $7 at websiterealitycheck.com. You'll also find that link in our show notes

It's Marginally Better. I'm Joe Taylor Jr. Let's close out this week with a lighter question that's been rattling around my notes all week. When did silence become a sale? And some of my clients will love this because I am fond of saying silence is acceptance on our Scrum calls. But somewhere in America right now, there is a customer who is staring at a chatbot answer, deciding it is no longer worth the effort, and they go ahead and they close the tab and leave to make dinner or whatever they're gonna do.

And somewhere, one of those billing systems that we talked about earlier is recording that moment as a success. "I'd like my ninety-nine cents, please," says the bot. Now, to be fair to everybody we've talked about this episode, all of those companies have thought through this, whether silence after a clarifying question doesn't count, and if the customer comes back for more help, well, that win should evaporate.

But the situation, I think, points at something that applies far beyond software. Every business runs on some kind of definition of done, and most of us have never really audited ours. So here is a simple ladder that I'll offer up borrowed from how ChargeFlow's model works. Level one is activity. We did the things.

We gathered the documents. We made the calls. We sent the emails. Level two is completion. We measured that the process finished, a dispute was submitted, ticket was closed, project shipped. And level three is the outcome. The money came back. The problem stayed solved. The customer got what they came for.

Most businesses are really good at measuring level one obsessively. A lot of us are good at celebrating level two. A lot of us are squinting at level three from a distance. Those outcome pricing companies, whatever their flaws, at least force themselves up the ladder because their revenue only lives at level three.

So now here's a fourth level for that ladder. ChargeFlow can prove it recovered your money. It can't prove the chargeback shouldn't have happened in the first place, or that the customer relationship survived the dispute. Winning the narrow outcome doesn't tell you the experience was good. A closed ticket can mean you solved it, but it can also mean I gave up on you.

From the inside, those can look identical, but the customer remembers which one really happened. So try this. Pick one metric that your business treats as success, the one that you would put in an investor update or in a performance review, and then ask the question that Zendesk built that second AI to ask.

Did the customer's problem actually get resolved, or did the interaction just end? If you can't tell the difference from your data, congratulations. You've made a useful discovery. It tells you where to get curious next, whether that's a follow-up call, a quick survey, look at who came back a week later through a different door.

Because the outcome economy is coming for more than software pricing. The businesses that do well under it will be the ones that defined done honestly long before anyone forced them to

And speaking of definition of done, our show is done for this week. And whether you're selling AI agents or jet engines or chargeback recoveries, I want you to remember, an ended conversation doesn't mean you've solved the problem. Your customers know the difference. I wanna make sure that you know that your metrics do, too.

Thanks for listening to Marginally Better. If you like what you heard, please help us out. Leave a quick review on Apple Podcasts. It'll help spread the word about the show to people just like you, who care deeply about great customer experiences. If you wanna get behind-the-scenes notes from me and the rest of the team, go to marginallybettershow.com or follow the link in our show notes.

Marginally Better is a Calufrax Radio production with research by Connie Evans. I'm Joe Taylor Jr.

View episode details


Subscribe

Listen to Marginally Better using one of many popular podcasting apps or directories.

Apple Podcasts Spotify Overcast Pocket Casts
← Previous · All Episodes