Price the Promise You Can Keep

Founder Lessons · No. 01

8 min read

In brief

What two failed aha moments and our own search for a value metric taught me about pricing AI between raw usage and business outcomes.

I was ready to spend more money on a product before it had delivered a single result.

I had been testing Listen Labs for a user-research study. Its pricing was one of the clearest systems I had seen in AI software. My trial account had 50 credits. In the experience I used, one participant cost 10 credits. That meant the trial bought one small study with five people.

Listen Labs recruitment setup showing five responses at ten credits each, fifty total credits, zero of five progress, and an eligible pool labeled Very Large

The unit was immediately legible: 10 credits per response and 50 credits for five. The same screen labeled the eligible pool “Very Large.” Screenshot from my project.

No separate charge for setting up the interview, follow-up questions, transcripts, video, or analysis. One participant, 10 credits. I understood the unit immediately.

I liked it enough to tell my team we should learn from it and find our version of one participant.

I set up several studies and contacted the company because I wanted to buy more credits. But I waited for the first study to run.

Nearly four weeks later, it had recruited no one.

Listen Labs active recruitment dashboard showing zero of five completed responses

The recruit remained active at 0 of 5 completed responses. Screenshot from my project.

Support told me the audience was difficult to find and that they would keep trying. I had completed the setup and felt the excitement of starting. But I never reached the moment that mattered: hearing from the first participant and seeing a useful result.

So I stopped before buying more.

Around the same time, I had the opposite experience with Framer. I was already a paying customer and tried its new Agent workflow on my site. In my workspace, two ambitious natural-language requests consumed the monthly allowance before producing a result I wanted to keep. Framer’s documentation explains that credit use depends on the complexity of the work, and that new AI requests pause when the allowance is gone. That is exactly where I ended up: upgrade, buy more, or wait. How Framer AI credits work

I use coding agents every day. I understand that ambitious requests cost more. But the meter became the main experience before the product gave me an aha moment.

These were two very different pricing systems. One felt beautifully simple. The other exposed the complexity of the machinery. Both lost me at the same place.

The value had not arrived yet.

That changed how I think about pricing and packaging for our own AI company.

The first mistake was pricing the machinery

We began where many AI companies begin: usage.

Tokens. Agent actions. Browser runs. Number of analyses. These are easy to count, and they directly create our costs.

They are also mostly meaningless to our customers.

Our customers do not wake up thinking, “I need 800,000 tokens today.” They ask different questions. Which customer journeys are we missing? What is hurting conversion? What should my team fix or test next? Did the change we shipped actually work?

Tokens are our unit of production, not their unit of value.

It is easy to forget that distinction when every model provider sends a token bill and every product dashboard tracks runs. Usage pricing protects the vendor’s margins because revenue rises with compute. But it can make the customer afraid to use the product when they cannot predict how much useful work one credit will buy.

Stripe’s current guide to AI pricing makes the distinction clearly: the customer-facing metric needs to track how the customer experiences value, while the business still has to account for how costs scale. Those are related problems, but they are not the same problem.

We should measure tokens, runs, model choice, retries, and human review to protect our unit economics.

We should not make the customer learn our factory.

The outcome can be too far away

Once we understood that customers did not care about tokens, the next answer seemed obvious: price the outcome.

In our case, the biggest outcome is revenue. UserApproved analyzes customer behavior, session replay, analytics, live journeys, and other evidence to find high-confidence growth opportunities. If a recommendation could create $500,000 in revenue, why not take a percentage?

Because discovery is not the same as revenue.

The customer still has to agree with the recommendation, prioritize it, implement it well, write the right content, preserve brand and compliance requirements, and run a valid experiment with the right audience, metric, and traffic. Then the result has to survive an internal decision.

We have seen tactics with a plausible online benefit get rejected because they created a conflict with an offline channel. That may be the right business decision. It also has nothing to do with whether our analysis was good.

If we take a share of revenue, we are making both sides negotiate over a result created by a long chain of people, systems, timing, and politics. We would either get paid for value we did not create or fail to get paid for work we did well.

I later found a useful public example in Intercom’s pricing work for Fin. The team considered taking a percentage of closed revenue for qualified sales leads. It rejected the idea for almost exactly this reason: after qualification, human sales execution, product reliability, customer budgets, and other factors take over. Intercom chose the qualified lead instead, the point where Fin has completed work it can clearly own and the customer can define. Why Intercom rejected revenue as the outcome

That gave me a better rule:

Do not price the grandest outcome in the story. Price the last valuable outcome you can reliably own.

A resolved support question can work. A qualified lead can work. A completed legal document can work. Revenue may work when the vendor also controls implementation, distribution, and measurement.

But “outcome-based” is not a shortcut. The outcome needs a boundary.

Every metric trains the company

Hours saved sounds customer-friendly, but it is often a subjective estimate. Whose hours? Compared with which process? What if the customer spends the saved time reviewing more work?

The number of experiment ideas is easy to count and dangerous for us.

Our product strategy is to investigate longer and show fewer, higher-confidence opportunities. If we charge per idea, every additional idea becomes revenue. The pricing model would quietly push the product toward the exact behavior we are trying to avoid: more cards, weaker evidence, and a busier backlog.

A value metric is a product instruction disguised as a price.

The metric influences what sales promises, what the product generates, what customer success celebrates, and what customers request. If it rewards volume, the company will create volume. If it rewards credible coverage and verified decisions, quality has room to survive.

I now test a pricing unit with five questions:

  1. Can the customer explain what they are buying without learning our internal machinery?

  2. Does more of this unit usually create more value?

  3. Can the customer forecast or control how much they will need?

  4. Can we prove that we delivered it without a six-month integration project?

  5. If we optimize this number, will the product become better or merely busier?

The fifth question is the one I missed at first.

The aha moment belongs in the package

Pricing decides more than the invoice. Packaging decides whether the user reaches value before the invoice becomes the story.

With Listen Labs, launching my study was activation. The aha moment would have been the first qualified interview and a useful pattern in the results. I reached the first and not the second.

With Framer, sending an agent request was activation. The aha moment would have been a page or change good enough to keep. I used the allowance before I got there.

This does not mean free trials need to be unlimited. AI has real marginal costs, and abuse is real. It means the trial or starter package should be designed backward from the first trusted result.

How many credits does the intended user need to reach it? If the model fails, recruitment is difficult, or the first attempt takes a wrong turn, does the product help the user recover? Or does it display a meter and leave them on the wrong side of the value?

A free allowance that ends before the aha moment is not generosity. It is an unfinished demonstration.

The best packaging absorbs enough uncertainty for the customer to learn whether the product works for them. Only then does an upgrade feel like expansion instead of a fee to continue the evaluation.

The model I am testing now

I do not think we have solved pricing for UserApproved. I do think we have a better direction.

Our internal costs will remain usage-based. We will track model calls, browser work, data processing, retries, and human review. The customer-facing package should be much simpler.

One model we are exploring is a base subscription with a defined level of journey coverage at a clear cadence. For example: we continuously watch a set of important customer journeys, investigate meaningful changes and opportunities, turn the evidence into qualified decisions, and verify important work after it ships.

Expansion can then follow customer need:

  • more journeys or business units;

  • more frequent analysis and verification;

  • more markets, segments, or data sources;

  • deeper implementation or experiment support; and

  • additional human review where the work requires it.

That is a hybrid model, but the word “hybrid” is not the insight. Lovable uses subscriptions with included credits and top-ups. Semrush lets customers assemble toolkits and add-ons. Those are useful packaging patterns because they combine a stable base with room to grow. See Lovable’s pricing and Semrush’s subscriptions and toolkits.

The harder work is choosing what the base promises and what expansion means.

Whatever unit we choose, coverage and frequency are better packaging axes than a guaranteed idea count. They give the customer budget predictability. They also give customer success one repeatable motion: define the important journeys, confirm the available evidence, reach the first trusted decision, establish a review cadence, and verify what happens after the team acts. Growth can follow broader responsibility, not customers accidentally burning more tokens.

We may still use performance pricing in a narrow engagement where implementation, baseline, traffic, measurement, and attribution are agreed in advance. But that should be an earned layer, not a substitute for defining the work we actually control.

The question we are testing now

We are still experimenting. The next question is more specific: can we find one unit as legible as Listen Labs’ participant?

For UserApproved, that unit may be a protected journey, a completed investigation, or a verified decision. Frequency and coverage may shape the package around it: how many journeys we cover, how often we investigate them, and how deeply we verify what changed.

That is what we need to learn from more customer conversations and real usage. Not which pricing label sounds most modern, but which single unit makes customer value, our unit economics, customer success, and expansion legible to both sides.

What is our version of one participant?

Continue reading

Let’s exchange ideas about technology that helps people live better.

© 2026 Reynold Wu. All rights reserved.

Let’s exchange ideas about technology that helps people live better.

© 2026 Reynold Wu. All rights reserved.

Let’s exchange ideas about technology that helps people live better.

© 2026 Reynold Wu. All rights reserved.