Almost two-thirds of early-stage products never reach product-market fit. Teams often build before proving real impact or willingness to pay. Founders, PMs, and product designers face many competing ideas and a tight runway.
They need a repeatable, evidence-driven way to prove purpose, pricing, and business viability without shipping a full product.
Process summary: validated purpose in 2–8 weeks
This section gives an ordered checklist to run immediately and get a decision. Follow each step in order and set timers for each experiment.
- Define one impact metric and hypothesis. Measure it with a proxy or brief pilot.
- Map that metric to a business KPI and projected ROI using a cohort model.
- Run three reproducible experiments: landing+pricing, targeted interviews, and a small concierge pilot.
- Predefine statistical pass/fail thresholds, sample sizes, and decision rules.
- Execute for 2–8 weeks and apply the decision checklist to accept, iterate, or kill.
Keep experiments short and focused for quick decision making.
A compact Purpose Validation framework makes decisions repeatable. Treat each hypothesis through the same five evidence tiers and a simple scorecard.
Tier A means paid commitments or refundable preorders. This is the strongest signal.
Tier B means measured per-user impact from a concierge pilot. This shows quantified retention or time saved.
Tier C means an A/B prototype lift on a main conversion metric. Tier D means high-quality survey or willingness-to-pay measures. Tier E means qualitative interviews that identify unmet needs.
Assign points per tier (A=5, B=4, C=3, D=2, E=1). Weight the primary metric and the ROI mapping. Sum points to get a decision score.
Use a threshold such as Accept ≥12, Iterate 8–11, Kill <8. This ties pass/iterate/kill logic to ranked evidence. Teams can collect tiered signals in parallel and aggregate them into one score.
Define one impact metric and map it to a business KPI and dollar-value projection
Pick a single measurable impact metric that ties to a business KPI. The metric must map to retention, revenue, or donor value. Do not pick vanity metrics like raw clicks.
Choose one metric that maps to a dollar or retention change. Without that mapping, purpose claims stay opinion, not business evidence.
What metric ties to revenue or retention?
- Examples: weekly retention lift, reduction in time-to-complete, churn reduction, or donated dollars per user.
- For each metric, map how the behavior change monetizes or cuts cost.
How to write a clear hypothesis
- Draft this: "If users get X feature, then metric M increases by Y% in Z weeks for target cohort C."
- Predefine Y (minimum detectable effect) and Z (time window). The most frequent error is vague benefit statements without a numeric target.
Translate the metric into a KPI and a dollar-value projection. Use conservative assumptions so the decision does not rely on optimistic estimates. Document every conversion step from metric to revenue or cost savings.
ROI per user skeleton:
ROI per user = (Δimpact → Δbehavior) × conversion rate × ARPU × expected user lifetime
Fill each term with conservative percentages and run a break-even scenario.
Example with numbers
Imagine a 15% weekly retention lift among 10,000 users. If ARPU is $10 per month and average life is 6 months, the cohort extra revenue is roughly $9,000 in the first six months. This example shows how retention maps to dollars.
Common errors to avoid
- Defining vanity metrics that do not map to dollars or retention.
- Writing hypotheses without numeric Y and Z.
- Assuming a metric automatically turns into revenue. Always document conversion steps and include conversion rates from comparable products or channels.
Step 3: match hypothesis to cheap, valid MVPs
Select the cheapest MVP that delivers the signal you need. Each hypothesis asks for a different signal: demand, willingness-to-pay, or measured impact.
Which MVP type should be used for each goal?
| Hypothesis |
Best MVP |
Signal |
Min sample |
| Demand and WTP |
Landing + pricing page |
Paid reservations |
50 paid reservations |
| Impact on behavior |
Concierge pilot |
Measured metric per user |
20–200 pilot users |
| Feature lift |
A/B prototype |
Conversion lift |
400 per variant (5 pp MDE) |
What sample sizes to use?
For binary conversion tests use these rules of thumb. Detect an absolute lift of about 5 percentage points with n ≈ 400 per variant. Detect smaller lifts like 2 percentage points with thousands per variant.
How to treat interviews and qualitative signals?
Run 5–12 interviews for directional clarity; run 30+ to reach saturation and enable reliable segmentation. Use the quick route when time is short and the robust route for deeper insight.
Step 4: reproducible experiments and templates
Run three parallel experiments: landing+pricing smoke test, targeted interviews, and a short concierge pilot. Each experiment must have clear metrics and a run window.
Landing + pricing smoke test template
Headline: [Value in one line]
Subhead: [Who benefits and by how much]
Price options: Basic $X / Premium $Y
FAQ: Refund, privacy, limited spots
Tracking: conversions, paid reservations, bounce rate
Run time: 7–14 days of targeted traffic
Interview script template
Intro: 60 seconds, confirm role and context
Problem: Describe the hardest part about X
Current solution: How do you handle it now?
Impact probe: How much time or money does this cost you?
Reaction to solution: Present idea, measure interest, ask for price reaction
Commitment test: Would you pay $X now? Why or why not?
Close: Ask for referral and permission to follow up
Concierge pilot checklist
Recruit: 20–100 users matching target persona
Deliver: Fulfill the solution manually for each user
Measure: Track individual impact metric weekly
Price: Offer discounted paid tier or refundable deposit
Duration: 4–8 weeks
Compute: Cohort uplift and incremental LTV
A single clear experiment result helps decide fast. Keep the plan tight and measurable.
Example experiment with numeric result
A landing+pricing test ran for 10 days with 5,000 targeted visitors. The page had three price points. Paid reservations numbered 62 at $29. Conversion per visitor was 1.24%.
The team had set 50 paid reservations as pass. The test passed the willingness-to-pay threshold. The team then ran a 40-user concierge pilot.
Expand pricing experiments beyond one page by using established WTP methods. Two compact options work well: Van Westendorp price sensitivity and Gabor‑Granger.
Van Westendorp asks four price perception questions and shows an optimal price interval. Gabor‑Granger asks if respondents would buy at specific price points and yields a demand curve.
In practice, run a landing page with randomized price exposures. Measure paid reservations or refundable deposits per price. Then combine that behavioral data with survey WTP outputs.
For example, if 5,000 visitors produce conversion rates of 1.2% at $29, 0.7% at $49, and 0.4% at $79, calculate expected revenue and margin after CAC. Pick the price that maximizes ROI, not the highest price with low conversion.
Adding these methods formalizes willingness-to-pay and pricing experiments into the MVP and impact loop.
Step 5: predefine pass, iterate, kill rules
Set statistical and business thresholds before starting experiments. Use alpha 0.05 and power 0.8 as defaults for A/B tests. Report confidence intervals.
Which business criteria decide move or no move?
Pass requires meeting the primary statistical threshold and achieving minimum projected ROI. Iterate if signals are mixed or fall short of ROI but show promise. Kill if tests fail both statistical and business thresholds.
How to avoid false positives
Register hypothesis, primary metric, and analysis plan before launch. Use a single primary metric. Require at least one secondary support signal such as paid conversions plus retention lift. The most common mistake is changing thresholds after seeing results.
A structured testing loop saves time and money but only if the team ties impact to dollars and enforces pre-registered rules. This method works for consumer and social products except when regulation blocks quick tests. Use conservative assumptions and prefer paid commitments over mere interest.
Tie the minimum detectable effect and sample-size calculation directly to dollars. Pick the smallest effect that makes the product viable and compute the needed sample with alpha=0.05 and power=0.8. For binary A/B tests the rule of thumb of about 400 per variant for a 5 percentage-point lift is useful.
If the required n is impractical, do one of three things: increase the expected effect, switch to a higher-signal MVP like a concierge pilot, or accept the test as directional only. Include a short worked calculation in every plan that shows baseline metric, target MDE, alpha, power, n per arm, and projected incremental revenue.
Experiment flow of decisions and tests
One impact metric
retention, time saved
KPI → ROI model
cohort math
Landing + Pricing
Interviews
Concierge pilot
Pass / Iterate / Kill
predefined rules
Run 2–8 weeks. Prefer paid reservations for WTP. Use cohort analysis to compute incremental LTV.
Errors that ruin purpose validation
Most teams measure vanity metrics instead of impact and revenue signals. Clicks, raw signups, or downloads do not prove purpose. The data show that conversion metrics often overstate lasting value.
The second frequent error is skipping real-money pricing tests. Many rely on email interest only. Email interest often overstates willingness-to-pay by a large margin.
A third common mistake is running underpowered tests without predefining MDE. Small tests produce noise and lead to wrong product bets.
When this method does not apply and alternatives
This loop does not apply to ideas that require long-term R&D, regulatory approval, or clinical trials where incremental MVPs and pricing tests are infeasible. For regulated health devices, pharmaceuticals, or certain public-sector contracts, use formal pilot partnerships, IRB oversight, and staged regulatory filings instead.
If the product needs heavy compliance, consider partnerships with institutions or run anonymized proxy studies. If contracts force features, focus validation on adoption patterns rather than pricing or impact testing.
Before running any experiment, check applicable laws such as FTC advertising rules, CCPA for California residents, COPPA for minors, HIPAA when handling health data, and FDA guidance if medical claims appear. For impact measurement, align with IRIS+ or SROI when seeking impact investors.
Y Combinator and Lean Startup thinking inform this method. Usability benchmarks from Nielsen Norman Group guide proto testing.
Clarification: run the landing+pricing smoke test and start a concierge pilot in parallel. Use the landing/pricing signal for a quick reality check within 7–14 days. Treat concierge pilot measurements as preliminary until the pilot reaches sufficient duration of 4–8 weeks.
Make short-term go or no-go decisions from the paid-reservation signal. Make scaling decisions after the concierge cohort reaches the pre-registered observation window.
Frequently asked questions
How long should a smoke test run?
A smoke test should run 7–14 days for paid reservations. Run longer if traffic volumes are low. The test must get enough conversions to reach the minimum sample target.
What is a safe minimum for paid reservations?
Fifty paid reservations are a practical heuristic for directional evidence, not a universal statistical threshold. The true minimum depends on baseline conversion, the minimum detectable effect, and desired confidence and power. For hypothesis tests, compute sample size with your chosen MDE, alpha=0.05 and power=0.8.
Use 50 only as an early directional checkpoint. Follow up with a powered experiment or larger concierge cohort if you need reliable WTP estimates.
How many interviews are enough?
Five to twelve interviews give directional insight. Thirty or more are needed for segmentation and saturation. Use the smaller range when time is short.
How to set the minimum detectable effect?
Define the smallest impact that makes the product economically viable. Then compute sample size with alpha 0.05 and power 0.8. If required n is impractical, prefer a different test type.
What to do if tests show interest but no paid?
Treat interest-only signals as weak. Run a small paid micro-test with a refundable deposit or a time-limited offer. If paid conversions stay low, redesign value or pricing.
Do pilot results scale to full product?
Pilot results provide directional estimates, not guarantees. Scale risks exist. Use cohort math to project conservatively and plan staged pilots during development.
How to measure social impact quantitatively?
Map impact to a measurable outcome such as dollars saved, hours regained, or donations raised. Use SROI or IRIS+ indicators and translate impact into funder or donor ROI for funding talks.
Closing resources and final checklist
Use this checklist before build:
- One impact metric with numeric hypothesis
- ROI cohort model
- Landing+pricing test with at least 50 paid reservations or meeting conversion MDE
- 5–12 interviews
- Concierge pilot with measured per-user impact and cohort LTV
- Pre-registered pass, iterate, kill rules
Nielsen Norman Group usability guidelines (2019) support rapid prototypes for early feedback. Kickstarter and Indiegogo campaign benchmarks (2022) show that committed backers reveal willingness-to-pay. CB Insights (2019) found that lack of market need causes many startup failures.
Templates above are ready to copy and run. The most common trap is changing success criteria after seeing results. Fix that now by pre-registering the primary metric, the MDE, and the decision rule before launch.