What's the core difference between "letting users make something first" and "explaining the product concept first"?
The core difference is what the user experiences within the first few seconds of the flow: an abstract promise, or concrete evidence. A flow that explains the concept first is essentially asking the user to believe "this product is worth my time" before that trust has been established — the user hasn't seen any actual result yet, only a Block of text or a few illustrative screenshots they have to use to guess what the product might actually help them do.
A flow that lets the user make something first skips the "convince the user to believe" step entirely, letting the user produce a real, visible result themselves — and that result is itself the most persuasive proof possible. There's no need to separately explain "this product is great" when the user just used it to make something.
Why did the original V3 version look logically sound, yet have a skip rate as high as 62%?
There was nothing wrong with V3's design logic — Research-Build-Measure-Learn genuinely is a mature product methodology, and working through it in sequence can genuinely help a user build a complete mental framework. The problem is that this logic only makes sense as a sequence for someone who already believes this product is useful. For a new user who's still evaluating "is this product even worth my time," being asked to do abstract problem definition right away is effectively being asked to invest cognitive cost before seeing any evidence of return.
The team's analysis points directly to this gap: "asking for clarity before showing possibility created friction" — a mismatch between the logical sequence and the user's psychological readiness is the real reason behind the 62% skip rate, not that the flow wasn't detailed or clear enough.
What's the specific logic behind dividing work across three models (ChatGPT, Claude Opus, Gemini Flash)?
The allocation logic assigns tasks according to what each model is relatively good at, rather than randomly splitting work or forcing a single model to do everything. Strategic analysis needs judgment that combines existing data context — that part went to ChatGPT as strategy advisor. Structural UX design for a multi-step flow needs systematic judgment about user behavior and interface logic — that part went to Claude Opus. Prototype generation speed directly affects whether a user loses patience from waiting too long during testing — that part went to speed-prioritized Gemini Flash, with generation time deliberately capped under 60 seconds.
The team's conclusion that "composition matters more than individual model preference" implies: rather than agonizing over which single model is "strongest overall," the more practical approach is to first inventory the subtasks within the whole job, then find the relatively best-suited model for each subtask, letting different models handle the part they're actually good at.
My team is small and can't operate three different AI models simultaneously — how else can this case's experience be applied?
You don't need to fully replicate the specific "three-model division of labor" approach to benefit from this case — the core transferable insight is the "reordering" itself, and that part can be experimented with using a single AI design tool. A concrete way to start: inventory how many steps your current onboarding flow requires before a user sees the product's core result. If the answer is "three or more steps before seeing anything," that's a signal worth prioritizing for a reordering test.
You can use an existing AI design tool to quickly build a simplified first-experience prototype where "the user types one description and immediately sees a result," then run a small-scale A/B comparison against your existing flow, watching whether Day 1 start rate and skip rate shift in a similar direction — rather than overhauling the entire system at once. Validate the effect of reordering first, then decide whether to invest further resources into the finer-grained optimization of splitting work across multiple models.
Many products' onboarding flows follow the same intuition: get the user to understand "what problem this product solves" first, then guide them step by step through setup and data entry, and only then let them actually reach the product's core feature. That order sounds reasonable, but a publicly documented AI-assisted redesign case shows that flipping it — letting users create a concrete result first, then filling in the methodology behind it afterward — produced a gap in retention and completion rates too large to ignore.
The original version in this case (referred to as V3) was fairly well-structured: it guided users sequentially through problem definition, hypothesis formation, prototyping, surveys, and dashboard setup, each step mapping to a standard Research-Build-Measure-Learn methodology framework. The problem was that this flow had a 62% skip rate and under 1% completion. The team's own post-hoc analysis pointed to the core cause: "asking for clarity before showing possibility created friction" — users were asked to do abstract problem definition before they even knew what this product could help them create, and most people gave up right there.
The redesigned version (V4) flipped the order entirely: instead of asking the user "what problem do you want to solve" up front, it let the user directly generate something concrete and visible — an app screen, a landing page, a dashboard — so the user could see their own description turn into a real visual result within seconds. Only after the user experienced that "oh, I can actually make something" moment did the flow gradually introduce the structured Research-Build-Measure-Learn methodology framework.
The change from this reordering was quite concrete: Day 1 retention after launch grew from 12.18% to 18.69%, a 53% lift; the onboarding start rate jumped from 27% to 85.5%; and the skip rate dropped sharply from 62% to 7.8%. Taken together, these three numbers don't say "users became more patient" — they say "users got a concrete sense of value right at the start of the flow, so they were more willing to keep going through the rest of the steps."
Another detail worth recording in this case is that the redesign process itself used multiple AI models divided by role, rather than betting on a single model to do everything: ChatGPT handled strategic analysis, pairing with existing analytics and experiment context to act as a strategy advisor; Claude Opus handled the structural UX design of the multi-step flow; and Gemini Flash handled generation speed, deliberately keeping generation time under 60 seconds to make sure users didn't lose momentum or focus while waiting during prototype testing. The team's conclusion was direct: "composition matters more than individual model preference."
Another detail that's easy to overlook is that this redesign didn't treat data tracking as a feature tacked on after the fact — the GA4 dashboard itself was generated directly from a natural-language description, letting the team run cohort analysis in real time and even have a direct conversation with an AI growth advisor about shifting retention patterns. That means "observing how users move through the flow" and "designing the flow" happened in parallel, rather than analytics being scheduled as a separate round of work after the flow went live.
If your product's onboarding flow currently asks users to fill in data, set preferences, and answer a string of questions before letting them see what the product can actually do, this case's data points to a concrete thing to check: letting users see a concrete, personally-theirs result as quickly as possible, then filling in the methodology afterward, is more likely to retain users than doing it the other way around. Using an AI design tool to quickly produce this kind of "see the result immediately" first-experience prototype costs far less than traditional development, and is worth using as a first step to test whether this reordering applies to your own product.