The bookings went up. Did the product get better?
At our imagined climbing gym, more beginners book after we update the beginner page and checkout, but more of them cancel without attending. We've got to decide what to trust and what to try next.

Remember the climbing gym we imagined in the last article? People were leaving before they booked, so checkout looked like the problem.
Our first guess was that the price or the booking flow was putting them off, until somebody who nearly booked told us the price wasn’t the problem. Every photo showed experienced climbers, and they couldn’t picture what happened in an intro session or shake the worry that they’d be the only beginner in the room.
One conversation wasn’t enough to change the page, so we checked whether instructors heard the same worry and whether it kept reaching the front desk. When it did, the concern stopped looking isolated. Booking and attendance stayed in view because the page was only one part of the whole first visit. That combination gave us a reason to improve it without pretending checkout had been cleared.
We kept the useful checkout clean-up and also changed the beginner page to help curious newcomers picture their first session, book it and still feel prepped enough to follow through and show up. After both changes went live, bookings were higher, but more people cancelled without attending too.
Now it’s Monday morning, and we’re the product manager opening the first report. It looks like good news because 140 of the 1,000 first-time adults who saw the revised page booked an intro session, up from 100 in the earlier group.
Before anyone gets too pleased, a colleague from the front desk brings over a second number. Thirty-five people in the new group cancelled and didn’t attend inside the comparison window, compared with ten earlier, and the front desk handled the refunds and rebooking questions around them. An instructor adds that sessions which looked full earlier in the week have begun with empty places.
We’ve got a page decision to make, so we drag the booking report and cancellation records onto the same screen and reopen the outcome we wrote down before launch. More beginners should book, and they should still feel comfortable enough to turn up. That’s our intended outcome. Booking conversion gives us evidence along the way, but it isn’t the definition of success by itself. The booking number looks better, but the first visit is still an open question, so to make sense of whether these product changes worked, we need to follow the same beginners a little further.
We follow one beginner through the gym
What still has to happen after someone clicks book for us to call this better?
Take one beginner with a Thursday evening booking. They’ve seen the page, pictured the session and chosen a time. From here, they might arrive as planned, change dates and arrive later, or cancel without attending. After a fixed window, we count whether they made it into an intro session. If they did, we can later see whether they choose to come back.
That gives each metric a different job. Booking conversion is our leading indicator because it arrives early and tells us whether more visitors took the next step. Attendance within 30 days is our nearer outcome metric because it measures whether beginners reached the first visit we were trying to make possible. Cancellations and front-desk work are guardrails because they warn us when that movement comes with a worse visit or more work than the gym can carry. Booking only earns its early-signal role while it stays connected to the outcome.
This thought experiment keeps the maths neat. We’re comparing two groups of 1,000 first-time adults. Everyone gets the same seven days to book, and each person counts once.
Anyone who books gets the same 30 days to attend an intro session, even if they change dates. Every original or replacement session falls inside that window, and both groups’ windows have closed before we compare them. Someone who rebooks and attends still counts as attended. Someone who cancels and doesn’t attend inside the 30 days goes into the cancelled group.
Product teams call each group a cohort and each fixed period an observation window. The names are useful because they remind us to keep the people and clocks consistent. If one cohort had different starting rules or more time, the comparison would lean its way.
We’re leaving no-shows out. If somebody simply didn’t turn up without cancelling, we’d keep that separate because the beginner and front desk are dealing with a different situation.
Before we interpret the movement, we check the report’s plumbing. Did each
cohort really see the page version used during its period? Did first-time adult
mean the same thing? Did a page view, booking, cancellation and attendance get
recorded in the same way? Product teams call these instrumentation checks. Here,
they’re simply how we make sure 140 and 100 were counted on the same terms.
Now the gap is easier to see. We’ve got 40 more bookings, 25 more people cancelling without attending and 15 more people walking into the gym.
The 35 is a count of people, not every message or call the front desk handled. Even so, each one can leave the front desk sorting a refund, looking for another time or trying to refill the place. The count tells us the volume of people and work the gym may need to carry.
Rates also make us name what each percentage is out of. That’s the denominator. Booking uses page visitors, cancellation uses booked people and attendance returns to page visitors. Ten of the earlier 100 booked people cancelled without attending, while this time it was 35 of 140. The rate rose from 10 per cent to 25 per cent, so the extra cancellations aren’t only the inevitable result of having more bookings.
Attendance was 1.5 percentage points higher in the second group, which gives us 15 extra arrivals. Reaching the gym is closer to the outcome we wanted, but it doesn’t tell us whether a beginner felt prepared, whether the session matched the page or whether the visit felt worthwhile. A difference of 15 attendees between two groups of 1,000 is small enough to arise through ordinary variation, so we can’t treat it as evidence of a stable improvement.
Those numbers can’t prove that the page caused it either. This is a before-and-after look at two groups, not a random split.
We can’t fill the return row yet. This is a second clock, starting after someone attends. The gym uses 30 days because that fits the way its beginners usually come back. Somebody who attended last Tuesday hasn’t had the same chance to return as somebody from the earlier group who attended six weeks ago.
We could fill the row now and make the report look finished, but we’d be comparing people who’ve had different amounts of time. So we leave it blank and write down when to come back.
Back at the screen, the numbers aren’t competing verdicts. Booking tells us more visitors crossed the booking step, attendance tells us a few more reached the gym, cancellations show that more booked people dropped out and return isn’t ready yet. Product success here would mean improving the first visit without a customer or operational trade-off we’d decided was unacceptable. We can’t say that happened, but we can ask where the additional cancellations are gathering.
The evening sessions keep coming up
The return row tells us to wait, but the outcome and front-desk records can still tell us where to look today. We sort the people who cancelled without attending by their original session time, and evenings climb to the top. Evening sessions regularly fill, and they made up a larger share of bookings in the new group. A higher share of the people booked into those sessions cancelled without attending than before, and their original places had been booked further ahead than in the earlier group.
We’re sorting the same result by original session time, which product teams call segmentation. Evening sessions rise to the top, so session time gives us a way into the problem. The cancellation rate within those bookings and how far ahead they were made become diagnostic metrics. They show us where the pattern sits and which explanation to check, not whether the whole launch succeeded.
One Thursday session appears full several days out. Then one person cancels. The Thursday place becomes available again, but another beginner may not see it or have enough time to take it. The instructor begins with a gap in a session that looked full on the earlier roster.
We still don’t know why that person cancelled.
Maybe the revised page set an expectation the evening session didn’t meet. Maybe the revised page helped people who were less certain cross the booking step, and the longer lead time for evening sessions left more room for plans to change. Maybe useful times are scarce, so people reserve a place before they’re sure. Maybe reminders arrive after plans have already changed. A local promotion may also have brought in a different mix of beginners.
The timestamp gets us only so far. It shows how far ahead the booking was made and when the cancellation arrived, but not why. So we go to the front desk and ask whether the person tried to rebook and whether the place filled again. Then we ask the instructor what confused beginners who arrived and whether the session matched the picture our page had set.
When those records run out of answers, we speak with a small mix of people who cancelled early, cancelled late and kept an evening booking.
We ask what each person thought they’d booked, when they realised Thursday wouldn’t work and whether another session felt possible. Three people mention that the reminder arrived after their plans had already changed. We’ve found a useful lead, but we still don’t know how common that explanation is across all 35 people who cancelled without attending.
We’re now checking the same problem from several angles, which product teams often call triangulation. The records show where the pattern sits, the front desk and instructor show what it creates, and the conversations suggest explanations worth testing. Three similar accounts make the reminder lead more interesting, but they don’t make it the answer for all 35 people.
The page still hasn’t proved it caused the change
We didn’t randomly decide who saw each version, so this is an observational comparison. Clean tracking helps us trust what moved, but it doesn’t give us causal attribution, the stronger claim that one particular change produced the movement.
Even with clean tracking, the second group isn’t a rerun of the first. The checkout clean-up went live with the page. The season, a promotion, tighter evening availability or a different mix of visitors could also explain some of what moved.
So the honest sentence we can write is that the booking rate and the share of booked people who cancelled without attending were higher after the combined launch. We can’t say the page caused either one.
With those checks done, another checkout change drops down the list. Evening availability and whether the page matches the session stay under investigation. Reminder timing moves first because it appeared in the conversations and gives us a small change we can compare without moving everything else at once.
We make a smaller call
The front desk says it can carry the extra requests for another two weeks. The page is easy to pull back, and 15 more people still made it through the door. So we keep the page for now and leave checkout alone. If empty places or front-desk work get worse before then, we reopen the decision sooner.
That reversibility matters. We can keep learning because the page is easy to pull back, the front desk can carry the work for now and we’ve got an earlier review if either starts causing trouble. A decision that could do more harm or would be difficult to undo would need a stronger case before we moved.
Instead of touching checkout, we use a set of upcoming evening sessions to compare the current reminder with the same reminder sent earlier. We assign each session at random, so everyone booked into it gets the same timing. The reminder timing is the only planned difference, and people can cancel or rebook in the same way either way.
The reminder comparison can do something the launch couldn’t. We’re running a small randomised experiment, assigning whole evening sessions at random, so the session is our unit of randomisation. We do that because our main result is the average share of places still empty when each session begins. Splitting reminder timings inside one session would mix both versions inside the same capacity and rebooking problem. That empty-place measure includes no-shows, even though we kept them separate when we were counting people earlier.
We keep each beginner with their original session’s reminder group when we read the result and record anyone who moves into the other timing after rebooking. If that happens often enough to muddy the comparison, we leave the result open.
Before we start, we decide how many completed sessions we need and how much lower the average share of empty places must be for the earlier reminder to be worthwhile. That reduction is our decision threshold. The 30-day attendance share and the average front-desk minutes spent during that same window on cancellation, refund and rebooking contacts tied to each completed session are the guardrail metrics. They stop us trading one problem for another.
A metric’s role belongs to the decision we’re making. Attendance is an outcome metric for the whole first visit, but a guardrail for this narrower reminder experiment. Together, the threshold and guardrails form its success criteria, our account of what needs to happen before we’ll keep the change. We write them now so we can’t choose a flattering answer after seeing the result.
We also check that the gym expects enough evening sessions for the comparison to become readable. Two weeks is our review point, not a promise that enough evidence will have arrived by then. While it runs, we keep checking evening availability and whether the session delivers what the page prepared beginners to expect.
At the review, we return to the choice we wrote before starting, which is our decision rule. We only apply it after enough sessions have finished, everyone’s 30-day attendance window has closed and rebooking hasn’t blurred the groups. Once it’s ready, we keep the earlier reminder only if the average empty-place share clears the decision threshold without lowering attendance or increasing the average front-desk minutes. Otherwise, reminder timing moves back down the list. This can guide future reminder timing, but it still can’t explain the earlier cancellations or tell us whether the page caused them.
Before the conversation ends, we leave ourselves a short note.
Keep the page for two weeks. Leave checkout alone. Assign upcoming evening sessions at random to the usual reminder or the same reminder sent earlier.
More visitors booked and a few more attended. More evening bookings are ending without an attendance, leaving empty places and extra work. Reminder timing is the smallest lead we can test while we keep checking availability and expectations.
Keep the earlier timing only if it lowers the average share of empty places by the amount we agreed without lowering 30-day attendance or increasing average front-desk minutes per completed session. Leave the result open if too many people move between reminder groups. Revisit the page if the session doesn't match what beginners were prepared to expect, or if later attendance and return visits don't improve while the cancelled share stays high.
Review empty places, crossover and front-desk minutes in two weeks. Wait for every booked beginner's 30-day attendance window before keeping the new timing. Reopen sooner if empty places or front-desk work get worse. Add return visits after every attendee has had the full 30 days.
What we carry into the next launch
So, did the product get better? We can answer part of it. On the black-and-white count, 105 beginners attended in the second cohort rather than 90 in the first. That’s 15 more people through the door. The difference may not hold, though, and the combined launch can’t tell us which change caused it. With 25 more people cancelling without attending and the return row still blank, we can’t call the whole first visit better yet.
The reports were measuring different parts of the same journey, so neither could supply the verdict alone. We still had to bring them together for one bounded decision. Next time, we’d rather settle their jobs before launch than reconstruct them on Monday. We’d write down what success looks like, what each metric is for, when it becomes readable and what result will change our decision.
If mixed numbers still land on our screen, five ordinary questions bring us back to the work.
- What were we trying to change, and which result sits closest to it?
- What job does each number have along that journey?
- Are we counting the same kinds of people, the same way, over the same amount of time?
- Where is the movement gathering? Could some of it be ordinary variation, and what do the records and conversations suggest without proving?
- What decision can we support now, and what result or review point would make us change it?
The booking report made the launch look successful, while the cancellation count made it look like a failure. Once we followed the whole visit, we stopped treating them as rival verdicts and made each metric useful. In product terms, we used booking conversion as a leading indicator, attendance as an outcome metric, and cancellations and front-desk work as guardrails. We checked the cohorts, observation windows, instrumentation and denominators. We used segmentation and diagnostic metrics to find the evening pattern, triangulated the records with what the front desk, instructor and customers knew, and stopped an observational comparison from carrying a causal claim it couldn’t support. That narrowed the problem into a testable hypothesis and let us decide upfront how much evidence we’d need, which metric would show progress, which guardrails would catch harm and what result would be strong enough to act on. Only then did we test reminder timing as a smaller, reversible bet, randomised by session and governed by a decision rule we’d agreed on before seeing the results.
These terms might sound complicated when they’re gathered in one paragraph, but each is really just a name for a move we made together when the next question called for it. Each move stopped us from drawing the wrong conclusion, helped us find where the problem was gathering or made the next decision safer. We still haven’t finished improving the whole visit, but we’ve gained far more than a yes or no. We know what moved, what it may be costing, what deserves a test, what evidence isn’t ready and what would change our minds. We’re clearer about what changed, more honest about what we still don’t know and much better equipped to keep making the product better.