Microsoft, Netflix, and Booking.com experiment metrics.
Which metrics are they considering and why.
Hi there,
Whilst I like to focus on D2C startup examples for the most part even I have to admit that there is a lot we can learn from these big players.
This is especially true when it comes to experiment metrics, where it is so easy to choose the wrong ones… like revenue, an all-too-common startup experiment metric.
The result? You end up chasing quick wins and the wrong experiments.
An Insights First Approach
I like Convert’s insight-first approach.
In which your focus is not on quick wins and revenue, and experimentation is not a goal in itself.
Rather, it is a way to better understand your customers and offer them consistent value throughout their lifecycle.
That doesn’t mean revenue isn’t important. You still need to communicate the value to the business of each learning to stakeholders.
Whilst that sounds great in theory, but what about in practice?
That is where our name-dropping comes in to play. Here are 5 key types of metrics you should be focusing on and how to start using these in three simple steps.
1. The North Star Metric
Your North Star is the overarching principle and goal of the entire business.
For example, Netflix’s NSM is the total hours watched. But their tests aren’t designed to just boost engagement.
They’re designed to increase retention, as if they’re retaining customers, it means subscribers engage with and find value in the product
Your North Star Metric should not be revenue.
2. Input Metrics
The early stages of experimentation should focus on input metrics.
Omniscient’s Alex Birkett described the main three input metrics he uses as:
Experiment velocity
Experiment win rate
Average win per experiment
But keep in mind a high win rate isn’t always better.
Sometimes, companies get obsessed with win rate. But too high a win rate could mean you aren’t taking risks / are getting caught on optimisations. It’s a balance.
3. Output Metrics
These are the outcomes you ultimately want to get, such as retention rate, NPS score, revenue etc.
4. Driver Metrics
These are metrics that you focus on in the short term, such as conversion rate.
5. Guardrail metrics
These are business metrics designed to indirectly measure business value and signpost any potentially misleading results and analysis.
But how do you start using these metrics in practice?
This is a map of different metrics and how they fit together:

They should all serve the overarching method on the far left and from there you break them down into the various output metrics, driver and guardrail metrics to understand how they impact each other.
From there, you can use the 3 steps from convert.com to kickstart your experimentation programme using metrics:
1. Choose the goal of your experimentation program
Booking.com, for example, uses ‘experiment quality’ as its overarching goal.
This is because they understand it is pointless to run bad-quality experiments for the sake of it.
2. Establish a log of acceptable Guardrail Metrics
Including test velocity, so long as you’re running good quality tests.
3. Choose your Driver Metrics on a case-by-case basis
As not all metrics are good metrics.
Microsoft has identified 6 key properties of a good A/B test to choose your metrics based on:

If you want to learn more on using the right metrics for experimenting check out the no jargon step by step guide by Convert.
Convert Experiences
If you want that extra guidance with your A/B testing, I recommend checking out Convert Experiences.
While I do partner with Convert transparently, I hugely recommend them for A/B testing too.
Convert Experiences is an experimentation platform that provides comprehensive features and support to run your tests across multiple channels. It’s great value for money and has a 4x faster support desk than average.
It’ll let you set up a personalised website experience in minutes, to test personalised messaging to different audiences and run multiple types of experiment on them including:
A/B testing
Split testing
Multivariate testing
Multipage experiments
Personalisations
Full stack experiments
You can get a 15-day Free Trial below.
Recommendation
Convert has tonnes of brilliant resources that will guide you through your testing struggles.
Here are some of my new favourite articles that they’ve shared
Riding out the Messy Middle of A/B testing - How to ensure you setup your A/B test for sucess from using the ALARM framework to common pitfalls.
Marketer’s Guide: Making Sense of Qualitative and Quantitative Data - How to combine the two to focus on the right experiments
How to Wrap Your Head Around Minimum Detectable Effect (MDE) - How to understand if you have enough traffic to run an A/B test and use MDE effectively.
Chat soon,
Daphne




