Measuring the experience where it actually happens.
This is how the EXP Index works: the index I designed to measure the experience process by process, without losing the detail.
Let’s be honest: most organizations don’t measure customer experience. They measure NPS and assume it’s the same thing.
That’s how it started for me. For a long time, experience reached stakeholders as two numbers that always traveled together: NPS, “How likely are you to recommend us?”, and next to it, overall satisfaction, “How satisfied are you with us?”. People used them to celebrate or to look for someone to blame.
The problem is that neither one actually talks about the experience. One measures an intention; the other, a general impression of the brand:
- They measure stated opinion, not what the customer actually experienced. Someone can recommend you and say they are satisfied because of your interest rate or your brand, while hating the process itself.
- They don’t tell you what to fix. If they drop 8 points, you don’t know whether it was the app, the branch, a fee, or a competitor’s campaign.
- NPS hides the distribution. By subtracting detractors from promoters, two very different realities can end up with the same score.
In the end, we had two numbers for the entire organization and none for each individual process. If, in one quarter, you launch a new onboarding flow, redesign document upload, and change how the status of an application is communicated, and satisfaction goes up 4 points, which of the three changes made it happen? Nobody knows. There’s no granularity.
So we went down to the level where experience actually happens: the flow. We started measuring effort, time, abandonment, and satisfaction at each process. We gained detail, but created a new problem: four numbers that nobody could relate to each other. In every stakeholder presentation, the same question came up: “So, is it good or bad?”
Everything in a single number
The answer I found was to bring them together into a single index. I didn’t invent any new metric: they have all existed for years in UX and product analytics. What I built is a synthesis layer on top of them, which I called the EXP Index: a number from 0 to 100 with a label next to it, something like “74, Good.” It works for me because any stakeholder can understand it without me having to explain what CES is, while anyone who wants the detail can still see the underlying metrics.
I measure it when each project is completed, to establish a baseline for the flow we delivered, and then continuously, to see whether it holds up or starts to deteriorate.
That said, it is not intended to measure business performance. That’s what conversion and retention are for; the EXP Index sits alongside those metrics, not in their place.
What it consists of, and why everything doesn’t weigh the same
| Component | What it measures | Weight | Original scale | How it’s collected |
|---|---|---|---|---|
| CES (Customer Effort Score) | How easy it was to complete the task | 35% | 1 to 7 (lower is better) | One question at the end of the flow |
| TCT (Task Completion Time) | Active time the customer spends completing the task, compared with a reference time | 10% | Minutes vs. benchmark | Analytics or stopwatch during testing |
| E2E (End-to-End Time) | Total time from start to outcome, including waiting time | 15% | Hours or days vs. benchmark | Case start and completion dates |
| Drop-off | Percentage of users who abandon before completing | 25% | 0 to 100% (lower is better) | Conversion funnel by step |
| CSAT (Customer Satisfaction) | Satisfaction with the experience of the flow | 15% | 1 to 5 (higher is better) | One question at the end of the flow |
[!NOTE] The weights reflect what we as a team want to drive in our customers’ experience and should be adjusted when implementing the metric.
Effort weighs more than anything else because, in financial processes, friction is the main enemy. A customer who had to struggle with a form won’t come back, even if they eventually completed it.
Time and abandonment make up half of the index because they are behavioral signals: they don’t depend on what the user says, but on what actually happened. Within time, end-to-end time weighs slightly more than active time because, in processes involving approvals, waiting is often what customers feel the most.
Satisfaction weighs less because it is largely a consequence of everything else. If the process was easy, fast, and nobody abandoned it, satisfaction usually follows. Lowering its weight is my way of avoiding counting the same thing twice, although it doesn’t solve the problem entirely (I’ll come back to this later).
How it’s calculated
The calculation has two steps: bringing all five components to the same 0–100 scale, and then applying the weights.
Step 1: Normalize
Two of the metrics are inverted because, for them, a lower number is better.
| Component | Normalization to 0–100 | Note |
|---|---|---|
| CES | ((7 − score) / 6) × 100 | Inverted: less effort, higher score |
| TCT | max(0, 1 − (T actual − T ideal) / (T max − T ideal)) × 100 | Requires an active-time benchmark for each flow |
| E2E | max(0, 1 − (T actual − T ideal) / (T max − T ideal)) × 100 | Same formula, with its own benchmark for the complete process |
| Drop-off | (1 − abandonment rate) × 100 | Inverted: less abandonment, higher score |
| CSAT | ((score − 1) / 4) × 100 | Direct |
Step 2: Apply the weights
EXP Index=0.35 CES+0.10 TCT+0.15 E2E+0.25 Drop-off+0.15 CSAT\text{EXP Index} = 0.35\,\text{CES} + 0.10\,\text{TCT} + 0.15\,\text{E2E} + 0.25\,\text{Drop-off} + 0.15\,\text{CSAT}
An example using fictional data. Imagine an online loan application flow:
| Component | Data | Normalized | Weight | Contribution |
|---|---|---|---|---|
| CES | 3.1 / 7 | 65 | 35% | 22.75 |
| TCT | 9 min (ideal 6, max 18) | 75 | 10% | 7.50 |
| E2E | 6 days (ideal 3, max 9) | 50 | 15% | 7.50 |
| Drop-off | 22% | 78 | 25% | 19.50 |
| CSAT | 3.8 / 5 | 70 | 15% | 10.50 |
| EXP Index | 67.75 → 68, Fair |
Instead of five numbers to debate, stakeholders get one: this flow needs attention. And the table tells us where to start looking: effort is what takes away the most points, followed by end-to-end waiting time. What it doesn’t tell us is why; that requires research.
Waiting time counts too
This is the part that generates the most discussion. At first, I put everything into TCT: the time the user spends in front of the screen and the time they spend waiting while someone approves, reviews, or validates something on the other side. But mixing them creates confusion, because in UX, task time is active time, not waiting time. So today I separate them into two: TCT, which is the time the customer spends actively completing the task, and E2E, which is everything that happens until they get their outcome.
Someone always says that this is operations, not design. Maybe it is, but the customer doesn’t care. If they filled out the form in 8 minutes and received a response five days later, the process took five days for them. Measuring only the 8 minutes means focusing on the part where we look good in the picture.
It also has a side effect that I like: it makes it clear that experience doesn’t depend only on the person designing the screens, but on the entire organization.
Separating them also makes the result more actionable. If TCT is good but E2E is bad, the problem isn’t in the screen: it’s in the operation, a manual approval, or an overloaded back office. And that changes who needs to be sitting at the table.
How we measure time
Both times are measured the same way: against an ideal and a maximum defined for each flow. The difference is what gets counted. TCT is measured in minutes and only runs while the customer is actively doing something. E2E is measured in hours or days and runs from the moment the application is submitted until the customer receives an answer, even if they do nothing in between.
If you don’t have historical data, this is how I define them:
- Map the complete flow, step by step.
- Estimate a reasonable amount of time for each step.
- Define T ideal as the frictionless total: how long a competent user takes, not the absolute minimum.
- Define T max as 2 to 3 times T ideal: the point where the experience is clearly becoming critical.
- Validate it through the first user sessions and adjust.
For a loan with multiple stages, the two times would look something like this (figures are illustrative):
| Time | Stage | T ideal | T max |
|---|---|---|---|
| TCT | Initial online application | 8 min | 20 min |
| TCT | Document upload | 5 min | 15 min |
| TCT | Signing and formalization | 30 min | 90 min |
| E2E | Complete process, from application to response | 3 days | 9 days |
What each result means
The number is only useful if it tells you what to do. That’s why each range comes with an associated action:
| Range | Label | What it means | What to do |
|---|---|---|---|
| 85–100 | Excellent | Flow is ready to scale | Document it as a reference |
| 70–84 | Good | Works, with clear opportunities | Iterate in the next cycle |
| 50–69 | Fair | Requires active intervention | Prioritize it in the current backlog |
| 0–49 | Critical | Broken experience | Prioritize an immediate redesign |
And it is measured at three points: when the project is completed (baseline), continuously (to detect deterioration), and with each iteration (to verify that the change actually improved something).
What the index does not do
First, a 68 is not a truth. It is what comes out of this formula, with these weights and these benchmarks. It helps me see whether a flow improves or deteriorates over time, but not as an absolute score, and not to compare an account opening with a mortgage: each flow has its own reference times.
Second, the weights don’t come from a statistical model. They come from what we as an organization want to drive in the experience: today, above all, reducing customer effort. They are a bet, and my plan is to review them after 3 to 6 months of real measurements, to see whether that bet is reflected in the results.
Third, the components overlap with each other. Friction increases time, time makes people abandon the flow, and all of that lowers satisfaction. Reducing the weight of CSAT helps, but I’m still partially counting the same thing more than once.
And most importantly: the index doesn’t tell you what to fix. What it gives you is the health of a particular flow: whether something is wrong and which component is driving it. The why comes from research and user interviews, and no number will replace that. That’s why I think of it as another layer, each with its own question:
| Layer | What question does it answer? |
|---|---|
| NPS | How is our relationship with the customer? |
| EXP Index | How is this flow performing? |
| Analytics and funnel | Where is it failing? |
| Research | Why is it failing? |
| Design and operations | What do we change? |
Try it with your own data
Demo
I built an open calculator so any team can test the index with their own numbers:
If you use it, I’d love to hear which weights work for you and what types of processes you use it on. I’d also love to know where you think the index falls short.