Home › Marketing Analytics › Cohort Analysis
MARKETING ANALYTICS · RETENTION

Cohort Analysis for Marketers: Read the Table, Not the Average

"Repeat rate: 18%" tells you nothing. Eighteen percent of whom, acquired when, from where? A cohort table answers all three at once — and it is the only report that shows whether your marketing is getting better or just bigger.

By the Digital Hangover team · Updated September 2026 · 10 min read
Quick answer: Cohort analysis groups customers by when (or how) they first bought, then tracks each group separately over the following months. The output is a triangle-shaped table: one row per cohort, one column per month since acquisition, each cell the share of that cohort still buying. Read down a column to see if retention is improving; read across a row to see how fast a cohort decays.
READ THE TABLE, NOT THE AVERAGE Each row is a month's cohort; each column, how many came back MONTHS SINCE FIRST ORDER → M0 M1 M2 M3 M4 M5 Jan cohort 100% 34% 26% 22% 20% 18% Feb cohort 100% 31% 24% 20% 17% Mar cohort 100% 38% 30% 26% Apr cohort 100% 42% 33% May cohort 100% 40% Jun cohort 100% down a column: is retention improving? (Jan 34% → Apr 42%) across a row: the decay curve Illustrative repeat rates for a D2C brand. The blended 'repeat rate' hides that April's customers are better than January's.

The triangle table, read two ways: down a column for whether retention is improving, across a row for the decay curve.

An average hides its own history. A blended repeat rate of 18% mixes customers acquired last week — who haven't had time to reorder — with customers from two years ago who stayed or left long ago.

Cohorts unmix it. They follow a customer past the first order, which is where most of the profit in ecommerce, subscriptions and apps sits.

This page is the depth behind the "CLV, cohorts and BigQuery" section of our marketing analytics guide: what a cohort is, how to read the table cell by cell, what GA4's cohort exploration gives you, how to build the table from an order export, and the four mistakes that make it lie.

What is a cohort?

A cohort is a group of customers who share a starting point. You pick the starting point; the analysis follows the group from there.

Three cohort definitions cover almost every marketing question:

  • Acquisition-month cohorts. Everyone whose first order was in January is the January cohort. This is the default, and the one the worked example below uses. It answers: "are the customers we acquire getting better or worse over time?"
  • Acquisition-channel cohorts. Everyone whose first order came from Meta ads, versus organic search, versus a marketplace listing. It answers: "which channel brings customers who come back?" — a very different question from "which channel is cheapest per first order".
  • First-product cohorts. Everyone whose first order was a trial pack, versus a full-size product, versus a bundle. It answers: "which entry product creates a repeat customer?"

You can combine them — the January cohort from Meta ads whose first product was the trial pack — but every split shrinks the group, and small cohorts are the first pitfall below.

One boundary up front: cohorts feed lifetime value and are compared against acquisition cost, but this page does neither calculation. Our customer lifetime value guide owns CLV; our CAC guide owns the cost side. Cohort analysis is the table both are built from.

The cohort table, cell by cell

Illustrative dataAcquisition-month cohortsSnapshot at 30 June

Below is a retention cohort table for a D2C brand, six acquisition months in, viewed at the end of June. Every number is invented to teach the shape — not a benchmark, not a client's data.

Each cell is the percentage of that cohort that placed at least one order in that month after acquisition. Month 0 is the acquisition month, always 100%, so it is left out.

Cohort (first order)New customersMonth 1Month 2Month 3Month 4Month 5
January1,20022%15%12%10%9%
February1,05024%16%13%11%—
March1,40021%14%11%——
April1,65027%19%———
May1,30029%————
June1,100—————

Illustrative cohort table for a D2C brand, snapshot at 30 June. A dash means that month hasn't happened yet for that cohort. Numbers are invented for the example.

Start with the triangle. January has lived through five months, so it has five cells. June was acquired in the snapshot month, so it has none yet. The staircase of dashes is not missing data — it is the future.

Now the cells. January, Month 1 = 22% means that of the 1,200 people whose first order was in January, 264 ordered again in February. January, Month 5 = 9% means 108 ordered in June — not necessarily the same people who ordered in February. This is a "standard" (per-period) view, not a cumulative one.

Across a row: the decay curve

Read January left to right: 22, 15, 12, 10, 9. That is a decay curve, and its shape matters more than any single number.

The Month 1 to Month 2 drop is the biggest. After Month 3 the curve flattens — roughly one in ten January customers is still buying every month. That flat tail is your core repeat base. A curve that flattens is a habit product; a curve that keeps falling towards zero is a business that has to re-buy every customer.

Down a column: is retention improving?

Read the Month 1 column top to bottom: 22, 24, 21, 27, 29. Something changed in April.

In our illustrative brand, April is when checkout started offering subscribe-and-save. The April and May cohorts come back in Month 1 at a visibly higher rate, and April's Month 2 (19%) already beats anything the earlier cohorts managed at the same age.

A blended average can never show this. One "repeat rate" for the half-year would land near 20% and hide the fact that the business got materially better in April.

The diagonal: what happened in one calendar month

Cells on the same diagonal share a calendar month: January's Month 5, February's Month 4, March's Month 3, April's Month 2 and May's Month 1 all describe June. If every cohort dips on one diagonal, the cause is that month — a stock-out, a site outage, a competitor's sale — not the cohorts.

How to build a cohort table from an order export in Google Sheets

You do not need a warehouse for the first version. A raw order export from Shopify, WooCommerce, Unicommerce or your own database is enough. The formulas are described in words so you can write them in Sheets or Excel.

  1. Export every order, one row each. Four columns: order ID, a stable customer identifier (email or phone — the same one every time), order date, order value. Add first-touch channel if your platform stores it; that is what makes channel cohorts possible later.
  2. Find each customer's first order date. In a new column, for every row, look up the earliest order date for that row's customer ID — in Sheets, a MINIFS over the date column filtered on the customer column. Every row for the same customer now carries the same first-order date.
  3. Turn it into a cohort label. Format the first-order date as year-and-month ("2026-01"). That is the row of the table. Do the same for the order date itself — the order's calendar month.
  4. Calculate the month index. (Order year × 12 + order month) minus (first-order year × 12 + first-order month). A first order gets 0, an order the next month gets 1. That is the column of the table.
  5. Build a pivot table. Rows: cohort label. Columns: month index. Value: count of unique customer IDs. The Month 0 column is your cohort size.
  6. Convert counts to percentages. Divide each cell by its row's Month 0 count. Apply a colour scale so the eye reads the table before the brain does.
  7. Blank the future. Any cell whose cohort month plus month index is later than the snapshot date has not happened yet. Leave it empty, not 0% — a zero there drags every average built on top of it.

Two refinements once it works. Swap "count of customers" for "sum of order value" and you have a revenue cohort table — the direct input to lifetime value. Add channel as a pivot filter and you have acquisition-channel cohorts without rebuilding anything.

Past a few lakh rows, Sheets slows down and the same logic belongs in a database; our BigQuery for marketers guide covers the query version.

What GA4's cohort exploration gives you (and what it doesn't)

FreeExplore › Cohort explorationVerified Sep 2026

GA4 has a built-in cohort table under Explore. Know what it does before deciding whether the Sheets version is still necessary. (Property setup is in our GA4 guide.)

Per Google's cohort exploration documentation, you define the cohort with an inclusion criterion — "First touch (acquisition date)", "Any event", "Any transaction", "Any conversion", or a specific event — and a return criterion a user must meet to count in later periods (any event, a transaction, a conversion, or a specific event). Granularity is daily, weekly (Sunday to Saturday) or monthly.

The same page lists three calculation modes: "Standard" (met the return criterion in that specific period — the view our example uses), "Rolling" (met it in that period and every period before) and "Cumulative" (met it in any period so far). It shows up to 60 cohorts, and a breakdown dimension splits each cohort by the top 15 values of that dimension — which is how you get channel cohorts inside GA4.

Now the limits, which matter more for a retention question:

  • It counts users, not customers. A GA4 "user" is a device or browser unless User-ID is set up with signed-in traffic. Someone who bought on a phone in January and a laptop in March is two users and zero repeat purchases.
  • It is bound by data retention. Google's data retention page states that standard properties default to 2 months of event-level data with a maximum of 14 months, and that the setting "only affects explorations and funnel reports". A cohort exploration cannot look further back than that window. Set 14 months on day one; it is not retroactive.
  • It has no order value per customer. The order export has every rupee. GA4's cohort table tells you who came back, not what they were worth.

Honest summary: GA4's cohort exploration is a good free check on engagement cohorts (did users from a campaign come back to the site?) and a weak tool for purchase cohorts. For "did the customers we bought in April buy again?", the order export wins.

What marketers actually do with a cohort table

The table is not the deliverable. The decisions are.

Judge channel quality, not channel cost

A channel that brings first orders at ₹600 each and a channel that brings them at ₹900 each are not ranked until you see their cohorts. If the ₹900 customers reorder at twice the rate, the "expensive" channel is the cheap one. That comparison — cohort repeat rate against acquisition cost — is the single most useful budget conversation a D2C brand can have, and it is the one we build into every performance marketing engagement once a client has enough orders to fill a table.

Judge campaign quality after the campaign ends

A flash sale that produced 2,000 first orders looks like a triumph in the campaign report. Its cohort, three months on, often tells a different story: discount-driven buyers reorder at a fraction of the rate of full-price buyers. Tag the campaign as a cohort and you can see this by Month 2. Without cohorts you find out only when the blended repeat rate sags a quarter later and nobody knows why.

The Diwali-cohort problem

Indian ecommerce has a version of this so common it deserves its own name. October and November cohorts are the largest of the year — festive gifting, marketplace sale events, the biggest ad budgets. They are also, for most categories, the worst-retaining cohorts of the year: gift buyers bought for someone else, deal hunters bought for the discount, and neither has any reason to return in December.

Read as a blended number, the December repeat rate looks like a collapse. Read as cohorts, it is one enormous, low-retention group sitting on top of a perfectly healthy base. The fix is not to skip Diwali — it is to plan for the festive cohort's retention separately: a post-festive reactivation flow, a first-full-price offer in January, and a budget that values a festive customer at what the festive cohort actually returns, not at the annual average. Our festive ecommerce marketing guide covers the acquisition side; the cohort table is how you keep the after-season honest.

Find where the journey breaks

The Month 1 to Month 2 drop is where most of a cohort disappears. Whatever happens to a customer in the four to eight weeks after the first order — delivery experience, the product running out, the first email — is what the cohort curve is measuring. Map that stretch of the customer journey and you are usually looking at the cheapest retention gain available.

Four pitfalls that make a cohort table lie

  • Small cohorts. A cohort of 40 customers where 9 reorder is 22.5%; if one more had reordered it would be 25%. Below a few hundred customers per cohort, month-to-month differences are mostly noise. Merge months into quarters, or merge channels, until the cells are big enough that a single customer cannot move them by a full percentage point.
  • Mixing new and returning customers. If the export counts every customer who ordered in January as the "January cohort", it includes people acquired two years ago who happened to buy in January. The cohort label must come from the first order ever, not the first order in your date range. Step 2 above is what prevents this — make sure the export goes back to the beginning.
  • Survivorship. Older cohorts have had more time to reorder, so a cumulative view will always make January look better than May. Compare cohorts at the same age (Month 1 to Month 1), never January's total to May's total.
  • Reading an incomplete month. The most recent diagonal is the current calendar month. If the snapshot is taken on the 10th, every cell on that diagonal is one-third of a month's data. Either exclude the current month or label it clearly; the standard mistake is to see a dip on the last diagonal and panic.

Where to go from here

Build the table before you build anything on top of it. One afternoon with the order export gives you a retention view that GA4 alone cannot, and every later question — what a customer is worth, what you can pay for one, which channel deserves the next budget increase — starts from this grid.

Then read it in three directions: across a row for the decay curve, down a column for improvement, along a diagonal for a calendar-month event. When the reading is stable, hand the revenue version to the lifetime value calculation and the acquisition-channel version to the budget meeting.

Key takeaways: A cohort groups customers by a shared starting point — acquisition month, channel or first product — and the table tracks each group separately over time. Across a row is the decay curve; down a column is whether retention is improving; a diagonal is one calendar month. GA4's cohort exploration is free but counts users rather than customers and is capped by data retention, so the order export in Sheets is the honest starting version. Keep cohorts large, define them by the first order ever, and compare at the same age.

Frequently asked questions

What is cohort analysis in marketing?

Cohort analysis groups customers by a shared starting point — most often the month of their first purchase, but also the channel they arrived from or the first product they bought — and then tracks each group separately over the months that follow. Instead of one blended repeat rate, you get a table showing how each group behaves at each age, which reveals whether your marketing is bringing in better customers over time.

How do you read a cohort table?

Rows are cohorts, columns are months since acquisition, and each cell is the share of that cohort still buying in that month. Read across a row to see the decay curve for one group. Read down a column to compare every cohort at the same age — that is how you tell whether retention is improving. Cells on the same diagonal share a calendar month, so a dip that runs along a diagonal points to a one-off event in that month rather than a change in customer quality.

Can I do cohort analysis in Google Analytics 4?

Yes, with limits. GA4's Explore section has a cohort exploration that defines cohorts by first touch, any event, transaction or conversion, tracks daily, weekly or monthly, and shows up to 60 cohorts with standard, rolling or cumulative calculation. It counts users (devices or browsers unless User-ID is set up) rather than customers, it cannot look further back than your event data retention setting (2 months by default, 14 months maximum on standard properties), and it does not show order value per customer. For purchase retention, build the table from your order export as well.

What is a good retention rate for a cohort?

There is no honest universal benchmark — it depends on the category, the purchase cycle and how the cohort was acquired. A consumable with a monthly refill cycle should show a much flatter curve than a durable product bought once a year. The useful comparison is internal: this month's cohort against last month's at the same age, or the Meta-ads cohort against the organic cohort. If the Month 1 column is rising across cohorts, retention is improving, whatever the absolute number.

How is cohort analysis different from customer lifetime value?

Cohort analysis is the table; lifetime value is a number calculated from it. The cohort table shows how many customers from each group come back, and how much they spend, at each age. Lifetime value takes that curve and turns it into the total revenue or margin a typical customer produces over a defined period. You need the cohort table first, because a lifetime value built from a blended average hides the difference between good and bad acquisition months.

Performance marketing, measured by cohort

Buy customers who come back, not just first orders

We build cohort reporting into paid campaigns so budget follows the channels whose customers reorder — not the ones with the cheapest first sale.

Explore performance marketing →

Get our posts in Google

Make Digital Hangover a preferred source

One tap tells Google to show more of our SEO and marketing coverage in your Top Stories.