Back to blogAI Creatives

Ad creative testing for your gym: the framework that fits within €600/month

Ad creative testing for your gym: the framework that fits within €600/month

You're not Coca-Cola. On €600/month you can't test 50 creatives, and the testing frameworks floating around online are written for accounts spending €50,000. What you can do is run one clean test every two weeks: 2-3 genuinely different concepts, around €50 per variant, and qualified lead CPL as the only judge. This article is that framework, nothing more.

Most gyms do one of two things: either they test nothing (upload an ad, let it run for six months, and wonder why it stopped working), or they test badly (launch 8 variations of the same ad with a different button color and declare the best CTR at day two the winner). Both burn money. The first through undetected creative fatigue; the second because it splits a small budget across variants that differ in nothing that matters.

The math in charge: how many tests you can afford

Before the framework, the numbers. To judge a variant you need it to generate enough data. With a CPL of €8-15 (the typical range according to CPL benchmarks for gyms), €50 per variant gives you 4-6 leads. That's the minimum to get a signal; less than that is noise.

Now divide your budget:

Monthly budget Testing spend (20-30%) Testable variants/month Real tests/month
€300 €60-90 1-2 1 every 3-4 weeks
€600 €120-180 2-3 1 every 2 weeks
€1,000 €200-300 4-6 2 per month
€1,500 €300-450 6-9 2-3 per month

The remaining 70-80% of your budget goes to what you already know works. Testing is funded from the margin, never from the core. If you're spending €600 and put €400 into experiments, you're not testing: you're gambling.

The uncomfortable conclusion from the table: on a gym budget you have ammunition for 2-6 variants per month. That's why each variant has to test something big. Which is exactly the next point.

What to test first: concept moves 10×, button moves nothing

Not all variables carry the same weight. Ranked by real impact on CPL:

  1. The concept or angle. Are you selling transformation, community, convenience, or safety for beginners? Changing angle can move CPL by 50-200%. This is the queen variable.
  2. The format. Video vs static image vs carousel. Moves 20-60%.
  3. The hook. The first 1-2 seconds of video or the first visual element of the image. Moves 15-40%, measured with the hook rate.
  4. The main copy. Moves 10-20%.
  5. Button color, font, emoji in the headline. Moves between nothing and 3%. At your data volume, you'll never reach statistical significance here. Forget it forever.

The usual trap is to start at the bottom because it's easy to produce: making 6 versions of the same ad with different backgrounds takes ten minutes in Canva. But you're burning your 4-6 leads per variant to measure a variable that moves 2%. It's like weighing flour on a truck scale.

Always start at the top. Until you have two or three angles compared with data, there's no point going down to formats. And until the format is clear, don't touch hooks.

The structure: 2-3 concepts × 2 formats, then stop

A well-built test for a gym looks like this:

  • 2-3 genuinely different concepts. Not "before/after with blue background" vs "before/after with gray background." Actually different: one on physical transformation, one on community ("nobody judges you here"), one on a concrete offer. If you're unsure which concepts to pick, the gym creatives guide walks through the angles that work in local fitness.
  • 2 formats per concept at most. Typically short video + static image. That's already 4-6 ads, the ceiling of your budget.
  • Everything else identical. Same audience, same offer, same landing page, same budget per variant. If you change two things at once, you won't know which one moved the result, and you'll have paid to learn nothing.

Where to set it up? One ad set with all 4-6 ads inside and letting Meta distribute the spend works reasonably well, with one caveat: Meta concentrates spend quickly on its favorite and kills the others before giving them a chance. For tests where you want a clean read, the A/B test tool in Ads Manager forces an equal budget split. It's less efficient short-term (you pay to show the loser too), but you're buying information. For budgets of €300, accept Meta's automatic distribution; from €600 up, the formal A/B test is worth it once a month.

How many ads to keep active in total, outside the test, is another question with a concrete answer: it's covered in how many creatives your campaign needs.

How much time and money per variant before judging

Non-negotiable minimums per variant:

  • €50 in spend. Below that, the 2-3 leads you have could be luck in either direction.
  • 3-4 full days. Ads perform differently on Monday versus Saturday, and the algorithm needs to exit its initial exploration phase. Judging at 24 hours is the most common way to kill the winner.
  • Ideally 5+ leads. If at €50 and 4 days a variant has 1 lead at €50 CPL, you can kill it early: the disaster signal arrives faster than the success signal.

Impatience is the number one enemy of testing on a small budget. Checking Ads Manager every three hours and "adjusting" turns a test into a lottery. Launch on Monday, don't touch anything until Friday.

The deciding metric: qualified CPL, not CTR

CTR is the favorite metric of badly-run tests because it comes quickly and feels good. But an ad can have a great CTR and bring in browsers who will never set foot in your gym. The "can you guess how many calories this dish has?" ad crushes on clicks and dies on sign-ups.

The metric hierarchy for declaring a winner:

Metric What it's for Does it decide the winner?
CTR Early diagnosis: does the creative grab attention? No
Gross CPL Quick comparison between variants Provisionally only
Qualified lead CPL Cost per lead who responds and fits your target customer Yes
Cost per sign-up The absolute truth Yes, but takes 4-6 weeks

In practice: decide with qualified CPL (leads who reply on WhatsApp and show real intent), then review at 30 days to see if cost per sign-up confirms the call. If your volume is too low to separate qualified leads by variant, use gross CPL but review the leads from each ad manually. Ten minutes reading conversations tells you more than any dashboard.

When to kill a creative

Three cut rules, in order:

  1. At €50 with zero leads: dead. No appeals.
  2. CPL 50%+ worse than your best variant after 4 days and 5+ leads each: dead.
  3. Similar CPL but visibly worse leads (no replies, only ask about price, outside your area): dead, even if the number looks fine.

And the inverse rule, which almost nobody applies: don't touch the winner. The temptation to "improve" it with a new variation each week sends it back into the learning phase and raises its CPL. The winner runs until fatigue degrades it, then it's replaced, not before.

Documenting: the 5-column table that's worth more than Ads Manager

Meta stores the numbers, but not the learnings. Six months from now you won't remember why you killed that spinning class video. A spreadsheet with five columns fixes that:

Date Concept tested Against what Result (CPL) Learning in one sentence
03/06 Community (real member video) Transformation (before/after) €7.20 vs €11.80 Here, belonging sells, not physique
17/06 Community in static image Community in video €10.10 vs €7.40 The concept needs video

The last column is the important one. A year of well-documented tests is a manual of what sells in your specific neighborhood, and no agency will give you that knowledge.

The realistic monthly cycle

For a gym spending €600-1,000/month, the sustainable pace is one test every two weeks:

  • Weeks 1-2: active test (2-3 new variants against the control). Produce the creatives for the next round.
  • Weeks 3-4: consolidate, move budget to the winner, launch the next test.

That's 24-26 tests a year. It sounds like little next to the "100 creatives a month" the gurus preach, but 24 clean concept comparisons leave you knowing exactly which angle, format, and hook work for your gym. The real alternative isn't testing more: it's testing badly or not testing at all.

The bottleneck, honestly, is usually not the budget but the production: someone has to make those 4-6 new creatives every two weeks. That's one of the reasons we built Pilotium, to generate and rotate creative variants automatically, matching the mix to what each club's budget can actually test.

This fortnight: pick two opposite concepts (not two versions of the same thing), put €50 on each, don't touch them for 4 days, and write down the learning in one sentence. That alone puts you ahead of 90% of gyms in your city.

Stay Ahead of the Game

Weekly AI marketing insights. No spam. Unsubscribe anytime.