Run as a programme, Hourly API Scaling was worth 30.8% and 19.0% of total account spend on the two brands measured — because each action starts from a floor the last one raised, and that floor stays raised.
Scored the strictest way instead — one action at a time, against the day before it — the number comes in lower, because that test throws away the raised floor every action leaves behind. It is a deliberate floor, not a verdict. Both readings are worked through below, including where the strict test runs out of comparison days and stops resolving anything at all.
What this study is trying to learn
The question is not whether spend went up. It is how much of the rise would have happened anyway.
Two brands running paid social through Meta were put on the same operating process: budgets are moved programmatically and many times a day, driven by what the account's own delivery is doing that hour, rather than adjusted by hand once a day. We call this Hourly API Scaling, and each individual move a scaling action.
How those moves are decided, sized and executed is our own method and is not described here. What follows is only the measurement — what the actions produced, and how much of it would have happened without them.
The counterfactual we are trying to price is simple to state and awkward to measure: if nobody had made the change, what would the account have spent? Everything below is an attempt to answer that without borrowing the answer from the result.
A second question sits underneath the first. Both brands received the same budget process, but only one received the Industry CAC work — the analysis that establishes what an acquisition should cost in a category and therefore how far an account can profitably be pushed. If the budget mechanism is worth the same on both and the outcomes differ by an order of magnitude, the difference is attributable to the baseline, not the mechanism.
About the data
Hourly delivery, campaign by campaign, joined to the account's own change log.
Delivery comes from the Meta Ads API at campaign × hour granularity. Order and revenue truth comes from the Shopify Admin API and from Statlas MCP, which holds the modelled and reconciled daily figures used for efficiency benchmarking.
The record of what was changed, and when, comes from the account's own activity log, read at day-level granularity across every day of both measurement windows.
177
scaling actions across 20 measured days, 2026-07-11 to 2026-08-08
108
scaling actions across 57 measured days, 2026-04-01 to 2026-08-03
€902,161
total delivery inside the two measurement windows
100%
complete activity record across every day in both windows; one day excluded on a platform server error
What is compared against what
Two baselines, deliberately. Reporting only one of them would be the whole trick.
Every measured day is scored from the hour of its first scaling action through to midnight. The actual spend in that window is compared against two different counterfactuals:
- Isolated action. The baseline is the same hours on the most recent day with no scaling action on it at all. This asks: how much more did this change produce than the state immediately before it? It is a strict floor, and it deliberately gives no credit for the fact that the day before was itself elevated by an earlier change.
- Compounded. The baseline is the last clean day before the programme began, held fixed. This asks: how far has the account travelled from where it started? It captures the ratchet — a floor that is raised tends to stay raised — which is the entire point of running budgets this way.
Neither is the true number on its own. The isolated reading understates because it treats an elevated starting point as though it were free. The compounded reading overstates because it credits the budget mechanism with everything that happened after the programme started, including creative and offer work. The honest answer is bracketed by the two, and we report both throughout.
Setup, controls and the variables we could not hold still
What was held constant, what moved, and what that costs the conclusion.
There is no clean holdout here and we are not going to pretend otherwise. A held-out set of days was considered and rejected: on Brand B the process ran from the first day of the engagement, so no untouched days exist to hold out. The comparison is therefore within-account and within-campaign — the same campaign, the same hours of the day, against its own recent history.
Held constant. Campaign identity, hour-of-day window, account, currency, and attribution setting. A campaign is only ever compared against itself.
Not held constant, and material:
- Creative volume. Both accounts launched new creative throughout. On Brand B the relationship runs the wrong way for a creative explanation — 84 new assets in April moved spend barely at all, 43 in May moved it more than fourfold — which is addressed directly in section 06.
- Offer and landing page. Brand B changed its offer page mid-period. Brand A changed feeds on 13 July.
- Auction conditions. Uncontrollable and unobservable at this granularity.
- Efficiency. Brand B's cost per purchase fell 22% over the window. That is the Industry CAC layer working, and it is a confound for the budget layer — which is precisely why the two brands are reported separately.
The consequence is stated plainly: the isolated figure is defensible as a causal floor, because a same-campaign, same-hour, recent-day comparison absorbs most of these. The compounded figure is a programme-level result and should be read as the value of the whole operating process, not of any single scaling action in isolation.
Instance one — Brand A, the budget mechanism on its own
A premium men's skincare account. No Industry CAC layer. The cleanest read on the mechanism itself.
30.8%
of account spend, €34,955 incremental
15.4%
of account spend, €17,458 incremental — the floor
20
each campaign compared against itself, same hours, most recent comparable day
€2,992
per day in August, from €1,499 in June
Each pair is the account measured against itself. Grey is what it spent from a given hour to midnight on its most recent day with no scaling action; cyan is what it spent in those same hours on a day one ran. Across the 20 pairs that is €37,610 against €55,068 — €17,458 more, a 46.4% difference.
The comparison is deliberately narrow. The account is only ever measured against itself, over the same hours of the day, on the most recent day that carried no scaling action at all. Nothing is compared across accounts or against a forecast, so the differences below cannot be produced by mix or by a model assumption.
Across the 20 pairs the account spent €37,610 in those hours on its comparison days and €55,068 on the days a scaling action ran — €17,458 more, a 46.4% difference. 5 of the 20 pairs are negative, which is the point of showing all of them rather than a summary statistic.
Instance two — Brand B, the mechanism on a raised baseline
A men's health supplement account running both layers. Spend rose 1,262%, and the two layers come apart cleanly: the efficiency work raised the ceiling, the scaling converted it.
19.0%
of account spend, €150,168 incremental
€11,659
per day, from €856 in April
0.91
return on ad spend, from 0.49 — the baseline nearly doubled
The account had two regimes, and the boundary is a single day
Scoring each day's spend against the budget actually available to it separates the period cleanly. For nine consecutive weeks the account spent between 89% and 115% of everything it was allowed to spend. Budget was the binding constraint, and every euro released was taken.
For nine weeks the two lines are the same line: the account spent essentially every euro it was allowed to. From 29 June the budget line jumps and the spend line does not follow — the constraint had moved.
Why the creative counts do not explain any of this
The obvious rival explanation is that creative volume drove the growth. The data runs against it, and not marginally:
- April — 84 new assets, spend broadly flat at €856/day.
- May — 43 new assets, half as many, and spend rose more than fourfold to €4,402/day.
- July — 110 new assets, the largest batch of the period, for an increase of roughly €3,400/day.
The rank order is inverted, so more creative did not buy more spend. What does move with spend is order value. Across the period it rises from €58.81 to €89.88 — a 53% improvement — and its week-by-week correlation with daily spend is r = +0.67, against +0.46 for creative volume on far too few points to mean anything. Cost per acquisition falls 6% over the same stretch while spend rises more than thirteenfold.
That is the relationship the Industry CAC work exists to find. A higher order value raises what an acquisition is worth, which raises what the account can afford to bid, which is what allows it to clear at far greater volume without the economics degrading. The account did not scale because it was handed more assets to run. It scaled because each order became worth more, and the ceiling on what could profitably be paid moved up with it — the same assets simply became affordable to run harder.
Both series trend upward across the period, so the correlation is supporting evidence rather than proof of direction. It is reported because it is the one relationship in this account that holds across the whole window, where creative volume does not.
Spend and efficiency moved together, not against each other. Return on ad spend rose from 0.49 to 0.91 — a 84% improvement — while cost per purchase fell from €128.41 to €99.65. That is the headroom the budget actions were spending into.
Stability, credibility and what we think is still wrong
The findings that survived, the ones that did not, and the errors we corrected along the way.
What holds up
- Two independent readings agree on Brand A. The strict pairing gives 15.4% of account spend and the compounded reading 30.8%. A third check that shares no assumptions with either — scoring each campaign-day against its own delivery pattern rather than a comparison day — lands inside the same band. Different failure modes, same conclusion.
- The per-action effect replicates across accounts. The size of the response to a single scaling action is close to identical on two unrelated advertisers in different categories. Account-level impact is therefore a question of dosage, not of effect size.
- The regime split is not a story imposed on the data. It falls out of a single ratio computed per day, and the boundary lands on one identifiable action.
What does not hold up
- The compounded figure cannot separate layers. On Brand B it credits the budget mechanism with growth that the offer, the creative and the Industry CAC work also produced. It is a programme result. Anyone quoting 19.0% as the value of one scaling action alone is misreading it.
Two ways this measurement goes wrong
- Account-level scoring. Measured across a whole account, the effect is diluted by campaigns that were never touched. Every figure here is scored per campaign.
- Multi-day windows. A window that spans midnight collapses to its start date unless it is expanded deliberately, which files the second day as an untouched comparison day. Left alone, the highest-spending day in the period ends up serving as a control for itself.
Confidence intervals
Resampled by day, because two hours of the same day share a promo, a creative rotation and the weather.
Measured as isolated actions, Brand B's interval crosses zero — a single budget action, judged against the day before it, is indistinguishable from noise. Measured as a compounding programme, both instances are comfortably clear of zero and of each other's floor.
| Scenario | Incremental | Lift on the window | % of account spend | 95% interval | Verdict |
|---|---|---|---|---|---|
| Brand A — isolated action | €17,458 | +46% | 15.39% | [7.8%, 23.4%] | clear of zero |
| Brand A — compounded | €34,955 | +174% | 30.81% | [23.2%, 38.5%] | clear of zero |
| Brand B — isolated action | €2,960 | +2% | 0.38% | [-1.0%, 1.9%] | crosses zero |
| Brand B — compounded | €150,168 | +1,120% | 19.04% | [13.8%, 25.4%] | clear of zero |
Read the two compounded intervals together and the pattern is the useful one: Brand A at 30.8% [23.2, 38.5] and Brand B at 19.0% [13.8, 25.4]. The intervals overlap. On two accounts of very different size, in different categories, at different stages, the budget programme lands in the same band — call it 19–31% of account spend.
The layers multiply rather than add. Brand B's efficiency work moved the base 1.84× on its own, and the budget layer then took 19.0% on top of that raised base — which is why the same mechanism, worth 30.8% on a flat baseline, sits underneath a 1,262% account move here.
Back-tested potential across the portfolio
What the same mechanism would have released on 37 brands that never received it.
The two instances above answer what happened where the process ran. This section asks the inverse: across a portfolio of brands that were never put on it, how much constrained spend was sitting unreleased?
The model has three parts, and each is separately falsifiable. A trigger identifies the campaign-hours where an account is demonstrably running into its own limit while performance is holding. The conditions are tuned per brand and are not published here. A value is attached to each trigger, calibrated so that a brand where the process genuinely ran reproduces the result it actually achieved — not assumed, and not chosen to flatter the outcome. Compounding is not a parameter at all: raises persist, so the raised floor becomes the base the next trigger fires from.
29.2%
spend-weighted modelled uplift on $32,292,369 of spend
28.0%
median · mean 27.4% · modal band 25–30% (10 brands)
$9,436,899
modelled incremental spend
Each bar is one brand; cyan sits at or above the median, grey below. The two reference lines are the spend-weighted portfolio figure and the plain average across brands. They sit within two points of each other, which matters: it means the result is broad-based rather than carried by a handful of large accounts.
Where brands actually land
Mean 27.4%, median 28.0%, modal band 25–30%. The three agree closely, and the spread is tighter than the headline range suggests: half of all brands sit between 23% and 34%.
Run over a single month the same analysis looked bimodal — brands piling up at both ends with almost nothing in between. That turned out to be an artefact of the short window. Over thirty days a brand either happens to catch a constrained stretch or it does not, and the result swings on that accident. Over ninety the runs average out and a single-peaked distribution shows through. It is worth stating because the bimodal version was the more striking chart, and it was the wrong one.
How much to trust this
Read the portfolio total with more confidence than any individual row. The aggregate is built from 37 brands and is stable; a single brand's figure carries roughly ±50%, because it rests on that account's own trigger pattern over one window.
The counterfactual is also unobservable by construction. On a brand that never ran the process there is no version of history where the budget was released, so a trigger is evidence of unmet demand rather than proof of what would have been delivered.
| Brand | Spend | Triggers | Modelled incremental | Uplift |
|---|---|---|---|---|
| Brand 01 | $4,015,149 | 163 | $1,167,243 | 29.07% |
| Brand 02 | $2,808,290 | 89 | $947,189 | 33.73% |
| Brand 03 | $3,185,036 | 112 | $795,443 | 24.97% |
| Brand 04 | $1,544,969 | 143 | $679,568 | 43.99% |
| Brand 05 | $2,010,101 | 153 | $614,575 | 30.57% |
| Brand 06 | $1,871,047 | 124 | $591,029 | 31.59% |
| Brand 07 | $1,425,490 | 63 | $517,905 | 36.33% |
| Brand 08 | $1,637,335 | 373 | $509,862 | 31.14% |
| Brand 09 | $1,267,769 | 119 | $411,207 | 32.44% |
| Brand 10 | $1,048,679 | 65 | $394,818 | 37.65% |
| Brand 11 | $1,206,563 | 56 | $338,136 | 28.02% |
| Brand 12 | $856,304 | 68 | $287,351 | 33.56% |
| Brand 13 | $811,898 | 22 | $273,387 | 33.67% |
| Brand 14 | $1,680,656 | 9 | $221,766 | 13.20% |
| Brand 15 | $648,930 | 33 | $154,806 | 23.86% |
| Brand 16 | $568,465 | 67 | $149,257 | 26.26% |
| Brand 17 | $963,645 | 23 | $142,632 | 14.80% |
| Brand 18 | $351,422 | 45 | $133,843 | 38.09% |
| Brand 19 | $342,055 | 76 | $132,344 | 38.69% |
| Brand 20 | $275,878 | 34 | $118,882 | 43.09% |
| Brand 21 | $409,161 | 93 | $105,629 | 25.82% |
| Brand 22 | $271,553 | 43 | $73,055 | 26.90% |
| Brand 23 | $241,938 | 36 | $72,451 | 29.95% |
| Brand 24 | $208,931 | 23 | $65,862 | 31.52% |
| Brand 25 | $237,796 | 74 | $63,119 | 26.54% |
| Brand 26 | $221,757 | 45 | $57,141 | 25.77% |
| Brand 27 | $322,008 | 34 | $55,674 | 17.29% |
| Brand 28 | $177,612 | 22 | $51,491 | 28.99% |
| Brand 29 | $132,572 | 32 | $46,427 | 35.02% |
| Brand 30 | $191,594 | 19 | $44,975 | 23.47% |
| Brand 31 | $308,371 | 50 | $41,891 | 13.58% |
| Brand 32 | $167,264 | 57 | $38,720 | 23.15% |
| Brand 33 | $223,387 | 32 | $38,514 | 17.24% |
| Brand 34 | $136,355 | 21 | $36,365 | 26.67% |
| Brand 35 | $210,482 | 28 | $34,565 | 16.42% |
| Brand 36 | $197,357 | 5 | $17,062 | 8.65% |
| Brand 37 | $114,552 | 10 | $12,716 | 11.10% |
| Portfolio | $32,292,369 | 2461 | $9,436,899 | 29.22% |
The mechanism is worth roughly a fifth to a third of an account, and it replicates. The order-of-magnitude difference between these two brands came from the ceiling, not the mechanism — Brand B's efficiency improved 84%, and a budget process is only ever worth as much as the headroom it has to spend into.
Which sets the operating rule: establish what an acquisition is worth in the category before touching anything. Where order value supports a higher bid, releasing budget converts close to one-for-one. Where it does not, the constraint sits somewhere else entirely and no amount of budget will move it.