FoxtrotAI · Research Notes

How We Test the Engine Before It Touches a Real Account

Inside FoxtrotAI's marketplace simulator and A/B testing harness: how we run a fair fight against a capable baseline seller, what we measured, and why a simulation matters.

FoxtrotAI · July 2026 · 6 min read

TL;DR

Why we don't test on your account

Any software that changes your prices and your ad bids should have to prove itself before it goes anywhere near your business. The obvious way to prove it, running it on live accounts and watching what happens, turns out to be a surprisingly bad experiment. Real accounts are noisy. Sales move with seasonality, with competitors' decisions, with review counts, with luck. If profit rises 10% the month after you switch tools, was that the tool, or was it seasonal traffic and a competitor running out of stock?

There is also no honest control group. No two products are alike, so tests where we manage one product manually and another with FoxtrotAI prove very little. The key question we want to answer is: what would my account have earned over the same timeframe without FoxtrotAI? This is unanswerable in the real world, since you only get to live each month once.

So before our optimization engine ever managed real money, we built a place where that question can be answered exactly: a simulated marketplace, and a testing harness that runs the same market twice.

A practice marketplace

The simulator is a working model of a marketplace like Amazon. Simulated shoppers search, compare, click, and buy, and they respond to price the way real shoppers do: a higher price wins more per sale but converts fewer clicks. Competing sellers bid against you in the same ad auctions. Organic ranking responds to sales momentum. Referral and fulfillment fees are charged on every order. Daily ad budgets run out. Demand shifts across the weeks.

Each simulated market is generated fresh, with its own randomized demand, competitors, and cost structure, so no single lucky market can carry a result. And the engine connects to the simulator the same way it connects to Amazon: it sees only the reports and results a real seller would see. It cannot peek at the simulator's hidden settings. It has to learn the market the same way it would have to learn in the real world.

A fair fight, run forty times

The harness stages a head-to-head competition. Each simulated market is run twice: once managed by a sensible baseline seller, and once managed by our engine. Same shoppers, same competitors, same fees, same budgets, same starting conditions. The only difference between the two runs is who is making the pricing and bidding decisions.

Simulated marketplace shoppers · competitors ad auctions · fees organic ranking · budgets 40 fresh markets Run A: Baseline seller manages prices and bids the exact same market, lived twice Run B: FoxtrotAI engine manages prices and bids Compare profit after fees and ads
The A/B harness. Each fresh simulated market is copied and run twice: once under the baseline seller, once under the FoxtrotAI engine, with everything else identical. The harness then compares the profit of the two runs. This is repeated across forty markets, and again under the stress scenarios below.

The baseline seller is not a strawman. They are competent. They manage the account the way a careful, hands-on seller does. They steer advertising by an ACoS target: they work out each product's break-even ACoS from that product's real margins and fees, then target safely below it. Every week the baseline seller reviews each keyword against that target using their actual ad spend and ad revenue, nudging bids up or down. They reset starving keywords to Amazon's suggested bid so they stay in the game, negate keywords that keep getting clicks but never convert, and reprice toward competitor prices every couple of weeks. They budget the way sellers budget, setting daily ad spend as a share of expected sales (a TACoS-style budget), and when a day's budget is spent, their ads stop for the day.

We rebuilt this baseline seller more than once specifically to make it stronger, because beating a weak opponent proves nothing.

We ran forty of these paired markets and measured one pre-registered outcome: profit dollars per product over the test window. Not sales, not clicks, not ad efficiency ratios. Profit, after fees and ad spend.

Profit per product, average across 40 paired simulated markets Same markets, same conditions; only the decision-maker differs Baseline seller $2,938 FoxtrotAI engine $3,163 +7.6%
Average profit per product over the simulated test window, across forty paired markets. The engine arm earned 7.6% more profit than the baseline arm managing the identical markets. A shuffle test puts the odds of a gap this consistent arising by chance at roughly one in seven hundred.

The result: the engine earned 7.6% more profit than the baseline seller across the forty paired markets. To check whether that could be luck, we use a standard statistical shuffle test: scramble which run was "engine" and which was "baseline" thousands of times and ask how often a gap this large appears by accident. The answer is about one in seven hundred. We also run the harness against itself, with the same strategy on both sides, to confirm it reports "no difference" when there is none. It does.

A 7.6% profit increase might sound modest next to the claims you may be used to seeing. But it is profit after all fees and ad spend and it was earned against a competent opponent in a controlled test rather than in a best-case story. On a business with thin margins, keeping 7.6% more of every profit dollar is rarely modest.

In these tests, the classic advertising ratios, ACoS and ROAS, were statistically identical between the two arms. Those metrics are useful for reporting, but they could not see the difference. The extra profit came from healthier margins and from getting more clicks at a lower average cost for the same spend. This is why FoxtrotAI steers by profit dollars rather than by ad ratios.

Then we tried to break it

A single favorable result is not robustness, so we stress-tested it. We re-ran the head-to-head with much tighter money, and with the simulated platform strictly enforcing daily budget caps: once a budget ran out, ads stopped serving for the day, the way they do on Amazon.

The engine's profit advantage under stress Profit lift vs. the baseline seller, each scenario statistically significant Normal budgets +7.6% Budgets cut in half +7.6% Budgets cut to a quarter +7.9% Budgets cut to a tenth +7.2%
The engine's profit advantage over the baseline seller with ad budgets at normal levels and cut to one half, one quarter, and one tenth. In the reduced-budget scenarios the simulated platform strictly enforced daily budget caps, stopping ads when a budget ran out, the way Amazon does. The advantage held between +7.2% and +7.9% in every scenario, each one statistically significant.

Two findings from the stress runs stand out. First, when money got tight, the engine's edge did not shrink. In the tightest scenarios its advantage held or grew, because deciding carefully where each dollar goes matters most when there are fewer dollars. Second, in a test focused on spend efficiency, the engine found an operating point that used 39% less ad spend for only 1.4% less profit. For a seller who cares about cash flow as much as top-line growth, that trade is often the more valuable discovery.

We also went hunting for conditions where the engine loses its edge, and we found some. Those findings changed how the engine is configured, and they define what we watch for on real accounts.

Why a simulator is a good test

A fair question: why should a simulated result give you any confidence at all? Three reasons.

It answers the question reality can't. The paired design, the same market lived twice with only the decision-maker changed, is the cleanest possible comparison, and it is impossible to run on a real account. Every dollar of difference in these tests is attributable to the engine's decisions and nothing else.

It compresses years into days. Forty markets, each simulated over many weeks, with thousands of shopper interactions per day, is more decision-making experience than any tool could accumulate on live accounts in years, and every mistake made along the way was made with simulated money.

It lets us test the bad days. No responsible company would deliberately slash a paying client's budgets to a tenth, or feed their account wrong assumptions, just to see what happens. In the simulator we do exactly that, on purpose, repeatedly. The engine you get has already been through the scenarios we hope your account never sees.

What it cannot tell you

A simulator is a model. It has its limits. Simulated shoppers are built to behave like real ones, but no model captures everything about a real category: your competitors' psychology, a sudden review swing, a supply shock, a fad. The strength of some effects in the simulator, like how paid sales momentum feeds organic ranking, will vary as Amazon's algorithms change.

So these numbers are evidence of capability, not a forecast of any seller account. Real-world results will vary, by category, by competition, by season, and by how much room your products have to improve. What the simulation establishes is narrower: that the engine's decision-making beats a competent human-style baseline in a controlled, repeatable, honestly-scored test, and keeps beating it when conditions get hard.

Methodology, briefly: forty paired simulated markets, generated with randomized demand, competition, and cost conditions fixed in advance of the run. Primary metric: profit dollars per product over the simulated test window, after fees and ad spend. Significance assessed with a paired permutation ("shuffle") test on the pre-registered primary metric (p ≈ 0.0015); an A/A run of the harness against itself showed no false difference. Stress scenarios re-ran the same head-to-head at one half, one quarter, and one tenth of normal ad budgets, with the simulated platform strictly enforcing daily budget caps in those runs. All figures in this post are simulation results and are not claims about results on any live account.
Disclaimer: The performance figures in this post come exclusively from the simulated marketplace environments described above. They are provided to illustrate FoxtrotAI's testing methodology and do not represent, and should not be relied on as, a prediction, promise, or guarantee of the results any seller will achieve using FoxtrotAI. Actual results depend on factors outside FoxtrotAI's control, including product category, competition, pricing, inventory, advertising budgets, and marketplace conditions, and will vary from seller to seller. Nothing in this post constitutes business, financial, or legal advice.