How We Test the Engine Before It Touches a Real Account
Inside FoxtrotAI's marketplace simulator and A/B testing harness: how we run a fair fight against a capable baseline seller, what we measured, and why a simulation matters.
- We built a simulated Amazon-like marketplace and ran the same market twice: once managed by a competent baseline seller, once by the FoxtrotAI engine. Only the decision-maker differed.
- Across 40 paired markets, the engine earned 7.6% more profit after fees and ad spend. The edge held (+7.2% to +7.9%) even when ad budgets were cut by up to 90%, and the engine found a setting that used 39% less ad spend for only 1.4% less profit.
- These are simulation results. They show the engine is robustly tested, not what any specific account will earn. Real-world results will vary.
Why we don't test on your account
Any software that changes your prices and your ad bids should have to prove itself before it goes anywhere near your business. The obvious way to prove it, running it on live accounts and watching what happens, turns out to be a surprisingly bad experiment. Real accounts are noisy. Sales move with seasonality, with competitors' decisions, with review counts, with luck. If profit rises 10% the month after you switch tools, was that the tool, or was it seasonal traffic and a competitor running out of stock?
There is also no honest control group. No two products are alike, so tests where we manage one product manually and another with FoxtrotAI prove very little. The key question we want to answer is: what would my account have earned over the same timeframe without FoxtrotAI? This is unanswerable in the real world, since you only get to live each month once.
So before our optimization engine ever managed real money, we built a place where that question can be answered exactly: a simulated marketplace, and a testing harness that runs the same market twice.
A practice marketplace
The simulator is a working model of a marketplace like Amazon. Simulated shoppers search, compare, click, and buy, and they respond to price the way real shoppers do: a higher price wins more per sale but converts fewer clicks. Competing sellers bid against you in the same ad auctions. Organic ranking responds to sales momentum. Referral and fulfillment fees are charged on every order. Daily ad budgets run out. Demand shifts across the weeks.
Each simulated market is generated fresh, with its own randomized demand, competitors, and cost structure, so no single lucky market can carry a result. And the engine connects to the simulator the same way it connects to Amazon: it sees only the reports and results a real seller would see. It cannot peek at the simulator's hidden settings. It has to learn the market the same way it would have to learn in the real world.
A fair fight, run forty times
The harness stages a head-to-head competition. Each simulated market is run twice: once managed by a sensible baseline seller, and once managed by our engine. Same shoppers, same competitors, same fees, same budgets, same starting conditions. The only difference between the two runs is who is making the pricing and bidding decisions.
The baseline seller is not a strawman. They are competent. They manage the account the way a careful, hands-on seller does. They steer advertising by an ACoS target: they work out each product's break-even ACoS from that product's real margins and fees, then target safely below it. Every week the baseline seller reviews each keyword against that target using their actual ad spend and ad revenue, nudging bids up or down. They reset starving keywords to Amazon's suggested bid so they stay in the game, negate keywords that keep getting clicks but never convert, and reprice toward competitor prices every couple of weeks. They budget the way sellers budget, setting daily ad spend as a share of expected sales (a TACoS-style budget), and when a day's budget is spent, their ads stop for the day.
We rebuilt this baseline seller more than once specifically to make it stronger, because beating a weak opponent proves nothing.
We ran forty of these paired markets and measured one pre-registered outcome: profit dollars per product over the test window. Not sales, not clicks, not ad efficiency ratios. Profit, after fees and ad spend.
The result: the engine earned 7.6% more profit than the baseline seller across the forty paired markets. To check whether that could be luck, we use a standard statistical shuffle test: scramble which run was "engine" and which was "baseline" thousands of times and ask how often a gap this large appears by accident. The answer is about one in seven hundred. We also run the harness against itself, with the same strategy on both sides, to confirm it reports "no difference" when there is none. It does.
A 7.6% profit increase might sound modest next to the claims you may be used to seeing. But it is profit after all fees and ad spend and it was earned against a competent opponent in a controlled test rather than in a best-case story. On a business with thin margins, keeping 7.6% more of every profit dollar is rarely modest.
In these tests, the classic advertising ratios, ACoS and ROAS, were statistically identical between the two arms. Those metrics are useful for reporting, but they could not see the difference. The extra profit came from healthier margins and from getting more clicks at a lower average cost for the same spend. This is why FoxtrotAI steers by profit dollars rather than by ad ratios.
Then we tried to break it
A single favorable result is not robustness, so we stress-tested it. We re-ran the head-to-head with much tighter money, and with the simulated platform strictly enforcing daily budget caps: once a budget ran out, ads stopped serving for the day, the way they do on Amazon.
Two findings from the stress runs stand out. First, when money got tight, the engine's edge did not shrink. In the tightest scenarios its advantage held or grew, because deciding carefully where each dollar goes matters most when there are fewer dollars. Second, in a test focused on spend efficiency, the engine found an operating point that used 39% less ad spend for only 1.4% less profit. For a seller who cares about cash flow as much as top-line growth, that trade is often the more valuable discovery.
We also went hunting for conditions where the engine loses its edge, and we found some. Those findings changed how the engine is configured, and they define what we watch for on real accounts.
Why a simulator is a good test
A fair question: why should a simulated result give you any confidence at all? Three reasons.
It answers the question reality can't. The paired design, the same market lived twice with only the decision-maker changed, is the cleanest possible comparison, and it is impossible to run on a real account. Every dollar of difference in these tests is attributable to the engine's decisions and nothing else.
It compresses years into days. Forty markets, each simulated over many weeks, with thousands of shopper interactions per day, is more decision-making experience than any tool could accumulate on live accounts in years, and every mistake made along the way was made with simulated money.
It lets us test the bad days. No responsible company would deliberately slash a paying client's budgets to a tenth, or feed their account wrong assumptions, just to see what happens. In the simulator we do exactly that, on purpose, repeatedly. The engine you get has already been through the scenarios we hope your account never sees.
What it cannot tell you
A simulator is a model. It has its limits. Simulated shoppers are built to behave like real ones, but no model captures everything about a real category: your competitors' psychology, a sudden review swing, a supply shock, a fad. The strength of some effects in the simulator, like how paid sales momentum feeds organic ranking, will vary as Amazon's algorithms change.
So these numbers are evidence of capability, not a forecast of any seller account. Real-world results will vary, by category, by competition, by season, and by how much room your products have to improve. What the simulation establishes is narrower: that the engine's decision-making beats a competent human-style baseline in a controlled, repeatable, honestly-scored test, and keeps beating it when conditions get hard.