Why Backtests Overstate Solana Bot Performance

A backtest replays a world where you were the only participant. The contention, latency, and impact that decide live results are the parts it cannot simulate.

BoltTx Team··9 min read
solanabacktestingsimulationtrading-botexecutiontransaction-landing

The backtest shows a strong edge. The bot goes live and earns a fraction of it, or nothing.

★This is not usually a modelling error in the strategy. It is that a backtest replays history as though you were not in it.★

What a Backtest Silently Assumes

Assumption Reality
★You get filled★ ★You compete for the same slot★
★At the observed price★ Your size moves it
★Instantly★ Slots pass between decision and execution
★Every time★ Reverts and expiries cost fees
★Without moving the market★ Others react to your flow

★Every one of these is optimistic in the same direction.★ They compound, which is why the gap between backtest and live is usually much larger than any single assumption would suggest.

The Contention Problem

★The opportunity you replayed was visible to everyone else at the same moment.★

A backtest asks "was there a profitable trade here?" Live trading asks "would I have won it?" — a different question with a different answer, and historical data contains no record of who else was trying.

// ★Backtest: the opportunity existed, so you took it.★
if (spread > threshold) profit += spread * size;

// ★Live: you took it only if you landed first.★
if (spread > threshold && landedBefore(competitors)) profit += ...;

There is no honest way to simulate the second line from historical data. The best available substitute is to assume you win some fraction of contested opportunities and see whether the strategy survives it.

★A strategy that only works at a 100% fill rate is not a strategy.★

Model the Execution Costs You Can

Three of the five assumptions above are computable, and fixing them removes most of the illusion:

function backtestTrade(opportunity, pool, size) {
  // ★1. Price impact — deterministic from reserves.★
  const out = amountOut(size, pool.reserveIn, pool.reserveOut, pool.feeBps);

  // ★2. Fees on every attempt, not just successes.★
  const fees = BASE_FEE + assumedPriorityFee;

  // ★3. Failure rate — attempts that cost fees and return nothing.★
  const landed = Math.random() > assumedFailureRate;

  return landed ? out - size - fees : -fees;
}

★The third line is the one most backtests omit entirely.★ A bot that lands 60% of its attempts pays fees on the other 40%, and at high frequency that is a substantial drag that never appears in a naive replay.

If your backtest accounts for execution and live results still differ, a free BoltTx key is one line to test the submission path.

Use Slot Time, Not Wall-Clock Time

// ★Wrong: seconds have no meaning for inclusion.★
const executed = decision.timestamp + 500;

// ★Right: model the delay in slots.★
const executionSlot = decisionSlot + assumedLandingDelay;
const state = reservesAtSlot(executionSlot);   // ★not at decision time★

★The state you execute against is the state at the landing slot, not the one you decided on.★ A backtest that prices fills at decision-time state is measuring a trade that could not have happened.

This single correction often removes most of a backtest's apparent edge, because opportunities that survive three slots are much rarer than opportunities that exist at an instant.

Where the Data Itself Misleads

Survivorship. Tokens that rugged or died are often missing from historical datasets, so a strategy tested on what remains was tested on the winners.

Aggregated candles. A one-minute bar hides the ordering within it. A wick your strategy "caught" may have occurred before your entry signal.

★Reconstructed reserves.★ ★Pool state derived from swap events is an approximation, and errors compound across a long replay.★

Missing failures. On-chain history contains landed transactions. ★Transactions that expired left no record anywhere★, so a replay built from chain data cannot see how often anyone failed to land.

What to Do Instead

★Backtesting is for rejecting strategies, not for validating them.★

A strategy that fails a backtest will fail live. That is a useful filter and it is cheap.

A strategy that passes tells you very little, because the assumptions above all favoured it.

The sequence that produces real information:

1. Backtest to reject. Cheap, fast, catches the obviously unprofitable.

2. Stress the assumptions. Rerun with a 50% fill rate, doubled fees, and three slots of delay. ★If the edge disappears, it was an artifact.★

3. Paper trade live. Compute decisions against real-time state without submitting. This catches detection and logic errors under real conditions.

4. Live at minimum size. ★The only way to measure your actual fill rate and slot distribution.★ Budget it as a measurement cost.

5. Compare live against backtest. The gap is your execution cost, and it is the number worth optimising.

Measure the Gap Deliberately

metrics.record({
  backtestExpected: expectedProfit,
  actualRealised: realisedProfit,
  slotDelay: landSlot - decisionSlot,
  outcome,                              // ★success | reverted | expired★
});

★A persistent gap that grows with slot delay tells you the problem is submission. A gap that is constant regardless of delay tells you the strategy itself was overfitted.★

That distinction is worth more than any backtest refinement, and you can only get it by running both.

What Landing Looks Like

Real transactions through our delivery nodes: median confirmation 336ms — under one slot.

★That is the delay to feed into a backtest instead of assuming instant execution.★ Replaying with a realistic landing delay is the single change that brings a backtest closest to what live trading actually produces.

Where BoltTx Fits

We handle submission. Backtesting and strategy design stay entirely in your code.

Submissions route through our own delivery nodes in four regions with stake-weighted routing and no public mempool exposure, so a transaction is not observable in transit before it lands — which reduces one of the execution costs a backtest cannot model.

You sign locally. We never hold funds, never sign, and never modify transaction contents. The tip travels inside the transaction, paid on chain from your own wallet, and reverts with the transaction if it fails, because that is how Solana handles atomic transactions. You pay only on transactions that reach the chain, which makes a small live measurement run inexpensive.

Get a free API key. No monthly fee:

const connection = new Connection("https://la.bolttx.io/?api-key=YOUR_KEY");

FAQ

Why do Solana bot backtests overstate performance? Because they replay history as though you were the only participant. Contention, execution delay, price impact, and failed attempts are all absent or optimistic.

What is the biggest backtest error? Assuming instant execution. Pricing fills at decision-time state measures a trade that could not have happened, since the state you execute against is the one at the landing slot.

How do I model competition in a backtest? You cannot model it directly, since historical data contains no record of who else was trying. Instead assume you win a fraction of contested opportunities and check whether the edge survives.

Should failed transactions be included in a backtest? Yes. A bot landing 60% of attempts pays fees on the other 40%, and at high frequency that drag is substantial. Naive replays count only successes.

How do I model price impact in a backtest? Compute the output from the pool reserves at your intended size rather than using the observed price. Impact is deterministic, so this is the easiest correction to make.

Should I use slots or seconds for timing? Slots. Seconds have no meaning for inclusion, and slot production varies. Model the delay as a number of slots and evaluate against state at the landing slot.

Why is survivorship bias a problem for token strategies? Because tokens that died are often missing from historical datasets. A strategy tested on what remains was tested on winners, which flatters any entry logic.

Can I reconstruct pool reserves from historical swaps? Approximately, but errors compound across a long replay. Treat reconstructed state as noisier than live reads, especially for strategies sensitive to small differences.

Why can't a backtest see failed transactions? On-chain history records landed transactions. Expired ones left no trace anywhere, so a replay built from chain data cannot measure how often anyone failed to land.

What is a realistic fill rate assumption? Whatever your live measurement shows, which is why step four exists. Before you have that, stress-testing at 50% is a reasonable way to check whether the edge is robust.

Should I trust a backtest that shows a large edge? Be more suspicious, not less. Every unmodelled assumption favours the backtest, so a large apparent edge often means more assumptions were doing the work.

What is paper trading good for? Catching detection and logic errors under real-time conditions without risking capital. It validates that your decisions are correct, though not that you would have won the race.

How much should I risk in the first live run? Minimum size, budgeted as a measurement cost rather than a profit attempt. Fill rate and slot distribution cannot be observed any other way.

How do I tell overfitting from an execution problem? Track the backtest-versus-live gap against slot delay. A gap that grows with delay is submission; a gap that is constant regardless of delay is the strategy.

Does a faster submission path change backtest accuracy? It narrows the gap between backtest and live by reducing execution delay, but the backtest still needs to model that delay honestly rather than assuming zero.

What should a backtest actually be used for? Rejecting strategies. A failing backtest is reliable information; a passing one mostly tells you the assumptions were favourable.

← Back to all posts