A program that works perfectly when called directly can be impossible to use inside an aggregator route. Nothing about it is broken — it just consumes budgets that its callers needed.
★Depth, account slots, and compute are shared across the whole transaction. A program that spends generously is a program that cannot be composed with.★
The Budgets You Are Spending
your instruction ← frame 1 (the caller's entry point)
└─ ★aggregator★ ← nested 1
└─ ★your program★ ← nested 2
└─ ★your CPI★ ← nested 3
└─ token program ← ★nested 4 — ceiling★
★If your program uses two nested CPIs, a caller invoking you through an aggregator has already spent the stack before your innermost call runs.★
The same applies to the other two budgets:
| Budget | Limit | What your program spends |
|---|---|---|
| ★CPI depth★ | ★4 nested levels★ | ★Every nested call★ |
| ★Accounts★ | ★1232 bytes total★ | ★32 bytes per account you require★ |
| ★Compute★ | ★Per transaction★ | ★Your usage plus CPI overhead★ |
None of these are per-program allowances. They belong to the transaction, and your callers are sharing them with you.
Flatten Where You Can
The most common avoidable depth cost is a helper program that exists for organisation:
// ★Three levels for what could be two.★
your_program
└─ your_helper_program
└─ token_program
// ★Two levels.★
your_program
└─ token_program
★A wrapper that adds no on-chain guarantee is depth your callers pay for.★ If the logic can live in your program rather than behind another invocation, the composability gain usually outweighs the code organisation loss.
Where the CPI is genuinely necessary — you need another program's state transition — the depth is the price of the feature, and that is a fair trade.
If your program composes well and transactions still miss, a free BoltTx key is one line for your users to test the submission path.
Ask for the Minimum Account Set
Every account in your instruction context costs 32 bytes in the caller's transaction:
#[derive(Accounts)]
pub struct Swap<'info> {
pub pool: Account<'info, Pool>,
pub user_source: Account<'info, TokenAccount>,
pub user_dest: Account<'info, TokenAccount>,
pub token_program: Program<'info, Token>,
// ★Do you actually need these?★
pub rent: Sysvar<'info, Rent>, // ★usually not, any more★
pub system_program: Program<'info, System>, // ★only if creating accounts★
}
★Sysvars like rent and clock are readable directly in most cases, without being passed as accounts.★ Requiring them is a legacy pattern that still costs 32 bytes each in every caller's transaction.
Derived accounts are worse than they look. If your program can derive a PDA itself, requiring the caller to pass it adds bytes for something you were going to compute anyway — though it does save the compute of deriving it on chain, which is a real tradeoff rather than a clear win either way.
Validate Before the Expensive Work
Ordering inside your instruction changes what a failure costs:
pub fn swap(ctx: Context<Swap>, amount: u64, min_out: u64) -> Result<()> {
// ★Cheap checks first — fail before spending compute.★
require!(amount > 0, SwapError::AmountZero);
require!(!ctx.accounts.pool.is_paused, SwapError::PoolPaused);
// ★Then the expensive parts.★
let out = compute_output(&ctx.accounts.pool, amount)?;
require!(out >= min_out, SwapError::SlippageExceeded);
token::transfer(ctx.accounts.transfer_ctx(), amount)?; // ★CPI last★
Ok(())
}
★A revert after the CPI wasted the CPI's compute; a revert before it did not.★ The transaction pays the base fee either way, but the caller's compute limit is a shared resource, and failing cheaply leaves them room to retry within the same budget.
Emit Events for What Callers Cannot Compute
#[event]
pub struct SwapExecuted {
pub amount_in: u64,
pub amount_out: u64, // ★after all internal adjustments★
pub fee_paid: u64,
}
★The transaction meta already reports balance changes, so events duplicating those add compute for nothing.★
What is worth emitting is internal state a caller cannot derive from balances — which fee tier applied, which route was taken internally, what the pool state was at execution. That information exists only inside your program, and without it callers reconstruct it by guessing.
Signing With PDAs
When your program needs to authorise a transfer from an account it owns:
let seeds = &[b"vault", mint.as_ref(), &[bump]];
let signer = &[&seeds[..]];
token::transfer(
CpiContext::new_with_signer(token_program, accounts, signer),
amount,
)?;
★Store the bump rather than calling find_program_address on chain.★ The search loop runs every invocation and consumes compute your callers pay for, to recompute a value you could have saved at initialisation.
A program can only sign for PDAs derived from its own program ID, which is the security boundary — you cannot forge a signature for a caller's wallet or another program's PDA.
Design for Being Called by a Bot
★The composability question that matters most: can your instruction be one leg of an atomic multi-step transaction?★
What makes that possible:
Modest depth, so there is room above you for a router.
A small account set, so the transaction still fits.
Predictable compute, published so callers can size their limit.
★Errors that distinguish retryable from permanent★, so a bot knows what to do next.
What makes it impossible is usually not one large problem but the accumulation — a wrapper here, a sysvar there, an unpublished compute figure — until a route that should fit does not.
What Landing Looks Like
Real transactions through our delivery nodes: median confirmation 336ms — under one slot.
★Composability decides whether a transaction can be built; routing decides when it lands.★ A program that leaves room for its callers is solving the first problem, which is the one no submission path can fix.
Where BoltTx Fits
We handle submission for the people calling your program. CPI structure and account requirements are decided in your program.
Submissions route through our own delivery nodes in four regions with stake-weighted routing and no public mempool exposure, so a transaction is not observable in transit before it lands. We never modify transaction contents, which includes never altering the instruction set or account list.
Your callers sign locally. We never hold funds, never sign, and never modify transaction contents. The tip travels inside the transaction, paid on chain from their own wallet, and reverts with the transaction if it fails, because that is how Solana handles atomic transactions. They pay only on transactions that reach the chain.
Get a free API key. No monthly fee:
const connection = new Connection("https://la.bolttx.io/?api-key=YOUR_KEY");
FAQ
How does my program's CPI usage affect callers? Depth, accounts, and compute are shared across the transaction rather than allocated per program. Every level you nest is one a caller routing through an aggregator no longer has.
What is the CPI depth limit on Solana? Four levels of invocation, counting the caller's instruction as the first. A program using two nested CPIs leaves almost no room for anyone invoking it indirectly.
Should I avoid helper programs? Where they add no on-chain guarantee, yes. A wrapper that exists for code organisation costs your callers a depth level on every transaction.
Do I still need to pass the rent sysvar? Usually not. Sysvars are readable directly in most cases, and requiring them costs 32 bytes in every caller's transaction for no benefit.
Should my program derive PDAs or require them as accounts?
It is a real tradeoff. Requiring them costs the caller bytes; deriving them costs compute. Storing the bump and using create_program_address avoids the expensive search either way.
Why store the PDA bump?
Because find_program_address runs a search loop on every invocation, consuming compute your callers pay for, to recompute a value you could have saved at initialisation.
Where should validation go in an instruction? Cheap checks first, expensive work and CPIs last. A revert before a CPI wastes less of the caller's compute budget than one after it.
What events are worth emitting? Internal state a caller cannot derive from balance changes, such as which fee tier applied. Events duplicating what the transaction meta already reports add compute for nothing.
How many accounts should an instruction require? As few as genuinely needed. Each costs 32 bytes against the 1232-byte transaction limit, which is usually what stops a multi-hop route from fitting.
Can my program sign for any account? Only PDAs derived from its own program ID. It cannot forge a signature for a caller's wallet or for another program's PDA, which is the security boundary.
Why does my program fail inside an aggregator route but work directly? Almost certainly a shared budget. The route has already consumed depth, accounts, or compute before reaching you, and your requirements exceed what is left.
How do I make my program composable? Modest depth, a small account set, published compute figures, and errors distinguishing retryable from permanent. Composability is usually lost to accumulation, not one large mistake.
Does batching CPIs reduce cost? Where the target program's interface allows it, yes. Each CPI carries setup overhead beyond what the called program spends, so fewer calls means less overhead.
Should I publish my program's compute usage? Yes. Callers who know the figure set an accurate limit, while those who do not default high and overpay, then attribute the cost to network conditions.
How do CPI accounts work? Every account any nested program touches must be present in the original transaction. Your program cannot invent them, which is why account requirements propagate up to your callers.
What is the most common composability mistake? Requiring accounts the program does not need, usually legacy sysvars, combined with an unnecessary wrapper program consuming a depth level.