Paper Trading Fills vs Live Execution: How to Review the Gap
Review paper trading fills against documented simulator limits, quote timing, costs, and order receipts without treating hypothetical results as live evidence.
Put this into practice with a watchlist
Build a watchlist, then review each signal’s entry, stop, target, and reasoning. Broker access is optional.
Paper trading fills are simulated outcomes. They can help rehearse an order workflow, but they do not establish the price, quantity, or timing a real order would receive. Review the simulator’s documented assumptions, the quote used, the order timeline, and excluded costs before drawing conclusions. Alpaca explicitly identifies limitations in its paper environment, including market impact, queue position, and latency-related slippage. Alpaca paper-trading documentation.
A useful comparison produces a list of measured differences and unresolved assumptions. It does not convert paper gains into expected live returns. Trading involves risk; this guide is educational information, not financial advice or permission to place a live order.
Separate a recorded decision from a simulated fill
The term paper trading can describe different activities. A decision journal records an idea and what happened afterward. A brokerage simulator accepts instructions and applies a model to produce simulated order states and fills. A historical backtest applies rules to a dataset. These activities answer different questions.
Before evaluating results, identify which activity produced them. A journal entry can show that you recorded an entry idea before a later price move. It does not prove that a broker would have filled an instruction at that entry. A simulator receipt shows what its model did, not what the live market did. A backtest can reproduce a rule under stated assumptions, but its fill model may be even simpler. Keep these labels in the report instead of merging them into one column called trades. The paper-trading workflow explains the decision-journal distinction in more detail.
Write an assumption register before reviewing returns
Begin with an assumption register: a short list of what the experiment includes, excludes, and leaves unknown. Record the source of market observations, fill model, available order types, quantity limits, session, and any modeled transaction costs. Add a date and a link to the provider documentation supporting each claim.
The register should use plain labels such as documented, observed in this test, and unknown. These labels prevent a demonstration from becoming an unsupported guarantee. If a simulator reports a fill but does not explain its treatment of resting orders, write that uncertainty down. If your own backtest assumes every eligible bar produces a complete fill, disclose that assumption directly. Neither more sophisticated code nor a longer test period removes an unknown mechanism. Readers need to understand which part of the result comes from observations and which part comes from the model you chose to apply.
Use one reference price consistently
Slippage comparisons need a defined reference. A fill can be compared with a decision-time quote, submission-time quote, or another explicitly chosen benchmark. These references are not interchangeable. Changing the benchmark after observing a fill changes the meaning of the comparison.
For a fictional buy, suppose the chosen reference is 100.00 and the recorded fill is 100.04. The difference is 0.04 per share, or four basis points relative to that reference. This is arithmetic, not a measured broker result. For a sell, define the sign convention so a worse price is represented consistently. Keep the raw reference, side, quantity, and fill price with every computed value. A number labeled slippage without its benchmark is difficult to interpret. Also keep fees separate: a price difference and an explicit charge are different components, even if both affect the total result.
Capture the full order timeline
Record when the decision was made, when the instruction was submitted, when it was acknowledged, and when each fill or status update appeared. Use the provider’s event timestamps where available and distinguish them from your application’s receipt timestamps. Clock differences should remain visible rather than being silently treated as zero.
A timeline helps locate a discrepancy without claiming to identify its cause automatically. A large gap between a decision and submission may come from your workflow. A gap between submission and acknowledgement may need network or provider investigation. A later fill may reflect the instruction’s price conditions or other behavior. These are hypotheses to test, not diagnoses from elapsed time alone. For paper testing, rehearse collecting this evidence while no capital is committed. A polished profit chart cannot reconstruct an execution sequence if the underlying event records were never retained.
Partial fills and nonfills belong in the sample
A performance report that keeps only completed fills omits important behavior. Record rejected instructions, expired instructions, remaining quantities, and opportunities that never produced a fill. State how each state enters the analysis. Otherwise, readers cannot determine whether an apparently favorable result depends on ignoring difficult cases.
Partial fills need both quantity and price records. If two fictional fills execute different quantities at different prices, calculate the quantity-weighted average rather than averaging the prices equally. Retain the individual events so the calculation can be checked. An unfilled remainder is not the same as an executed position, and a canceled instruction is not evidence of a zero-cost trade. Do not invent fills to complete the journal. The missing or incomplete cases may reveal more about the workflow than the easiest simulated executions, especially when a strategy relies on prompt entry and exit.
A comparison worksheet that keeps evidence clear
| Field | Paper record | Live record, if separately available |
|---|---|---|
| Environment | Named simulator and configuration | Broker account environment |
| Evidence type | Simulated receipt or journal decision | Authoritative order and fill record |
| Reference | Defined quote and timestamp | Same benchmark definition |
| Quantity | Requested, filled, remaining | Requested, filled, remaining |
| Costs | Modeled, excluded, or unknown | Observed charges and known omissions |
| Status | Full simulated timeline | Full observed timeline |
| Limits | Documented model assumptions | Missing observations and uncertainties |
A live column is optional. This guide does not require placing real trades to populate it. If authorized historical receipts already exist, they can be reviewed separately under an explicit methodology. Without them, label live execution unmeasured. Do not use another person’s screenshot or a provider’s aggregate claim as a substitute for matching records.
Put the setup on a watchlist first
Use the rules in this guide to evaluate a signal’s entry, stop, target, and reasoning before deciding what, if anything, to do.
Stress-test assumptions without claiming realism
You can study how a hypothetical result changes under less favorable assumptions. For example, recompute a backtest with a higher assumed execution cost, delayed entry, or a rule that leaves some eligible opportunities unfilled. Describe each scenario as a sensitivity test rather than an observed live result.
Choose the scenarios before inspecting which one preserves a favorable conclusion. State the rationale for their ranges, and disclose when the values are illustrative rather than empirically calibrated. A strategy that depends on one optimistic assumption deserves further investigation. A strategy that survives several scenarios still has not demonstrated live profitability. Sensitivity analysis answers whether a conclusion is fragile under the specified changes; it does not prove that those changes span every possible market condition. The backtesting guide provides context for historical rule evaluation and the need to preserve held-out evidence.
Do not fit the simulator after every disagreement
When a simulated outcome surprises you, changing the model may be appropriate, but it creates a new experiment. Version the assumptions and retain the original result. Otherwise, an evolving simulation can make yesterday’s conclusion appear more robust than it actually was.
Keep a change log that identifies the discrepancy, proposed explanation, evidence supporting the change, and next test. A new rule should explain a class of behavior rather than merely make one selected outcome look plausible. Evaluate it on separate cases after the change. Avoid adjusting the model until it matches the most favorable observed fills. That would replace validation with fitting. The same principle applies to signal thresholds, order timing, and cost assumptions: freeze the version under review, state when it changed, and distinguish evidence collected before and after the revision.
Exit prices need the same scrutiny as entries
A simulated entry is only part of the result. Exit assumptions can materially change the hypothetical outcome. Record what triggers an exit, what instruction follows, which quote is used, and how the model handles gaps or incomplete quantity. Do not treat a stop level as an automatic guaranteed fill price.
Investor.gov explains that a stop order becomes a market order when its stop is reached; execution price is therefore a separate question. Investor.gov order types. A historical bar that touches both stop and target also requires an explicit sequencing rule if you lack finer observations. Keep that ambiguity visible. The stop-loss guide discusses related planning concepts, but any research result still depends on the actual exit model used. A realistic-looking entry cannot compensate for an undisclosed exit assumption.
How to use the review with Tradewink
Start with the signal’s rationale and risk context, then record why you chose to follow or dismiss it on paper. Return to the original record when reviewing the outcome. This supports a decision-review habit without claiming a brokerage fill occurred.
Public subscriptions are paper-only; separately approved private beta accounts may submit live broker orders. A private paper-track decision does not create an order or broker position. If you also use an external brokerage simulator, keep its receipt and assumptions in a separate execution record. This article does not claim that Tradewink provides every field in the worksheet, calculates the benchmark used in your experiment, or exposes all simulator mechanics. Confirm what the actual interface and documentation provide. Treat missing evidence as a research limitation rather than filling the gap with an assumed product capability.
Limitations and evidence thresholds
No fixed number of paper sessions proves that a strategy is ready for a live market. The relevant evidence depends on the rule, instruments, operating conditions, execution model, and unresolved risks. A long favorable paper history can still omit the mechanism that matters most to a particular strategy.
This guide does not compare broker execution quality, prescribe a live transition, or report a tested edge. Its purpose is to make assumptions and receipts reviewable. Keep your conclusion specific: the workflow handled a partial fill; the model excluded queue position; the chosen sensitivity scenario changed the result; live execution remains unmeasured. These statements are more useful than calling a simulator realistic without defining the claim. Provider documentation was checked on October 4, 2026 and may change. Verify the current source before relying on a specific feature or rule in a later experiment.
Frequently Asked Questions
Why can paper trading results differ from live results?
A simulator applies a model with documented or unknown omissions. Live outcomes also depend on the actual instruction, market observations, timing, and available execution conditions.
Does a paper fill prove a live order would fill?
No. A simulated receipt establishes the model’s outcome, not the live market’s treatment of a real instruction.
How should I measure paper slippage?
Define the benchmark, timestamp, side convention, quantity, and fill record first. Label the result simulated and separate explicit modeled fees.
Do I need real trades to complete this worksheet?
No. Keep live execution unmeasured unless separately authorized authoritative records already exist. Do not place a live trade solely to complete a comparison.
Can more paper sessions eliminate execution uncertainty?
No. More observations can improve a defined test, but they do not remove a mechanism that the model omits or the records cannot measure.
Read next
Keep learning with a related guide before putting an idea on your watchlist.
Paper Trading App Workflow: Review Stock Signals
Learn to paper trade stock signals by reviewing entry, stop, target, and rationale, then recording and revisiting each decision.
How to Backtest Trading Strategies: A Practical Guide for 2026
Learn how to backtest trading strategies properly -- avoid common pitfalls like overfitting, survivorship bias, and look-ahead bias. Includes frameworks, metrics, and validation techniques.
Stop-Loss Strategies: 7 Methods to Protect Your Trading Capital
Learn the best stop-loss strategies for day trading and swing trading. From ATR-based stops to trailing stops, percentage stops, and AI-driven dynamic exits.
Ready to evaluate a signal?
Start free with a watchlist and inspect the context before you consider a broker connection.
Try AI signals on your watchlist
Send yourself a signal preview, then add tickers to see ranked entries, exits, and risk notes in Tradewink.
Tradewink builds explainable market research for self-directed traders. Build a watchlist, inspect signal reasoning and risk context, and paper-track ideas before you decide. Public subscriptions are paper-only; separately approved private beta accounts may submit live broker orders.
How this guide is reviewed
Tradewink reviews educational content against its documented market-data sources, risk controls, and product methodology. See our data sources and evaluation methodology for the evidence and limitations behind the platform.