Activity
Mon
Wed
Fri
Sun
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
What is this?
Less
More
1 contribution to ZeroOne · Your First AI Agent
How to check whether a trading strategy actually has an edge before you deploy it
Most of us can now build a working trading bot in an afternoon. What almost none of the tutorials give you is a way to answer the only question that matters: **does this strategy make money, or does it just run without crashing?** Those are different questions. Joseph made this split really clearly in his posts here — strategy evidence and operational readiness are independent, and a bot that never errors while trading a worthless strategy is a very efficient way to lose money. I want to show the missing half: how to get strategy evidence cheaply, in minutes, before anything touches an exchange. I'll use a real worked example, including the part where I found a bug in my own test. ### The trap that started this I deployed a bot on a 4-hour schedule and planned to paper trade for 72 hours to "see how it does." Then I did the arithmetic: - 72 hours ÷ 4-hour cron = **18 evaluations** - Historically, only **0.9%** of evaluations on that timeframe produced a trade - Expected trades in 72 hours: **0.16** I was going to wait three days to collect approximately zero trades, and then probably deploy anyway. Forward testing at low frequency is almost useless for judging a strategy. A backtest produced **2,581 trades in three minutes**. ### The one rule that makes a backtest mean anything **The backtest must run the exact same code the bot runs.** If you reimplement the strategy in your backtest, you are validating code that never trades, and trading code that was never validated. They drift, silently. The fix is simple: extract the decision logic into a shared module both import. ``` strategy.js <- indicators + entry/exit decisions, pure functions | +-- bot.js (live/paper: prints, places orders) +-- backtest.js (replays history) My backtest asserts indicator parity on every run and **refuses to print results** if the numbers diverge from what the live bot computes. That assertion is worth more than any statistic in the output. ## Bias controls you need or the numbers lie
@Joseph Manion Joseph — thank you for reading the actual code rather than the README. You were right, and it's worse than an edge case, so I want to be specific. I reproduced it before changing anything: `calcVWAP` read the wall clock, so on cached history it returned null, and the assertion treated null as a pass and left VWAP out of the throw. Every run printed a green VWAP tick without comparing anything, while the README said it refused to report on divergence. I then ran the old assertion against injected faults: it didn't throw for a null VWAP, and didn't throw for a VWAP that was 50% wrong. (The underlying math turned out to be fine — 250/250 sampled points agree to floating-point precision — so the check was empty, not the calculation. But that's not much comfort for a guarantee I'd advertised.) Fixed: VWAP takes an as-of time, null counts as failure, every indicator can throw, and there are now mutation tests that break each implementation on purpose and assert the harness notices. Also fixed from your list: mark-to-market drawdown including intrabar extremes, open-position handling at end of data, timestamp-based portfolio alignment, and dividend-adjusted ETF data. On the last one, a small precision: Yahoo's plain close is already split-adjusted, so it was dividends that were missing — ~3.5pp/yr for TLT, which changed my ETF numbers materially. Neither test universe had an actual date gap, so the alignment fix changed no published figure. Your hash suggestion is in too: strategy.js's sha256 is now printed by every backtest and stamped on every live decision log, so the shared-code claim is checkable without publishing the bot. Not addressed yet: close-based stops (matches a cron-driven bot, not a resting broker stop), lot sizes, and a formal walk-forward split — all listed in the README, along with a correction log.
0 likes • 5d
@Maryan Kondratyuk Thanks Maryan — walk-forward across regimes, benchmark comparison in the main engine and Monte Carlo drawdowns are all on the open list in the README. "Running code isn't evidence; surviving hostile tests is" is going in as the standard.
1-1 of 1
Elizabeth Wilderose
2
7 points to level up
@elizabeth-wilderose-6546
SheCEO, entrepreneur, dominatrix

Active 5d ago
Joined Sep 1, 2026
Powered by