Dan, the fact that the strategies fail after realistic spread and trading costs may actually mean your evaluation system is working correctly. Many common strategies appear profitable only before execution friction is included. Before searching for more strategies, I would separate three possibilities: 1. The strategy has no durable edge. 2. The edge exists, but the timeframe or instrument cannot support the trading frequency after costs. 3. The test contains a timing, pricing, spread, position-sizing, or look-ahead problem. I would examine gross expectancy versus total cost per trade, performance by market regime, trade frequency, profit concentration, maximum adverse excursion, and sensitivity to small parameter changes. Also compare results across instruments with different spreads and liquidity. I would resist lowering the test’s standard simply to produce a passing strategy. A self-improving system should be allowed to conclude that no strategy currently qualifies. That refusal is valuable evidence. The direction I’m taking with Threshold is similar, although its first job is supervised execution rather than autonomous strategy creation. It records evaluated opportunities, accepted and rejected setups, market conditions, strategy decisions, and outcomes. The longer-term research layer will use that evidence to recommend strategy changes, but it will not be permitted to rewrite or deploy live rules without separate validation and human approval. I’d be interested in seeing the exact requirements a strategy must pass in your evaluation cycle. That may tell us whether you have discovered a strategy problem, a market-selection problem, or an implementation problem.