← Dashboard

OmmAlpha · ML Lab

Phase 1 · leakage-safe dataset foundation for XGBoost

Dataset readiness · BACKTEST

7,386Historical signals
415Stocks represented
2,500Rows usable for first success model
2,500Rows usable for false-breakout model
2,512Rows with 20-session return
99.9%Historical RS coverage
Strong starting dataset
The first XGBoost classifier uses only resolved triggered outcomes: TARGET = 1, STOP/TIMEOUT = 0. PENDING, AMBIGUOUS and NOT_TRIGGERED are excluded.
NIFTY 500 historical breadth table detected. Breadth columns are exported with a membership-quality flag. Approximate historical membership is not silently treated as pristine data.

Outcome inventory

1,037TARGET
1,050STOP
413TIMEOUT
4,619NOT_TRIGGERED
49AMBIGUOUS
218PENDING

Time coverage

2024-10-09 → 2026-09-25. Training will be split chronologically, never by random row shuffle.

2024 · 1,419 2025 · 3,279 2026 · 2,688

What happens next

1. Complete Phase 1B historical RS / breadth enrichment 2. Export ML Dataset v1B CSV 3. Train Logistic Regression baseline 4. Train XGBoost classifier on the same chronological split 5. Compare ROC-AUC, PR-AUC, Brier score, precision and top-decile performance 6. Save model + feature importance 7. Only if XGBoost beats the baseline → integrate predictions into OmmAlpha