Detecting Price Anomalies: Time-Series Outlier Detection in High-Frequency Market Data
How statistical filters and isolation forests clean volatility noise, bot scrapes, and platform glitch prices from downstream analytics.
High-frequency price scrapers capture millions of data points daily, but raw extraction data is subject to glitches: sellers setting placeholder prices (e.g., ₹99,999), temporary discount coupon glitches, and scraping parser drops. Unfiltered data ruins downstream pricing intelligence.
Comparison of Statistical Anomaly Techniques
| Methodology | Latency Profile | Ideal Use-Case | False Positive Rate |
|---|---|---|---|
| Rolling Z-Score / IQR | Sub-millisecond | Stationary catalog pricing with steady baseline history | High during seasonal discount events |
| Seasonal ESD (S-ESD) | 5–10ms per SKU | Detecting flash anomalies during regular diurnal promotional cycles | Low across cyclical FMCG retail |
| Isolation Forests | 15–25ms batch | Multi-feature outliers (combining price, seller rating, and stock level) | Extremely low across diverse marketplaces |
Operational Implementation
- Two-Pass Pipeline: Applying real-time heuristic validation at the ingestion layer followed by hourly unsupervised machine learning model clustering.
- Contextual Thresholding: Dynamically relaxing anomaly barriers during known festive windows (e.g., Diwali or Big Billion Days) to capture authentic price cuts.
Ensure 100% Reliable Retail Intelligence
Integrate our automated anomaly detection and data cleansing pipelines to power your downstream BI with confidence.
