Stock Markets Analytics Zoomcamp 2026

Homework 1: Intro and Data Sources Statistics

Distribution of scores and reported study time for this homework.

Submissions

148

Median total score

10

Average total score

9

Score distribution

All values are points.

Questions score

Min
-
Median
10.0
Max
10
Q1
8.0
Avg
8.5
Q3
10.0

Learning in public score

Min
-
Median
0.0
Max
3
Q1
0.0
Avg
0.3
Q3
0.0

Total score

Min
-
Median
10.0
Max
13
Q1
8.0
Avg
8.8
Q3
10.0

Time distribution

All values are hours reported by students.

Lectures

Min
1.0
Median
2.0
Max
30.0
Q1
2.0
Avg
3.6
Q3
3.0

Homework

Min
0.1
Median
3.0
Max
29.0
Q1
2.0
Avg
3.7
Q3
4.0

Question breakdown

Correctness and answer distribution per question.

1. Which year had the highest number of additions (starting from 2020)?

137 / 148 correct (92.6%)

1 2025 137 (92.6%)
2 2024 7 (4.7%)
3 2023 3 (2.0%)
4 2022 1 (0.7%)

2. How many indexes (out of 10) have better year-to-date returns than the US (S&P 500) as of August 21, 2026?

132 / 148 correct (89.2%)

1 1 4 (2.7%)
2 2 132 (89.2%)
3 3 7 (4.7%)
4 4 4 (2.7%)

3. Median drawdown (in %) of significant market corrections in the S&P 500 index

129 / 148 correct (87.2%)

1 8 129 (87.2%)
2 16 9 (6.1%)
3 24 4 (2.7%)
4 32 3 (2.0%)

4. Calculate the median 2-day percentage change in stock prices following positive earnings surprise days.

103 / 148 correct (69.6%)

1 3.35 5 (3.4%)
2 2.35 18 (12.2%)
3 1.35 18 (12.2%)
4 0.35 103 (69.6%)

5. Idea for your capstone project

0 / 148 correct (0.0%)

Answer Count
For my capstone project, I would like to build a machine learning model that predicts short-term price movements of major U.S. technology and telecommunications stocks. I plan to combine historical price and volume data, technical indicators such as RSI and MACD, and macroeconomic indicators such as interest rates and inflation. The model would predict whether a stock is likely to have a positive or negative return over the next 5–30 trading days. I would also like to compare different machine learning models and evaluate whether macroeconomic and technical indicators improve prediction accuracy. 1
Predict 1-month relative returns for large US technology stocks. 1
Intermarket & Market Breadth Dashboard (ML-Driven Market Regime Detection & Downside Risk Timing): Multi-layer system integrating leading indices (SPY, QQQ, DIA, SMH), market breadth (% > 20/50/100/200 EMA, RSP/SPY), macro gauges (^TNX, DXY, VIX/VVIX), and FBMA ribbons with HMM / XGBoost regime classification and Traffic Light (Green / Yellow / Red) tactical allocation. 1
I would like to build an earnings reaction prediction project for S&P 500 companies. The goal would be to predict whether a company’s stock price will increase or decrease during the two trading days after an earnings announcement. I would use data such as earnings surprise percentage, trading volume, recent stock returns, volatility, company sector, and overall market performance. I would collect and clean the data using Python, store it in SQL, and compare models such as logistic regression and random forest. 1
Build a machine learning model to predict short-term commodity and stock price movements using historical prices, technical indicators, earnings surprises, and financial news sentiment to generate trading signals and compare ML models such as XGBoost and LSTM 1
i am thinking to AI agent to manage a porfolio 1
Not quite defined yet, but I'd like to pick a few assets to invest in in the shorter term as "moonshots" to give my current "boring" portfolio a higher potential upside. 1
I want to build a short-term recommendation system for the Mexican REITs (FIBRAS), focusing on a subset of FIBRAS in the industrial segment. I plan to use macro indicators and technical indicators from financial reports. 1
I want to study how European stock markets react to major geopolitical shocks — the COVID-19 crash (2020), the Russia-Ukraine war (2022-), and the Israel-Iran-US escalations (June 2025 and the 2026 escalation) — using company-level equities from the Euro Stoxx 50 and STOXX Europe 600, with sector granularity, complemented by oil & gas commodities as a secondary asset class (fixed income deferred for a later iteration). Within the equity universe, I'll focus on **Defense/Aerospace** as the main industry vertical — the sector that historically benefits from rising military spending during conflict (e.g. Rheinmetall, BAE Systems, Thales, Leonardo, Saab) — and, scope permitting, add **Energy** and **Travel & Leisure** as secondary verticals to capture contrasting exposures (commodity-driven upside vs. conflict-driven disruption). The investment strategy is **event-driven sector rotation**: using dated geopolitical escalations as triggers, I'll measure the differential/abnormal return reaction across Defense vs. Energy vs. Travel & Leisure in the days/weeks following each shock, using oil & gas prices as a complementary macro context layer. Longer term, this could extend into an ML model that flags geopolitical shock windows and predicts the direction of sector-level abnormal returns, potentially using a geopolitical risk index (e.g. the Caldara & Iacoviello GPR Index) as a feature. 1
My capstone idea is an end-to-end short-horizon momentum + earnings-drift strategy for US mega-cap tech stocks, delivered as a Power BI analytics product. I will rank the ~50 largest US-listed tech companies by blended momentum (closing price vs. 20/50/200-day moving averages) combined with post-earnings announcement drift (PEAD): after each earnings event I compute the surprise magnitude and the 2-day reaction (building on the AMZN event study from Q4), and use both as features. A gradient-boosted classifier (LightGBM) will predict the probability of a positive 10-day forward return; I will backtest a top-quintile monthly-rebalanced portfolio (Sharpe, max drawdown, hit rate) over 2015–2026. All outputs — daily prices, earnings calendars, surprise events, portfolio states — will be modeled into a star schema in Power BI, with DAX measures for momentum scores, drawdown, Sharpe, and median post-earnings drift, and an interactive dashboard (portfolio curve, event-study heatmap, per-stock drill-through). This doubles as my PL-300 (Power BI Data Analyst) hands-on project: data modeling, DAX, and visualization are exactly the certification domains, and the auto-refresh of the semantic model keeps the report live. The project combines technical analysis, earnings fundamentals, and ML prediction into one reproducible pipeline with a business-grade BI front end. 1
I'd like to build an early-warning model for S&P 500 corrections — predicting the odds of a 5%+ drawdown starting in the next 30 days. My initial idea is to use macro features like the yield curve spread (DGS5 - DGS1), core CPI inflation, and recent market volatility, and train a simple classifier on the correction history I built in Question 3. Still rough, but I think the yield curve inversion signal could be a strong starting point. 1
I want to investigate market regimes in cryptoassets and build a model to predict the direction and risk of future Bitcoin and Ethereum returns (possibly extending later to Solana). The model would estimate whether returns over a given horizon (starting at 7 days, later expanding to 30) will be positive or negative, with probabilities that can be converted into buy/sell/hold signals, plus risk metrics such as volatility and drawdown. The approach has two stages. First, an unsupervised pipeline identifies market regimes: engineered features (returns, volatility, volume, and later macro data from FRED) are standardized, reduced via PCA, and clustered with KMeans into discrete regimes. A Markov chain then models how these regimes persist and transition over time, not to predict prices directly, but to describe regime dynamics. Second, supervised models use these market indicators plus the current regime (and its transition probabilities) to predict direction and risk. All steps are evaluated chronologically (training strictly before validation/test) to avoid data leakage. Data starts with daily OHLCV from Yahoo Finance, later expanding to CoinGecko/Binance and macro variables from FRED. As a stretch goal, I'd like to add an LLM/RAG assistant to explain predictions and retrieve similar historical periods, potentially feeding a Telegram alert system. The project will be built incrementally across the course modules. 1
This is an initial project idea, and I may refine or change the direction of the project as I progress through the course and learn more about financial data and market analysis, So. I would like to build a machine learning project for short-term stock market analysis, focusing on major US stocks and the S&P 500. My goal is to identify different market regimes, such as bullish, bearish, and high-volatility periods, using historical market data and clustering techniques. I would then combine these market regimes with technical indicators such as RSI, MACD, moving averages, returns, and trading volume, along with global news sentiment. The goal would be to predict the short-term direction of selected stocks over the next 1–5 trading days and evaluate whether adding market-regime and news information improves the predictions compared with using technical indicators alone. 1
I want to build a capstone project centered on the UK stock market (FTSE 100) in comparison with the US (S&P 500), using Brexit as the central event anchor rather than treating it as a background fact. The core question is: how did Brexit change the behavior of the UK market, viewed through the lens of market corrections and earnings-surprise reactions, relative to the US over the same period? Specifically, the project would extend the methodology from this homework in two connected directions: Corrections around Brexit (extending Q3's approach): identify whether the 2016 referendum and the final 2020 exit triggered measurable corrections in the FTSE 100, and compare their depth and duration against both FTSE 100's own "typical" corrections and S&P 500 corrections over the same window — to see whether Brexit-driven drawdowns were statistically distinct from ordinary market volatility, or whether the effect was overstated in media coverage. Earnings surprise reaction before vs. after Brexit (extending Q4's approach): using a handful of large FTSE 100 companies, test whether the market's reaction to earnings surprises (the correlation between surprise magnitude and 2-day return) changed after Brexit — i.e., did the UK market become more or less predictable/rational in how it prices corporate news once the political uncertainty resolved? As an optional third layer, if time allows, I'd also like to check whether the index inclusion effect (from Q1) weakened for FTSE 100 additions post-Brexit, potentially reflecting reduced international index-fund flows into UK equities. The unifying thread is Brexit as a natural experiment: rather than comparing UK and US markets in the abstract, the project uses a single, precisely dated political event to test whether — and how — it measurably reshaped market behavior in the UK relative to a US benchmark that experienced no equivalent shock. 1
I am interested in building a short-term prediction model for large U.S. healthcare stocks. The goal would be to predict whether an individual healthcare stock will outperform the healthcare sector benchmark (XLV) over the following 5 trading days. I plan to combine historical stock prices with technical indicators, market conditions, macroeconomic indicators, and earnings-related information. Potential features include recent returns, moving averages, RSI, MACD, volatility, trading volume, S&P 500 and healthcare-sector returns, VIX, interest rates, and earnings surprise metrics. I would initially compare models such as Decision Tree and Random Forest and potentially experiment with models such as XGBoost later. The predictions would then be used to construct a weekly trading strategy that selects the stocks with the highest predicted probability of outperforming the sector. I would backtest the strategy and compare its return, Sharpe ratio, and drawdown against benchmarks such as XLV and the S&P 500. 1
I would like to build a Supply Chain Disruption Early-Warning and Risk Intelligence System. The goal is to detect potential supply chain disruptions before they occur by combining time-series forecasting, anomaly detection, and risk modeling. The system would analyze data such as demand, inventory levels, supplier performance, lead times, shipment delays, logistics activity, weather events, commodity prices, and other external signals. I plan to develop models for demand and lead-time forecasting, anomaly detection, and disruption-risk prediction using machine-learning and deep-learning approaches such as SVR, LSTM, and other suitable models. The final system would generate a disruption risk score, provide an early warning when abnormal conditions are detected, and identify the main factors contributing to the risk. I would evaluate the system using time-aware and walk-forward validation to avoid data leakage and measure not only prediction accuracy but also the quality and timeliness of disruption warnings. The project would be developed as an end-to-end ML system, including data ingestion, feature engineering, model training, evaluation, monitoring, and a dashboard or API for visualizing supply-chain risks and alerts. 1
ai analyze public data for fraud. 1
I want to build a dual-market 5–20 day prediction and rotation model for Germany’s DAX 40 vs Nepal’s NEPSE Top-15 (hydropower/banks), forecasting return terciles to overweight the stronger market hedged for EUR/NPR for remittance timing from Berlin. I’ll use RSI/MACD/Bollinger/volatility/illiquidity plus Nepal remittance, monsoon/hydro seasonality, ZEW and FinBERT news sentiment, trained as LightGBM with a 20-day cost-aware backtest (2018-26 split). 1
Modeling Of Stock Market using Markov Equations 1
I want to build a 15-day trend classification model (predicting Up or Down) for emerging markets, specifically Brazil and India. I will focus strictly on quantitative data, combining technical indicators (MACD, RSI, Moving Averages), trading volume, and correlations with global indexes. My goal is to build an end-to-end machine learning pipeline using XGBoost to see if global market movements can predict emerging market trends. 1
Modelo de momentum/reversal a corto plazo en megacaps tecnológicas 1
Determine top 10 cryptocurrencies by lowest reaction time on global economic indicators 1
Indian market 1
I want to build a short-term prediction model for the London stock markets, focusing on the largest stocks over a 30-day investment horizon. I plan to use RSI and MACD technical indicators and news coverage data to generate predictions. 1
I want to build an automated end-to-end quantitative pipeline that uses machine learning to detect market volatility regimes, dynamically rotates capital between top-performing sector ETFs and defensive assets to minimize drawdowns, and automatically sends weekly portfolio rebalancing alerts logged to an SQLite database. 1
Build a full-stack, automated quantitative trading system specifically tailored for the Stock Exchange of Thailand (SET). Designing an end-to-end pipeline that handles everything from initial data ingestion to advanced risk management 1
Long-term investing into ETFs, focusing on risk assessment to minimize probability of selling while optimizing for returns by choosing right timing of purchase 1
I want to build an algorithmic equity prediction model for the US-STOCK MARKET, focusing on S&P 500 Large-cap Technology stocks across an investment horizon of 10 trading days.I plan to feed a Random Forest model with feature vectors combining rolling14-days RSI, MACD histograms and alternative sentiment trend lines from financial news scraping. 1
A pipeline that scrapes public job posting counts for a basket of companies (via their careers-page APIs), tracks month-over-month hiring velocity, and tests it as a leading feature for earnings-surprise direction ; an alt-data signal that requires real data engineering (scraping, scheduling, storage) rather than just pulling a pre-packaged price series. 1
Market: Philippine equities. I want to evaluate the initial reaction to a company disclosure, and build a model that predict the potential influence on price in the next 5 to 20 days. I will use RVOL and ROC to quantify the market's initial reaction. I'm thinking both RVOL and ROC should be +/- 1.5x its standard deviation for the disclosure to qualify as being "material". What I want to see: Does +/- ROC&RVOL → +/- 5DChg, 20DChg? Or inverse to no relation? Also, nuances like shakeouts after the initial reaction. Can I use the initial reaction to shape my bias for the short- to medium-term? Things I will need: 1. OHLCV market data, 2. Disclosures — filing dates, subject, content, fundamentals. 1
I plan to build a personal quantitative analysis and portfolio monitoring system for high-volatility assets (including cryptocurrencies like BTC, SOL, and LTC). The system will initially compute rolling asset correlation matrices, beta relative to market benchmarks, and key risk metrics (VaR and Sharpe ratio). Building on this foundation, I will implement a Machine Learning pipeline using ensemble models (Random Forest, XGBoost) and technical indicator features (EMA, RSI, MACD) to predict short-term price direction and improve strategy entry points over a 2-to-5-day horizon. 1
Use valatility to inform trading decisions and potentially identify market corrections. 1
Analyse historical stock-price data to examine returns, volatility, trading volume and performance relative to a benchmark such as the S&P 500. 1
Project: Regime-Aware US–Malaysia Equity Ranking and Portfolio Dashboard. I want to build an end-to-end system that ranks liquid US and Malaysian stocks by their expected risk-adjusted return over the next 20 trading days while separately estimating downside risk. 1
Create a poc of an automated trading system. 1
: I want to build a system that identifies undervalued stocks using financial and market data. I would compare metrics such as earnings, revenue growth, P/E ratios, and stock performance. The goal would be to rank companies based on their potential investment value. 1
I would like to build a machine learning model to predict short-term stock market movements in the Indian stock market, focusing on Nifty 50 companies. I plan to use historical prices, trading volume, RSI, MACD, moving averages, earnings surprises, and news sentiment as input features. I will compare models such as Random Forest and XGBoost to predict whether a stock's price will increase or decrease over the next 5–20 trading days. The goal is to identify useful patterns in financial data and evaluate whether these features can improve short-term market predictions. 1
I aim to develop a short-term predictive modeling pipeline for Indian equities, specifically targeting Nifty 50 constituent stocks over a 30-day holding period. The strategy will integrate quantitative technical signals—specifically Relative Strength Index (RSI) and Moving Average Convergence Divergence (MACD)—with NLP-derived market news sentiment scores to forecast price trends. Finally, the model's output signals, feature metrics, and technical visualizations will be deployed to an interactive Streamlit dashboard for real-time monitoring. 1
High-Frequency/Medium-Horizon Tech Sector Momentum Predictor 1
I would like to build a short-term day-trading prediction model focused on major U.S. stocks and ETFs such as SPY, QQQ, NVDA, and TSLA. The goal would be to predict short-term price movements using 1-minute or 5-minute market data. I would use technical indicators such as VWAP, moving averages, RSI, MACD, volume, volatility, and price action, along with candlestick patterns such as hammers, shooting stars, and hanging man patterns. I would also explore whether the model performs differently in trending and ranging markets. Finally, I would backtest a simple day-trading strategy based on the model's predictions and compare its results with a buy-and-hold approach. 1
Quantitative trading system for the Indonesian stock market (IDX), focused on IDX80 constituents for swing/position trading. Combines momentum screening, fundamental filtering, and an XGBoost model with TimeSeriesSplit CV, plus an IHSG market regime filter and realistic backtesting with slippage/fees. 1
Predict whether KONE (Nasdaq Helsinki) stock will gain or lose value over a 5-day horizon. Backtest long-only trading strategy against a benchmark Buy & Hold index allocation. 1
I want to build a stock market report focused on the Nigerian stock exchange that gives real-time info on the best stocks to trade within a certain cycle period. 1
I want a short-term model for the US S&P 500 (and a few large names like AMZN) about 30 trading days after a dip. I would start from the 5% corrections in Q3 (“buy the dip”), plus Q4-style earnings surprises, and Fed funds / core CPI from the lesson notebook. Not sure about the model yet. 1
Multi-Timeframe Machine Learning Bot for Short-Term Commodity Trading (XAU/USD & XAG/USD) Executive Summary & Core Objective I aim to design and back test an automated algorithmic trading model optimized for high-liquidity precious metals (Gold / XAU/USD and Silver / XAG/USD). The system will focus on capturing short-term intraday and swing price momentum across 1-minute to 1-hour timeframes. The goal is to move beyond simple heuristic rules by training machine learning classifiers to evaluate price action patterns (such as opening range breakouts and engulfing candlestick structures) in conjunction with dynamic technical indicators, filtering out low-probability trade signals while dynamically sizing risk. Asset Class & Target Market Asset Class: Spot Commodities / Commodities Futures Primary Tickers: Gold (XAU/USD) and Silver (XAG/USD) Execution Environment: Meta Trader 5 (MT5) via Python Integration / Trading View Pine Script for rapid visualization Holding Horizon: Scalping to short-term intraday holding periods (ranging from minutes to a few hours) Data Sources & Feature Engineering Market Data (OHLCV): High-frequency tick and minute-level data extracted via the Meta Trader 5 Python API (MetaTrader5 package) or financial data providers (e.g., Yahoo Finance / Alpha Vantage). Technical Indicators: Momentum & Trend: Exponential Moving Averages (EMA 9, 20, 200) tailored to lower timeframes, Relative Strength Index (RSI 14). Volatility & Range: Average True Range (ATR) for dynamic Stop-Loss and Take-Profit calculations. Pattern Recognition & Structural Features: Automated detection of key price structures: Bullish & Bearish Engulfing Candlesticks mapped with exact open/close boundary conditions. Opening Range Breakout (ORB) high/low levels established during key market session opens (London and New York sessions). Modeling Approach & Architecture Target Variable: Binary classification predicting directional movement (Y∈{+1,-1}) over a defined forward horizon (e.g., next 5 to 15 bars), conditioned on reaching a target Risk-to-Reward ratio (e.g., 1:2 RR using ATR) before hitting the stop loss. Model Pipeline: Baseline / Benchmark: Logistic Regression and Decision Trees. Ensemble Models: XGBoost and LightGBM to capture non-linear relationships between multi-timeframe EMA distances, RSI levels, and breakout candle sizing. Signal Execution & Sizing: A risk management module that dynamically adjusts position size (e.g., fractional lot sizing starting at 0.01 lots per baseline unit of account equity) based on model confidence score and volatility. Performance Evaluation & Validation Walk-Forward Optimization: Time-series split cross-validation to eliminate lookahead bias Key Metrics: ML Metrics: Precision, Precision-Recall AUC (prioritizing high precision over high recall to reduce false entries). Financial Metrics: Net Return, Win Rate %, Profit Factor, Max Drawdown %, and Sharpe Ratio. Back testing Framework: Python (Back trader or custom vectorized back tester) followed by live paper-trading execution test in Meta Trader 5. 1
I have an MSc in AI but no finance foundations. I would like to focus on the US or UK market and build a model for predicting future stock prices of S&P or FTSE indexes. 1
Macro-aware sector rotation. Which U.S. sectors should perform best next month under current inflation, rate, and growth conditions? Hold the top 2–3 predicted sector ETFs 1
My long-term goal is to become an options trader, so I would like to build an options trading decision-support tool using machine learning and quantitative risk measures. The tool would help evaluate potential options trades by analyzing the Greeks, including Delta, Gamma, Theta, Vega, and Rho, together with implied volatility, historical volatility, option price, volume, open interest, and the underlying asset's price behavior. # The objective would be to develop a model that estimates the probability and potential magnitude of an option trade being profitable over a specific time horizon. 1
a swing trading system focused on SPY using Machine learning to define probability of up /down day on next open , input from past trading days, volume, VIX 1
Earnings momentum: a short-term strategy for large US tech stocks 1
Done in the corresponding notebook 1
I want to map where the next dollar of AI-infrastructure capital expenditure actually lands. The universe has four books: (A) 50 listed global builders and buyers outside mainland China; (B) a separate 20-name China hardware/cloud book; (C) listed power companies plus the ten largest dedicated generation projects; (D) the companies behind OpenRouter Value Leaders (MiniMax, GLM/Z.ai, Kimi/Moonshot, Qwen/Alibaba, Tencent, Xiaomi, DeepSeek). Public China model labs are tracked with HK tickers (0100.HK MiniMax, 2513.HK Z.AI, plus Alibaba, Tencent, Xiaomi). The investment question is which layer is the bottleneck this quarter (GPU, HBM, power, or the model lab) and whether that sleeve is invest, hold, or underweight. 1
I want to build bear/bull prediction model for the US stock market. I plan to run a model to identify the main factors that predict the market and then build a model to predict the future. 1
I want to build an automated trading strategy for stocks in the S&P 500, using historical stock market and economic data to predict short-term returns. 1
Global Equity Rotation ML Model 1
I want to build a retrieval-based analysis tool that predicts how the UK market (gilts, FTSE 100, GBP) will react to Bank of England interest rate decisions and MPC statements. The project has two connected stages: (1) a factor-attribution model that uses historical UK macro data (inflation, unemployment, wage growth, GDP) to explain and weight what drove each past BoE statement's hawkish/dovish tone, and how that weighting has shifted across different economic periods; (2) a retrieval component that finds historically similar past BoE statements (by language/tone) to the current one, and uses how markets reacted in those analogous cases to predict the likely reaction this time. This combines causal macro-factor analysis with an IR/retrieval-based market-reaction model — most existing public research and tools of this kind focus on the US Federal Reserve, so applying it to the UK with a factor-weighted retrieval approach is a fairly novel angle. Data sources: ONS for macro time series, Bank of England for MPC minutes/statements/speeches, and market data (gilt yields, FTSE, GBP) via yfinance or similar. 1
Capstone: build a 30-day return prediction model for liquid Brazilian equities, combining RSI, MACD, volatility, macroeconomic data, and news sentiment. I will benchmark the model against buy-and-hold using time-series validation and transaction-cost controls. 1
I want to build an early-warning model for US financial stocks, predicting significant price declines over the next 30 days. The model will combine market indicators, interest rates, yield-curve data, volatility, valuation metrics, and financial news sentiment. I will evaluate whether macroeconomic and news-based features improve predictions compared with a baseline using only historical prices and technical indicators. 1
Evaluation of historical data to predict future performace of stock in a dedicated market 1
Build a model predicting short‑term movements in Indian large‑cap stocks using sentiment from financial news + technical indicators (RSI, MACD). 1
I want to build a Multi-Chain Micro-Cap Crypto Quantitative Backtesting & Strategy Comparison Engine to systematically identify the most profitable intraday trading strategy for pump-and-dump and breakout dynamics in decentralized exchange (DEX) tokens. 1
I want to build a drawdown risk classifier for the S&P 500 index (^GSPC). For each trading day, the model predicts the probability that the index closes 5% or more below that day's close at some point within the next 21 trading days. On daily data from 1962 to the present, about 16% of days meet this condition, so the class imbalance is mild and the model can be trained on the true class balance, with probability outputs rather than resampling. Features are limited to information available on the prediction day: recent returns, realized volatility, distance from the 52-week high, and market stress indicators. Models are logistic regression as an interpretable baseline and RandomForest, evaluated with walk-forward splits and a 21-day gap to prevent label leakage, scored on precision-recall AUC against the 0.16 base rate, with Brier score and calibration checks. The strategy holds the S&P 500 and reduces exposure when predicted drawdown probability exceeds a threshold, compared against buy-and-hold on CAGR, Sharpe ratio, and maximum drawdown 1
I would like to explore short-term stock market movements using historical prices, technical indicators, and broader market conditions. 1
I want to build a short-term prediction model for the 15 stocks with the most volume going into the day that are in the S&P 500 & QQQ index. I would like to use the MACD, VWAP, and EMA indicators in order to generate predictions for day trades that I should enter. I would also like to use historical data to suggest the price levels I should enter the trades. 1
I want to build a short-term (5–10 day) prediction model for a small set of frontier/emerging-market equities — likely a mix of Nigerian (NGX), Kenyan (NSE), or other East African-listed stocks, or a broader Sub-Saharan Africa ETF if single-name liquidity/data is too thin. I'll combine standard technical indicators (RSI, MACD, moving averages) with an LLM-based news/event scoring layer reused from my LLM Zoomcamp RAG capstone — ingesting local financial news headlines, chunking and embedding them, and scoring sentiment or event relevance as an additional feature alongside price data. The goal is to see whether adding a lightweight retrieval/LLM signal improves on a pure technical-indicator baseline in a market that's less efficiently priced (and less covered by existing quant research) than US large-caps. 1
Capstone idea 1
I would like to build a stock price prediction project for the Kenyan and US stock markets. I plan to use historical stock prices and technical indicators to predict short-term stock price movements and identify potentially good investment opportunities. 1
I want to build a model that predicts the likelihood of the S&P 500 entering a correction (5%+ drawdown from its all-time high) over a 20–30 day horizon. I'll combine technical indicators (RSI, MACD, volatility) with macro signals (VIX, yield curve) and earnings surprise breadth across constituents. I'll also test whether corrections in the US lead or lag corrections in other markets (e.g., Hong Kong, Germany, Brazil) to explore a tactical diversification strategy. 1
I want to build an app which will the agent will act as an experienced quant researcher and help us to do the stock research using the latest news, basically using the latest workshop Ivan has showed and I want to build an app using some of the AI engineering concepts from Alexey's AI-dev-tool zoomcamp, and do more function on it, for example, the depth of the research and how i should define the success of the answer generated is what I will focus on 1
**My capstone project will develop a quantitative risk and asset allocation engine for an Argentine multi-asset portfolio (`YPF`, `ZS=F`, `BZ=F`) integrated with crypto-native risk factors (`BTC-USD`, `USDT-ARS`).** 1
I want to build a predictive classification model focusing on the US mid-cap technology sector to forecast whether a stock will outperform the Russell 2000 index over a 60-day investment horizon. Instead of purely relying on technical indicators like RSI or moving averages, I want to incorporate alternative data via NLP sentiment analysis. By processing the management discussion sections of quarterly SEC 10-Q filings and cross-referencing them with insider trading volumes, I hope to identify companies that are signaling quiet confidence before the broader market catches on. 1
My professional background includes trading soft commodity markets, dating back approximately 15 years to the early stages of my career. Building on that foundation, I am looking to develop a systematic trading strategy operating on a multi-week to longer-term horizon, designed to capture directional price movements in Raw Sugar No. 11 futures. Given the structural drivers of this market, the framework would need to integrate several key fundamental inputs: production-side news flow (crop conditions, harvest reports, and mill output data, particularly from Brazil and India), supply-demand balance dynamics, currency effects — principally BRL and INR, given their influence on export competitiveness and grower incentives — and crude oil price movements, owing to the well-established linkage between sugar and ethanol markets in Brazil's flex-fuel economy. 1
I want to develop a trading strategy benchmarking platform that backtests and compares strategies such as Buy & Hold, Moving Average Crossover, RSI, MACD, and Momentum using historical market data. The strategies are evaluated using performance metrics like returns, Sharpe ratio, volatility, and maximum drawdown, while also analyzing their behavior in bull, bear, and sideways markets. The goal is to identify the strengths, weaknesses, and best use cases of each strategy and create a framework for ranking them based on risk-adjusted performance. 1
I would like to develop a short-term stock market prediction model focusing on large-cap US technology stocks. The project will use historical stock prices, technical indicators such as RSI, MACD, moving averages, and trading volume, together with earnings surprise data. The goal is to predict short-term price movements following earnings announcements and evaluate whether these indicators can help identify potential investment opportunities. 1
I want to build a 20-trading-day post-earnings drift model for large-cap US stocks (S&P 500 names, starting with mega-cap tech and then the full index). 1
Bursa Malaysia ML screener. Extend a rules-based fundamental screener into a predictive model classifying next-quarter outperformers on Bursa Malaysia-listed equities, using fundamental + technical features. Less-efficient market, existing data pipeline, clear evaluation via forward returns. 1
A capstone project idea focused on building an allocation model for country index ETFs. 1
I would like to build a machine learning model for predicting short-term movements and risk in the Nigerian stock market, focusing on the NGX All-Share Index and selected large-cap Nigerian stocks. 1
For my capstone project, I would like to analyze the stock performance of my current company and compare it with relevant market and sector indexes. I am interested in using historical price data and financial indicators to explore whether the company is likely to grow or outperform its benchmark over the medium to long term. 1
Short-term directional model for the S&P 500 based on multi-horizon OHLC levels. I want to build a 1–10 day directional model for the S&P 500 (SPX index and SPY ETF) where the main features are not classic technical indicators, but the position of price relative to OHLC levels from higher timeframes: the open, high, low and close of the daily, weekly, monthly, quarterly and yearly bars. The hypothesis I want to test is that the normalised distance from price to those levels, and whether price touched them during the session, carry directional information beyond chance. My prior is negative: I expect most of the effect to disappear once I correct for volatility and for multiple comparisons. That is precisely the point of the project — putting a very widely used family of technical signals through a serious statistical test, which is rarely done. Design: features are ATR-normalised distances to each OHLC level plus touch rates, with a regime block (VIX, market breadth, yield curve slope). Models are a regularised logistic regression as baseline and gradient boosting as the alternative, always compared against controls (label shuffling, randomly generated fake levels, block bootstrap). Validation is walk-forward with a strict train/test split, Benjamini–Hochberg FDR correction, α = 0.01 and a minimum effect of interest of 3 percentage points, all fixed before looking at results. Evaluation is not only accuracy but net profitability after transaction costs and slippage, which is where this type of signal usually dies. The deliverable I care about is a pre-registered, reproducible research framework that separates confirmed findings from open hypotheses and purely descriptive observations. 1
I want to explore the Asunción Stock Exchange (BVA) and gain a thorough understanding of trading concepts and processes within the Paraguayan market, focusing on the most liquid instruments—such as corporate and government bonds and repo transactions—across short- and medium-term horizons. I plan to analyze trends in yields and trading volumes, as well as the impact of local macroeconomic factors (such as the Central Bank of Paraguay's monetary policy rate and data from the agricultural export sector), in order to develop analytical models that facilitate investment decision-making. 1
"I want to build a short-term (1–2 week) return prediction model for large-cap US tech stocks (AAPL, MSFT, GOOGL, AMZN, NVDA, META). I will combine technical indicators (RSI, MACD, Bollinger Bands, 20-day momentum), earnings surprise magnitude, VIX levels, and sector ETF flows as features. I'll train an XGBoost classifier to predict the direction of 5-day forward returns and evaluate via walk-forward validation, then compare against a simple buy-and-hold benchmark." 1
I want to build a short-to-medium term (30-day) directional prediction model for the S&P 500 Information Technology sector, focusing on the top 20 largest stocks by market cap. My goal is to predict whether the stock will outperform the sector ETF over the next month. I plan to use a combination of traditional technical indicators, fundamental sentiment derived from last two earnings call transcripts, and macroeconomic regime indicators. Target variable will be binnary classification, if it will or not. 1
Capture short-to-medium-term post-earnings announcement drift (PEAD) or mean reversion in high-liquidity US tech stocks (e.g., Apple, Microsoft, Nvidia) over a 5-day to 15-day investment horizon. 1
I'm going to build ranking model for the global gambling/betting sector over a 60day horizon — ranking stocks daily by return relative to the sector median rather than predicting absolute returns, combining price-based technicals with iGaming-specific fundamentals drawn from my professional background 1
A Cross-Sectional Return-Ranking Model for the Mexican Equity Market (Reto Actinver Universitario) 1
Short-term prediction model for the Korean stock market (KOSPI), focusing on large-cap exporters like Samsung Electronics and SK Hynix, over a horizon of a few weeks — using price/technical indicators (moving averages, RSI) plus the USD/KRW exchange rate for macro context. 1
I’d like to look at semiconductor supply chain bottlenecks. My goal is to forecast monthly inventory-to-sales ratios and stock volatility for chipmakers (like TSM, NVDA, and ASML) by combining lead-time tracking data, global shipping rates, and quarterly capex figures with XGBoost to see how upstream delays impact downstream stock dips. 1
I want to build a machine learning framework to evaluate short-term price momentum following structural macro index rebalancing events in the US and Indian equity markets. I plan to use historical S&P 500 and Nifty 50 constituent addition/deletion dates as my event windows. The target variables will track 5-day, 14-day, and 30-day post-announcement investment horizons. Features will combine technical indicators (RSI, MACD, and Bollinger Band deviations) along with daily volume spikes to predict optimal exit windows for event-driven trading strategies. 1
trading bot (ideally) or at least decision driver 1
Sector Liquidity Rotation & Institutional Flow Strategy for the Vietnam Stock Market (VN-Index / VN30). Focuses on the top 50 liquid stocks on HOSE across core sectors (Banking, Securities, Real Estate, Steel) with a 10-15 trading-day holding horizon matching Vietnam's T+2.5 settlement cycle. Combines sector liquidity shifts (% turnover of total HOSE) with 5-day cumulative Foreign Net Trading and Domestic Proprietary Trading (Khối ngoại & Tự doanh). Uses momentum filters (Price > 20-day SMA, 14-day RSI in 50-65, volume breakout) and a LightGBM classifier to predict VN-Index outperformance probability. 1
Create an Elliot Wave tool for analysing a stock's performance. 1
Analyzing the impact of earnings surprises on stock price movements and predicting post-earnings returns using machine learning. 1
I want to build a machine learning-driven asset allocation model for US Equities that strictly avoids overfitting (false positives) using Walk-Forward Optimization and Purged Cross-Validation. 1
I want to build a machine learning system for short-term stock market prediction, focusing on large-cap U.S. technology companies such as Apple, Microsoft, Amazon, Nvidia and Google. The model will predict whether the stock price will increase over the next 5 trading days. I plan to use technical indicators such as RSI, MACD, moving averages and volatility, together with market variables such as the S&P 500, VIX and earnings surprises. I will backtest the predictions against a buy-and-hold strategy and evaluate the model using accuracy, precision, cumulative return, Sharpe ratio and maximum drawdown. Finally, I would like to present the results in an interactive Streamlit dashboard 1
I propose developing an algorithmic trading strategy that capitalizes on "on-chain surprises" for top cryptocurrencies, serving as a direct analog to traditional earnings beats. 1
Build an AI-powered financial market analysis tool that combines stock price movements, earnings surprises, market corrections, and other financial indicators to identify patterns and generate insights about market performance. 1
I'd like to build a prediction model for Thailand stock market. 1
[not yet decide] I want to build short-term classification model to predict whether the daily closing price of top 5 US stock (AAPL, MSFT, NVDA, GOOGL, AMZN) will close higher or lower on following trading day. I plan to use scikit learn tot rain a random forest classifier on past 5-10 years via yfinance API. So the result would be as a binary classification (up/down) instead predicting the exact numerical price target 1
My goal will be to create a short-term earnings reaction prediction model for the largest U.S. tech firms (for example, AMZN, GOOGL, MSFT, AAPL) where I try to predict both the direction and size of the earnings reaction in the 2-3 days following the announcement.The target variable will be either a classification (positive, negative or neutral move in two days) or a regression model on the two day returns. 1
Predicting short-term stock returns using financial and market indicators 1
I would like to build an ML-driven crypto trading bot that predicts short-term price moves for coins using OHLCV data, technical indicators, on-chain metrics, sentiment, and order-book features, then backtests strategies with risk management and eventually paper-trades before going live. 1
I want to build a volatility forecasting system for the US market, focusing on the S&P 500 and a handful of large-cap stocks over a weekly horizon. I'll target realized volatility rather than returns. I plan to use classical time-series models (ARIMA/SARIMA and GARCH) as the core, benchmarked against a naive random-walk baseline and compared to a machine-learning model, with walk-forward validation and no data leakage. As an application, I'll turn the forecast into a vol-targeting position-sizing overlay evaluated against SPY buy-and-hold, using yfinance price data plus a few FRED macro indicators, with ntfy alerts when forecasted volatility spikes. 1
I'll extend my startup VeriHub (auditing how public LLMs represent organisations) into a domain with objective, free ground truth: verifiable facts about S&P 500 companies. I'll ask several public LLMs factual questions per ticker (latest dividend, P/E, last earnings date, 52-week high, index membership) in English, French and Dutch, and score each answer against Yahoo Finance / FRED / Wikipedia. My ML target is not price direction but a binary classifier predicting when an LLM answer is wrong, using features like the stock's volatility, fact recency (days since last earnings or index change), ticker popularity (market cap, volume) and query language. 1
I want to build a real-time data engineering pipeline and machine learning system to predict 24-hour volatility regimes for high-liquidity cryptocurrencies (BTC, ETH, SOL). The system will stream live trade data, order book depth, and derivatives metrics (Futures Funding Rates, Open Interest) via WebSockets into Apache Kafka. A stream processing engine will compute rolling technical features (Parkinson Volatility, ATR) and store them in a time-series database. An integrated ML service will classify market regimes (High vs. Low Volatility) to dynamically trigger risk-hedging signals, serving results onto an automated monitoring dashboard. 1
I am not sure yet, probably the CAPE / Burger Costs aggregation project with some funny metrics to correlate with stock market 1
ML for Market Microstructure & Liquidity Prediction — Develop a quantitative research platform that uses limit order book and trade data to model short-horizon liquidity, order-flow dynamics, and price impact. Compare classical market microstructure models with machine learning approaches such as XGBoost, temporal CNNs, and Transformers, and evaluate their performance through rigorous backtesting, risk analysis, and realistic transaction-cost simulations. 1
develop a trading strategy that beats "buy and hold" for Gold market 1
I aim to build a volatility regime detection model for WTI crude oil and gold futures—using a 30–60 day timeframe—that employs unsupervised clustering (GMM/HMM) on technical features and GARCH-based volatility, combined with sentiment scores extracted from X and commodity-focused Telegram channels, to improve volatility prediction accuracy compared to a standalone GARCH model. 1
I have recently worked on predicting the financial risk of the company. For the sake of example, I recently used multi modal data like time series data, news, financial ratios and candle stick charts to predict the financial risk of the company. Now I want to use my skills in the portfolio management. So yeah! I will be finding any idea and will be working on portfolio and asset management 1
Build an ML model to predict significant post-earnings price movements in US stocks using market, earnings, volatility, and sentiment data. 1
I would like to build a machine learning model for predicting residential house prices in Sweden. The project would focus on predicting house price changes over a 3–12 month horizon, using historical transaction data together with economic and housing-market indicators such as interest rates, inflation, unemployment, location, property characteristics, and regional housing trends. I would like to investigate whether machine learning models can improve house price predictions compared with simpler statistical approaches, and identify which factors have the strongest influence on prices. If sufficient data is available, I would also like to compare predictions across different Swedish regions, such as Stockholm, Gothenburg, and Malmö. 1
I want to build a simple model that predicts whether to buy, hold, or avoid a selected company’s stock over the next month using its recent price trend, trading volume, volatility, and earnings results. 1
Indian equities, Nifty Midcap 150 for Indian mid-cap stocks, parameters to consider: Bull markets. Recessions. High- and low-interest-rate periods. Crises and recoveries. Periods of changing regulation or market structure. 1
Earnings-Driven Momentum Strategy for US Large-Cap Tech Stocks 1
I would like to explore a project focused on financial data analysis, specifically using the US30 index. The main idea is to apply basic machine learning techniques to analyze historical price and volume data. I want to investigate if a model can learn to identify certain behavioral patterns using common indicators, such as moving averages. The result would be a simple tool that evaluates this data and generates useful summaries or statistics about how the index is moving. 1
i want to create ML system to predict the exact percent change of fed interest rate at the comming meetings 1
I want to build a weekly market direction prediction model for the S&P 500, predicting whether the index will close higher or lower 5 trading days from today. Data: S&P 500 daily OHLCV history from 1990 to present via yfinance Features I plan to use: RSI (14-day) — momentum / overbought-oversold signal MACD and MACD signal line crossover — trend change detection 50-day and 200-day moving averages — short and long-term trend Weekly return of the past 1, 2, and 4 weeks — recent momentum VIX level — market fear/uncertainty Model: Binary classification using Random Forest (up/down label), with a strict train/test time split to avoid data leakage. I will evaluate using accuracy and a simple backtest comparing cumulative returns vs a passive buy-and-hold strategy, measuring Sharpe Ratio. 1
AI-powered financial report analyst for Indonesian stocks Build a system that reads annual reports, quarterly reports, and corporate disclosures, then extracts financial metrics and explains what changed. Automatically extract revenue, margins, debt, cash flow, and segment performance. Compare a company's latest report with previous periods. Detect unusual changes, such as revenue growing while operating cash flow declines. Build a graph of companies, major shareholders, subsidiaries, business groups, and related entities to study ownership concentration and corporate relationships. 1
I would build a cross-market equity risk dashboard and 30-day return model for large US and Poland companies. It would combine price/volume features (momentum, RSI, MACD, volatility, drawdown), index-level regime features, earnings surprises, and news sentiment. The output would be calibrated probabilities of positive returns and a paper-trading portfolio with transaction-cost and drawdown reporting. 1
I want to build an end-to-end machine learning pipeline predicting short-term stock performance specifically for the SaaS and EdTech sectors. I plan to use Scikit-Learn and TensorFlow to train the prediction models based on technical indicators (RSI, MACD) alongside fundamental data. Given my background in software architecture, I intend to track the model's training iterations and hyperparameter tuning using MLflow. Ultimately, I would like to serve these predictions through a full-stack dashboard, building a REST API with Node.js or Spring Boot and designing the frontend interface in React. 1
I want to build a short-term prediction model for pure-play and enterprise Quantum Computing stocks (such as IONQ, RGTI, QBTS, IBM, and the QTUM ETF constituents) over a 14-day investment horizon. I plan to combine technical indicators with sentiment scores extracted from quantum technology news releases and patent filing announcements to forecast stock price volatility and post-earnings drift. 1
I do not have yet a very well defined topic for my project. I work as a Packaging buyer. Most my suppliers are huge corporations. There have been major "M&A" in the Packaging industry in the last decade and the trend is increase. So this is a great project topic. A potential subject would be "Industry Consolidation and M&A Value Creation in the Global Packaging Sector". However I still need to explore a potential project topic 1
Global Stock Market Performance Dashboard: Build a dashboard comparing major global stock indexes and individual stocks using returns, volatility, drawdowns, correlations, and other risk-adjusted performance metrics. The goal is to identify which markets and sectors have the best risk/return characteristics over different time periods. 1
I want to prepare an analysis on commodities to understand the relation between certain news and the changes in the market 1
Construir un modelo de Machine Learning (ej. Random Forest / XGBoost) para predecir la dirección y magnitud del movimiento de precios a 3-5 días en acciones del S&P 500 tras la publicación de resultados trimestrales, utilizando variables como la sorpresa en BPA (EPS surprise), el volumen previo al anuncio y ratios de volatilidad histórica. 1
I want to build a production-oriented ML platform for cross-sectional stock return prediction and portfolio ranking. The system will ingest historical market prices, earnings surprises, macroeconomic indicators, and potentially financial-news sentiment for a universe such as the S&P 500. I plan to engineer momentum, volatility, drawdown, technical, earnings, and market-regime features and train time-series models such as XGBoost/LightGBM to predict relative stock performance over a 20-trading-day horizon. I will evaluate the predictions through a realistic backtesting framework using metrics such as Sharpe ratio, maximum drawdown, CAGR, turnover, and transaction costs rather than prediction accuracy alone. From an engineering perspective, I want to build this as a production-style data and ML platform rather than a single notebook. Data ingestion, feature engineering, model training, and inference will run as containerized workloads on Kubernetes, with scheduled pipelines, experiment tracking, model versioning, automated testing, data-quality checks, monitoring, and a FastAPI/dashboard layer for displaying predictions and backtest results. The goal is to demonstrate both quantitative ML research and end-to-end data/ML engineering. 1
I want to build a medium to long-term prediction model for commodities such as gold in the global and Indian markets, alongside some macro drivers like exchange rate, inflation, etc. 1
I want to build a short-term stock return prediction model for large US companies using RSI, MACD, moving averages, trading volume, volatility, and market index data. I would train a machine learning model to predict 5-day returns and backtest a simple investment strategy. 1
I would like to build a machine learning model for predicting short-term stock price movements in the Nigerian stock market, with a focus on large and actively traded companies. The project would use historical price and volume data together with technical indicators such as RSI, MACD, moving averages, and volatility. I would also like to explore whether market sentiment from financial news can improve the predictions. My goal would be to predict the direction of a stock over a short investment horizon, such as 5 to 30 trading days. I plan to compare different ML models, such as Logistic Regression, Random Forest, and XGBoost, and evaluate their performance using appropriate time-series validation techniques. The final objective would be to determine whether the model can generate useful signals for a simple investment strategy while considering risk and transaction costs. 1
EnergyRisk Colombia — Electricity Spot Price Forecast I would like to build a simple machine learning project to forecast the next-day average electricity spot price in Colombia and identify periods of unusually high volatility. The project will use public data from XM, the operator and market administrator of Colombia’s interconnected electricity system. The main prediction target will be the next day’s average national spot price in COP/kWh. The explanatory variables will include recent spot prices, electricity demand, useful reservoir volume and hydrological inflows. Hourly observations will be aggregated into daily data to keep the first version simple and interpretable. I will compare the model against basic benchmarks such as the previous day’s price and a seven-day moving average. A linear regression and one tree-based regression model may be evaluated using time-based train and test periods. High-volatility periods will be identified using the rolling standard deviation of daily prices. The project could help electricity retailers and large consumers understand short-term purchasing risk. It is intended as an analytical prototype rather than a trading or production forecasting system. 1
A prediction model for the us and Japan market 1
I want to develop a machine learning model that predicts short-term abnormal stock-price reactions to clinical trial results and regulatory events in biotechnology and pharmaceutical companies. I plan to combine biomedical features such as trial phase, sample size, therapeutic area, primary endpoint outcome, efficacy, and safety results with financial features such as historical returns, volatility, trading volume, and market performance. I will compare several machine learning models and evaluate whether incorporating biomedical information improves prediction compared with using financial data alone. 1

6. Investigate new metrics

0 / 148 correct (0.0%)

Answer Count
- 2
Commodities,FX,Macro,News(Bloomberg) 1
For my capstone project, I would explore several additional metrics that may help predict short-term stock price movements. First, I would collect trading volume and volatility from Yahoo Finance using yfinance, since unusual volume or increasing volatility may indicate changing investor sentiment. I would also calculate technical indicators such as RSI, MACD, and moving averages from historical price data to measure momentum and identify possible trend changes. I would include macroeconomic indicators from FRED, particularly the Federal Funds Rate, 10-year Treasury yield, inflation (CPI), and unemployment rate, because changes in interest rates and economic conditions can significantly affect technology and telecommunications companies. These series can be retrieved in Python using pandas_datareader. Finally, I would explore the VIX volatility index as a measure of overall market uncertainty and compare individual stock performance with the S&P 500 and Nasdaq to determine whether movements are company-specific or driven by the broader market. In Python, I would retrieve stock and index data with yfinance and macroeconomic data from FRED with pandas_datareader, then combine the time series by date and use them as features for my machine learning model. 1
stock returns, trading volume, RSI, and volatility, and compared to S&P500 1
"Dólar Cripto": "ARS=X","Bitcoin Liquidez": "BTC-USD", "Estabilidad Tether": "USDT-USD" 1
1) % of stocks above 20, 50, 100, 200 EMAs; 2) Breadth & Equal-Weight divergences (RSP/SPY, SMH/SPY); 3) Macro spreads & volatility spikes (10Y Yields, DXY, VIX/VVIX); 4) Multi-timeframe moving averages & FBMA bandwidth. 1
1. Trading volume: High volume after earnings may indicate a stronger investor reaction. 2. Stock volatility: Volatility shows how risky or unstable a stock was before earnings. I would calculate it from daily returns. 1
Useful additional metrics include the VIX as a measure of market risk, the 10-year minus 2-year Treasury yield spread as a recession indicator, credit spreads as a measure of financial stress, trading volume for detecting unusual market activity, and news sentiment for capturing new information. These series can be retrieved using Python from FRED, Yahoo Finance, and financial news APIs. 1
VIX 1990-01-02 2026-09-11 9243 BAA10Y 1986-01-02 2026-09-09 10172 T10Y3M 1982-01-04 2026-09-10 11175 1
An investigation of additional metrics, including the VIX, US 10-year Treasury yield, US Dollar Index, and crude oil prices, with code for downloading and analyzing the data. 1
checking on the historical prices using yfinance 1
For the AI-infrastructure capstone I will track how the large spenders allocate capital — Microsoft (working assumption: about $20 billion of infrastructure programmes, always checked against the latest 10-Q), plus Amazon, Google, Meta, Oracle, Apple and NVIDIA — and which listed contractors receive that spend (TSMC, SK Hynix, Vertiv, Constellation, Entergy, and the rest of the Q5 books). Metrics: trailing and quarterly CapEx, CapEx/sales, CapEx/depreciation, the share labelled cloud or data-center, a spender-to-contractor map, announced GW and FID status for power projects, plus ^SOX, FRED IPB53122S / VIXCLS / T10Y2Y / FEDFUNDS, earnings-surprise 2-day returns, Wikipedia GICS, and Yahoo info valuation gates. China model labs stay in a separate book (MiniMax 0100.HK, Z.AI 2513.HK; DeepSeek and Kimi via company blogs and HKEX/news, not a price series). Python: yfinance (history, get_earnings_dates, info, cashflow), pandas_datareader for FRED after truststore.inject_into_ssl(), and pandas.read_html for Wikipedia. 1
Metrics for the project can be calculated from financial reports, here you can check, debt level, occupation, geographical presence, properties, and more specific metrics like FFO, AFFO, NOI. 1
Beyond the equity price data already collected, three additional metrics are directly relevant to a geopolitical event-study project: **1. Geopolitical Risk Index (GPR)** — a daily index built by Caldara & Iacoviello (Federal Reserve Board) by automatically scanning major newspapers for geopolitical-risk-related terms. It gives a quantitative, ready-made measure of geopolitical tension, useful as an independent variable to correlate against sector-level returns (Defense, Energy, Travel & Leisure) without having to build a custom NLP/sentiment pipeline. ```python import pandas as pd gpr = pd.read_excel('https://www.matteoiacoviello.com/gpr_files/data_gpr_daily_recent.xls') ``` **2. CBOE Volatility Index (VIX)** — the standard market-wide "fear gauge." Useful as a control variable: it lets me separate a sector-specific reaction (e.g. Defense stocks rallying) from a broad risk-off move affecting the whole market at once. ```python import yfinance as yf vix = yf.download('^VIX', start='2018-01-01', end='2026-08-21') ``` **3. Brent Crude Oil futures (BZ=F)** — the direct commodity counterpart to the Energy equity vertical. Comparing the daily return of oil futures against the daily return of Energy-sector equities (e.g. Shell, TotalEnergies, Eni) around each geopolitical shock tests whether the equity market reacts in lockstep with the commodity or with a lag/divergence — directly addressing the original motivation for using company-level data instead of the sector index alone. ```python oil = yf.download('BZ=F', start='2018-01-01', end='2026-08-21') ``` Together, these three series let the project distinguish between *market-wide* risk sentiment (VIX), a *quantified geopolitical driver* (GPR), and the *commodity transmission channel* (Brent), against which the sector-level equity reactions (Defense, Energy, Travel & Leisure) can be benchmarked. 1
To support my capstone project, I would pull three additional time series: CBOE Volatility Index (VIX): Retrieved via yf.download('^VIX'). This metric gauges broader market fear. I believe high VIX environments might mute the positive price impact of strong earnings reports, serving as a vital macro-filter for my strategy. US Treasury Yield Curve (10-Year minus 2-Year): Accessed using the pandas_datareader library linked to the FRED (Federal Reserve Economic Data) API (T10Y2Y). The slope of the yield curve is a reliable proxy for economic expansion or contraction, helping contextualize whether the market favors growth or value stocks at any given time. Insider Net Buying Volume: While trickier to scrape natively through yfinance, this can be retrieved via the SEC EDGAR database (Form 4 filings). Tracking the ratio of insider buys to sells within a specific mid-cap company provides a strong signal of internal conviction that technical indicators cannot capture. 1
VIX Index (^VIX): Also known as the 'fear index,' it measures the market's expectation of future stock market volatility over the next 30 days, serving as a forward-looking sentiment indicator. Producer Price Index for All Commodities (PPIACO): This index measures the average change in selling prices received by domestic producers for their output, acting as a crucial indicator of inflationary pressures at the producer level. Crude Oil Prices (WTI - DCOILWTICO, Brent - DCOILBRENTEU): These prices reflect global supply and demand for oil, impacting energy costs and inflation worldwide. Corporate Financials (NVDA financials, balance sheet): These are detailed financial statements (income statements, balance sheets) that provide deep insights into a company's historical performance, financial health, and profitability. 1
A few metrics that could strengthen my correction early-warning model: the VIX (^VIX via yfinance) as a more direct fear/volatility signal than realized volatility; additional Treasury yields like DGS10 and DGS2 (via pandas_datareader + FRED) to build out the full yield curve instead of just one spread; and news sentiment around market-moving events (Polygon.io's news endpoint) to capture shocks that macro indicators like CPI or GDP only reflect months later. 1
To complement daily OHLCV data, I explored four additional metrics on Bitcoin and Ethereum (2018-present): retrospective returns over ranges of 1, 7 and 30 days; 30-day rolling volatility; 30_day rolling drawdown and finally 30-day volume z-score. 1
Metrics I would add, and why each one earns its place: From Yahoo Finance: ^VIX (30-day implied volatility) as a regime filter, since technical signals behave very differently in high- and low-VIX environments. ^VVIX (volatility of volatility), which tends to signal regime shifts before the VIX does. ^TNX (US 10-year yield), because the level and slope of rates drive equity valuation. HYG (high yield bonds), since credit usually turns before equities and works as an early stress warning. DX-Y.NYB (dollar index), which affects S&P foreign earnings and flows into emerging markets. GC=F (gold) to classify risk-off episodes. And RSP (equal-weighted S&P 500), which compared against SPY measures market breadth and how concentrated the index has become in a handful of names. From FRED (pandas_datareader.data.DataReader(series, 'fred')): T10Y2Y for the 10y–2y curve slope, the most widely followed recession lead indicator; BAMLH0A0HYM2 for the high yield spread; DGS3MO for the short end; plus inflation and employment series to build a macro surprise index. From CBOE: the daily put/call ratio, downloadable as CSV. It measures positioning, which cannot be derived from price alone. From the CFTC: weekly COT net positioning in S&P futures by participant type, via their public data API. From the options chain (yf.Ticker('SPY').option_chain(date)): open interest and volume by strike, to approximate market maker gamma exposure — which is the mechanism that would explain why certain price levels act as a magnet or a barrier in the first place. My inclusion criterion is the same for all of them: a metric only goes in if it carries information that is not already contained in the price series. Otherwise it just adds noise and multiplies the risk of finding spurious relationships. 1
For my initial capstone idea, I would like to explore several additional metrics that could be useful for short-term stock market analysis. These may include the VIX index to measure market volatility and fear, trading volume to identify the strength of price movements, S&P 500 returns to represent the overall market direction, and news sentiment to measure the potential impact of positive or negative global news. I may also explore technical indicators such as RSI and MACD. For the market and price-related data, I could use Python with yfinance to retrieve historical time series. For news sentiment, I could explore available news data sources and apply sentiment analysis in Python. This is an initial exploration based on my current project idea, and I may add, remove, or change these metrics and data sources as I progress through the course and better understand the available data. 1
For my capstone project on Brexit's impact on the UK market (FTSE 100 vs. S&P 500), I would add the following metrics beyond closing price: GBP/USD exchange rate. Since Brexit is fundamentally a political/currency-sensitive event, the pound's exchange rate may be an even more direct signal of political shock than the equity index itself — currency markets often react faster and more sharply to political uncertainty. This is retrievable via yfinance using the ticker GBPUSD=X, downloaded the same way as the equity indices (yf.download('GBPUSD=X', start=..., end=...)). Trading volume. This data is already included in every yfinance download (the Volume column) but hasn't been analyzed yet. Volume measures how many market participants were actively involved in a given price move — a sharp drop on high volume is a stronger, more convincing signal than the same drop on low volume, and a volume spike without a matching price move can indicate confusion or indecision among traders. Comparing FTSE 100 volume on key Brexit dates (the 2016 referendum, the 2020 exit) to typical volume would help gauge how significant those events actually were to market participants, not just how the price moved. UK government bond yields (Gilts). Bond yields are often more sensitive to political uncertainty than equities and provide independent, cross-market confirmation of a shock. If equities, currency, and bonds all react sharply to the same Brexit news, that's stronger evidence of a genuine market shock than a move in the FTSE 100 alone. This can be retrieved via yfinance using a UK 10-year gilt yield ticker (e.g. ^TNX-equivalent for UK, or via a gilt ETF/futures ticker) or, more reliably, from the Bank of England's data API / FRED. Realized volatility. Rather than sourcing this externally, it can be computed directly from the price data already collected — using a rolling standard deviation of daily returns (returns.rolling(window=21).std() for a ~1-month rolling window, for example). This adds a risk/uncertainty dimension that neither price nor volume captures directly, and requires no additional data source. Together, these four metrics (price, volume, currency, and bond yields/volatility) approach the same Brexit shock from different market angles — price direction, participation, currency risk, and cross-market confirmation 1
For the healthcare-stock prediction project, I would explore several additional metrics. First, I would calculate technical indicators such as RSI, MACD, moving-average ratios, rolling volatility, trading-volume changes, and distance from the 52-week high using historical Yahoo Finance data. These may help capture momentum, volatility, and overbought/oversold conditions. I would also include market-regime variables such as the S&P 500 return, the XLV healthcare ETF return, VIX, and U.S. Treasury yields. These variables could help distinguish stock-specific movements from broader market or sector conditions. Earnings-related features could include EPS surprise percentage, recent earnings growth, and the number of days before or after an earnings announcement. The Amazon analysis in this homework suggests that earnings surprises may contain information about short-term price reactions. Finally, I would create relative-performance features such as the stock's 5-day return minus XLV's 5-day return. This would be useful because the target of the project is sector-relative outperformance rather than simply whether the stock price increases. 1
To support my quantitative portfolio management project, I would investigate the following time-series metrics to help the model assess risk and market regimes, rather than just predicting price direction: 1. The VIX (CBOE Volatility Index) • Why it's useful: The VIX acts as a macroeconomic "fear gauge." Tracking this allows the ML model to identify the current market regime. If volatility is spiking, the system knows it must automatically reduce equity exposure or tighten stop-losses to avoid heavy portfolio drawdowns. • How to retrieve: I would use the yfinance library in Python to pull the daily closing prices using the ticker symbol: yf.Ticker("^VIX").history(period="max") 2. 10-Year US Treasury Yield (^TNX) • Why it's useful: The risk-free interest rate heavily dictates stock valuations, especially for high-growth tech stocks. By tracking bond yields, the model can dynamically adjust the portfolio's allocation between aggressive equities and defensive assets based on whether the macro environment is tightening or easing. • How to retrieve: Using Python: yf.Ticker("^TNX").history(period="max") 3. Asset Beta (Systematic Risk) • Why it's useful: For a portfolio risk-management dashboard, I need to know how sensitive a specific stock is compared to the overall S&P 500. Filtering for low-beta stocks during bear markets is a classic quantitative strategy to preserve capital. • How to retrieve: I would retrieve the stock's summary info dictionary using yfinance and extract the fundamental beta value: yf.Ticker("AAPL").info.get('beta') 1
For this project, I would explore several additional time series and metrics that could help detect and predict supply chain disruptions. 1. Lead time: useful for measuring how long suppliers or shipments take to fulfill orders and detecting abnormal increases. 2. Supplier reliability: useful for measuring the historical frequency of late deliveries, shortages, or failed orders for each supplier. 3. Inventory level and inventory turnover: useful for identifying potential stockout risks and abnormal inventory accumulation. 4. Demand volatility: useful for identifying unusual changes in customer demand that could create supply-demand imbalances. 5. Shipment delay: useful for detecting transportation problems and measuring the severity and persistence of logistics disruptions. 6. Weather and external-event indicators: useful because severe weather, natural disasters, or other external events can cause transportation and supply disruptions. 7. Commodity prices: useful for identifying changes in input costs that may affect procurement, production, or supplier behavior. 8. Stockout and disruption indicators: useful as target variables for training an early-warning model that predicts whether a disruption is likely to occur within a future time window. These metrics could be retrieved from supply-chain datasets, public APIs, financial-data providers, weather APIs, logistics data sources, or other relevant open datasets. Using Python, I would retrieve and align the different time series based on timestamps, engineer features, and evaluate which signals provide useful information before a disruption occurs. 1
CBOE Volatility Index (^VIX) serves as a macro risk-on/risk-off regime filter. High market-wide volatility often degrades short-term technical signals, requiring tighter stop-losses or position-size scaling. 1
I would explore VIX, trading volume, volatility, Treasury yields, earnings data, and news sentiment to improve stock movement predictions. 1
For my DAX vs NEPSE rotation, I’ll add EUR/NPR (EURNPR=X) for remittance FX-hedging, German 10Y Bund (FRED IRLTLT01DEM156N) vs US 10Y for DAX discount-rate regime, and Nepal remittances (World Bank BX.TRF.PWKR.CD.DT, $11.25B 2024) as NEPSE liquidity lead. They’re fetched via yfinance for FX, pandas_datareader for FRED yields, and requests to World Bank/NRB APIs, then merged as 20-day MoM/YoY features to the LightGBM label. 1
Stock Markets 1
I will use the yfinance library in Python to extract: Exchange Rates (USD/BRL, USD/INR): To evaluate currency risk impacts. VIX Index (^VIX): To measure global market volatility and fear. S&P 500 (^GSPC): As a baseline to calculate correlation features (e.g., does a drop in the US market today predict a drop in Brazil tomorrow?). 1
VIX, curva 10-2 años, rotación sectorial, sentimiento de noticias y rendimiento overnight 1
I will explore new metrics on Market Breadth and Risk (Volatility, 1
Overall Market Volatility (e.g., VIX Index): The CBOE Volatility Index (VIX) measures the market's expectation of future volatility. 1
Any from https://portfoliocharts.com/charts/ 1
To enhance feature density, I plan to integrate the CBOE Volatility Index(^VIX) to measure systemic fear levels and the 10 year Treasury Yield (^TNX) to track macroeconomic shifts in capital allocations.These can be programmatically added using yf.download(['^VIX' ,'TNX']) and joined directly with core asset frames. 1
I looked at the Baltic Dry Index ; a shipping-cost index unrelated to stock market sentiment, which tends to lead global industrial activity and GDP by weeks to months, making it a useful macro cross-check independent of equity-market psychology. 1
I am already well acquainted with the sources mentioned in class. Sadly they are not suited for my local market. Philippines fall under emerging markets, but that group is heavily weighed by Taiwan and Hong Kong/China equities. This makes FRED emerging indicators not as applicable. Even yFinance does not support Philippine equities. My only free source for technical and fundamental data is the local exchange itself. It regularly publishes end-of-day market reports which contains the OHLCV for all tickers. I already built scripts with AI that extract data from these documents. I had it build the scraper using requests and BeautifulSoup, but I studied the website and looked for patterns in the reports and url. I also made a script utilizing pandas and DuckDB to process the data. I employed the same pattern downloading the company fundamentals and disclosures. No EOD reports for this, making it a lot more complex. I had AI build a separate script utilizing Base64 and Selenium, and later pdfplumber for extraction. I designed it to first create a catalogue of all the disclosures along with its filing dates. I looked for clues in the page source and inspected the elements, and saw a unique code that I can use to cycle through the disclosures. I had this data added to the catalogue. The script now cycles through the website using these keys, and saves the snapshot as pdf files. They follow a template and this is how pdfplumber collect the fundamental and disclosures data. In summary, the steps: 1. Create catalogue, 2. Compare website to catalogue, 3. Download new data, 4. Process data using DuckDB, 5. Export to Amibroker (a charting software) so I can use it now for signal generation, 6. [Future] Perform rigorous analysis in Python using methods I will learn in class. As for external metrics like local industry and macro data, these are not well-maintained unlike its foreign counterparts. I already checked the statistics department, and a lot of the surveys are 5 to 10 years old. They are obsolete and I have no guarantee when they get updated. I want the data sources to be readily available and regularly updated. Choices are truly limited, and I have to address this challenge by being creative with how I incorporate the data. I am also looking for potential proxies. 1
For my project, I want to explore technical indicator time series calculated over different rolling windows to capture momentum, trend strength, and volatility regimes: Moving Average Convergence Divergence (MACD): Measures momentum changes by comparing short-term and long-term exponential moving averages. It helps identify trend direction and potential crossover buy/sell signals. Relative Strength Index (RSI): Measures the speed and change of price movements on a scale from 0 to 100 to identify overbought (>70) or oversold (<30) conditions. Bollinger Bands Width: Measures price volatility. Narrowing bands indicate low volatility (potential upcoming breakout), while wide bands signal high volatility. 1
vix = yf.download("^VIX", start="2015-01-01", progress=False) 1
Investigate financial metrics such as daily returns, annualised volatility, Sharpe ratio, maximum drawdown, beta and cumulative returns to evaluate stock performance and risk. 1
Price, momentum, and realized risk, Market stress and regime, Interest rates and the yield curve, Currency and cross-market conditions, Breadth and liquidity, Earnings and fundamentals 1
GPR at https://www.matteoiacoviello.com/gpr.htm 1
I would explore trading volume because unusually high volume can indicate strong buying or selling activity. I could compare volume changes with stock price movements. This may help identify periods when investors are reacting strongly to new information. 1
I would investigate RSI, MACD, trading volume, historical volatility, and earnings surprise. RSI and MACD can help identify momentum and trend changes, while trading volume can show the strength of price movements. Historical volatility can measure market risk, and earnings surprise can indicate whether a company's results exceeded expectations. These metrics could be used as features in a machine-learning model to predict stock returns. Research on financial ML also commonly uses technical, volatility, volume, and fundamental indicators as predictive variables. 1
I've retrieved data for three additional metrics that can enhance our market analysis: 1)VIX Index (CBOE Volatility Index): Why useful: Known as the 'fear gauge,' VIX measures market expectations of future volatility. High VIX values indicate increased uncertainty and potential market downturns, serving as a crucial indicator for market sentiment and risk. Retrieval method: Using yfinance for ticker ^VIX. 2) Federal Funds Rate (Effective Federal Funds Rate - FRED): Why useful: This key interest rate, set by the U.S. Federal Reserve, influences borrowing costs and investor risk appetite. Its changes provide vital macroeconomic context for market movements and economic cycles. Retrieval method: Using pandas_datareader from the FRED database for series FEDFUNDS. 3) Gold Prices (Futures): Why useful: Gold is a traditional safe-haven asset, often performing well during economic uncertainty or high inflation. Tracking gold prices can indicate investor risk aversion and shifts in asset allocation during crises. Retrieval method: Using yfinance for ticker GC=F (Gold Futures). These metrics provide valuable insights into market sentiment, macroeconomic conditions, and investor behavior, offering a more holistic view for our project. 1
Measure implied volatility and general market fear to adjust dynamic position risk. 1
IHSG (^JKSE) as benchmark/regime filter, USD/IDR exchange rate, Bank Indonesia rate & CPI, sector breadth (% of IDX80 above 50-day MA), IDX foreign net buy/sell flow, and corporate actions calendar (splits/rights issues) to avoid distorting the price series. 1
Additional metrics from yfinance: 14-day RSI, 20/50-day Moving Average ratios, and Bollinger Band percentage bandwidth. 1
I reused FRED series from the Module 1 notebook via pandas_datareader: FEDFUNDS (does buying the dip look different when rates are high), CPILFESL (core CPI / inflation), and DGS5 (5-year Treasury; we also had DGS1 in class). They are not all daily, so I would take the last value before the dip. 1
Building on an intraday/short-term prediction strategy for liquid, multi-asset markets (such as Gold/Silver commodities and major FX pairs), incorporating high-frequency market microstructure and macroeconomic liquidity metrics is essential for capturing non-linear volatility shifts. Here are three high-value data metrics to enhance the strategy, along with their strategic rationale and complete Python retrieval code. Key Data Metrics & Strategic Value 1. CME Market Depth & Order Book Imbalance (Level 2 / Depth of Market) • What it measures: The net difference between bid and ask liquidity (volume resting at the top 5–10 price levels) for COMEX Gold (GC) and Silver (SI) futures. • Why it’s useful: Traditional technical indicators like RSI or MACD lag behind real-time price action. Order book imbalance acts as a leading indicator for short-term liquidity sweeps and institutional absorption. Spikes in bid depth relative to ask depth frequently precede upward breakout continuations on lower timeframes (1-minute to 5-minute charts). 2. U.S. Dollar Index (DXY) & Real Yield Differentials • What it measures: High-frequency price movement of the U.S. Dollar Index alongside TIPS (Treasury Inflation-Protected Securities) yields. • Why it’s useful: Gold and Silver are priced in USD and function as zero-yield inflation hedges. Intraday price movements in XAU/USD are strongly inversely correlated with sudden intraday surges in real rates and the USD. Including real-time USD momentum metrics allows the ML model to filter out false technical breakouts during major currency re-pricings. 3. Economic Calendar News Releases & Volatility Surges • What it measures: Scheduled macro news events (NFP, CPI, FOMC rate decisions) combined with real-time actual vs. forecast economic surprise values. • Why it’s useful: Standard statistical and ML models fail during high-volatility news events due to extreme spread expansion and slippage. Modeling the time-to-event countdown alongside historical "economic surprise" standard deviations enables dynamic position sizing, widening stop-loss thresholds, or flatting positions prior to major announcements. Data Retrieval Implementation in Python Below is a Python script demonstrating how to pull historical daily/intraday prices (Gold, Silver, DXY) using yahoo finance, fetch macro yield data from FRED via pandas data reader, and compute custom order-flow/volatility features. 1
I am not sure about my capstone project yet so as to be able to provide new metrics, but I would like to explore with bloomberg or yfinance datasets. 1
We analyzed four conditions: - Market uncertainty — VIX: How much stock-market movement investors expect. - Credit conditions — HYG/LQD: Whether investors prefer riskier or safer company bonds. - Interest-rate expectations — T10Y2Y: The difference between 10-year and 2-year US Treasury rates. - Economic activity — INDPRO_YOY: Annual growth in US industrial production. We compared these conditions with the month’s return for all 11 S&P 500 sectors. 1
VIX and Fed funds 1
I would like to use the VIX data 1
VIX (^VIX) — the market’s expected 30-day volatility. The regime matters a lot for post-earnings drift: if VIX is calm and expected vol is low, index-fund flows are calmer and surprises should be priced more efficiently. US 10-year Treasury yield (^TNX) — the risk-free discount rate. Rising yields compress equity multiples, especially for long-duration tech/growth stocks such as AMZN — a natural second factor. 1
Done in the corresponding notebook 1
I will cross different metrics to build ones that fit more the data 1
news, sentiment analyses 1
VIX, Treasury Yield, DXY, Gold, Oil, Credit Spreads, Market, Breadth, Earnings Revisions 1
Building on the Q1–Q4 analyses in this homework, I explored a few additional metrics relevant to my capstone direction: (1) correction duration alongside drawdown magnitude (median 40 days for S&P 500 corrections, vs. 8% median depth) — duration and depth don't move together, so both matter for characterizing risk; (2) the correlation between earnings surprise magnitude and 2-day price reaction (found to be weak, ~0.24, for AMZN), suggesting surprise size alone is a poor predictor and that "how much was already priced in" is likely a stronger signal — this directly motivates the capstone's factor-attribution approach; (3) for the capstone itself, I plan to explore hawkish/dovish sentiment scoring of BoE statement language as a derived metric, and factor-importance weights (e.g., via SHAP) linking macro variables to that sentiment score over time. 1
I will use adjusted prices and volume from Yahoo Finance for returns, realized volatility, and liquidity; the Selic rate and inflation from the Central Bank for macroeconomic regimes; BRL/USD exchange rates for external exposure; and news-sentiment series. These indicators will be retrieved through APIs or CSV downloads in Python and aligned to trading dates. 1
Examples: Volatility Index (VIX), 10‑Year US Treasury yields, commodity prices (oil, gold), ETF flows. 1
DEX Volume-to-Liquidity Ratio ($V/L$), Liquidity Pool Lock / Burn Percentage 1
News sentiment/event scores — an LLM pipeline over headlines (reusing my RAG chunking/retrieval stack) via a source like NewsAPI or GDELT; useful because it captures information price data alone misses, especially in thinly-covered markets. Implied volatility (VIX, or a proxy since most frontier markets lack one) — pulled via yfinance (^VIX) as a macro risk-regime feature, to test whether my model behaves differently in high- vs low-volatility regimes. Sector/industry relative strength — a stock's return vs. its GICS sector ETF, via yfinance, to separate stock-specific signal from broad sector moves. Insider transactions / institutional ownership changes — SEC EDGAR Form 4 filings (for any US-cross-listed names) as a slower-moving conviction signal to complement short-term technicals. 1
distance from 52-week high, 30-day realised volatility vs. its 1-year average, earnings-surprise magnitude 1
I would explore trading volume, moving averages, RSI, and market volatility. These metrics can help show how actively a stock is traded, its price trend, and how much its price changes over time. I can retrieve these data using Python and yfinance. 1
VIX (^VIX) - measures market fear; spikes often precede corrections. 10Y-2Y Treasury yield spread - an inverted curve has historically preceded downturns. Put/Call ratio - options sentiment indicator; elevated puts often signal hedging before drops. 1
I think there are quite a few metrics that i like to use when I am doing my day to day trading. RSI, Boll bands, call wall and put wall from the option market, fibonacci 0.618 and also fear and greed index from the macro enviromenment. For most of them I can retrieve it from yahoo finance directly and with some additional calculations. for the Fear and greed index, I can get it from Alternative.me’s Crypto Fear & Greed Index API 1
For my capstone (momentum + earnings-drift strategy on US mega-cap tech), I explored three additional data sources beyond price history: 1. CBOE Volatility Index (VIX) — a market-regime feature. Momentum and post-earnings drift behave differently in high-volatility regimes, so I will use VIX levels and its 20-day change as a state variable (and later as a regime filter in the model). Retrieval is trivial via yfinance: import yfinance as yf vix = yf.Ticker("^VIX").history(start="2015-01-01", end="2026-08-21")[["Close"]] 2. Earnings calendar + reported EPS vs. analyst estimate — the raw material of the PEAD signal (already prototyped in Q4 for AMZN). I will fetch it per ticker with get_earnings_dates() to build surprise magnitude and the 2-day reaction features, then join onto the price frame: cal = yf.Ticker("AAPL").get_earnings_dates(limit=30) 3. Macro rates (Fed Funds Rate, 10Y Treasury yield) — long-duration tech valuations are rate-sensitive; a rising-rate regime should change how momentum/PEAD alpha behaves. I will pull the Fed Funds series from FRED via pandas_datareader (or the fredapi package with a personal API key): from pandas_datareader import data as pdr dff = pdr.DataReader("DFF", "fred", "2015-01-01", "2026-08-21") All three will be modeled as date-dimension tables in the Power BI star schema, with DAX measures (e.g., VIX regime flag) so the same semantic model serves both the ML pipeline and the live dashboard. 1
UNICA (the Brazilian sugarcane industry association) publishes fortnightly crop production reports covering cane crushing volumes, sugar/ethanol production mix, and regional output for Brazil's Center-South region — the dominant global sugar-producing area. Since Brazilian Center-South supply is a key driver of Raw Sugar No. 11 pricing, this series would serve as a leading fundamental indicator for anticipating shifts in global supply.The reports are only available as PDFs on UNICA's website, with no structured API, so retrieval would require scraping the report archive, downloading the PDFs, and parsing the embedded tables into a clean, time-indexed dataset for analysis. 1
For my project, I would explore additional metrics such as trading volume, volatility, RSI, MACD, and earnings surprise percentage. Trading volume can indicate the strength of market activity, while volatility measures the level of price uncertainty. RSI and MACD can help identify momentum and potential trend changes. Earnings surprise percentage may also be useful for measuring how unexpected company results affect short-term stock returns. These metrics can be obtained and calculated in Python using historical market and earnings data from Yahoo Finance through the yfinance library. 1
I pulled additional Yahoo Finance series I would use in that project, all via yfinance.download, as of 21 Aug 2026. 1
Post-earnings-announcement drift (PEAD) 1
interest rates + MDAX + company KPIs + exchange rates 1
VIX volatility index, commodity futures prices, macroeconomic indicators (interest rates and inflation), options implied volatility, and financial news sentiment collected using Python from yfinance, FRED, and NewsAPI 1
VIX (^VIX) — volatility regime affects post-earnings drift and drawdown depth. US 10-Year Treasury Yield (^TNX) — rising rates compress growth/tech valuations. DXY US Dollar Index — impacts multinational revenue translation (relevant for AMZN, AAPL). Sector ETFs (XLK, XLF, XLE) — capture sector rotation / macro regime. FRED macro series (UNRATE, CPI, FEDFUNDS) via pandas_datareader for macro context. Put/Call ratio & AAII sentiment — sentiment contrarian signals. Retrieve all via yfinance (tickers like ^VIX, ^TNX, DX-Y.NYB) or pandas_datareader.get_data_fred(). 1
I might investigate CBOE volatility index, Put/Call Ratio or Short Interest as % of Float. 1
data sources for my project (besides price-based technicals): BETZ (Roundhill Sports Betting & iGaming ETF) — a sector-specific benchmark, pulled via yfinance, used both as a robustness check for my median-based target and as a relative-strength feature (spread of each stock's return vs. the ETF). Google Trends search interest for "online casino" / "sports betting" by regulated market (Brazil, Netherlands, Germany, US-NJ), retrieved via pytrends. This proxies acquisition-driven demand ahead of official active-player disclosures, directly testing my acquisition-vs-monetisation hypothesis. 10-Year Treasury yield (FRED, fredapi) as a macro control — several names in my universe (Evolution, DraftKings) are high-multiple growth stocks sensitive to discount-rate moves, and I want to separate that macro effect from sector-specific signal. 1
Three simple related metrics: USD/KRW exchange rate (exporter revenue sensitivity), Philadelphia Semiconductor Index ^SOX (US chip-sector spillover into Samsung/SK Hynix), and VIX ^VIX (risk-appetite proxy affecting emerging-market selloffs) — all pullable via the same yfinance calls used elsewhere in the homework. 1
SOX Index (^SOX): Benchmark to isolate industry-wide semiconductor swings from individual company moves. Retrieved via yf.download('^SOX'). Taiwan Export Orders (Electronics): Early indicator of global chip demand since Taiwan leads manufacturing. Retrieved via the Taiwan Ministry of Economic Affairs API using requests. Days of Inventory Outstanding (DIO): Signals inventory build-ups or shortages. Calculated from quarterly balance sheets (ticker.quarterly_balance_sheet) as (Inventory / COGS) * 90. 1
The CBOE Volatility Index (VIX): This metric represents market expectations of near-term volatility. It is useful because stock price reactions to individual events (like earnings surprises or index additions) scale differently during low-volatility regimes vs. high-fear bear market regimes. It can be retrieved easily in Python via yf.download("^VIX"). 1
1. Foreign & Proprietary Net Flow (Khối ngoại & Tự doanh mua/bán ròng): In an 85%+ retail-driven market, foreign funds and prop desks act as anchor price drivers. Tracking 5-day and 20-day cumulative net flow by sector detects institutional accumulation vs. distribution ahead of retail FOMO (retrieved via python package 'vnstock3' or SSI iBoard API). 2. Sector Liquidity Concentration Ratio (% Thanh khoản ngành / Total HOSE Turnover): Extreme turnover concentration (>30-35% flowing into a single sector like Banking or Securities) signals momentum inflection points or retail climax tops (calculated via 'vnstock3' daily industry transaction series). 1
Volatility, Sharpe Ratio, Maximum Drawdown, Earnings Surprise %, and abnormal returns. 1
For my capstone project, I would explore additional market indicators such as the VIX volatility index, U.S. 10-year Treasury yield, trading volume and realized volatility. The VIX could help identify periods of market fear and uncertainty, while Treasury yields may capture changes in interest-rate expectations that affect technology stock valuations. Trading volume could help identify unusually strong market movements. I would also calculate 20-day realized volatility from daily stock returns. These datasets can be retrieved with Python using yfinance, for example with the tickers ^VIX and ^TNX, while realized volatility can be calculated directly from historical stock prices using pandas 1
Investigate new metrics for evaluating stock market performance and identifying patterns in market movements. 1
The key valuation and performance metrics for the Stock Exchange of Thailand (SET) stand at a Price-to-Earnings (P/E) ratio of 15.84 and a Price-to-Book Value (P/BV) of 1.48. 1
VIX, TNX & XLK are the metrics I plan to integrate 1
VIX 1
I would add on-chain metrics (active addresses, exchange netflows, hash rate) and sentiment/order-book data (Twitter sentiment, bid-ask spread, order-book imbalance) to improve crypto prediction. I would fetch these in Python using APIs like Glassnode, Santiment, ccxt, and yfinance. 1
For my volatility-forecasting project, a few useful additions: VIX (^VIX, yfinance): implied volatility, the forward-looking counterpart to my realized-vol target and likely my strongest feature. Realized volatility (computed from ^GSPC OHLCV, yfinance): rolling std of log returns plus a high/low range estimator (Garman-Klass); both my target and key predictor, since vol clusters. Term spread (DGS10 minus TB3MS, FRED) and Fed Funds Rate (FEDFUNDS, FRED): risk-regime indicators, since curves invert and rates shift around the periods when volatility spikes. Retrieval: yf.download(['^GSPC','^VIX'], auto_adjust=True) for prices and pandas_datareader's FRED reader for the macro series. I'd resample the mixed-frequency macro data to weekly and forward-fill only past values to avoid leakage. 1
Metrics I'd add, each as a feature for predicting LLM factual errors: realized volatility (how fast a company's facts go stale), market cap and average volume (popularity → better represented in training data), earnings dates and surprise (event recency), S&P 500 membership changes (models lag on recent additions/removals), and per-ticker news-coverage volume (density of published text about the company). All are retrievable with yfinance, pandas.read_html on the Wikipedia list, and a news/RSS API such as GDELT — showing how I'd turn the project description into concrete data requests. 1
I will skip this one 1
Earlier i tried to build a bloomberg kinda system with data like news, filings, macro&micro market conditions, fundamental data, ohlcv data and ai generated text, https://medium.com/@moh1tt/i-built-a-quantitative-stock-research-platform-from-scratch-heres-everything-i-learned-b429b64caaad 1
RIR (Real Interest Rate) - inflation-adjusted interest rate - I would like to check whether it has a better predictive power than nominal rate + inflation separately 1
Mortgage interest rate, inflation (CPIF), unemployment, wages, housing transactions. Can download data from "https://www.statistikdatabasen.scb.se/pxweb/en/ssd/". There is a working API as well, so can get data with request similar to this: "import requests url = ( "https://api.scb.se/OV0104/v1/doris/en/ssd/" "BO/BO0501/BO0501B/SmahusT1M" ) response = requests.get(url) response.raise_for_status() metadata = response.json()" 1
I will investigate momentum, volatility, trading volume, RSI, and earnings surprises to predict whether to buy, hold, or avoid a selected company’s stock. 1
Controls: delisted stocks, survivorship bias, corporate actions, transaction costs, and publication dates. 1
Metric 1: VIX (CBOE Volatility Index) Why it's useful: The VIX measures market-implied 30-day volatility and serves as a "fear gauge." High VIX periods often correspond to market corrections (as we found in Q3), and stocks may react differently to earnings surprises during high vs. low volatility environments. 1
ppi. import pandas_datareader as pdr; ppi = pdr.Data_Reader('PPI', 'fred', start='2020-01-01'); 1
VIX, MA-200, MA-50, RSI 1
To support my S&P 500 weekly direction prediction project, I explored the following additional metrics: 1. VIX (CBOE Volatility Index) The VIX measures the market's expectation of near-term volatility derived from S&P 500 options. It is often called the "fear index." When VIX spikes above 30, markets are usually near a bottom, making it a useful contrarian signal. 2. RSI — Relative Strength Index (14-day) RSI measures the speed and magnitude of recent price changes to identify overbought (above 70) or oversold (below 30) conditions. It is one of the most widely used momentum indicators in both retail and institutional trading. 3. MACD — Moving Average Convergence Divergence MACD captures trend momentum by comparing two exponential moving averages (12-day and 26-day). The crossover between the MACD line and its 9-day signal line is a classic entry and exit signal used in quantitative trading strategies. 4. Golden Cross and Death Cross (50-day vs 200-day Moving Average) When the 50-day MA crosses above the 200-day MA, it is called a Golden Cross and signals a long-term bullish trend. The reverse is a Death Cross. These are among the most well-known trend regime indicators used by professional fund managers and are easily computed from daily price history. 5. Rolling Momentum Returns (1-week, 2-week, 4-week) Past short-term returns are one of the strongest predictors of near-term stock performance according to academic finance research. I will calculate 5-day, 10-day, and 20-day rolling percentage changes from the daily close price to capture short and medium-term momentum signals. 1
To support my capstone, I plan to investigate macroeconomic indicators and sector-specific ETF flows that heavily influence tech valuations. Specifically, I will track the 10-Year U.S. Treasury Yield using the FRED API, as rising interest rates typically pressure SaaS stock multiples. Additionally, I will monitor the volume of tech-focused ETFs like the Global X Cloud Computing ETF (CLOU) using yfinance. I will build a Python ingestion script using pandas and requests to fetch this time-series data, and I plan to model and store the output in a relational database so my backend services can query the historical data efficiently. 1
To support the short-term quantum computing stock prediction model, I explored two additional metrics via yfinance: - Volatility Index (^VIX): Speculative quantum stocks carry high beta. Monitoring overall market volatility helps filter out macroeconomic risk-off noise from genuine company-specific price drivers. - Semiconductor ETF (SOXX): Pure-play quantum hardware companies are tightly linked to the broader chip industry ecosystem. Sector-level price movements offer an essential benchmark for relative strength. 1
Sharpe ratio, maximum drawdown, Sortino ratio, beta, and rolling volatility. These metrics can provide a better view of risk-adjusted performance than returns alone. 1
Commodities deppend on different markets, and there are different providers for each commodity. I don't have a list of sources yet. Probably weather has an effect, but I haven't decided which commodity to analyse, thus I don't have a concrete use case in mind yet. 1
Sortino Ratio: Mide el rendimiento ajustado al riesgo evaluando únicamente la volatilidad bajista (downside risk), en lugar de la volatilidad total. A diferencia del Sharpe Ratio —que penaliza variaciones tanto al alza como a la baja—, el Sortino Ratio permite medir si el exceso de retorno compensa adecuadamente solo el riesgo de perder capital. 1
For my gold prediction model, I'd explore these additional metrics such as USD/INR exchange rates, domestic gold ETF prices, and VIX (^VIX) - A fear/risk-sentiment gauge. 1
I would explore VIX, trading volume, rolling volatility, RSI, MACD, and the 10Y-2Y Treasury yield spread. VIX can measure market fear, the yield spread can provide information about economic conditions, and volume and technical indicators can help identify momentum and risk. I would retrieve these data using Python, yfinance, and FRED 1
For my proposed Nigerian stock-market prediction project, I explored three additional time series that could be useful as machine-learning features: the NGX All-Share Index, the USD/NGN exchange rate, and the Nigerian Monetary Policy Rate. I downloaded daily NGX All-Share Index data and calculated daily returns. The index is useful because it represents the overall direction of the Nigerian stock market and could help distinguish market-wide movements from movements specific to individual companies. I also downloaded USD/NGN exchange-rate data and calculated daily changes. This could be useful because changes in the naira can affect companies differently depending on their exposure to imports, exports, foreign-currency revenue, and foreign-currency debt. Finally, I explored the Nigerian Monetary Policy Rate. Interest-rate changes could be useful for the model because they can affect borrowing costs, investment decisions, and stock valuations. The CBN provides monetary and financial time-series data that can be used for this type of analysis. As an initial exploration, I compared NGX All-Share Index returns with changes in the USD/NGN exchange rate and calculated their correlation. These variables could be added as external features alongside stock-specific technical indicators such as RSI, MACD, moving averages, and trading volume. This exercise helped me identify additional data sources and demonstrated how I could retrieve, clean, align, and investigate external time series before using them in a machine-learning model. 1
For **EnergyRisk Colombia**, I downloaded and explored 31 days of public data from XM’s Sinergox platform, ending on 21 August 2026. The objective was to identify variables that could help forecast Colombia’s next-day electricity spot price and detect periods of high price volatility. I selected four time series: * **Weighted national spot price (`PPPrecBolsNaci`)**: This is the main variable I want to predict. I also used it to calculate seven-day rolling volatility and identify unusually unstable periods. * **Real electricity demand (`DemaReal`)**: Demand measures consumption pressure on the electricity system. Higher demand could contribute to higher prices when available generation is limited. * **Useful reservoir volume (`PorcVoluUtilDiar`)**: Colombia depends heavily on hydroelectric generation, so reservoir levels provide information about the amount of stored water available for producing electricity. Lower reservoir levels may increase scarcity and price risk. * **Hydrological energy inflows (`AporEner`)**: This measures the water-energy entering the hydroelectric system. It may help anticipate future reservoir conditions and electricity-generation capacity. I retrieved the data from XM’s public API using Python’s `requests` library and converted the JSON responses into pandas DataFrames. Hourly electricity demand was aggregated into daily GWh, while the other variables were already available at daily frequency. I then joined the series by date and examined missing values, descriptive statistics, time-series graphs and correlations. In this short sample, the spot price ranged from approximately **569 to 1,032 COP/kWh**. Reservoir volume had a **−0.73 correlation** with the spot price, meaning that higher reservoir levels coincided with lower electricity prices during the period. Demand had only a weak positive correlation with price, while same-day hydrological inflows had almost no correlation. This does not mean that inflows are irrelevant, because their effect may appear later through changes in reservoir levels. I also calculated the rolling standard deviation of prices over seven days. The 90th-percentile volatility threshold was **152.95 COP/kWh**, and three days were classified as high-volatility periods. These findings are exploratory and do not prove causation. A complete version of EnergyRisk Colombia would use several years of data, include lagged variables and evaluate forecasts using chronological training and testing periods. 1
Additional metrics I would investigate include clinical trial phase, sample size, primary endpoint success, statistical significance, adverse-event rates, therapeutic area, FDA approval or rejection outcomes, and the abnormal return of the company's stock relative to the S&P 500. These variables are useful because they capture the scientific and regulatory importance of a biomedical event and may explain differences in the market's reaction. Clinical trial and regulatory data could be retrieved from ClinicalTrials.gov and FDA sources, while stock and market data could be retrieved using yfinance. 1

Calculated: 15 September 2026, 15:52