Course Overview
Course
Time Series Analysis and Forecasting
Bucharest University of Economic Studies
Duration
Lectures: 2 hours/week
Seminars: 2 hours/week
Prerequisites
Statistics, Probability Theory
Linear Algebra, Python Programming
Assessment
Final Exam: 70%
Projects & Assignments: 30%
Learning Objectives
By the end of this course, you will be able to:
- Understand stationarity, autocorrelation (ACF), and partial autocorrelation (PACF)
- Apply decomposition methods and exponential smoothing (SES, Holt, Holt-Winters)
- Build and diagnose ARMA and ARIMA models using Box-Jenkins methodology
- Model seasonal time series with SARIMA
- Model volatility clustering with GARCH, EGARCH, and GJR-GARCH
- Calculate Value at Risk (VaR) using volatility forecasts
- Analyze multivariate systems with VAR models and Granger causality tests
- Interpret Impulse Response Functions and Forecast Error Variance Decomposition
- Test for cointegration and estimate VECM models
- Implement time series analysis and forecasting in Python
Key Formulas
Decomposition
Additive: $X_t = T_t + S_t + \varepsilon_t$
Multiplicative: $X_t = T_t \times S_t \times \varepsilon_t$
Stationarity
$\mathbb{E}[X_t] = \mu$ (constant)
$\text{Var}(X_t) = \sigma^2$ (constant)
$\text{Cov}(X_t, X_{t+h}) = \gamma(h)$
AR(p) Model
$X_t = c + \sum_{i=1}^{p}\phi_i X_{t-i} + \varepsilon_t$
Stationary if roots outside unit circle
MA(q) Model
$X_t = \mu + \sum_{j=1}^{q}\theta_j\varepsilon_{t-j} + \varepsilon_t$
Always stationary
ARIMA(p,d,q)
$\phi(L)(1-L)^d X_t = \theta(L)\varepsilon_t$
$d$ = differencing order
SARIMA
$(p,d,q) \times (P,D,Q)_s$
Seasonal + non-seasonal ARIMA
VAR(p) Model
$\mathbf{Y}_t = \mathbf{c} + \sum_{i=1}^{p}\mathbf{A}_i\mathbf{Y}_{t-i} + \boldsymbol{\varepsilon}_t$
Multivariate AR extension
Granger Causality
$X \rightarrow Y$: past $X$ predicts $Y$
$H_0$: lagged $X$ coefficients = 0
GARCH(1,1)
$\sigma_t^2 = \omega + \alpha \varepsilon_{t-1}^2 + \beta \sigma_{t-1}^2$
Volatility clustering model
Value at Risk
$\text{VaR}_\alpha = \mu + z_\alpha \cdot \sigma_t$
Risk measure at confidence $\alpha$
Cointegration
$Y_t = \beta X_t + u_t$, $u_t \sim I(0)$
Long-run equilibrium relationship
Model Selection
AIC $= -2\log L + 2k$
BIC $= -2\log L + k\log n$
Forecast Metrics
RMSE $= \sqrt{\frac{1}{n}\sum(y_t - \hat{y}_t)^2}$
MAPE $= \frac{100}{n}\sum|\frac{y_t - \hat{y}_t}{y_t}|$
Prophet Model
$y(t) = g(t) + s(t) + h(t) + \varepsilon_t$
Trend + Seasonality + Holidays
TBATS
Trigonometric + Box-Cox + ARMA
Multiple seasonality via Fourier
Course Chapters
Chapter 0: Fundamentals
- What is a Time Series?
- Time Series Decomposition
- Exponential Smoothing Methods
- Forecast Evaluation
- Modeling Seasonality
- Handling Trend and Seasonality
Chapter 1: Stochastic Processes & Stationarity
- Stochastic Processes
- Stationarity (Strict & Weak)
- White Noise and Random Walk
- Autocorrelation Functions (ACF/PACF)
- Lag Operator and Differencing
- Testing for Stationarity (ADF, KPSS)
Chapter 2: ARMA Models
- Autoregressive (AR) Models
- Moving Average (MA) Models
- ARMA Model Identification
- Parameter Estimation
- Model Diagnostics
- Forecasting with ARMA
Chapter 3: ARIMA Models
- Non-Stationarity & Unit Roots
- Differencing & Integration
- ARIMA(p,d,q) Models
- Unit Root Tests (ADF, KPSS)
- Box-Jenkins Methodology
- Forecasting with ARIMA
Chapter 4: SARIMA Models
- Seasonality in Time Series
- Seasonal Decomposition
- Seasonal Differencing
- SARIMA$(p,d,q) \times (P,D,Q)_s$
- Airline Model
- Seasonal Forecasting
Chapter 5: Volatility Models ARCH/GARCH
- Stylized Facts of Financial Returns
- Volatility Clustering & Heteroskedasticity
- ARCH Model (Engle, 1982)
- GARCH, IGARCH, EWMA, CGARCH, FIGARCH
- Asymmetric Models: EGARCH, GJR-GARCH, TGARCH, APARCH
- Markov-Switching GARCH
- Volatility Forecasting, Value at Risk & Expected Shortfall
Chapter 5b: Multivariate GARCH Models
- Multivariate Volatility Modeling
- VECH and BEKK Models
- DCC-GARCH (Dynamic Conditional Correlation)
- CCC-GARCH (Constant Conditional Correlation)
- Portfolio Risk & Hedging Applications
Chapter 6: VAR Models & Granger Causality
- Vector Autoregression (VAR)
- Granger Causality Testing
- Impulse Response Functions
- Forecast Error Variance Decomposition
- VAR Diagnostics & Forecasting
- Structural VAR (SVAR) Introduction
Chapter 7: Cointegration & VECM
- Spurious Regression Problem
- Cointegration Concept
- Engle-Granger Two-Step Method
- Johansen Cointegration Test
- Vector Error Correction Model (VECM)
- VECM Estimation & Interpretation
Chapter 8: Modern Extensions
- ARFIMA Models & Long Memory
- Hurst Exponent & Fractional Differencing
- Machine Learning for Time Series
- Random Forest with Lag Features
- LSTM Networks for Sequential Data
- Time Series Cross-Validation
Chapter 9: Prophet & TBATS
- Multiple Seasonality Challenge
- TBATS: Trigonometric Seasonality
- Fourier Terms & Box-Cox Transform
- Prophet: Decomposable Models
- Trend Changepoints & Holidays
- Forecasting with Uncertainty
Chapter 10: Review
- Complete Analysis Workflow
- Case Study: Bitcoin (ARIMA-GARCH Volatility)
- Case Study: Sunspots (11-Year Cycle, SARIMA)
- Case Study: US Unemployment (Structural Breaks)
- Model Selection Decision Guide
- Forecast Evaluation Metrics
Chapter 11: LLMs and Foundation Models
- Transformer Architecture and Self-Attention
- Foundation Models: Chronos, TimesFM, TimeGPT, Lag-Llama, Moirai
- Tokenization Strategies: Patching, Quantization
- Zero-Shot vs Fine-Tuning
- Use Cases: EUR/RON, Energy, Volatility
- Benchmarks, Limitations, and Future Directions
Chapter 12: Spectral Analysis
- Fourier Analysis and the Discrete Fourier Transform
- Periodogram and Spectral Density Function
- Spectral Estimation: Multitaper, Welch, Blackman-Tukey
- Cross-Spectral Analysis: Coherence, Phase, Gain
- Band-Pass Filters: Baxter-King, Christiano-Fitzgerald, HP
- Wavelet Analysis and Scalograms
Chapter 13: LPPL Models for Bubble Detection
- Financial Bubbles and Critical Phenomena
- Log-Periodic Power Law (LPPL) Model Derivation
- Crash Hazard Rate and Phase Transitions
- LPPL Estimation: Slaving and Differential Evolution
- Lomb–Scargle Confidence Indicator
- Case Studies: Dot-com, Bitcoin, Shanghai, Oil, COVID
Interactive Quizzes
Please log in with your GitHub account to access the quizzes.
Click on a chapter tab above to switch between quizzes
Chapter 0: Fundamentals
Chapter 1: Stochastic Processes & Stationarity
Question 1: Stochastic Process
What does a stochastic process $\{X_t\}$ describe?
- A) A deterministic function of time
- B) A fixed numerical value
- C) A family of random variables indexed by time
- D) A sequence of numbers in ascending order
Question 2: Autocovariance Function
What defines the autocovariance function $\gamma(t,s)$?
- A) $\mathbb{E}[X_t] \cdot \mathbb{E}[X_s]$
- B) $\text{Var}(X_t + X_s)$
- C) $\mathbb{E}[(X_t - \mu_t)(X_s - \mu_s)]$
- D) $\mathbb{E}[X_t^2]$
Question 3: Strict Stationarity
What does strict stationarity imply?
- A) Only the mean is constant
- B) Only the variance is constant
- C) All finite-dimensional distributions are invariant to time shifts
- D) The series has no visible trend
Question 4: Strict vs. Weak Stationarity
Which statement about the relationship between strict and weak stationarity is correct?
- A) Weak stationarity implies strict stationarity
- B) The two concepts are identical
- C) Strict stationarity is not relevant in practice
- D) Strict stationarity (with finite moments) implies weak stationarity
Question 5: Ergodicity
What does ergodicity allow in practice?
- A) Estimating parameters from multiple realizations of the process
- B) Estimating population parameters from a single realization
- C) Exact prediction of future values
- D) Complete elimination of uncertainty
Question 6: Wold's Theorem
What does the Wold decomposition theorem state?
- A) Any time series is stationary
- B) Any weakly stationary process decomposes into MA($\infty$) + deterministic component
- C) Every time series has a trend
- D) Only AR processes are stationary
Question 7: Lag Operator
What is $(1-L)X_t$, where $L$ is the lag operator?
- A) $X_t + X_{t-1}$
- B) $X_t - X_{t-1}$ (first difference)
- C) $X_t \cdot X_{t-1}$
- D) $X_t / X_{t-1}$
Question 8: Random Walk with Drift
For $X_t = \mu + X_{t-1} + \varepsilon_t$ with $\mu > 0$ and $X_0 = 0$, which statement is correct?
- A) The series is stationary
- B) $\mathbb{E}[X_t] = \mu$ (constant)
- C) $\mathbb{E}[X_t] = \mu \cdot t$ (grows linearly)
- D) The variance is constant
Question 9: ACF and Non-Stationarity
An Autocorrelation Function (ACF) that decays very slowly toward zero suggests:
- A) The series is white noise
- B) The series is stationary with short memory
- C) The series is non-stationary (possible unit root)
- D) The series has seasonality
Question 10: Autocorrelation Function Properties
For a weakly stationary process, $\rho(h) = \gamma(h)/\gamma(0)$ has the property:
- A) $\rho(0) = 0$
- B) $\rho(h) = \rho(-h)$ and $|\rho(h)| \leq 1$
- C) $\rho(h)$ increases with $h$
- D) $\rho(h) = 1$ for all $h$
Question 11: Log Returns
How do we obtain a stationary series from financial prices $P_t$?
- A) Apply the transformation $\sqrt{P_t}$
- B) Multiply by a constant
- C) Compute log returns: $r_t = \ln(P_t / P_{t-1})$
- D) Apply a moving average
Question 12: Weak Stationarity
Which condition is NOT required for weak stationarity?
- A) Constant mean: $\mathbb{E}[X_t] = \mu$
- B) Constant variance: $\text{Var}(X_t) = \sigma^2$
- C) Normal distribution
- D) Covariance depends only on lag
Question 13: Random Walk
For a random walk $X_t = X_{t-1} + \varepsilon_t$ with $X_0 = 100$ and $\sigma^2 = 4$, what is $\text{Var}(X_{25})$?
- A) 4
- B) 25
- C) 100
- D) 625
Question 14: White Noise
Which statement about white noise is FALSE?
- A) White noise has zero mean
- B) White noise has constant variance
- C) White noise must be normally distributed
- D) White noise has no autocorrelation
Question 15: ACF Patterns
For an AR(1) process with $\phi = 0.8$, the ACF:
- A) Cuts off after lag 1
- B) Decays exponentially
- C) Oscillates around zero
- D) Is zero for all lags
Question 16: PACF Interpretation
The PACF of an MA(1) process:
- A) Cuts off after lag 1
- B) Decays exponentially
- C) Is zero for all lags
- D) Shows a seasonal pattern
Question 17: Differencing
If $X_t$ is a random walk, then $\Delta X_t = X_t - X_{t-1}$ is:
- A) Still a random walk
- B) White noise (stationary)
- C) An AR(1) process
- D) A trending series
Question 18: ADF Test
In the ADF test, the null hypothesis is:
- A) The series is stationary
- B) The series has a unit root
- C) The series is white noise
- D) The series has no trend
Question 19: KPSS Test
If KPSS rejects $H_0$, we conclude:
- A) The series is stationary
- B) The series is non-stationary
- C) The series is white noise
- D) We need more data
Question 20: Deterministic vs Stochastic Trend
How should you remove a stochastic trend (unit root)?
- A) Fit a linear regression on time
- B) Apply differencing
- C) Use a moving average
- D) Apply seasonal adjustment
Chapter 2: ARMA Models
Question 1: AR Stationarity
For which value of $\phi$ is the AR(1) process $X_t = c + \phi X_{t-1} + \varepsilon_t$ stationary?
- A) $\phi = 1.2$
- B) $\phi = 1.0$
- C) $\phi = -0.8$
- D) $\phi = -1.5$
Question 2: ACF/PACF Pattern Recognition
You observe: ACF has a spike at lag 1, then cuts off. PACF decays gradually. What model is suggested?
- A) AR(1)
- B) MA(1)
- C) ARMA(1,1)
- D) White noise
Question 3: MA Invertibility
Is the MA(1) process $X_t = \varepsilon_t + 1.5\varepsilon_{t-1}$ invertible?
- A) Yes, MA processes are always invertible
- B) Yes, because $\theta = 1.5 > 0$
- C) No, because $|\theta| = 1.5 > 1$
- D) No, MA processes are never invertible
Question 4: ARMA Representation
The compact form $\phi(L)X_t = \theta(L)\varepsilon_t$ represents which model?
- A) Pure AR model
- B) Pure MA model
- C) ARMA model
- D) None of the above
Question 5: Lag Operator
What is $(1-L)^2 X_t$?
- A) $X_t - X_{t-1}$
- B) $X_t - 2X_{t-1} + X_{t-2}$
- C) $X_t + X_{t-1} + X_{t-2}$
- D) $X_t - X_{t-2}$
Question 6: Information Criteria
Comparing ARMA(1,1) vs ARMA(2,1) using BIC, which statement is correct?
- A) Lower BIC always means better forecasts
- B) BIC penalizes complexity less than AIC
- C) The model with lower BIC is preferred
- D) BIC can only compare models with same number of parameters
Question 7: Ljung-Box Test
After fitting an ARMA model, you run the Ljung-Box test on residuals and get p-value = 0.03. What does this mean?
- A) Model is adequate, residuals are white noise
- B) Model is inadequate, residuals have autocorrelation
- C) Need to increase sample size
- D) Test is inconclusive
Question 8: Forecast Properties
For a stationary AR(1) model, what happens to forecasts as horizon $h \to \infty$?
- A) Forecasts grow without bound
- B) Forecasts oscillate forever
- C) Forecasts converge to the unconditional mean $\mu$
- D) Forecasts become more accurate
Question 9: AR(1) Variance
For the AR(1) process $X_t = \phi X_{t-1} + \varepsilon_t$ with $\text{Var}(\varepsilon_t) = \sigma^2$, what is $\text{Var}(X_t)$?
- A) $\sigma^2$
- B) $\sigma^2 / (1 - \phi)$
- C) $\sigma^2 / (1 - \phi^2)$
- D) $\sigma^2 (1 + \phi^2)$
Question 10: MA(1) Autocorrelation
For an MA(1) process $X_t = \varepsilon_t + \theta \varepsilon_{t-1}$, what is $\rho(2)$?
- A) $\theta / (1 + \theta^2)$
- B) $\theta^2 / (1 + \theta^2)$
- C) $\theta^2$
- D) 0
Question 11: AR(1) Mean
For the AR(1) process $X_t = c + \phi X_{t-1} + \varepsilon_t$ with $c = 2$ and $\phi = 0.6$, what is $E[X_t]$?
- A) 2
- B) 3.33
- C) 5
- D) 0.8
Question 12: Wold's Theorem
According to Wold's decomposition theorem, any stationary process can be written as:
- A) A finite AR process
- B) An infinite MA process plus a deterministic component
- C) A random walk
- D) White noise only
Question 13: AR(2) Stationarity
For AR(2): $X_t = \phi_1 X_{t-1} + \phi_2 X_{t-2} + \varepsilon_t$, which conditions ensure stationarity?
- A) $|\phi_1| < 1$ and $|\phi_2| < 1$
- B) $\phi_1 + \phi_2 < 1$, $\phi_2 - \phi_1 < 1$, $|\phi_2| < 1$
- C) $\phi_1^2 + \phi_2^2 < 1$
- D) $|\phi_1 + \phi_2| < 1$
Question 14: Yule-Walker Equations
The Yule-Walker equations are used to:
- A) Test for stationarity
- B) Estimate AR parameters from autocorrelations
- C) Estimate MA parameters
- D) Compute forecasts
Question 15: PACF Interpretation
The PACF at lag $k$ measures:
- A) Total correlation between $X_t$ and $X_{t-k}$
- B) Correlation between $X_t$ and $X_{t-k}$ after removing effects of intermediate lags
- C) The MA coefficient at lag $k$
- D) The variance at lag $k$
Question 16: Parsimony Principle
The parsimony principle in ARMA modeling suggests:
- A) Always use the highest order model
- B) Choose the simplest adequate model
- C) Only use AR models
- D) Always include seasonal terms
Question 17: Characteristic Roots
For AR(1) with $\phi = 0.9$, the characteristic root is:
- A) 0.9
- B) 1/0.9 ≈ 1.11
- C) -0.9
- D) 0.81
Question 18: Estimation Methods
Which estimation method is generally preferred for ARMA models?
- A) Ordinary Least Squares (OLS)
- B) Maximum Likelihood Estimation (MLE)
- C) Method of Moments only
- D) Simple averaging
Question 19: Forecast Error Variance
For a stationary ARMA model, as forecast horizon increases, the forecast error variance:
- A) Decreases to zero
- B) Increases without bound
- C) Converges to the unconditional variance
- D) Remains constant
Question 20: Box-Jenkins Methodology
The correct order of steps in Box-Jenkins methodology is:
- A) Estimation → Identification → Diagnostics
- B) Identification → Estimation → Diagnostics
- C) Diagnostics → Identification → Estimation
- D) Estimation → Diagnostics → Identification
Chapter 3: ARIMA Models
Question 1: Order of Integration
A time series requires two differences to become stationary. What is its order of integration?
- A) $I(0)$
- B) $I(1)$
- C) $I(2)$
- D) Cannot be determined
Question 2: Random Walk Variance
For a random walk $Y_t = Y_{t-1} + \varepsilon_t$ with $\text{Var}(\varepsilon_t) = \sigma^2$, what is $\text{Var}(Y_t)$?
- A) $\sigma^2$
- B) $t \cdot \sigma^2$
- C) $\sigma^2 / t$
- D) $\sigma^2 / (1-\phi^2)$
Question 3: ADF Test Hypothesis
In the Augmented Dickey-Fuller test, what is the null hypothesis?
- A) The series is stationary
- B) The series has a unit root
- C) The series has no autocorrelation
- D) The series is normally distributed
Question 4: ARIMA Notation
What does ARIMA(2,1,1) represent?
- A) AR(2) on differenced data with MA(1) errors
- B) AR(1) with 2 differences and MA(1)
- C) MA(2) with 1 difference and AR(1)
- D) 2 lags, 1 trend, 1 seasonal component
Question 5: Second Difference
What is $(1-L)^2 Y_t$ expanded?
- A) $Y_t - Y_{t-1}$
- B) $Y_t - 2Y_{t-1} + Y_{t-2}$
- C) $Y_t + 2Y_{t-1} + Y_{t-2}$
- D) $Y_t - Y_{t-2}$
Question 6: KPSS vs ADF
How does the KPSS test differ from the ADF test?
- A) KPSS tests for seasonality, ADF tests for trends
- B) KPSS has stationarity as null, ADF has unit root as null
- C) KPSS is more powerful than ADF
- D) There is no difference
Question 7: Overdifferencing
If $Y_t \sim I(1)$ and we compute $\Delta^2 Y_t$, what happens?
- A) We get a better stationary series
- B) We introduce artificial negative autocorrelation
- C) The variance decreases
- D) Nothing changes
Question 8: Forecast Variance
For ARIMA(0,1,0) (random walk), how does forecast variance behave as horizon $h$ increases?
- A) Stays constant
- B) Decreases to zero
- C) Grows linearly with $h$
- D) Converges to a finite limit
Question 9: Deterministic vs Stochastic Trend
What distinguishes a stochastic trend from a deterministic trend?
- A) Stochastic trends can be removed by detrending
- B) Deterministic trends have permanent shock effects
- C) Stochastic trends have permanent shock effects
- D) They are mathematically equivalent
Question 10: ADF Critical Values
In the ADF test, you reject H0 (unit root) when:
- A) Test statistic > critical value
- B) Test statistic < critical value (more negative)
- C) p-value > 0.05
- D) Test statistic equals zero
Question 11: IMA(1,1) Model
ARIMA(0,1,1) is also known as:
- A) Random walk
- B) Simple exponential smoothing
- C) Holt's method
- D) Pure AR model
Question 12: Constant in ARIMA
In ARIMA(1,1,0) with constant $c$, what does $c$ represent?
- A) The mean of $Y_t$
- B) The drift (average change)
- C) The variance
- D) The AR coefficient
Question 13: Auto-ARIMA
What does auto_arima primarily use to select the best model?
- A) Only visual inspection
- B) Information criteria (AIC/BIC)
- C) Only residual analysis
- D) Random selection
Question 14: Box-Jenkins Step 1
What is the first step in Box-Jenkins methodology?
- A) Estimate parameters
- B) Identify model order (p,d,q)
- C) Diagnose residuals
- D) Generate forecasts
Question 15: ACF of I(1) Series
What is characteristic of the ACF of an I(1) series?
- A) Cuts off sharply at lag 1
- B) Decays very slowly
- C) Shows no significant lags
- D) Alternates between positive and negative
Question 16: Ljung-Box on ARIMA Residuals
If Ljung-Box test p-value is 0.02 on ARIMA residuals, what should you conclude?
- A) The model is adequate
- B) The model is inadequate, consider re-specifying
- C) The series is stationary
- D) The series has a unit root
Question 17: MLE for ARIMA
What is the standard estimation method for ARIMA parameters?
- A) Ordinary Least Squares
- B) Maximum Likelihood Estimation
- C) Method of Moments
- D) Bayesian estimation only
Question 18: Typical $d$ for Economic Data
For most macroeconomic time series (GDP, prices), what is the typical order of integration?
- A) $I(0)$ - stationary
- B) $I(1)$ - unit root
- C) $I(2)$ or higher
- D) $I(-1)$ - overdifferenced
Question 19: Conflicting Test Results
If ADF does not reject but KPSS rejects, what does this suggest?
- A) The series is definitely stationary
- B) The series is definitely non-stationary
- C) Both tests agree the series has a unit root
- D) Results are inconclusive, need more analysis
Question 20: ARIMA Limitations
Which is NOT a limitation of ARIMA models?
- A) Assumes linearity
- B) Cannot capture volatility clustering
- C) Can handle non-stationary data
- D) Assumes constant variance
Chapter 4: SARIMA Models
Question 1: Seasonal Differencing
For monthly data with annual seasonality, what does $(1-L^{12})Y_t$ compute?
- A) $Y_t - Y_{t-1}$
- B) $Y_t - Y_{t-12}$
- C) Average of 12 months
- D) $Y_t / Y_{t-12}$
Question 2: SARIMA Notation
In SARIMA$(1,1,1) \times (1,1,1)_{12}$, what does the subscript 12 represent?
- A) Number of parameters
- B) Seasonal period
- C) Maximum lag order
- D) Sample size requirement
Question 3: Airline Model
The classic "airline model" SARIMA$(0,1,1) \times (0,1,1)_{12}$ has how many parameters (excluding $\sigma^2$)?
- A) 2
- B) 4
- C) 6
- D) 12
Question 4: Seasonal ACF Pattern
For monthly data with strong annual seasonality, where would you expect significant ACF spikes?
- A) Only at lag 1
- B) At lags 12, 24, 36, ...
- C) Randomly distributed
- D) No spikes at all
Question 5: Multiplicative Structure
In SARIMA, "multiplicative structure" means:
- A) Seasonal amplitude grows proportionally with level
- B) Regular and seasonal polynomials are multiplied
- C) Data must be multiplied by seasonal factors
- D) Model is estimated using multiplication
Question 6: Full Differencing
When should you apply both regular ($d=1$) and seasonal ($D=1$) differencing?
- A) When data has only a trend
- B) When data has only seasonality
- C) When data has both trend and seasonal non-stationarity
- D) Never - they cancel each other
Question 7: Log Transformation
When should you apply log transformation before fitting SARIMA?
- A) When data contains zeros or negative values
- B) When seasonal fluctuations grow with the level
- C) Always, regardless of data characteristics
- D) Never for seasonal data
Question 8: Seasonal Period
What seasonal period ($s$) would you use for quarterly GDP data with annual patterns?
- A) $s = 1$
- B) $s = 4$
- C) $s = 12$
- D) $s = 52$
Question 9: Expansion Formula
What is $(1-L)(1-L^{12})$ expanded?
- A) $1 - L - L^{12}$
- B) $1 - L - L^{12} + L^{13}$
- C) $1 - 2L + L^{12}$
- D) $1 - L^{13}$
Question 10: Deterministic vs Stochastic
What distinguishes stochastic seasonality (SARIMA) from deterministic seasonality (seasonal dummies)?
- A) Stochastic seasonality requires more data
- B) Stochastic seasonality evolves over time
- C) Deterministic seasonality is more flexible
- D) There is no difference
Chapter 5: Volatility Models ARCH/GARCH
Note: Quiz content is being updated. Current quizzes cover VAR topics.
Question 1: VAR Definition
In a VAR(2) model with 3 variables, how many coefficient matrices $\mathbf{A}_i$ are there?
- A) 2
- B) 3
- C) 6
- D) 9
Question 2: Number of Parameters
A VAR(2) with $K=3$ variables (including constants) has how many parameters per equation?
- A) 3
- B) 6
- C) 7
- D) 9
Question 3: Granger Causality
"$X$ Granger-causes $Y$" means:
- A) $X$ is the economic cause of $Y$
- B) Past $X$ helps predict future $Y$
- C) $X$ and $Y$ are contemporaneously correlated
- D) $X$ always increases when $Y$ increases
Question 4: Granger Causality Test
To test if $Y_2$ Granger-causes $Y_1$ in a VAR(p), we test:
- A) All coefficients in the $Y_1$ equation equal zero
- B) Coefficients on lagged $Y_2$ in the $Y_1$ equation equal zero
- C) Coefficients on lagged $Y_1$ in the $Y_2$ equation equal zero
- D) The error covariance equals zero
Question 5: VAR Stability
A VAR(1) model is stable (stationary) if:
- A) All diagonal elements of $\mathbf{A}_1$ are less than 1
- B) The determinant of $\mathbf{A}_1$ is less than 1
- C) All eigenvalues of $\mathbf{A}_1$ are less than 1 in absolute value
- D) The trace of $\mathbf{A}_1$ equals zero
Question 6: Impulse Response Functions
An impulse response function shows:
- A) The correlation between two variables
- B) The effect of a shock to one variable on all variables over time
- C) The forecast accuracy of the model
- D) The p-values of coefficient tests
Question 7: Lag Order Selection
Which criterion typically selects the most parsimonious VAR model?
- A) AIC (Akaike Information Criterion)
- B) BIC (Bayesian Information Criterion)
- C) FPE (Final Prediction Error)
- D) Adjusted $R^2$
Question 8: FEVD
Forecast Error Variance Decomposition (FEVD) tells us:
- A) The correlation between variables
- B) What proportion of forecast error variance comes from each shock
- C) The optimal forecast horizon
- D) Which variables to include in the model
Question 9: Cholesky Ordering
Cholesky ordering in IRF analysis assumes:
- A) All variables are equally important
- B) Variables ordered first affect later variables contemporaneously, not vice versa
- C) Shocks are uncorrelated
- D) No restrictions are needed
Question 10: Cointegration
If variables are I(1) and cointegrated, you should use:
- A) VAR in levels
- B) VAR in first differences
- C) Vector Error Correction Model (VECM)
- D) Univariate ARIMA models
Question 11: VAR Residuals
In a well-specified VAR, residuals should be:
- A) Autocorrelated but homoskedastic
- B) Serially uncorrelated (white noise)
- C) Perfectly normally distributed
- D) Zero for all observations
Question 12: Structural VAR
The main difference between SVAR and reduced-form VAR is:
- A) SVAR uses more lags
- B) SVAR identifies structural shocks with economic interpretation
- C) SVAR requires more data
- D) SVAR cannot forecast
Question 13: Granger vs True Causality
Finding that X Granger-causes Y means:
- A) X definitely causes Y economically
- B) X has predictive power for Y, but may not be true causation
- C) Y causes X
- D) X and Y are unrelated
Question 14: VAR Estimation
VAR models can be estimated by:
- A) OLS on each equation separately
- B) Only maximum likelihood
- C) Only Bayesian methods
- D) Weighted least squares only
Question 15: IRF Convergence
In a stable VAR, impulse responses as $h \to \infty$:
- A) Explode to infinity
- B) Converge to zero
- C) Oscillate forever
- D) Stay constant
Question 16: Bidirectional Causality
If both "X Granger-causes Y" and "Y Granger-causes X", this is called:
- A) No causality
- B) Unidirectional causality
- C) Bidirectional (feedback) causality
- D) Instantaneous causality
Question 17: Companion Form
The companion matrix of a VAR(p) is used to:
- A) Reduce the number of parameters
- B) Convert VAR(p) to VAR(1) form for stability analysis
- C) Estimate the model faster
- D) Test for Granger causality
Question 18: FEVD Over Horizons
In FEVD, as the horizon $h$ increases:
- A) Own shocks always dominate
- B) The proportions converge to long-run values
- C) All shocks contribute equally
- D) FEVD becomes undefined
Question 19: Instantaneous Causality
Instantaneous causality tests whether:
- A) Lagged X predicts Y
- B) Shocks to X and Y are correlated within the same period
- C) X and Y share a common trend
- D) The VAR is stable
Question 20: Pre-Estimation Checks
Before estimating a VAR, you should always check:
- A) That all variables are I(2)
- B) The stationarity of each variable
- C) That variables are perfectly correlated
- D) That the sample size is exactly 100
Chapter 6: VAR Models & Granger Causality
Question 1: VAR Model Structure
In a bivariate VAR(1), the equation for \\( Y_{1t} \\) includes:
- A) Only its own lagged values \\( Y_{1,t-1} \\)
- B) Only the lagged values of \\( Y_{2,t-1} \\)
- C) Lagged values of all variables: \\( Y_{1,t-1} \\) and \\( Y_{2,t-1} \\)
- D) Current and lagged values of all variables
Question 2: Curse of Dimensionality
A VAR(4) model with K=5 variables has how many parameters?
- A) 20
- B) 100
- C) 105
- D) 125
Question 3: Stability Condition
For a general VAR(p), the stability condition requires that all roots of \\( \\det(\\mathbf{I}_K - \\mathbf{A}_1 z - \\cdots - \\mathbf{A}_p z^p) = 0 \\) lie:
- A) Inside the unit circle
- B) On the unit circle
- C) Outside the unit circle
- D) At the origin
Question 4: Granger Causality Test
The Granger causality test uses which type of test statistic?
- A) Durbin-Watson statistic
- B) Wald/F-test comparing restricted and unrestricted models
- C) Augmented Dickey-Fuller statistic
- D) Jarque-Bera statistic
Question 5: Autocovariance Matrix Property
For a weakly stationary multivariate time series, the autocovariance matrix satisfies:
- A) \\( \\boldsymbol{\\Gamma}(-h) = \\boldsymbol{\\Gamma}(h) \\)
- B) \\( \\boldsymbol{\\Gamma}(-h) = \\boldsymbol{\\Gamma}(h)' \\) (transpose)
- C) \\( \\boldsymbol{\\Gamma}(-h) = -\\boldsymbol{\\Gamma}(h) \\)
- D) \\( \\boldsymbol{\\Gamma}(-h) = \\boldsymbol{\\Gamma}(h)^{-1} \\)
Question 6: MA(infinity) Representation
The impulse response matrices \\( \\boldsymbol{\\Phi}_h \\) in the MA(\\( \\infty \\)) representation of a VAR(1) are given by:
- A) \\( \\boldsymbol{\\Phi}_h = h \\cdot \\mathbf{A} \\)
- B) \\( \\boldsymbol{\\Phi}_h = \\mathbf{A}^h \\)
- C) \\( \\boldsymbol{\\Phi}_h = \\mathbf{A}^{-h} \\)
- D) \\( \\boldsymbol{\\Phi}_h = \\boldsymbol{\\Sigma}^h \\)
Question 7: Orthogonalized IRFs
The purpose of orthogonalizing IRFs using the Cholesky decomposition is to:
- A) Increase the number of parameters in the model
- B) Make the model stationary
- C) Produce uncorrelated shocks so that the effect of each shock can be isolated
- D) Remove the constant term from the VAR
Question 8: FEVD Interpretation
If FEVD shows that 70% of the forecast error variance of unemployment at horizon 8 is due to GDP shocks, this means:
- A) GDP directly causes 70% of unemployment
- B) 70% of the uncertainty in forecasting unemployment 8 periods ahead is attributable to GDP shocks
- C) Unemployment will increase by 70% after 8 periods
- D) The correlation between GDP and unemployment is 0.70
Question 9: Granger Causality Numerical Example
Given T=100, RSS_U=45.2, RSS_R=52.8, and p=2 lags, the F-statistic for the Granger causality test is approximately:
- A) 3.09
- B) 5.42
- C) 7.98
- D) 12.35
Question 10: SVAR Identification
Which of the following is NOT a common identification scheme for Structural VAR models?
- A) Short-run (Cholesky) restrictions
- B) Long-run (Blanchard-Quah) restrictions
- C) OLS residual minimization
- D) Sign restrictions
Question 11: Companion Form Purpose
The companion form is useful because it:
- A) Reduces the number of parameters to estimate
- B) Converts any VAR(p) into a VAR(1), simplifying stationarity analysis, forecasting, and IRF computation
- C) Eliminates the need for the Cholesky decomposition
- D) Makes the error covariance matrix diagonal
Question 12: Diagnostic Tests
If the Portmanteau test for a VAR model's residuals is rejected, the most appropriate action is to:
- A) Apply a logarithmic transformation to the data
- B) Increase the lag order p or add additional variables
- C) Switch to a univariate ARIMA model
- D) Remove the constant term from the model
Question 13: Cross-Covariance Properties
Unlike the univariate case, the cross-covariance function satisfies:
- A) \\( \\gamma_{ij}(h) = \\gamma_{ij}(-h) \\) always
- B) \\( \\gamma_{ij}(h) = \\gamma_{ji}(-h) \\), but generally \\( \\gamma_{ij}(h) \ eq \\gamma_{ij}(-h) \\)
- C) \\( \\gamma_{ij}(h) = 0 \\) for all \\( h \ eq 0 \\)
- D) \\( \\gamma_{ij}(h) = \\gamma_{ij}(h+1) \\) for all h
Question 14: Forecast Error MSE
For a stable VAR(1), as the forecast horizon h increases, the mean squared error matrix converges to:
- A) The zero matrix
- B) The innovation covariance matrix \\( \\boldsymbol{\\Sigma} \\)
- C) The unconditional variance \\( \\boldsymbol{\\Gamma}(0) \\)
- D) Infinity
Question 15: Restricted VAR
The likelihood ratio test for restrictions in a VAR model follows which distribution under the null?
- A) Normal distribution
- B) F-distribution
- C) Chi-squared distribution with r degrees of freedom
- D) Student's t-distribution
Question 16: Eigenvalue Example
For the matrix \\( \\mathbf{A} = \\begin{pmatrix} 0.7 & 0.2 \\\\ -0.1 & 0.6 \\end{pmatrix} \\), the eigenvalues are \\( \\lambda = 0.65 \\pm 0.132i \\). The modulus \\( |\\lambda| \\) equals:
- A) 0.782
- B) 0.663
- C) 1.30
- D) 0.44
Question 17: Granger Causality Pitfalls
Which of the following is a common pitfall when applying Granger causality tests?
- A) Using too many variables in the VAR
- B) A third omitted variable Z may cause both X and Y, creating a spurious Granger causal relationship
- C) The test cannot be applied to quarterly data
- D) Granger causality always implies true economic causation
Question 18: Bootstrap Confidence Intervals
The bootstrap procedure for IRF confidence intervals involves:
- A) Estimating the model only once and using analytical formulas
- B) Resampling residuals with replacement, re-estimating the VAR, and computing IRFs repeatedly
- C) Increasing the sample size by adding artificial data points
- D) Removing outliers from the dataset and re-estimating
Question 19: Diebold-Mariano Test
The Diebold-Mariano test is used to:
- A) Test for Granger causality between two variables
- B) Test for the presence of unit roots
- C) Test whether one forecasting model is significantly more accurate than another
- D) Test for normality of VAR residuals
Question 20: Cumulative IRF
The long-run multiplier (cumulative IRF as \\( H \\to \\infty \\)) for a stable VAR is given by:
- A) \\( \\boldsymbol{\\Psi}_\\infty = \\mathbf{A}_1 + \\mathbf{A}_2 + \\cdots + \\mathbf{A}_p \\)
- B) \\( \\boldsymbol{\\Psi}_\\infty = (\\mathbf{I}_K - \\mathbf{A}_1 - \\mathbf{A}_2 - \\cdots - \\mathbf{A}_p)^{-1} \\)
- C) \\( \\boldsymbol{\\Psi}_\\infty = \\mathbf{I}_K \\)
- D) \\( \\boldsymbol{\\Psi}_\\infty = \\mathbf{0} \\)
Chapter 7: Cointegration & VECM
Quiz content coming soon...
Chapter 8: Modern Extensions (ARFIMA, RF, LSTM)
Question 1: Hurst Exponent Interpretation
A time series has Hurst exponent $H = 0.8$. What does this indicate?
- A) The series is a pure random walk
- B) The series has long memory and is persistent (trend-following)
- C) The series is anti-persistent (mean-reverting)
- D) The series is stationary I(0)
Question 2: ARFIMA Parameter d
In the ARFIMA(p, d, q) model, the parameter $d$ can take values:
- A) Only integer values (0, 1, 2, ...)
- B) Only $d = 0$ or $d = 1$
- C) Any real value, including fractional
- D) Only negative values
Question 3: Long Memory in Finance
In which financial series is long memory most commonly documented?
- A) Stock prices
- B) Daily returns
- C) Volatility (squared returns)
- D) Trading volume
Question 4: Feature Engineering for ML
To apply Random Forest to time series forecasting, we must create:
- A) Dummy variables for each observation
- B) Lag features and rolling statistics
- C) Fourier transforms of the series
- D) Only the first difference of the series
Question 5: Time Series Cross-Validation
Why can't we use standard k-fold cross-validation for time series?
- A) It's too slow for long series
- B) It violates temporal order and causes data leakage
- C) It only works for classification
- D) It requires too much data
Question 6: LSTM Advantage
What is the main advantage of LSTM over simple RNNs?
- A) It's faster to train
- B) It solves the vanishing/exploding gradient problem
- C) It requires less data
- D) It's easier to interpret
Question 7: Data Normalization
Before training an LSTM, data should be:
- A) Log-transformed
- B) Normalized/scaled to [0,1] or [-1,1]
- C) Differenced twice
- D) Converted to integers
Question 8: LSTM Hyperparameters
Which is NOT a typical LSTM hyperparameter?
- A) Number of units (neurons) per layer
- B) Input sequence length
- C) Learning rate
- D) Differencing parameter $d$
Question 9: Feature Importance
Feature importance in Random Forest for time series helps us:
- A) Eliminate all low-importance variables
- B) Identify which lags and features are most predictive
- C) Determine Granger causality
- D) Calculate confidence intervals
Question 10: Model Selection
When comparing ARFIMA, Random Forest, and LSTM for forecasting:
- A) LSTM always wins because it's deep learning
- B) ARFIMA is always best for financial data
- C) The best model depends on data characteristics and requirements
- D) Random Forest can't be used for time series
Question 11: ACF Decay Pattern
Long memory processes have ACF that decays:
- A) Exponentially fast
- B) Hyperbolically (slowly)
- C) Linearly
- D) Immediately to zero
Question 12: ARFIMA Stationarity
An ARFIMA process with $d = 0.3$ is:
- A) Non-stationary
- B) Stationary with long memory
- C) Stationary with short memory
- D) A unit root process
Question 13: Random Forest Ensemble
Random Forest reduces overfitting compared to a single decision tree by:
- A) Using deeper trees
- B) Averaging predictions from many trees trained on bootstrap samples
- C) Using fewer features
- D) Training on the full dataset only
Question 14: Data Leakage
Which is an example of data leakage in time series ML?
- A) Using lag features
- B) Scaling data using statistics from the entire dataset (train+test)
- C) Using rolling window features
- D) Splitting data chronologically
Question 15: LSTM Gates
Which gate in LSTM decides what information to discard from the cell state?
- A) Input gate
- B) Forget gate
- C) Output gate
- D) Update gate
Question 16: Dropout Regularization
Dropout in LSTM helps to:
- A) Speed up training
- B) Prevent overfitting by randomly zeroing neurons
- C) Increase model capacity
- D) Handle missing data
Question 17: Early Stopping
Early stopping monitors validation loss to:
- A) Speed up computation
- B) Stop training when the model starts overfitting
- C) Select the best features
- D) Adjust the learning rate
Question 18: Sequence Length
Choosing the LSTM input sequence length (lookback) involves:
- A) Always using the maximum available data
- B) Balancing memory capacity vs. computational cost
- C) Using exactly 30 time steps
- D) Matching the seasonal period only
Question 19: R/S Analysis
The R/S (Rescaled Range) statistic is used to estimate:
- A) The ARIMA order
- B) The Hurst exponent
- C) The GARCH parameters
- D) The seasonal period
Question 20: Direction Accuracy
Direction accuracy measures:
- A) The magnitude of forecast errors
- B) How often the model correctly predicts up/down movements
- C) The correlation between forecast and actual
- D) The variance of forecast errors
Chapter 9: Prophet & TBATS for Multiple Seasonality
Question 1: Multiple Seasonality Challenge
Why can't standard SARIMA handle hourly electricity demand data?
- A) SARIMA can only handle monthly data
- B) SARIMA allows only one seasonal period ($m$ parameter)
- C) SARIMA doesn't support trend components
- D) SARIMA requires normally distributed data
Question 2: TBATS Acronym
What does TBATS stand for?
- A) Trend, Baseline, ARMA, Transform, Seasonal
- B) Trigonometric, Box-Cox, ARMA, Trend, Seasonal
- C) Time-Based Automatic Time Series
- D) Temporal Bayesian Adaptive Trend System
Question 3: Fourier Terms
In TBATS, increasing the number of Fourier harmonics ($K$) for a seasonal pattern:
- A) Always improves forecast accuracy
- B) Allows more flexible (complex) seasonal shapes
- C) Reduces the model complexity
- D) Eliminates the need for Box-Cox transformation
Question 4: Prophet Decomposition
Prophet decomposes a time series into which components?
- A) AR, MA, and seasonal components
- B) Trend, seasonality, holidays, and error
- C) Mean, variance, and autocorrelation
- D) Level, slope, and curvature
Question 5: Prophet vs TBATS
When would you choose Prophet over TBATS?
- A) When you need automatic model selection
- B) When you have known holidays and changepoints to incorporate
- C) When you need the most parsimonious model
- D) When your data has no trend
Question 6: Seasonality Mode
For retail sales where December sales are 3x the monthly average, which seasonality mode is more appropriate in Prophet?
- A) Additive seasonality
- B) Multiplicative seasonality
- C) Both work equally well
- D) Neither---use ARIMA instead
Question 7: Prophet Changepoints
In Prophet, changepoints allow the model to:
- A) Change the seasonal period automatically
- B) Adjust the trend slope at specific points in time
- C) Switch between additive and multiplicative modes
- D) Detect and remove outliers
Question 8: Model Selection
You have daily call center data with weekly seasonality only. Which model is most appropriate?
- A) TBATS (designed for multiple seasonality)
- B) Prophet (handles any seasonality well)
- C) Standard SARIMA (simpler and sufficient)
- D) LSTM neural network (most flexible)
Question 9: Prophet Uncertainty
Prophet generates prediction intervals by:
- A) Assuming normally distributed residuals
- B) Sampling from the posterior distribution of parameters
- C) Using bootstrap resampling of historical errors
- D) Applying a fixed multiplier to point forecasts
Question 10: Energy Demand Forecasting
For hourly energy demand with daily, weekly, and annual patterns plus holidays, which approach is best?
- A) SARIMA with $m=24$
- B) TBATS with three seasonal periods
- C) Prophet with custom holidays
- D) Either TBATS or Prophet, depending on holiday importance
Question 11: Box-Cox Transformation
In TBATS, the Box-Cox transformation is used to:
- A) Remove the trend
- B) Stabilize variance
- C) Remove seasonality
- D) Speed up computation
Question 12: Prophet Default Trend
In Prophet, the default trend type is:
- A) Logistic growth
- B) Piecewise linear
- C) Exponential
- D) Polynomial
Question 13: Fourier Parameters
For weekly seasonality ($m=7$) with $K=3$ Fourier terms, how many parameters are added?
- A) 3
- B) 6
- C) 7
- D) 14
Question 14: Missing Data
Which statement about missing data handling is correct?
- A) SARIMA handles missing data better than Prophet
- B) Prophet handles missing data and irregular timestamps gracefully
- C) TBATS automatically imputes missing values
- D) All methods require complete data
Question 15: Holiday Effects
Prophet's holiday effects are modeled as:
- A) Multiplicative adjustments to trend
- B) Additive effects with optional windows
- C) Separate ARIMA components
- D) Fourier terms at holiday frequencies
Question 16: Overfitting with Fourier
A symptom of overfitting with too many Fourier terms is:
- A) Smooth, realistic seasonal patterns
- B) Jagged seasonality and poor out-of-sample performance
- C) Faster model fitting
- D) Better handling of outliers
Question 17: TBATS Selection
TBATS automatically selects the number of Fourier terms using:
- A) Cross-validation
- B) AIC (Akaike Information Criterion)
- C) Visual inspection
- D) Fixed rules based on frequency
Question 18: Prophet Cross-Validation
Prophet's built-in cross_validation function uses:
- A) Standard k-fold CV
- B) Rolling origin (time series) CV
- C) Leave-one-out CV
- D) Random holdout
Question 19: Changepoint Prior
In Prophet, increasing changepoint_prior_scale makes the trend:
- A) Smoother and more stable
- B) More flexible and responsive to changes
- C) Linear without changepoints
- D) Logistic instead of linear
Question 20: Model Interpretability
A key advantage of Prophet over black-box ML models is:
- A) Always higher accuracy
- B) Interpretable component decomposition (trend, seasonality, holidays)
- C) Faster training time
- D) No hyperparameters to tune
Chapter 10: Review - From Data to Forecast
Question 1: Analysis Workflow
What is the correct order of the time series analysis workflow?
- A) Model fitting, Data exploration, Diagnostics, Forecasting
- B) Data exploration, Stationarity testing, Model selection, Diagnostics, Forecasting
- C) Forecasting, Data exploration, Model fitting, Diagnostics
- D) Diagnostics, Model selection, Data exploration, Forecasting
Question 2: RMSE vs MAE
If RMSE is much larger than MAE, this suggests:
- A) The model is overfitting
- B) There are some large outlying errors
- C) The model is underfitting
- D) The data is stationary
Question 3: MAPE Limitation
MAPE is problematic when:
- A) The data has a trend
- B) Actual values are close to or equal to zero
- C) The forecast horizon is long
- D) Multiple models are being compared
Question 4: Volatility Clustering
If S&P 500 returns show periods of high volatility followed by high volatility, you should consider:
- A) ARIMA with higher order
- B) Exponential smoothing
- C) GARCH model for variance
- D) Seasonal differencing
Question 5: ADF vs KPSS
If ADF test fails to reject and KPSS rejects, the series is likely:
- A) Stationary
- B) Non-stationary
- C) Trend-stationary
- D) White noise
Question 6: Multiplicative Seasonality
Air passengers data shows increasing seasonal amplitude over time. Which decomposition is appropriate?
- A) Additive: $Y_t = T_t + S_t + R_t$
- B) Multiplicative: $Y_t = T_t \times S_t \times R_t$
- C) Both work equally well
- D) Neither---use differencing instead
Question 7: SARIMA for Monthly Data
For monthly data with yearly seasonality, the seasonal period $m$ in SARIMA is:
- A) 4
- B) 7
- C) 12
- D) 52
Question 8: Model Comparison
When comparing SARIMA and Prophet forecasts, which metric is scale-independent?
- A) RMSE
- B) MAE
- C) MAPE
- D) MSE
Question 9: Structural Breaks
US Retail Sales experienced a structural break during COVID-19. Prophet handles this via:
- A) Automatic differencing
- B) Changepoint detection in the trend
- C) Seasonal adjustment
- D) GARCH modeling
Question 10: Ljung-Box Test
After fitting an ARIMA model, the Ljung-Box test on residuals tests for:
- A) Normality
- B) Remaining autocorrelation
- C) Heteroscedasticity
- D) Stationarity
Question 11: ACF/PACF Interpretation
ACF decays exponentially and PACF cuts off after lag 2. This suggests:
- A) MA(2)
- B) AR(2)
- C) ARMA(2,2)
- D) Random walk
Question 12: Log Returns
For S&P 500, we use log returns $r_t = \ln(P_t/P_{t-1})$ instead of prices because:
- A) Prices are always stationary
- B) Returns are approximately stationary; prices are not
- C) Log returns are easier to compute
- D) Returns have stronger autocorrelation
Question 13: AIC vs BIC
Compared to AIC, BIC typically selects:
- A) More complex models
- B) Simpler (more parsimonious) models
- C) Identical models to AIC
- D) Models with better in-sample fit
Question 14: Cross-Validation
For time series cross-validation, we use:
- A) Random k-fold CV
- B) Leave-one-out CV
- C) Rolling origin (expanding window) CV
- D) Stratified CV
Question 15: Multiple Seasonality
Hourly data with daily, weekly, and yearly patterns is best handled by:
- A) SARIMA with $m=24$
- B) Simple exponential smoothing
- C) TBATS or Prophet
- D) ARIMA with differencing
Question 16: GARCH Persistence
In GARCH(1,1), high volatility persistence means $\alpha + \beta$ is:
- A) Close to 0
- B) Close to 1
- C) Greater than 1
- D) Negative
Question 17: Forecast Uncertainty
As forecast horizon increases, prediction intervals typically:
- A) Narrow
- B) Stay constant
- C) Widen
- D) Oscillate
Question 18: Direction Accuracy
A model with 45% direction accuracy for stock returns is:
- A) Better than random guessing
- B) Worse than random guessing
- C) Optimal for trading
- D) Statistically significant
Question 19: Model Selection Principle
The principle of parsimony suggests choosing:
- A) The model with the best in-sample fit
- B) The simplest model that adequately fits the data
- C) The most complex model available
- D) The model with the most parameters
Question 20: Complete Workflow Summary
Which statement best summarizes the time series analysis workflow?
- A) Fit the most complex model first, then simplify
- B) Start with visualization, test stationarity, fit models, validate on held-out data
- C) Use the same model for all datasets
- D) Skip diagnostics if in-sample fit is good
Chapter 11: LLMs and Foundation Models for Time Series
Question 1: Self-Attention Mechanism
What is the formula for the Scaled Dot-Product Attention in the Transformer (Vaswani et al., 2017)?
- A) $\text{Attention}(Q,K,V) = \text{softmax}(QK^T) \cdot V$
- B) $\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^T}{\sqrt{d_k}}\right) V$
- C) $\text{Attention}(Q,K,V) = \sigma(QK^T + b) \cdot V$
- D) $\text{Attention}(Q,K,V) = Q \cdot \text{softmax}(K^T V)$
Question 2: Architectural Paradigms
What type of Transformer architecture does Chronos (Amazon) use?
- A) Decoder-only (GPT-style)
- B) Encoder-decoder (T5)
- C) Encoder-only (BERT-style)
- D) Recurrent network (LSTM)
Question 3: Patching Tokenization
If a time series has 512 time steps and we use patch size $P = 32$, how many tokens result?
- A) 512 tokens
- B) 16 tokens
- C) 32 tokens
- D) 256 tokens
Question 4: Zero-Shot vs Fine-Tuning
When is zero-shot forecasting preferable to fine-tuning?
- A) When you have millions of observations from the target domain
- B) When data is limited, rapid prototyping is needed, or cross-domain deployment is required
- C) When causal interpretability is needed
- D) When the series has fewer than 10 observations
Question 5: Chronos — Tokenization
How does Chronos convert continuous time series values into discrete tokens?
- A) Direct embedding of raw values
- B) Mean-scale normalization followed by quantization into 4096 bins via normal CDF
- C) Differencing the series and one-hot encoding
- D) Converting the series to text and using the standard GPT tokenizer
Question 6: LoRA (Low-Rank Adaptation)
What percentage of parameters are trainable in LoRA fine-tuning?
- A) 100% — all parameters are updated
- B) 0.1–1% — only small low-rank matrices ($W' = W + BA$, $r \ll d$)
- C) 50% — half of the network layers
- D) 10–20% — only the last layers
Question 7: Moirai — Universal Design
What makes Moirai (Salesforce) unique among time series foundation models?
- A) It is the only model with API access
- B) It handles any number of variates, any frequency, and any prediction length
- C) It has the most parameters (710M)
- D) It uses exclusively synthetic data for pre-training
Question 8: Foundation Model Limitations
Which of the following is NOT a major limitation of current time series foundation models?
- A) Most are univariate only
- B) They cannot produce probabilistic forecasts
- C) They lack causal reasoning
- D) They are black boxes that are hard to interpret
Question 9: Lag-Llama — Output Distribution
What output distribution does Lag-Llama use instead of discrete token probabilities?
- A) Normal distribution $N(\mu, \sigma^2)$
- B) Student-t distribution with parameters $(\mu, \sigma, \nu)$
- C) Uniform distribution over an interval
- D) Categorical distribution over 4096 bins
Question 10: Scaling Hypothesis
What has been found about scaling laws for time series foundation models compared to NLP?
- A) Larger models are always significantly better
- B) Evidence is mixed — data quality matters more than parameter count
- C) Scaling works identically to NLP
- D) Smaller models are always superior
Chapter 12: Spectral Analysis and Frequency-Domain Methods
Question 1: Discrete Fourier Transform
What does the Discrete Fourier Transform (DFT) of a time series measure?
- A) The linear trend of the series
- B) The decomposition of the series into sinusoidal components at different frequencies
- C) The number of missing observations
- D) The correlation between different series
Question 2: Nyquist Frequency
What happens when a signal contains frequencies above the Nyquist limit ($f_N = 1/2$)?
- A) The frequencies are detected correctly
- B) Aliasing occurs — high frequencies appear as spurious low frequencies
- C) The signal is automatically filtered
- D) The amplitude becomes zero
Question 3: Periodogram — Inconsistency
Why is the periodogram an inconsistent estimator of the spectral density?
- A) Bias increases with sample size
- B) Variance does NOT decrease as $T \to \infty$
- C) It cannot detect low frequencies
- D) It requires stationary data
Question 4: Welch Method
How does Welch's method reduce spectral estimation variance?
- A) It uses only the central half of the data
- B) It averages periodograms over overlapping windowed segments
- C) It applies differencing to the series
- D) It removes frequencies above a threshold
Question 5: Long Memory
How does long memory manifest in the frequency domain?
- A) The spectral density is constant (flat)
- B) The spectral density diverges at frequency zero: $f(\omega) \propto |\omega|^{-2d}$
- C) The spectrum has a single peak at the Nyquist frequency
- D) All autocorrelations are zero
Question 6: Spectral Coherence
What does the squared coherence $C^2_{xy}(\omega)$ between two series measure?
- A) The total correlation between the series
- B) The linear association at each frequency, analogous to frequency-specific $R^2$
- C) The amplitude difference between the series
- D) The number of common cycles
Question 7: Hodrick-Prescott Filter
What is the main limitation of the HP filter ($\lambda = 1600$ for quarterly data)?
- A) It cannot be applied to quarterly data
- B) It is a high-pass filter that leaks low-frequency noise
- C) It requires stationary data
- D) It only works for series with linear trends
Question 8: Wavelet vs Fourier Analysis
What is the main advantage of wavelet analysis over Fourier for non-stationary series?
- A) Wavelets have better frequency resolution
- B) Wavelets provide time-frequency representation — showing WHEN frequencies occur
- C) Wavelets require less data
- D) Wavelets automatically remove the trend
Question 9: Spectral Leakage
What causes spectral leakage and how is it reduced?
- A) Signal frequency is too high; reduced by subsampling
- B) Frequency doesn't fall exactly on the DFT grid; reduced by windowing (tapering, e.g. Hann)
- C) Series has too few observations; reduced by zero-padding
- D) Data contains missing values; reduced by interpolation
Question 10: Business Cycle
What is the standard NBER band for business cycles in quarterly data?
- A) 2–4 quarters (6 months – 1 year)
- B) 6–32 quarters (1.5–8 years)
- C) 40–100 quarters (10–25 years)
- D) 1–2 quarters (3–6 months)
Chapter 13: LPPL Models for Bubble Detection
Question 1: Super-Exponential Growth
What is the main mathematical signature of a financial bubble?
- A) Price grows linearly in time
- B) Log-price grows faster than linearly (super-exponential growth)
- C) Volatility is constant
- D) Returns are normally distributed
Question 2: Ising Model — Ordered Phase
In the financial analogy of the Ising model, what does the regime $T < T_c$ (ordered phase) correspond to?
- A) Efficient market with random trading
- B) Strong herding — bubble regime, one group dominates
- C) High-frequency trading with no directional bias
- D) Bear market with declining prices
Question 3: Susceptibility at the Critical Point
Why does the susceptibility $\chi \to \infty$ at the critical point $T_c$ matter for financial markets?
- A) The market becomes insensitive to news
- B) The market is perfectly efficient
- C) Even a small perturbation can trigger a system-wide cascade
- D) Volatility drops to zero
Question 4: Discrete Scale Invariance
What does discrete scale invariance (DSI) produce in the LPPL model?
- A) Constant exponential growth
- B) Log-periodic oscillations with a preferred scaling ratio $\lambda$
- C) Gaussian-distributed returns
- D) Zero correlation between returns
Question 5: LPPL Equation — Parameter B
Why is the condition $B < 0$ essential in the LPPL model?
- A) It ensures the price decreases over time
- B) It ensures super-exponential price growth as $t \to t_c$
- C) It eliminates log-periodic oscillations
- D) It forces $t_c$ to be in the past
Question 6: Partial Linearization
What advantage does partial linearization (slaving) provide in LPPL estimation?
- A) It eliminates all 7 parameters from optimization
- B) It reduces the search from 7D to 3D $(t_c, m, \omega)$ + OLS for the rest
- C) It guarantees a unique global optimum
- D) It does not require real market data
Question 7: Filter Conditions
Why is the condition $0.1 \leq m \leq 0.9$ necessary for a valid bubble signal?
- A) It ensures growth is slower than exponential
- B) It ensures growth faster than exponential but not instantaneous
- C) It eliminates oscillations from the model
- D) It forces the price to mean-revert
Question 8: LPPLS Confidence Indicator
How is the LPPLS Confidence Indicator (CI) constructed?
- A) By computing daily volatility
- B) By fitting LPPL over multiple windows and counting the fraction passing all 8 filter conditions
- C) By comparing the current price to a moving average
- D) By applying the ADF test on returns
Question 9: Negative Control
Why is negative control testing (e.g., the COVID-19 crash) important for validating LPPL?
- A) It shows LPPL can predict any type of crash
- B) It confirms LPPL does NOT produce false alarms for exogenous (non-bubble) shocks
- C) It shows the market was efficient in 2020
- D) It validates that COVID-19 was an endogenous bubble
Question 10: Phase Transition
What happens to the magnetization $|M|$ at the critical temperature $T_c$ in the 2D Ising model?
- A) $|M|$ remains at 1 (perfect order persists)
- B) $|M|$ drops continuously to 0 (second-order phase transition)
- C) $|M|$ jumps discontinuously from 1 to 0 (first-order transition)
- D) $|M|$ oscillates between 0 and 1
Resources
Recommended Textbooks
- Hyndman, R.J. & Athanasopoulos, G. (2021). Forecasting: Principles and Practice, 3rd ed. OTexts (free online)
- Hamilton, J.D. (1994). Time Series Analysis. Princeton University Press.
- Tsay, R.S. (2010). Analysis of Financial Time Series, 3rd ed. Wiley.
- Shumway, R.H. & Stoffer, D.S. (2017). Time Series Analysis and Its Applications, 4th ed. Springer.
Online Resources
Contact Information
Instructor
Prof. dr. Daniel Traian Pele
Department of Statistics and Econometrics
Location
Bucharest University of Economic Studies
Faculty of Cybernetics, Statistics and Economic Informatics

