Volatility \(\sigma\) is the standard deviation (the square root of the variance \(\sigma^2\))
Later:\(\mu_{\color{red}{t}} := E(r_{t+1}|\mathcal{F}_t)\) and \(\Sigma_{\color{red}{t}} = \text{Cov}(r_{t+1}|\mathcal{F}_t)\) where \(\mathcal{F}_t\) denotes the available information at \(t\)
Sample moments
Suppose you have \(T\) observations of the \((N\times1)\) vector, \(r_1, \ldots, r_{t}, \ldots, r_T\)
The sample counterparts \(\hat{\mu}\) and \(\hat{\Sigma}\) are \[
\hat\mu = \frac{1}{T}\sum\limits_{t=1}^Tr_t \qquad \hat{\Sigma} = \frac{1}{T-1}\sum\limits_{t=1}^T\left((r_t - \hat{\mu})(r_t - \hat{\mu})'\right)
\]
Aim: Choose \((N \times 1)\) vector \(\omega\) such that \(\sum\limits_{i=1}^N \omega_i = \iota'\omega = 1\) where \(\iota\) is an \((N\times 1)\) vector of ones
Portfolio returns \(r^{pf}_{t} = \omega'r_t\)
Common properties of utility functions (concave) \[
U'(r_t) > 0 \text{ and } U(E(r_t)) > E(U(r_t))
\]
Preference for higher expected return and lower volatility \(\sigma^{pf} =\sqrt{\text{Var}\left(r^{pf}_t\right)}\)
The minimum variance portfolio weights are given by the solution to \[
\omega_\text{mvp} = \arg\min \omega'\Sigma \omega \text{ s.t. } \iota'\omega= 1
\]
As a result, the efficient portfolio weight takes the form (for \(\bar{\mu} \geq D/C = \mu'\omega_\text{mvp}\)) \[
\omega_\text{eff}\left(\bar\mu\right) = \omega_\text{mvp} + \frac{\tilde\lambda}{2}\left(\Sigma^{-1}\mu -\frac{D}{C}\Sigma^{-1}\iota \right)
\]
The efficient portfolio
Note that \[
\iota'\left(\Sigma^{-1}\mu -\frac{D}{C}\Sigma^{-1}\iota \right) = D - D = 0\text{ so }\iota'\omega_\text{eff} = \iota'\omega_\text{mvp} = 1
\]
The efficient portfolio allocates wealth in the minimum variance portfolio \(\omega_\text{mvp}\) and a levered (self-financing) portfolio to increase the expected return
The efficient frontier
Assume you computed \(\omega_\text{eff}(\bar\mu)\) and \(\omega_\text{eff}(\tilde\mu)\) for \(\bar\mu > \tilde \mu \geq D/C\), then any linear combination with \(c\in\mathbb{R}_+\) can be represented as \[
\omega^* = c\omega_\text{eff}(\bar\mu) + (1-c)\omega_\text{eff}(\tilde\mu) = \omega_\text{mvp} + \frac{\lambda^*}{2}\left(\Sigma^{-1}\mu -\frac{D}{C}\Sigma^{-1}\iota \right)
\]
with \(\lambda^* = 2\frac{c\bar\mu + (1-c)\tilde\mu - D/C}{E-D^2/C}\).
portfolio weight vector \(\omega\in\mathbb{R}^N\) which denotes investments into the available \(N\) risky assets.
Now: Instead of \(\sum_{i=1}^N \omega_i=1\), assume that all the remaining wealth, \(1-\iota'\omega\), is invested in a risk-free asset which pays a constant interest \(r_f > 0\)
The expected portfolio return for the portfolio of risky assets \(\omega\) is then \[
\mu_\omega = \omega^{\prime}\mu + (1-\iota^{\prime}\omega)r_f = r_f + \omega^{\prime}\underbrace{(\mu-r_f)}_{\tilde\mu}
\]
We refer to \(\tilde\mu\) as the vector of expected excess returns.
the volatility of the portfolio is given by \[
\sigma_\omega = \sqrt{\omega^{\prime}\Sigma\omega}
\]
Optimal decision problem
earn a desired level of expected portfolio (excess) returns (\(\bar\mu-r_f\)) with the lowest possible variance leads to \[
\min_\omega Z(\omega) = \min_\omega \omega^{\prime}\Sigma\omega - \lambda \left(\omega^{\prime}\tilde\mu-\bar\mu\right).
\]
The first-order conditions for this optimization problem yield: \[
\frac{\partial Z}{\partial \omega} = 2\Sigma\omega - \lambda \tilde\mu = 0 \Leftrightarrow \omega^* = \frac{\lambda}{2}\Sigma^{-1}\tilde\mu
\]
the optimal portfolio weights are given by \[
\omega^* = \frac{\bar\mu}{\tilde\mu^{\prime}\Sigma^{-1}\tilde\mu}\Sigma^{-1}\tilde\mu \Rightarrow \omega_\text{tan} := \frac{\omega^*}{\iota'\omega^*}= \frac{\Sigma^{-1}(\mu-r_f)}{\iota^{\prime}\Sigma^{-1}(\mu-r_f)}.
\]
Why does \(\bar\mu\) not show up in \(\omega_\text{tan}\)?
Taking another look at the efficient tangency portfolio \(\omega_\text{tan}\) reveals that expected asset excess returns \(\tilde\mu\) cannot be arbitrarily large or small
From the first order condition of the optimization problem above we get \[
\frac{\partial Z}{\partial \omega} = 2\Sigma\omega - \lambda \tilde\mu =0 \\\Leftrightarrow \tilde\mu = \frac{2}{\lambda}\Sigma\omega^* = \frac{2}{\lambda}\underbrace{\iota'\omega^*}_{=\frac{\lambda}{2}\iota^{\prime}\Sigma^{-1}\tilde\mu}\Sigma\omega_\text{tan} \\ = \iota^{\prime}\Sigma^{-1}\tilde\mu\Sigma\omega_\text{tan}
\]
Putting everything together yields for the expected excess return of asset \(i\):
Calculating the tangency portfolio can be cumbersome
What is the correct asset universe?
How to estimate \(\mu\) and \(\Sigma\) for many assets?
In the CAPM: market portfolio = tangency portfolio
Skip calculation of tangency portfolio weights
Use portfolios weighted by market capitalization
Capital asset pricing model
If all investors share beliefs about \(\Sigma\) and \(\mu\)and can borrow and invest without limits at \(r_f\), everybody invests a fraction of her wealth in \(\omega_\text{tan}\) and in the risk-free rate (two mutual fund theorem)
The tangency portfolio is the market portfolio \(\omega_\text{m}\) and the individual weights of asset \(i\) are just \[
\omega_\text{tan, i} = \omega_\text{m, i} = \frac{P_iSCO_i}{\sum\limits_{j=1}^N P_jSCO_j}
\]
where \(SCO_j\) is the number of shares outstanding of stock \(j\). - To align the two results, the capital asset pricing model imposes constraints for the expected return of stock \(i\)
For the econometric analysis of the model, we assume that innovations \(\varepsilon_{i, t}\) are independently and identically distributed (IID) through time and jointly multivariate normal
Expected returns are entirely determined by the price of risk (market risk premium) and the co-movement of asset \(i\) with the market, \(\beta_i\)
To determine whether the portfolio/asset generates abnormal returns relative to the CAPM risk model, we evaluate whether the fitted intercept coefficient \(\alpha_i\), which serves as the estimate of the average abnormal return per period, is statistically distinguishable from zero
Discuss: What are the questionable assumptions behind the baseline regression framework?
CAPM - Example with simple linear regression
Accessing financial data
The US Stock market
Large parts of the academic literature focus on US stock markets
Stocks are listed on US exchanges (NYSE, AMEX, NASDAQ, and some smaller ones)
Extensive data on prices and trading activity is provided by the Center for Research in Security Prices (CRSP), maintained by the University of Chicago, Booth School of Business
Full sample starts from December 1925 and is continuously updated
Note that in the textbook, we explain how to get CRSP data from the WRDS interface. As a KU student, you do not have access to data beyond CRSP. You will receive the preprocessed data directly via Absalon
Composition of the CRSP sample
Familiarize yourself with the cleaning steps in the exercise on the CRSP sample (Chapter 7, Bali, Engle, and Murray)
Monthly processed (!) data is available in crsp_monthly.parquet file in Absalon
Only contains US stocks (shrcd%in%c(10,11))
Variable exchcd determines listing exchanges, siccd lists industry
We adjust the market capitalization values for inflation using the Consumer Price Index (CPI, All Urban Consumer series) from the Bureau of Labor Statistics website.
Computing beta for the CRSP universe
Market beta for month \(t\) is estimated with data from prior to and including \(t\), e.g. 5 years of monthly data
Rolling-window regressions are straightforward from a methodological perspective but tricky to implement
Exercises: conduct rolling window regression of beta for the entire CRSP universe
Figure shows decile portfolio sorts based on lag beta
Portfolio 10 corresponds to the highest beta decile
Each bar corresponds to the CAPM alpha of the value-weighted portfolio performance
How to use any factor structure for portfolio optimization
Suppose the CAPM holds: implied factor structure for expected excess returns is \[
\begin{aligned}\mu &= E(r) = r_f + \beta E(r_m - r_f)\\\Sigma &=\sigma^2_m\beta\beta' + \Sigma_\varepsilon
\end{aligned}
\]
If \(\Sigma_\varepsilon\) is a diagonal matrix, estimation of \(\Sigma\) requires only \(2N + 1\) instead of \(N(N-1)/2\) parameters
Practical recipe for portfolio optimization with general factor structure:
Estimate \(\beta_i\) for each asset
Estimate market risk premium \(E(r_m - r_f)\)
(Univariate) estimation of the elements of \(\Sigma_{\varepsilon}\) based on residualized returns \[
\hat\varepsilon_{i,t} = r_{t,i} - r_f - \hat\beta_i(r_{m,t} - r_f)
\]
Replace sample estimates \(\hat\mu\) and \(\hat\Sigma\) with the theoretically implied values computed above to choose \[
\omega = \arg\max \hat\mu'\omega - \frac{\gamma}{2}\omega'\hat\Sigma\omega
\]
Testing the CAPM
Suppose we want to test if the CAPM holds jointly for \(N\) assets
The CAPM implies that all elements of the \((N \times 1)\) vector \(\alpha\) are zero in the joint regression framework \[
Z_t = \alpha + \beta Z_{m,t} + \varepsilon_t
\]
where \(\beta\) is the \((N \times 1)\) vector of market betas (if \(\alpha\) is zero, then the market portfolio is the tangency portfolio) and \(Z_{i,t}\) denotes excess returns - Standard ordinary least squares (OLS) estimation delivers \[
\hat\alpha = \hat\mu - \hat\beta\hat\mu_m
\]
The Wald-test statistic of the null hypothesis \(H_0: \alpha = 0\) is \(J = \hat\alpha'\left(Var(\hat\alpha)\right)^{-1}\hat\alpha\)
Testing the CAPM
MacKinlay (1987) and Gibbons, Ross, and Shanken (1989) developed the finite-sample distribution of \(J\) which yields \[
J = \frac{T-N-1}{N}\left(1 + \frac{\hat \mu_m ^2}{\hat\sigma_m^2}\right)^{-1}\hat\alpha'\hat\Sigma^{-1}\hat\alpha
\]
where \(\hat\Sigma = Cov(\hat\varepsilon)\) and \(\hat\sigma\) is the standard deviation of the market excess returns.
Under the null hypothesis, \(J\) is unconditionally distributed central \(F\) with \(N\) degrees of freedom in the numerator and \((T-N-1)\) degrees of freedom in the denominator
For more information, see Chapter 4 of The Econometrics of Financial Markets
Fama-MacBeth regressions
Instead of focusing on the mean-variance efficiency of the market portfolio, the CAPM also implies a linear relationship between expected returns and market betas which completely explains the cross-section of expected returns
Portfolio sorts already revealed that this may not be the case
These implications can also be tested using a cross-sectional regression methodology (Fama and MacBeth, 1973)
Basic idea: For each cross-section of returns (e.g., each month), project asset returns on factor exposures or characteristics that resemble exposure to a risk factor and then aggregate the estimates in the time dimension
E.g., first, for each month \(t\), estimate \[
Z_t = \gamma_{0,t}\iota + \gamma_{1,t}\hat\beta + \eta_t
\]
Then, we analyse time series of \(\hat\gamma_{0,t}\) and \(\hat\gamma_{1,t}\).
CAPM implies that \(E(\gamma_{0,t}) = 0\) (no mispricing) and \(E(\gamma_{1,t})>0\) (positive market premium)
In most applications we use \(\hat\beta \Rightarrow\) errors-in-variables problem (Shanken, 1992)
Unobservability of the market portfolio
Gibbons, Ross, and Shanken’s test focuses on the mean-variance efficiency of the market portfolio
Most tests use a value- or equal-weighted basket of NYSE and AMEX stocks as the market proxy, whereas theoretically, the market portfolio contains all assets
Roll (1977) emphasizes that tests of the CAPM only reject the mean-variance efficiency of the proxy and that the model might not be rejected if the return on the true market portfolio were used
Implications of the overwhelming evidence against the CAPM?
Replace CAPM with multifactor models with several sources of risk
Maybe the evidence against the CAPM is overstated because of mismeasurement of the market portfolio, improper neglect of conditioning information, data-snooping, or sample-selection bias
What if no risk-based model can explain the anomalies of the stock market behavior (behavioral finance)?
Theoretical drawbacks of the CAPM
The CAPM assumes that the average investor cares only about the performance of the investment portfolio but wealth could emerge from other sources and higher-order risks could play a role
The CAPM assumes a static one-period model. In Merton’s (1973) ICAPM, the demand for risky assets is attributed not only to the mean-variance component, as in the CAPM, but also to hedging against unfavorable shifts in the investment opportunity set
Empirically, the poor performance of the single factor CAPM motivated a search for multifactor models
Empirical problems (Discussion)
Let’s move from theory to practice:
Which decisions do you as a portfolio manager have to take?
What issues may occur and lead to deviations from the theoretically optimal portfolio?
How do you evaluate your choices?
Portfolio backtesting
Portfolio backtesting is often perceived as a quest to find the best strategy or at least a solidly profitable one (downside: data snooping, p-hacking)
Out-of-sample test to analyze the (hypothetical) performance of a strategy
Procedure
Fix sample period \(T\), estimation window size \(h\), and out-of-sample horizon \(K = T - h\)
Respect the information constraint: never use data an investor would not have at the time of decision
For each period, recompute \(\hat\Sigma\) and \(\hat\mu\) and reallocate wealth
At the end of each period store the portfolio performance (e.g. return)
Compare different strategies by evaluating the average out-of-sample performance
Rolling estimation window of size \(h\); portfolio weights are recomputed each period and performance is measured out-of-sample.
Evaluation metrics
Typical metrics are the out-of-sample portfolio return (eventually risk-adjusted) \[
\hat E(r^{pf}) = \frac{1}{T-h}\sum\limits_{t=h+1}^T r^{pf}_t
\]
Note that \(r_t^{pf} = \omega_{t-1}'r_t\) where \(\omega_{t-1}\) denotes weights formed using information available up to time \(t-1\)
Later we will also consider transaction costs which depend on rebalancing from \(\omega_{t-1}\) to \(\omega_{t}\)
Rolling windows in R and Python
You will need for loops to mimic sliding through time
for loops provide a way to tell, “Do this for every value of that.” In R and Python syntax, this looks like this:
for (value inc("My", "first", "for", "loop")) {print(value)}
[1] "My"
[1] "first"
[1] "for"
[1] "loop"
for value in ["My", "first", "for", "loop"]:print(value)
My
first
for
loop
Task: Develop pseudo-code for portfolio backtesting
Plug-in weights are highly unstable
Inverting a noisy \(\hat\Sigma\) and scaling by \(\hat\mu\) amplifies estimation errors into extreme weights
Rolling-window tangency portfolio weights can be extremely volatile, often well exceeding \(\pm 100\%\)
High turnover and extreme short positions are a direct consequence of unconstrained plug-in estimation
This motivates imposing constraints even at the cost of some theoretical inefficiency
Parameter uncertainty
Consider a quadratic utility function with certainty equivalent \[CE(\omega) = \omega'\mu - \frac{\gamma}{2} \omega'\Sigma \omega \] where \(\gamma > 0\) is the coefficient of risk aversion
Maximum expected utility portfolio (under the constraint \(\iota'\omega = 1\)) is equivalent to the framework above (minimize volatility for a given level of return \(\bar\mu\)) (the proof is an exercise: show that there is a bijective mapping from \(\gamma\) to \(\bar\mu\))
Econometrician can only invest in \(\omega_\gamma(\hat\mu, \hat\Sigma)\)
As a result: Inefficient wealth allocation
How important is parameter uncertainty?
Certainty equivalent loss is defined as \[
CEL = CE(\omega_\gamma(\mu, \Sigma)) - E\left(CE(\omega_\gamma(\hat\mu, \hat\Sigma))\right) \geq 0
\]
Writing \(\hat\omega := \omega_\gamma(\hat\mu, \hat\Sigma)\) for the plug-in portfolio, the loss can be approximated by \[
CEL \approx \frac{\gamma}{2}\times tr\left(\text{Var}(\hat\omega)\,\hat\Sigma\right)
\]
Depends on the risk-aversion \(\gamma\), the \((N\times N)\) sampling variance of \(\hat\omega\), and the return covariance
Kan and Zhou (2007) and DeMiguel, Garlappi, and Uppal (2009) consider different portfolio weights that reduce the certainty equivalent loss
It can be shown that the naive portfolio \(\frac{1}{N}\iota\) can be optimal if estimation (and model) uncertainty is huge
Efficient frontier uncertainty: a bootstrap view
Plug-in estimates \(\hat\mu\) and \(\hat\Sigma\) imply a single frontier — but it is just one noisy estimate
Resample the return data and compute the efficient frontier for each bootstrap draw
Each bootstrapped frontier, when evaluated at the true (full-sample) parameters, shows how much the frontier moves with sampling variation
The in-sample frontier is the best-looking one by construction — a severe form of overfitting (Michaud, 1998)
The cost of plug-in estimation
Simulation: treat full-sample \((\mu_0, \Sigma_0)\) as the truth; simulate samples of length \(T = 60\) months from \(\mathcal{N}(\mu_0, \Sigma_0)\)
For each simulation compute \(\hat\omega_s^\text{tan}\) from plug-in estimates and evaluate its Sharpe ratio under the true parameters
Compare the distribution of \(\widehat{\text{SR}}_s\) to the oracle \(\text{SR}^* = \sqrt{\mu_0'\Sigma_0^{-1}\mu_0}\)
How imposing restrictions helps
Consider the Kuhn-Tucker conditions for the minimum variance portfolio with short-selling constraints \[
\Sigma\omega - \lambda = \lambda^0\iota \text{ where } \lambda_i \geq0 \text{ and } \lambda_i = 0 \text{ if } \omega_i >0
\]
Let \(\omega_\text{no short}\) denote the solution to the problem above
What is so special about \(\tilde\Sigma\) instead of \(\Sigma\)? Consider the (unconstrained) first-order condition again. If stock \(i\)’s marginal contribution to the portfolio variance is large, the weights \(\omega_i\) are reduced (become negative)
With \(\tilde\Sigma\), assets which would imply negative weights are treated as if the variance is reduced by \(2\lambda_i\) and covariance with asset \(j\) is reduced by \(\lambda_i + \lambda_j\)
How imposing restrictions helps
Two commonly used approaches are naive diversification \(\omega_\text{naive} = \frac{1}{N}\iota\) and short-sell constrained portfolios
No closed-form solution exists to compute \(\omega_\text{no short}\) but solve.QP (R) from package quadprog or minimize from library scipy.optimize deliver the numerical solution to quadratic programming problems of the form