學術論文
Research Papers · 學術論文與研究成果
14 papers
Forecast-Timing Conventions and the Value of Overnight Information in Volatility Forecasting
Yi-Hao Lai
11 pages | 18 references
Uploaded: 2026-04-05 19:19 · Revised: 2026-09-21 17:04
Abstract▶
Session-level volatility models produce overnight and intraday forecast components that are naturally issued at different times: the overnight component at the previous close, the intraday component at the market open, after the overnight return has realized. We show that the forecast-timing convention—when each component is issued and what the benchmark observes—drives the apparent value of session-level modeling. Using a parsimonious Periodic Realized GARCH (PRG, 6–8 parameters) on pinned data for six equity, commodity, and futures markets, we evaluate the same model under three conventions. Under a mixed-timing design that session-level evaluation easily slips into—model components issued at two different times, benchmark held at the previous close—PRG beats GJR-GARCH with Diebold–Mariano statistics of +4.3 to +6.4 in all six markets. Under a coherent close-time convention, with every component issued at the previous close, the advantage vanishes: zero of six markets clear the conservative |t|>3 threshold, and the only nominally significant market points against PRG. Under a coherent open-time convention, where an information-matched GJR-X benchmark also observes the realized overnight return, PRG's advantage is positive in all six markets and on the pinned vintage exceeds the conservative |t|>3 threshold in five; the exception has the smallest overnight variance share. The same model and data thus deliver DM statistics from -2.3 to +10.1 depending only on the evaluation clock. Overnight information has genuine forecasting value at the open horizon and none at the day-ahead horizon; mixed-timing comparisons attribute to the model what belongs to the clock.
Earnings-Announcement Volatility Amplification: A Cross-Market Regularity with Magnitude Ordering Evidence from Taiwan, U.S., and Japan Equity Markets
Yi-Hao Lai
30 pages | 18 references
Uploaded: 2026-05-18 18:00 · Revised: 2026-09-01 16:26
Abstract▶
We document a robust cross-market regularity in earnings-announcement volatility amplification using a multiplicative GARCH framework that decomposes conditional variance into a firm-specific GJR component (g_i,t) and a pooled market-level factor (τ_t). Estimating a shared τ coefficient on a binary earnings indicator across three independent equity markets—Taiwan (N=31 TWSE bluechips), the United States (N=30 S&P 500 large-caps), and Japan (N=30 TOPIX large-caps), all 2014–2025—we obtain θ_EAV > 0 with cluster-bootstrap |t| of 5.24 (TW), 4.50 (US), and 11.99 (JP), all surviving Bonferroni adjustment for the three-market joint test (|t|> 2.39). Within-stock permutation placebos (n=60 per market) reject the null at 0/60 in every market, with observed θ_EAV standing 13.3 (TW), 70.7 (US), and 38.6 (JP) placebo standard errors above zero. Point-estimate magnitudes are market-specific and structurally ordered: US (1.9110^-4)> JP (1.4110^-4)> TW (6.3610^-5), consistent with cross-market differences in analyst coverage density, earnings-call culture, and institutional pre-announcement positioning. A PCA-based market-factor absorption test confirms the effect is orthogonal to systematic stress in both markets (Scenario A), with stress-interaction asymmetric across markets (US amplified; TW null). Within each market, however, no observable firm-attribute predictor of θ_EAV survives multiple-testing correction across four pre-registered tests, supporting interpretation of θ_EAV as a market-level constant whose cross-market variation reflects structural rather than idiosyncratic forces. Extending the analysis to a 13-market institutional panel (N = 172 stocks across Brazil, Canada, China, the EU, Hong Kong, Indonesia, India, Japan, Korea, Mexico, Taiwan, and the U.S., 2011–2025) identifies three structural drivers of cross-market θ_EAV heterogeneity: (i) within-market analyst attention (panel-OLS Harvey |t| on (analyst) rising monotonically 3.236 3.808 across five N-extension iterations, all clearing the |t| > 3 bar); (ii) within-market GICS sector composition (incremental adj.-R^2 = 0.148, joint F = 689.5, p = 7.9 × 10^-14); and (iii) cross-market institutional ownership (refined Spearman ρ = +0.379, N = 13). A pooled-MLE multistart audit (≥ 100 random initialisations per market) further establishes that the US>JP>TW headline magnitude ordering is preserved between canonical single-init and refined 100-start estimates (Taiwan diagnosed STABLE/flat-ridge; U.S. and Japan diagnosed FRAGILE with refined θ_EAV 10–28× canonical), providing a third contribution: a small-S multistart diagnostic protocol for cross-market shared-MIDAS+stock-FE-GJR specifications.
The Crypto Fear Channel: Asymmetric and Regime-Dependent Volatility Spillover between Bitcoin and Equity Markets
Yi-Hao Lai
19 pages | 45 references
Uploaded: 2026-05-11 18:09 · Revised: 2026-09-26 08:10
Abstract▶
Using daily returns on SPY, BTC-USD, and VIX from 2015-02 to 2026-04 (N = 2,812), we document two stylized facts about the crypto-equity volatility spillover and one negative result. The channel is bidirectional: symmetric Granger tests reject non-causality in the VIX BTC RV direction at every lag from 1 to 10, and at lag 1 that reverse leg is the stronger of the two (F = 7.77, p = 0.00535, against a forward F = 0.59, p = 0.444); the BTC VIX leg dominates from lag 2 onward. Our findings characterize that dominant leg and do not license a one-way reading. First, the spillover is asymmetric: Bitcoin downside volatility Granger-causes VIX at lags 1–5, whereas Bitcoin upside volatility does not (symmetric Granger tests covering lags 1–10 corroborate the forward leg at lags 2–10). Second, the spillover is regime-dependent: in a five-subperiod breakdown, only 2020 survives Bonferroni correction (F = 12.31, p < 10^-6); 2015–2017 is marginally significant at the uncorrected 5% level (p = 0.014) and the remaining three subperiods are clearly non-significant. The negative result concerns tail concentration. A quantile regression of VIX on lagged BTC realized variance alone shows a slope that is negative in the lower quantiles and large and positive in the upper tail, but that pattern does not survive once lagged VIX is controlled: the upper-tail slope falls to +0.42 at τ = 0.95 with a 95% moving-block bootstrap interval of [-0.61, +1.14], and what remains is a small negative partial slope at τ ≤ 0.5. The unconditional pattern is therefore absorbed by the persistence of VIX itself and carries no evidence of tail amplification. A generalized (order-invariant) connectedness decomposition places Bitcoin only marginally on the receiving side (net -0.95 percentage points, negative in 72.5% of rolling windows), too weakly to carry a directional claim on its own; the "fear amplifier conditional on pre-existing equity stress" reading therefore rests on the Granger evidence alone. Out-of-sample forecasts show that adding BTC realized volatility does not improve an AR(3) model of VIX (Diebold-Mariano t = -0.99, p = 0.32; Clark-West t = -0.12 for the nested comparison), highlighting that Granger causality is a necessary but not sufficient condition for forecastability. Taken together, our findings describe the crypto-equity spillover as a bidirectional, crisis-concentrated channel whose dominant leg runs from Bitcoin downside volatility to equity fear: informative for ex-post risk attribution, but inadequate as an ex-ante predictor of equity fear.
Why GAS-t Fails on Pre-Institutional Bitcoin: A Robust Deficit with a Loss-Conditional Diagnosis
Yi-Hao Lai
66 pages
Uploaded: 2026-07-27 12:00 · Revised: 2026-09-24 19:07
Abstract▶
We document a previously unreported reversal in Generalized Autoregressive Score (GAS) volatility models with Student-t innovations applied to Bitcoin: over the pre-institutional period (January 2017 to December 2020, n=1,441 out-of-sample days), GAS-t produces 11.79% worse quasi-likelihood (QLIKE) loss than the GJR-Normal benchmark (Diebold-Mariano-HLN t = -4.40, p = 1.2 × 10^-5). Two natural hypotheses—that score-driven GAS dynamics fail, or that Student-t innovations fail—are typically conflated in the GAS literature. We resolve this via a factorial five-model decomposition (GJR-N, GJR-t, GAS-N, GAS-t, GJR-N-standardized) which shows that GAS-Normal recovers to near-parity with GJR-Normal (DM-HLN = -1.97, p = 0.049: marginal at the conventional 5% level and far below the |t| > 3 bar we apply throughout) while GAS-t and GJR-t both reverse, pointing at the Student-t innovation rather than at the GAS dynamics. We report the limits of that diagnosis rather than only its point estimate. The evidence against the dynamics is the weaker half, because the dynamics contrast does not reach the |t| > 3 bar. The GJR-Normal benchmark was estimated from a single start and the GAS cells from 100, but in each of the 23 Period 1 refit windows the single start reaches the best objective of 100 starts to within 10^-6, so the difference in search effort does not move the benchmark's fitted objective. The deficit itself is robust: it survives all three Patton (2011) losses we score it under (t = -4.40, -5.32, -6.37), a 60-day shift of the period boundary, and a leave-one-year-out jackknife. The attribution is not: the GAS-Normal versus GAS-t contrast is +3.30 under QLIKE but changes sign and loses significance under MSE (-1.01) and under b=-1 (-0.31), and in a leave-one-year-out jackknife it keeps its sign in every fold but clears |t| > 3 in only one. A Generalized Error Distribution innovation—also fat-tailed—recovers to parity, which narrows the culprit from heavy tails in general to the Student-t score's specific down-weighting of extremes. A Markov-Switching GAS-t extension rescues GAS-t to statistical parity with GJR-Normal under both estimation protocols we report (DM-HLN vs single-state GAS-t = +4.60 in the original estimation and +3.35 under a common 100-start re-estimation; vs GJR-Normal +0.27 and -1.36, neither significant), at a cost of ten free parameters against four—parity, not superiority. In sample, both the Akaike and the Bayesian information criteria prefer the switching specification in each of the six Period 1 estimation windows of the switching model. Because neither protocol leaves a statistically detectable residual gap, a regime-mixture explanation of the Period 1 reversal is not rejected by this test, and we do not claim a structural one. The reversal disappears in the FTX-Luna evaluation window (January–December 2023), where GAS-t is 1.09% better than the benchmark; in the spot-ETF evaluation window (January–April 2026) the point estimate is still 4.10% adverse but rests on 100 evaluation days and is statistically indistinguishable from zero. Institutional flows plausibly altered Bitcoin's return distribution. We provide practitioner guidance on when GAS-t is and is not appropriate for crypto volatility forecasting.
Leverage Direction Matters: Cross-Asset Evidence on GARCH Model Selection and Volatility Targeting
Yi-Hao Lai
44 pages | 57 references
Uploaded: 2026-03-18 02:30 · Revised: 2026-09-01 21:23
Abstract▶
The direction of the asymmetric volatility response—measured by the GJR-GARCH γ parameter—carries distinct economic content that varies systematically across asset classes. Using daily data for seven primary assets (equities, gold, bonds, emerging markets, cryptocurrency) over 2017–2025 (with 2026 reserved for out-of-sample validation), extended to 26 assets in validation analyses, we establish a single central proposition: the sign of γ identifies the dominant price-driving mechanism of an asset—balance-sheet leverage for equities (γ > 0), fear-driven flight-to-quality for gold (γ < 0 conditional on stress regimes), and coupon-dominated pricing for bonds (γ 0). Gold's inverted leverage is regime-dependent rather than unconditional: the canonical rolling-window mean is statistically indistinguishable from zero (mean γ = +0.002, HAC t = +0.15), while a regime decomposition concentrates inversion in fear-driven rallies and reverts to standard leverage during liquidation-driven declines (t = -3.79, p < 0.001). This economic content has two empirical manifestations that we test rather than assume. First, γ-based model selection delivers out-of-sample forecasting gains only when asymmetry is statistically significant (t > 1.65 on the quarterly-mean HAC statistic), yielding 6/6 correct classifications on the disjoint 2024–2025 sample using 2023 γ estimates (small cross-section acknowledged, N = 6). Second, within homogeneous equity markets, γ predicts whether volatility targeting (VT) acts as trend-following or contrarian (Spearman ρ = 0.886, p = 0.019, N = 6); this mapping is a genuine domain restriction that does not survive extension to heterogeneous asset classes (ρ = -0.448, p = 0.14, N = 12). VT's drawdown-reduction benefit is meanwhile universal and driven by base volatility level (ρ = 0.944, N = 5; ρ = 0.83, N = 14), independent of γ. Together these results position leverage direction as an economically interpretable asset-classification device—not a general-purpose forecasting rule.
The True Cost of Volatility Targeting: Decomposing the Insurance Premium into Opportunity and Transaction Components
Yi-Hao Lai Department of Finance, Da-Yeh University 168 University Road, Dacun, Changhua 515006, Taiwan Corresponding author. Email: yhlai@mail.dyu.edu.tw
16 pages | 18 references
Uploaded: 2026-04-04 14:05 · Revised: 2026-09-07 14:24
Abstract▶
Volatility targeting (VT) strategies scale equity exposure inversely to expected volatility, producing well-documented tail-risk reductions but persistent return shortfalls relative to buy-and-hold. We decompose this return gap—the "insurance premium"—into opportunity cost (returns foregone by reducing equity participation) and direct cost (transaction expenses from rebalancing). Using daily S&P 500 data conditioned on the CBOE VVIX index over 2012–2024 (3,262 trading days, N=3,261 returns), we find that opportunity cost accounts for 90% of the total insurance premium under unconditional VT (4.05% vs. 0.43% per annum). The share is imprecisely estimated—its stationary-bootstrap 95% interval is [37%, 95%], which does not exclude direct cost exceeding opportunity cost—so what survives resampling is the level of the premium ([0.66%, 8.01%] per annum), not the split. VVIX-conditional targeting—activating VT only when the volatility-of-volatility z-score exceeds 1.0—reduces total cost by 74% in the full sample (from 4.48% to 1.18%), primarily by removing hedging in calm markets. This reduction, however, does not survive uncertainty quantification: the conditional strategy's premium is not bounded away from zero under resampling, and cross-OOS analysis over six non-overlapping windows shows it beats buy-and-hold on Sharpe in only 2. We therefore treat VVIX conditioning as a hypothesis-generating in-sample accounting result rather than a validated trading rule. The 50/50 SPY/GLD benchmark earns a 54 basis-point average rebalancing premium (negative in two of four sub-periods), raising the bar for VT to clear. These findings redirect the VT design problem from minimizing transaction costs to measuring when opportunity cost dominates the insurance premium.
Scale Before Shape: A Distribution-Free Diagnostic for Plug-In VaR Failure
Yi-Hao Lai
10 pages | 8 references
Uploaded: 2026-09-06 12:00 · not revised
Abstract▶
A plug-in Value-at-Risk built on a volatility forecast can fail coverage tests for two very different reasons: the forecast is right about shape but wrong about scale, or it is wrong about the shape of the tail itself. We give a distribution-free diagnostic that separates the two. For a plug-in VaR at level α, the implied scale (α) is the (1-α) empirical quantile of the standardized exceedance ratio u_t = r_t / VaR_t; a scale failure appears as 1 at both levels with Δ c = (1%) - (5%) indistinguishable from zero, while a shape failure appears as a significant Δ c. Inference uses a paired moving-block bootstrap and assumes no parametric innovation law. Applied to a HAR-RV plug-in VaR on 0050.TW, the diagnostic classifies the failure as scale, and a scale correction estimated from the same information set moves the affected specification back inside the calibrated region. We also report a cautionary construction result. The forecast-accuracy-versus-tail-adequacy divergence that motivated this line was an artifact of scoring a point-forecast loss against a realized-variance proxy different from the return series on which the VaR was evaluated; on the aligned target the loss ranking is not merely weaker but reverses sign and is statistically indistinguishable from zero. The practical rule is ordered: check target alignment, then scale, then shape.
Monotone Strategy-Specific Erosion under Volatility-Targeting Crowding: Matched-Control Identification via Agent-Based Simulation
Yi-Hao Lai
34 pages | 22 references
Uploaded: 2026-04-04 14:05 · not revised
Abstract▶
Volatility targeting (VT), trend-following (TF), and short-horizon mean-reversion (MR) share a procyclical structure: each strategy's position update is a non-trivial function of recent prices or volatility, so widespread adoption can in principle generate self-reinforcing trading flows. Whether the resulting strategy-performance degradation reflects a discrete crowding tipping point, and whether it is VT-specific or attributable to crowded flow in general, remains unsettled. We build an agent-based model with 1,000 agents of heterogeneous types, a Kyle (1985) market maker, endogenous VIX dynamics, and 200 fixed noise traders, and re-examine the threshold structure under matched microstructure with an exogenous sup-Wald detector and turnover-matched random-direction controls. Across 94,500 Monte Carlo simulations in the exogenous-detector redesign layer (27 cell–adoption combinations—seven adoption levels in the baseline cell and five in each of four perturbed cells—× 7 treatments/controls × 500 MC), three findings emerge. First, VT Sharpe declines monotonically with adoption rather than crossing a discrete threshold: in the canonical cell (λ = 0.005, γ = 200), Sharpe falls from 0.510 at 10% adoption to -0.271 at 100%, with adjacent path-bootstrap 95% confidence intervals separating from 40% onward; the sup-Wald test rejects flatness in all five cells (p = 0.001 in every cell) but identifies no interior break (the single breakpoint localizes to the saturation boundary), and the maximum marginal degradation concentrates in (70%, 100%]. Second, a turnover-matched random-direction control (RR\_VT) matching VT's trading-footprint distribution without its vol-feedback direction exhibits no degradation in any cell; at VT's footprint scale, the erosion is therefore identified to the systematic direction of vol-feedback trading rather than to crowded flow, providing a cleaner mechanism counterfactual than a no-strategy baseline (at trend-following footprint scales two orders of magnitude larger, direction-randomized controls do erode, so the identification is footprint-scale-dependent). Third, the canonical "70% threshold" reported under a Sharpe-only detector survives only as a descriptive relative-drop level-crossing (drop >70%) in 3 of 5 cells; under the exogenous, treatment-aware detector the interpretation is monotone erosion without a tipping point. The applicability gate excludes the TF treatment in four of five cells and MR in all five (structurally loss-making regimes), so cross-strategy family claims are withdrawn and the paper's mechanism evidence is confined to VT; our results sharpen the policy-relevant question from "where is the tipping point" to "how steep is the monotone erosion and what identifies its mechanism."
Volatility Targeting in the Taiwan Stock Market: Leverage Amplification, Model Selection, and Practical Implementation
Yi-Hao Lai
51 pages | 34 references
Uploaded: 2026-03-18 02:30 · Revised: 2026-09-03 15:22
Abstract▶
We examine volatility targeting (VT) strategies in the Taiwan stock market using the 0050.TW ETF over 2008–2026. Two contributions emerge. First, the TAIEX exhibits a diversification amplification of the leverage effect: under the canonical full-sample specification, the broad-index GJR-GARCH asymmetry parameter exceeds the individual-stock average by approximately 4.3× (90% bootstrap CI: [2.28, 6.58]; N = 9 stocks; the calendar-aligned rolling-window specification yields 6.12×, though the rolling index estimate is not itself individually significant). Despite this strong leverage effect, GJR-GARCH does not improve average forecast accuracy over standard GARCH (DM p = 0.86), yet GJR-based VT produces higher Sharpe ratios, suggesting that tail-specific accuracy drives VT performance. Second, over the 2010–2026 sample (canonical replication), EWMA-based VT changes the Sharpe ratio from 0.799 (buy-and-hold) to 0.701 while reducing maximum drawdown from -33.8% to -21.2% (a 12.6 percentage point reduction); over a common 2020–2026 evaluation period, all VT strategies outperform buy-and-hold. A VIX-proxy formula (8.63/VIX), calibrated via the VIXTWN-to-VIX ratio, enables practical implementation without a local implied volatility index. Combining VIX scaling with a leading indicator direction signal further improves risk-adjusted returns (DM p = 0.0005). A systematic cross-market validation shows that three of four U.S. VT findings generalize or partially generalize to Taiwan (VIX sufficiency, allocation near vol-parity, calendar null results) while one does not (rebalancing frequency). Results are robust to TSMC concentration risk. VaR and Expected Shortfall (ES) backtesting confirms that GJR+HistSim and GJR+Student-t specifications pass both the VaR trinity and Acerbi–Szekely ES tests; VT strategies achieve VaR compliance at 1% but ES severity ratios remain elevated during crisis episodes, motivating tail-risk reserves beyond standard VaR frameworks.
Is Volatility Targeting Just Trend Following? Alpha Absorption and the Leverage Effect Across 22 Assets
Yi-Hao Lai
36 pages | 24 references
Uploaded: 2026-03-21 22:30 · not revised
Abstract▶
Volatility targeting (VT) is often described as repackaged trend following. We examine this claim across 22 assets by regressing the excess returns of a VIX-based VT strategy on the market and an orthogonalized time-series momentum (TSMOM) factor. Seventeen of the 22 assets load significantly on TSMOM. For the four equity assets with a positive CAPM alpha (SPY, QQQ, DIA, XLF), adding TSMOM absorbs between 26.9% and 54.0% of that alpha, although no asset has a CAPM alpha significantly greater than zero (largest t = 1.44). For SPY the loading is significant at every VIX threshold tested (8 to 20) and survives Fama–French five-factor, momentum and betting-against-beta controls. The cross-sectional pattern is tied to the leverage effect: the GJR-GARCH asymmetry parameter γ predicts the TSMOM loading (r = 0.564, p = 0.006, N = 22), and the relation weakens but persists when we control for equity versus non-equity assets (γ_1 = 0.457, HC3 t = 2.03). These are associations, not causal effects, because γ and the loading come from the same return series. We do not test whether VT's drawdown protection survives TSMOM removal. The evidence supports a narrower statement than "VT is trend following": VT carries asset-specific trend exposure that scales with the leverage effect, and that exposure accounts for part of its measured alpha in equity assets.
Volatility Absorption: The Diminishing Marginal Impact of Market Fear
Yi-Hao Lai
50 pages | 45 references
Uploaded: 2026-03-30 18:09 · Revised: 2026-09-06 18:02
Abstract▶
We document a novel empirical regularity: the proportional impact of fear shocks on realized asset returns diminishes as the ambient fear level rises—a phenomenon we term volatility absorption. The qualifier is load-bearing. Estimating the conditional response on all 5,092 days rather than on shock days alone, the level increment β_j is flat across the first four ambient regimes and larger in crisis (1.07, 1.05, 1.04, 1.12, 1.72 percentage points from calm to crisis; calm-minus-high -0.06, HC3 t = -0.38), while the proportional increment β_j/μ_j falls from 2.80 to 1.33 by the high regime. Absorption in this paper is an elasticity, not a marginal effect. Using daily data on four assets (SPY, GLD, TLT, 0050.TW) and VIX from 2006 to 2026, we measure absorption primarily through the shock amplification ratio (SAR)—the ratio of shock-day to normal-day absolute returns within each VIX regime—which removes VIX from the denominator of the dependent variable. We are explicit about the limit of that step: the SAR is still a ratio to a regime baseline that itself rises with VIX, so it measures multiplicative amplification relative to that baseline and does not by itself identify a declining marginal impact. The SAR declines from 3.15× in calm markets (VIX < 15) to 2.33× in high-fear regimes (VIX 25–30), with a slight reversal to 2.45× in crisis regimes (VIX ≥ 30); the calm-to-high decline of 0.82 carries a paired moving-block bootstrap p-value of 0.003. A GARCH null simulation with a contemporaneous volatility proxy shows that pooled null comparisons for this statistic are inconclusive—no single calibration of a GARCH null matches both the level and the daily increments of VIX—while a relative-threshold variant attributes roughly 58% of the pooled headline decline to fixed-threshold shock selection. The absorption claim therefore rests on a conditioning-and-sign decomposition: sorting days by the pre-shock ambient VIX level and restricting to genuine fear shocks (ΔVIX > +2), the calm-to-high SAR decline is +1.05 with a 20-day paired moving-block bootstrap interval of [0.31, 1.76] that excludes zero—but that figure is specific to the fixed 2-point threshold, which selects 2.68% of calm days and 19.15% of crisis days as upward shocks, a 7.2-fold within-regime selection gradient. Under frequency-matched relative and standardized thresholds, whose within-regime selection rates are nearly flat, the same statistic is +0.16 ([-0.45, 0.76]) and +0.58 ([-0.08, 1.20]), both covering zero. Cross-asset regressions under the 2026-04-19 pinned snapshot show negative absorption coefficients for SPY, GLD, and TLT, with SPY only marginally significant and GLD/TLT significant at the 5% level. Event-level evidence from nonfarm payroll releases is directionally consistent, though we note the limitation of using a binary NFP dummy rather than surprise-magnitude controls. The paper is strongest as a reduced-form characterization of absorption and its robustness to alternative normalizers, while richer shock-type and portfolio-design extensions remain deferred pending fully reproducible supporting experiments.
Predictability at the Execution Boundary: Threshold-Crossing Dynamics and the Capacity Ceiling in Taiwan Index Futures
Yi-Hao Lai
18 pages | 4 figures | 7 tables | 7 references
Uploaded: 2026-08-22 20:27 · not revised
Abstract▶
Working paper. Not under submission and not planned for submission — see the closing note. Using the complete tick record of 3,575 TAIFEX TX futures sessions (207,105,249 day-session trades, 2012-2026), we show that intraday threshold-crossing predictability is real and large under the right conditioning, yet unavailable to a liquidity taker for two separately measurable reasons. Execution absorbs 90-107% of the gross edge across thresholds spanning 1 to 160 index points: reversals travel farther than continuations, so the true gross edge is 1.7x to 5.3x a symmetric-payoff approximation, but next-tick fills realise only -0.016 to +0.326 points. Conditioning on incident momentum raises the reversal probability from an unconditional 51% to 89% at a 40-point threshold, yet the median fill available when the signal fires is 2 lots, so realisable value is edge x fillable size x frequency rather than edge x frequency. Decomposing the reversal rate over fourteen years separates a bid-ask bounce component that oscillates without trend (annual range 52.58-57.92%) from a genuinely predictable component that declined monotonically from 55.67% to 50.63%. The ratio of gross edge to execution loss is pinned between 0.66 and 1.07 across a 160-fold range of thresholds, locating a Grossman-Stiglitz equilibrium in a market whose transaction tax is proportional to contract value and therefore quadrupled in point terms as the index rose. No profitable strategy is claimed; the contribution is the measurement. Why it is not being submitted (2026-08-22 owner decision): every angle that would carry a paper-level contribution is either already occupied — the threshold-crossing framework and its count scaling law by Guillaume et al. (1997) and Glattfelder et al. (2011), transient impact and liquidity refill by the propagator literature, post-push reversal in SPY by Vlasiuk and Smirnov (2025), the transaction-tax overlay being that paper's stated future work, and Taiwan's futures transaction tax and market quality by Chou and Wang — or was withdrawn by this study's own tests. The strongest remaining candidate, that the mechanical component of reversal is pinned by the one-point tick size, was refuted by our own day-versus-night identification test: the day-night gap collapses as the threshold widens, the signature of a threshold-scale confound rather than a liquidity effect, and h(1) is not fixed but oscillates between 52.58% and 57.92% year to year. An earlier version of this abstract carried the "held fixed by tick size" wording; it is withdrawn here. The manuscript, the four retractions made during the investigation, and the full eleven-phase evidence base are kept as they stand so the measurement remains citable and reproducible.
Can Anything Beat VIX? A Systematic Out-of-Sample Evaluation of Thirteen Signal Families for Equity Volatility Forecasting and Volatility Timing
Yi-Hao Lai
69 pages | 49 references
Uploaded: 2026-03-31 00:18 · Revised: 2026-10-02 00:03
Abstract▶
We conduct a systematic out-of-sample horse race among thirteen pre-specified signal families—cross-asset volatility momentum, VIX term structure, behavioral sentiment, the variance risk premium, multi-asset portfolio optimization, equal risk contribution, Bitcoin volatility, the yield curve slope, Google Trends fear proxies, overnight VIX changes, calendar anomalies, economic-policy uncertainty, and financial-stress indices—evaluating whether any can improve upon VIX alone for equity volatility forecasting and volatility-timing portfolio construction. Using daily S&P 500 data spanning 1993–2026, we apply a unified forecast pipeline with strict lag enforcement, transaction-cost accounting, publication-delay-corrected treatment of weekly macroeconomic series, the Harvey et al.(2016) multiple-testing threshold (|t|>3.0), the Clark-West test for nested comparisons, and proxy-robust evaluation following Patton(2011). The main finding is strongly negative: not a single signal family produces a statistically significant out-of-sample improvement at the Harvey threshold—whether measured by the standard Campbell and Thompson(2008) R^2_OOS relative to the historical mean, by incremental R^2 relative to VIX-only forecasts, or by Sharpe ratio gains for volatility-timing strategies—and no nested daily comparison approaches the Harvey threshold under the Clark-West test; the nearest, the VIX term-structure family, has a one-sided p-value of 0.0435 at Newey-West lag 21 and 0.0580 at the canonical lag 35 (0.0452 at lag 21 with the Harvey-Leybourne-Newbold small-sample factor), so it straddles the 5% line, and it does not survive a Bonferroni correction across the seven nested families. A power check shows that the Clark-West statistic stays near 2 even for a look-ahead signal and its noise-degraded versions, so for those signal classes the Harvey threshold cannot separate content from noise, and the null rests on the point estimates and the Clark-West p-values rather than on the threshold alone. The VIX–realized volatility R^2 ranges from 0.24 to 0.64 across five non-overlapping eras spanning 33 years (coefficient of variation = 0.33), demonstrating time-invariance. Volatility model rankings are criterion-dependent—GJR-GARCH dominates under proxy-robust QLIKE on r_t^2 (Model Confidence Set sole member) while AMEM dominates for VaR/ES risk management (composite score 1.94 vs. 1.56)—yet VIX sufficiency holds regardless of which model or criterion is used. We frame VIX sufficiency as a rational outcome of option-market information aggregation and show that volatility timing functions as drawdown insurance—costly in Sharpe terms (average drag 3.49%/year for 12/VIX) but welfare-improving for investors with CRRA risk aversion γ ≥ 4.5. The informative null sharpens the research frontier: future gains require either higher-frequency realized measures or genuinely exogenous information not already embedded in option prices.
Multiplicative GARCH-X with VIX: A Parsimonious Alternative to GARCH-MIDAS for Volatility Forecasting
Yi-Hao Lai
73 pages | 31 references
Uploaded: 2026-04-10 15:11 · Revised: 2026-09-03 16:13
Abstract▶
We propose a daily multiplicative GARCH-X model, σ_t^2 = τ_t × g_t, where τ_t is a simple function of lagged VIX and g_t follows a GJR-GARCH process on standardized returns. In a systematic horse race among 17 specifications—including ten daily GARCH-X variants, six GARCH-MIDAS alternatives with Beta-polynomial weighting, and a GJR-GARCH benchmark—the simplest daily model with τ_t = θ_0 + θ_1 VIX_t-1^2 and freely estimated intercept (A4f) achieves the lowest QLIKE loss function, outperforming GJR-GARCH by a margin that clears our conservative |t| > 3.0 reporting screen (Diebold-Mariano t = 4.47; the screen is adopted following the multiple-testing concerns of Harvey et al., 2016 and is not a critical value from any cited source), and outperforming all GARCH-MIDAS specifications in point estimates. The result extends to QQQ (t = 3.96); on the four remaining assets—GLD with its own fear index GVZ (t = 2.93), FEZ (t = 2.74), EEM (t = 2.69) and 0050.TW (t = 1.54)—the improvement is consistent in sign but does not clear our |t| > 3.0 reporting screen. We show that τ_t reflects the option-market fear level while g_t captures residual GARCH dynamics. We also report a null result on the variance risk premium (VRP): across all four specifications examined, g_t shows no economically meaningful contemporaneous correlation with VRP (|ρ| ≤ 0.09), and no out-of-sample predictive power for it. The multiplicative structure is a decomposition of volatility, not a recoverable premium signal. Residual diagnostics on the same pinned sample indicate that VIX absorption reduces excess kurtosis by 42% and shifts the median optimal Student-t degrees of freedom from 5.2 to 8.0; this accompanies the calibration ordering reported below rather than establishing a tail-risk advantage. For risk management the evidence is weaker than the forecasting evidence: under a three-test pass criterion (UC, CC, DQ) A4f-t, A4f-Normal and GJR-Normal each pass at one of three confidence levels and GJR-t at none, so the backtests do not establish a tail-risk advantage, although A4f-t's violation rate is the closest to nominal at every level. These findings demonstrate that a parsimonious multiplicative GARCH-X specification with daily VIX is statistically indistinguishable from the best GARCH-MIDAS alternative (B1, K=22; DM t = -0.90) while outranking longer-lag MIDAS specifications by margins that clear the |t| > 3.0 screen. The case for it therefore rests on parsimony and point-forecast accuracy; the backtests do not add a tail-risk advantage to that case.