{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I'd like to share some additional thoughts on why the current stability metric is not a reliable way to estimate the stability of a model. I know it's too late to change anything, but I believe it may still be of interest to understand the underlying issues.\n\n### TL;DR\n\n- **The penalty term in the metric is not a reliable way to estimate how model performance gets worse over time.** It assumes a predictable, deterministic trend, which can be extrapolated into the future, but the trend in ginis can be stochastic = unpredictable and still greatly penalize a model.\n- **Negative slope can appear without a deterministic trend.** A simple example is a Brownian motion - it has a stochastic trend and can produce a negative slope due to random shocks, but we can't expect the slope to persist in the future. In this case the penalty is not justified. \n- **It's possible to distinguish between a stochastic and a deterministic trend, but the current metric fails to do so.** The host wants to avoid a deterioration in performance outside the evaluation sample, which they believe will occur if they observe a downward slope during the evaluation sample, but that is not true in case of a stochastic trend, which is still penalized by the current metric.\n- **A better approach is to check if the average rate of change in ginis is negative.** Penalties should only be applied if this rate is statistically significant, which would ensure we account for real trends and not random changes.\n\n\nThe corrected approach would be:\n1. Calculate ginis $gini_t$ for each week $t$, and their average value $\\overline{gini_t}$.\n2. Estimate regression $gini_t - gini_{t-1} = a + \\varepsilon_t$, calculate p-value of $a$ and compare it with a chosen confidence level $\\alpha$.\n3. Calculate the final score\n$$\n\\begin{equation}\n\\begin{cases}\n\\text{stability metric} & = \\overline{gini_t} - 0.5\\cdot \\hat{\\sigma} + p\\cdot min(0, \\hat{a}),& \\text{if p-value} \\le \\alpha \\\\\n\\text{stability metric} & = \\overline{gini_t} - 0.5\\cdot \\hat{\\sigma},& \\text{if p-value} > \\alpha \\\\\n\\end{cases}\n\\end{equation}\n$$\nwhere $p$ is a chosen hyperparameter.\n\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport datetime as dt\nimport matplotlib.pyplot as plt\nfrom scipy import stats\nfrom tqdm.auto import tqdm\nfrom sklearn.linear_model import LinearRegression\nimport statsmodels.api as sm\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":false,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:19:40.013471Z","iopub.execute_input":"2024-05-10T19:19:40.013857Z","iopub.status.idle":"2024-05-10T19:19:41.473393Z","shell.execute_reply.started":"2024-05-10T19:19:40.013827Z","shell.execute_reply":"2024-05-10T19:19:41.472178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def calculate_stability_metric(ginis, weeks=None):\n    if weeks is None:\n        weeks = np.arange(1, len(ginis) + 1) + 100\n    \n    ginis = np.array(ginis)\n\n    reg = LinearRegression().fit(\n        weeks.reshape((len(weeks), 1)), ginis\n    )\n\n    a = reg.coef_[0]\n    b = reg.intercept_\n    residuals = ginis - (a * weeks + b)\n    score = np.mean(ginis) + 88 * min(0, a) - 0.5 * np.std(residuals)\n    print(f\"Score: {score:.4f}, mean {np.mean(ginis):.4f}, 88*min(slope, 0) {88*min(0,a):.4f}, 0.5*std {0.5*np.std(residuals):.4f}\")\n    return score","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:19:41.475885Z","iopub.execute_input":"2024-05-10T19:19:41.476599Z","iopub.status.idle":"2024-05-10T19:19:41.487678Z","shell.execute_reply.started":"2024-05-10T19:19:41.476554Z","shell.execute_reply":"2024-05-10T19:19:41.485912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. Intro\n\nThe goal of this competition is to create a model that is stable over time. \n> When looking at the performance of models in production the predictable stable decrease in performance is undesirable even if it is not yet observed, but expected to come. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/482474\n\n\nTo evaluate the stability of the model, a special metric has been proposed, with a penalty added if there is a negative slope in the regression of ginis over time. The question is: **if the slope is negative, does it mean that there is a *predictable stable decrease in performance***?\n\nLet's formalize this.\n\n\n## 2. Determenistic trend\nAssume that the ginis over time, $X_t$, follow Brownian motion with a drift $\\mu$. The drift term represents a predictable stable decrease in ginis over time.\n\n$$\n\\begin{equation}\n\\begin{split}\nX_t & = \\mu t + \\sigma B_t \\\\\ndX_t & = \\mu dt + \\sigma dB_t \\\\\n\\end{split}\n\\end{equation}\n$$\n\nWhere:\n\n- $dX_t$ is the change in the process $X$ at time $t$\n- $\\mu$ is the drift coefficient\n- $\\sigma$ is the diffusion coefficient (or volatility)\n- $W_t$ is a standard Brownian motion\n\nThat is, differences of $X_t$ are normally distributed with mean $\\mu$ and standard deviation $\\sigma$.\n\nLet's simulate it with \n- Number of weeks $n=100$ (as we have approximately 90 weeks in the train dataset, and the test dataset is about 90% of the train dataset. However, the number of weeks may differ from 80 due to the varying distribution of observations across weeks).\n- Initial value $X_0=0.8$, to resemble real ginis.\n- Drift $\\mu = -0.001$, indicating a negative slope.","metadata":{}},{"cell_type":"code","source":"np.random.seed(2)\n\nX0 = 0.8\nn = 100\n\nX_diff = np.random.normal(-0.001, 0.005, size=n) # normal distribution with negative mean\n\nX = np.cumsum(X_diff)+X0\nt = np.arange(n)\n\nfig, ax = plt.subplots(1, 2, figsize=(14, 5))\nax[0].plot(t, X)\nax[0].set_title(\"Ginis with a linear trend ($X_t$)\")\n\nax[1].plot(t, X_diff)\nax[1].set_title(\"Differences of ginis ($X_t-X_{t-1}$).\" + f\" Mean {np.mean(X_diff):.4f}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:19:41.489257Z","iopub.execute_input":"2024-05-10T19:19:41.490603Z","iopub.status.idle":"2024-05-10T19:19:42.24488Z","shell.execute_reply.started":"2024-05-10T19:19:41.490543Z","shell.execute_reply":"2024-05-10T19:19:42.243517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's estimate the slope and calculate the stability metric value:","metadata":{}},{"cell_type":"code","source":"score = calculate_stability_metric(X, t)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:19:42.247676Z","iopub.execute_input":"2024-05-10T19:19:42.248202Z","iopub.status.idle":"2024-05-10T19:19:42.257427Z","shell.execute_reply.started":"2024-05-10T19:19:42.248155Z","shell.execute_reply":"2024-05-10T19:19:42.255966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The average gini value is $0.71$, and the final score is $0.58$, mainly due to the slope term equal to $−0.13$.","metadata":{}},{"cell_type":"markdown","source":"### 2.1 Estimation of a trend via regression $X_t = \\alpha + \\beta t + \\varepsilon_t$\nLet’s examine the slope coefficient we obtained. I'll use the 'statsmodels' library to generate a report on a regression of our simulated ginis on weeks.","metadata":{}},{"cell_type":"code","source":"t_df = pd.DataFrame(sm.add_constant(t), columns=[\"const\", \"slope\"])\nols = sm.OLS(X, t_df)\nresults = ols.fit()\nresults.summary()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:19:42.259298Z","iopub.execute_input":"2024-05-10T19:19:42.260106Z","iopub.status.idle":"2024-05-10T19:19:42.299337Z","shell.execute_reply.started":"2024-05-10T19:19:42.259978Z","shell.execute_reply":"2024-05-10T19:19:42.298026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The slope coefficient is significantly negative ($P>|t| = 0.000$). \n\n### 2.2 Estimation of a trend via regression $dX_t = \\alpha dt + \\varepsilon_t$\n\nWe know that if there's a drift in the time series (i.e., a linear deterministic trend), then the differences of our series should have an average value statistically different from zero. Let's regress ginis first differences on a constant and check the significance of the coefficient:","metadata":{}},{"cell_type":"code","source":"ols = sm.OLS(X_diff, np.ones(len(X_diff)))\nresults = ols.fit()\nresults.summary()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:19:42.300959Z","iopub.execute_input":"2024-05-10T19:19:42.301847Z","iopub.status.idle":"2024-05-10T19:19:42.335581Z","shell.execute_reply.started":"2024-05-10T19:19:42.301813Z","shell.execute_reply":"2024-05-10T19:19:42.334214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The coefficient is significant (p-value $0.004$), as expected, because we simulated our series with $\\mu < 0$.\n\nTherefore, we have a **deterministic trend**: a consistent, *predictable change over time*. If we are to predict $X_{t+n}$, it will be lower than $X_t$, due to the negative linear trend significantly influencing $X_t$ dynamics: \n$$\nE(X_{t+n} | X_{t}) < X_{t}\n$$\nIn simpler terms, if there’s a deterministic linear trend, then if we observe a decrease in ginis over time, we can confidently say that ginis will continue to decrease outside of our evaluation sample. **This is what the host aims to estimate: they want a model that doesn’t produce a determenistic trend in ginis over time**, so that if we predict the gini value 100 weeks ahead, it won't differ from the gini value in the evaluation sample.\n\n## 3. Stochastic trend\n\nNow, let's move to another important concept, the **stochastic trend**: a trend arising from the *accumulation of random shocks*. There’s a crucial difference between stochastic and deterministic trends: a deterministic trend is predictable, while a stochastic trend is **not**. If there’s a stochastic trend and we observe a decrease in scores, we cannot extrapolate and say that the scores will continue to decrease, because the scores change due to the influence of random shocks. That is, for $B_t$ following Brownian motion without drift, \n$$\nE(B_{t+n} | B_{t}) = B_{t}\n$$\n\nLet's simulate this$^1$.","metadata":{}},{"cell_type":"code","source":"seed = 30\nnp.random.seed(seed)\n\nB_diff = np.random.normal(0, 0.013, size=n) # Zero drift\nB = np.cumsum(B_diff)+X0\n\nfig, ax = plt.subplots(1, 2, figsize=(14, 5))\nax[0].plot(t, B)\nax[0].set_title(\"Ginis with a stochastic trend ($B_t$)\")\n\nax[1].plot(t, B_diff)\nax[1].set_title(\"Differences of ginis ($B_t-B_{t-1}$).\" + f\" Mean {np.mean(B_diff):.4f}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:24:37.879867Z","iopub.execute_input":"2024-05-10T19:24:37.880367Z","iopub.status.idle":"2024-05-10T19:24:38.475443Z","shell.execute_reply.started":"2024-05-10T19:24:37.88033Z","shell.execute_reply":"2024-05-10T19:24:38.474204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score = calculate_stability_metric(B, t)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:24:42.865454Z","iopub.execute_input":"2024-05-10T19:24:42.86588Z","iopub.status.idle":"2024-05-10T19:24:42.874646Z","shell.execute_reply.started":"2024-05-10T19:24:42.865845Z","shell.execute_reply":"2024-05-10T19:24:42.873268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The average gini is almost the same as in our previous example, $0.73$, and we have an even larger negative slope of $−0.18$ and a final score of $0.54$. But there’s no deterministic trend! We shouldn’t penalize the metric because the trend is stochastic and not predictable! Let’s extend our sample and generate 100 more observations with the same seed to see what happens to our scores in the \"future.\" Do they reduce further?","metadata":{}},{"cell_type":"code","source":"np.random.seed(seed)\n\nn2 = 200\n\nt2 = np.arange(n2)\nB2_diff = np.random.normal(0, 0.013, size=n2) # Zero drift\nB2 = np.cumsum(B2_diff)+X0\n\nfig, ax = plt.subplots(1, 2, figsize=(14, 5))\nax[0].plot(t2, B2)\nax[0].set_title(\"Ginis with a stochastic trend ($B_t$)\")\n\nax[1].plot(t2, B2_diff)\nax[1].set_title(\"Differences of ginis ($B_t-B_{t-1}$).\" + f\" Mean {np.mean(B2_diff):.4f}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:25:37.014475Z","iopub.execute_input":"2024-05-10T19:25:37.014936Z","iopub.status.idle":"2024-05-10T19:25:37.657282Z","shell.execute_reply.started":"2024-05-10T19:25:37.014902Z","shell.execute_reply":"2024-05-10T19:25:37.656088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"No, gini value returns back to $0.8$ after a drop to $0.6$. **The trend we observed in the first 100 weeks was of stochastic nature, so we can’t predict it**. The penalty can surely be added, but not for the slope, rather for the high variance of the scores!\n\n### 3.1 Estimation of a trend via regression $X_t = \\alpha + \\beta t + \\varepsilon_t$\n\nBut how can we determine if the trend is predictable or not? Let’s check the statistical significance of the slope in our current example:","metadata":{}},{"cell_type":"code","source":"ols = sm.OLS(B, t_df)\nresults = ols.fit()\nresults.summary()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:28:14.118653Z","iopub.execute_input":"2024-05-10T19:28:14.119085Z","iopub.status.idle":"2024-05-10T19:28:14.156623Z","shell.execute_reply.started":"2024-05-10T19:28:14.119052Z","shell.execute_reply":"2024-05-10T19:28:14.155201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The slope is significant ($P>∣t∣=0.000$). That’s because we fitted the wrong model: we are trying to fit a series with a stochastic trend with a linear model. \n\n### 3.2 Estimation of a trend via regression $dX_t = \\alpha dt + \\varepsilon_t$\n\nIf there is a determenistic trend, the ginis should change with a negative rate, i.e. the differences should have an average value statistically different from zero. Our first simulated series had it. Let’s run a regression of differences on a constant for our new series:","metadata":{}},{"cell_type":"code","source":"ols = sm.OLS(B_diff, np.ones(len(B_diff)))\nresults = ols.fit()\nresults.summary()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-10T19:28:59.365286Z","iopub.execute_input":"2024-05-10T19:28:59.365723Z","iopub.status.idle":"2024-05-10T19:28:59.399943Z","shell.execute_reply.started":"2024-05-10T19:28:59.365689Z","shell.execute_reply":"2024-05-10T19:28:59.398779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The p-value is $0.304$, indicating no significant linear trend in the data.\n\n## 4. Conclusion\n\nIf one wants to estimate whether the performance of a model deteriorates over time, they should **add a penalty only if the gini differences over time have a statistically significant average value**. As we don’t have access to the test data, it's impossible to determine if there is one, but as shown, it is entirely possible that the slope is significant and the penalty in the current gini stability score is substantial ($0.1–0.2$), yet the performance of the model may return to normal after a drop.\n\nThe corrected approach would be:\n1. Calculate ginis $gini_t$ for each week $t$, and their average value $\\overline{gini_t}$.\n2. Estimate regression $gini_t - gini_{t-1} = a + \\varepsilon_t$, calculate p-value of $a$ and compare it with a chosen confidence level $\\alpha$.\n3. Calculate the final score\n$$\n\\begin{equation}\n\\begin{cases}\n\\text{stability metric} & = \\overline{gini_t} - 0.5\\cdot \\hat{\\sigma} + p\\cdot min(0, \\hat{a}),& \\text{if p-value} \\le \\alpha \\\\\n\\text{stability metric} & = \\overline{gini_t} - 0.5\\cdot \\hat{\\sigma},& \\text{if p-value} > \\alpha \\\\\n\\end{cases}\n\\end{equation}\n$$\nwhere $p$ is a chosen hyperparameter.\n\n\n\n$^1$ I've chosen a random seed and a variance so that the slope is negative and the numbers are close to the ones from the first example to highlight the issue. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}