{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9849268,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Credits\n\nThis analysis is built upon the foundational work by [@carlmcbrideellis](https://www.kaggle.com/carlmcbrideellis) from the JS 2021 competition, who published the acclaimed notebook titled **\"Jane Street: EDA of Day 0 and Feature Importance.\"** I want to extend my gratitude to Carl for his insightful contribution.\n\nIn this notebook, I have adapted and utilized Carl's original code to analyze this year's Jane Street competition data. I am employing the **Dask** library instead of the original **PyArrow** framework used by Carl. ","metadata":{}},{"cell_type":"markdown","source":"![](http://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcTp5YjQGHIHm4i2W_5ES846YKjaLcNL78DJLw&s)\n\n# Jane Street Market Prediction: A simple EDA\n\n> \"*Machine learning (ML) at Jane Street begins, unsurprisingly, with data. We collect and store around 2.3TB of market data every day. Hidden in those petabytes of data are the relationships and statistical regularities which inform the models inside our strategies. But it’s not just awesome models. ML work in a production environment like Jane Street’s involves many interconnected pieces.*\" -- [Jane Street Tech Blog \"*Real world machine learning*\"](https://blog.janestreet.com/real-world-machine-learning-part-1/).\n\nThis notebook is a simple exploratory data analysis (EDA) of the files provided for the kaggle [Jane Street Market Prediction](https://www.kaggle.com/c/jane-street-market-prediction) competition. Here we shall...\n\n> \"**Explore the data:** *It’s hard to know what techniques to throw at a problem before we understand what the data looks like, and indeed figure out what data to use. Spending the time to visualize and understand the structure of the problem helps pick the right modeling tools for the job. Plus, pretty plots are catnip to traders and researchers!*\"\n\n## <center style=\"background-color:Gainsboro; width:40%;\">Contents</center>\n* [The train.csv file is big](#train_csv)\n* [responders](#resp)\n* [weight](#weight)\n* [Cumulative return](#return)\n* [Time](#time)\n* [The `features.csv` file](#features_file)\n* [The first day (\"day 0\")](#day_0)\n* [Are there any missing values?](#missing_values)\n* [Is there any missing data: Days 2 and 294](#missing_data)\n* [Permutation Importance using the Random Forest](#permutation)\n* [Is there any correlation between day 100 and day 200?](#Pearson)","metadata":{}},{"cell_type":"code","source":"!pip install dask","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:22:18.620910Z","iopub.execute_input":"2024-10-15T05:22:18.621391Z","iopub.status.idle":"2024-10-15T05:22:34.344798Z","shell.execute_reply.started":"2024-10-15T05:22:18.621318Z","shell.execute_reply":"2024-10-15T05:22:34.343247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# numpy\nimport numpy as np\n\n# pandas stuff\nimport pandas as pd\npd.set_option('display.max_rows', None)\npd.set_option('display.max_columns', None)\n\n# plotting stuff\nfrom pandas.plotting import lag_plot\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.graph_objects as go\ncolorMap = sns.light_palette(\"blue\", as_cmap=True)\n#plt.rcParams.update({'font.size': 12})\n\n\n# install dabl\n#!pip install dabl > /dev/null\n#import dabl\n# install datatable\n#!pip install datatable > /dev/null\n#import datatable as dt\n\n# misc\nimport missingno as msno\n\n# system\nimport warnings\nwarnings.filterwarnings('ignore')\n# for the image import\nimport os\nfrom IPython.display import Image\n# garbage collector to keep RAM in check\nimport gc  ","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:22:34.347377Z","iopub.execute_input":"2024-10-15T05:22:34.347794Z","iopub.status.idle":"2024-10-15T05:22:36.476707Z","shell.execute_reply.started":"2024-10-15T05:22:34.347744Z","shell.execute_reply":"2024-10-15T05:22:36.475504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport dask.dataframe as dd\ntrain_data = dd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet')\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:22:36.477995Z","iopub.execute_input":"2024-10-15T05:22:36.478570Z","iopub.status.idle":"2024-10-15T05:22:41.347193Z","shell.execute_reply.started":"2024-10-15T05:22:36.478529Z","shell.execute_reply":"2024-10-15T05:22:41.345848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compute the number of rows\nrow_count = len(train_data)\n\nprint(f\"Total number of rows: {row_count}\")","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:22:41.348695Z","iopub.execute_input":"2024-10-15T05:22:41.349322Z","iopub.status.idle":"2024-10-15T05:22:41.376331Z","shell.execute_reply.started":"2024-10-15T05:22:41.349278Z","shell.execute_reply":"2024-10-15T05:22:41.375022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that it has a total of 4,71,27,338 rows.\n\nI have used dask to load in the `train.parquet` and it took almost 5.31 sec.","metadata":{}},{"cell_type":"code","source":"# Count the unique 'date_id' values\nunique_dates = train_data['date_id'].nunique().compute()\n\nprint(f\"Number of unique days (date_id): {unique_dates}\")","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:22:41.379261Z","iopub.execute_input":"2024-10-15T05:22:41.379693Z","iopub.status.idle":"2024-10-15T05:22:43.126763Z","shell.execute_reply.started":"2024-10-15T05:22:41.379651Z","shell.execute_reply":"2024-10-15T05:22:43.125457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"return\"></a>\n## <center style=\"background-color:Gainsboro; width:40%;\">resp</center>\n\nThere are a total of 1699 days of data in `train.parquet` (*i.e.* more than four and a half years of trading data). Let us take a look at the cumulative values of `responders` over time. \n\n> \"*The longer the Time Horizon, the more aggressive, or riskier portfolio, an investor can build. The shorter the Time Horizon, the more conservative, or less risky, the investor may want to adopt.*\"","metadata":{}},{"cell_type":"markdown","source":"The cumulative sums appear to be starting from a negative value (or a much lower cumulative sum), indicating that the total responses have been negative right from the beginning. This suggests that there may have been more negative trades or less positive performance overall.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15, 5))\n#balance= pd.Series(train_data['resp']).cumsum()\nresp_1= pd.Series(train_data['responder_1']).cumsum()\nresp_2= pd.Series(train_data['responder_2']).cumsum()\nresp_3= pd.Series(train_data['responder_3']).cumsum()\nresp_4= pd.Series(train_data['responder_4']).cumsum()\nresp_5 = pd.Series(train_data['responder_5']).cumsum()\nresp_6 = pd.Series(train_data['responder_6']).cumsum()\nresp_7 = pd.Series(train_data['responder_7']).cumsum()\nresp_8 = pd.Series(train_data['responder_8']).cumsum()\nax.set_xlabel (\"Trade\", fontsize=18)\nax.set_title (\"Cumulative resp and time horizons 1, 2, 3,4,5,6,7 and 8 (1699 days)\", fontsize=18)\n#balance.plot(lw=3)\nresp_1.plot(lw=3, label ='responder_1')\nresp_2.plot(lw=3, label ='responder_2')\nresp_3.plot(lw=3, label ='responder_3')\nresp_4.plot(lw=3, label ='responder_4')\nresp_5.plot(lw=3, label ='responder_5')\nresp_6.plot(lw=3,label ='responder_6')\nresp_7.plot(lw=3, label ='responder_7')\nresp_8.plot(lw=3, label ='responder_8')\nplt.legend(loc=\"upper left\");\ndel resp_1\ndel resp_2\ndel resp_3\ndel resp_4\ndel resp_5\ndel resp_6\ndel resp_7\ndel resp_8\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:25:09.173447Z","iopub.execute_input":"2024-10-15T05:25:09.173890Z","iopub.status.idle":"2024-10-15T05:26:32.128576Z","shell.execute_reply.started":"2024-10-15T05:25:09.173847Z","shell.execute_reply":"2024-10-15T05:26:32.127181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that `responder_6` (in brown) most closely follows time horizon 8 (`responder_8` is the uppermost curve, in grey). \n","metadata":{}},{"cell_type":"markdown","source":"Let us now plot a histogram of the `responder_6` values (here only shown for values between -5 and 5)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (12,5))\nax = sns.distplot(train_data['responder_6'], \n             bins=3000, \n             kde_kws={\"clip\":(-5,5)}, \n             hist_kws={\"range\":(-5,5)},\n             color='darkcyan', \n             kde=False);\nvalues = np.array([rec.get_height() for rec in ax.patches])\nnorm = plt.Normalize(values.min(), values.max())\ncolors = plt.cm.jet(norm(values))\nfor rec, col in zip(ax.patches, colors):\n    rec.set_color(col)\nplt.xlabel(\"Histogram of the responder_6 values\", size=14)\nplt.show();\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:27:44.362569Z","iopub.execute_input":"2024-10-15T05:27:44.363067Z","iopub.status.idle":"2024-10-15T05:27:52.785532Z","shell.execute_reply.started":"2024-10-15T05:27:44.363019Z","shell.execute_reply":"2024-10-15T05:27:52.784231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assuming train_data is a Dask DataFrame\n# Compute the min and max values for responder_6\nmin_resp = train_data['responder_6'].min().compute()\nprint('The minimum value for resp is: %.5f' % min_resp)\n\nmax_resp = train_data['responder_6'].max().compute()\nprint('The maximum value for resp is: %.5f' % max_resp)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:31:02.653123Z","iopub.execute_input":"2024-10-15T05:31:02.653740Z","iopub.status.idle":"2024-10-15T05:31:04.668332Z","shell.execute_reply.started":"2024-10-15T05:31:02.653688Z","shell.execute_reply":"2024-10-15T05:31:04.667103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us also calculate the [skew](https://en.wikipedia.org/wiki/Skewness) and [kurtosis](https://en.wikipedia.org/wiki/Kurtosis) of this distribution:","metadata":{}},{"cell_type":"code","source":"# Compute and print the skewness and kurtosis of responder_6\nskew_resp = train_data['responder_6'].skew().compute()\nkurt_resp = train_data['responder_6'].kurtosis().compute()\n\nprint(\"Skew of resp is:      %.2f\" % skew_resp)\nprint(\"Kurtosis of resp is: %.2f\"  % kurt_resp)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:33:11.202056Z","iopub.execute_input":"2024-10-15T05:33:11.202569Z","iopub.status.idle":"2024-10-15T05:33:15.581702Z","shell.execute_reply.started":"2024-10-15T05:33:11.202522Z","shell.execute_reply":"2024-10-15T05:33:15.580443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, let us fit a [Cauchy distribution](https://en.wikipedia.org/wiki/Cauchy_distribution) to this data","metadata":{}},{"cell_type":"code","source":"from scipy.optimize import curve_fit\n# the values\nx = list(range(len(values)))\nx = [((i)-1500)/30000 for i in x]\ny = values\n\ndef Lorentzian(x, x0, gamma, A):\n    return A * gamma**2/(gamma**2+( x - x0 )**2)\n\n# seed guess\ninitial_guess=(0, 0.001, 3000)\n\n# the fit\nparameters,covariance=curve_fit(Lorentzian,x,y,initial_guess)\nsigma=np.sqrt(np.diag(covariance))\n\n# and plot\nplt.figure(figsize = (12,5))\nax = sns.distplot(train_data['responder_6'], \n             bins=3000, \n             kde_kws={\"clip\":(-5,5)}, \n             hist_kws={\"range\":(-5,5)},\n             color='darkcyan', \n             kde=False);\nvalues = np.array([rec.get_height() for rec in ax.patches])\n#norm = plt.Normalize(values.min(), values.max())\n#colors = plt.cm.jet(norm(values))\n#for rec, col in zip(ax.patches, colors):\n#    rec.set_color(col)\nplt.xlabel(\"Histogram of the resp values\", size=14)\nplt.plot(x,Lorentzian(x,*parameters),'--',color='black',lw=3)\nplt.show();\ndel values\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:36:00.000585Z","iopub.execute_input":"2024-10-15T05:36:00.001133Z","iopub.status.idle":"2024-10-15T05:36:08.443184Z","shell.execute_reply.started":"2024-10-15T05:36:00.001084Z","shell.execute_reply":"2024-10-15T05:36:08.441941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that a Cauchy distribution can be generated from the ratio of two independent normally distributed random variables with mean zero. The paper by [David E. Harris \"*The Distribution of Returns*\"](https://www.scirp.org/pdf/JMF_2017083015172459.pdf) goes into detail regarding the use of a Cauchy distribution to model returns.\n\n<a class=\"anchor\" id=\"weight\"></a>\n## <center style=\"background-color:Gainsboro; width:40%;\">weight</center>\n\n> *Each trade has an associated `weight` and `resp`, which together represents a return on the trade.*","metadata":{}},{"cell_type":"code","source":"# Calculate the number of zero weights\nnum_zero_weights = (train_data['weight'] == 0).sum().compute()  # Compute the total number of zero weights\ntotal_weights = train_data.shape[0].compute()  # Compute the total number of rows\n\n# Calculate the percentage of zero weights\npercent_zeros = (100 * num_zero_weights) / total_weights\n\nprint('Percentage of zero weights is: %.2f%%' % percent_zeros)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:38:16.730319Z","iopub.execute_input":"2024-10-15T05:38:16.731471Z","iopub.status.idle":"2024-10-15T05:38:18.234929Z","shell.execute_reply.started":"2024-10-15T05:38:16.731338Z","shell.execute_reply":"2024-10-15T05:38:18.233745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us see if there are any negative weights. A negative weight would be meaningless, but you never know...","metadata":{}},{"cell_type":"code","source":"min_weight = train_data['weight'].min().compute()\nprint('The minimum weight is: %.2f' % min_weight)","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:40:17.438397Z","iopub.execute_input":"2024-10-15T05:40:17.438864Z","iopub.status.idle":"2024-10-15T05:40:18.309225Z","shell.execute_reply.started":"2024-10-15T05:40:17.438819Z","shell.execute_reply":"2024-10-15T05:40:18.307970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And now to find the maximum weight used","metadata":{}},{"cell_type":"code","source":"max_weight = train_data['weight'].max().compute()\nprint('The maximum weight was: %.2f' % max_weight)","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:40:57.259853Z","iopub.execute_input":"2024-10-15T05:40:57.260304Z","iopub.status.idle":"2024-10-15T05:40:58.130063Z","shell.execute_reply.started":"2024-10-15T05:40:57.260264Z","shell.execute_reply":"2024-10-15T05:40:58.128488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us take a look at a histogram of the non-zero weights","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (12,5))\nax = sns.distplot(train_data['weight'], \n             bins=1400, \n             kde_kws={\"clip\":(0.15,10.24)}, \n             hist_kws={\"range\":(0.15,10.24)},\n             color='darkcyan', \n             kde=False);\nvalues = np.array([rec.get_height() for rec in ax.patches])\nnorm = plt.Normalize(values.min(), values.max())\ncolors = plt.cm.jet(norm(values))\nfor rec, col in zip(ax.patches, colors):\n    rec.set_color(col)\nplt.xlabel(\"Histogram of non-zero weights\", size=14)\nplt.show();\ndel values\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:45:36.212333Z","iopub.execute_input":"2024-10-15T05:45:36.212863Z","iopub.status.idle":"2024-10-15T05:45:41.827617Z","shell.execute_reply.started":"2024-10-15T05:45:36.212802Z","shell.execute_reply":"2024-10-15T05:45:41.826441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There appear to be one peak. Could this be indicative of a single underlying distribution that we see here. \nWe can plot the logarithm of the weights (*Credit*: [\"*Target Engineering; CV; ⚡ Multi-Target*\"](https://www.kaggle.com/marketneutral/target-engineering-cv-multi-target) by [marketneutral](https://www.kaggle.com/marketneutral))","metadata":{}},{"cell_type":"code","source":"train_data_nonZero = train_data.query('weight > 0').reset_index(drop = True)\nplt.figure(figsize = (10,4))\nax = sns.distplot(np.log(train_data_nonZero['weight']), \n             bins=1000, \n             kde_kws={\"clip\":(-4,5)}, \n             hist_kws={\"range\":(-4,5)},\n             color='darkcyan', \n             kde=False);\nvalues = np.array([rec.get_height() for rec in ax.patches])\nnorm = plt.Normalize(values.min(), values.max())\ncolors = plt.cm.jet(norm(values))\nfor rec, col in zip(ax.patches, colors):\n    rec.set_color(col)\nplt.xlabel(\"Histogram of the logarithm of the non-zero weights\", size=14)\nplt.show();\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:48:05.516906Z","iopub.execute_input":"2024-10-15T05:48:05.517429Z","iopub.status.idle":"2024-10-15T05:49:32.590448Z","shell.execute_reply.started":"2024-10-15T05:48:05.517386Z","shell.execute_reply":"2024-10-15T05:49:32.589111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"and we can now try to fit a pair of Gaussian functions to this distribution","metadata":{}},{"cell_type":"code","source":"from scipy.optimize import curve_fit\n# the values\nx = list(range(len(values)))\nx = [(i/110)-4 for i in x]\ny = values\n\n# define a Gaussian function\ndef Gaussian(x,mu,sigma,A):\n    return A*np.exp(-0.5 * ((x-mu)/sigma)**2)\n\ndef bimodal(x,mu_1,sigma_1,A_1,mu_2,sigma_2,A_2):\n    return Gaussian(x,mu_1,sigma_1,A_1) + Gaussian(x,mu_2,sigma_2,A_2)\n\n# seed guess\ninitial_guess=(1, 1 , 1,    1, 1, 1)\n\n# the fit\nparameters,covariance=curve_fit(bimodal,x,y,initial_guess)\nsigma=np.sqrt(np.diag(covariance))\n\n# the plot\nplt.figure(figsize = (10,4))\nax = sns.distplot(np.log(train_data_nonZero['weight']), \n             bins=1000, \n             kde_kws={\"clip\":(-5,5)}, \n             hist_kws={\"range\":(-5,5)},\n             color='darkcyan', \n             kde=False);\nvalues = np.array([rec.get_height() for rec in ax.patches])\nnorm = plt.Normalize(values.min(), values.max())\ncolors = plt.cm.jet(norm(values))\nfor rec, col in zip(ax.patches, colors):\n    rec.set_color(col)\nplt.xlabel(\"Histogram of the logarithm of the non-zero weights\", size=14)\n# plot gaussian #1\nplt.plot(x,Gaussian(x,parameters[0],parameters[1],parameters[2]),':',color='black',lw=2,label='Gaussian #1', alpha=0.8)\n# plot gaussian #2\nplt.plot(x,Gaussian(x,parameters[3],parameters[4],parameters[5]),'--',color='black',lw=2,label='Gaussian #2', alpha=0.8)\n# plot the two gaussians together\nplt.plot(x,bimodal(x,*parameters),color='black',lw=2, alpha=0.7)\nplt.legend(loc=\"upper left\");\nplt.show();\ndel values\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:50:37.207025Z","iopub.execute_input":"2024-10-15T05:50:37.208135Z","iopub.status.idle":"2024-10-15T05:51:58.675183Z","shell.execute_reply.started":"2024-10-15T05:50:37.208086Z","shell.execute_reply":"2024-10-15T05:51:58.674053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"return\"></a>\n## <center style=\"background-color:Gainsboro; width:40%;\">Cumulative return</center>\n\nLet us take a look at the cumulative daily return over time, which is given by `weight` multiplied by the value of `responders`","metadata":{}},{"cell_type":"code","source":"train_data['weight_resp_1'] = train_data['weight']*train_data['responder_1']\ntrain_data['weight_resp_2'] = train_data['weight']*train_data['responder_2']\ntrain_data['weight_resp_3'] = train_data['weight']*train_data['responder_3']\ntrain_data['weight_resp_4'] = train_data['weight']*train_data['responder_4']\ntrain_data['weight_resp_5'] = train_data['weight']*train_data['responder_5']\ntrain_data['weight_resp_6'] = train_data['weight']*train_data['responder_6']\ntrain_data['weight_resp_7'] = train_data['weight']*train_data['responder_7']\ntrain_data['weight_resp_8'] = train_data['weight']*train_data['responder_8']\nfig, ax = plt.subplots(figsize=(15, 5))\nresp_1  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_1'].mean())).cumprod()\nresp_2  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_2'].mean())).cumprod()\nresp_3  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_3'].mean())).cumprod()\nresp_4  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_4'].mean())).cumprod()\nresp_5  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_5'].mean())).cumprod()\nresp_6  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_6'].mean())).cumprod()\nresp_7  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_7'].mean())).cumprod()\nresp_8  = pd.Series(1+(train_data.groupby('date_id')['weight_resp_8'].mean())).cumprod()\nax.set_xlabel (\"Day\", fontsize=18)\nax.set_title (\"Cumulative daily return for resp and time horizons 1, 2, 3, 4, 5, 6, 7 and 8 (1699 days)\", fontsize=18)\n#resp.plot(lw=3, label='resp x weight')\nresp_1.plot(lw=3, label='resp_1 x weight')\nresp_2.plot(lw=3, label='resp_2 x weight')\nresp_3.plot(lw=3, label='resp_3 x weight')\nresp_4.plot(lw=3, label='resp_4 x weight')\nresp_5.plot(lw=3, label='resp_5 x weight')\nresp_6.plot(lw=3, label='resp_6 x weight')\nresp_7.plot(lw=3, label='resp_7 x weight')\nresp_8.plot(lw=3, label='resp_8 x weight')\n# day 85 marker\nax.axvline(x=85, linestyle='--', alpha=0.3, c='red', lw=1)\nax.axvspan(0, 85 , color=sns.xkcd_rgb['grey'], alpha=0.1)\nplt.legend(loc=\"lower left\");","metadata":{"execution":{"iopub.status.busy":"2024-10-15T05:57:44.121524Z","iopub.execute_input":"2024-10-15T05:57:44.122020Z","iopub.status.idle":"2024-10-15T05:58:55.441079Z","shell.execute_reply.started":"2024-10-15T05:57:44.121978Z","shell.execute_reply":"2024-10-15T05:58:55.439821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that `resp_5`, representing a more conservative strategy, result in the highest return.\n\nWe shall now plot a histogram of the `weight` multiplied by the value of `responder 6`","metadata":{}},{"cell_type":"code","source":"train_data_no_0 = train_data.query('weight > 0').reset_index(drop = True)\ntrain_data_no_0['wAbsResp'] = train_data_no_0['weight'] * (train_data_no_0['responder_6'])\n#plot\nplt.figure(figsize = (12,5))\nax = sns.distplot(train_data_no_0['wAbsResp'], \n             bins=1500, \n             kde_kws={\"clip\":(-5,5)}, \n             hist_kws={\"range\":(-5,5)},\n             color='darkcyan', \n             kde=False);\nvalues = np.array([rec.get_height() for rec in ax.patches])\nnorm = plt.Normalize(values.min(), values.max())\ncolors = plt.cm.jet(norm(values))\nfor rec, col in zip(ax.patches, colors):\n    rec.set_color(col)\nplt.xlabel(\"Histogram of the weights * responder_6\", size=14)\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:03:34.490061Z","iopub.execute_input":"2024-10-15T06:03:34.490589Z","iopub.status.idle":"2024-10-15T06:04:18.413903Z","shell.execute_reply.started":"2024-10-15T06:03:34.490544Z","shell.execute_reply":"2024-10-15T06:04:18.412701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"time\"></a>\n## <center style=\"background-color:Gainsboro; width:40%;\">Time</center>\nLet us plot the number of `time_id` per day.","metadata":{}},{"cell_type":"code","source":"# Group by 'date_id' and count 'time_id'\ntrades_per_day = train_data.groupby(['date_id'])['time_id'].count().compute()  # Compute the result\n\n# Create the plot\nfig, ax = plt.subplots(figsize=(15, 5))\nplt.plot(trades_per_day.index, trades_per_day.values)  # Plot with x (day) and y (count)\nax.set_xlabel(\"Day\", fontsize=18)\nax.set_title(\"Total number of time_ids for each day\", fontsize=18)\n\n# Add a vertical line and shaded region for day 85\nax.axvline(x=85, linestyle='--', alpha=0.3, c='red', lw=1)\nax.axvspan(0, 85, color='gray', alpha=0.1)  # Changed from sns.xkcd_rgb['grey'] to 'gray' for simplicity\n\n# Set x-axis limits\nax.set_xlim(xmin=0, xmax=1800)\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:06:21.359259Z","iopub.execute_input":"2024-10-15T06:06:21.360421Z","iopub.status.idle":"2024-10-15T06:06:23.050336Z","shell.execute_reply.started":"2024-10-15T06:06:21.360368Z","shell.execute_reply":"2024-10-15T06:06:23.048919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we assume a [trading day](https://en.wikipedia.org/wiki/Trading_day) is 6½ hours long (*i.e.* 23400 seconds) then","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15, 5))\nplt.plot(23400/trades_per_day)\nax.set_xlabel (\"Day\", fontsize=18)\nax.set_ylabel (\"Av. time between trades (s)\", fontsize=18)\nax.set_title (\"Average time between trades for each day\", fontsize=18)\nax.axvline(x=85, linestyle='--', alpha=0.3, c='red', lw=1)\nax.axvspan(0, 85 , color=sns.xkcd_rgb['grey'], alpha=0.1)\nax.set_xlim(xmin=0)\nax.set_xlim(xmax=500)\nax.set_ylim(ymin=0)\nax.set_ylim(ymax=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:22:29.840051Z","iopub.execute_input":"2024-10-15T06:22:29.840671Z","iopub.status.idle":"2024-10-15T06:22:30.138123Z","shell.execute_reply.started":"2024-10-15T06:22:29.840622Z","shell.execute_reply":"2024-10-15T06:22:30.136933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nHere is a histogram of the number of trades per day (it has been [suggested](https://www.kaggle.com/c/jane-street-market-prediction/discussion/201930#1125847) that the number of trades per day is an indication of the [volatility](https://www.investopedia.com/terms/v/volatility.asp) that day)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (10,4))\n# the minimum has been set to 1000 so as not to draw the partial days like day 2 and day 294\n# the maximum number of trades per day is 18884\n# I have used 100 bins for the 1699 days\nax = sns.distplot(trades_per_day, \n             bins=100, \n             kde_kws={\"clip\":(2500,350000)}, \n             hist_kws={\"range\":(2500,350000)},\n             color='darkcyan', \n             kde=True);\nvalues = np.array([rec.get_height() for rec in ax.patches])\nnorm = plt.Normalize(values.min(), values.max())\ncolors = plt.cm.jet(norm(values))\nfor rec, col in zip(ax.patches, colors):\n    rec.set_color(col)\nplt.xlabel(\"Number of trades per day\", size=14)\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:26:41.206320Z","iopub.execute_input":"2024-10-15T06:26:41.206944Z","iopub.status.idle":"2024-10-15T06:26:41.661739Z","shell.execute_reply.started":"2024-10-15T06:26:41.206894Z","shell.execute_reply":"2024-10-15T06:26:41.660488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If that is the case, then 'volatile' days, say with more than 9k trades (*i.e.* `time_id`) per day, are the following ","metadata":{}},{"cell_type":"code","source":"volatile_days = pd.DataFrame(trades_per_day[trades_per_day > 9000])\nvolatile_days.T","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:27:48.090327Z","iopub.execute_input":"2024-10-15T06:27:48.091798Z","iopub.status.idle":"2024-10-15T06:27:48.589616Z","shell.execute_reply.started":"2024-10-15T06:27:48.091746Z","shell.execute_reply":"2024-10-15T06:27:48.588408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"features_file\"></a>\n## <center style=\"background-color:Gainsboro; width:40%;\">The features.csv file</center>\n​\nWe are also provided with a `features.csv` file which contains \"*metadata pertaining to the anonymized features*\". Let us take a quick look at it, where `1` is `True` and `0` is `False`. The file has 17 \"tags\" associated with each feature (00 - 78).","metadata":{}},{"cell_type":"code","source":"feature_tags = pd.read_csv(\"/kaggle/input/jane-street-real-time-market-data-forecasting/features.csv\" ,index_col=0)\n# convert to binary\nfeature_tags = feature_tags*1\n# plot a transposed dataframe\nfeature_tags.T.style.background_gradient(cmap='Oranges')","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:56:55.861535Z","iopub.execute_input":"2024-10-15T06:56:55.862630Z","iopub.status.idle":"2024-10-15T06:56:56.173592Z","shell.execute_reply.started":"2024-10-15T06:56:55.862579Z","shell.execute_reply":"2024-10-15T06:56:56.172398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just for fun let us re-plot the above data, but now in '8-bit' mode; totally illegible, but may perhaps serve as an overall visual aid...","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(32,14))\nsns.heatmap(feature_tags.T,\n            cbar=False,\n            xticklabels=False,\n            yticklabels=False,\n            cmap=\"Oranges\");","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:57:24.318090Z","iopub.execute_input":"2024-10-15T06:57:24.318581Z","iopub.status.idle":"2024-10-15T06:57:24.668640Z","shell.execute_reply.started":"2024-10-15T06:57:24.318537Z","shell.execute_reply":"2024-10-15T06:57:24.667338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us sum the number of tags for each feature:","metadata":{}},{"cell_type":"code","source":"tag_sum = pd.DataFrame(feature_tags.T.sum(axis=0),columns=['Number of tags'])\ntag_sum.T","metadata":{"execution":{"iopub.status.busy":"2024-10-15T06:57:52.878575Z","iopub.execute_input":"2024-10-15T06:57:52.879134Z","iopub.status.idle":"2024-10-15T06:57:52.919730Z","shell.execute_reply.started":"2024-10-15T06:57:52.879086Z","shell.execute_reply":"2024-10-15T06:57:52.918558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that all of the features have at least one tag, and some as many as four. ","metadata":{}},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"day_0\"></a>\n## <center style=\"background-color:Gainsboro; width:90%;\">Now let us take a look at the first day (\"day 0\")</center>\nTo do this we shall make a new dataframe called `day_0`","metadata":{}},{"cell_type":"code","source":"day_0 = train_data.loc[train_data['date_id'] == 0]","metadata":{"execution":{"iopub.status.busy":"2024-10-15T07:04:35.130049Z","iopub.execute_input":"2024-10-15T07:04:35.130520Z","iopub.status.idle":"2024-10-15T07:04:35.139322Z","shell.execute_reply.started":"2024-10-15T07:04:35.130477Z","shell.execute_reply":"2024-10-15T07:04:35.137736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15, 5))\nresp_1= pd.Series(day_0['responder_1']).cumsum()\nresp_2= pd.Series(day_0['responder_2']).cumsum()\nresp_3= pd.Series(day_0['responder_3']).cumsum()\nresp_4= pd.Series(day_0['responder_4']).cumsum()\nresp_5= pd.Series(day_0['responder_5']).cumsum()\nresp_6= pd.Series(day_0['responder_6']).cumsum()\nresp_7= pd.Series(day_0['responder_7']).cumsum()\nresp_8= pd.Series(day_0['responder_8']).cumsum()\nax.set_xlabel (\"Trade\", fontsize=18)\nax.set_title (\"Cumulative values for resp and time horizons 1, 2, 3, 4, 5, 6, 7 and 8 for day 0\", fontsize=18)\nresp_1.plot(lw=3, label = 'resp_1')\nresp_2.plot(lw=3, label = 'resp_2')\nresp_3.plot(lw=3, label = 'resp_3')\nresp_4.plot(lw=3, label = 'resp_4')\nresp_5.plot(lw=3, label = 'resp_5')\nresp_6.plot(lw=3, label = 'resp_6')\nresp_7.plot(lw=3, label = 'resp_7')\nresp_8.plot(lw=3, label = 'resp_8')\nplt.legend(loc=\"upper left\");","metadata":{"execution":{"iopub.status.busy":"2024-10-15T07:07:28.580749Z","iopub.execute_input":"2024-10-15T07:07:28.581309Z","iopub.status.idle":"2024-10-15T07:22:25.939684Z","shell.execute_reply.started":"2024-10-15T07:07:28.581261Z","shell.execute_reply":"2024-10-15T07:22:25.938228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del resp_1\ndel resp_2\ndel resp_3\ndel resp_4\ndel resp_5\ndel resp_6\ndel resp_7\ndel resp_8\ngc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T07:22:25.943591Z","iopub.execute_input":"2024-10-15T07:22:25.944011Z","iopub.status.idle":"2024-10-15T07:22:26.604018Z","shell.execute_reply.started":"2024-10-15T07:22:25.943969Z","shell.execute_reply":"2024-10-15T07:22:26.602839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <center style=\"background-color:Gainsboro; width:80%;\">Descriptive statistics of the `train.csv` file for day 0</center>\nSome simple [descriptive statistics](https://en.wikipedia.org/wiki/Descriptive_statistics) of the day 0 data:","metadata":{}},{"cell_type":"code","source":"# Convert Dask DataFrame to Pandas\nday_0_pandas = day_0.describe().compute()\n\n# Apply background gradient styling (Pandas operation)\nstyled_day_0 = day_0_pandas.style.background_gradient(cmap=colorMap)\n\n# Display the styled DataFrame\nstyled_day_0\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T07:24:02.578771Z","iopub.execute_input":"2024-10-15T07:24:02.579291Z","iopub.status.idle":"2024-10-15T07:25:01.006419Z","shell.execute_reply.started":"2024-10-15T07:24:02.579245Z","shell.execute_reply":"2024-10-15T07:25:01.005190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"missing_values\"></a>\n## <center style=\"background-color:Gainsboro; width:60%;\">Are there any missing values?</center>\n\nTo start with let us look at day 0 ","metadata":{}},{"cell_type":"code","source":"# Convert Dask DataFrame to Pandas\nday_0_pandas = day_0.compute()\n\n# Visualize missing data using missingno matrix\nmsno.matrix(day_0_pandas, color=(0.35, 0.35, 0.75))","metadata":{"execution":{"iopub.status.busy":"2024-10-15T07:25:45.486471Z","iopub.execute_input":"2024-10-15T07:25:45.486943Z","iopub.status.idle":"2024-10-15T07:26:18.077209Z","shell.execute_reply.started":"2024-10-15T07:25:45.486900Z","shell.execute_reply":"2024-10-15T07:26:18.075789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"graphically we can see that indeed there are complete chunks of missing data (in white) in some of the columns, and there appears to be numerous columns like this.","metadata":{}},{"cell_type":"markdown","source":"Now let us look at the sum of the number of missing data in each column for the whole `train.csv` file:","metadata":{}},{"cell_type":"code","source":"#missing_data = pd.DataFrame(train_data.isna().sum().sort_values(ascending=False),columns=['Total missing'])\n#missing_data.T\n\ngone = train_data.isnull().sum().compute()\npx.bar(gone, color=gone.values, title=\"Total number of missing values for each column\").show()","metadata":{"execution":{"iopub.status.busy":"2024-10-15T07:33:10.279254Z","iopub.execute_input":"2024-10-15T07:33:10.279754Z","iopub.status.idle":"2024-10-15T07:33:47.061386Z","shell.execute_reply.started":"2024-10-15T07:33:10.279713Z","shell.execute_reply":"2024-10-15T07:33:47.060180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Overall Observations on Missing Values\n\n1. **Repeated Missing Patterns**: Several features exhibit identical missing value counts. For instance:\n   1. Features 00, 01, 02, 03, and 04 share the same number of missing values (3.18M).\n   2. Features 21, 26, 27, and 31 also have identical missing values (8.43M).\n\n2. **Pairs of Features with Similar Missing Values**: Some features tend to appear in pairs with the same number of missing values, such as: Features 39 and 42 (4.3M), 40 and 43 (67.85k), and 41 and 44 (1.09M).\n   \n3. **Other Groupings of Identical Missing Counts**: Features 50 and 53 (4.25M), 52 and 55 (1.04M), 65 and 66 (317.16k), 73 and 74 (483.76k), and 75 and 76 (58.43k) show identical missing value counts.\n\n4. **Possible Data Collection Issues**: These repeated patterns may suggest that certain groups of features were collected together, or that specific sensors, processes, or systems responsible for generating these features faced common disruptions or failures.\n\n\n| Features          | Missing Values      |\n|-------------------|---------------------|\n| 00, 01, 02, 03, 04 | 3.182052 M          |\n| 21, 26, 27, 31     | 8.43 M              |\n| 39, 42             | 4.3 M               |\n| 40, 43             | 67.856 K            |\n| 41, 44             | 1.093012 M          |\n| 50, 53             | 4.25 M              |\n| 52, 55             | 1.04 M              |\n| 65, 66             | 317.163 K           |\n| 73, 74             | 483.759 K           |\n| 75, 76             | 58.43 K             |\n\n\nThere are more features with even less missing values. I think the interesting thing is not so much the quantity of missing values in so much as it may tell us which features represent similar measures/metrics.\n\nIs day 0 special, or does every day have missing data?","metadata":{}},{"cell_type":"code","source":"missing_features = train_data.iloc[:, 7:137].isnull().sum(axis=1).groupby(train_data['date_id']).sum().compute().to_frame()\n\n# Now plot\nfig, ax = plt.subplots(figsize=(15, 5))\nplt.plot(missing_features)\nax.set_xlabel(\"Day\", fontsize=18)\nax.set_title(\"Total number of missing values in all features for each day\", fontsize=18)\nax.axvline(x=85, linestyle='--', alpha=0.3, c='red', lw=2)\nax.axvspan(0, 85, color=sns.xkcd_rgb['grey'], alpha=0.1)\nax.set_xlim(xmin=0)\nax.set_xlim(xmax=1699)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T08:43:43.991616Z","iopub.execute_input":"2024-10-15T08:43:43.992134Z","iopub.status.idle":"2024-10-15T08:45:24.078989Z","shell.execute_reply.started":"2024-10-15T08:43:43.992091Z","shell.execute_reply":"2024-10-15T08:45:24.077612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Indeed we can see that there is missing data *almost* every day. We can see 5 linear line converging there indicating a pattern. \n\nIn the notebook [\"*Jane Street EDA Market Regime*\"](https://www.kaggle.com/marketneutral/jane-street-eda-market-regime) written by [marketneutral](https://www.kaggle.com/marketneutral) a plot is made of the number of trades per day, and is strikingly similar to the above plot. In view of this, for curiosity, we shall plot the number of missing values in the features with respect to the number of trades, for each day.","metadata":{}},{"cell_type":"code","source":"# Calculate the count of weights per date_id and flatten the column index\ncount_weights = train_data[['date_id', 'weight']].groupby('date_id').agg(['count']).compute()\ncount_weights.columns = ['weights']  # Flatten the column names\n\n# Compute the missing features DataFrame\nmissing_features = train_data.iloc[:, 7:137].isnull().sum(axis=1).groupby(train_data['date_id']).sum().compute().to_frame()\nmissing_features.columns = ['missing']  # Rename the column for clarity\n\n# Merge the DataFrames\nresult = pd.merge(count_weights, missing_features, on=\"date_id\", how=\"inner\")\n\n# Calculate the ratio and the mean of missing features per trade\nresult['ratio'] = result['missing'] / result['weights']\nmissing_per_trade = result['ratio'].mean()\n\n# Now make a plot\nfig, ax = plt.subplots(figsize=(15, 5))\nplt.plot(result['ratio'])\nplt.axhline(missing_per_trade, linestyle='--', alpha=0.85, c='r')\nax.set_xlabel(\"Day\", fontsize=18)\nax.set_title(\"Average number of missing feature values per trade, for each day\", fontsize=18)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T08:54:28.092679Z","iopub.execute_input":"2024-10-15T08:54:28.093700Z","iopub.status.idle":"2024-10-15T08:55:35.584068Z","shell.execute_reply.started":"2024-10-15T08:54:28.093645Z","shell.execute_reply":"2024-10-15T08:55:35.582727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the average comes out to be $\\approx$ 2 missing feature values per trade, per day, but the plot suggests different patterns. The number of missing values decrease after 500 days and majority of missing values are prior to 500th day. Prior to day 250, the number of missing values exceed 8 on an average. \nThis raises the question of [what to do with missing data in the unseen test data?](https://www.kaggle.com/c/jane-street-market-prediction/discussion/200691). Whatever one decides to do, in this competition time is of the essence, so we have to do it fast, and [Yirun Zhang](https://www.kaggle.com/gogo827jz) has made an exhaustive study of the time taken in various filling methods in the notebook [\"*Optimise Speed of Filling-NaN Function*\"](https://www.kaggle.com/gogo827jz/optimise-speed-of-filling-nan-function).\n\n<a class=\"anchor\" id=\"missing_data\"></a>\n## <center style=\"background-color:Gainsboro; width:80%;\">Is there any missing data: Days 2 and 294</center>\nIf we produce scatter plots of `feature_64` we see that each day has the same sweeping pattern. Here is a plot of day 1 (in blue), day 2 (in red) and day 3 (blue again). Day 2 has been encircled as a visual aid.","metadata":{}},{"cell_type":"code","source":"# Convert Dask DataFrames to Pandas DataFrames\nday_1 = train_data.loc[train_data['date_id'] == 1].compute()\nday_2 = train_data.loc[train_data['date_id'] == 2].compute()\nday_3 = train_data.loc[train_data['date_id'] == 3].compute()\n\n# Concatenate the Pandas DataFrames\nthree_days = pd.concat([day_1, day_2, day_3])\n\n# Create the scatter plot\nfig, ax = plt.subplots(figsize=(15, 3))\nax.scatter(three_days.time_id, three_days.feature_65, s=0.5, color='b', label='Days 1, 2, 3')\nax.scatter(day_2.time_id, day_2.feature_65, s=0.5, color='r', label='Day 2')\nax.scatter(1515, 5.2, s=1800, facecolors='none', edgecolors='black', linestyle='--', lw=2)\nax.set_xlabel('time_id')\nax.set_ylabel('feature_65')\nax.set_title('feature_65 for days 1, 2, and 3')\nax.legend()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T09:21:19.783018Z","iopub.execute_input":"2024-10-15T09:21:19.783588Z","iopub.status.idle":"2024-10-15T09:22:52.329585Z","shell.execute_reply.started":"2024-10-15T09:21:19.783540Z","shell.execute_reply":"2024-10-15T09:22:52.328265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot of `responder_6` values with respect to time (`time_id`) for day 0","metadata":{}},{"cell_type":"code","source":"# Convert Dask DataFrame to Pandas DataFrame\nday_0 = day_0.compute()\n\n# Now create the scatter plot\nfig_1 = px.scatter(day_0, x='time_id', y='responder_6', \n                   trendline=\"ols\", marginal_y=\"violin\",\n                   title=\"Scatter plot of resp with respect to ts_id for day 0\")\nfig_1.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T09:28:50.134784Z","iopub.execute_input":"2024-10-15T09:28:50.135271Z","iopub.status.idle":"2024-10-15T09:29:47.839641Z","shell.execute_reply.started":"2024-10-15T09:28:50.135218Z","shell.execute_reply":"2024-10-15T09:29:47.838389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"permutation\"></a>\n## <center style=\"background-color:Gainsboro; width:100%;\">Very quick Permutation Importance using the Random Forest</center>\nWe shall now perform a simple [permutation importance](https://www.kaggle.com/dansbecker/permutation-importance) calculation, a basic way of seeing which features may be important. We shall perform a regression, with `responder_6` as the target.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.ensemble import RandomForestRegressor\n\n# Assuming day_0 is already defined\n# Select only the columns that contain 'feature'\nX_train = day_0.loc[:, day_0.columns.str.contains('feature')]\n\n# Step 1: Remove columns that are entirely NaN\nX_train.dropna(axis=1, how='all', inplace=True)\n\n# Step 2: Replace infinite values with NaN\nX_train.replace([float('inf'), float('-inf')], pd.NA, inplace=True)\n\n# Fill remaining NaN values with the mean of each column\nX_train.fillna(X_train.mean(), inplace=True)\n\n# Ensure our target variable is correctly defined\ny_train = day_0['responder_6']\n\n# Step 3: Check if y_train has any missing values\nif y_train.isnull().any():\n    print(\"Target variable y_train contains NaN values. Please handle them before fitting the model.\")\nelse:\n    # Create and fit the RandomForestRegressor model\n    regressor = RandomForestRegressor(max_features='auto')\n    \n    # Fit the model\n    try:\n        regressor.fit(X_train, y_train)\n        print(\"Model trained successfully.\")\n    except ValueError as e:\n        print(f\"Error occurred while fitting the model: {e}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T09:37:18.512284Z","iopub.execute_input":"2024-10-15T09:37:18.512753Z","iopub.status.idle":"2024-10-15T09:38:09.600660Z","shell.execute_reply.started":"2024-10-15T09:37:18.512712Z","shell.execute_reply":"2024-10-15T09:38:09.599212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import eli5\nfrom eli5.sklearn import PermutationImportance\nperm_import = PermutationImportance(regressor, random_state=1).fit(X_train, y_train)\n# visualize the results\neli5.show_weights(perm_import, top=15, feature_names = X_train.columns.tolist())","metadata":{"execution":{"iopub.status.busy":"2024-10-15T09:39:12.327863Z","iopub.execute_input":"2024-10-15T09:39:12.328423Z","iopub.status.idle":"2024-10-15T09:40:20.516297Z","shell.execute_reply.started":"2024-10-15T09:39:12.328373Z","shell.execute_reply":"2024-10-15T09:40:20.515068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see in this very simple example that for the first day (**day 0**) the top 5 most important features appear to be 58, 05, 37, 07 and 49.\n\nNote:\n\n> Features that are deemed of low importance for a bad model (low cross-validation score) could be very important for a good model. Therefore it is always important to evaluate the predictive power of a model using a held-out set (or better with cross-validation) prior to computing importances. Permutation importance does not reflect to the intrinsic predictive value of a feature by itself but how important this feature is for a particular model. (Source: [scikit-learn permutation importance](https://scikit-learn.org/stable/modules/permutation_importance.html)).\n\nIt goes without saying that a serious study of the feature importance is essential (and will use *a lot* of CPU). For a much more advanced approach may I suggest taking a look at [\"*Feature selection using the Boruta-SHAP package*\"](https://www.kaggle.com/carlmcbrideellis/feature-selection-using-the-boruta-shap-package). \n\nHowever, the global feature importance ranking does not tell the whole story, and in the notebook [\"*TabNet and interpretability: Jane Street example*\"](https://www.kaggle.com/carlmcbrideellis/tabnet-and-interpretability-jane-street-example) thanks to [TabNet](https://www.kaggle.com/carlmcbrideellis/jane-street-tabnet-3-0-0-starter-notebook) one can inspect which features were important for each and every calculation, and one can see that the process is much more dynamic than a static overall ranking would suggest.\n\n<a class=\"anchor\" id=\"Pearson\"></a>\n## <center style=\"background-color:Gainsboro; width:90%;\">Is there any correlation between day 100 and day 200?</center>\nAre the days independent? For the moment let us take a look at day(100) and day(200) using a [Pearson pairwise correlation](https://en.wikipedia.org/wiki/Pearson_correlation_coefficient) matrix (this is a **big** matrix!). Why days 100 and 200? Because they are far apart in time, thus reducing any temporal leakage.\nWe shall use a [diverging colormap](https://matplotlib.org/3.1.0/tutorials/colors/colormaps.html) where red indicates positive linear correlation, and blue indicates linear anti-correlation:","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Assuming day_100 and day_200 are already defined\nday_100 = train_data.loc[train_data['date_id'] == 100].compute()\nday_200 = train_data.loc[train_data['date_id'] == 200].compute()\n\n# Concatenate the two DataFrames\nday_100_and_200 = pd.concat([day_100, day_200])\n\n# Calculate the correlation matrix\ncorrelation_matrix = day_100_and_200.corr(method='pearson')\n\n# Use the Styler to apply formatting and background gradient\nstyled_corr = correlation_matrix.style.background_gradient(cmap='coolwarm', axis=None).format(\"{:.2f}\")\n\n# Display the styled correlation matrix\nstyled_corr\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T09:51:57.591774Z","iopub.execute_input":"2024-10-15T09:51:57.592341Z","iopub.status.idle":"2024-10-15T09:53:00.621432Z","shell.execute_reply.started":"2024-10-15T09:51:57.592293Z","shell.execute_reply":"2024-10-15T09:53:00.620084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It can be seen, these features have nan values and needs to be handled properly: \n1. feature_00\n2. feature_01\n3. feature_02\n4. feature_03\n5. feature_04\n6. feature_21\n7. feature_26\n8. feature_27\n9. feature_31","metadata":{}},{"cell_type":"markdown","source":"### <center style=\"background-color:LightGreen; width:40%;\">Thank you for your time.</center>\n\nPostscript:\n\n> \"*If you have been asked to develop ML strategies on your own, the odds are stacked\nagainst you. It takes almost as much effort to produce one true investment strategy\nas to produce a hundred, and the complexities are overwhelming*\". ([Marcos Lopez de Prado in \"*Advances in Financial Machine Learning*\"](https://www.wiley.com/en-es/Advances+in+Financial+Machine+Learning-p-9781119482109))\n\n> \"*...It looks just a little more mathematical and regular than it is; its exactitude is obvious, but its inexactitude is hidden; its wildness lies in wait.*\" (G. K. Chesterton (1908))\n\n### <center style=\"background-color:LightGreen; width:40%;\">Good luck to everyone!</center>\n","metadata":{}}]}