{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"### This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-11-19T06:42:26.141364Z","iopub.execute_input":"2024-11-19T06:42:26.141737Z","iopub.status.idle":"2024-11-19T06:42:26.573557Z","shell.execute_reply.started":"2024-11-19T06:42:26.141673Z","shell.execute_reply":"2024-11-19T06:42:26.572492Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Understanding","metadata":{}},{"cell_type":"markdown","source":"### Exploring training features","metadata":{}},{"cell_type":"code","source":"# Read a single partition of data\n\npartition_number = 3\ndf = pd.read_parquet(f'/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id={partition_number-1}/part-0.parquet')\n\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:42:26.575563Z","iopub.execute_input":"2024-11-19T06:42:26.576033Z","iopub.status.idle":"2024-11-19T06:42:33.309592Z","shell.execute_reply.started":"2024-11-19T06:42:26.575994Z","shell.execute_reply":"2024-11-19T06:42:33.308431Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"a = {1, 2, 3}\na.update(np.array([3,4]))\na","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:42:33.311090Z","iopub.execute_input":"2024-11-19T06:42:33.311407Z","iopub.status.idle":"2024-11-19T06:42:33.318334Z","shell.execute_reply.started":"2024-11-19T06:42:33.311376Z","shell.execute_reply":"2024-11-19T06:42:33.317202Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:42:33.319573Z","iopub.execute_input":"2024-11-19T06:42:33.319900Z","iopub.status.idle":"2024-11-19T06:42:33.411667Z","shell.execute_reply.started":"2024-11-19T06:42:33.319869Z","shell.execute_reply":"2024-11-19T06:42:33.410498Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# We will try to find read all the available partitions (one at a time)\n# and get the unique values of some of the features\n\n\nimport time\n\n# Measure time for the whole loop\nglobal_start = time.time()\n\n\nunique_date_id = set()\nunique_time_id = set()\nunique_symbol_id = set()\n\nlocal_start = None\npartitions = range(10)\nfor p in partitions:\n    print(f\"Reading partition {p}\")\n    local_start = time.time()\n    df = pd.read_parquet(f'/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id={p}/part-0.parquet')\n\n    unique_date_id.update(df['date_id'])\n    unique_time_id.update(df['time_id'])\n    unique_symbol_id.update(df['symbol_id'])\n    print(f\"Finished analyzing partition {p} in {time.time() - local_start} seconds\")\n    gc.collect()\n\nprint(f\"Finished the whole loop in {time.time() - global_start} seconds\")\n\nunique_date_id = sorted(list(unique_date_id))\nunique_time_id = sorted(list(unique_time_id))\nunique_symbol_id = sorted(list(unique_symbol_id))\nprint(f\"The number of unique date_id is {len(unique_date_id)}, its min: {min(unique_date_id)} and its max: {max(unique_date_id)}\")\nprint(f\"The number of unique date_id is {len(unique_time_id)}, its min: {min(unique_time_id)} and its max: {max(unique_time_id)}\")\nprint(f\"The number of unique date_id is {len(unique_symbol_id)}, its min: {min(unique_symbol_id)} and its max: {max(unique_symbol_id)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:42:33.414639Z","iopub.execute_input":"2024-11-19T06:42:33.415154Z","iopub.status.idle":"2024-11-19T06:44:29.866258Z","shell.execute_reply.started":"2024-11-19T06:42:33.415103Z","shell.execute_reply":"2024-11-19T06:44:29.864680Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We notice that the three features increment from zero to a certain maximum value without any missing value in between. ","metadata":{}},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:44:29.867645Z","iopub.execute_input":"2024-11-19T06:44:29.868019Z","iopub.status.idle":"2024-11-19T06:44:29.947809Z","shell.execute_reply.started":"2024-11-19T06:44:29.867983Z","shell.execute_reply":"2024-11-19T06:44:29.946560Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Exploring symbols","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:44:51.850869Z","iopub.execute_input":"2024-11-19T06:44:51.851269Z","iopub.status.idle":"2024-11-19T06:44:51.857022Z","shell.execute_reply.started":"2024-11-19T06:44:51.851234Z","shell.execute_reply":"2024-11-19T06:44:51.855579Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Let us first fix a symbol_id\nfirst_symbol_data = df[df['symbol_id']==0]\n\n# Let us now try to see how the responders move for this symbol\n\nfirst_responder= 2\nlast_responder=5\n\nfig, ax = plt.subplots(nrows=last_responder - first_responder, \n                       ncols=1, \n                       sharex=True)\n\ncurrent_responder = first_responder\nfor row in ax:\n    row.plot(first_symbol_data[f\"responder_{current_responder}\"])\n    current_responder += 1\n\n    if current_responder > last_responder:\n        break\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:53:33.724180Z","iopub.execute_input":"2024-11-19T06:53:33.725387Z","iopub.status.idle":"2024-11-19T06:53:35.104526Z","shell.execute_reply.started":"2024-11-19T06:53:33.725220Z","shell.execute_reply":"2024-11-19T06:53:35.102250Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We notice that all the responders are concentrated around 0.","metadata":{}},{"cell_type":"code","source":"# Let us now try to see how the FEATURES move for this symbol\n\nfirst_feature= 12\nlast_feature=15\n\nfig, ax = plt.subplots(nrows=last_feature - first_feature, \n                       ncols=1, \n                       sharex=True)\n\ncurrent_feature = first_feature\nfor row in ax:\n    row.plot(first_symbol_data[\"feature_\"+str(current_feature).zfill(2)])\n    current_feature += 1\n\n    if current_feature > last_feature:\n        break\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T06:57:09.089061Z","iopub.execute_input":"2024-11-19T06:57:09.089519Z","iopub.status.idle":"2024-11-19T06:57:09.933995Z","shell.execute_reply.started":"2024-11-19T06:57:09.089481Z","shell.execute_reply":"2024-11-19T06:57:09.932563Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From the behavior of the features, we can notice that it resembles to a **Returns behavior** of certain assets, the symbol_id are thus just portfolios (or some ETFs for those familiar with financial instruments).\n\nTo check if this is the case,we can compare the same feature have the same behavior in different portfolios.","metadata":{}},{"cell_type":"code","source":"feature = \"feature_08\"\n\n# Let us now try to see how the responders move for this symbol\n\nfirst_symbol= 0\nlast_symbol=4\n\nfig, ax = plt.subplots(nrows=last_symbol - first_symbol, \n                       ncols=1, \n                       sharex=True)\n\ncurrent_symbol = first_symbol\nfor row in ax:\n    row.plot(df[df[\"symbol_id\"]==current_symbol][feature])\n    row.set_title(f\"Behavior of {feature} for symbol: {current_symbol}\")\n    current_symbol += 1\n\n    if current_symbol > last_symbol:\n        break\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-19T07:07:49.211235Z","iopub.execute_input":"2024-11-19T07:07:49.211657Z","iopub.status.idle":"2024-11-19T07:07:50.779680Z","shell.execute_reply.started":"2024-11-19T07:07:49.211619Z","shell.execute_reply":"2024-11-19T07:07:50.778433Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We notice that the overall behavior of a single feature is almost the same for every symbol, however there may be some difference that can be due to the liquidity. \n\nAgain for people familiar with Financial Markets, this can be due to the fact that the volume of that feature is different from one portfolio to another, which presents more transaction costs.\n\nTo sum up:\n* Symbol_id represent the identifier of a portfolio (which is a basket of assets).\n\n* Features represents the returns of those assets.","metadata":{}}]}