{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Dataset Description\n\nThe competition dataset comprises a set of timeseries with 79 features and 9 responders, anonymized but representing real market data. The goal of the competition is to forecast one of these responders, i.e., `responder_6`, for up to six months in the future.\n\nYou must submit to this competition using the provided Python evaluation API, which serves test set data one timestep by timestep. To use the API, follow the example in [this notebook](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission). (Note that this API is different from our legacy timeseries API used in past forecasting competitions.)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport polars as pl\nfrom matplotlib import pyplot as plt\nfrom matplotlib.ticker import MaxNLocator, FormatStrFormatter, PercentFormatter\n#提供了用于配置刻度定位和格式的类，包括通用和特定于域的定位器和格式化程序。\nimport seaborn as sns \nimport gc","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-11-04T14:34:06.129446Z","iopub.execute_input":"2024-11-04T14:34:06.129955Z","iopub.status.idle":"2024-11-04T14:34:09.808399Z","shell.execute_reply.started":"2024-11-04T14:34:06.129887Z","shell.execute_reply":"2024-11-04T14:34:09.807140Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"ROOT_DIR = \"/kaggle/input/jane-street-real-time-market-data-forecasting\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T14:34:12.272991Z","iopub.execute_input":"2024-11-04T14:34:12.273569Z","iopub.status.idle":"2024-11-04T14:34:12.279452Z","shell.execute_reply.started":"2024-11-04T14:34:12.273526Z","shell.execute_reply":"2024-11-04T14:34:12.278189Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Features\n\n- features.csv - metadata pertaining to the anonymized features","metadata":{}},{"cell_type":"code","source":"features = pd.read_csv(f\"{ROOT_DIR}/features.csv\")\nfeatures  ## 17列","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 找到所有列都是 False 的行索引\nfeatures[features.all(axis = 1)]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(20, 10))\nplt.imshow(features.iloc[:, 1:].T.values, cmap=\"gray\")\nplt.xlabel(\"feature_00  ~  feature_78\")\nplt.ylabel(\"tag_0  ~  tag_16\")\nplt.yticks(np.arange(17))\nplt.xticks(np.arange(79))\nplt.grid()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# corr between feature_XX and feature_YY （基于tag的值）\nplt.figure(figsize=(10, 10))\nsns.heatmap(features[[ f\"tag_{no}\" for no in range(0,17,1) ] ].T.corr(), square=True, cmap=\"jet\")\n## cmap：指定一个colormap对象，用于热力图的填充色\n## square：bool类型参数，是否使热力图的每个单元格为正方形，默认为False\n\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Responders\n\n- responders.csv - metadata pertaining to the anonymized responders","metadata":{}},{"cell_type":"code","source":"responders = pd.read_csv(f\"{ROOT_DIR}/responders.csv\")\nresponders","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:34:26.745090Z","iopub.execute_input":"2024-11-04T14:34:26.745513Z","iopub.status.idle":"2024-11-04T14:34:26.782428Z","shell.execute_reply.started":"2024-11-04T14:34:26.745474Z","shell.execute_reply":"2024-11-04T14:34:26.781259Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# corr between responder_XX and responder_YY\nsns.heatmap(responders[[ f\"tag_{no}\" for no in range(0,5,1) ] ].T.corr(),  annot=True, square=True, cmap=\"jet\")\n## annot：指定一个bool类型的值或与data参数形状一样的数组，如果为True，就在热力图的每个单元上显示数值\nplt.xlabel(\"responder_0  ~  responder_8\")\nplt.ylabel(\"responder_0  ~  responder_8\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:34:29.100492Z","iopub.execute_input":"2024-11-04T14:34:29.100945Z","iopub.status.idle":"2024-11-04T14:34:29.712289Z","shell.execute_reply.started":"2024-11-04T14:34:29.100901Z","shell.execute_reply":"2024-11-04T14:34:29.711193Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Sample submission\n\n- **sample_submission.csv** - This file illustrates the format of the predictions your model should make.","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(f\"{ROOT_DIR}/sample_submission.csv\")\nprint( f\"sub.shape = {sub.shape}\" )\nsub","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:34:37.199241Z","iopub.execute_input":"2024-11-04T14:34:37.199716Z","iopub.status.idle":"2024-11-04T14:34:37.231144Z","shell.execute_reply.started":"2024-11-04T14:34:37.199674Z","shell.execute_reply":"2024-11-04T14:34:37.229841Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Train.parquet\n\n- **train.parquet** - The training set, contains historical data and returns. For convenience, the training set has been partitioned into ten parts.\n  - `date_id` and `time_id` - Integer values that are ordinally sorted, providing a chronological structure to the data, although the actual time intervals between `time_id` values may vary.\n  - `symbol_id` - Identifies a unique financial instrument.\n  - `weight` - The weighting used for calculating the scoring function.\n  - `feature_{00...78}` - Anonymized market data.\n  - `responder_{0...8}` - Anonymized responders clipped between -5 and 5. The responder_6 field is what you are trying to predict.\n  \n  \nEach row in the `{train/test}.parquet` dataset corresponds to a unique combination of a symbol (identified by `symbol_id`) and a timestamp (represented by `date_id` and `time_id`). You will be provided with multiple responders, with `responder_6` being the only responder used for scoring. The date_id column is an integer which represents the day of the event, while time_id represents a time ordering. It's important to note that the real time differences between each time_id are not guaranteed to be consistent.\n\nThe `symbol_id` column contains encrypted identifiers. Each `symbol_id` is not guaranteed to appear in all `time_id` and `date_id` combinations. Additionally, new `symbol_id` values may appear in future test sets.","metadata":{}},{"cell_type":"code","source":"!tree /kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:34:52.467992Z","iopub.execute_input":"2024-11-04T14:34:52.468435Z","iopub.status.idle":"2024-11-04T14:34:53.673045Z","shell.execute_reply.started":"2024-11-04T14:34:52.468392Z","shell.execute_reply":"2024-11-04T14:34:53.671546Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = (\n    pl.read_parquet(f\"{ROOT_DIR}/train.parquet/partition_id=0/part-0.parquet\")\n)\ntrain.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(str(train.columns))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.null_count()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## 最后一个parquet的缺失值情况:\n\ntrain_9 = (\n    pl.read_parquet(f\"{ROOT_DIR}/train.parquet/partition_id=9/part-0.parquet\")\n)\nprint(train_9.shape)\ntrain_9.null_count()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"|    |   date_id |   time_id |   symbol_id |   weight |   feature_00 |   feature_01 |   feature_02 |   feature_03 |   feature_04 |   feature_05 |   feature_06 |   feature_07 |   feature_08 |   feature_09 |   feature_10 |   feature_11 |   feature_12 |   feature_13 |   feature_14 |   feature_15 |   feature_16 |   feature_17 |   feature_18 |   feature_19 |   feature_20 |   feature_21 |   feature_22 |   feature_23 |   feature_24 |   feature_25 |   feature_26 |   feature_27 |   feature_28 |   feature_29 |   feature_30 |   feature_31 |   feature_32 |   feature_33 |   feature_34 |   feature_35 |   feature_36 |   feature_37 |   feature_38 |   feature_39 |   feature_40 |   feature_41 |   feature_42 |   feature_43 |   feature_44 |   feature_45 |   feature_46 |   feature_47 |   feature_48 |   feature_49 |   feature_50 |   feature_51 |   feature_52 |   feature_53 |   feature_54 |   feature_55 |   feature_56 |   feature_57 |   feature_58 |   feature_59 |   feature_60 |   feature_61 |   feature_62 |   feature_63 |   feature_64 |   feature_65 |   feature_66 |   feature_67 |   feature_68 |   feature_69 |   feature_70 |   feature_71 |   feature_72 |   feature_73 |   feature_74 |   feature_75 |   feature_76 |   feature_77 |   feature_78 |   responder_0 |   responder_1 |   responder_2 |   responder_3 |   responder_4 |   responder_5 |   responder_6 |   responder_7 |   responder_8 |\n|---:|----------:|----------:|------------:|---------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|-------------:|--------------:|--------------:|--------------:|--------------:|--------------:|--------------:|--------------:|--------------:|--------------:|\n|  0 |         0 |         0 |           0 |        0 |            0 |            0 |            0 |            0 |            0 |            0 |            0 |            0 |        37752 |            0 |            0 |            0 |            0 |            0 |            0 |       155568 |            0 |        25928 |            0 |            0 |            0 |        77440 |            0 |            0 |            0 |            0 |        77440 |        77440 |            0 |            0 |            0 |        77440 |        61545 |        61545 |            0 |            0 |            0 |            0 |            0 |       440778 |            2 |       116677 |       440778 |            2 |       116677 |          139 |          139 |            0 |            0 |            0 |       440776 |            0 |       116676 |       440776 |            0 |       116676 |            0 |            0 |        61545 |            0 |            0 |            0 |          126 |           96 |          101 |          139 |          139 |            0 |            0 |            0 |            0 |            0 |            0 |        61707 |        61707 |        10480 |        10480 |         1945 |         1945 |             0 |             0 |             0 |             0 |             0 |             0 |             0 |             0 |             0 |","metadata":{}},{"cell_type":"markdown","source":"## Missing values","metadata":{}},{"cell_type":"code","source":"supervised_usable = (\n    train\n    .filter(pl.col('responder_6').is_not_null())\n)\n\nmissing_count = (\n    supervised_usable\n    .null_count()\n    .transpose(include_header=True,\n               header_name='feature',\n               column_names=['null_count'])\n    .sort('null_count', descending=True)\n    .with_columns((pl.col('null_count') / len(supervised_usable)).alias('null_ratio'))\n)\n\nplt.figure(figsize=(6, 20))\nplt.title(f'Missing values over the {len(supervised_usable)} samples which have a target')\nplt.barh(np.arange(len(missing_count)), missing_count.get_column('null_ratio'), color='coral', label='missing')\nplt.barh(np.arange(len(missing_count)), \n         1 - missing_count.get_column('null_ratio'),\n         left=missing_count.get_column('null_ratio'),\n         color='darkseagreen', label='available')\nplt.yticks(np.arange(len(missing_count)), missing_count.get_column('feature'))\nplt.gca().xaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))\nplt.xlim(0, 1)\nplt.legend()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"################ （没事别运行！！！！！！！！）\n\n# 可以再对别的parquet进行缺失值的分析:\nfor partition_id in range(10):\n    print(f\"> train.parquet/partition_id={partition_id}/part-0.parquet\")\n    train = pl.read_parquet(f\"{ROOT_DIR}/train.parquet/partition_id={partition_id}/part-0.parquet\")\n    supervised_usable = (\n        train\n        .filter(pl.col('responder_6').is_not_null())\n    )\n\n    missing_count = (\n        supervised_usable\n        .null_count()\n        .transpose(include_header=True,\n                   header_name='feature',\n                   column_names=['null_count'])\n        .sort('null_count', descending=True)\n        .with_columns((pl.col('null_count') / len(supervised_usable)).alias('null_ratio'))\n    )\n\n    plt.figure(figsize=(6, 20))\n    plt.title(f'Missing values in {partition_id} over the {len(supervised_usable)} samples which have a target')\n    plt.barh(np.arange(len(missing_count)), missing_count.get_column('null_ratio'), color='coral', label='missing')\n    plt.barh(np.arange(len(missing_count)), \n             1 - missing_count.get_column('null_ratio'),\n             left=missing_count.get_column('null_ratio'),\n             color='darkseagreen', label='available')\n    plt.yticks(np.arange(len(missing_count)), missing_count.get_column('feature'))\n    plt.gca().xaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))\n    plt.xlim(0, 1)\n    plt.legend()\n    plt.show()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"看不同instructment的缺失情况:","metadata":{}},{"cell_type":"code","source":"from memory_profiler import profile\n\nfor partition_id in range(10):\n    print(f\"> train.parquet/partition_id={partition_id}/part-0.parquet\")\n    train = pl.read_parquet(f\"{ROOT_DIR}/train.parquet/partition_id={partition_id}/part-0.parquet\")\n    tot_num_of_instrument = train.select(\"symbol_id\").unique().sort(\"symbol_id\")\n    print(f\"在数据集{partition_id}中，共有{tot_num_of_instrument.shape[0]}个instrument\")\n    if tot_num_of_instrument.shape[0] < 39:\n        print(f\"缺失的instrument为:{np.setdiff1d(np.arange(39), tot_num_of_instrument.to_numpy().squeeze())}\")","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:37:04.194307Z","iopub.execute_input":"2024-11-04T14:37:04.194770Z","iopub.status.idle":"2024-11-04T14:37:49.805620Z","shell.execute_reply.started":"2024-11-04T14:37:04.194723Z","shell.execute_reply":"2024-11-04T14:37:49.804326Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"id=5以后就齐了，进一步看feature的缺失情况:","metadata":{}},{"cell_type":"code","source":"def dui_bi_bu_tong_partition_id_xia_de_que_shi_zhi(ins_id):\n    for partition_id in range(10):\n        print(f\"> train.parquet/partition_id={partition_id}/part-0.parquet\")\n        train = pl.read_parquet(f\"{ROOT_DIR}/train.parquet/partition_id={partition_id}/part-0.parquet\")\n        #print(f\"在第{partition_id}个数据集中第{ins_id}个instrument的缺失情况\")\n        mis_set_new = set(train.filter(pl.col(\"symbol_id\") == ins_id).null_count().\n              unpivot().filter(pl.col(\"value\") != 0).to_pandas()['variable'].tolist())\n        if partition_id > 0:\n            print(\"与上一个partition_id的对比：\")\n            print(\"多的缺失值:\",mis_set_new - mis_set)\n            print(\"少的缺失值:\",mis_set - mis_set_new)\n        mis_set = mis_set_new\n        \ndui_bi_bu_tong_partition_id_xia_de_que_shi_zhi(7)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dui_bi_bu_tong_partition_id_xia_de_que_shi_zhi(30)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## feature_00-78","metadata":{}},{"cell_type":"code","source":"train_09_nonull = train.drop_nulls()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(15, 15))\nsns.heatmap(train_09_nonull[[ f\"feature_{target:02d}\" for target in range(79)]].corr(), square=True, cmap=\"jet\")\nplt.xlabel(\"feature_00  ~  feature_78\")\nplt.ylabel(\"feature_00  ~  feature_78\")\nplt.grid()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"对tags分组看看：","metadata":{}},{"cell_type":"code","source":"features = pd.read_csv(f\"{ROOT_DIR}/features.csv\",index_col=0)\nfeatures  ## 17列","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T14:36:39.640668Z","iopub.execute_input":"2024-11-04T14:36:39.641136Z","iopub.status.idle":"2024-11-04T14:36:39.680502Z","shell.execute_reply.started":"2024-11-04T14:36:39.641092Z","shell.execute_reply":"2024-11-04T14:36:39.679014Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"######## 这个尽量不要运行\n\n\nfor tags_id in range(17):\n    # 想要索引（feature names）\n    feature_names = features[features[f'tag_{tags_id}']].index#.tolist()\n    plt.figure(figsize=(15, 15))\n    print(tags_id,\":\")\n    sns.heatmap(train_09_nonull[[ f\"{i}\" for i in feature_names]].corr(), square=True, cmap=\"jet\",\n            xticklabels=feature_names,  # 添加 x 轴标签\n            yticklabels=feature_names)\n    plt.xlabel(f\"feature in tag_{tags_id}\")\n    plt.ylabel(f\"feature in tag_{tags_id}\")\n    plt.grid()\n    plt.show()","metadata":{"trusted":true,"_kg_hide-output":false},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## responder_0 - 8","metadata":{}},{"cell_type":"code","source":"for target in range(9):\n    col = f\"responder_{target}\"\n    mean_, sgm_ = train[col].mean(), np.sqrt(train[col].var())\n    min_, max_ = train[col].min(), train[col].max()\n    print(\"- \" * 30)\n    print( f\"column = {col}\" )\n    print( f\" - mean  : {mean_:.4f}\",  )\n    print( f\" - sigma : {sgm_:.4f}\",  )\n    print( f\" - min  : {min_:.4f}\",  )\n    print( f\" - max  : {max_:.4f}\",  )\n    \n    plt.hist(train[col], bins=20)\n    plt.xlabel(col)\n    plt.ylabel(\"frequency / records\")\n    #plt.yscale(\"log\")\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:38:09.843801Z","iopub.execute_input":"2024-11-04T14:38:09.844294Z","iopub.status.idle":"2024-11-04T14:38:13.987430Z","shell.execute_reply.started":"2024-11-04T14:38:09.844248Z","shell.execute_reply":"2024-11-04T14:38:13.986050Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\nimport scipy.stats as stats\n\nplt.figure(figsize=(10, 6))\nsns.histplot(train[col].to_pandas(), stat='density', kde=True)\n# 添加正态分布参考线\nx = np.linspace(min(train[col]), max(train[col].to_pandas()), 100)\nplt.plot(x, stats.norm.pdf(x, np.mean(train[col].to_pandas()), np.std(train[col].to_pandas())))\nplt.title('Distribution vs Normal')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T15:06:24.889648Z","iopub.execute_input":"2024-11-04T15:06:24.890112Z","iopub.status.idle":"2024-11-04T15:07:00.887746Z","shell.execute_reply.started":"2024-11-04T15:06:24.890070Z","shell.execute_reply":"2024-11-04T15:07:00.886486Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def check_tail(data):\n    # 计算峰度（正态分布的峰度为3）\n    kurtosis = stats.kurtosis(data)\n    # 计算偏度\n    skewness = stats.skew(data)\n    \n    print(f\"Kurtosis: {kurtosis:.2f}\")  # >3 表示厚尾\n    print(f\"Skewness: {skewness:.2f}\")  # !=0 表示分布不对称\n\ncheck_tail(train[col])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T15:16:19.512506Z","iopub.execute_input":"2024-11-04T15:16:19.513189Z","iopub.status.idle":"2024-11-04T15:16:19.685521Z","shell.execute_reply.started":"2024-11-04T15:16:19.513141Z","shell.execute_reply":"2024-11-04T15:16:19.684258Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import statsmodels.api as sm\nimport matplotlib.pyplot as plt\n\n####### responder 6 的 QQ-图\n\ncol = f\"responder_{6}\"\nplt.figure(figsize=(10, 6))\nsm.qqplot(train[col], line='45')  # 45度参考线\nplt.title('Q-Q Plot')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T14:38:19.112562Z","iopub.execute_input":"2024-11-04T14:38:19.113081Z","iopub.status.idle":"2024-11-04T14:38:36.054656Z","shell.execute_reply.started":"2024-11-04T14:38:19.113034Z","shell.execute_reply":"2024-11-04T14:38:36.053411Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# responder的自相关性平稳性检验\nimport matplotlib.pyplot as plt\nfrom statsmodels.graphics.tsaplots import plot_acf\n\nplt.rcParams.update({'figure.figsize':(8,6), 'figure.dpi':100}) #设置图片大小\nplot_acf(train[col][-2000:],lags=10) #生成自相关图\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T14:40:52.059076Z","iopub.execute_input":"2024-11-04T14:40:52.059542Z","iopub.status.idle":"2024-11-04T14:40:52.390065Z","shell.execute_reply.started":"2024-11-04T14:40:52.059502Z","shell.execute_reply":"2024-11-04T14:40:52.388681Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 平稳性ADF检验\nfrom statsmodels.tsa.stattools import adfuller\n\nADF_result = adfuller(train[col][-2000:])\n \nprint('The ADF Statistic of responder 6: %f' % ADF_result[0])\nprint('The p value of responder 6: %f' % ADF_result[1])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T14:41:47.707107Z","iopub.execute_input":"2024-11-04T14:41:47.709070Z","iopub.status.idle":"2024-11-04T14:41:47.832109Z","shell.execute_reply.started":"2024-11-04T14:41:47.708999Z","shell.execute_reply":"2024-11-04T14:41:47.830839Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"######给定 symbol_id下：\nfor sym_id in range(39):\n    train_sym_0 = train.filter(pl.col(\"symbol_id\") == sym_id)\n    # responder的自相关性平稳性检验\n    plt.rcParams.update({'figure.figsize':(8,6), 'figure.dpi':100}) #设置图片大小\n    plot_acf(train_sym_0[col][-10000:],lags=25) #生成自相关图\n    plt.xlabel(f'symbol_id = {sym_id}')\n    plt.show()\n    \n    # 平稳性ADF检验\n    ADF_result = adfuller(train_sym_0[col][-2000:])\n     \n    print('The ADF Statistic of responder 6: %f' % ADF_result[0])\n    print('The p value of responder 6: %f' % ADF_result[1])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-04T14:54:40.498704Z","iopub.execute_input":"2024-11-04T14:54:40.499175Z","iopub.status.idle":"2024-11-04T14:55:03.981578Z","shell.execute_reply.started":"2024-11-04T14:54:40.499134Z","shell.execute_reply":"2024-11-04T14:55:03.979715Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(8, 8))\nsns.heatmap(train[[ f\"responder_{target}\" for target in range(9)]].corr(),  annot=True, square=True, cmap=\"jet\")\nplt.xlabel(\"responder_0  ~  responder_8\")\nplt.ylabel(\"responder_0  ~  responder_8\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:42:24.779739Z","iopub.execute_input":"2024-11-04T14:42:24.780248Z","iopub.status.idle":"2024-11-04T14:42:26.016314Z","shell.execute_reply.started":"2024-11-04T14:42:24.780203Z","shell.execute_reply":"2024-11-04T14:42:26.014841Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## symbol_id","metadata":{}},{"cell_type":"code","source":"for partition_id in range(10):\n    print(f\"> train.parquet/partition_id={partition_id}/part-0.parquet\")\n    train_data = pl.read_parquet(f\"{ROOT_DIR}/train.parquet/partition_id={partition_id}/part-0.parquet\")\n\n    print( f\"symbol_id: \", train_data[\"symbol_id\"].min(), \"-\", train_data[\"symbol_id\"].max())\n    bins = train_data[\"symbol_id\"].max() - train_data[\"symbol_id\"].min() + 1\n    plt.hist(train_data[\"symbol_id\"], bins=bins)\n    plt.xlabel(\"symbol_id\")\n    plt.ylabel(\"frequency / records\")\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:58:08.271025Z","iopub.execute_input":"2024-11-04T14:58:08.271627Z","iopub.status.idle":"2024-11-04T14:58:46.315350Z","shell.execute_reply.started":"2024-11-04T14:58:08.271568Z","shell.execute_reply":"2024-11-04T14:58:46.314078Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## date_id","metadata":{}},{"cell_type":"code","source":"for partition_id in range(10):\n    print(f\"> train.parquet/partition_id={partition_id}/part-0.parquet\")\n    train_data = pl.read_parquet(f\"{ROOT_DIR}/train.parquet/partition_id={partition_id}/part-0.parquet\")\n\n    print( f\"date_id: \", train_data[\"date_id\"].min(), \"-\", train_data[\"date_id\"].max())\n    bins = train_data[\"date_id\"].max() - train_data[\"date_id\"].min() + 1\n    plt.hist(train_data[\"date_id\"], bins=bins)\n    plt.xlabel(\"date_id\")\n    plt.ylabel(\"frequency / records\")\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:58:51.847829Z","iopub.execute_input":"2024-11-04T14:58:51.848445Z","iopub.status.idle":"2024-11-04T14:59:13.935632Z","shell.execute_reply.started":"2024-11-04T14:58:51.848390Z","shell.execute_reply":"2024-11-04T14:59:13.934176Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Test.parquet\n\n- **test.parquet** - A mock test set which represents the structure of the unseen test set. This example set demonstrates a single batch served by the evaluation API, that is, data from a single `date_id, time_id` pair. The test set contains columns including `date_id`, `time_id`, `symbol_id`, `weight` and `feature_{00...78}`. You will not be directly using the test set or sample submission in this competition, as the evaluation API will get/set the test set and predictions.","metadata":{}},{"cell_type":"code","source":"!tree {ROOT_DIR}/test.parquet/","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:35:04.272259Z","iopub.execute_input":"2024-11-04T14:35:04.272725Z","iopub.status.idle":"2024-11-04T14:35:05.433811Z","shell.execute_reply.started":"2024-11-04T14:35:04.272678Z","shell.execute_reply":"2024-11-04T14:35:05.432220Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test = (\n    pl.read_parquet(f\"{ROOT_DIR}/test.parquet/date_id=0/part-0.parquet\")\n)\ntest.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:35:08.187213Z","iopub.execute_input":"2024-11-04T14:35:08.187686Z","iopub.status.idle":"2024-11-04T14:35:08.333085Z","shell.execute_reply.started":"2024-11-04T14:35:08.187637Z","shell.execute_reply":"2024-11-04T14:35:08.331824Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:35:10.851021Z","iopub.execute_input":"2024-11-04T14:35:10.851479Z","iopub.status.idle":"2024-11-04T14:35:10.888622Z","shell.execute_reply.started":"2024-11-04T14:35:10.851434Z","shell.execute_reply":"2024-11-04T14:35:10.887255Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Missing values","metadata":{}},{"cell_type":"code","source":"supervised_usable = (\n    test\n)\n\nmissing_count = (\n    supervised_usable\n    .null_count()\n    .transpose(include_header=True,\n               header_name='feature',\n               column_names=['null_count'])\n    .sort('null_count', descending=True)\n    .with_columns((pl.col('null_count') / len(supervised_usable)).alias('null_ratio'))\n)\n\nplt.figure(figsize=(6, 20))\nplt.title(f'Missing values over the {len(supervised_usable)} samples which have a target')\nplt.barh(np.arange(len(missing_count)), missing_count.get_column('null_ratio'), color='coral', label='missing')\nplt.barh(np.arange(len(missing_count)), \n         1 - missing_count.get_column('null_ratio'),\n         left=missing_count.get_column('null_ratio'),\n         color='darkseagreen', label='available')\nplt.yticks(np.arange(len(missing_count)), missing_count.get_column('feature'))\nplt.gca().xaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))\nplt.xlim(0, 1)\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:35:59.144452Z","iopub.execute_input":"2024-11-04T14:35:59.145001Z","iopub.status.idle":"2024-11-04T14:36:00.651278Z","shell.execute_reply.started":"2024-11-04T14:35:59.144953Z","shell.execute_reply":"2024-11-04T14:36:00.649998Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# lags.parquet\n\n- `lags.parquet` - Values of `responder_{0...8}` lagged by one `date_id`. The evaluation API serves the entirety of the lagged responders for a `date_id` on that date_id's first `time_id`. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.","metadata":{}},{"cell_type":"code","source":"!tree {ROOT_DIR}/lags.parquet","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:36:05.754178Z","iopub.execute_input":"2024-11-04T14:36:05.754615Z","iopub.status.idle":"2024-11-04T14:36:06.930292Z","shell.execute_reply.started":"2024-11-04T14:36:05.754560Z","shell.execute_reply":"2024-11-04T14:36:06.928922Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags = (\n    pl.read_parquet(f\"{ROOT_DIR}/lags.parquet/date_id=0/part-0.parquet\")\n)\nlags.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:36:10.535042Z","iopub.execute_input":"2024-11-04T14:36:10.536452Z","iopub.status.idle":"2024-11-04T14:36:10.549585Z","shell.execute_reply.started":"2024-11-04T14:36:10.536390Z","shell.execute_reply":"2024-11-04T14:36:10.548293Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags.columns","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:36:15.035143Z","iopub.execute_input":"2024-11-04T14:36:15.035594Z","iopub.status.idle":"2024-11-04T14:36:15.044127Z","shell.execute_reply.started":"2024-11-04T14:36:15.035552Z","shell.execute_reply":"2024-11-04T14:36:15.042843Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:36:17.559313Z","iopub.execute_input":"2024-11-04T14:36:17.559784Z","iopub.status.idle":"2024-11-04T14:36:17.571644Z","shell.execute_reply.started":"2024-11-04T14:36:17.559737Z","shell.execute_reply":"2024-11-04T14:36:17.570201Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.plot(lags[\"responder_6_lag_1\"])\nplt.grid()\nplt.xlabel(\"symbol_id\")\nplt.ylabel(\"responder_6_lag_1\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-04T14:36:20.529969Z","iopub.execute_input":"2024-11-04T14:36:20.530383Z","iopub.status.idle":"2024-11-04T14:36:20.805006Z","shell.execute_reply.started":"2024-11-04T14:36:20.530345Z","shell.execute_reply":"2024-11-04T14:36:20.803630Z"},"trusted":true},"outputs":[],"execution_count":null}]}