{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# INTRODUCTION\nCalculating statistics on text-based features such as \"text\" or \"fqid\" may be effective in predicting student performance. However, since these features contain over 100 unique values (over 500 for \"text\"), using all of them as features is impractical because of the computational cost.\n\nTherefore, we thought it would be better to calculate and compare the feature importance of each value, and use the top few values as the features. In a previous post, linked below, we determined the importance of the \"text\", \"fqid\", and \"text_fqid\" features. The prediction accuracy (CV-Score) using only each feature was 0.688, 0.692, and 0.684, respectively, suggesting that the importance of \"fqid\" was relatively high. In this study, we generated more features for \"fqid\" and investigated the feature importance in more detail. This notebook is the Analysis part of a series of investigations.\n\n---\nThe train part is [here](http://www.kaggle.com/code/tsuyoshifujii/importance-analysis-about-fqid-features-train/notebook)\n\nThe input table data (pickle file) can be downloaded [here](http://www.kaggle.com/datasets/tsuyoshifujii/importance-analysis-about-fqid-features-fi-data).\n\n## Previous post\n- About \"text\" feature: https://www.kaggle.com/code/tsuyoshifujii/xgboost-using-only-text-which-text-is-important\n- About \"fqid\" feature: https://www.kaggle.com/code/tsuyoshifujii/xgboost-using-only-fqid-feature-importance\n- About \"text_fqid\" feature: https://www.kaggle.com/code/tsuyoshifujii/xgboost-using-only-text-fqid-feature-importance\n\n## Reference\nThis code is based on the following amazing notebooks.\n- https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680\n- https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\n- https://www.kaggle.com/code/pourchot/simple-xgb\n\nThe idea of training on GPU and Inference on CPU was based on the following discussion.\n- https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218","metadata":{}},{"cell_type":"markdown","source":"# SUMMARY\n## THIS NOTEBOOK IS ONGOING!","metadata":{}},{"cell_type":"markdown","source":"# HOW TO CREATE FEATURES IN TRAIN PART\nEach element of \"fqid\" was aggregated to create the following 30 + 1 features.\n\n## 30 features\nFor the following 5 base features (= original_feature), the difference between the corresponding click and the previous one was calculated.\n- elapsed_time\n- room_coor_x\n- room_coor_y\n- screen_coor_x\n- screen_coor_y\n\nFor each of the above, the following 6 statistics were calculated and used as features.\n- average\n- median\n- standard deviation\n- sum\n- min\n- max\n\n## +1 feature\nThe number of times each element of \"fqid\" was recorded was tabulated and used as a feature.\n- frequency","metadata":{}},{"cell_type":"markdown","source":"# IMPORTS AND LOAD DATA","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom sklearn.metrics import f1_score\nfrom xgboost import XGBClassifier\n\nfrom tqdm.notebook import tqdm\nfrom collections import defaultdict\nfrom itertools import combinations\nimport pickle\nimport warnings\nimport gc\n\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:50.843275Z","iopub.execute_input":"2023-04-14T04:35:50.843992Z","iopub.status.idle":"2023-04-14T04:35:52.348500Z","shell.execute_reply.started":"2023-04-14T04:35:50.843940Z","shell.execute_reply":"2023-04-14T04:35:52.347529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f_read = open('/kaggle/input/importance-analysis-about-fqid-features-fi-data/feature_importance_dict.pkl', 'rb')\nfeature_importance_dict = pickle.load(f_read)\nf_read.close()","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.353874Z","iopub.execute_input":"2023-04-14T04:35:52.356175Z","iopub.status.idle":"2023-04-14T04:35:52.385187Z","shell.execute_reply.started":"2023-04-14T04:35:52.356131Z","shell.execute_reply":"2023-04-14T04:35:52.384150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DATA FORMATTING\nIn this notebook, the top 10% of feature importance was extracted for each question and analyzed for its content.","metadata":{}},{"cell_type":"code","source":"def data_format(x):\n    tmp = x.reset_index()\n    tmp['fqid'] = tmp['feature'].apply(lambda x: str(x)[:-6])\n    tmp['cat_tmp'] = tmp['feature'].apply(lambda x: str(x)[-5:-3])\n    tmp['stat'] = tmp['feature'].apply(lambda x: str(x)[-3:])\n    \n    tmp['original_feature'] = 0\n    tmp.loc[tmp['cat_tmp'] == 'cn', 'original_feature'] = 'frequency'\n    tmp.loc[tmp['cat_tmp'] == 'cn', 'stat'] = 'Count'\n    tmp.loc[tmp['cat_tmp'] == 'et', 'original_feature'] = 'elapsed_time'\n    tmp.loc[tmp['cat_tmp'] == 'rx', 'original_feature'] = 'room_coor_x'\n    tmp.loc[tmp['cat_tmp'] == 'ry', 'original_feature'] = 'room_coor_y'\n    tmp.loc[tmp['cat_tmp'] == 'sx', 'original_feature'] = 'screen_coor_x'\n    tmp.loc[tmp['cat_tmp'] == 'sy', 'original_feature'] = 'screen_coor_y'\n    \n    tmp.drop(columns='cat_tmp', axis=1, inplace=True)\n    \n    return tmp","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.389852Z","iopub.execute_input":"2023-04-14T04:35:52.391820Z","iopub.status.idle":"2023-04-14T04:35:52.403503Z","shell.execute_reply.started":"2023-04-14T04:35:52.391772Z","shell.execute_reply":"2023-04-14T04:35:52.402402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fi_dict = {}\nfor t in range(1, 19):\n    fi_dict[str(t)] = data_format(feature_importance_dict[str(t)])\n    \nfi_dict['1'].head(10)","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.409381Z","iopub.execute_input":"2023-04-14T04:35:52.410461Z","iopub.status.idle":"2023-04-14T04:35:52.669596Z","shell.execute_reply.started":"2023-04-14T04:35:52.410399Z","shell.execute_reply":"2023-04-14T04:35:52.668707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TOTAL NUMBER OF FEATURES\nprint('TOTAL NUMBER OF FEATURES')\nfor t in range(1, 19):\n    tmp = fi_dict[str(t)]\n    print(f'Question {t}: {tmp.shape[0]}')","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.673607Z","iopub.execute_input":"2023-04-14T04:35:52.675881Z","iopub.status.idle":"2023-04-14T04:35:52.685126Z","shell.execute_reply.started":"2023-04-14T04:35:52.675840Z","shell.execute_reply":"2023-04-14T04:35:52.684157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# EXTRACT THE TOP 10%\nfi_top_dict = {}\nprint('NUMBER OF TOP 10% FEATURES')\nfor t in range(1, 19):\n    n_pick = int(np.ceil(len(fi_dict[str(t)]) * 0.1))\n    fi_top_dict[str(t)] = fi_dict[str(t)][:n_pick]\n    print(f'Question {t}: {fi_top_dict[str(t)].shape[0]}')\n    \nfi_top_dict['1'].head(10)","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.689412Z","iopub.execute_input":"2023-04-14T04:35:52.691857Z","iopub.status.idle":"2023-04-14T04:35:52.713669Z","shell.execute_reply.started":"2023-04-14T04:35:52.691816Z","shell.execute_reply":"2023-04-14T04:35:52.712774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ANALYSIS FOR EACH \"ORIGINAL_FEATURE\"","metadata":{}},{"cell_type":"code","source":"# SUMMARIZE DATA\noriginal_features = {}\nfor t in range(1, 19):\n    df = fi_dict[str(t)].groupby('original_feature')['feature'].agg('count')\n    df.name = 'n_feature(all)'\n    df = df.reset_index()\n    df = df.set_index('original_feature')\n    \n    df['n_feature(top10%)'] = fi_top_dict[str(t)].groupby('original_feature')['feature'].agg('count')\n    df['extracted_ratio'] = df['n_feature(top10%)'] / df['n_feature(all)']\n    \n    df['importance_mean'] = fi_top_dict[str(t)].groupby('original_feature')['mean'].agg('mean')\n    df['importance_med'] = fi_top_dict[str(t)].groupby('original_feature')['mean'].agg('median')\n    df['importance_max'] = fi_top_dict[str(t)].groupby('original_feature')['mean'].agg('max')\n    df['importance_min'] = fi_top_dict[str(t)].groupby('original_feature')['mean'].agg('min')\n    \n    df = df.reindex(index=['frequency', 'elapsed_time', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y'])\n    original_features[str(t)] = df","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.717722Z","iopub.execute_input":"2023-04-14T04:35:52.719835Z","iopub.status.idle":"2023-04-14T04:35:52.895574Z","shell.execute_reply.started":"2023-04-14T04:35:52.719793Z","shell.execute_reply":"2023-04-14T04:35:52.894520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"original_features['1']","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.897243Z","iopub.execute_input":"2023-04-14T04:35:52.898096Z","iopub.status.idle":"2023-04-14T04:35:52.912342Z","shell.execute_reply.started":"2023-04-14T04:35:52.898011Z","shell.execute_reply":"2023-04-14T04:35:52.911483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"original_features['18']","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.913813Z","iopub.execute_input":"2023-04-14T04:35:52.914881Z","iopub.status.idle":"2023-04-14T04:35:52.928845Z","shell.execute_reply.started":"2023-04-14T04:35:52.914820Z","shell.execute_reply":"2023-04-14T04:35:52.927711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Percentage included in the top 10% (extracted_ratio)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 35))\nplt.subplots_adjust(wspace=0.3, hspace=0.5)\nfor t in range(1, 19):\n    data = original_features[str(t)]\n    plt.subplot(6, 3, t)\n    left = np.arange(data.shape[0])\n    plt.bar(left, data['extracted_ratio'], tick_label=data.index)\n    plt.xticks(rotation=90)\n    plt.ylim([0, 0.45])\n    plt.xlabel('original_feature')\n    plt.ylabel('extracted_ratio')\n    plt.title(f'Question {t}')\n    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:52.934788Z","iopub.execute_input":"2023-04-14T04:35:52.935213Z","iopub.status.idle":"2023-04-14T04:35:55.546284Z","shell.execute_reply.started":"2023-04-14T04:35:52.935180Z","shell.execute_reply":"2023-04-14T04:35:55.545334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (all data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 40))\nplt.subplots_adjust(wspace=0.3, hspace=0.5)\nx_label = ['frequency', 'elapsed_time', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y']\nfor t in range(1, 19):\n    data = fi_dict[str(t)]\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='original_feature', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:55.547320Z","iopub.execute_input":"2023-04-14T04:35:55.547650Z","iopub.status.idle":"2023-04-14T04:35:59.051523Z","shell.execute_reply.started":"2023-04-14T04:35:55.547619Z","shell.execute_reply":"2023-04-14T04:35:59.050683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (top 10% data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 40))\nplt.subplots_adjust(wspace=0.3, hspace=0.5)\nx_label = ['frequency', 'elapsed_time', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y']\nfor t in range(1, 19):\n    data = fi_top_dict[str(t)]\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='original_feature', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-14T04:35:59.053065Z","iopub.execute_input":"2023-04-14T04:35:59.053716Z","iopub.status.idle":"2023-04-14T04:36:02.616803Z","shell.execute_reply.started":"2023-04-14T04:35:59.053679Z","shell.execute_reply":"2023-04-14T04:36:02.615854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In original_feature, \"frequency\" and \"elapsed_time\" seems to be relatively important and are included in the top 10%. On the other hand, the other four are relatively less important with some exceptions.","metadata":{}},{"cell_type":"markdown","source":"# FURTHER BREAKDOWN BY STATISTIC","metadata":{}},{"cell_type":"markdown","source":"## elapsed_time","metadata":{}},{"cell_type":"code","source":"# SUMMARIZE DATA\nelapsed_times = {}\nfor t in range(1, 19):\n    df = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'elapsed_time'].groupby('stat')['feature'].agg('count')\n    df.name = 'n_feature(all)'\n    df = df.reset_index()\n    df = df.set_index('stat')\n    \n    df['n_feature(top10%)'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'elapsed_time'].groupby('stat')['feature'].agg('count')\n    df['extracted_ratio'] = df['n_feature(top10%)'] / df['n_feature(all)']\n    \n    df['importance_mean'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'elapsed_time'].groupby('stat')['mean'].agg('mean')\n    df['importance_med'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'elapsed_time'].groupby('stat')['mean'].agg('median')\n    df['importance_max'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'elapsed_time'].groupby('stat')['mean'].agg('max')\n    df['importance_min'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'elapsed_time'].groupby('stat')['mean'].agg('min')\n    \n    df = df.reindex(index=['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max'])\n    elapsed_times[str(t)] = df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:02.618139Z","iopub.execute_input":"2023-04-14T04:36:02.618674Z","iopub.status.idle":"2023-04-14T04:36:02.838177Z","shell.execute_reply.started":"2023-04-14T04:36:02.618639Z","shell.execute_reply":"2023-04-14T04:36:02.837273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Percentage included in the top 10% (extracted_ratio)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nfor t in range(1, 19):\n    data = elapsed_times[str(t)]\n    plt.subplot(6, 3, t)\n    left = np.arange(data.shape[0])\n    plt.bar(left, data['extracted_ratio'], tick_label=data.index)\n    plt.xticks(rotation=90)\n    plt.ylim([0, 0.5])\n    plt.xlabel('stat')\n    plt.ylabel('extracted_ratio')\n    plt.title(f'Question {t}')\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:02.839445Z","iopub.execute_input":"2023-04-14T04:36:02.840009Z","iopub.status.idle":"2023-04-14T04:36:04.924948Z","shell.execute_reply.started":"2023-04-14T04:36:02.839973Z","shell.execute_reply":"2023-04-14T04:36:04.924125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (all data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'elapsed_time']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2023-04-14T04:36:04.926232Z","iopub.execute_input":"2023-04-14T04:36:04.926727Z","iopub.status.idle":"2023-04-14T04:36:08.228549Z","shell.execute_reply.started":"2023-04-14T04:36:04.926694Z","shell.execute_reply":"2023-04-14T04:36:08.227649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (top 10% data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'elapsed_time']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:08.229878Z","iopub.execute_input":"2023-04-14T04:36:08.230403Z","iopub.status.idle":"2023-04-14T04:36:11.719248Z","shell.execute_reply.started":"2023-04-14T04:36:08.230370Z","shell.execute_reply":"2023-04-14T04:36:11.715652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Sum > Med > Max > other, it seems.","metadata":{}},{"cell_type":"markdown","source":"## room_coor_x","metadata":{}},{"cell_type":"code","source":"# SUMMARIZE DATA\nroom_coor_xs = {}\nfor t in range(1, 19):\n    df = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'room_coor_x'].groupby('stat')['feature'].agg('count')\n    df.name = 'n_feature(all)'\n    df = df.reset_index()\n    df = df.set_index('stat')\n    \n    df['n_feature(top10%)'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_x'].groupby('stat')['feature'].agg('count')\n    df['extracted_ratio'] = df['n_feature(top10%)'] / df['n_feature(all)']\n    \n    df['importance_mean'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_x'].groupby('stat')['mean'].agg('mean')\n    df['importance_med'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_x'].groupby('stat')['mean'].agg('median')\n    df['importance_max'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_x'].groupby('stat')['mean'].agg('max')\n    df['importance_min'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_x'].groupby('stat')['mean'].agg('min')\n    \n    df = df.reindex(index=['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max'])\n    room_coor_xs[str(t)] = df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:11.720423Z","iopub.execute_input":"2023-04-14T04:36:11.721028Z","iopub.status.idle":"2023-04-14T04:36:11.941853Z","shell.execute_reply.started":"2023-04-14T04:36:11.720988Z","shell.execute_reply":"2023-04-14T04:36:11.940649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Percentage included in the top 10% (extracted_ratio)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nfor t in range(1, 19):\n    data = room_coor_xs[str(t)]\n    plt.subplot(6, 3, t)\n    left = np.arange(data.shape[0])\n    plt.bar(left, data['extracted_ratio'], tick_label=data.index)\n    plt.xticks(rotation=90)\n    plt.ylim([0, 0.2])\n    plt.xlabel('stat')\n    plt.ylabel('extracted_ratio')\n    plt.title(f'Question {t}')\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:11.943080Z","iopub.execute_input":"2023-04-14T04:36:11.943379Z","iopub.status.idle":"2023-04-14T04:36:14.250005Z","shell.execute_reply.started":"2023-04-14T04:36:11.943349Z","shell.execute_reply":"2023-04-14T04:36:14.249035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (all data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'room_coor_x']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:14.251267Z","iopub.execute_input":"2023-04-14T04:36:14.251785Z","iopub.status.idle":"2023-04-14T04:36:17.702987Z","shell.execute_reply.started":"2023-04-14T04:36:14.251751Z","shell.execute_reply":"2023-04-14T04:36:17.700972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (top 10% data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_x']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:17.704736Z","iopub.execute_input":"2023-04-14T04:36:17.705109Z","iopub.status.idle":"2023-04-14T04:36:21.201222Z","shell.execute_reply.started":"2023-04-14T04:36:17.705072Z","shell.execute_reply":"2023-04-14T04:36:21.200263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unlike \"elapsed_time\", no clear trend was obtained. If pressed, \"Std\" may be more important than the others.","metadata":{}},{"cell_type":"markdown","source":"## room_coor_y","metadata":{}},{"cell_type":"code","source":"# SUMMARIZE DATA\nroom_coor_ys = {}\nfor t in range(1, 19):\n    df = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'room_coor_y'].groupby('stat')['feature'].agg('count')\n    df.name = 'n_feature(all)'\n    df = df.reset_index()\n    df = df.set_index('stat')\n    \n    df['n_feature(top10%)'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_y'].groupby('stat')['feature'].agg('count')\n    df['extracted_ratio'] = df['n_feature(top10%)'] / df['n_feature(all)']\n    \n    df['importance_mean'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_y'].groupby('stat')['mean'].agg('mean')\n    df['importance_med'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_y'].groupby('stat')['mean'].agg('median')\n    df['importance_max'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_y'].groupby('stat')['mean'].agg('max')\n    df['importance_min'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_y'].groupby('stat')['mean'].agg('min')\n    \n    df = df.reindex(index=['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max'])\n    room_coor_ys[str(t)] = df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:21.202534Z","iopub.execute_input":"2023-04-14T04:36:21.203702Z","iopub.status.idle":"2023-04-14T04:36:21.427646Z","shell.execute_reply.started":"2023-04-14T04:36:21.203663Z","shell.execute_reply":"2023-04-14T04:36:21.426546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Percentage included in the top 10% (extracted_ratio)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nfor t in range(1, 19):\n    data = room_coor_ys[str(t)]\n    plt.subplot(6, 3, t)\n    left = np.arange(data.shape[0])\n    plt.bar(left, data['extracted_ratio'], tick_label=data.index)\n    plt.xticks(rotation=90)\n    plt.ylim([0, 0.2])\n    plt.xlabel('stat')\n    plt.ylabel('extracted_ratio')\n    plt.title(f'Question {t}')\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:21.429251Z","iopub.execute_input":"2023-04-14T04:36:21.429684Z","iopub.status.idle":"2023-04-14T04:36:23.730428Z","shell.execute_reply.started":"2023-04-14T04:36:21.429649Z","shell.execute_reply":"2023-04-14T04:36:23.729460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (all data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'room_coor_y']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:23.731983Z","iopub.execute_input":"2023-04-14T04:36:23.732645Z","iopub.status.idle":"2023-04-14T04:36:27.107690Z","shell.execute_reply.started":"2023-04-14T04:36:23.732592Z","shell.execute_reply":"2023-04-14T04:36:27.106669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (top 10% data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'room_coor_y']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:27.109264Z","iopub.execute_input":"2023-04-14T04:36:27.109901Z","iopub.status.idle":"2023-04-14T04:36:30.522738Z","shell.execute_reply.started":"2023-04-14T04:36:27.109865Z","shell.execute_reply":"2023-04-14T04:36:30.521798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Same as \"room_coor_x\". Only \"Std\" and \"Max\" stand out a little.","metadata":{}},{"cell_type":"markdown","source":"## screen_coor_x","metadata":{}},{"cell_type":"code","source":"# SUMMARIZE DATA\nscreen_coor_xs = {}\nfor t in range(1, 19):\n    df = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'screen_coor_x'].groupby('stat')['feature'].agg('count')\n    df.name = 'n_feature(all)'\n    df = df.reset_index()\n    df = df.set_index('stat')\n    \n    df['n_feature(top10%)'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_x'].groupby('stat')['feature'].agg('count')\n    df['extracted_ratio'] = df['n_feature(top10%)'] / df['n_feature(all)']\n    \n    df['importance_mean'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_x'].groupby('stat')['mean'].agg('mean')\n    df['importance_med'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_x'].groupby('stat')['mean'].agg('median')\n    df['importance_max'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_x'].groupby('stat')['mean'].agg('max')\n    df['importance_min'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_x'].groupby('stat')['mean'].agg('min')\n    \n    df = df.reindex(index=['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max'])\n    screen_coor_xs[str(t)] = df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:30.524277Z","iopub.execute_input":"2023-04-14T04:36:30.524946Z","iopub.status.idle":"2023-04-14T04:36:30.743992Z","shell.execute_reply.started":"2023-04-14T04:36:30.524909Z","shell.execute_reply":"2023-04-14T04:36:30.743035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Percentage included in the top 10% (extracted_ratio)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nfor t in range(1, 19):\n    data = screen_coor_xs[str(t)]\n    plt.subplot(6, 3, t)\n    left = np.arange(data.shape[0])\n    plt.bar(left, data['extracted_ratio'], tick_label=data.index)\n    plt.xticks(rotation=90)\n    plt.ylim([0, 0.2])\n    plt.xlabel('stat')\n    plt.ylabel('extracted_ratio')\n    plt.title(f'Question {t}')\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:30.745548Z","iopub.execute_input":"2023-04-14T04:36:30.746217Z","iopub.status.idle":"2023-04-14T04:36:33.045151Z","shell.execute_reply.started":"2023-04-14T04:36:30.746181Z","shell.execute_reply":"2023-04-14T04:36:33.044084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (all data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'screen_coor_x']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:33.046508Z","iopub.execute_input":"2023-04-14T04:36:33.046915Z","iopub.status.idle":"2023-04-14T04:36:36.347064Z","shell.execute_reply.started":"2023-04-14T04:36:33.046883Z","shell.execute_reply":"2023-04-14T04:36:36.345999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (top 10% data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_x']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:36.351984Z","iopub.execute_input":"2023-04-14T04:36:36.352308Z","iopub.status.idle":"2023-04-14T04:36:39.870468Z","shell.execute_reply.started":"2023-04-14T04:36:36.352277Z","shell.execute_reply":"2023-04-14T04:36:39.869344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\"Std\", \"Sum\", and \"Max\" seem to be more important than the others.","metadata":{}},{"cell_type":"markdown","source":"## screen_coor_y","metadata":{}},{"cell_type":"code","source":"# SUMMARIZE DATA\nscreen_coor_ys = {}\nfor t in range(1, 19):\n    df = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'screen_coor_y'].groupby('stat')['feature'].agg('count')\n    df.name = 'n_feature(all)'\n    df = df.reset_index()\n    df = df.set_index('stat')\n    \n    df['n_feature(top10%)'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_y'].groupby('stat')['feature'].agg('count')\n    df['extracted_ratio'] = df['n_feature(top10%)'] / df['n_feature(all)']\n    \n    df['importance_mean'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_y'].groupby('stat')['mean'].agg('mean')\n    df['importance_med'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_y'].groupby('stat')['mean'].agg('median')\n    df['importance_max'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_y'].groupby('stat')['mean'].agg('max')\n    df['importance_min'] = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_y'].groupby('stat')['mean'].agg('min')\n    \n    df = df.reindex(index=['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max'])\n    screen_coor_ys[str(t)] = df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:39.871957Z","iopub.execute_input":"2023-04-14T04:36:39.872281Z","iopub.status.idle":"2023-04-14T04:36:40.098131Z","shell.execute_reply.started":"2023-04-14T04:36:39.872250Z","shell.execute_reply":"2023-04-14T04:36:40.096026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Percentage included in the top 10% (extracted_ratio)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nfor t in range(1, 19):\n    data = screen_coor_ys[str(t)]\n    plt.subplot(6, 3, t)\n    left = np.arange(data.shape[0])\n    plt.bar(left, data['extracted_ratio'], tick_label=data.index)\n    plt.xticks(rotation=90)\n    plt.ylim([0, 0.2])\n    plt.xlabel('stat')\n    plt.ylabel('extracted_ratio')\n    plt.title(f'Question {t}')\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:40.100437Z","iopub.execute_input":"2023-04-14T04:36:40.101083Z","iopub.status.idle":"2023-04-14T04:36:42.384358Z","shell.execute_reply.started":"2023-04-14T04:36:40.101044Z","shell.execute_reply":"2023-04-14T04:36:42.383494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (all data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_dict[str(t)][fi_dict[str(t)]['original_feature'] == 'screen_coor_y']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:42.385617Z","iopub.execute_input":"2023-04-14T04:36:42.386137Z","iopub.status.idle":"2023-04-14T04:36:45.907128Z","shell.execute_reply.started":"2023-04-14T04:36:42.386102Z","shell.execute_reply":"2023-04-14T04:36:45.906211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Box-plots for importance (top 10% data)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 30))\nplt.subplots_adjust(wspace=0.3, hspace=0.3)\nx_label = ['Ave', 'Med', 'Std', 'Sum', 'Min', 'Max']\nfor t in range(1, 19):\n    data = fi_top_dict[str(t)][fi_top_dict[str(t)]['original_feature'] == 'screen_coor_y']\n    plt.subplot(6, 3, t)\n    sns.boxplot(data=data, x='stat', y='mean', order=x_label)\n    plt.title(f'Question {t}')\n    plt.xticks(rotation=90)\n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-14T04:36:45.908435Z","iopub.execute_input":"2023-04-14T04:36:45.908966Z","iopub.status.idle":"2023-04-14T04:36:49.480093Z","shell.execute_reply.started":"2023-04-14T04:36:45.908923Z","shell.execute_reply":"2023-04-14T04:36:49.479149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Is \"Std\", \"Min\", \"Max\" relatively important here?","metadata":{}}]}