{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"このnotebookは以下のスコア0.699のnotebookに日本語訳を追加したものです。\n日本人Kagglerの手助けになれば幸いです。\nコメントか投票をいただけると嬉しいです。\n\nThis notebook is a Japanese translation of the following notebook with a score of 0.699. I hope it helps Japanese Kaggler. I would appreciate your comments or votes.\n\nhttps://www.kaggle.com/code/vadimkamaev/catboost","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport pickle\nimport sys\n\nfrom catboost import CatBoostClassifier","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Train Data and Labels","metadata":{"papermill":{"duration":0.004542,"end_time":"2023-02-07T00:59:59.189777","exception":false,"start_time":"2023-02-07T00:59:59.185235","status":"completed"},"tags":[]}},{"cell_type":"code","source":"dtypes = {\"session_id\": 'int64',\n          \"index\": np.int16,\n          \"elapsed_time\": np.int32,\n          \"event_name\": 'category',\n          \"name\": 'category',\n          \"level\": np.int8,\n          \"page\": np.float16,\n          \"room_coor_x\": np.float16,\n          \"room_coor_y\": np.float16,\n          \"screen_coor_x\": np.float16,\n          \"screen_coor_y\": np.float16,\n          \"hover_duration\": np.float32,\n          \"text\": 'category',\n          \"fqid\": 'category',\n          \"room_fqid\": 'category',\n          \"text_fqid\": 'category',\n          \"fullscreen\": np.int8,\n          \"hq\": np.int8,\n          \"music\": np.int8,\n          \"level_group\": 'category'\n          }\nuse_col = ['session_id', 'index', 'elapsed_time', 'event_name', 'name', 'level', 'page',\n           'room_coor_x', 'room_coor_y', 'hover_duration', 'text', 'fqid', 'room_fqid', 'text_fqid', 'level_group']","metadata":{"papermill":{"duration":59.284316,"end_time":"2023-02-07T01:00:58.478743","exception":false,"start_time":"2023-02-07T00:59:59.194427","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T12:14:07.014291Z","iopub.execute_input":"2023-06-05T12:14:07.014712Z","iopub.status.idle":"2023-06-05T12:14:07.024972Z","shell.execute_reply.started":"2023-06-05T12:14:07.014674Z","shell.execute_reply":"2023-06-05T12:14:07.023895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )\ntargets['q'] = targets.session_id.apply(lambda x: int(x.split('_')[-1][1:]) )\n# print( targets.shape )\n# targets.head()","metadata":{"papermill":{"duration":0.598155,"end_time":"2023-02-07T01:00:59.082015","exception":false,"start_time":"2023-02-07T01:00:58.48386","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T12:14:11.292167Z","iopub.execute_input":"2023-06-05T12:14:11.292619Z","iopub.status.idle":"2023-06-05T12:14:12.904083Z","shell.execute_reply.started":"2023-06-05T12:14:11.292582Z","shell.execute_reply":"2023-06-05T12:14:12.902983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_df = pd.read_csv('/kaggle/input/featur/feature_sort.csv')","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:17.196753Z","iopub.execute_input":"2023-06-05T12:14:17.197175Z","iopub.status.idle":"2023-06-05T12:14:17.345285Z","shell.execute_reply.started":"2023-06-05T12:14:17.197138Z","shell.execute_reply":"2023-06-05T12:14:17.344319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineer","metadata":{"papermill":{"duration":0.005196,"end_time":"2023-02-07T01:00:59.092865","exception":false,"start_time":"2023-02-07T01:00:59.087669","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"df を 'session_id' 列と 'elapsed_time' 列でソートしています。inplace=True は、ソート結果を元の DataFrame に反映することを意味します。\n\n'elapsed_time' 列の前の行との差分を計算して、'd_time' 列に代入しています。diff(1) は、前の行との差分を計算するメソッドです。\n\n'd_time' 列の欠損値（NaN）を 0 で埋めています。fillna(0) は、欠損値を指定した値で埋めるメソッドです。\n\n'd_time' 列の値を 0 から 103000 の範囲にクリップして、'delt_time' 列に代入しています。clip(0, 103000) は、値を指定した範囲内にクリップするメソッドです。\n\n'delt_time' 列の値を 1 行後ろにシフトして、'delt_time_next' 列に代入しています。shift(-1) は、値を指定した数だけ行方向にシフトするメソッドです。\n\n変更した DataFrame df を関数の結果として返します。","metadata":{}},{"cell_type":"code","source":"def delt_time_def(df):\n    df.sort_values(by=['session_id', 'elapsed_time'], inplace=True)\n    df['d_time'] = df['elapsed_time'].diff(1)\n    df['d_time'].fillna(0, inplace=True)\n    df['delt_time'] = df['d_time'].clip(0)\n    df['delt_time'] = df['d_time'].clip(0, 103000)\n    df['delt_time_next'] = df['delt_time'].shift(-1)\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:20.463967Z","iopub.execute_input":"2023-06-05T12:14:20.464527Z","iopub.status.idle":"2023-06-05T12:14:20.472450Z","shell.execute_reply.started":"2023-06-05T12:14:20.464484Z","shell.execute_reply":"2023-06-05T12:14:20.471478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"kol_col と kol_col_max というグローバル変数を定義しています。それぞれの値は、kol_col は 9 に、kol_col_max は 11+kol_f*2 に設定されます。kol_f は引数として渡される値です。\n\n新しいデータフレーム new_train を作成しています。カラムの数は kol_col_max で指定された値になります。index=train['session_id'].unique() は、元のデータフレーム train の 'session_id' 列のユニークな値をインデックスとして使用します。\n\nnew_train の 10 列に 'session_id' を設定しています。\n\ntrain データフレームを 'session_id' でグループ化し、'd_time' 列の値のパーセンタイルを計算して、new_train の 0～3 列に代入しています。\n\ntrain データフレームを 'session_id' でグループ化し、'd_time' 列の値の平均値と標準偏差を計算して、new_train の 4～5 列に代入しています。\n\nnew_train の 10 列に 'session_id'が格納されているので、year, month, day, hourに分割して、new_train の 6～9 列に代入しています。\n\nnew_train の 10 列に0を代入しています。\n\nnew_train の 欠損値を-1で埋めます。","metadata":{}},{"cell_type":"code","source":"def feature_engineer(train, kol_f):\n    global kol_col, kol_col_max\n    kol_col = 9\n    kol_col_max = 11+kol_f*2\n    col = [i for i in range(0,kol_col_max)]\n    new_train = pd.DataFrame(index=train['session_id'].unique(), columns=col, dtype=np.float16)  \n    new_train[10] = new_train.index # \"session_id\"    \n\n    new_train[0] = train.groupby(['session_id'])['d_time'].quantile(q=0.3)\n    new_train[1] = train.groupby(['session_id'])['d_time'].quantile(q=0.8)\n    new_train[2] = train.groupby(['session_id'])['d_time'].quantile(q=0.5)\n    new_train[3] = train.groupby(['session_id'])['d_time'].quantile(q=0.65)\n    new_train[4] = train.groupby(['session_id'])['hover_duration'].agg('mean')\n    new_train[5] = train.groupby(['session_id'])['hover_duration'].agg('std')    \n    new_train[6] = new_train[10].apply(lambda x: int(str(x)[:2])).astype(np.uint8) # \"year\"\n    new_train[7] = new_train[10].apply(lambda x: int(str(x)[2:4])+1).astype(np.uint8) # \"month\"\n    new_train[8] = new_train[10].apply(lambda x: int(str(x)[4:6])).astype(np.uint8) # \"day\"\n    new_train[9] = new_train[10].apply(lambda x: int(str(x)[6:8])).astype(np.uint8) + new_train[10].apply(lambda x: int(str(x)[8:10])).astype(np.uint8)/60\n    new_train[10] = 0\n    new_train = new_train.fillna(-1)\n    \n    return new_train","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:21.691219Z","iopub.execute_input":"2023-06-05T12:14:21.691654Z","iopub.status.idle":"2023-06-05T12:14:21.713737Z","shell.execute_reply.started":"2023-06-05T12:14:21.691617Z","shell.execute_reply":"2023-06-05T12:14:21.712709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"グローバル変数 kol_col の値を1増やしています。\n\nrow_f から 'col1' と 'val1' の値を取得し、条件式 train[col1] == val1 を用いて maska を作成しています。\n\nrow_f の 'kol_col' の値に応じて条件分岐を行っています。'kol_col' の値が 1 の場合は、train データフレームを maska でフィルタリングし、'delt_time_next' 列を 'session_id' でグループ化して合計値を計算し、それを new_train の kol_col 列に代入しています。さらに、gran_1 が真の場合は、同様に平均値を計算し、gran_2 が真の場合は、'index' 列の個数を計算して代入しています。\n\n'kol_col' の値が 2 の場合も同様に処理が行われますが、追加で 'col2' と 'val2' の値を取得し、maska を更新しています。\n\n最後に、new_train を返して関数の処理を終了します。","metadata":{}},{"cell_type":"code","source":"def feature_next_t(row_f, new_train, train, gran_1, gran_2, i):\n    global kol_col\n    kol_col +=1\n    col1 = row_f['col1']\n    val1 = row_f['val1']\n    maska = (train[col1] == val1)\n    if row_f['kol_col'] == 1:       \n        new_train[kol_col] = train[maska].groupby(['session_id'])['delt_time_next'].sum()\n        if gran_1:\n            kol_col +=1\n            new_train[kol_col] = train[maska].groupby(['session_id'])['delt_time'].mean()\n        if gran_2:\n            kol_col +=1\n            new_train[kol_col] = train[maska].groupby(['session_id'])['index'].count()          \n    elif row_f['kol_col'] == 2: \n        col2 = row_f['col2']\n        val2 = row_f['val2']\n        maska = maska & (train[col2] == val2)        \n        new_train[kol_col] = train[maska].groupby(['session_id'])['delt_time_next'].sum()\n        if gran_1:\n            kol_col +=1\n            new_train[kol_col] = train[maska].groupby(['session_id'])['delt_time'].mean()\n        if gran_2:\n            kol_col +=1\n            new_train[kol_col] = train[maska].groupby(['session_id'])['index'].count()\n    return new_train","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:22.610256Z","iopub.execute_input":"2023-06-05T12:14:22.611078Z","iopub.status.idle":"2023-06-05T12:14:22.627464Z","shell.execute_reply.started":"2023-06-05T12:14:22.611035Z","shell.execute_reply":"2023-06-05T12:14:22.626415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"グローバル変数 kol_col の値を1増やしています。\n\nこの行では、row_f から 'col1' と 'val1' の値を取得し、条件式 train[col1] == val1 を用いて maska を作成しています。\n\nrow_f の 'kol_col' の値に応じて条件分岐を行っています。'kol_col' の値が 1 の場合は、train データフレームを maska でフィルタリングし、'delt_time_next' 列の合計値を計算し、それを new_train の kol_col 列に代入しています。さらに、gran_1 が真の場合は、同様に平均値を計算し、gran_2 が真の場合は、'index' 列の個数を計算して代入しています。\n\n'kol_col' の値が 2 の場合も同様に処理が行われますが、追加で 'col2' と 'val2' の値を取得し、maska を更新しています。\n\n最後に、new_train を返して関数の処理を終了します。","metadata":{}},{"cell_type":"code","source":"def feature_next_t_otvet(row_f, new_train, train, gran_1, gran_2, i):\n    global kol_col\n    kol_col +=1\n    col1 = row_f['col1']\n    val1 = row_f['val1']\n    maska = (train[col1] == val1)\n    if row_f['kol_col'] == 1:      \n        new_train[kol_col] = train[maska]['delt_time_next'].sum()\n        if gran_1:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['delt_time'].mean()\n        if gran_2:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['index'].count()          \n    elif row_f['kol_col'] == 2: \n        col2 = row_f['col2']\n        val2 = row_f['val2']\n        maska = maska & (train[col2] == val2)        \n        new_train[kol_col] = train[maska]['delt_time_next'].sum()\n        if gran_1:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['delt_time'].mean()\n        if gran_2:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['index'].count()\n    return new_train","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:24.200787Z","iopub.execute_input":"2023-06-05T12:14:24.201190Z","iopub.status.idle":"2023-06-05T12:14:24.216960Z","shell.execute_reply.started":"2023-06-05T12:14:24.201156Z","shell.execute_reply":"2023-06-05T12:14:24.215927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"グローバル変数 kol_col の値を1増やしています。\n\nrow_f の 'kol_col' の値に応じて条件分岐を行っています。'kol_col' の値が 1 の場合は、row_f の 'col1' 列と 'val1' の値に基づいて maska を作成しています。その後、train データフレームを maska でフィルタリングし、'delt_time_next' 列の合計値を計算して new_train の kol_col 列に代入しています。さらに、gran_1 が真の場合は、同様に平均値を計算し、gran_2 が真の場合は、'index' 列の個数を計算して代入しています。\n\n'kol_col' の値が 2 の場合も同様に処理が行われますが、追加で 'col2' と 'val2' の値を取得し、maska を更新しています。\n\n最後に、new_train を返して関数の処理を終了します。","metadata":{}},{"cell_type":"code","source":"def experiment_feature_next_t_otvet(row_f, new_train, train, gran_1, gran_2, i):\n    global kol_col\n    kol_col +=1\n    if row_f['kol_col'] == 1: \n        maska = train[row_f['col1']] == row_f['val1']\n        new_train[kol_col] = train[maska]['delt_time_next'].sum()\n        if gran_1:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['delt_time'].mean()\n        if gran_2:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['index'].count()          \n    elif row_f['kol_col'] == 2: \n        col2 = row_f['col2']\n        val2 = row_f['val2']\n        maska = (train[col1] == val1) & (train[col2] == val2)        \n        new_train[kol_col] = train[maska]['delt_time_next'].sum()\n        if gran_1:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['delt_time'].mean()\n        if gran_2:\n            kol_col +=1\n            new_train[kol_col] = train[maska]['index'].count()\n    return new_train","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:29.378667Z","iopub.execute_input":"2023-06-05T12:14:29.379081Z","iopub.status.idle":"2023-06-05T12:14:29.392080Z","shell.execute_reply.started":"2023-06-05T12:14:29.379045Z","shell.execute_reply":"2023-06-05T12:14:29.391055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"グローバル変数 kol_col の値を 9 に設定し、g1 と g2 にそれぞれ 0.7 と 0.3 の値を割り当てています。\n\nfeature_df データフレームから 'quest' 列が quest と一致する行をフィルタリングして feature_q にコピーし、インデックスをリセットしています。\n\nkol_f を g1 と g2 に乗じて四捨五入した値を gran1 と gran2 に設定しています。そして、0から kol_f-1 の範囲でループを回しています。各ループでは、feature_q の i 番目の行を row_f に代入し、feature_next_t_otvet 関数を呼び出しています。この関数には、row_f と各種パラメータを渡しており、それに基づいて new_train を更新しています。\n\n最後に、new_train のカラム番号が 0 から kol_col までの範囲の部分を抽出して返しています。\n\nこのコードは、与えられた quest に対して特徴量の計算と追加を行う関数です。関数内で、feature_next_t_otvet 関数が呼び出されており、各行の特徴量の追加処理が行われています。また、関数の最後では特徴量が追加された new_train の一部を抽出して返しています。","metadata":{}},{"cell_type":"code","source":"def feature_quest_otvet(new_train, train, quest, kol_f):\n    global kol_col\n    kol_col = 9\n    g1 = 0.7 \n    g2 = 0.3 \n\n    feature_q = feature_df[feature_df['quest'] == quest].copy()\n    feature_q.reset_index(drop=True, inplace=True)\n    \n    gran1 = round(kol_f * g1)\n    gran2 = round(kol_f * g2)    \n    for i in range(0, kol_f):         \n        row_f = feature_q.loc[i]\n        new_train = feature_next_t_otvet(row_f, new_train, train, i < gran1, i <  gran2, i) \n    col = [i for i in range(0,kol_col+1)]\n    return new_train[col]","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:32.929753Z","iopub.execute_input":"2023-06-05T12:14:32.930167Z","iopub.status.idle":"2023-06-05T12:14:32.939962Z","shell.execute_reply.started":"2023-06-05T12:14:32.930130Z","shell.execute_reply":"2023-06-05T12:14:32.938382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"g1 に 0.7、g2 に 0.3 の値を割り当て、kol_f をそれぞれ乗じて四捨五入した結果を gran1 と gran2 に代入しています。\n\n0から kol_f-1 の範囲でループを回しています。各ループでは、feature_q の i 番目の行を row_f に代入し、feature_next_t 関数を呼び出しています。この関数には、row_f と各種パラメータを渡しており、それに基づいて new_train を更新しています。\n\n最後に、更新された new_train を返しています。\n\nこのコードは、与えられた new_train と train データフレームに対して特徴量エンジニアリングを行う関数です。feature_q データフレームから行を取り出し、feature_next_t 関数を呼び出して特徴量の追加を行います。最終的に、更新された new_train を返します。","metadata":{}},{"cell_type":"code","source":"def feature_engineer_new(new_train, train, feature_q, kol_f):\n    g1 = 0.7 \n    g2 = 0.3 \n    gran1 = round(kol_f * g1)\n    gran2 = round(kol_f * g2)    \n    for i in range(0, kol_f): \n        row_f = feature_q.loc[i]       \n        new_train = feature_next_t(row_f, new_train, train, i < gran1, i <  gran2, i)         \n    return new_train","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:36.563973Z","iopub.execute_input":"2023-06-05T12:14:36.564407Z","iopub.status.idle":"2023-06-05T12:14:36.572855Z","shell.execute_reply.started":"2023-06-05T12:14:36.564355Z","shell.execute_reply":"2023-06-05T12:14:36.571307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"グローバル変数 kol_col に値 9 を代入しています。\n\nfeature_df データフレームから 'quest' 列の値が quest と等しい行のみを抽出して、feature_q にコピーします。そして、インデックスをリセットしています。\n\nfeature_engineer_new 関数を呼び出して特徴量エンジニアリングを行い、その結果を new_train に代入しています。new_train と train データフレーム、feature_q、および kol_f を引数として渡しています。\n\nインデックスが 0 から kol_col までの範囲を持つリストを作成し、new_train のそれらの列だけを抽出して返します。\n\nこのコードは、与えられた new_train と train データフレームに対して、指定された 'quest' 値に対応する特徴量エンジニアリングを行う関数です。feature_q データフレームから必要な行を抽出し、feature_engineer_new 関数を呼び出して特徴量の追加を行います。最後に、特定の列だけを持つ new_train を返します。","metadata":{}},{"cell_type":"code","source":"def feature_quest(new_train, train, quest, kol_f):\n    global kol_col\n    kol_col = 9\n    feature_q = feature_df[feature_df['quest'] == quest].copy()\n    feature_q.reset_index(drop=True, inplace=True)\n    new_train = feature_engineer_new(new_train, train, feature_q, kol_f)\n    col = [i for i in range(0,kol_col+1)]\n    return new_train[col]","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:40.644569Z","iopub.execute_input":"2023-06-05T12:14:40.644981Z","iopub.status.idle":"2023-06-05T12:14:40.652809Z","shell.execute_reply.started":"2023-06-05T12:14:40.644946Z","shell.execute_reply":"2023-06-05T12:14:40.651562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"questsリストの要素数を kol_quest に代入しています。\n\nquestsリストをイテレーションしながら、各要素 q について処理を行います。また、q を出力しています。\n\nfeature_engineer 関数を使用して、old_train データフレームを特徴量エンジニアリングし、その結果を new_train に代入しています。list_kol_f[q] を feature_engineer の引数として渡しています。\n\nfeature_quest 関数を使用して、new_train と old_train データフレームを特徴量エンジニアリングし、その結果を train_x に代入しています。q と list_kol_f[q] を feature_quest の引数として渡しています。\n\ntrain_x の形状を出力しています。\n\ntrain_x のインデックス値を train_users に代入し、targets データフレームから q に対応する行を抽出して train_y に代入しています。\n\nCatBoostClassifier モデルをインスタンス化して model に代入しています。\n\ntrain_x を float32 型にキャストし、train_y['correct'] を正解ラベルとしてモデルを学習させます。\n\nmodels ディクショナリに q をキーとして学習したモデルを追加します。\n\n処理が終了したことを示す出力を行い、models を返します。\n\nこのコードは、与えられたデータと設定に基づいて複数のモデルを作成する関数です。各質問ごとに特徴量エンジニアリングを行い、トレーニングデータを作成し、CatBoostClassifier モデルを学習させます。学習済みモデルは models ディクショナリに格納され、最終的に返されます。","metadata":{}},{"cell_type":"code","source":"def create_model(old_train, quests, models, list_kol_f):\n    \n    kol_quest = len(quests)\n    # ITERATE THRU QUESTIONS\n    for q in quests:\n        print('### quest ', q, end='')\n        new_train = feature_engineer(old_train, list_kol_f[q])\n        train_x = feature_quest(new_train, old_train, q, list_kol_f[q])\n        print (' ---- ', 'train_q.shape = ', train_x.shape)\n           \n        # TRAIN DATA\n        train_users = train_x.index.values\n        train_y = targets.loc[targets.q==q].set_index('session').loc[train_users]\n\n        # TRAIN MODEL \n\n        #model = CatBoostClassifier(\n        #    n_estimators = 300,\n        #    learning_rate= 0.045,\n        #    depth = 6\n        #)\n        \n        model = CatBoostClassifier(\n            n_estimators = 300,\n            learning_rate= 0.045,\n            depth = 5\n        )\n        \n        model.fit(train_x.astype('float32'), train_y['correct'], verbose=False)\n\n        # SAVE MODEL, PREDICT VALID OOF\n        models[f'{q}'] = model\n    print('***')\n    \n    return models","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:14:45.582655Z","iopub.execute_input":"2023-06-05T12:14:45.583062Z","iopub.status.idle":"2023-06-05T12:14:45.593661Z","shell.execute_reply.started":"2023-06-05T12:14:45.583027Z","shell.execute_reply":"2023-06-05T12:14:45.592289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models = {}\nbest_threshold = 0.63","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:16:02.051303Z","iopub.execute_input":"2023-06-05T12:16:02.051775Z","iopub.status.idle":"2023-06-05T12:16:02.057073Z","shell.execute_reply.started":"2023-06-05T12:16:02.051735Z","shell.execute_reply":"2023-06-05T12:16:02.055922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_kol_f = {\n    1:140,3:110,\n    4:110, 5:220, 6:120, 7:110, 8:110, 9:100, 10:120, 11:120,\n    14: 110, 15:160, 16:105, 17:140             \n             }","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:16:09.113655Z","iopub.execute_input":"2023-06-05T12:16:09.114054Z","iopub.status.idle":"2023-06-05T12:16:09.120679Z","shell.execute_reply.started":"2023-06-05T12:16:09.114020Z","shell.execute_reply":"2023-06-05T12:16:09.119572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"'/kaggle/input/featur/train_0_4t.csv' ファイルからデータを読み込み、df0_4 にデータフレームとして格納しています。dtypes を使用してデータ型を指定しています。\n\ndf0_4 データフレームを 'session_id' でグループ化し、各グループの 'level' のユニークな値の数が5未満かどうかを判定し、結果を kol_lvl に代入しています。\n\nkol_lvl の値が True であるインデックスを抽出し、list_session に格納しています。\n\ndf0_4 データフレームから list_session に含まれる 'session_id' を持つ行を削除しています。\n\ndelt_time_def 関数を使用して df0_4 データフレームに対して時間の差分を計算し、新しい列を追加しています。\n\nquests_0_4 リストに [1, 3] を代入しています。\n\ncreate_model 関数を使用して df0_4 データフレームに対してモデルを作成し、結果を models に代入しています。\n\ndf0_4 を削除してメモリを解放しています。\n\nこのコードは、データの読み込みから特徴量の前処理、モデルの作成までの一連の処理を示しています。","metadata":{}},{"cell_type":"code","source":"df0_4 = pd.read_csv('/kaggle/input/featur/train_0_4t.csv', dtype=dtypes) \nkol_lvl = (df0_4 .groupby(['session_id'])['level'].agg('nunique') < 5)\nlist_session = kol_lvl[kol_lvl].index\ndf0_4  = df0_4 [~df0_4 ['session_id'].isin(list_session)]\ndf0_4 = delt_time_def(df0_4)\n\nquests_0_4 = [1, 3] \n# list_kol_f = {1:140,3:110}\n\nmodels = create_model(df0_4, quests_0_4, models, list_kol_f)\ndel df0_4","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:16:16.395752Z","iopub.execute_input":"2023-06-05T12:16:16.396160Z","iopub.status.idle":"2023-06-05T12:17:42.120964Z","shell.execute_reply.started":"2023-06-05T12:16:16.396125Z","shell.execute_reply":"2023-06-05T12:17:42.120011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df5_12 = pd.read_csv('/kaggle/input/featur/train_5_12t.csv', dtype=dtypes)\nkol_lvl = (df5_12.groupby(['session_id'])['level'].agg('nunique') < 8)\nlist_session = kol_lvl[kol_lvl].index\ndf5_12 = df5_12[~df5_12['session_id'].isin(list_session)]\ndf5_12 = delt_time_def(df5_12)\nquests_5_12 = [4, 5, 6, 7, 8, 9, 10, 11] \n\n# list_kol_f = {4:110, 5:220, 6:120, 7:110, 8:110, 9:100, 10:120, 11:120}\n\nmodels = create_model(df5_12, quests_5_12, models, list_kol_f)\ndel df5_12","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:18:06.683024Z","iopub.execute_input":"2023-06-05T12:18:06.683479Z","iopub.status.idle":"2023-06-05T12:25:27.315636Z","shell.execute_reply.started":"2023-06-05T12:18:06.683441Z","shell.execute_reply":"2023-06-05T12:25:27.314474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df13_22 = pd.read_csv('/kaggle/input/featur/train_13_22t.csv', dtype=dtypes) \nkol_lvl = (df13_22 .groupby(['session_id'])['level'].agg('nunique') < 10)\nlist_session = kol_lvl[kol_lvl].index\ndf13_22  = df13_22 [~df13_22 ['session_id'].isin(list_session)]\ndf13_22 = delt_time_def(df13_22)\n\nquests_13_22 = [14, 15, 16, 17] \n# list_kol_f = {14: 110, 15:160, 16:105, 17:140}\n\nmodels = create_model(df13_22, quests_13_22, models, list_kol_f)\ndel df13_22   #本notebookで追加","metadata":{"execution":{"iopub.status.busy":"2023-06-05T12:48:06.461471Z","iopub.execute_input":"2023-06-05T12:48:06.461882Z","iopub.status.idle":"2023-06-05T12:53:12.141774Z","shell.execute_reply.started":"2023-06-05T12:48:06.461847Z","shell.execute_reply":"2023-06-05T12:53:12.139805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Infer Test Data**","metadata":{}},{"cell_type":"markdown","source":"特定の変数や属性をリセットするための処理が行われています。jo_wilder.make_env.__called__、env.__called__、type(env)._state の値がリセットされます。例外が発生した場合は、処理をスキップします。\n\njo_wilder.make_env() を呼び出して env を作成します。make_env は jo_wilder モジュール内の関数で、環境を初期化するための処理が含まれています。\n\nenv の iter_test() メソッドを呼び出して、イテレーターを取得しています。これにより、テストデータに対して反復処理を行うことができます。\n\nこのコードは、jo_wilder モジュールを使用して環境を初期化し、テストデータに対して反復処理を行うための準備を行っています。","metadata":{}},{"cell_type":"code","source":"import jo_wilder\n\ntry:\n    jo_wilder.make_env.__called__ = False\n    env.__called__ = False\n    type(env)._state = type(type(env)._state).__dict__['INIT']\nexcept:\n    pass\n\nenv = jo_wilder.make_env()\niter_test = env.iter_test()    ","metadata":{"execution":{"iopub.status.busy":"2023-05-31T11:43:00.889708Z","iopub.execute_input":"2023-05-31T11:43:00.890324Z","iopub.status.idle":"2023-05-31T11:43:00.930872Z","shell.execute_reply.started":"2023-05-31T11:43:00.890285Z","shell.execute_reply":"2023-05-31T11:43:00.929385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time","metadata":{"execution":{"iopub.status.busy":"2023-05-31T11:43:00.932742Z","iopub.execute_input":"2023-05-31T11:43:00.933132Z","iopub.status.idle":"2023-05-31T11:43:00.938533Z","shell.execute_reply.started":"2023-05-31T11:43:00.933093Z","shell.execute_reply":"2023-05-31T11:43:00.937034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"変数 g_end4 と g_end5 を初期化しています。\n\nキーと値のペアを持つ辞書 list_q を作成しています。キーは文字列であり、値は quests_0_4、quests_5_12、quests_13_22 のいずれかのリストです。\n\niter_test を使ってテストデータに対して反復処理を行います。iter_test はテストデータのイテレーターであり、各反復では (test, sam_sub) のタプルが取得されます。\n\nsam_sub の 'session_id' 列の値を処理して 'question' 列を作成します。'session_id' の値から数字を抽出し、'question' 列に格納します。\n\ntest データフレームの 'level_group' 列の最初の値を grp に代入します。\n\nsam_sub データフレームの 'correct' 列を 1 で初期化し、特定の質問番号（5、8、10、13、15）に対応する行の 'correct' 列の値を 0 に設定します。\n\ntest データフレームから 'level_group' 列が grp と等しい行を選択し、その部分データフレームを delt_time_def 関数に渡して old_train を作成します。\n\ngrp に対応するキーを使って list_q 辞書から対応するリストを取得し、q として反復処理します。\n\n特徴エンジニアリングの関数を使用して old_train から new_train を作成します。feature_engineer と feature_quest_otvet の呼び出しによって特徴量が追加され、計算時間 end4 が測定されます。\n\nモデルを使用して new_train の予測を行います。モデルは models 辞書から取得され、predict_proba メソッドを使って確率の予測結果を取得します。計算時間 end5 も測定され、g_end5 に加算されます。\n\n予測結果をもとに 'correct' 列の値を更新します。sam_sub.question == q の条件に一致する行の 'correct' 列の値を、p[0] と best_threshold の比較結果に基づいて更新します。\n\nsam_sub データフレームから 'session_id' 列と 'correct' 列を選択し、それを使って env.predict() を呼び出します。これにより、予測結果が提出用に環境に送信されます。\n\nこのコードは、テストデータに対して予測を行い、その結果を提出用に環境に送信するための処理を実行しています。計算時間も測定され、g_end4 と g_end5 に加算されます。","metadata":{}},{"cell_type":"code","source":"g_end4 = 0\ng_end5 = 0\n\nlist_q = {'0-4':quests_0_4, '5-12':quests_5_12, '13-22':quests_13_22}\nfor (test, sam_sub) in iter_test:\n    sam_sub['question'] = [int(label.split('_')[1][1:]) for label in sam_sub['session_id']]    \n    grp = test.level_group.values[0]   \n    sam_sub['correct'] = 1\n    sam_sub.loc[sam_sub.question.isin([5, 8, 10, 13, 15]), 'correct'] = 0  \n    old_train = delt_time_def(test[test.level_group == grp])\n       \n    for q in list_q[grp]:\n        \n        start4 = time.time()\n        new_train = feature_engineer(old_train, list_kol_f[q])\n        new_train = feature_quest_otvet(new_train, old_train, q, list_kol_f[q])\n#         new_train = feature_quest(new_train, old_train, q, kol_f)\n        \n        end4 = time.time() - start4\n        g_end4 += end4\n        \n        start5 = time.time()        \n        \n        clf = models[f'{q}']\n        p = clf.predict_proba(new_train.astype('float32'))[:,1]        \n        \n        end5 = time.time() - start5\n        g_end5 += end5\n             \n        \n        mask = sam_sub.question == q \n        x = int(p[0]>best_threshold)\n        sam_sub.loc[mask,'correct'] = x      \n        \n        \n    sam_sub = sam_sub[['session_id', 'correct']]      \n    env.predict(sam_sub)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T11:43:00.940337Z","iopub.execute_input":"2023-05-31T11:43:00.940773Z","iopub.status.idle":"2023-05-31T11:43:12.790753Z","shell.execute_reply.started":"2023-05-31T11:43:00.940715Z","shell.execute_reply":"2023-05-31T11:43:12.78929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA submission.csv","metadata":{"papermill":{"duration":0.011427,"end_time":"2023-02-07T01:02:45.502331","exception":false,"start_time":"2023-02-07T01:02:45.490904","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# df = pd.read_csv('submission.csv')\n# print( df.shape )\n# df.head(60)","metadata":{"papermill":{"duration":0.027432,"end_time":"2023-02-07T01:02:45.541022","exception":false,"start_time":"2023-02-07T01:02:45.51359","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-31T11:43:12.792736Z","iopub.execute_input":"2023-05-31T11:43:12.793095Z","iopub.status.idle":"2023-05-31T11:43:12.797888Z","shell.execute_reply.started":"2023-05-31T11:43:12.793059Z","shell.execute_reply":"2023-05-31T11:43:12.79671Z"},"trusted":true},"execution_count":null,"outputs":[]}]}