{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Test Data Error Check - Multiple Games\n\nI made some notebooks to check if the data errors present in the training data are also present in the test data. I am making public the codes for the elapsed time column check and mutilple games present check. You can further modify these notebooks to check for other errors also. \n\nIf this notebook fails when running or during submission, then that means the problem of multiple games being present in a single session is present in the test data also. This will mean that the kaggle api provides a leak into the future during testing also.\n\nTo get a better understanding of how to check for multiple games, you can see [this][1] notebook I created.\n\n[1]: https://www.kaggle.com/code/shashwatraman/dealing-with-multiple-games-in-the-data\n\n**Update**: This notebook fails during submission. This means Multiple Games within a single session are also present in the new test data!! (Kaggle API does provide a leak into the future in the new test data also)\n\n#### If you find this notebook helpful, please show support!!","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport gc\nimport pickle\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score\nfrom tqdm.notebook import tqdm\nfrom collections import defaultdict\nimport warnings\nfrom itertools import combinations\n\nwarnings.filterwarnings('ignore')\npd.set_option(\"display.max_columns\", None)\npd.set_option(\"display.max_rows\", 200)","metadata":{"execution":{"iopub.status.busy":"2023-03-18T08:54:49.757026Z","iopub.execute_input":"2023-03-18T08:54:49.757440Z","iopub.status.idle":"2023-03-18T08:54:51.200346Z","shell.execute_reply.started":"2023-03-18T08:54:49.757407Z","shell.execute_reply":"2023-03-18T08:54:51.199254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_engineer(x, grp):\n    \n    #session num_events\n    df_final = x.groupby('session_id')['index'].agg('count')\n    df_final.name = 'num_events'\n    df_final = df_final.reset_index()\n    df_final = df_final.set_index('session_id')\n    \n    return df_final","metadata":{"execution":{"iopub.status.busy":"2023-03-18T08:54:51.202109Z","iopub.execute_input":"2023-03-18T08:54:51.202786Z","iopub.status.idle":"2023-03-18T08:54:51.207225Z","shell.execute_reply.started":"2023-03-18T08:54:51.202740Z","shell.execute_reply":"2023-03-18T08:54:51.206517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f_read = open('/kaggle/input/try-train/importance_dict.pkl', 'rb')\nimportance_dict = pickle.load(f_read)\nf_read.close()","metadata":{"execution":{"iopub.status.busy":"2023-03-18T08:54:51.208445Z","iopub.execute_input":"2023-03-18T08:54:51.208941Z","iopub.status.idle":"2023-03-18T08:54:51.225764Z","shell.execute_reply.started":"2023-03-18T08:54:51.208911Z","shell.execute_reply":"2023-03-18T08:54:51.224770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Loading the models\nQUESTION_MODELS = []\nfor t in range(1,19):\n    clf = XGBClassifier()\n    clf.load_model(f'../input/try-train/XGB_question{t}.xgb')\n    QUESTION_MODELS.append( clf )","metadata":{"execution":{"iopub.status.busy":"2023-03-18T08:54:51.226626Z","iopub.execute_input":"2023-03-18T08:54:51.226927Z","iopub.status.idle":"2023-03-18T08:54:51.387138Z","shell.execute_reply.started":"2023-03-18T08:54:51.226900Z","shell.execute_reply":"2023-03-18T08:54:51.386191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Infer Test Data","metadata":{}},{"cell_type":"code","source":"# IMPORT KAGGLE API\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-03-18T08:54:51.390376Z","iopub.execute_input":"2023-03-18T08:54:51.392043Z","iopub.status.idle":"2023-03-18T08:54:51.416059Z","shell.execute_reply.started":"2023-03-18T08:54:51.391978Z","shell.execute_reply":"2023-03-18T08:54:51.415034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"limits = {'0-4':(1,4), '5-12':(4,14),'13-22':(14,19)}\nbest_threshold = 0.625\n\nhistorical_meta = defaultdict(list)\n\nfor (test, sample_submission) in iter_test:\n    \n    grp = test.level_group.values[0]\n    session_id = test.session_id.values[0]\n    \n    #Make level_diff column\n    test['cons'] = 1\n    test['index1'] = test.groupby('session_id')['cons'].agg('cumsum') - 1\n    test.drop(['cons'], axis=1, inplace=True)\n    \n    test['level_diff'] = test.level.diff()\n    test['level_diff'] = test['level_diff'].replace(0, np.nan)\n    test.loc[test['index1']==0,'level_diff'] = 0\n    test['level_diff'] = test['level_diff'].ffill()\n    test['level_diff'] = test['level_diff'].fillna(0)\n    \n    #Check if multiple games are present in the test data also\n    if test[test['level_diff']<0]['level_diff'].sum() !=0:\n        print()\n        print('Test Failed!!')\n        break\n    \n    df = feature_engineer(test, grp)\n    \n    a,b = limits[grp]\n    for t in range(a,b):\n        FEATURES = importance_dict[str(t)]\n        \n        clf = QUESTION_MODELS[t-1]\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int(p.item()>best_threshold)\n    \n    env.predict(sample_submission)","metadata":{"execution":{"iopub.status.busy":"2023-03-18T08:54:51.421681Z","iopub.execute_input":"2023-03-18T08:54:51.424008Z","iopub.status.idle":"2023-03-18T08:54:51.954208Z","shell.execute_reply.started":"2023-03-18T08:54:51.423954Z","shell.execute_reply":"2023-03-18T08:54:51.952700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### The test does not fail in the infer notebook!!\n\n### Now we submit and wait!!!","metadata":{}}]}