{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Context\nI had issues with scoring run time due to having too many models, so I wanted a way to mimic the Kaggle API and troubleshoot my pipeline.\n\n### Things to keep in mind\nA few things I've noticed:\n- Kaggle API fetches one `level_group` of one `session_id` per batch\n- The `level_group` in scoring was originally fetched in the wrong order: `13-22` to `5-12` to `0-4`. Thanks to [Chris for pointing that out](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479), Kaggle team has fixed/will fix it to the correct order: `0-4` --> `5-12` --> `13-22`.\n    - As of the time of this writing (2023-02-18), the fix seems to be live.\n<br>\n\n### Option\nTherefore, in this notebook, I've created this dataset which (1) contains 5% of the real train data and (2) splits each `level_group` of each user into a single `.csv`. To troubleshoot a submitted notebook, we will load these `.csv` into a `data_cache` (a `list` for example), then replace the\n```\nfor (sample_submission, test) in iter_test:\n\n    <<<your pipeline>>>\n    \n    env.predict(sample_submission)\n```\nin your original notebook with\n```\nfor (sample_submission, test) in [sample_submission_cache, data_cache]:\n\n    <<<your pipeline>>>\n    \n    # env.predict(sample_submission) --> remove/comment out\n```\nThen, you can time how long each components in `<<<your pipeline>>>` runs and identify the troublemaker. This is how I found out my issue was that I had too many models (model-loading & prediction-concatenating) and not feature engineering as I originally thought.\n<br>\n<br>\nA few notes\n- You'll need to run the notebook without internet and GPU\n- This method doesn't account for the API's run time, so it can serve as an imperfect estimate for scoring run time. To get the scoring time estimate, multiply the troubleshooting notebook's runtime by 20. The result is the **least** amount of time you'll need when submitting a similar pipeline for scoring.","metadata":{}},{"cell_type":"code","source":"import gc, os\nfrom datetime import datetime as dt\nimport numpy as np\nimport pandas as pd\nimport cudf","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-18T08:46:16.493951Z","iopub.execute_input":"2023-02-18T08:46:16.495839Z","iopub.status.idle":"2023-02-18T08:46:20.092921Z","shell.execute_reply.started":"2023-02-18T08:46:16.495761Z","shell.execute_reply":"2023-02-18T08:46:20.091850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Parameters","metadata":{}},{"cell_type":"code","source":"# General parameters\nRS = 719\nSUBSET_FRAC = 5e-2 # Fraction of session_id to be used\n\n# URL\nORIG_PTH = '/kaggle/input/predict-student-performance-from-game-play/'\n\nGRP_ORDER = ['0-4', '5-12', '13-22']\n# GRP_ORDER = ['13-22', '5-12', '0-4']\nNEXT_GRP = {\n    GRP_ORDER[0]: GRP_ORDER[1]\n    , GRP_ORDER[1]: GRP_ORDER[2]\n    , GRP_ORDER[2]: 'Mangekyo Sharingan'\n}","metadata":{"execution":{"iopub.status.busy":"2023-02-18T08:46:20.098142Z","iopub.execute_input":"2023-02-18T08:46:20.098568Z","iopub.status.idle":"2023-02-18T08:46:20.108887Z","shell.execute_reply.started":"2023-02-18T08:46:20.098521Z","shell.execute_reply":"2023-02-18T08:46:20.106201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Load","metadata":{}},{"cell_type":"code","source":"def load_csv(path, dtype=None, usecols=None):\n    return cudf.read_csv(path, dtype=dtype, usecols=usecols)\n\ndef timer(start):\n    return (dt.now() - start).seconds","metadata":{"execution":{"iopub.status.busy":"2023-02-18T08:46:20.111533Z","iopub.execute_input":"2023-02-18T08:46:20.111913Z","iopub.status.idle":"2023-02-18T08:46:20.121113Z","shell.execute_reply.started":"2023-02-18T08:46:20.111877Z","shell.execute_reply":"2023-02-18T08:46:20.120028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrain_df = load_csv(ORIG_PTH + 'train.csv')\nssids = train_df.session_id.unique().to_arrow().to_pylist()\nssids_subset = np.random.choice(ssids, size=int(len(ssids)*SUBSET_FRAC), replace=False)\nssids_ss_len = len(ssids_subset)\nprint(f'Total: {len(ssids)} sessions')\nprint(f'Writing: {ssids_ss_len} sessions ({SUBSET_FRAC} of total)')","metadata":{"execution":{"iopub.status.busy":"2023-02-18T08:46:20.123331Z","iopub.execute_input":"2023-02-18T08:46:20.124231Z","iopub.status.idle":"2023-02-18T08:47:03.410403Z","shell.execute_reply.started":"2023-02-18T08:46:20.124194Z","shell.execute_reply":"2023-02-18T08:47:03.409260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create CSVs","metadata":{}},{"cell_type":"code","source":"%%time\n\nfor grp in GRP_ORDER:\n    counter = 0\n    print(f'Generating files for group {grp}')\n    if not os.path.isdir(grp): os.mkdir(grp)\n    st = dt.now()\n    for ssid in ssids_subset:\n        tmp = train_df.loc[(train_df.session_id == ssid) & (train_df.level_group == grp)]\n        tmp = tmp.reset_index(drop=True)\n        if len(tmp) > 0: tmp.to_pandas().to_csv(f'{grp}/{ssid}.csv', index=False)\n        del tmp\n        _ = gc.collect()\n        counter += 1\n        if counter == ssids_ss_len:\n            print(f'{counter} ({timer(st)}s)')\n        elif (counter%50)==0:\n            print(counter, end=', ')","metadata":{"execution":{"iopub.status.busy":"2023-02-18T08:47:03.412508Z","iopub.execute_input":"2023-02-18T08:47:03.412873Z","iopub.status.idle":"2023-02-18T08:47:03.471887Z","shell.execute_reply.started":"2023-02-18T08:47:03.412834Z","shell.execute_reply":"2023-02-18T08:47:03.470838Z"},"trusted":true},"execution_count":null,"outputs":[]}]}