{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"Hi everybody 🖐🖐\nThe preliminary work for this notebook is located here [Cleaning Riiid! dataset & add part of TOEIC test](https://www.kaggle.com/adrian1903/cleaning-riiid-dataset-add-part-of-toeic-test)"},{"metadata":{},"cell_type":"markdown","source":"I am considering a cross validation ✅ on this project, with ⏲ TimeSeriesSplit(). \n\nWe have a time series which is timestamp. We can consider a second variable, which can be considered as a time series, parts of TOEIC test. We can isolate every block. \nAt least for the written (5, 6, 7) and oral parts (1, 2, 3, 4).\n\nBelow my approach :"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nplt.style.use(\"default\")\nfrom sklearn.model_selection import TimeSeriesSplit","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\npath = \"../input/cleaning-riiid-dataset-add-part-of-toeic-test/reduced_riiid_train.pkl.gzip\"\ntrain = pd.read_pickle(path)\ntrain = train.astype({'row_id': 'int64',\n                      'timestamp': 'int64',\n                      'user_id': 'int32',\n                      'content_id': 'int16',\n                      'content_type_id': 'int8',\n                      'task_container_id': 'int16',\n                      'user_answer': 'int8',\n                      'answered_correctly': 'int8',\n                      'prior_question_elapsed_time': 'float32',\n                      'prior_question_had_explanation': 'boolean'})\ntrain.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train.sort_values(by=['part', 'timestamp'])\n\ncol = ['part', 'timestamp', 'row_id', 'user_id', 'content_id', 'content_type_id',\n       'task_container_id', 'user_answer', 'answered_correctly',\n       'prior_question_elapsed_time', 'prior_question_had_explanation',\n       'question_id']\ntrain[col].set_index('row_id')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I want to split like this below ⬇️⬇️"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Initialisation du TSCV\nn_iter = 5\ntscv = TimeSeriesSplit(n_splits=n_iter)\n\n# Echantillonnage\ntrain_lst = []\nval_lst = []\n\nfor train_index, val_index in tscv.split(train):\n    train_lst.append(train_index[-1])\n    val_lst.append(val_index[-1] - val_index[0])\n    \n    # X_train, X_test = X.iloc[train_index], X.iloc[val_index]\n    # y_train, y_test = y.iloc[train_index], y.iloc[val_index]\n\ntrain_tpl = tuple(train_lst)\nval_tpl = tuple(val_lst)\n\n# Tracage de l'échantillon\nind = np.arange(n_iter) + 1\nwidth = 0.35\n\nbtrain = plt.barh(ind, train_tpl, width)\nbval = plt.barh(ind, val_tpl, width, left=train_tpl)\n\nplt.yticks(ind)\nplt.gca().invert_yaxis()\nplt.legend((btrain, bval), (\"train\", \"val\"))\nplt.title('Sample Cross Validation')\nplt.ylabel('Interaction CV')\nplt.xlabel('Index')\nplt.savefig('./img_cross_validation.png', transparent=True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In my opinion 💭 it is necessary to treat the parts individually in CV, because :\n- there are several levels of difficulty,\n- the sizes of the parts are heterogeneous,\n- these are exercises to listen to and read.\n\nAnd you ❓ What do you think ❓"},{"metadata":{},"cell_type":"markdown","source":"If you made it to the point thank you for reading and 🆙 vote if it helped you. I will read your comments with pleasure.\n\nThanks"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}