{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hey!\n\nA couple of days into the competition I decided to update my EDA with an answer to a question I have been asked on a couple of occasions -- **how do we split our data into train and validation?**\n\nI am reframing this EDA to now consist of two parts. In the first part I will answer the question above.\n\nThe second part will be the original deep dive into the data I shared at the start of the competiton.\n\n## Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n\nLet's get started! 🚀","metadata":{}},{"cell_type":"markdown","source":"# How to split the data for training","metadata":{}},{"cell_type":"markdown","source":"This is likely to be one of the most important questions in this competition.\n\nOur `train set` consists of 4 weeks of unmodified session data. The `test set` however contains sessions who are self-contained. They don't start before the first day of the test set! This is a significant difference that is not accounted for in the split provided to us by the organizers.\n\nAnyhow, there are many ways to aproach splitting the data for training. Here, I will share with you a setup for training a two-stage recommender that is *probably* okay. It is theoretically sound but that doesn't guarantee it will give you the best results.\n\nTherefore, I would also suggest you experiment with other approaches, as there is a chance they might work better!","metadata":{}},{"cell_type":"markdown","source":"## How to think about splitting data","metadata":{}},{"cell_type":"markdown","source":"This [video](https://www.youtube.com/watch?v=Q0QmziFcfU0) is by far the best resource on training with cross-validation I have ever come across.\n\nIn this competition I will not advocate training with cross-validatin per se, however that video discusses a very important aspect of splitting data, that is avoiding leakage.\n\nThe idea of leakage is fundamental to thinking about the splits in this competition.\n\nIn short, we don't want to make the life of our model artificially easier in training than it is going to be at inference! Our model might train more poorly due to us giving it unfair advantage (leagage) in train.\n\nHere is a leak-free scheme for training a two-stage recommender.","metadata":{}},{"cell_type":"markdown","source":"## Leak free training of a two-stage recommender","metadata":{}},{"cell_type":"markdown","source":"Here is my suggestion: use [the script from organizer's repo](https://github.com/otto-de/recsys-dataset) to split the train data of this competition (four weeks of train) into chunks of 3 weeks and the last week.\n\n**If you split the data like this, you will get the following:**\n\n* `train1` - 3 weeks of train\n* `train2` - last week of original train THAT will be just like the test set (this is important!)\n\nUsing this setup, we can:\n\n* Construct co-visitation matrices on `train1` and train data from `train2` (not the `test_labels`).\n* Train a ranking model on `train2`\n* Predict on competition test\n\nBut how do we train with validation? There is no validation set here!\n\nThe solution is simple.\n\nKeep going back in weeks using the script from organizer's repo. Run it again on our current `train1`! It will get split into two weeks of original train + the last week of `train1`. You can train your ranker on that week and validation on `train2` from above!","metadata":{}},{"cell_type":"markdown","source":"## Is this the best approach though?","metadata":{}},{"cell_type":"markdown","source":"The above approach has one problem -- some data will be discarded that we could otherwise train on. Also, as the Kaggle GM explains in the video I linked to, not all leaks are bad! They are bad from a theoretical stand point, but might still lead to improved results.\n\nFor instance, one question that comes to mind is this: 'would the benefit of constructing a covisiation matrix on more data not outweight the cost incurred due to leakage of information from the future?'\n\nNo one has an answer to this apriori. This is just one of the many considerations here that would require testing.\n\n**BONUS:** Instead of fumbling with the script from organizer's repo to create `train1` and `train2` from above, I uploaded a dataset obtained via this technique (along with the minimization of the memory footprint) [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation).","metadata":{}},{"cell_type":"markdown","source":"\n**If you are finding this notebook useful, please upvote! 🙏 Thank you, appreciate your support!**","metadata":{}},{"cell_type":"markdown","source":"# Deep dive into competition data","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom matplotlib import pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-03T11:58:17.382132Z","iopub.execute_input":"2022-11-03T11:58:17.382696Z","iopub.status.idle":"2022-11-03T11:58:17.411380Z","shell.execute_reply.started":"2022-11-03T11:58:17.382601Z","shell.execute_reply":"2022-11-03T11:58:17.410437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us read in the `train` and `test` datasets.","metadata":{}},{"cell_type":"code","source":"train = pd.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\ntest = pd.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:17.413243Z","iopub.execute_input":"2022-11-03T11:58:17.413916Z","iopub.status.idle":"2022-11-03T11:58:39.654805Z","shell.execute_reply.started":"2022-11-03T11:58:17.413879Z","shell.execute_reply":"2022-11-03T11:58:39.653845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us also read in the pickle file that will allow us to decipher the `type` information that has been encoded as integers to conserve memory.","metadata":{}},{"cell_type":"code","source":"!pip install pickle5\n\nimport pickle5 as pickle\n\nwith open('../input/otto-full-optimized-memory-footprint/id2type.pkl', \"rb\") as fh:\n    id2type = pickle.load(fh)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:39.657734Z","iopub.execute_input":"2022-11-03T11:58:39.658561Z","iopub.status.idle":"2022-11-03T11:58:51.153440Z","shell.execute_reply.started":"2022-11-03T11:58:39.658511Z","shell.execute_reply":"2022-11-03T11:58:51.152204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape, test.shape","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:51.156823Z","iopub.execute_input":"2022-11-03T11:58:51.157232Z","iopub.status.idle":"2022-11-03T11:58:51.168640Z","shell.execute_reply.started":"2022-11-03T11:58:51.157193Z","shell.execute_reply":"2022-11-03T11:58:51.167393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `train` dataset contains 216716096 datapoints with `test` containing only 6928123.\n\nProportion of `test` to `train`:","metadata":{}},{"cell_type":"code","source":"test.shape[0]/train.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:51.170260Z","iopub.execute_input":"2022-11-03T11:58:51.171282Z","iopub.status.idle":"2022-11-03T11:58:51.181636Z","shell.execute_reply.started":"2022-11-03T11:58:51.171239Z","shell.execute_reply":"2022-11-03T11:58:51.180524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The size of the test set is ~3.1% of the train set. This can give us an idea of how lightweight the inference is likely to be compared to training.","metadata":{}},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:51.183134Z","iopub.execute_input":"2022-11-03T11:58:51.183521Z","iopub.status.idle":"2022-11-03T11:58:51.203872Z","shell.execute_reply.started":"2022-11-03T11:58:51.183478Z","shell.execute_reply":"2022-11-03T11:58:51.202988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many session are there in `train` and `test`?","metadata":{}},{"cell_type":"code","source":"train.session.unique().shape[0], test.session.unique().shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:51.204890Z","iopub.execute_input":"2022-11-03T11:58:51.205219Z","iopub.status.idle":"2022-11-03T11:58:53.928207Z","shell.execute_reply.started":"2022-11-03T11:58:51.205189Z","shell.execute_reply":"2022-11-03T11:58:53.927043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.session.unique().shape[0]/train.session.unique().shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:53.929573Z","iopub.execute_input":"2022-11-03T11:58:53.929942Z","iopub.status.idle":"2022-11-03T11:58:56.594351Z","shell.execute_reply.started":"2022-11-03T11:58:53.929909Z","shell.execute_reply":"2022-11-03T11:58:56.593025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Seems the sessions in the test set are much shorter! Let's confirm this.","metadata":{}},{"cell_type":"code","source":"test.groupby('session')['aid'].count().apply(np.log1p).hist()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:56.595981Z","iopub.execute_input":"2022-11-03T11:58:56.596991Z","iopub.status.idle":"2022-11-03T11:58:57.317989Z","shell.execute_reply.started":"2022-11-03T11:58:56.596953Z","shell.execute_reply":"2022-11-03T11:58:57.316853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.groupby('session')['aid'].count().apply(np.log1p).hist()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:58:57.319248Z","iopub.execute_input":"2022-11-03T11:58:57.319603Z","iopub.status.idle":"2022-11-03T11:59:05.645550Z","shell.execute_reply.started":"2022-11-03T11:58:57.319569Z","shell.execute_reply":"2022-11-03T11:59:05.644408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There might be something at play here. Could the organizers have thrown us a curve ball and the train and test data are not from the same distribution?\n\nLet's quickly look at timestamps.","metadata":{}},{"cell_type":"code","source":"train.ts","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:59:05.646847Z","iopub.execute_input":"2022-11-03T11:59:05.647179Z","iopub.status.idle":"2022-11-03T11:59:05.656789Z","shell.execute_reply.started":"2022-11-03T11:59:05.647150Z","shell.execute_reply":"2022-11-03T11:59:05.655553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import datetime\n\ndatetime.datetime.fromtimestamp(train.ts.min()/1000), datetime.datetime.fromtimestamp(train.ts.max()/1000)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:59:05.658133Z","iopub.execute_input":"2022-11-03T11:59:05.658578Z","iopub.status.idle":"2022-11-03T11:59:06.137617Z","shell.execute_reply.started":"2022-11-03T11:59:05.658536Z","shell.execute_reply":"2022-11-03T11:59:06.136313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import datetime\n\ndatetime.datetime.fromtimestamp(test.ts.min()/1000), datetime.datetime.fromtimestamp(test.ts.max()/1000)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:59:06.141955Z","iopub.execute_input":"2022-11-03T11:59:06.142342Z","iopub.status.idle":"2022-11-03T11:59:06.166649Z","shell.execute_reply.started":"2022-11-03T11:59:06.142306Z","shell.execute_reply":"2022-11-03T11:59:06.165388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like we have temporally split data. The problem is that the data doesn't come from the same period. In most geographies the beginning of September is the start of the school year!\n\nThat is the period where people are coming back from vacation, commerce resumes after a slowdown during the vacation season.\n\nThe organizers are not making this easy for us 🙂\n\nLet's see if there are any new items in the test set that were not see in train.","metadata":{}},{"cell_type":"code","source":"len(set(test.aid.tolist()) - set(train.aid.tolist()))","metadata":{"execution":{"iopub.status.busy":"2022-11-03T11:59:59.027205Z","iopub.execute_input":"2022-11-03T11:59:59.027667Z","iopub.status.idle":"2022-11-03T12:00:41.634513Z","shell.execute_reply.started":"2022-11-03T11:59:59.027629Z","shell.execute_reply":"2022-11-03T12:00:41.633315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So at least we have that going for us, no new items in the test set! 😊\n\nI just scanned the forums really quickly and seems we have an answer as to why the session length differs between `train` and `test`!\n\nFirst, let's look at the data more closely.","metadata":{}},{"cell_type":"code","source":"train.groupby('session')['aid'].count().describe()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:03:03.565408Z","iopub.execute_input":"2022-11-03T12:03:03.565820Z","iopub.status.idle":"2022-11-03T12:03:11.561797Z","shell.execute_reply.started":"2022-11-03T12:03:03.565788Z","shell.execute_reply":"2022-11-03T12:03:11.560356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.groupby('session')['aid'].count().describe()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:02:54.352163Z","iopub.execute_input":"2022-11-03T12:02:54.353447Z","iopub.status.idle":"2022-11-03T12:02:54.816036Z","shell.execute_reply.started":"2022-11-03T12:02:54.353397Z","shell.execute_reply":"2022-11-03T12:02:54.814802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And here is the [key piece of information on the forums](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486).\n\nApparently, a session are all actions by a user in the tracking period. So naturally, if the tracking period is shorter, the sessions will also be shorter.\n\nMaybe there is nothing amiss happening here.","metadata":{}},{"cell_type":"code","source":"train.session.max(), test.session.min()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:06:35.164544Z","iopub.execute_input":"2022-11-03T12:06:35.164987Z","iopub.status.idle":"2022-11-03T12:06:35.412944Z","shell.execute_reply.started":"2022-11-03T12:06:35.164947Z","shell.execute_reply":"2022-11-03T12:06:35.411541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"An we see that the `session_ids` are not overlapping between `train` and `test` so it will be impossible to map the users (even if we have seen them before in train). We have to assume each session is from a different user.\n\nNow that I know a bit more about the dataset, I can't wait to start playing around with it. This is shaping up to be a very interesting problem! 😊\n\n**If you found the notebook useful, please upvote it! 🙏 Thank you**","metadata":{}}]}