{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook builds on [the work](https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix) by [@vslaykovsky](https://www.kaggle.com/vslaykovsky). As suggested by [@cdeotte](https://www.kaggle.com/cdeotte) [here](https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost) I am including test information in clculating the co-visitation matrix.\n\nHere I provide a simplified implementation that is easier to follow and that can also be more easily extended!\n\nPlease find some additional discussion on this notebook in this [thread](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364210).\n\nBecause the code is simpler, it was straightforward to modify the logic and make a couple of different decisions to achieve a better result.\n\nAdditionally, I am also using a version of the dataset that I shared [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843). This further simplifies matters as we no longer have to read the data from `jasonl` files!\n\n## Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n\n\n### Update:\n\n* added sorting of `AIDs` in test based on number of clicks from [Multiple clicks vs latest items](https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items) by [pietromaldini1](https://www.kaggle.com/pietromaldini1)\n* added speed improvements from [Fast Co-Visitation Matrix](https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items) by [dpalbrecht](https://www.kaggle.com/dpalbrecht)\n* added weigting by type of `AIDs` in test from [Item type vs multiple clicks vs latest items](https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items) by [ingvarasgalinskas](https://www.kaggle.com/ingvarasgalinskas)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"fraction_of_sessions_to_use = 1\n\nimport pandas as pd\nimport numpy as np\n\ntrain = pd.read_parquet('../input/otto-full-optimized-memory-footprint//train.parquet')\ntest = pd.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')\n\nsample_sub = pd.read_csv('../input/otto-recommender-system/sample_submission.csv')\n\nif fraction_of_sessions_to_use != 1:\n    lucky_sessions_train = train.drop_duplicates(['session']).sample(frac=fraction_of_sessions_to_use, random_state=42)['session']\n    subset_of_train = train[train.session.isin(lucky_sessions_train)]\n    \n    lucky_sessions_test = test.drop_duplicates(['session']).sample(frac=fraction_of_sessions_to_use, random_state=42)['session']\n    subset_of_test = test[test.session.isin(lucky_sessions_test)]\nelse:\n    subset_of_train = train\n    subset_of_test = test\n\nsubset_of_train.index = pd.MultiIndex.from_frame(subset_of_train[['session']])\nsubset_of_test.index = pd.MultiIndex.from_frame(subset_of_test[['session']])\n\nchunk_size = 30_000\nmin_ts = train.ts.min()\nmax_ts = test.ts.max()\n\nfrom collections import defaultdict, Counter\nnext_AIDs = defaultdict(Counter)\n\nsubsets = pd.concat([subset_of_train, subset_of_test])\nsessions = subsets.session.unique()","metadata":{"execution":{"iopub.status.busy":"2022-11-09T09:11:33.947472Z","iopub.execute_input":"2022-11-09T09:11:33.947949Z","iopub.status.idle":"2022-11-09T09:12:06.609530Z","shell.execute_reply.started":"2022-11-09T09:11:33.947915Z","shell.execute_reply":"2022-11-09T09:12:06.608379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subsets","metadata":{"execution":{"iopub.status.busy":"2022-11-09T09:14:35.297253Z","iopub.execute_input":"2022-11-09T09:14:35.297644Z","iopub.status.idle":"2022-11-09T09:14:35.314248Z","shell.execute_reply.started":"2022-11-09T09:14:35.297613Z","shell.execute_reply":"2022-11-09T09:14:35.312973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(0, sessions.shape[0], chunk_size):\n    current_chunk = subsets.loc[sessions[i]:sessions[min(sessions.shape[0]-1, i+chunk_size-1)]].reset_index(drop=True)\n    current_chunk = current_chunk.groupby('session', as_index=False).nth(list(range(-30,0))).reset_index(drop=True)\n    consecutive_AIDs = current_chunk.merge(current_chunk, on='session')\n    consecutive_AIDs = consecutive_AIDs[consecutive_AIDs.aid_x != consecutive_AIDs.aid_y]\n    consecutive_AIDs['days_elapsed'] = (consecutive_AIDs.ts_y - consecutive_AIDs.ts_x) / (24 * 60 * 60)\n    consecutive_AIDs = consecutive_AIDs[(consecutive_AIDs.days_elapsed >= 0) & (consecutive_AIDs.days_elapsed <= 1)]\n    \n    for aid_x, aid_y in zip(consecutive_AIDs['aid_x'], consecutive_AIDs['aid_y']):\n        next_AIDs[aid_x][aid_y] += 1\n    \ndel train, subset_of_train, subsets\n\nsession_types = ['clicks', 'carts', 'orders']\ntest_session_AIDs = test.reset_index(drop=True).groupby('session')['aid'].apply(list)\ntest_session_types = test.reset_index(drop=True).groupby('session')['type'].apply(list)\n\nlabels = []\n\nno_data = 0\nno_data_all_aids = 0\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\nfor AIDs, types in zip(test_session_AIDs, test_session_types):\n    if len(AIDs) >= 20:\n        weights=np.logspace(0.1,1,len(AIDs),base=2, endpoint=True)-1\n        aids_temp=defaultdict(lambda: 0)\n        for aid,w,t in zip(AIDs,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers[t]\n            \n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n        labels.append(sorted_aids[:20])\n    else:\n        AIDs = list(dict.fromkeys(AIDs[::-1]))\n        AIDs_len_start = len(AIDs)\n        \n        candidates = []\n        for AID in AIDs:\n            if AID in next_AIDs: candidates += [aid for aid, count in next_AIDs[AID].most_common(20)]\n        AIDs += [AID for AID, cnt in Counter(candidates).most_common(40) if AID not in AIDs]\n        \n        labels.append(AIDs[:20])\n        if candidates == []: no_data += 1\n        if AIDs_len_start == len(AIDs): no_data_all_aids += 1\n\n# >>> outputting results to CSV\n\nlabels_as_strings = [' '.join([str(l) for l in lls]) for lls in labels]\n\npredictions = pd.DataFrame(data={'session_type': test_session_AIDs.index, 'labels': labels_as_strings})\n\nprediction_dfs = []\n\nfor st in session_types:\n    modified_predictions = predictions.copy()\n    modified_predictions.session_type = modified_predictions.session_type.astype('str') + f'_{st}'\n    prediction_dfs.append(modified_predictions)\n\nsubmission = pd.concat(prediction_dfs).reset_index(drop=True)\nsubmission.to_csv('submission.csv', index=False)\n\nprint(f'Test sessions that we did not manage to extend based on the co-visitation matrix: {no_data_all_aids}')","metadata":{"execution":{"iopub.status.busy":"2022-11-09T05:51:14.512293Z","iopub.execute_input":"2022-11-09T05:51:14.513198Z","iopub.status.idle":"2022-11-09T05:53:47.875400Z","shell.execute_reply.started":"2022-11-09T05:51:14.513143Z","shell.execute_reply":"2022-11-09T05:53:47.874069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following plot (combined with the information printed above) is quite significant -- it show us how much data we are still missing, how many predictions are not at their maximum allowable length.\n\nAnd there is never a point in not outputting all 20 AIDs for any given prediction!","metadata":{}},{"cell_type":"code","source":"from matplotlib import pyplot as plt\n\nplt.hist([len(l) for l in labels]);\nplt.suptitle('Distribution of predicted sequence lengths');","metadata":{"execution":{"iopub.status.busy":"2022-11-09T05:53:47.878091Z","iopub.execute_input":"2022-11-09T05:53:47.878566Z","iopub.status.idle":"2022-11-09T05:53:53.226261Z","shell.execute_reply.started":"2022-11-09T05:53:47.878513Z","shell.execute_reply":"2022-11-09T05:53:53.224793Z"},"trusted":true},"execution_count":null,"outputs":[]}]}