{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Test Dataset Is All We Need?\nThis notebook is a fork of Cdeotte's notebook [here][2] and Vladimir's notebook [here][1]. Please upvote their original notebook.\n\nHere, I trained the model ONLY with test dataset. ~~The score is 0.541, but slightly lower than that of Cdeotte's notebook (see LB ranking) but higher than Vladimir's notebook. So, test dataset is very important but not all we need.~~ \n\nUPDATE: I mistakingly submitted the same result as Cdeotte's notebook. The true score is 0.522, which is lower than that of Vladimir's notebook. So, the test set is NOT all we need. Thank you Radek Osmulski for pointing it out!\n\n# Test Data Leak - LB Boost\nThis notebook is a fork of Vladimir's notebook [here][1]. Please upvote his original notebook.\n\nIn Kaggle's OTTO – Multi-Objective Recommender System Competition, we are given both the train data (from present) and the test data (from future). The test data contains partial sequences of sessions. Therefore we can use these partial sequences to train our models in addition to using the train data. In this notebook we begin with Vladimir's Co-visitation Matrix notebook [here][1] and add test data to the training data. We make one change to his code below. Where he loads train parquets with `'../input/otto-chunk-data-inparquet-format/train_parquet/*'` , we change this to include test with `'../input/otto-chunk-data-inparquet-format/*_parquet/*'`.\n\nThis is an experiment to see if using test data during training will boost LB score. The original notebook achieved LB 0.539. Let's see what this notebook achieves. Note this method of using test data cannot be used in real life because some of the data we are training with occurs in the future of some of the inference data. The inference test data is from a one week period. Therefore when we infer the first day of the week, we have the advantage of using data from the last 6 days of the week which is impossible in real life.\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost","metadata":{}},{"cell_type":"markdown","source":"# OTTO: Co-visitation Matrix\n\nThere exist products that are frequently viewed and bought together. Here we leverage this idea by computing a co-visitation matrix of products. It's done in the following way:\n\n1. First we look at all pairs of events within the same session that are close to each other in time (< 1 day). We compute co-visitation matrix $M_{aid1,aid2}$ by counting global number of event pairs for each pair across all sessions.\n2. For each $aid1$ we find top 20 most frequent aid2:  `aid2=argsort(M[aid])[-20:]`\n3. We produce test results by concatenating `tail(20)` of test session events (see https://www.kaggle.com/code/simamumu/old-test-data-last-20-aid-get-lb0-947) with the most likely recommendations from co-visitation matrix. These recommendations are generated from session AIDs and `aid2` from the step 2\n\n\n**Please, smash that thumbs up button and subscribe if you like this notebook!**","metadata":{}},{"cell_type":"markdown","source":"## Utils, imports","metadata":{}},{"cell_type":"code","source":"### import numpy as np\nfrom collections import defaultdict\nimport pandas as pd\nfrom tqdm.notebook import tqdm\nimport glob\nimport numpy as np\nimport multiprocessing\nimport os\nimport pickle\nimport gc\nimport glob\nfrom collections import Counter\n\nDEBUG=False   \nSAMPLING = 1  # Reduce it to improve performance","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:34:16.185399Z","iopub.execute_input":"2022-12-05T10:34:16.185852Z","iopub.status.idle":"2022-12-05T10:34:16.308199Z","shell.execute_reply.started":"2022-12-05T10:34:16.185756Z","shell.execute_reply":"2022-12-05T10:34:16.306982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TOP_20_CACHE = '../input/otto-pickles/top_20_aids_v1.pkl'\nTOP_20_CACHE = 'NO_DATA'\n\ntry:\n    from kaggle_secrets import UserSecretsClient\n    user_secrets = UserSecretsClient()\n    secret_value_0 = user_secrets.get_secret(\"gcloud\")\n\n    with open('/tmp/json', 'w+') as f:\n        f.write(secret_value_0)\n        \n    !gcloud auth login --cred-file /tmp/json    \n    !gsutil cp gs://nesp/top_20_aids.pkl .        \n        \nexcept Exception  as ex:\n    pass","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:34:16.312485Z","iopub.execute_input":"2022-12-05T10:34:16.313231Z","iopub.status.idle":"2022-12-05T10:34:16.420219Z","shell.execute_reply.started":"2022-12-05T10:34:16.313197Z","shell.execute_reply":"2022-12-05T10:34:16.419375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generate AID pairs","metadata":{}},{"cell_type":"code","source":"import sys\ndef gen_pairs(df):\n    df = df.query('session % @SAMPLING == 0').groupby('session', as_index=False, sort=False).apply(lambda g: g.tail(30)).reset_index(drop=True)\n    df = pd.merge(df, df, on='session')\n    pairs = df.query('abs(ts_x - ts_y) < 24 * 60 * 60 * 1000 and aid_x != aid_y')[['session', 'aid_x', 'aid_y']].drop_duplicates()\n    return pairs[['aid_x', 'aid_y']].values\n    \n\ndef gen_aid_pairs():\n    all_pairs = defaultdict(lambda: Counter())\n    with tqdm(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*'), desc='Chunks') as prog:\n        with multiprocessing.Pool(4) as p:\n            for idx, chunk_file in enumerate(prog):\n                chunk = pd.read_parquet(chunk_file).drop(columns=['type'])\n                pair_chunks = p.map(gen_pairs, np.array_split(chunk.head(100000000 if not DEBUG else 10000), 120))            \n                for pairs in pair_chunks:\n                    for aid1, aid2 in pairs:\n                        all_pairs[aid1][aid2] +=1 \n                prog.set_description(f'Mem: {sys.getsizeof(object) // (2 ** 20)}MB')\n\n                if DEBUG and idx >= 2:\n                    break\n                del chunk, pair_chunks\n                gc.collect()\n    return all_pairs\n        \nif os.path.exists(TOP_20_CACHE):\n    print('Reading top20 AIDs from cache')\n    top_20 = pickle.load(open(TOP_20_CACHE, 'rb'))\nelse:\n    all_pairs = gen_aid_pairs()\n    df_top_20 = []\n    for aid, cnt in tqdm(all_pairs.items()):\n        df_top_20.append({'aid1': aid, 'aid2': [aid2 for aid2, freq in cnt.most_common(20)]})\n\n    df_top_20 = pd.DataFrame(df_top_20).set_index('aid1')\n    top_20 = df_top_20.aid2.to_dict()\n    import pickle\n    with open('top_20_aids.pkl', 'wb') as f:\n        pickle.dump(top_20, f)\n        \nlen(top_20)","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:34:16.421705Z","iopub.execute_input":"2022-12-05T10:34:16.422381Z","iopub.status.idle":"2022-12-05T10:39:32.50454Z","shell.execute_reply.started":"2022-12-05T10:34:16.42234Z","shell.execute_reply":"2022-12-05T10:39:32.503315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, (k, v) in enumerate(top_20.items()):\n    print(k, v)\n    if i > 10:\n        break","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:39:32.506902Z","iopub.execute_input":"2022-12-05T10:39:32.507574Z","iopub.status.idle":"2022-12-05T10:39:32.513652Z","shell.execute_reply.started":"2022-12-05T10:39:32.50754Z","shell.execute_reply":"2022-12-05T10:39:32.512887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test set inference","metadata":{}},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(tqdm(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*'))):\n        chunk = pd.read_parquet(chunk_file)\n        dfs.append(chunk)\n\n    return pd.concat(dfs).reset_index(drop=True).astype({\"ts\": \"datetime64[ms]\"})","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:39:32.514694Z","iopub.execute_input":"2022-12-05T10:39:32.515562Z","iopub.status.idle":"2022-12-05T10:39:32.523819Z","shell.execute_reply.started":"2022-12-05T10:39:32.51552Z","shell.execute_reply":"2022-12-05T10:39:32.523035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = load_test()","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:39:32.524835Z","iopub.execute_input":"2022-12-05T10:39:32.525595Z","iopub.status.idle":"2022-12-05T10:39:34.54614Z","shell.execute_reply.started":"2022-12-05T10:39:32.525564Z","shell.execute_reply":"2022-12-05T10:39:34.544933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import itertools\n\ndef suggest_aids(df):\n    aids = df.tail(20).aid.tolist()\n    \n    if len(aids) >= 20:\n        # We have enough events in the test session\n        return aids\n    \n    # Append it with AIDs from the co-visitation matrix. \n    aids = set(aids)\n    aids2 = list(itertools.chain(*[top_20[aid] for aid in aids if aid in top_20]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in aids]        \n    return list(aids) + top_aids2[:20 - len(aids)]\n\n        ","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:39:34.547491Z","iopub.execute_input":"2022-12-05T10:39:34.547825Z","iopub.status.idle":"2022-12-05T10:39:34.555246Z","shell.execute_reply.started":"2022-12-05T10:39:34.547795Z","shell.execute_reply":"2022-12-05T10:39:34.554229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = test_df.sort_values([\"session\", \"type\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_aids(x)\n)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-05T10:39:34.55659Z","iopub.execute_input":"2022-12-05T10:39:34.557445Z","iopub.status.idle":"2022-12-05T10:45:53.922934Z","shell.execute_reply.started":"2022-12-05T10:39:34.557413Z","shell.execute_reply":"2022-12-05T10:45:53.921944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:45:53.924767Z","iopub.execute_input":"2022-12-05T10:45:53.926157Z","iopub.status.idle":"2022-12-05T10:45:57.326368Z","shell.execute_reply.started":"2022-12-05T10:45:53.926106Z","shell.execute_reply":"2022-12-05T10:45:57.32507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:45:57.329408Z","iopub.execute_input":"2022-12-05T10:45:57.329767Z","iopub.status.idle":"2022-12-05T10:45:57.341664Z","shell.execute_reply.started":"2022-12-05T10:45:57.329734Z","shell.execute_reply":"2022-12-05T10:45:57.340575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat(\n    [clicks_pred_df, orders_pred_df, carts_pred_df]\n)\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-12-05T10:45:57.343025Z","iopub.execute_input":"2022-12-05T10:45:57.343354Z","iopub.status.idle":"2022-12-05T10:47:11.633743Z","shell.execute_reply.started":"2022-12-05T10:45:57.343327Z","shell.execute_reply":"2022-12-05T10:47:11.632638Z"},"trusted":true},"execution_count":null,"outputs":[]}]}