{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Do you love Polars, I got Radek's simple Covisitation matrix in polars for you\n\nI have implemented Radek's covisitation matrix model in polars for 3 reasons. \n\n- I know rewriting the notebook can help me understand the logic behind better and write helpful comments of the codes for beginners.\n- I just love the api and flow of polars, writing it in polars is more readable to me. I have even rewritten some pretty cool python codes of Radek's in polars too. I probably take it too far, but I can't help it.\n    - However, sadly, some implementation in polars results in super slower than Radek's pandas + python version, so I reverted those slow parts back to Radek's version.\n\n- I want to have a proper understanding of co-visitation matrix so that further tweaking could be easier for me. (see the sub-sections for details)","metadata":{}},{"cell_type":"markdown","source":"You can read Radek's insights of the original notebook from [here](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic?scriptVersionId=115190040&cellId=1)","metadata":{}},{"cell_type":"markdown","source":"You can set the number to 0.00001 and the whole notebook should finish within 1 minute.","metadata":{}},{"cell_type":"code","source":"DEBUG = False","metadata":{"execution":{"iopub.status.busy":"2023-01-17T15:58:54.916810Z","iopub.execute_input":"2023-01-17T15:58:54.917532Z","iopub.status.idle":"2023-01-17T15:58:54.944264Z","shell.execute_reply.started":"2023-01-17T15:58:54.917447Z","shell.execute_reply":"2023-01-17T15:58:54.942827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install polars\n!pip install snoop\nfrom collections import defaultdict, Counter\nimport gc\nfrom snoop import pp\nimport polars as pl\nimport pandas as pd\nimport numpy as np\nimport random\nfrom polars.testing import assert_frame_equal, assert_series_equal\nfrom datetime import datetime\n\nfrom IPython.core.interactiveshell import InteractiveShell\nInteractiveShell.ast_node_interactivity = \"all\"\npd.set_option('display.max_colwidth', None)\ncfg = pl.Config.restore_defaults()  \npl.Config.set_tbl_rows(50)  \npl.Config.set_fmt_str_lengths(1000)","metadata":{"execution":{"iopub.status.busy":"2023-01-17T15:58:54.946793Z","iopub.execute_input":"2023-01-17T15:58:54.947177Z","iopub.status.idle":"2023-01-17T15:59:17.785618Z","shell.execute_reply.started":"2023-01-17T15:58:54.947153Z","shell.execute_reply":"2023-01-17T15:59:17.784046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if DEBUG: fraction_of_sessions_to_use = 0.00001\nelse: fraction_of_sessions_to_use = 1","metadata":{"execution":{"iopub.status.busy":"2023-01-17T15:59:17.787143Z","iopub.execute_input":"2023-01-17T15:59:17.787513Z","iopub.status.idle":"2023-01-17T15:59:17.795581Z","shell.execute_reply.started":"2023-01-17T15:59:17.787484Z","shell.execute_reply":"2023-01-17T15:59:17.793628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ms = pl.scan_parquet('/kaggle/input/otto-radek-style-polars/train_ms.parquet')\ntest_ms = pl.scan_parquet('/kaggle/input/otto-radek-style-polars/test_ms.parquet')\nsample_sub = pl.scan_csv('/kaggle/input/otto-recommender-system/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2023-01-17T15:59:17.798758Z","iopub.execute_input":"2023-01-17T15:59:17.799145Z","iopub.status.idle":"2023-01-17T15:59:17.827094Z","shell.execute_reply.started":"2023-01-17T15:59:17.799119Z","shell.execute_reply":"2023-01-17T15:59:17.825887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Subset train and test and join them together","metadata":{}},{"cell_type":"code","source":"%%time\nlucky_sessions_train = (\n    train_ms\n    .select([\n        pl.col('session').unique().sample(frac=fraction_of_sessions_to_use, seed=42)\n    ])\n    .collect()\n    .to_series().to_list()\n)\n\nlucky_sessions_test = (\n    test_ms\n    .select([\n        pl.col('session').unique().sample(frac=fraction_of_sessions_to_use, seed=42)\n    ])\n    .collect()\n    .to_series().to_list()\n)\n\n\nsubset_of_train = (\n    train_ms\n    .filter(pl.col('session').is_in(lucky_sessions_train))\n\n)\n\nsubset_of_test = (\n    test_ms\n    .filter(pl.col('session').is_in(lucky_sessions_test))\n\n)\n\nsubsets = pl.concat([subset_of_train, subset_of_test]).collect()\nsessions = subsets.select('session').unique().to_series().to_list()","metadata":{"execution":{"iopub.status.busy":"2023-01-17T15:59:17.828035Z","iopub.execute_input":"2023-01-17T15:59:17.828537Z","iopub.status.idle":"2023-01-17T15:59:54.338767Z","shell.execute_reply.started":"2023-01-17T15:59:17.828511Z","shell.execute_reply":"2023-01-17T15:59:54.337196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pp(lucky_sessions_train[:3], len(lucky_sessions_train), lucky_sessions_test[:3], len(lucky_sessions_test),\n   subset_of_train.collect().height, subset_of_test.collect().height, subsets.height)","metadata":{"execution":{"iopub.status.busy":"2023-01-17T15:59:54.340854Z","iopub.execute_input":"2023-01-17T15:59:54.341395Z","iopub.status.idle":"2023-01-17T16:00:03.472115Z","shell.execute_reply.started":"2023-01-17T15:59:54.341355Z","shell.execute_reply":"2023-01-17T16:00:03.469755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating Co-visitation Matrix - `nextAIDs`\n\n### co-visitation matrix is just a name for exploring the following ideas/questions\n\n- Is there any relationship between one aid/product with other aids/products in the same session or across all sessions? \n- Are there some aids more similar to some and more different to others? \n- Could it be possible when this aid is viewed, some aids are more likely to be clicked/carted/ordered than other aids?\n- Could we pair aids together for each and every session and count the occurrences of pairs?\n- Since one aid (eg., '122') could have many pair-partners, by counting the occurrences of the pairs ('122', pair-partner), could we find the most common pair-partners of aid '122'?\n- Could the next clicks or carts or orders be the most common pair-partners of the last aid (or all aids) of a test session?\n\n### The first challenge of building a co-visitation matrix is how to pairing\n\nIn Radek's notebook, the pairing logic is the following\n\n- use only the last 30 aids of each session to pair on each other\n- remove the pairs of same partners\n- keep pairs whose right-partner is after left-partner within a day\n\n**We can tweak the pairing logic to change our co-visitation matrix**\n\n","metadata":{}},{"cell_type":"code","source":"%%time \nnext_AIDs = defaultdict(Counter)\nchunk_size = 300000\nfor i in range(0, len(sessions), chunk_size):\n    current_chunk = (\n        subsets\n        .filter(pl.col('session').is_between(sessions[i], sessions[np.min([i+chunk_size-1, len(sessions)-1])], closed='both'))\n        .unique() # no duplicates\n        .groupby('session').tail(30)\n    )\n    current_chunk = (\n        current_chunk\n        .join(current_chunk, on='session', suffix='_right')\n        .sort(['session', 'aid', 'aid_right']) # nice view\n        .filter(pl.col('aid') != pl.col('aid_right')) # no need for pairs of themselves\n        .with_columns([\n            ((pl.col('ts_right') - pl.col('ts'))/(24*60*60*1000)).alias('days_elapsed') # differentiate aid_right is after or before aid in days\n        ])\n        .filter((pl.col('days_elapsed')>=0) & (pl.col('days_elapsed') <=1)) # only pairs whose aid_rights are after aid within 24 hrs\n    )\n\n    # defaultdict + Counter is super faster than pure polars solution\n    for aid_x, aid_y in zip(current_chunk.select('aid').to_series().to_list(), current_chunk.select('aid_right').to_series().to_list()):\n        next_AIDs[aid_x][aid_y] += 1\n\n    print(f'{int(np.ceil(i/chunk_size))} out of {int(np.ceil(len(sessions)/chunk_size))} - {np.min([i+chunk_size-1, len(sessions)-1])} sessions are done')\nlen(next_AIDs)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:00:03.474656Z","iopub.execute_input":"2023-01-17T16:00:03.475054Z","iopub.status.idle":"2023-01-17T16:00:03.511115Z","shell.execute_reply.started":"2023-01-17T16:00:03.475028Z","shell.execute_reply":"2023-01-17T16:00:03.509067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"My polars version is about the same to pandas version, see Radek's [speed here](https://www.kaggle.com/code/danielliao/check-the-number-of-next-aids/)","metadata":{}},{"cell_type":"code","source":"del train_ms, subset_of_train, subsets\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:00:03.512473Z","iopub.execute_input":"2023-01-17T16:00:03.512791Z","iopub.status.idle":"2023-01-17T16:00:03.718811Z","shell.execute_reply.started":"2023-01-17T16:00:03.512767Z","shell.execute_reply":"2023-01-17T16:00:03.716695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Using co-visitation matrix to provide aid candidates when there is less than 20 aids in a test session\n\nRadek has shown us two things here: \n\n- how to create a feature (weight based on time, type and occurrences) to select the 20 aids from many in a test session\n- how to select candidates from co-visitation matrix? \n    - take 20 most common aids for each aid in a test session and put them into candidate list\n    - take 40 most common aids from the candidate list and add them to the aids of the test session if they are new to the session\n    - then select the first 20 aids","metadata":{}},{"cell_type":"code","source":"%%time\nlists_aids_types = (\n    test_ms\n    .unique() #\n    .groupby('session')\n    .agg([\n        pl.col('aid').list().alias('test_session_AIDs'),\n        pl.col('type').list().alias('test_session_types'),        \n    ])\n    .collect()\n)\n\nlists_aids_types.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:00:03.723139Z","iopub.execute_input":"2023-01-17T16:00:03.723515Z","iopub.status.idle":"2023-01-17T16:00:06.574667Z","shell.execute_reply.started":"2023-01-17T16:00:03.723490Z","shell.execute_reply":"2023-01-17T16:00:06.573800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlabels = []\nsession_types = ['clicks', 'carts', 'orders']\nno_data = 0\nno_data_all_aids = 0\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\ntest_session_AIDs = lists_aids_types.select('test_session_AIDs').to_series().to_list()\ntest_session_types = lists_aids_types.select('test_session_types').to_series().to_list()\n\n# take each session's aids and types\nfor AIDs, types in zip(test_session_AIDs, test_session_types):\n\n    # if the session has more than 20 aids\n    if len(AIDs) >= 20: \n        # np.logspace: Return numbers spaced evenly on a log scale.\n        # `-1` is to ensure the weights ranges between [0,1]\n        # the weights is given to AIDs based on the time order or chronological order\n        weights=np.logspace(start=0.1,stop=1,num=len(AIDs),base=2, endpoint=True)-1 \n        \n        # create a defaultdict for this session only\n        # anything added into this dict will have a default value 0\n        # try `aids_temp[1]` and `aids_temp`\n        aids_temp=defaultdict(lambda: 0)\n        \n        # in each sess, an aid may occur multiples in multiple types at different time, \n        # the line below is to take all 3 factors into account to value the importance of this aid to the session\n        # each unique aid and its aggregated weight are stored in a defaultdict\n        for aid,w,t in zip(AIDs,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers[t]\n          \n        # let's \n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n\n        # when using the polars below to replace the line above, it is actually 2 times slower\n        # aid = [key for (key, value) in aids_temp.items()]\n        # adwt = [value for (key, value) in aids_temp.items()]\n        # sorted_aids = (\n        #     pl.DataFrame([aid, adwt], columns=['aid', 'weight'])\n        #     .sort('weight', reverse=True)\n        #     .select('aid').to_series().to_list()\n        # )\n\n        # take the 20 aids with the largest weights from this session as one list and append it into a new list `labels`\n        labels.append(sorted_aids[:20])\n    \n    # when this session has less than 20 aids\n    else:\n        # reverse the order of AIDs (a list of aids of this session) and remove the duplicated aids\n        AIDs = list(dict.fromkeys(AIDs[::-1])) # python version\n        \n        # If using this polars below to replace the line above, it is infinitely slower\n        # AIDs = pl.Series('aid', AIDs).unique().reverse().to_list() # polars version\n\n        # keep track of the length of new AIDs above\n        AIDs_len_start = len(AIDs)\n        \n\n        candidates = []\n        # take each unique aid of this session, access its the 20 most common pair-partners and their counts\n        # insert the list of the 20 most common pair-partner aids into another list `candidates` (only a pure list )\n        # in the end, this `candidates` list is a lot and has many duplicates too\n        for AID in AIDs:\n            if AID in next_AIDs: candidates += [aid for aid, count in next_AIDs[AID].most_common(20)]\n                \n        # take the 40 most common aids from `candidates`, and if they are already inside AIDs of this session, \n        # then insert them into AIDs (still a pure list because of `+`, and `append` can't do it)\n        AIDs += [AID for AID, cnt in Counter(candidates).most_common(40) if AID not in AIDs]\n        \n        # but we still only take the first 20 aids from AIDs as this session's prediction and store it in `labels`\n        labels.append(AIDs[:20])\n        \n        # if no candidates are generated, count 1 to `no_data`\n        # if candidates == []: no_data += 1 # this variable is actually not used by Radek\n        \n        # keep an account of the num of aids in this session and all sessions which adding no candidates\n        if AIDs_len_start == len(AIDs): no_data_all_aids += 1\n","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:00:06.576266Z","iopub.execute_input":"2023-01-17T16:00:06.576880Z","iopub.status.idle":"2023-01-17T16:00:26.634231Z","shell.execute_reply.started":"2023-01-17T16:00:06.576847Z","shell.execute_reply":"2023-01-17T16:00:26.633144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub.fetch().head()","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:00:26.635836Z","iopub.execute_input":"2023-01-17T16:00:26.636189Z","iopub.status.idle":"2023-01-17T16:00:26.653887Z","shell.execute_reply.started":"2023-01-17T16:00:26.636156Z","shell.execute_reply":"2023-01-17T16:00:26.652062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating submission from labels","metadata":{}},{"cell_type":"code","source":"%%time\n(\n    pl.DataFrame({'session': lists_aids_types.select('session').to_series().to_list(), \n                  'labels': labels})\n    .with_columns([\n        pl.col('labels').arr.eval(pl.element().cast(pl.Utf8)).arr.join(' '),\n        (pl.col('session')+\"_clicks\").alias('clicks'),\n        (pl.col('session')+\"_carts\").alias('carts'),\n        (pl.col('session')+\"_orders\").alias('orders'),        \n    ])\n    .select([\n        'session',\n        pl.concat_list(['clicks', 'carts', 'orders']).alias('session_type'),\n        'labels'\n    ])\n    .explode('session_type')\n    .sort('session')\n    .select(pl.exclude('session'))\n    .write_csv('submission.csv')\n)\nprint(f'Test sessions that we did not manage to extend based on the co-visitation matrix: {no_data_all_aids}')","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:00:26.656188Z","iopub.execute_input":"2023-01-17T16:00:26.656705Z","iopub.status.idle":"2023-01-17T16:00:32.453949Z","shell.execute_reply.started":"2023-01-17T16:00:26.656663Z","shell.execute_reply":"2023-01-17T16:00:32.453195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pl.read_csv('submission.csv').shape\nsample_sub.collect().shape","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:01:10.012043Z","iopub.execute_input":"2023-01-17T16:01:10.012484Z","iopub.status.idle":"2023-01-17T16:01:11.085288Z","shell.execute_reply.started":"2023-01-17T16:01:10.012447Z","shell.execute_reply":"2023-01-17T16:01:11.084046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following plot (combined with the information printed above) is quite significant -- it show us how much data we are still missing, how many predictions are not at their maximum allowable length.\n\nAnd there is never a point in not outputting all 20 AIDs for any given prediction!","metadata":{}},{"cell_type":"code","source":"# from matplotlib import pyplot as plt\n\n# plt.hist([len(l) for l in labels]);\n# plt.suptitle('Distribution of predicted sequence lengths');","metadata":{"execution":{"iopub.status.busy":"2023-01-17T16:00:32.986408Z","iopub.status.idle":"2023-01-17T16:00:32.986771Z","shell.execute_reply.started":"2023-01-17T16:00:32.986606Z","shell.execute_reply":"2023-01-17T16:00:32.986621Z"},"trusted":true},"execution_count":null,"outputs":[]}]}