{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":38760,"databundleVersionId":4493939},{"sourceType":"datasetVersion","sourceId":4474043,"datasetId":2601572,"databundleVersionId":4534159}],"dockerImageVersionId":30301,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this notebook we will train a `Word2Vec` model.\n\nWe will use the `gensim` library which offers extremely fast training on the CPU.\n\nWe will again rely on `polars` and its small memory footprint to load and process the data. To speed things up, let's use the dataset in a parquet format (we won't have to deal with `jasonl` files anymore). [I shared the dataset here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint). \n\nWhy are we training word2vec embeddings in the first place?\n\nA session where one action follows another action is very much like... a sentence! In sentences, words that are related appear together. We don't necessarily expect to see the word 'spaceship' in a sentence discussing various ways to cook a steak. The word \"steak\" is more likely to appear close to words such as rosemary, salt, pepper, oil and butter. In this sense, these words are thematically related. And with a large enough corpus we can start making further distinctions! Maybe butter will appear closer in the embedding space to milk than for instance to orange juice, even though both are drinks you can have with your breakfast! (that might be due to milk having the property of being a substance used to produce butter, which might tip the embeddings for \"milk\" and \"butter\" closer together assuming our corpus would contain texts on butter production!).\n\nSimilarly here we can exploit the fact that `aids` appearing in a sequence close together likely share some similarity. A person browsing for gardening equipment is probably not looking at surfboards and vice versa.\n\nOnce we train our model, what will we be able to use it for? First and foremost, candidate generation! Though one might also imagine using it for scoring. Essentially, a model such as this can be very handy in the context of session-based recommendation models!\n\nLet's get to work! 🙂\n\n## Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Data Preprocessing","metadata":{}},{"cell_type":"code","source":"!pip install polars\n\nimport polars as pl\nfrom gensim.test.utils import common_texts\nfrom gensim.models import Word2Vec\n\ntrain = pl.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\ntest = pl.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')","metadata":{"execution":{"iopub.status.busy":"2026-04-29T19:44:50.061342Z","iopub.execute_input":"2026-04-29T19:44:50.061840Z","iopub.status.idle":"2026-04-29T19:45:12.192506Z","shell.execute_reply.started":"2026-04-29T19:44:50.061801Z","shell.execute_reply":"2026-04-29T19:45:12.191006Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let us now transform the data into a format that the `gensim` library can work with. Thanks to `polars` we can do so very efficiently and very quickly.\n\nThere are various ways we could feed our data to our model, however doing so straight from RAM in the form of Python lists is probably one of the fastest! As we have enough resources on Kaggle to do so, let us take this approach!","metadata":{}},{"cell_type":"code","source":"sentences_df = pl.concat([train, test]).groupby('session').agg(\n    pl.col('aid').alias('sentence')\n)","metadata":{"execution":{"iopub.status.busy":"2026-04-29T19:45:12.194897Z","iopub.execute_input":"2026-04-29T19:45:12.195324Z","iopub.status.idle":"2026-04-29T19:45:30.037915Z","shell.execute_reply.started":"2026-04-29T19:45:12.195298Z","shell.execute_reply":"2026-04-29T19:45:30.036416Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sentences = sentences_df['sentence'].to_list()","metadata":{"execution":{"iopub.status.busy":"2026-04-29T19:45:30.039396Z","iopub.execute_input":"2026-04-29T19:45:30.039793Z","iopub.status.idle":"2026-04-29T19:46:21.896681Z","shell.execute_reply.started":"2026-04-29T19:45:30.039757Z","shell.execute_reply":"2026-04-29T19:46:21.894885Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Time to train our model.","metadata":{}},{"cell_type":"markdown","source":"# Training a word2vec model","metadata":{}},{"cell_type":"code","source":"%%time\n\nw2vec = Word2Vec(sentences=sentences, vector_size=32, min_count=1, workers=4)","metadata":{"execution":{"iopub.status.busy":"2026-04-29T19:46:21.898806Z","iopub.execute_input":"2026-04-29T19:46:21.899529Z","iopub.status.idle":"2026-04-29T20:18:56.457202Z","shell.execute_reply.started":"2026-04-29T19:46:21.899411Z","shell.execute_reply":"2026-04-29T20:18:56.455607Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"With the model fully train, let us use similarity between trained representations of our `aids` to create a submission.\n\nThe search functionality where we look for nearest neighbors in the embedding space is built into `gensim`, but it is unfortunately super slow. Let's use `annoy` which is much faster (it performs approximate nearest neigbor search).","metadata":{}},{"cell_type":"code","source":"%%time\n\nfrom annoy import AnnoyIndex\n\naid2idx = {aid: i for i, aid in enumerate(w2vec.wv.index_to_key)}\nindex = AnnoyIndex(32, 'euclidean')\n\nfor aid, idx in aid2idx.items():\n    index.add_item(idx, w2vec.wv.vectors[idx])\n    \nindex.build(10)","metadata":{"execution":{"iopub.status.busy":"2026-04-29T20:18:56.462526Z","iopub.execute_input":"2026-04-29T20:18:56.462972Z","iopub.status.idle":"2026-04-29T20:19:12.810845Z","shell.execute_reply.started":"2026-04-29T20:18:56.462921Z","shell.execute_reply":"2026-04-29T20:19:12.809618Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's create a submission! 🙂","metadata":{}},{"cell_type":"markdown","source":"# Outputting a submission","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nfrom collections import defaultdict\n\nsample_sub = pd.read_csv('../input/otto-recommender-system//sample_submission.csv')\n\nsession_types = ['clicks', 'carts', 'orders']\ntest_session_AIDs = test.to_pandas().reset_index(drop=True).groupby('session')['aid'].apply(list)\ntest_session_types = test.to_pandas().reset_index(drop=True).groupby('session')['type'].apply(list)\n\nlabels = []\n\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\nfor AIDs, types in zip(test_session_AIDs, test_session_types):\n    if len(AIDs) >= 20:\n        # if we have enough aids (over equals 20) we don't need to look for candidates! we just use the old logic\n        weights=np.logspace(0.1,1,len(AIDs),base=2, endpoint=True)-1\n        aids_temp=defaultdict(lambda: 0)\n        for aid,w,t in zip(AIDs,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers[t]\n            \n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n        labels.append(sorted_aids[:20])\n    else:\n        # here we don't have 20 aids to output -- we will use word2vec embeddings to generate candidates!\n        AIDs = list(dict.fromkeys(AIDs[::-1]))\n        \n        # let's grab the most recent aid\n        most_recent_aid = AIDs[0]\n        \n        # and look for some neighbors!\n        nns = [w2vec.wv.index_to_key[i] for i in index.get_nns_by_item(aid2idx[most_recent_aid], 21)[1:]]\n                        \n        labels.append((AIDs+nns)[:20])","metadata":{"execution":{"iopub.status.busy":"2026-04-29T20:19:12.812502Z","iopub.execute_input":"2026-04-29T20:19:12.812834Z","iopub.status.idle":"2026-04-29T20:22:11.544934Z","shell.execute_reply.started":"2026-04-29T20:19:12.812807Z","shell.execute_reply":"2026-04-29T20:22:11.543906Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's now pull it all together and write to a file,","metadata":{}},{"cell_type":"code","source":"labels_as_strings = [' '.join([str(l) for l in lls]) for lls in labels]\n\npredictions = pd.DataFrame(data={'session_type': test_session_AIDs.index, 'labels': labels_as_strings})\n\nprediction_dfs = []\n\nfor st in session_types:\n    modified_predictions = predictions.copy()\n    modified_predictions.session_type = modified_predictions.session_type.astype('str') + f'_{st}'\n    prediction_dfs.append(modified_predictions)\n\nsubmission = pd.concat(prediction_dfs).reset_index(drop=True)\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2026-04-29T20:22:11.546058Z","iopub.execute_input":"2026-04-29T20:22:11.546266Z","iopub.status.idle":"2026-04-29T20:22:43.756857Z","shell.execute_reply.started":"2026-04-29T20:22:11.546247Z","shell.execute_reply":"2026-04-29T20:22:43.755832Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"And we are done!\n\nThank you for reading! Happy Kaggling! 🙌","metadata":{"execution":{"iopub.status.busy":"2022-11-18T02:49:02.940358Z","iopub.execute_input":"2022-11-18T02:49:02.940858Z","iopub.status.idle":"2022-11-18T02:49:02.973867Z","shell.execute_reply.started":"2022-11-18T02:49:02.94076Z","shell.execute_reply":"2022-11-18T02:49:02.972223Z"}}},{"cell_type":"markdown","source":"# BONUS: How to use word2vec to generate candidates/features for training a 2-stage recommender","metadata":{}},{"cell_type":"markdown","source":"Just like a covisiation matrix, for any `AID` `word2vec` can give us a list of `AIDs` resembling our query `AID`. The output will be ordered starting with `AIDs` that are most alike.\n\nIn order for us to visualize what is happening, let me give you a simplified example.","metadata":{}},{"cell_type":"markdown","source":"## Mock data","metadata":{}},{"cell_type":"markdown","source":"In our data we have `aids` organized by `sessions`.","metadata":{}},{"cell_type":"code","source":"data = pl.DataFrame(data={'session': [0, 0, 1, 1], 'aid': [10, 20, 20, 30], 'type': [0, 0, 1, 0]})\ndata","metadata":{"execution":{"iopub.status.busy":"2026-04-29T20:22:43.757941Z","iopub.execute_input":"2026-04-29T20:22:43.758239Z","iopub.status.idle":"2026-04-29T20:22:43.767748Z","shell.execute_reply.started":"2026-04-29T20:22:43.758210Z","shell.execute_reply":"2026-04-29T20:22:43.766529Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can use word2vec to generate candidates. For instance, maybe using word2vec we would generate the following candidates for the sessions in our data:\n\n`{0: [11, 20], 1: [25, 6]}`\n\nWe can reshape our candidates to look as follows:","metadata":{}},{"cell_type":"code","source":"candidates = pl.DataFrame(data={'session': [0, 0, 1, 1], 'aid': [11, 20, 25, 6]})\ncandidates","metadata":{"execution":{"iopub.status.busy":"2026-04-29T20:22:43.769134Z","iopub.execute_input":"2026-04-29T20:22:43.769474Z","iopub.status.idle":"2026-04-29T20:22:43.799700Z","shell.execute_reply.started":"2026-04-29T20:22:43.769428Z","shell.execute_reply":"2026-04-29T20:22:43.798544Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As you see, for our canddiates we don't have too much information apart from `session` and `aid`! This is exactly like the out put `word2vec` can give us!\n\nAnd that is okay. Our ranker can deal with that. For some rows we will have information in this or that column, for another we won't. This is not an issue to a ranking model, `NAs` or `nulls` (depending on the framework) also can carry signal (absence of some type of data).\n\nHere, our ranker will see that we don't have `type` information for candidates... but we will create another important column that will allow it to uniquely identify our candidates as coming from `word2vec`.\n\nHere are the key steps.","metadata":{}},{"cell_type":"markdown","source":"### 1. Add ordering information to our candidates.\n\nThe order is important! A candidate appearing earlier in the list of candidates in some sense has a higher score, is more similar to the AIDs in a session (remember, `word2vec` can give us output ordered by similarity in descending order).","metadata":{}},{"cell_type":"code","source":"candidates = candidates.with_columns(pl.col('aid').cumcount().over('session').alias('word2vec_rank') + 1)\ncandidates","metadata":{"execution":{"iopub.status.busy":"2026-04-29T20:22:43.801119Z","iopub.execute_input":"2026-04-29T20:22:43.801357Z","iopub.status.idle":"2026-04-29T20:22:43.821722Z","shell.execute_reply.started":"2026-04-29T20:22:43.801337Z","shell.execute_reply":"2026-04-29T20:22:43.820356Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2. Merge this information onto candidates.","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"Now, we need to take this information and add this onto our original data.\n\nBut how do we add candidates?! If we just concat these dataframes together, we will have a duplicate entry for session `0` for `aid` of 20.\n\nWhat we need to do is a join but of a specific kind! (hello SQL ideas 👋).\n\nWe want to keep the rows that are already there in `data`, append information to them where there is a match AND create new rows if there isn't.\n\n`outer join` to the rescue!","metadata":{}},{"cell_type":"code","source":"data = data.join(candidates, on=['session', 'aid'], how='outer')\ndata","metadata":{"execution":{"iopub.status.busy":"2026-04-29T20:22:43.823077Z","iopub.execute_input":"2026-04-29T20:22:43.823412Z","iopub.status.idle":"2026-04-29T20:22:43.845047Z","shell.execute_reply.started":"2026-04-29T20:22:43.823390Z","shell.execute_reply":"2026-04-29T20:22:43.843590Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Beautiful! We added candidates, we didn't duplicate rows, information got appended to wehre it should go!\n\nPlus, we added a reflection of the ordering coming from the `word2vec` model in the form of the `word2vec_rank` column.\n\nAnd the code to do it all is provided above!","metadata":{}},{"cell_type":"markdown","source":"### 3. How to practice the above / put it to good use to improve your LB standing.\n\n1. Generate candidates using word2vec from this notebook.\n2. Go to [🏆 Training an XGBoost Ranker on the GPU 🔥🔥🔥](https://www.kaggle.com/code/radek1/training-an-xgboost-ranker-on-the-gpu) and add the features to the training and test data.\n3. Rerun training and see your score improve.","metadata":{}},{"cell_type":"markdown","source":"Hope this can be of help! 🙂 If you found this useful, please upvote the notebook! Thank you 🙏\n\nAlso, if you have any questions, let me know. I'll do my best to try to address them (this is how this bonus section came about, it answers the questions several people kept asking me! 🙂)\n\nThank you for reading!","metadata":{}}]}