{
  "id": 374229,
  "title": "One Month Left - Here is what you need to know!",
  "url": "/competitions/otto-recommender-system/discussion/374229",
  "author_name": "",
  "post_date": "2022-12-26T04:42:06.261120900Z",
  "votes": 89,
  "comment_count": 8,
  "views": 0,
  "content": "<h3>Here is what you need to know</h3>\n<h5>One month to go!</h5>\n<hr>\n<h5>Recommendation Systems for Large Datasets (<a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a>)</h5>\n<p><a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a> discussed the use of recommendation systems for large datasets, such as the one used for this competition. The train dataset consists of: </p>\n<ul>\n<li>12,899,779 sessions</li>\n<li>1,855,603 items</li>\n<li>216,716,096 events</li>\n<li>194,720,954 clicks</li>\n<li>16,896,191 carts</li>\n<li>5,098,951 orders</li>\n</ul>\n<p>He expand on this and add information about the stages for solving this problem:<br>\n<strong>Candidate Generation</strong></p>\n<p>Example criteria you can use to select you candidates:</p>\n<ul>\n<li>previously purchased items</li>\n<li>repurchased items</li>\n<li>overall most popular items</li>\n<li>similar items based on some sort of clustering technique</li>\n<li>similar items based on something such as a co-visitation matrix</li>\n</ul>\n<p>By now we should have much fewer items for each session, so we should be able to input these into a ranker model.</p>\n<p><strong>Ranking Model Examples</strong></p>\n<ul>\n<li><a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html\" target=\"_blank\">LGBMRanker</a></li>\n<li><a href=\"https://medium.com/predictly-on-tech/learning-to-rank-using-xgboost-83de0166229d\" target=\"_blank\">XGBRanker</a></li>\n<li><a href=\"https://towardsdatascience.com/learning-to-rank-with-python-scikit-learn-327a5cfd81f\" target=\"_blank\">Ranking with sklearn</a></li>\n<li><a href=\"https://maroo.cs.umass.edu/getpdf.php?id=1373\" target=\"_blank\">Neural Network Ranker</a></li>\n</ul>\n<hr>\n<h5>How To Build a GBT Ranker Model (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> provided an easy way to build a solution for the competition by creating candidate rerank models. He explained how to create train data, train and infer a GBT ranker model.</p>\n<hr>\n<h5>local validation tracks public LB perfecty -- here is the setup (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has found that local validation tracks public leaderboard (LB) performance perfectly. They have shared their setup <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> to allow others to replicate their results. The setup includes a scoring function that uses the same public LB metric and a test set of 500 images that corresponds to the public LB.</p>\n<hr>\n<h5>💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> Shares an <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">EDA</a> notebook, some [dataset] he preprocessed to parquet and a <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">robust validation framework</a>.</p>\n<hr>\n<h5>30x Faster Co-Visitation Matrices using RAPIDS cuDF! (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> published a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369\" target=\"_blank\">notebook</a> which uses RAPIDS cuDF to compute co-visitation matrices 30x faster. These matrices help to provide models with \"candidates\" which can then be reranked and selected for submission CSV. This process is useful for improving prediction accuracy.</p>\n<hr>\n<h5>Surprising LB 0.587 !!!Share some experimental results ~~ (<a href=\"https://www.kaggle.com/jiahongxie\" target=\"_blank\">Jiahong Xie</a>)</h5>\n<p><a href=\"https://www.kaggle.com/jiahongxie\" target=\"_blank\">Jiahong Xie</a> has made an amazing discovery with a surprising LB score of 0.587! They shared their experimental results which showed that they had a recall rate of 200 candidates for each user. This was an impressive result and could provide valuable insights for other Kagglers.</p>\n<p>In short: And generate 300+ Features for each user and item pair.In training stage and downsample positive:negative as 1:20 for training. Also small improvement was made by deleting aid col in features.</p>\n<hr>\n<h5>📈 What do we know so far? ⚡Summary with  links to relevant resources (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> created an incredibly useful post which contains links to relevant resources related to the Santa 2022 competition. He also shared a repo on GitHub which contains data for the competition, including preprocessing code and information not available on Kaggle. This repo can be accessed <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">here</a>.</p>\n<hr>\n<h5>📑 [Step-by-step guide] How I got to my current standing on the LB and how to improve going forward (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> shared <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278\" target=\"_blank\">this post</a> on how he achieved his current standing on the leaderboard and how to improve it going forward.</p>\n<ul>\n<li>The main idea is to use a <strong>divide-and-conquer approach to search for small improvements</strong>, by breaking down the problem into smaller pieces and looking for more efficient solutions.</li>\n<li>He advises to use a combination of different methods (random search, minimum spanning tree, etc.), as well as to take into account the constraints of the problem, such as the maximum link length.</li>\n<li>He also suggests revisiting points, as well as to experiment with different strategies for finding the optimum solution. Lastly, he recommends to use the baseline functions provided but with modifications to make them faster and more efficient.</li>\n</ul>\n<hr>\n<h5>Full dataset processed to CSV/parquet files with optimized memory footprint (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> shared a post discussing how to process the full dataset to CSV/parquet files with an optimized memory footprint.</p>\n<ul>\n<li>The main idea is to use dask to read the raw data, filter it, and then save it to parquet files using the dask.dataframe.to_parquet method.</li>\n<li>The resulting files are much smaller than the original and can be loaded into memory much faster. Additionally, the post provides a code example of how to use dask to process the data.</li>\n</ul>\n<hr>\n<h5>Locate Real Users and Real Sessions EDA (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> discussed the use of the term \"session\" in Kaggle's Otto competition in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138\" target=\"_blank\">this post</a>.</p>\n<ul>\n<li>In the competition, \"session\" actually means \"user\". We are given train data for 12,899,779 users over a 4 week period, and test data for users during 1 week (in the future).</li>\n<li>We must predict what a user will do in the remainder of the 1 week that we do not have information about.</li>\n</ul>\n<hr>\n<h5>There are some bugs in evaluation (<a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">Tawara</a>)</h5>\n<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">Tawara</a> discovered a bug in the evaluation which resulted in a score of 4.848 on the competition.</p>\n<ul>\n<li>He shared the <a href=\"https://www.kaggle.com/code/ttahara/last-aid-20/\" target=\"_blank\">notebook</a> he used to achieve this score.</li>\n<li>It is important to keep in mind that the bug may have been fixed and the score may have changed.</li>\n</ul>\n<hr>\n<h5>Andrew Ng Recommender Systems (<a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">Andrada Olteanu</a>)</h5>\n<p><a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">Andrada Olteanu</a> discussed the fundamentals of Recommender Systems as outlined by Andrew Ng in his course on Machine Learning.<br>\n <strong>The main topics covered include:</strong></p>\n<ul>\n<li>Collaborative Filtering and its two main approaches: User-User and Item-Item.</li>\n<li>Content-Based Filtering which uses the attributes of the items being recommended.</li>\n<li>Hybrid Methods which combine the two approaches.</li>\n<li>Matrix Factorization which is used to classify users and items into latent features.</li>\n<li>Evaluation of Recommender Systems which includes methods such as Mean Average Precision (MAP) and Root Mean Squared Error (RMSE).</li>\n<li>Challenges of Recommender Systems such as scalability and cold-start problem.</li>\n</ul>\n<hr>\n<h5>How to thrive in this competition without going crazy -- 1 out of 2 important truths ❤️‍🔥 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has shared <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503\" target=\"_blank\">this post</a> as a cautionary reminder of the complexity of this competition.</p>\n<ul>\n<li>He has highlighted two important truths to keep in mind in order to thrive:</li>\n<li><strong>Complexity:</strong> Once you start to dig deeper, the competition becomes increasingly complex.</li>\n<li><strong>Mindset:</strong> Having the right mindset is crucial to success. Be patient and optimistic, while also being mindful of the complexity.</li>\n</ul>\n<hr>\n<h5>The Story of User 13479136 and User 13710374 (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p>In the Kaggle's Otto competition, two particular users stood out:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678\" target=\"_blank\">User 13479136</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678\" target=\"_blank\">User 13710374</a></li>\n<li><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> managed to extract item categories using RAPIDS TSNE such different categories of items may be clothing and electronics.</li>\n<li>He then continue to deconstruct the individual user behaviour in the data.</li>\n</ul>\n<p><strong>Very interesting read!</strong></p>\n<hr>\n<h5>Yes, We Can Use Test Data Leakage. (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> discussed the possibility of using test data leakage in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939\" target=\"_blank\">this post</a>.</p>\n<ul>\n<li>Since we are predicting test events occurring within a one week period, it is possible to use the test data to predict future events.</li>\n<li>This could be done by using a rolling window approach, where the training data is used to predict the test data in the following steps.</li>\n<li>The test data can then be used to update the model and predict the next step.</li>\n<li>This approach is beneficial as it allows for more accurate predictions in the future and could be used to improve the model's performance.</li>\n</ul>\n<hr>\n<h5>Say NO to GPU: Nx Faster Co-Visitation Matrices using single CPU! (<a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">Carno Zhao</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">Carno Zhao</a> published a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365873\" target=\"_blank\">notebook</a> which uses <code>numba.jit</code> to compute Nx faster co-visitation matrices.</li>\n<li>This is an incredibly useful tool to make computations faster and more efficient.</li>\n</ul>\n<hr>\n<h5>💡 What is the co-visitation matrix, really? (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discussed a concept known as the co-visitation matrix in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358\" target=\"_blank\">this post</a> and how it relates to modern techniques. </p>\n<p><strong>Interesting starter reading for this competition</strong></p>\n<hr>\n<h5>A top-down perspective on the current metric values (<a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">narsil</a>)</h5>\n<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">narsil</a> shared a useful post discussing the current metric values from a top-down perspective.</p>\n<p><strong>Sessions in the test set are truncated in a random moment of the session life.</strong></p>\n<ul>\n<li>For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart &amp; purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.</li>\n<li>For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat (see the winning scores for H&amp;M competition - they are close to 0). This is different metric, but shows how difficult such a problem is.</li>\n</ul>\n<p><strong>Some intuitions:</strong></p>\n<ul>\n<li>If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard</li>\n<li>It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.</li>\n</ul>\n<hr>\n<h5>How to train a Word2Vec model for item embeddings - a simple code example 📖 (<a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a>)</h5>\n<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a> shared a useful post on how to train a Word2Vec model to create item embeddings.<br>\nThis includes a simple code example and detailed explanation of the process.</p>\n<pre><code> pandas  pd\n gensim.models  Word2Vec\ncs = \ncount = \ntotal = \n (, )  f:\n     df  pd.read_json(, lines=, chunksize=cs):\n         val  df[].apply( x: [event[]  event  x]).values:\n            f.write(.join((, val)) + )\n        count += cs\n        (, end=)\nmodel = Word2Vec(corpus_file=, vector_size=, window=, min_count=, workers=)\nmodel.save()\n</code></pre>\n<p>And after training the model we can use it like this:</p>\n<pre><code>model = Word2Vec.load()\nmodel.wv.most_similar(, topn=)\n</code></pre>\n<hr>\n<h5>[Starter pack] LGBMRanker with polars 🚀🚀🚀 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>, he explain how re-ranking models are the industry standard for dealing with datasets like those in the Otto Recommender System competition, which have high cardinality categories.</li>\n<li>He also show how to use LightGBM's ranker to create a baseline model that is competitive with the leaderboard. The post also provides a starter pack to help people getstarted, with instructions on how to use the ranker, what the parameters mean and how to tune them, as well as how to use the model for recommendations. Finally, the post also provides a notebook to help users to get their own model up and running.</li>\n</ul>\n<hr>\n<h5>[Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> shared a great <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893\" target=\"_blank\">post</a> on improving candidate generation using a matrix factorization model!</p>\n<ul>\n<li>This post provides a starter pack with a Pytorch implementation and a Merlin Dataloader to get started. With this starter pack, you can load data, define the model, and start training.</li>\n<li>This will allow you to quickly get up and running with matrix factorization and improve candidate generation.</li>\n</ul>\n<hr>\n<h5>Saving GPU memory when processing features == (Speedup + Efficiency) (<a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">Giba</a>)</h5>\n<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">Giba</a> discussed a new functionality of Cudf that can help save GPU memory when processing features.</p>\n<ul>\n<li>By setting the default data type to 32-bit instead of 64-bit, the memory taken up by the GPU can be reduced. The code for this can be found below:</li>\n</ul>\n<pre><code> cudf\ncudf.set_allocator(, default_dtype=np.float32)\n</code></pre>\n<hr>\n<h5>I think word2vec is a nice choice. (<a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">taku_sid</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">taku_sid</a> shared a post discussing the use of word2vec to generate candidates.</li>\n<li>Managed to get to a score of lb 0.578 which is great for word2vec.</li>\n<li>Some of the comments discussions includes important information about differt models used on this competition.</li>\n</ul>\n<p><strong>Recommended!</strong></p>\n<hr>\n<h5>Processed dataset (<a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">Konrad Banachewicz</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">Konrad Banachewicz</a> created a post <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363617\" target=\"_blank\">here</a> providing a training dataset in both .csv and .parquet formats for those who love pandas.</li>\n<li>This can be a useful resource for anyone looking to quickly get started with the dataset.</li>\n</ul>\n<hr>\n<h5>💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition) (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>, he discussed several strategies for working on ML projects more efficiently.</p>\n<p><strong>He include:</strong></p>\n<ul>\n<li>Breaking projects into smaller tasks and setting up checkpoints</li>\n<li>Using domain knowledge to reduce the size of the search space</li>\n<li>Leveraging existing frameworks, libraries, and tools</li>\n<li>Looking for pre-existing solutions to problems</li>\n<li>Focusing on one task at a time</li>\n<li>Streamlining the data collection and labelling process</li>\n<li>Setting up automated evaluations and feedback loops</li>\n<li>Taking breaks to reflect and reassess goals</li>\n<li>Working in collaboration with others</li>\n</ul>\n<hr>\n<h5>Which metric is correct? (<a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a>)</h5>\n<p><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a> raised a question on the evaluation metric for this competition: Recall@k.</p>\n<p><strong><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a> asked which of the following two definitions is correct:</strong></p>\n<ul>\n<li>(1) Recall@k is the fraction of the top k pixels of an image that are the same as the target image.</li>\n<li>(2) Recall@k is the fraction of the target image that is in the top k pixels of an image.</li>\n</ul>\n<p>The discussion that followed suggested that the <strong>second definition</strong> is the correct one.</p>\n<hr>\n<h5>💡How to improve the results of your Approximate Nearest Neighbor search! (annoy) (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> gave an insightful post on <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368385\" target=\"_blank\">how to improve the results of your Approximate Nearest Neighbor search</a>!</p>\n<p><strong>Main Idea:</strong> Use an Approximate Nearest Neighbor (ANN) search to improve the results of your search.</p>\n<p><strong>Benefits:</strong></p>\n<ul>\n<li>ANN searches are much faster than traditional nearest neighbor searches and can be used to quickly find similar items.</li>\n<li>ANN searches can also be used to find items with similar characteristics, even if they don’t have an exact match.</li>\n<li>ANNs are also more scalable, as they can be easily parallelized and distributed across multiple nodes.</li>\n</ul>\n<p><strong>Implementation:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggests using the <a href=\"https://github.com/spotify/annoy\" target=\"_blank\">Annoy library</a> to implement the ANN search.</li>\n<li>Annoy is a Python library that allows you to quickly index and search for points in a vector space.</li>\n<li>Annoy also includes a few useful features such as the ability to easily control the size of the search pool, as well as the ability to specify the number of neighbors to search for.</li>\n</ul>\n<p><strong>Conclusion:</strong></p>\n<p>Using an Approximate Nearest Neighbor search can provide a number of benefits, such as increased search speed and scalability. <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggests using the Annoy library to implement the ANN search, which provides a number of useful features for controlling the size of the search pool and specifying the number of neighbors to search for.</p>\n<hr>\n<h5>💡 Training an XGBoost Ranker on the GPU with Merlin Models 🔥🔥🔥 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368848\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discussed the use of Merlin Models to train an XGBoost Ranker on the GPU.<br>\nThis can greatly reduce the training time for XGBoost models, allowing for faster and more efficient training. The main benefits are that it does not require a Tensorflow backend, and allows for faster training. It also allows for better hyperparameter optimization, since it is possible to train multiple models in parallel. Additionally, the model can be deployed to the cloud for scalability and higher performance. Finally, the model can also be used for inference tasks such as ranking and recommendation.</p>\n<hr>\n<h5>Matrix Factorization with GPU: 6.5x faster! (<a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a> shared a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371166\" target=\"_blank\">very interesting notebook</a> to compute matrix factorization, using polar, annoy, Merlin data loader and pytorch.</li>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a> then created a GPU version of the notebook which is <a href=\"https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\" target=\"_blank\">available here</a> and is 6.5x faster than the original version!</li>\n</ul>\n<hr>\n<h5>🐘 the elephant in the room -- high cardinality of targets and what to do about this (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discusses an issue that is often overlooked when working with Machine Learning models, high cardinality of targets and how to handle it. The main idea presented is that, in order to build an effective model, it is important to accurately measure the importance of each feature. The author suggests two main approaches: </p>\n<ul>\n<li><strong>Feature Engineering:</strong> transforming existing features into a different representation that can be used to better understand the underlying data. </li>\n<li><strong>Dimensionality Reduction:</strong> reducing the number of features to a more manageable amount.</li>\n</ul>\n<p>The author also outlines some potential pitfalls, such as overfitting and underfitting, that can arise when dealing with high cardinality. Finally, the post provides some tips on how to effectively use the two main approaches to tackle this problem.</p>\n<hr>\n<h5>Co-Visitation Matrices and Matrix Factorization (<a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a>)</h5>\n<p><a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a> discussed the use of co-visitation matrices and matrix factorization for recommendation systems in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589\" target=\"_blank\">this post</a>. The main idea is to create features based on items which are frequently viewed together. This technique has been shown to be effective for creating better recommendation systems.</p>\n<hr>\n<h5>💡What is a good initial goal in the competition? How to improve beyond it? 📈 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368685\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggested a good initial goal for the competition is to get everything up and running on your end. This includes setting up your environment and running the notebooks. Additionally, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggested that to improve beyond this goal, it is important to understand the underlying techniques and concepts, such as using the right data structure and algorithm for the task, as well as understanding the limitations of the algorithms. Finally, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> recommended debugging your code and understanding the underlying algorithms, such as a greedy approach, to improve beyond the initial goal.</p>\n<hr>\n<h5>To the people still hacking away on this competition… (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has noticed that there are fewer and fewer submissions on the Leaderboard of <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/369789\" target=\"_blank\">this competition</a>. He encourages competitors to keep working on their solutions and suggests they explore the following ideas to get better results: </p>\n<ul>\n<li>Make sure you are using algorithms that are tailored to the problem, such as the Minimum Spanning Tree algorithm.</li>\n<li>Try out different optimization techniques.</li>\n<li>Take a close look at your code and identify possible improvements.</li>\n<li>Use the insights and ideas of other Kagglers. </li>\n<li>Share your progress and ideas with others.</li>\n</ul>\n<hr>\n<h5>recommenders - Best Practices on Recommendation Systems by Microsoft (<a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a>)</h5>\n<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a> from Microsoft shared a great post discussing best practices when building Recommender Systems. The main points covered include:</p>\n<ul>\n<li>Data Preprocessing: including data cleaning, normalization, and feature engineering.</li>\n<li>Model Selection: including selecting the right model for the task, hyperparameter optimization and model validation.</li>\n<li>Evaluation Metrics: including the use of standard metrics (such as RMSE) as well as business metrics.</li>\n<li>Deployment &amp; Maintenance: including integrating the model into applications, model retraining and monitoring.</li>\n<li>User Experience: including recommendations personalization and explainability. <br>\nCheck out the full <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363730\" target=\"_blank\">post</a> for further details.</li>\n</ul>\n<hr>\n<h5>Why best model might not win RecSys competition (<a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">Roman</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">Roman</a> discussed the difference between offline and online metrics in Recommender System (RecSys) competitions.</li>\n<li>The ground truth labels for RecSys competitions are derived from the organizers' current model. This may mean that the best model might not win the competition, as the model which scores the highest on the offline metric might not be the same as the one which scores the highest on the online metric.</li>\n</ul>\n<hr>\n<h5>Important information regarding test data from competition repository (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has shared a helpful repository containing important information regarding the test data from the competition.</li>\n<li>The repository includes a comprehensive README as well as some related code. Check it out <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965\" target=\"_blank\">here</a>!</li>\n</ul>\n<hr>\n<h5>Otto's Talk on Transformer Recommendation Systems (<a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">BenediktSchifferer</a>)</h5>\n<p><a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">BenediktSchifferer</a> presented on the topic of Transformer Recommendation Systems in a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/373224\" target=\"_blank\">talk</a>.</p>\n<ul>\n<li>The main idea of the talk was to explain the use of transformers in recommendation systems and to discuss their potential applications.</li>\n<li>The talk discussed how transformers can be used to learn user preference and then use them to generate recommendations. It also discussed how transformers can be used to build an end-to-end recommendation pipeline, as well as how they can be used to improve existing models.</li>\n<li>The talk discussed how transformers can be used to solve cold-start and long-tail problems. The talk highlighted the advantages of using transformers, such as their capability to capture long-term and short-term user preferences, and their ability to handle both structured and unstructured data. It also discussed the challenges of using transformers, such as their large memory requirement and computational cost.</li>\n</ul>\n<hr>\n<h5>RePlay - opensource RecSys constructor (<a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a>)</h5>\n<p><a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a> has created an opensource RecSys constructor called <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365917\" target=\"_blank\">RePlay</a>. It is a library that provides tools for all stages of creating a recommendation system, such as data preprocessing, model evaluation, and comparison. RePlay uses PySpark to handle big data.</p>\n<hr>\n<h5>💡 ANN -&gt; NN: get better results and run faster with NN search on the GPU 🔥🔥🔥 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> recently shared an amazing post on how to get better results and run faster with NN search on the GPU.</li>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggested using Matrix Factorization with GPU which is 6.5x faster. Furthermore, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> mentioned that this solution was already implemented by their colleagues.</li>\n</ul>\n<hr>\n<h5>Ground-truth? (<a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a> asked a question about what \"test set is truncated\" means and if it is possible to create ground-truth for the training set by only one event.</li>\n<li>The post explains that it is not truncated by a single event, instead it is truncated by the total number of frames in the test set.</li>\n<li>Therefore, it is not possible to create the ground-truth by a single event, as the test set contains more frames than the training set.</li>\n</ul>\n<hr>\n<h5>The time zone in Germany is UTC+2 in August (<a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a> discussed the time zone in Germany being UTC+2 in August.</li>\n<li>This is due to Daylight Saving Time (DST) which is observed in Germany from the last Sunday in March until the last Sunday in October.</li>\n<li>During DST, the clocks are moved forward by one hour, resulting in the time zone shifting from UTC+1 to UTC+2. This means that the daylight hours will be longer in the summer months, with the sun setting later in the evening.</li>\n</ul>\n<hr>\n<h5>Ranker models vs. Binary classification models (<a href=\"https://www.kaggle.com/andrejzuba\" target=\"_blank\">Andrej Zubaľ</a>)</h5>\n<p><a href=\"https://www.kaggle.com/andrejzuba\" target=\"_blank\">Andrej Zubaľ</a> discussed the differences between ranker models and binary classification models in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370502\" target=\"_blank\">this post</a>.</p>\n<ul>\n<li><strong>Ranker models</strong> were found to be more accurate for predicting the probability of a certain event or outcome, as they can rank a list of outcomes from most to least likely.</li>\n<li><strong>Binary classification</strong> models are simpler and can be used to classify events into two categories, such as yes or no, true or false.</li>\n<li><strong>Ranker models</strong> are better suited for prediction tasks such as predicting the relevance of a search query, while binary classification models are better suited for categorizing events.</li>\n</ul>\n<hr>\n<h5>How do you train Ranking model? (<a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a>)</h5>\n<p><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a> asked a question about how to build a ranking model. In the post, they mentioned that they have tried various methods but with not much success.</p>\n<p><strong>This post provides useful information on how to train a ranking model, such as:</strong></p>\n<ul>\n<li>Understanding how Ranking works</li>\n<li>Normalising features</li>\n<li>Feature selection</li>\n<li>Data Augmentation</li>\n<li>Training and Tuning</li>\n<li>Evaluation Metric Selection</li>\n<li>Deployment and Monitoring</li>\n</ul>\n<hr>\n<h5>💡How to ensemble predictions -- a key component to every strong solution 🏅 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> in this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368747\" target=\"_blank\">post</a> discussed the importance of ensembling for a strong solution in Kaggle competitions. Ensembling is combining multiple models to get a better result than the individual models. </p>\n<p>The main idea is to create multiple models that are independent from each other and then combine them to get better results.<br>\nThe author suggests combining models that have different architectures, different pre-processing techniques and different hyperparameters.</p>\n<p><strong>The author also outlines some techniques for ensembling such as:</strong></p>\n<ul>\n<li>Averaging</li>\n<li>Weighted Averaging</li>\n<li>Model Stacking</li>\n<li>Bagging and Boosting</li>\n</ul>\n<p><strong>The author also provides some useful tips to keep in mind while ensembling such as:</strong></p>\n<ul>\n<li>Identify the best models and ensemble them</li>\n<li>Try using different weights for each model</li>\n<li>Use different seed values for each model</li>\n<li>Use different input data to train each model</li>\n<li>Use regularization to reduce overfitting</li>\n<li>Use cross-validation to evaluate the models.</li>\n</ul>\n<p>Ensembling is an important technique to improve the accuracy of models and can be used to achieve higher scores in Kaggle competitions.</p>\n<hr>\n<h5>Guidance on what algos to try from one of the authors of Microsoft Recommenders repo (<a href=\"https://www.kaggle.com/hoaphumanoid\" target=\"_blank\">Miguel Fierro</a>)</h5>\n<p><a href=\"https://www.kaggle.com/hoaphumanoid\" target=\"_blank\">Miguel Fierro</a> is one of the authors of the <a href=\"https://github.com/microsoft/recommenders/\" target=\"_blank\">Microsoft Recommenders repo</a>.</p>\n<ul>\n<li>In this post they provide guidance on which algorithms to try depending on the problem.</li>\n<li>They mention to start with simple algorithms like popularity, content-based filtering, and matrix factorization.</li>\n<li>They also suggest to look into user-based collaborative filtering and embeddings. Finally, they suggest to try deep learning methods such as deep matrix factorization, deep autoencoders, and convolutional neural networks.</li>\n</ul>\n<hr>\n<h5>Graph Neural Network, is complex to model but efficient for this task (<a href=\"https://www.kaggle.com/younesselbrag\" target=\"_blank\">Youness EL BRAG</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/younesselbrag\" target=\"_blank\">Youness EL BRAG</a> presented a new approach, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368355\" target=\"_blank\">Session-based Recommendation with Graph Neural Networks</a>, to the task of session-based recommendation.</li>\n<li>This method models a session as a complex transition of items and estimates user representations in addition to item representations. An accompanying <a href=\"https://github.com/CRIPAC-DIG/SR-GNN?utm_source=catalyzex.com\" target=\"_blank\">code</a> was also released which enables users to implement the method in their own projects.</li>\n</ul>\n<hr>\n<h5>cuDF vs Pandas Vs Modin for faster data processing and there comparison (<a href=\"https://www.kaggle.com/satyaprakashshukl\" target=\"_blank\">Satya</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/satyaprakashshukl\" target=\"_blank\">Satya</a> discussed different methods for faster data processing and their comparison and speed on this competition.</li>\n<li>Really useful post for those who are looking for faster data processing methods.</li>\n<li>Check out the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366258\" target=\"_blank\">post</a> for more information.</li>\n</ul>\n<hr>\n<h5>Some Interesting Times Series on Products (<a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a>)</h5>\n<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a> recently <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368373\" target=\"_blank\">shared</a> an interesting discussion topic about product times series.<br>\n<strong>The main idea behind it is to use the sales data of products to make predictions and forecast future sales.</strong></p>\n<p><strong>The post contains a few insights such as:</strong></p>\n<ul>\n<li>Times series can be used to forecast future sales more accurately. </li>\n<li>Seasonality can be taken into account to make more precise predictions.</li>\n<li>The need to use the most recent data available to make more reliable predictions.</li>\n<li>Machine Learning algorithms can be used to model times series data.</li>\n<li>Using multiple linear models can improve accuracy.</li>\n<li>Different methods of smoothing can be applied to times series data.</li>\n</ul>\n<hr>\n<h5>How can a session start with an order? (<a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a> posed a question about how to start a session with an order.</li>\n<li>He suggested that it could be done by creating a queue of tasks and using a scheduler to set the order of tasks and limit the number of tasks run in parallel.</li>\n<li>Additionally, they suggested the use of a job server that can manage the scheduling. Also, he discussed the use of a separate service for handling long-running tasks, such as background tasks.</li>\n</ul>\n<hr>\n<h5>Difference between Co-visitation Matrix v.s. Matrix Factorization (<a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">bilzard</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">bilzard</a> discussed the difference between the co-visitation matrix approach, which can be found <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725851\" target=\"_blank\">here</a>, and the matrix decomposition approach, which can be found <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725852\" target=\"_blank\">here</a>.</li>\n<li>The decisive difference between them is the dependence of events, which could be called the context in consideration. This could explain why the co-visitation matrix approach scores higher than the matrix decomposition approach.</li>\n</ul>\n<hr>\n<h5>Some concerns about validation (<a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a> noticed that most people were using the test set split code from the organizers' <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370384\" target=\"_blank\">repository</a></li>\n<li>But has some concerns that it might have flaws.</li>\n<li>He suggests using an alternative method to split the data, such as using a train/validation/test split, which would allow for better validation. This could help avoid overfitting and could potentially lead to better results. He also encourages people to think about the validation method and consider different approaches.</li>\n</ul>\n<hr>\n<h5>We can try more NLP methods in sequence recomendation, like …. (<a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">KKY</a>)</h5>\n<ul>\n<li>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367870\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">KKY</a> discussed the possibility of using NLP methods to improve sequence recommendation in this user-anonymized dataset.</li>\n<li>They suggested trying out various natural language processing (NLP) techniques to get a better representation of the action/item sequence.</li>\n<li>Some ideas they proposed include using Long Short Term Memory (LSTM) networks, convolutional neural networks (CNNs), and transformer networks.</li>\n<li>Additionally, they suggested using methods such as attention mechanisms, bidirectional networks, and sequence-to-sequence learning.</li>\n</ul>\n<hr>\n<h5>Summary about Loading and Preprocessing Big Jsonl Data File (<a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">Lei Wang</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">Lei Wang</a> shared two notebooks to help with loading and preprocessing data for this competition.</li>\n<li><strong>Loading:</strong> The notebook explains how to load big jsonl data file into a pandas dataframe with different methods and also provides a comparison between them. </li>\n<li><strong>Preprocessing:</strong> This notebook explains how to convert the data into csv, parquet or a dataframe, as well as the advantages and disadvantages of each.</li>\n</ul>\n<hr>\n<h5>Since there are only few features(only time info), any chance to use ML algorithm? (<a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">cocoshe</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">cocoshe</a> recently posted a discussion about the possibility of using a Machine Learning (ML) algorithm for this competition.</li>\n<li>They observed that all the recall methods currently used involve \"top-k frequency aids of different session\" and \"top-k frequency aids of different type\", or a \"top-k frequency of all aids\" counting method (co-occurrence matrix). Additionally, padding is used to fill the recall@20.</li>\n<li><a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">cocoshe</a> then asked if any chance to use ML algorithm existed, despite the few features that are available.</li>\n</ul>\n<hr>\n<h5>Q about the Data (<a href=\"https://www.kaggle.com/avieldanin\" target=\"_blank\">Aviel Danin</a>)</h5>\n<p><a href=\"https://www.kaggle.com/avieldanin\" target=\"_blank\">Aviel Danin</a> asked a question about the data in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366478\" target=\"_blank\">this post</a>:</p>\n<ul>\n<li>In some cases there is a \"cart event\" of a certain aid followed by a \"click\", and then another \"cart event\" of the same aid followed by another \"click\". </li>\n<li>In other cases, there is a \"cart event\" without a click before it. </li>\n</ul>\n<p>These scenarios could indicate that there are multiple processes related to the data, and it's important to understand them in order to get the most accurate results.</p>\n<p><strong>Answer (By (<a href=\"https://www.kaggle.com/danieleroncaglioni\" target=\"_blank\">roncadr</a>):</strong></p>\n<pre><code>Yeah I noticed the same, also in the full train dataset there are 16.896.191 carts, while there are less (12.142.933) instances of an aid of some type being followed by the same aid and type \"cart\", so definetely, apparently it is possible to add an item to the cart without having to click that item immediately beforehand. Perhaps when you are browsing on multiple browser tabs and switch between them….?\n</code></pre>\n<hr>\n<h5>💡 How to split the data for training a two-stage recommender? (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discussed the best way to split data when training a two-stage recommender system.</li>\n<li>He shared his work on <a href=\"https://www.kaggle.com/code/radek1/eda-a-look-at-the-data-training-splits\" target=\"_blank\">this</a> notebook for anyone interested in learning more about the topic.</li>\n</ul>\n<hr>\n<h5>How did you read the data? (<a href=\"https://www.kaggle.com/takuma0306\" target=\"_blank\">takuma0306</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/372990\" target=\"_blank\">post</a> <a href=\"https://www.kaggle.com/takuma0306\" target=\"_blank\">takuma0306</a> asked how others were loading the training data for the competition.<br>\nThey mentioned that when using VS Code, their application was terminating halfway through and when using Google Colab they were running out of system RAM and losing the session. They were wondering how others were successfully loading the data and asked for help.</p>\n<p><strong>Answer (By (<a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">Giannis</a>)):</strong></p>\n<pre><code>As an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n- https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\n</code></pre>\n<hr>\n<h5>Test Set Clarification Questions (<a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">dpalbrecht</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">dpalbrecht</a> raised some questions regarding the test set on the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365554\" target=\"_blank\">post</a> in the repo.</li>\n<li>The sentence that confused them was: \"The test set contains 1000 user queries and the ground truth is not known.\"</li>\n</ul>\n<p><strong>The questions raised were:</strong></p>\n<ul>\n<li>How the queries are produced?</li>\n<li>What does \"ground truth\" refer to?</li>\n<li>Does it mean that for each user query there is a corresponding ground truth image?</li>\n</ul>\n<p><strong>Answer (By <a href=\"https://www.kaggle.com/pietromaldini1\" target=\"_blank\">Pietro Maldini</a>):</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2F957ed2d68815e3db8abef7b03304a135%2Fground_truth.png?generation=1668199923227677&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h5>The relatively less data for each test session (<a href=\"https://www.kaggle.com/samsonfha\" target=\"_blank\">samson fha</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/samsonfha\" target=\"_blank\">samson fha</a> discussed the issue of the relatively low amount of data for each test session compared to training sessions, noting that it is difficult to identify user features and generate labels for them.</li>\n<li>They suggested focusing more on generating labels for goods/aids instead.</li>\n</ul>\n<p><strong>Answer (By <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</strong><br>\nThe reason for the shorter sessions is explained <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486\" target=\"_blank\">here</a> by the organizer.</p>\n<hr>\n<h5>How many candidates you used? (<a href=\"https://www.kaggle.com/baekseungyun\" target=\"_blank\">L0Z1K</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/baekseungyun\" target=\"_blank\">L0Z1K</a> suggests that sharing the number of candidates and recall score can help the community in their analysis of the Santa 2020 competition.</li>\n<li>This could provide valuable insights into the effectiveness of certain strategies and help to identify potential improvements.</li>\n</ul>\n<hr>\n<h5>Unseen Aids in Test Set (<a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">moth</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">moth</a> made an interesting observation in their <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156\" target=\"_blank\">post</a> about the test set for the competition.</li>\n<li>They noticed that many items in the test set do not appear in the train set, which can make it difficult to accurately predict the outcomes.</li>\n<li><strong>This is an important reminder to be mindful of the unseen data that may exist in the test set when creating models.</strong></li>\n</ul>\n<hr>\n<h5>Otto recommender Kernel Stats (<a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a> shared an amazing post about notebook analytics!</li>\n<li>Such as the number of upvotes or a fork graph.</li>\n</ul>\n<p><strong>Really cool idea!</strong><br>\nmount of data and also reduces the amount of time needed to train a model. Additionally, it can also reduce the likelihood of overfitting due to the smaller data size.</p>",
  "messages": [
    {
      "id": "2076039",
      "postDate": "12/26/2022 04:42:06",
      "content": "<h3>Here is what you need to know</h3>\n<h5>One month to go!</h5>\n<hr>\n<h5>Recommendation Systems for Large Datasets (<a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a>)</h5>\n<p><a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a> discussed the use of recommendation systems for large datasets, such as the one used for this competition. The train dataset consists of: </p>\n<ul>\n<li>12,899,779 sessions</li>\n<li>1,855,603 items</li>\n<li>216,716,096 events</li>\n<li>194,720,954 clicks</li>\n<li>16,896,191 carts</li>\n<li>5,098,951 orders</li>\n</ul>\n<p>He expand on this and add information about the stages for solving this problem:<br>\n<strong>Candidate Generation</strong></p>\n<p>Example criteria you can use to select you candidates:</p>\n<ul>\n<li>previously purchased items</li>\n<li>repurchased items</li>\n<li>overall most popular items</li>\n<li>similar items based on some sort of clustering technique</li>\n<li>similar items based on something such as a co-visitation matrix</li>\n</ul>\n<p>By now we should have much fewer items for each session, so we should be able to input these into a ranker model.</p>\n<p><strong>Ranking Model Examples</strong></p>\n<ul>\n<li><a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html\" target=\"_blank\">LGBMRanker</a></li>\n<li><a href=\"https://medium.com/predictly-on-tech/learning-to-rank-using-xgboost-83de0166229d\" target=\"_blank\">XGBRanker</a></li>\n<li><a href=\"https://towardsdatascience.com/learning-to-rank-with-python-scikit-learn-327a5cfd81f\" target=\"_blank\">Ranking with sklearn</a></li>\n<li><a href=\"https://maroo.cs.umass.edu/getpdf.php?id=1373\" target=\"_blank\">Neural Network Ranker</a></li>\n</ul>\n<hr>\n<h5>How To Build a GBT Ranker Model (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> provided an easy way to build a solution for the competition by creating candidate rerank models. He explained how to create train data, train and infer a GBT ranker model.</p>\n<hr>\n<h5>local validation tracks public LB perfecty -- here is the setup (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has found that local validation tracks public leaderboard (LB) performance perfectly. They have shared their setup <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a> to allow others to replicate their results. The setup includes a scoring function that uses the same public LB metric and a test set of 500 images that corresponds to the public LB.</p>\n<hr>\n<h5>💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> Shares an <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">EDA</a> notebook, some [dataset] he preprocessed to parquet and a <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">robust validation framework</a>.</p>\n<hr>\n<h5>30x Faster Co-Visitation Matrices using RAPIDS cuDF! (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> published a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369\" target=\"_blank\">notebook</a> which uses RAPIDS cuDF to compute co-visitation matrices 30x faster. These matrices help to provide models with \"candidates\" which can then be reranked and selected for submission CSV. This process is useful for improving prediction accuracy.</p>\n<hr>\n<h5>Surprising LB 0.587 !!!Share some experimental results ~~ (<a href=\"https://www.kaggle.com/jiahongxie\" target=\"_blank\">Jiahong Xie</a>)</h5>\n<p><a href=\"https://www.kaggle.com/jiahongxie\" target=\"_blank\">Jiahong Xie</a> has made an amazing discovery with a surprising LB score of 0.587! They shared their experimental results which showed that they had a recall rate of 200 candidates for each user. This was an impressive result and could provide valuable insights for other Kagglers.</p>\n<p>In short: And generate 300+ Features for each user and item pair.In training stage and downsample positive:negative as 1:20 for training. Also small improvement was made by deleting aid col in features.</p>\n<hr>\n<h5>📈 What do we know so far? ⚡Summary with  links to relevant resources (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> created an incredibly useful post which contains links to relevant resources related to the Santa 2022 competition. He also shared a repo on GitHub which contains data for the competition, including preprocessing code and information not available on Kaggle. This repo can be accessed <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">here</a>.</p>\n<hr>\n<h5>📑 [Step-by-step guide] How I got to my current standing on the LB and how to improve going forward (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> shared <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278\" target=\"_blank\">this post</a> on how he achieved his current standing on the leaderboard and how to improve it going forward.</p>\n<ul>\n<li>The main idea is to use a <strong>divide-and-conquer approach to search for small improvements</strong>, by breaking down the problem into smaller pieces and looking for more efficient solutions.</li>\n<li>He advises to use a combination of different methods (random search, minimum spanning tree, etc.), as well as to take into account the constraints of the problem, such as the maximum link length.</li>\n<li>He also suggests revisiting points, as well as to experiment with different strategies for finding the optimum solution. Lastly, he recommends to use the baseline functions provided but with modifications to make them faster and more efficient.</li>\n</ul>\n<hr>\n<h5>Full dataset processed to CSV/parquet files with optimized memory footprint (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> shared a post discussing how to process the full dataset to CSV/parquet files with an optimized memory footprint.</p>\n<ul>\n<li>The main idea is to use dask to read the raw data, filter it, and then save it to parquet files using the dask.dataframe.to_parquet method.</li>\n<li>The resulting files are much smaller than the original and can be loaded into memory much faster. Additionally, the post provides a code example of how to use dask to process the data.</li>\n</ul>\n<hr>\n<h5>Locate Real Users and Real Sessions EDA (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> discussed the use of the term \"session\" in Kaggle's Otto competition in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138\" target=\"_blank\">this post</a>.</p>\n<ul>\n<li>In the competition, \"session\" actually means \"user\". We are given train data for 12,899,779 users over a 4 week period, and test data for users during 1 week (in the future).</li>\n<li>We must predict what a user will do in the remainder of the 1 week that we do not have information about.</li>\n</ul>\n<hr>\n<h5>There are some bugs in evaluation (<a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">Tawara</a>)</h5>\n<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">Tawara</a> discovered a bug in the evaluation which resulted in a score of 4.848 on the competition.</p>\n<ul>\n<li>He shared the <a href=\"https://www.kaggle.com/code/ttahara/last-aid-20/\" target=\"_blank\">notebook</a> he used to achieve this score.</li>\n<li>It is important to keep in mind that the bug may have been fixed and the score may have changed.</li>\n</ul>\n<hr>\n<h5>Andrew Ng Recommender Systems (<a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">Andrada Olteanu</a>)</h5>\n<p><a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">Andrada Olteanu</a> discussed the fundamentals of Recommender Systems as outlined by Andrew Ng in his course on Machine Learning.<br>\n <strong>The main topics covered include:</strong></p>\n<ul>\n<li>Collaborative Filtering and its two main approaches: User-User and Item-Item.</li>\n<li>Content-Based Filtering which uses the attributes of the items being recommended.</li>\n<li>Hybrid Methods which combine the two approaches.</li>\n<li>Matrix Factorization which is used to classify users and items into latent features.</li>\n<li>Evaluation of Recommender Systems which includes methods such as Mean Average Precision (MAP) and Root Mean Squared Error (RMSE).</li>\n<li>Challenges of Recommender Systems such as scalability and cold-start problem.</li>\n</ul>\n<hr>\n<h5>How to thrive in this competition without going crazy -- 1 out of 2 important truths ❤️‍🔥 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has shared <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503\" target=\"_blank\">this post</a> as a cautionary reminder of the complexity of this competition.</p>\n<ul>\n<li>He has highlighted two important truths to keep in mind in order to thrive:</li>\n<li><strong>Complexity:</strong> Once you start to dig deeper, the competition becomes increasingly complex.</li>\n<li><strong>Mindset:</strong> Having the right mindset is crucial to success. Be patient and optimistic, while also being mindful of the complexity.</li>\n</ul>\n<hr>\n<h5>The Story of User 13479136 and User 13710374 (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p>In the Kaggle's Otto competition, two particular users stood out:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678\" target=\"_blank\">User 13479136</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678\" target=\"_blank\">User 13710374</a></li>\n<li><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> managed to extract item categories using RAPIDS TSNE such different categories of items may be clothing and electronics.</li>\n<li>He then continue to deconstruct the individual user behaviour in the data.</li>\n</ul>\n<p><strong>Very interesting read!</strong></p>\n<hr>\n<h5>Yes, We Can Use Test Data Leakage. (<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a>)</h5>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> discussed the possibility of using test data leakage in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939\" target=\"_blank\">this post</a>.</p>\n<ul>\n<li>Since we are predicting test events occurring within a one week period, it is possible to use the test data to predict future events.</li>\n<li>This could be done by using a rolling window approach, where the training data is used to predict the test data in the following steps.</li>\n<li>The test data can then be used to update the model and predict the next step.</li>\n<li>This approach is beneficial as it allows for more accurate predictions in the future and could be used to improve the model's performance.</li>\n</ul>\n<hr>\n<h5>Say NO to GPU: Nx Faster Co-Visitation Matrices using single CPU! (<a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">Carno Zhao</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">Carno Zhao</a> published a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365873\" target=\"_blank\">notebook</a> which uses <code>numba.jit</code> to compute Nx faster co-visitation matrices.</li>\n<li>This is an incredibly useful tool to make computations faster and more efficient.</li>\n</ul>\n<hr>\n<h5>💡 What is the co-visitation matrix, really? (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discussed a concept known as the co-visitation matrix in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358\" target=\"_blank\">this post</a> and how it relates to modern techniques. </p>\n<p><strong>Interesting starter reading for this competition</strong></p>\n<hr>\n<h5>A top-down perspective on the current metric values (<a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">narsil</a>)</h5>\n<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">narsil</a> shared a useful post discussing the current metric values from a top-down perspective.</p>\n<p><strong>Sessions in the test set are truncated in a random moment of the session life.</strong></p>\n<ul>\n<li>For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart &amp; purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.</li>\n<li>For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat (see the winning scores for H&amp;M competition - they are close to 0). This is different metric, but shows how difficult such a problem is.</li>\n</ul>\n<p><strong>Some intuitions:</strong></p>\n<ul>\n<li>If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard</li>\n<li>It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.</li>\n</ul>\n<hr>\n<h5>How to train a Word2Vec model for item embeddings - a simple code example 📖 (<a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a>)</h5>\n<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a> shared a useful post on how to train a Word2Vec model to create item embeddings.<br>\nThis includes a simple code example and detailed explanation of the process.</p>\n<pre><code> pandas  pd\n gensim.models  Word2Vec\ncs = \ncount = \ntotal = \n (, )  f:\n     df  pd.read_json(, lines=, chunksize=cs):\n         val  df[].apply( x: [event[]  event  x]).values:\n            f.write(.join((, val)) + )\n        count += cs\n        (, end=)\nmodel = Word2Vec(corpus_file=, vector_size=, window=, min_count=, workers=)\nmodel.save()\n</code></pre>\n<p>And after training the model we can use it like this:</p>\n<pre><code>model = Word2Vec.load()\nmodel.wv.most_similar(, topn=)\n</code></pre>\n<hr>\n<h5>[Starter pack] LGBMRanker with polars 🚀🚀🚀 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>, he explain how re-ranking models are the industry standard for dealing with datasets like those in the Otto Recommender System competition, which have high cardinality categories.</li>\n<li>He also show how to use LightGBM's ranker to create a baseline model that is competitive with the leaderboard. The post also provides a starter pack to help people getstarted, with instructions on how to use the ranker, what the parameters mean and how to tune them, as well as how to use the model for recommendations. Finally, the post also provides a notebook to help users to get their own model up and running.</li>\n</ul>\n<hr>\n<h5>[Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> shared a great <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893\" target=\"_blank\">post</a> on improving candidate generation using a matrix factorization model!</p>\n<ul>\n<li>This post provides a starter pack with a Pytorch implementation and a Merlin Dataloader to get started. With this starter pack, you can load data, define the model, and start training.</li>\n<li>This will allow you to quickly get up and running with matrix factorization and improve candidate generation.</li>\n</ul>\n<hr>\n<h5>Saving GPU memory when processing features == (Speedup + Efficiency) (<a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">Giba</a>)</h5>\n<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">Giba</a> discussed a new functionality of Cudf that can help save GPU memory when processing features.</p>\n<ul>\n<li>By setting the default data type to 32-bit instead of 64-bit, the memory taken up by the GPU can be reduced. The code for this can be found below:</li>\n</ul>\n<pre><code> cudf\ncudf.set_allocator(, default_dtype=np.float32)\n</code></pre>\n<hr>\n<h5>I think word2vec is a nice choice. (<a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">taku_sid</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">taku_sid</a> shared a post discussing the use of word2vec to generate candidates.</li>\n<li>Managed to get to a score of lb 0.578 which is great for word2vec.</li>\n<li>Some of the comments discussions includes important information about differt models used on this competition.</li>\n</ul>\n<p><strong>Recommended!</strong></p>\n<hr>\n<h5>Processed dataset (<a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">Konrad Banachewicz</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">Konrad Banachewicz</a> created a post <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363617\" target=\"_blank\">here</a> providing a training dataset in both .csv and .parquet formats for those who love pandas.</li>\n<li>This can be a useful resource for anyone looking to quickly get started with the dataset.</li>\n</ul>\n<hr>\n<h5>💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition) (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>, he discussed several strategies for working on ML projects more efficiently.</p>\n<p><strong>He include:</strong></p>\n<ul>\n<li>Breaking projects into smaller tasks and setting up checkpoints</li>\n<li>Using domain knowledge to reduce the size of the search space</li>\n<li>Leveraging existing frameworks, libraries, and tools</li>\n<li>Looking for pre-existing solutions to problems</li>\n<li>Focusing on one task at a time</li>\n<li>Streamlining the data collection and labelling process</li>\n<li>Setting up automated evaluations and feedback loops</li>\n<li>Taking breaks to reflect and reassess goals</li>\n<li>Working in collaboration with others</li>\n</ul>\n<hr>\n<h5>Which metric is correct? (<a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a>)</h5>\n<p><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a> raised a question on the evaluation metric for this competition: Recall@k.</p>\n<p><strong><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a> asked which of the following two definitions is correct:</strong></p>\n<ul>\n<li>(1) Recall@k is the fraction of the top k pixels of an image that are the same as the target image.</li>\n<li>(2) Recall@k is the fraction of the target image that is in the top k pixels of an image.</li>\n</ul>\n<p>The discussion that followed suggested that the <strong>second definition</strong> is the correct one.</p>\n<hr>\n<h5>💡How to improve the results of your Approximate Nearest Neighbor search! (annoy) (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> gave an insightful post on <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368385\" target=\"_blank\">how to improve the results of your Approximate Nearest Neighbor search</a>!</p>\n<p><strong>Main Idea:</strong> Use an Approximate Nearest Neighbor (ANN) search to improve the results of your search.</p>\n<p><strong>Benefits:</strong></p>\n<ul>\n<li>ANN searches are much faster than traditional nearest neighbor searches and can be used to quickly find similar items.</li>\n<li>ANN searches can also be used to find items with similar characteristics, even if they don’t have an exact match.</li>\n<li>ANNs are also more scalable, as they can be easily parallelized and distributed across multiple nodes.</li>\n</ul>\n<p><strong>Implementation:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggests using the <a href=\"https://github.com/spotify/annoy\" target=\"_blank\">Annoy library</a> to implement the ANN search.</li>\n<li>Annoy is a Python library that allows you to quickly index and search for points in a vector space.</li>\n<li>Annoy also includes a few useful features such as the ability to easily control the size of the search pool, as well as the ability to specify the number of neighbors to search for.</li>\n</ul>\n<p><strong>Conclusion:</strong></p>\n<p>Using an Approximate Nearest Neighbor search can provide a number of benefits, such as increased search speed and scalability. <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggests using the Annoy library to implement the ANN search, which provides a number of useful features for controlling the size of the search pool and specifying the number of neighbors to search for.</p>\n<hr>\n<h5>💡 Training an XGBoost Ranker on the GPU with Merlin Models 🔥🔥🔥 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368848\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discussed the use of Merlin Models to train an XGBoost Ranker on the GPU.<br>\nThis can greatly reduce the training time for XGBoost models, allowing for faster and more efficient training. The main benefits are that it does not require a Tensorflow backend, and allows for faster training. It also allows for better hyperparameter optimization, since it is possible to train multiple models in parallel. Additionally, the model can be deployed to the cloud for scalability and higher performance. Finally, the model can also be used for inference tasks such as ranking and recommendation.</p>\n<hr>\n<h5>Matrix Factorization with GPU: 6.5x faster! (<a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a> shared a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/371166\" target=\"_blank\">very interesting notebook</a> to compute matrix factorization, using polar, annoy, Merlin data loader and pytorch.</li>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a> then created a GPU version of the notebook which is <a href=\"https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu\" target=\"_blank\">available here</a> and is 6.5x faster than the original version!</li>\n</ul>\n<hr>\n<h5>🐘 the elephant in the room -- high cardinality of targets and what to do about this (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discusses an issue that is often overlooked when working with Machine Learning models, high cardinality of targets and how to handle it. The main idea presented is that, in order to build an effective model, it is important to accurately measure the importance of each feature. The author suggests two main approaches: </p>\n<ul>\n<li><strong>Feature Engineering:</strong> transforming existing features into a different representation that can be used to better understand the underlying data. </li>\n<li><strong>Dimensionality Reduction:</strong> reducing the number of features to a more manageable amount.</li>\n</ul>\n<p>The author also outlines some potential pitfalls, such as overfitting and underfitting, that can arise when dealing with high cardinality. Finally, the post provides some tips on how to effectively use the two main approaches to tackle this problem.</p>\n<hr>\n<h5>Co-Visitation Matrices and Matrix Factorization (<a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a>)</h5>\n<p><a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a> discussed the use of co-visitation matrices and matrix factorization for recommendation systems in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589\" target=\"_blank\">this post</a>. The main idea is to create features based on items which are frequently viewed together. This technique has been shown to be effective for creating better recommendation systems.</p>\n<hr>\n<h5>💡What is a good initial goal in the competition? How to improve beyond it? 📈 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368685\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggested a good initial goal for the competition is to get everything up and running on your end. This includes setting up your environment and running the notebooks. Additionally, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggested that to improve beyond this goal, it is important to understand the underlying techniques and concepts, such as using the right data structure and algorithm for the task, as well as understanding the limitations of the algorithms. Finally, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> recommended debugging your code and understanding the underlying algorithms, such as a greedy approach, to improve beyond the initial goal.</p>\n<hr>\n<h5>To the people still hacking away on this competition… (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has noticed that there are fewer and fewer submissions on the Leaderboard of <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/369789\" target=\"_blank\">this competition</a>. He encourages competitors to keep working on their solutions and suggests they explore the following ideas to get better results: </p>\n<ul>\n<li>Make sure you are using algorithms that are tailored to the problem, such as the Minimum Spanning Tree algorithm.</li>\n<li>Try out different optimization techniques.</li>\n<li>Take a close look at your code and identify possible improvements.</li>\n<li>Use the insights and ideas of other Kagglers. </li>\n<li>Share your progress and ideas with others.</li>\n</ul>\n<hr>\n<h5>recommenders - Best Practices on Recommendation Systems by Microsoft (<a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a>)</h5>\n<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">Sinan Calisir</a> from Microsoft shared a great post discussing best practices when building Recommender Systems. The main points covered include:</p>\n<ul>\n<li>Data Preprocessing: including data cleaning, normalization, and feature engineering.</li>\n<li>Model Selection: including selecting the right model for the task, hyperparameter optimization and model validation.</li>\n<li>Evaluation Metrics: including the use of standard metrics (such as RMSE) as well as business metrics.</li>\n<li>Deployment &amp; Maintenance: including integrating the model into applications, model retraining and monitoring.</li>\n<li>User Experience: including recommendations personalization and explainability. <br>\nCheck out the full <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363730\" target=\"_blank\">post</a> for further details.</li>\n</ul>\n<hr>\n<h5>Why best model might not win RecSys competition (<a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">Roman</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">Roman</a> discussed the difference between offline and online metrics in Recommender System (RecSys) competitions.</li>\n<li>The ground truth labels for RecSys competitions are derived from the organizers' current model. This may mean that the best model might not win the competition, as the model which scores the highest on the offline metric might not be the same as the one which scores the highest on the online metric.</li>\n</ul>\n<hr>\n<h5>Important information regarding test data from competition repository (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> has shared a helpful repository containing important information regarding the test data from the competition.</li>\n<li>The repository includes a comprehensive README as well as some related code. Check it out <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965\" target=\"_blank\">here</a>!</li>\n</ul>\n<hr>\n<h5>Otto's Talk on Transformer Recommendation Systems (<a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">BenediktSchifferer</a>)</h5>\n<p><a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">BenediktSchifferer</a> presented on the topic of Transformer Recommendation Systems in a <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/373224\" target=\"_blank\">talk</a>.</p>\n<ul>\n<li>The main idea of the talk was to explain the use of transformers in recommendation systems and to discuss their potential applications.</li>\n<li>The talk discussed how transformers can be used to learn user preference and then use them to generate recommendations. It also discussed how transformers can be used to build an end-to-end recommendation pipeline, as well as how they can be used to improve existing models.</li>\n<li>The talk discussed how transformers can be used to solve cold-start and long-tail problems. The talk highlighted the advantages of using transformers, such as their capability to capture long-term and short-term user preferences, and their ability to handle both structured and unstructured data. It also discussed the challenges of using transformers, such as their large memory requirement and computational cost.</li>\n</ul>\n<hr>\n<h5>RePlay - opensource RecSys constructor (<a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a>)</h5>\n<p><a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a> has created an opensource RecSys constructor called <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365917\" target=\"_blank\">RePlay</a>. It is a library that provides tools for all stages of creating a recommendation system, such as data preprocessing, model evaluation, and comparison. RePlay uses PySpark to handle big data.</p>\n<hr>\n<h5>💡 ANN -&gt; NN: get better results and run faster with NN search on the GPU 🔥🔥🔥 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> recently shared an amazing post on how to get better results and run faster with NN search on the GPU.</li>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> suggested using Matrix Factorization with GPU which is 6.5x faster. Furthermore, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> mentioned that this solution was already implemented by their colleagues.</li>\n</ul>\n<hr>\n<h5>Ground-truth? (<a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a> asked a question about what \"test set is truncated\" means and if it is possible to create ground-truth for the training set by only one event.</li>\n<li>The post explains that it is not truncated by a single event, instead it is truncated by the total number of frames in the test set.</li>\n<li>Therefore, it is not possible to create the ground-truth by a single event, as the test set contains more frames than the training set.</li>\n</ul>\n<hr>\n<h5>The time zone in Germany is UTC+2 in August (<a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a> discussed the time zone in Germany being UTC+2 in August.</li>\n<li>This is due to Daylight Saving Time (DST) which is observed in Germany from the last Sunday in March until the last Sunday in October.</li>\n<li>During DST, the clocks are moved forward by one hour, resulting in the time zone shifting from UTC+1 to UTC+2. This means that the daylight hours will be longer in the summer months, with the sun setting later in the evening.</li>\n</ul>\n<hr>\n<h5>Ranker models vs. Binary classification models (<a href=\"https://www.kaggle.com/andrejzuba\" target=\"_blank\">Andrej Zubaľ</a>)</h5>\n<p><a href=\"https://www.kaggle.com/andrejzuba\" target=\"_blank\">Andrej Zubaľ</a> discussed the differences between ranker models and binary classification models in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370502\" target=\"_blank\">this post</a>.</p>\n<ul>\n<li><strong>Ranker models</strong> were found to be more accurate for predicting the probability of a certain event or outcome, as they can rank a list of outcomes from most to least likely.</li>\n<li><strong>Binary classification</strong> models are simpler and can be used to classify events into two categories, such as yes or no, true or false.</li>\n<li><strong>Ranker models</strong> are better suited for prediction tasks such as predicting the relevance of a search query, while binary classification models are better suited for categorizing events.</li>\n</ul>\n<hr>\n<h5>How do you train Ranking model? (<a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a>)</h5>\n<p><a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">dehokanta</a> asked a question about how to build a ranking model. In the post, they mentioned that they have tried various methods but with not much success.</p>\n<p><strong>This post provides useful information on how to train a ranking model, such as:</strong></p>\n<ul>\n<li>Understanding how Ranking works</li>\n<li>Normalising features</li>\n<li>Feature selection</li>\n<li>Data Augmentation</li>\n<li>Training and Tuning</li>\n<li>Evaluation Metric Selection</li>\n<li>Deployment and Monitoring</li>\n</ul>\n<hr>\n<h5>💡How to ensemble predictions -- a key component to every strong solution 🏅 (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> in this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368747\" target=\"_blank\">post</a> discussed the importance of ensembling for a strong solution in Kaggle competitions. Ensembling is combining multiple models to get a better result than the individual models. </p>\n<p>The main idea is to create multiple models that are independent from each other and then combine them to get better results.<br>\nThe author suggests combining models that have different architectures, different pre-processing techniques and different hyperparameters.</p>\n<p><strong>The author also outlines some techniques for ensembling such as:</strong></p>\n<ul>\n<li>Averaging</li>\n<li>Weighted Averaging</li>\n<li>Model Stacking</li>\n<li>Bagging and Boosting</li>\n</ul>\n<p><strong>The author also provides some useful tips to keep in mind while ensembling such as:</strong></p>\n<ul>\n<li>Identify the best models and ensemble them</li>\n<li>Try using different weights for each model</li>\n<li>Use different seed values for each model</li>\n<li>Use different input data to train each model</li>\n<li>Use regularization to reduce overfitting</li>\n<li>Use cross-validation to evaluate the models.</li>\n</ul>\n<p>Ensembling is an important technique to improve the accuracy of models and can be used to achieve higher scores in Kaggle competitions.</p>\n<hr>\n<h5>Guidance on what algos to try from one of the authors of Microsoft Recommenders repo (<a href=\"https://www.kaggle.com/hoaphumanoid\" target=\"_blank\">Miguel Fierro</a>)</h5>\n<p><a href=\"https://www.kaggle.com/hoaphumanoid\" target=\"_blank\">Miguel Fierro</a> is one of the authors of the <a href=\"https://github.com/microsoft/recommenders/\" target=\"_blank\">Microsoft Recommenders repo</a>.</p>\n<ul>\n<li>In this post they provide guidance on which algorithms to try depending on the problem.</li>\n<li>They mention to start with simple algorithms like popularity, content-based filtering, and matrix factorization.</li>\n<li>They also suggest to look into user-based collaborative filtering and embeddings. Finally, they suggest to try deep learning methods such as deep matrix factorization, deep autoencoders, and convolutional neural networks.</li>\n</ul>\n<hr>\n<h5>Graph Neural Network, is complex to model but efficient for this task (<a href=\"https://www.kaggle.com/younesselbrag\" target=\"_blank\">Youness EL BRAG</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/younesselbrag\" target=\"_blank\">Youness EL BRAG</a> presented a new approach, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368355\" target=\"_blank\">Session-based Recommendation with Graph Neural Networks</a>, to the task of session-based recommendation.</li>\n<li>This method models a session as a complex transition of items and estimates user representations in addition to item representations. An accompanying <a href=\"https://github.com/CRIPAC-DIG/SR-GNN?utm_source=catalyzex.com\" target=\"_blank\">code</a> was also released which enables users to implement the method in their own projects.</li>\n</ul>\n<hr>\n<h5>cuDF vs Pandas Vs Modin for faster data processing and there comparison (<a href=\"https://www.kaggle.com/satyaprakashshukl\" target=\"_blank\">Satya</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/satyaprakashshukl\" target=\"_blank\">Satya</a> discussed different methods for faster data processing and their comparison and speed on this competition.</li>\n<li>Really useful post for those who are looking for faster data processing methods.</li>\n<li>Check out the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366258\" target=\"_blank\">post</a> for more information.</li>\n</ul>\n<hr>\n<h5>Some Interesting Times Series on Products (<a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a>)</h5>\n<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">aldparis</a> recently <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368373\" target=\"_blank\">shared</a> an interesting discussion topic about product times series.<br>\n<strong>The main idea behind it is to use the sales data of products to make predictions and forecast future sales.</strong></p>\n<p><strong>The post contains a few insights such as:</strong></p>\n<ul>\n<li>Times series can be used to forecast future sales more accurately. </li>\n<li>Seasonality can be taken into account to make more precise predictions.</li>\n<li>The need to use the most recent data available to make more reliable predictions.</li>\n<li>Machine Learning algorithms can be used to model times series data.</li>\n<li>Using multiple linear models can improve accuracy.</li>\n<li>Different methods of smoothing can be applied to times series data.</li>\n</ul>\n<hr>\n<h5>How can a session start with an order? (<a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">CPMP</a> posed a question about how to start a session with an order.</li>\n<li>He suggested that it could be done by creating a queue of tasks and using a scheduler to set the order of tasks and limit the number of tasks run in parallel.</li>\n<li>Additionally, they suggested the use of a job server that can manage the scheduling. Also, he discussed the use of a separate service for handling long-running tasks, such as background tasks.</li>\n</ul>\n<hr>\n<h5>Difference between Co-visitation Matrix v.s. Matrix Factorization (<a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">bilzard</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">bilzard</a> discussed the difference between the co-visitation matrix approach, which can be found <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725851\" target=\"_blank\">here</a>, and the matrix decomposition approach, which can be found <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725852\" target=\"_blank\">here</a>.</li>\n<li>The decisive difference between them is the dependence of events, which could be called the context in consideration. This could explain why the co-visitation matrix approach scores higher than the matrix decomposition approach.</li>\n</ul>\n<hr>\n<h5>Some concerns about validation (<a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">Gunes Evitan</a> noticed that most people were using the test set split code from the organizers' <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370384\" target=\"_blank\">repository</a></li>\n<li>But has some concerns that it might have flaws.</li>\n<li>He suggests using an alternative method to split the data, such as using a train/validation/test split, which would allow for better validation. This could help avoid overfitting and could potentially lead to better results. He also encourages people to think about the validation method and consider different approaches.</li>\n</ul>\n<hr>\n<h5>We can try more NLP methods in sequence recomendation, like …. (<a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">KKY</a>)</h5>\n<ul>\n<li>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/367870\" target=\"_blank\">post</a>, <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">KKY</a> discussed the possibility of using NLP methods to improve sequence recommendation in this user-anonymized dataset.</li>\n<li>They suggested trying out various natural language processing (NLP) techniques to get a better representation of the action/item sequence.</li>\n<li>Some ideas they proposed include using Long Short Term Memory (LSTM) networks, convolutional neural networks (CNNs), and transformer networks.</li>\n<li>Additionally, they suggested using methods such as attention mechanisms, bidirectional networks, and sequence-to-sequence learning.</li>\n</ul>\n<hr>\n<h5>Summary about Loading and Preprocessing Big Jsonl Data File (<a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">Lei Wang</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/leiwong\" target=\"_blank\">Lei Wang</a> shared two notebooks to help with loading and preprocessing data for this competition.</li>\n<li><strong>Loading:</strong> The notebook explains how to load big jsonl data file into a pandas dataframe with different methods and also provides a comparison between them. </li>\n<li><strong>Preprocessing:</strong> This notebook explains how to convert the data into csv, parquet or a dataframe, as well as the advantages and disadvantages of each.</li>\n</ul>\n<hr>\n<h5>Since there are only few features(only time info), any chance to use ML algorithm? (<a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">cocoshe</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">cocoshe</a> recently posted a discussion about the possibility of using a Machine Learning (ML) algorithm for this competition.</li>\n<li>They observed that all the recall methods currently used involve \"top-k frequency aids of different session\" and \"top-k frequency aids of different type\", or a \"top-k frequency of all aids\" counting method (co-occurrence matrix). Additionally, padding is used to fill the recall@20.</li>\n<li><a href=\"https://www.kaggle.com/cocoshe\" target=\"_blank\">cocoshe</a> then asked if any chance to use ML algorithm existed, despite the few features that are available.</li>\n</ul>\n<hr>\n<h5>Q about the Data (<a href=\"https://www.kaggle.com/avieldanin\" target=\"_blank\">Aviel Danin</a>)</h5>\n<p><a href=\"https://www.kaggle.com/avieldanin\" target=\"_blank\">Aviel Danin</a> asked a question about the data in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366478\" target=\"_blank\">this post</a>:</p>\n<ul>\n<li>In some cases there is a \"cart event\" of a certain aid followed by a \"click\", and then another \"cart event\" of the same aid followed by another \"click\". </li>\n<li>In other cases, there is a \"cart event\" without a click before it. </li>\n</ul>\n<p>These scenarios could indicate that there are multiple processes related to the data, and it's important to understand them in order to get the most accurate results.</p>\n<p><strong>Answer (By (<a href=\"https://www.kaggle.com/danieleroncaglioni\" target=\"_blank\">roncadr</a>):</strong></p>\n<pre><code>Yeah I noticed the same, also in the full train dataset there are 16.896.191 carts, while there are less (12.142.933) instances of an aid of some type being followed by the same aid and type \"cart\", so definetely, apparently it is possible to add an item to the cart without having to click that item immediately beforehand. Perhaps when you are browsing on multiple browser tabs and switch between them….?\n</code></pre>\n<hr>\n<h5>💡 How to split the data for training a two-stage recommender? (<a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a> discussed the best way to split data when training a two-stage recommender system.</li>\n<li>He shared his work on <a href=\"https://www.kaggle.com/code/radek1/eda-a-look-at-the-data-training-splits\" target=\"_blank\">this</a> notebook for anyone interested in learning more about the topic.</li>\n</ul>\n<hr>\n<h5>How did you read the data? (<a href=\"https://www.kaggle.com/takuma0306\" target=\"_blank\">takuma0306</a>)</h5>\n<p>In this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/372990\" target=\"_blank\">post</a> <a href=\"https://www.kaggle.com/takuma0306\" target=\"_blank\">takuma0306</a> asked how others were loading the training data for the competition.<br>\nThey mentioned that when using VS Code, their application was terminating halfway through and when using Google Colab they were running out of system RAM and losing the session. They were wondering how others were successfully loading the data and asked for help.</p>\n<p><strong>Answer (By (<a href=\"https://www.kaggle.com/ikogias\" target=\"_blank\">Giannis</a>)):</strong></p>\n<pre><code>As an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n- https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\n</code></pre>\n<hr>\n<h5>Test Set Clarification Questions (<a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">dpalbrecht</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">dpalbrecht</a> raised some questions regarding the test set on the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365554\" target=\"_blank\">post</a> in the repo.</li>\n<li>The sentence that confused them was: \"The test set contains 1000 user queries and the ground truth is not known.\"</li>\n</ul>\n<p><strong>The questions raised were:</strong></p>\n<ul>\n<li>How the queries are produced?</li>\n<li>What does \"ground truth\" refer to?</li>\n<li>Does it mean that for each user query there is a corresponding ground truth image?</li>\n</ul>\n<p><strong>Answer (By <a href=\"https://www.kaggle.com/pietromaldini1\" target=\"_blank\">Pietro Maldini</a>):</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2F957ed2d68815e3db8abef7b03304a135%2Fground_truth.png?generation=1668199923227677&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h5>The relatively less data for each test session (<a href=\"https://www.kaggle.com/samsonfha\" target=\"_blank\">samson fha</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/samsonfha\" target=\"_blank\">samson fha</a> discussed the issue of the relatively low amount of data for each test session compared to training sessions, noting that it is difficult to identify user features and generate labels for them.</li>\n<li>They suggested focusing more on generating labels for goods/aids instead.</li>\n</ul>\n<p><strong>Answer (By <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a>)</strong><br>\nThe reason for the shorter sessions is explained <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486\" target=\"_blank\">here</a> by the organizer.</p>\n<hr>\n<h5>How many candidates you used? (<a href=\"https://www.kaggle.com/baekseungyun\" target=\"_blank\">L0Z1K</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/baekseungyun\" target=\"_blank\">L0Z1K</a> suggests that sharing the number of candidates and recall score can help the community in their analysis of the Santa 2020 competition.</li>\n<li>This could provide valuable insights into the effectiveness of certain strategies and help to identify potential improvements.</li>\n</ul>\n<hr>\n<h5>Unseen Aids in Test Set (<a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">moth</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">moth</a> made an interesting observation in their <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156\" target=\"_blank\">post</a> about the test set for the competition.</li>\n<li>They noticed that many items in the test set do not appear in the train set, which can make it difficult to accurately predict the outcomes.</li>\n<li><strong>This is an important reminder to be mindful of the unseen data that may exist in the test set when creating models.</strong></li>\n</ul>\n<hr>\n<h5>Otto recommender Kernel Stats (<a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">Alexander Ryzhkov</a> shared an amazing post about notebook analytics!</li>\n<li>Such as the number of upvotes or a fork graph.</li>\n</ul>\n<p><strong>Really cool idea!</strong><br>\nmount of data and also reduces the amount of time needed to train a model. Additionally, it can also reduce the likelihood of overfitting due to the smaller data size.</p>",
      "rawMarkdown": "###  Here is what you need to know\n#####  One month to go!\n_____\n\n##### Recommendation Systems for Large Datasets ([Ravi Shah](https://www.kaggle.com/ravishah1))\n\n[Ravi Shah](https://www.kaggle.com/ravishah1) discussed the use of recommendation systems for large datasets, such as the one used for this competition. The train dataset consists of: \n- 12,899,779 sessions\n- 1,855,603 items\n- 216,716,096 events\n- 194,720,954 clicks\n- 16,896,191 carts\n- 5,098,951 orders\n\nHe expand on this and add information about the stages for solving this problem:\n**Candidate Generation**\n\nExample criteria you can use to select you candidates:\n- previously purchased items\n- repurchased items\n- overall most popular items\n- similar items based on some sort of clustering technique\n- similar items based on something such as a co-visitation matrix\n\nBy now we should have much fewer items for each session, so we should be able to input these into a ranker model.\n\n**Ranking Model Examples**\n\n- [LGBMRanker](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html)\n- [XGBRanker](https://medium.com/predictly-on-tech/learning-to-rank-using-xgboost-83de0166229d)\n- [Ranking with sklearn](https://towardsdatascience.com/learning-to-rank-with-python-scikit-learn-327a5cfd81f)\n- [Neural Network Ranker](https://maroo.cs.umass.edu/getpdf.php?id=1373)\n\n_____\n\n##### How To Build a GBT Ranker Model ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) provided an easy way to build a solution for the competition by creating candidate rerank models. He explained how to create train data, train and infer a GBT ranker model.\n\n_____\n\n##### local validation tracks public LB perfecty -- here is the setup ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) has found that local validation tracks public leaderboard (LB) performance perfectly. They have shared their setup [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991) to allow others to replicate their results. The setup includes a scoring function that uses the same public LB metric and a test set of 500 images that corresponds to the public LB.\n\n_____\n\n##### 💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) Shares an [EDA](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset) notebook, some [dataset] he preprocessed to parquet and a [robust validation framework](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework).\n\n_____\n\n##### 30x Faster Co-Visitation Matrices using RAPIDS cuDF! ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) published a [notebook](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369) which uses RAPIDS cuDF to compute co-visitation matrices 30x faster. These matrices help to provide models with \"candidates\" which can then be reranked and selected for submission CSV. This process is useful for improving prediction accuracy.\n\n_____\n\n##### Surprising LB 0.587 !!!Share some experimental results ~~ ([Jiahong Xie](https://www.kaggle.com/jiahongxie))\n\n[Jiahong Xie](https://www.kaggle.com/jiahongxie) has made an amazing discovery with a surprising LB score of 0.587! They shared their experimental results which showed that they had a recall rate of 200 candidates for each user. This was an impressive result and could provide valuable insights for other Kagglers.\n\nIn short: And generate 300+ Features for each user and item pair.In training stage and downsample positive:negative as 1:20 for training. Also small improvement was made by deleting aid col in features.\n\n\n_____\n\n##### 📈 What do we know so far? ⚡Summary with  links to relevant resources ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) created an incredibly useful post which contains links to relevant resources related to the Santa 2022 competition. He also shared a repo on GitHub which contains data for the competition, including preprocessing code and information not available on Kaggle. This repo can be accessed [here](https://github.com/otto-de/recsys-dataset).\n\n_____\n\n##### 📑 [Step-by-step guide] How I got to my current standing on the LB and how to improve going forward ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) shared [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278) on how he achieved his current standing on the leaderboard and how to improve it going forward.\n\n- The main idea is to use a **divide-and-conquer approach to search for small improvements**, by breaking down the problem into smaller pieces and looking for more efficient solutions.\n- He advises to use a combination of different methods (random search, minimum spanning tree, etc.), as well as to take into account the constraints of the problem, such as the maximum link length.\n- He also suggests revisiting points, as well as to experiment with different strategies for finding the optimum solution. Lastly, he recommends to use the baseline functions provided but with modifications to make them faster and more efficient.\n\n_____\n\n##### Full dataset processed to CSV/parquet files with optimized memory footprint ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) shared a post discussing how to process the full dataset to CSV/parquet files with an optimized memory footprint.\n- The main idea is to use dask to read the raw data, filter it, and then save it to parquet files using the dask.dataframe.to_parquet method.\n- The resulting files are much smaller than the original and can be loaded into memory much faster. Additionally, the post provides a code example of how to use dask to process the data.\n\n_____\n\n##### Locate Real Users and Real Sessions EDA ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) discussed the use of the term \"session\" in Kaggle's Otto competition in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138).\n- In the competition, \"session\" actually means \"user\". We are given train data for 12,899,779 users over a 4 week period, and test data for users during 1 week (in the future).\n- We must predict what a user will do in the remainder of the 1 week that we do not have information about.\n\n_____\n\n##### There are some bugs in evaluation ([Tawara](https://www.kaggle.com/ttahara))\n\n[Tawara](https://www.kaggle.com/ttahara) discovered a bug in the evaluation which resulted in a score of 4.848 on the competition.\n- He shared the [notebook](https://www.kaggle.com/code/ttahara/last-aid-20/) he used to achieve this score.\n- It is important to keep in mind that the bug may have been fixed and the score may have changed.\n\n_____\n\n##### Andrew Ng Recommender Systems ([Andrada Olteanu](https://www.kaggle.com/andradaolteanu))\n\n[Andrada Olteanu](https://www.kaggle.com/andradaolteanu) discussed the fundamentals of Recommender Systems as outlined by Andrew Ng in his course on Machine Learning.\n **The main topics covered include:**\n- Collaborative Filtering and its two main approaches: User-User and Item-Item.\n- Content-Based Filtering which uses the attributes of the items being recommended.\n- Hybrid Methods which combine the two approaches.\n- Matrix Factorization which is used to classify users and items into latent features.\n- Evaluation of Recommender Systems which includes methods such as Mean Average Precision (MAP) and Root Mean Squared Error (RMSE).\n- Challenges of Recommender Systems such as scalability and cold-start problem.\n\n_____\n\n##### How to thrive in this competition without going crazy -- 1 out of 2 important truths ❤️‍🔥 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) has shared [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503) as a cautionary reminder of the complexity of this competition.\n- He has highlighted two important truths to keep in mind in order to thrive:\n- **Complexity:** Once you start to dig deeper, the competition becomes increasingly complex.\n- **Mindset:** Having the right mindset is crucial to success. Be patient and optimistic, while also being mindful of the complexity.\n\n_____\n\n##### The Story of User 13479136 and User 13710374 ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\nIn the Kaggle's Otto competition, two particular users stood out:\n- [User 13479136](https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678)\n- [User 13710374](https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678)\n- [Chris Deotte](https://www.kaggle.com/cdeotte) managed to extract item categories using RAPIDS TSNE such different categories of items may be clothing and electronics.\n- He then continue to deconstruct the individual user behaviour in the data.\n\n**Very interesting read!**\n\n_____\n\n##### Yes, We Can Use Test Data Leakage. ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) discussed the possibility of using test data leakage in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939).\n\n- Since we are predicting test events occurring within a one week period, it is possible to use the test data to predict future events.\n- This could be done by using a rolling window approach, where the training data is used to predict the test data in the following steps.\n- The test data can then be used to update the model and predict the next step.\n- This approach is beneficial as it allows for more accurate predictions in the future and could be used to improve the model's performance.\n\n_____\n\n##### Say NO to GPU: Nx Faster Co-Visitation Matrices using single CPU! ([Carno Zhao](https://www.kaggle.com/carnozhao))\n\n- [Carno Zhao](https://www.kaggle.com/carnozhao) published a [notebook](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365873) which uses `numba.jit` to compute Nx faster co-visitation matrices.\n- This is an incredibly useful tool to make computations faster and more efficient.\n\n_____\n\n##### 💡 What is the co-visitation matrix, really? ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) discussed a concept known as the co-visitation matrix in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358) and how it relates to modern techniques. \n\n**Interesting starter reading for this competition**\n\n_____\n\n##### A top-down perspective on the current metric values ([narsil](https://www.kaggle.com/narsil))\n\n[narsil](https://www.kaggle.com/narsil) shared a useful post discussing the current metric values from a top-down perspective.\n\n**Sessions in the test set are truncated in a random moment of the session life.**\n- For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart & purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.\n- For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat (see the winning scores for H&M competition - they are close to 0). This is different metric, but shows how difficult such a problem is.\n\n**Some intuitions:**\n- If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard\n- It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.\n_____\n\n\n##### How to train a Word2Vec model for item embeddings - a simple code example 📖 ([Sinan Calisir](https://www.kaggle.com/snnclsr))\n\n[Sinan Calisir](https://www.kaggle.com/snnclsr) shared a useful post on how to train a Word2Vec model to create item embeddings.\nThis includes a simple code example and detailed explanation of the process.\n\n```python\nimport pandas as pd\nfrom gensim.models import Word2Vec\ncs = 100000\ncount = 0\ntotal = 12899779\nwith open(\"w2v_input.txt\", \"w\") as f:\n    for df in pd.read_json(\"../input/otto-recommender-system/train.jsonl\", lines=True, chunksize=cs):\n        for val in df[\"events\"].apply(lambda x: [event[\"aid\"] for event in x]).values:\n            f.write(\" \".join(map(str, val)) + \"\\n\")\n        count += cs\n        print(f\"{count}/{total}\\r\", end=\"\")\nmodel = Word2Vec(corpus_file=\"w2v_input.txt\", vector_size=50, window=5, min_count=1, workers=4)\nmodel.save(\"word2vec.model\")\n```\n\nAnd after training the model we can use it like this:\n\n```python\nmodel = Word2Vec.load(\"word2vec.model\")\nmodel.wv.most_similar(\"1460571\", topn=20)\n```\n_____\n\n\n##### [Starter pack] LGBMRanker with polars 🚀🚀🚀 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- In this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194) by [Radek Osmulski](https://www.kaggle.com/radek1), he explain how re-ranking models are the industry standard for dealing with datasets like those in the Otto Recommender System competition, which have high cardinality categories.\n- He also show how to use LightGBM's ranker to create a baseline model that is competitive with the leaderboard. The post also provides a starter pack to help people getstarted, with instructions on how to use the ranker, what the parameters mean and how to tune them, as well as how to use the model for recommendations. Finally, the post also provides a notebook to help users to get their own model up and running.\n\n_____\n\n##### [Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) shared a great [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893) on improving candidate generation using a matrix factorization model!\n- This post provides a starter pack with a Pytorch implementation and a Merlin Dataloader to get started. With this starter pack, you can load data, define the model, and start training.\n- This will allow you to quickly get up and running with matrix factorization and improve candidate generation.\n\n_____\n\n##### Saving GPU memory when processing features == (Speedup + Efficiency) ([Giba](https://www.kaggle.com/titericz))\n\n[Giba](https://www.kaggle.com/titericz) discussed a new functionality of Cudf that can help save GPU memory when processing features.\n- By setting the default data type to 32-bit instead of 64-bit, the memory taken up by the GPU can be reduced. The code for this can be found below:\n\n```python\nimport cudf\ncudf.set_allocator('managed', default_dtype=np.float32)\n```\n\n_____\n\n##### I think word2vec is a nice choice. ([taku_sid](https://www.kaggle.com/takusid))\n\n- [taku_sid](https://www.kaggle.com/takusid) shared a post discussing the use of word2vec to generate candidates.\n- Managed to get to a score of lb 0.578 which is great for word2vec.\n- Some of the comments discussions includes important information about differt models used on this competition.\n\n**Recommended!**\n_____\n\n##### Processed dataset ([Konrad Banachewicz](https://www.kaggle.com/konradb))\n\n- [Konrad Banachewicz](https://www.kaggle.com/konradb) created a post [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363617) providing a training dataset in both .csv and .parquet formats for those who love pandas.\n- This can be a useful resource for anyone looking to quickly get started with the dataset.\n\n_____\n\n##### 💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition) ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000) by [Radek Osmulski](https://www.kaggle.com/radek1), he discussed several strategies for working on ML projects more efficiently.\n\n**He include:**\n- Breaking projects into smaller tasks and setting up checkpoints\n- Using domain knowledge to reduce the size of the search space\n- Leveraging existing frameworks, libraries, and tools\n- Looking for pre-existing solutions to problems\n- Focusing on one task at a time\n- Streamlining the data collection and labelling process\n- Setting up automated evaluations and feedback loops\n- Taking breaks to reflect and reassess goals\n- Working in collaboration with others\n\n_____\n\n##### Which metric is correct? ([dehokanta](https://www.kaggle.com/dehokanta))\n\n[dehokanta](https://www.kaggle.com/dehokanta) raised a question on the evaluation metric for this competition: Recall@k.\n\n**[dehokanta](https://www.kaggle.com/dehokanta) asked which of the following two definitions is correct:**\n\n- (1) Recall@k is the fraction of the top k pixels of an image that are the same as the target image.\n- (2) Recall@k is the fraction of the target image that is in the top k pixels of an image.\n\nThe discussion that followed suggested that the **second definition** is the correct one.\n_____\n\n##### 💡How to improve the results of your Approximate Nearest Neighbor search! (annoy) ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) gave an insightful post on [how to improve the results of your Approximate Nearest Neighbor search](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368385)!\n\n**Main Idea:** Use an Approximate Nearest Neighbor (ANN) search to improve the results of your search.\n\n**Benefits:**\n\n- ANN searches are much faster than traditional nearest neighbor searches and can be used to quickly find similar items.\n- ANN searches can also be used to find items with similar characteristics, even if they don’t have an exact match.\n- ANNs are also more scalable, as they can be easily parallelized and distributed across multiple nodes.\n\n**Implementation:**\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) suggests using the [Annoy library](https://github.com/spotify/annoy) to implement the ANN search.\n- Annoy is a Python library that allows you to quickly index and search for points in a vector space.\n- Annoy also includes a few useful features such as the ability to easily control the size of the search pool, as well as the ability to specify the number of neighbors to search for.\n\n**Conclusion:**\n\nUsing an Approximate Nearest Neighbor search can provide a number of benefits, such as increased search speed and scalability. [Radek Osmulski](https://www.kaggle.com/radek1) suggests using the Annoy library to implement the ANN search, which provides a number of useful features for controlling the size of the search pool and specifying the number of neighbors to search for.\n\n_____\n\n##### 💡 Training an XGBoost Ranker on the GPU with Merlin Models 🔥🔥🔥 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368848), [Radek Osmulski](https://www.kaggle.com/radek1) discussed the use of Merlin Models to train an XGBoost Ranker on the GPU.\nThis can greatly reduce the training time for XGBoost models, allowing for faster and more efficient training. The main benefits are that it does not require a Tensorflow backend, and allows for faster training. It also allows for better hyperparameter optimization, since it is possible to train multiple models in parallel. Additionally, the model can be deployed to the cloud for scalability and higher performance. Finally, the model can also be used for inference tasks such as ranking and recommendation.\n\n_____\n\n##### Matrix Factorization with GPU: 6.5x faster! ([CPMP](https://www.kaggle.com/cpmpml))\n\n- [CPMP](https://www.kaggle.com/cpmpml) shared a [very interesting notebook](https://www.kaggle.com/competitions/otto-recommender-system/discussion/371166) to compute matrix factorization, using polar, annoy, Merlin data loader and pytorch.\n- [CPMP](https://www.kaggle.com/cpmpml) then created a GPU version of the notebook which is [available here](https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu) and is 6.5x faster than the original version!\n\n_____\n\n##### 🐘 the elephant in the room -- high cardinality of targets and what to do about this ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722), [Radek Osmulski](https://www.kaggle.com/radek1) discusses an issue that is often overlooked when working with Machine Learning models, high cardinality of targets and how to handle it. The main idea presented is that, in order to build an effective model, it is important to accurately measure the importance of each feature. The author suggests two main approaches: \n\n- **Feature Engineering:** transforming existing features into a different representation that can be used to better understand the underlying data. \n- **Dimensionality Reduction:** reducing the number of features to a more manageable amount.\n\nThe author also outlines some potential pitfalls, such as overfitting and underfitting, that can arise when dealing with high cardinality. Finally, the post provides some tips on how to effectively use the two main approaches to tackle this problem.\n\n_____\n\n##### Co-Visitation Matrices and Matrix Factorization ([Ravi Shah](https://www.kaggle.com/ravishah1))\n\n[Ravi Shah](https://www.kaggle.com/ravishah1) discussed the use of co-visitation matrices and matrix factorization for recommendation systems in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589). The main idea is to create features based on items which are frequently viewed together. This technique has been shown to be effective for creating better recommendation systems.\n\n_____\n\n##### 💡What is a good initial goal in the competition? How to improve beyond it? 📈 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368685), [Radek Osmulski](https://www.kaggle.com/radek1) suggested a good initial goal for the competition is to get everything up and running on your end. This includes setting up your environment and running the notebooks. Additionally, [Radek Osmulski](https://www.kaggle.com/radek1) suggested that to improve beyond this goal, it is important to understand the underlying techniques and concepts, such as using the right data structure and algorithm for the task, as well as understanding the limitations of the algorithms. Finally, [Radek Osmulski](https://www.kaggle.com/radek1) recommended debugging your code and understanding the underlying algorithms, such as a greedy approach, to improve beyond the initial goal.\n\n_____\n\n##### To the people still hacking away on this competition... ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) has noticed that there are fewer and fewer submissions on the Leaderboard of [this competition](https://www.kaggle.com/competitions/otto-recommender-system/discussion/369789). He encourages competitors to keep working on their solutions and suggests they explore the following ideas to get better results: \n- Make sure you are using algorithms that are tailored to the problem, such as the Minimum Spanning Tree algorithm.\n- Try out different optimization techniques.\n- Take a close look at your code and identify possible improvements.\n- Use the insights and ideas of other Kagglers. \n- Share your progress and ideas with others.\n\n_____\n\n##### recommenders - Best Practices on Recommendation Systems by Microsoft ([Sinan Calisir](https://www.kaggle.com/snnclsr))\n\n[Sinan Calisir](https://www.kaggle.com/snnclsr) from Microsoft shared a great post discussing best practices when building Recommender Systems. The main points covered include:\n- Data Preprocessing: including data cleaning, normalization, and feature engineering.\n- Model Selection: including selecting the right model for the task, hyperparameter optimization and model validation.\n- Evaluation Metrics: including the use of standard metrics (such as RMSE) as well as business metrics.\n- Deployment & Maintenance: including integrating the model into applications, model retraining and monitoring.\n- User Experience: including recommendations personalization and explainability. \nCheck out the full [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363730) for further details.\n\n_____\n\n##### Why best model might not win RecSys competition ([Roman](https://www.kaggle.com/nroman))\n\n- [Roman](https://www.kaggle.com/nroman) discussed the difference between offline and online metrics in Recommender System (RecSys) competitions.\n- The ground truth labels for RecSys competitions are derived from the organizers' current model. This may mean that the best model might not win the competition, as the model which scores the highest on the offline metric might not be the same as the one which scores the highest on the online metric.\n\n_____\n\n##### Important information regarding test data from competition repository ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) has shared a helpful repository containing important information regarding the test data from the competition.\n- The repository includes a comprehensive README as well as some related code. Check it out [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965)!\n\n_____\n\n##### Otto's Talk on Transformer Recommendation Systems ([BenediktSchifferer](https://www.kaggle.com/benediktschifferer))\n\n[BenediktSchifferer](https://www.kaggle.com/benediktschifferer) presented on the topic of Transformer Recommendation Systems in a [talk](https://www.kaggle.com/competitions/otto-recommender-system/discussion/373224).\n- The main idea of the talk was to explain the use of transformers in recommendation systems and to discuss their potential applications.\n- The talk discussed how transformers can be used to learn user preference and then use them to generate recommendations. It also discussed how transformers can be used to build an end-to-end recommendation pipeline, as well as how they can be used to improve existing models.\n- The talk discussed how transformers can be used to solve cold-start and long-tail problems. The talk highlighted the advantages of using transformers, such as their capability to capture long-term and short-term user preferences, and their ability to handle both structured and unstructured data. It also discussed the challenges of using transformers, such as their large memory requirement and computational cost.\n\n_____\n\n##### RePlay - opensource RecSys constructor ([Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov))\n\n[Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov) has created an opensource RecSys constructor called [RePlay](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365917). It is a library that provides tools for all stages of creating a recommendation system, such as data preprocessing, model evaluation, and comparison. RePlay uses PySpark to handle big data.\n\n_____\n\n##### 💡 ANN -&gt; NN: get better results and run faster with NN search on the GPU 🔥🔥🔥 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) recently shared an amazing post on how to get better results and run faster with NN search on the GPU.\n- [Radek Osmulski](https://www.kaggle.com/radek1) suggested using Matrix Factorization with GPU which is 6.5x faster. Furthermore, [Radek Osmulski](https://www.kaggle.com/radek1) mentioned that this solution was already implemented by their colleagues.\n\n_____\n\n##### Ground-truth? ([Gunes Evitan](https://www.kaggle.com/gunesevitan))\n\n- [Gunes Evitan](https://www.kaggle.com/gunesevitan) asked a question about what \"test set is truncated\" means and if it is possible to create ground-truth for the training set by only one event.\n- The post explains that it is not truncated by a single event, instead it is truncated by the total number of frames in the test set.\n- Therefore, it is not possible to create the ground-truth by a single event, as the test set contains more frames than the training set.\n\n_____\n\n##### The time zone in Germany is UTC+2 in August ([aldparis](https://www.kaggle.com/adaubas))\n\n- [aldparis](https://www.kaggle.com/adaubas) discussed the time zone in Germany being UTC+2 in August.\n- This is due to Daylight Saving Time (DST) which is observed in Germany from the last Sunday in March until the last Sunday in October.\n- During DST, the clocks are moved forward by one hour, resulting in the time zone shifting from UTC+1 to UTC+2. This means that the daylight hours will be longer in the summer months, with the sun setting later in the evening.\n\n_____\n\n##### Ranker models vs. Binary classification models ([Andrej Zubaľ](https://www.kaggle.com/andrejzuba))\n\n[Andrej Zubaľ](https://www.kaggle.com/andrejzuba) discussed the differences between ranker models and binary classification models in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370502).\n- **Ranker models** were found to be more accurate for predicting the probability of a certain event or outcome, as they can rank a list of outcomes from most to least likely.\n- **Binary classification** models are simpler and can be used to classify events into two categories, such as yes or no, true or false.\n- **Ranker models** are better suited for prediction tasks such as predicting the relevance of a search query, while binary classification models are better suited for categorizing events.\n_____\n\n##### How do you train Ranking model? ([dehokanta](https://www.kaggle.com/dehokanta))\n\n[dehokanta](https://www.kaggle.com/dehokanta) asked a question about how to build a ranking model. In the post, they mentioned that they have tried various methods but with not much success.\n\n**This post provides useful information on how to train a ranking model, such as:**\n- Understanding how Ranking works\n- Normalising features\n- Feature selection\n- Data Augmentation\n- Training and Tuning\n- Evaluation Metric Selection\n- Deployment and Monitoring\n\n_____\n\n##### 💡How to ensemble predictions -- a key component to every strong solution 🏅 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) in this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368747) discussed the importance of ensembling for a strong solution in Kaggle competitions. Ensembling is combining multiple models to get a better result than the individual models. \n\nThe main idea is to create multiple models that are independent from each other and then combine them to get better results.\nThe author suggests combining models that have different architectures, different pre-processing techniques and different hyperparameters.\n\n**The author also outlines some techniques for ensembling such as:**\n- Averaging\n- Weighted Averaging\n- Model Stacking\n- Bagging and Boosting\n\n**The author also provides some useful tips to keep in mind while ensembling such as:**\n- Identify the best models and ensemble them\n- Try using different weights for each model\n- Use different seed values for each model\n- Use different input data to train each model\n- Use regularization to reduce overfitting\n- Use cross-validation to evaluate the models.\n\nEnsembling is an important technique to improve the accuracy of models and can be used to achieve higher scores in Kaggle competitions.\n\n_____\n\n##### Guidance on what algos to try from one of the authors of Microsoft Recommenders repo ([Miguel Fierro](https://www.kaggle.com/hoaphumanoid))\n\n[Miguel Fierro](https://www.kaggle.com/hoaphumanoid) is one of the authors of the [Microsoft Recommenders repo](https://github.com/microsoft/recommenders/).\n- In this post they provide guidance on which algorithms to try depending on the problem.\n- They mention to start with simple algorithms like popularity, content-based filtering, and matrix factorization.\n- They also suggest to look into user-based collaborative filtering and embeddings. Finally, they suggest to try deep learning methods such as deep matrix factorization, deep autoencoders, and convolutional neural networks.\n\n_____\n\n##### Graph Neural Network, is complex to model but efficient for this task ([Youness EL BRAG](https://www.kaggle.com/younesselbrag))\n\n- [Youness EL BRAG](https://www.kaggle.com/younesselbrag) presented a new approach, [Session-based Recommendation with Graph Neural Networks](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368355), to the task of session-based recommendation.\n- This method models a session as a complex transition of items and estimates user representations in addition to item representations. An accompanying [code](https://github.com/CRIPAC-DIG/SR-GNN?utm_source=catalyzex.com) was also released which enables users to implement the method in their own projects.\n\n_____\n\n##### cuDF vs Pandas Vs Modin for faster data processing and there comparison ([Satya](https://www.kaggle.com/satyaprakashshukl))\n\n- [Satya](https://www.kaggle.com/satyaprakashshukl) discussed different methods for faster data processing and their comparison and speed on this competition.\n- Really useful post for those who are looking for faster data processing methods.\n- Check out the [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366258) for more information.\n\n_____\n\n##### Some Interesting Times Series on Products ([aldparis](https://www.kaggle.com/adaubas))\n\n[aldparis](https://www.kaggle.com/adaubas) recently [shared](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368373) an interesting discussion topic about product times series.\n**The main idea behind it is to use the sales data of products to make predictions and forecast future sales.**\n\n**The post contains a few insights such as:**\n- Times series can be used to forecast future sales more accurately. \n- Seasonality can be taken into account to make more precise predictions.\n- The need to use the most recent data available to make more reliable predictions.\n- Machine Learning algorithms can be used to model times series data.\n- Using multiple linear models can improve accuracy.\n- Different methods of smoothing can be applied to times series data.\n\n_____\n\n##### How can a session start with an order? ([CPMP](https://www.kaggle.com/cpmpml))\n\n- [CPMP](https://www.kaggle.com/cpmpml) posed a question about how to start a session with an order.\n- He suggested that it could be done by creating a queue of tasks and using a scheduler to set the order of tasks and limit the number of tasks run in parallel.\n- Additionally, they suggested the use of a job server that can manage the scheduling. Also, he discussed the use of a separate service for handling long-running tasks, such as background tasks.\n_____\n\n##### Difference between Co-visitation Matrix v.s. Matrix Factorization ([bilzard](https://www.kaggle.com/tatamikenn))\n- [bilzard](https://www.kaggle.com/tatamikenn) discussed the difference between the co-visitation matrix approach, which can be found [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725851), and the matrix decomposition approach, which can be found [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725852).\n- The decisive difference between them is the dependence of events, which could be called the context in consideration. This could explain why the co-visitation matrix approach scores higher than the matrix decomposition approach.\n\n_____\n\n##### Some concerns about validation ([Gunes Evitan](https://www.kaggle.com/gunesevitan))\n\n- [Gunes Evitan](https://www.kaggle.com/gunesevitan) noticed that most people were using the test set split code from the organizers' [repository](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370384)\n- But has some concerns that it might have flaws.\n-  He suggests using an alternative method to split the data, such as using a train/validation/test split, which would allow for better validation. This could help avoid overfitting and could potentially lead to better results. He also encourages people to think about the validation method and consider different approaches.\n\n_____\n\n##### We can try more NLP methods in sequence recomendation, like .... ([KKY](https://www.kaggle.com/evilpsycho42))\n\n- In this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367870), [KKY](https://www.kaggle.com/evilpsycho42) discussed the possibility of using NLP methods to improve sequence recommendation in this user-anonymized dataset.\n- They suggested trying out various natural language processing (NLP) techniques to get a better representation of the action/item sequence.\n- Some ideas they proposed include using Long Short Term Memory (LSTM) networks, convolutional neural networks (CNNs), and transformer networks.\n- Additionally, they suggested using methods such as attention mechanisms, bidirectional networks, and sequence-to-sequence learning.\n\n_____\n\n##### Summary about Loading and Preprocessing Big Jsonl Data File ([Lei Wang](https://www.kaggle.com/leiwong))\n\n- [Lei Wang](https://www.kaggle.com/leiwong) shared two notebooks to help with loading and preprocessing data for this competition.\n- **Loading:** The notebook explains how to load big jsonl data file into a pandas dataframe with different methods and also provides a comparison between them. \n- **Preprocessing:** This notebook explains how to convert the data into csv, parquet or a dataframe, as well as the advantages and disadvantages of each.\n\n_____\n\n##### Since there are only few features(only time info), any chance to use ML algorithm? ([cocoshe](https://www.kaggle.com/cocoshe))\n\n- [cocoshe](https://www.kaggle.com/cocoshe) recently posted a discussion about the possibility of using a Machine Learning (ML) algorithm for this competition.\n- They observed that all the recall methods currently used involve \"top-k frequency aids of different session\" and \"top-k frequency aids of different type\", or a \"top-k frequency of all aids\" counting method (co-occurrence matrix). Additionally, padding is used to fill the recall@20.\n- [cocoshe](https://www.kaggle.com/cocoshe) then asked if any chance to use ML algorithm existed, despite the few features that are available.\n\n_____\n\n##### Q about the Data ([Aviel Danin](https://www.kaggle.com/avieldanin))\n\n[Aviel Danin](https://www.kaggle.com/avieldanin) asked a question about the data in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366478):\n\n- In some cases there is a \"cart event\" of a certain aid followed by a \"click\", and then another \"cart event\" of the same aid followed by another \"click\". \n- In other cases, there is a \"cart event\" without a click before it. \n\nThese scenarios could indicate that there are multiple processes related to the data, and it's important to understand them in order to get the most accurate results.\n\n**Answer (By ([roncadr](https://www.kaggle.com/danieleroncaglioni)):**\n```\nYeah I noticed the same, also in the full train dataset there are 16.896.191 carts, while there are less (12.142.933) instances of an aid of some type being followed by the same aid and type \"cart\", so definetely, apparently it is possible to add an item to the cart without having to click that item immediately beforehand. Perhaps when you are browsing on multiple browser tabs and switch between them….?\n```\n_____\n\n##### 💡 How to split the data for training a two-stage recommender? ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) discussed the best way to split data when training a two-stage recommender system.\n- He shared his work on [this](https://www.kaggle.com/code/radek1/eda-a-look-at-the-data-training-splits) notebook for anyone interested in learning more about the topic.\n_____\n\n##### How did you read the data? ([takuma0306](https://www.kaggle.com/takuma0306))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/372990) [takuma0306](https://www.kaggle.com/takuma0306) asked how others were loading the training data for the competition.\nThey mentioned that when using VS Code, their application was terminating halfway through and when using Google Colab they were running out of system RAM and losing the session. They were wondering how others were successfully loading the data and asked for help.\n\n**Answer (By ([Giannis](https://www.kaggle.com/ikogias))):**\n```\nAs an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n- https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\n```\n_____\n\n##### Test Set Clarification Questions ([dpalbrecht](https://www.kaggle.com/dpalbrecht))\n\n- [dpalbrecht](https://www.kaggle.com/dpalbrecht) raised some questions regarding the test set on the [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365554) in the repo.\n- The sentence that confused them was: \"The test set contains 1000 user queries and the ground truth is not known.\"\n\n**The questions raised were:**\n- How the queries are produced?\n- What does \"ground truth\" refer to?\n- Does it mean that for each user query there is a corresponding ground truth image?\n\n**Answer (By [Pietro Maldini](https://www.kaggle.com/pietromaldini1)):**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2F957ed2d68815e3db8abef7b03304a135%2Fground_truth.png?generation=1668199923227677&alt=media)\n\n_____\n\n\n##### The relatively less data for each test session ([samson fha](https://www.kaggle.com/samsonfha))\n\n- [samson fha](https://www.kaggle.com/samsonfha) discussed the issue of the relatively low amount of data for each test session compared to training sessions, noting that it is difficult to identify user features and generate labels for them.\n- They suggested focusing more on generating labels for goods/aids instead.\n\n**Answer (By [Radek Osmulski](https://www.kaggle.com/radek1))**\nThe reason for the shorter sessions is explained [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486) by the organizer.\n\n_____\n\n##### How many candidates you used? ([L0Z1K](https://www.kaggle.com/baekseungyun))\n\n- [L0Z1K](https://www.kaggle.com/baekseungyun) suggests that sharing the number of candidates and recall score can help the community in their analysis of the Santa 2020 competition.\n- This could provide valuable insights into the effectiveness of certain strategies and help to identify potential improvements.\n\n_____\n\n##### Unseen Aids in Test Set ([moth](https://www.kaggle.com/alejopaullier))\n\n- [moth](https://www.kaggle.com/alejopaullier) made an interesting observation in their [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156) about the test set for the competition.\n- They noticed that many items in the test set do not appear in the train set, which can make it difficult to accurately predict the outcomes.\n- **This is an important reminder to be mindful of the unseen data that may exist in the test set when creating models.**\n\n_____\n\n##### Otto recommender Kernel Stats ([Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov))\n\n- [Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov) shared an amazing post about notebook analytics!\n- Such as the number of upvotes or a fork graph.\n\n**Really cool idea!**\nmount of data and also reduces the amount of time needed to train a model. Additionally, it can also reduce the likelihood of overfitting due to the smaller data size.",
      "votes": null
    },
    {
      "id": "2076639",
      "postDate": "12/26/2022 16:56:02",
      "content": "<p>Amazing! Thanks for sharing this with us!</p>",
      "rawMarkdown": "Amazing! Thanks for sharing this with us!",
      "votes": null
    },
    {
      "id": "2076928",
      "postDate": "12/27/2022 03:08:41",
      "content": "<p>Great, so detailed !! Thanks for sharing!</p>",
      "rawMarkdown": "Great, so detailed !! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2086845",
      "postDate": "01/05/2023 05:18:18",
      "content": "<p>This is amazing, Thanks a lot for sharing!</p>",
      "rawMarkdown": "This is amazing, Thanks a lot for sharing!",
      "votes": null
    },
    {
      "id": "2087467",
      "postDate": "01/05/2023 16:02:22",
      "content": "<p>So detailed and helpful!!! Thanks for your sharing!</p>",
      "rawMarkdown": "So detailed and helpful!!! Thanks for your sharing!",
      "votes": null
    },
    {
      "id": "2088258",
      "postDate": "01/06/2023 07:54:12",
      "content": "<p>Great information! Thanks for your summary!</p>",
      "rawMarkdown": "Great information! Thanks for your summary!",
      "votes": null
    },
    {
      "id": "2093188",
      "postDate": "01/09/2023 21:12:09",
      "content": "<p>This is amazing and helpful, thanks a lot for sharing!</p>",
      "rawMarkdown": "This is amazing and helpful, thanks a lot for sharing!",
      "votes": null
    },
    {
      "id": "2095636",
      "postDate": "01/11/2023 14:28:43",
      "content": "<p>This is amazing and helpful! Thanks for sharing!</p>",
      "rawMarkdown": "This is amazing and helpful! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2105455",
      "postDate": "01/18/2023 14:17:34",
      "content": "<p><a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> Wow, super cool! Thank you so much for sharing this wonderful post with summarizing all these wonderful resources! I was intended to a similar summary like this in the beginning but failed. It's wonderful you are sharing this with all of us!!! 🙏🙏🙏</p>",
      "rawMarkdown": "thedevastator Wow, super cool! Thank you so much for sharing this wonderful post with summarizing all these wonderful resources! I was intended to a similar summary like this in the beginning but failed. It's wonderful you are sharing this with all of us!!! 🙏🙏🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2076639,
      "author_name": "d17129765",
      "author_url": "",
      "post_date": "12/26/2022 16:56:02",
      "content": "<p>Amazing! Thanks for sharing this with us!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2076928,
      "author_name": "vincent108",
      "author_url": "",
      "post_date": "12/27/2022 03:08:41",
      "content": "<p>Great, so detailed !! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2086845,
      "author_name": "jigyasasinghal",
      "author_url": "",
      "post_date": "01/05/2023 05:18:18",
      "content": "<p>This is amazing, Thanks a lot for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2087467,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "01/05/2023 16:02:22",
      "content": "<p>So detailed and helpful!!! Thanks for your sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2088258,
      "author_name": "toyou2u",
      "author_url": "",
      "post_date": "01/06/2023 07:54:12",
      "content": "<p>Great information! Thanks for your summary!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2093188,
      "author_name": "stephanjenzen",
      "author_url": "",
      "post_date": "01/09/2023 21:12:09",
      "content": "<p>This is amazing and helpful, thanks a lot for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2095636,
      "author_name": "jinnng0016",
      "author_url": "",
      "post_date": "01/11/2023 14:28:43",
      "content": "<p>This is amazing and helpful! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2105455,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "01/18/2023 14:17:34",
      "content": "<p><a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> Wow, super cool! Thank you so much for sharing this wonderful post with summarizing all these wonderful resources! I was intended to a similar summary like this in the beginning but failed. It's wonderful you are sharing this with all of us!!! 🙏🙏🙏</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2076039": "###  Here is what you need to know\n#####  One month to go!\n_____\n\n##### Recommendation Systems for Large Datasets ([Ravi Shah](https://www.kaggle.com/ravishah1))\n\n[Ravi Shah](https://www.kaggle.com/ravishah1) discussed the use of recommendation systems for large datasets, such as the one used for this competition. The train dataset consists of: \n- 12,899,779 sessions\n- 1,855,603 items\n- 216,716,096 events\n- 194,720,954 clicks\n- 16,896,191 carts\n- 5,098,951 orders\n\nHe expand on this and add information about the stages for solving this problem:\n**Candidate Generation**\n\nExample criteria you can use to select you candidates:\n- previously purchased items\n- repurchased items\n- overall most popular items\n- similar items based on some sort of clustering technique\n- similar items based on something such as a co-visitation matrix\n\nBy now we should have much fewer items for each session, so we should be able to input these into a ranker model.\n\n**Ranking Model Examples**\n\n- [LGBMRanker](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html)\n- [XGBRanker](https://medium.com/predictly-on-tech/learning-to-rank-using-xgboost-83de0166229d)\n- [Ranking with sklearn](https://towardsdatascience.com/learning-to-rank-with-python-scikit-learn-327a5cfd81f)\n- [Neural Network Ranker](https://maroo.cs.umass.edu/getpdf.php?id=1373)\n\n_____\n\n##### How To Build a GBT Ranker Model ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) provided an easy way to build a solution for the competition by creating candidate rerank models. He explained how to create train data, train and infer a GBT ranker model.\n\n_____\n\n##### local validation tracks public LB perfecty -- here is the setup ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) has found that local validation tracks public leaderboard (LB) performance perfectly. They have shared their setup [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991) to allow others to replicate their results. The setup includes a scoring function that uses the same public LB metric and a test set of 500 images that corresponds to the public LB.\n\n_____\n\n##### 💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) Shares an [EDA](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset) notebook, some [dataset] he preprocessed to parquet and a [robust validation framework](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework).\n\n_____\n\n##### 30x Faster Co-Visitation Matrices using RAPIDS cuDF! ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) published a [notebook](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365369) which uses RAPIDS cuDF to compute co-visitation matrices 30x faster. These matrices help to provide models with \"candidates\" which can then be reranked and selected for submission CSV. This process is useful for improving prediction accuracy.\n\n_____\n\n##### Surprising LB 0.587 !!!Share some experimental results ~~ ([Jiahong Xie](https://www.kaggle.com/jiahongxie))\n\n[Jiahong Xie](https://www.kaggle.com/jiahongxie) has made an amazing discovery with a surprising LB score of 0.587! They shared their experimental results which showed that they had a recall rate of 200 candidates for each user. This was an impressive result and could provide valuable insights for other Kagglers.\n\nIn short: And generate 300+ Features for each user and item pair.In training stage and downsample positive:negative as 1:20 for training. Also small improvement was made by deleting aid col in features.\n\n\n_____\n\n##### 📈 What do we know so far? ⚡Summary with  links to relevant resources ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) created an incredibly useful post which contains links to relevant resources related to the Santa 2022 competition. He also shared a repo on GitHub which contains data for the competition, including preprocessing code and information not available on Kaggle. This repo can be accessed [here](https://github.com/otto-de/recsys-dataset).\n\n_____\n\n##### 📑 [Step-by-step guide] How I got to my current standing on the LB and how to improve going forward ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) shared [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278) on how he achieved his current standing on the leaderboard and how to improve it going forward.\n\n- The main idea is to use a **divide-and-conquer approach to search for small improvements**, by breaking down the problem into smaller pieces and looking for more efficient solutions.\n- He advises to use a combination of different methods (random search, minimum spanning tree, etc.), as well as to take into account the constraints of the problem, such as the maximum link length.\n- He also suggests revisiting points, as well as to experiment with different strategies for finding the optimum solution. Lastly, he recommends to use the baseline functions provided but with modifications to make them faster and more efficient.\n\n_____\n\n##### Full dataset processed to CSV/parquet files with optimized memory footprint ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) shared a post discussing how to process the full dataset to CSV/parquet files with an optimized memory footprint.\n- The main idea is to use dask to read the raw data, filter it, and then save it to parquet files using the dask.dataframe.to_parquet method.\n- The resulting files are much smaller than the original and can be loaded into memory much faster. Additionally, the post provides a code example of how to use dask to process the data.\n\n_____\n\n##### Locate Real Users and Real Sessions EDA ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) discussed the use of the term \"session\" in Kaggle's Otto competition in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366138).\n- In the competition, \"session\" actually means \"user\". We are given train data for 12,899,779 users over a 4 week period, and test data for users during 1 week (in the future).\n- We must predict what a user will do in the remainder of the 1 week that we do not have information about.\n\n_____\n\n##### There are some bugs in evaluation ([Tawara](https://www.kaggle.com/ttahara))\n\n[Tawara](https://www.kaggle.com/ttahara) discovered a bug in the evaluation which resulted in a score of 4.848 on the competition.\n- He shared the [notebook](https://www.kaggle.com/code/ttahara/last-aid-20/) he used to achieve this score.\n- It is important to keep in mind that the bug may have been fixed and the score may have changed.\n\n_____\n\n##### Andrew Ng Recommender Systems ([Andrada Olteanu](https://www.kaggle.com/andradaolteanu))\n\n[Andrada Olteanu](https://www.kaggle.com/andradaolteanu) discussed the fundamentals of Recommender Systems as outlined by Andrew Ng in his course on Machine Learning.\n **The main topics covered include:**\n- Collaborative Filtering and its two main approaches: User-User and Item-Item.\n- Content-Based Filtering which uses the attributes of the items being recommended.\n- Hybrid Methods which combine the two approaches.\n- Matrix Factorization which is used to classify users and items into latent features.\n- Evaluation of Recommender Systems which includes methods such as Mean Average Precision (MAP) and Root Mean Squared Error (RMSE).\n- Challenges of Recommender Systems such as scalability and cold-start problem.\n\n_____\n\n##### How to thrive in this competition without going crazy -- 1 out of 2 important truths ❤️‍🔥 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) has shared [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367503) as a cautionary reminder of the complexity of this competition.\n- He has highlighted two important truths to keep in mind in order to thrive:\n- **Complexity:** Once you start to dig deeper, the competition becomes increasingly complex.\n- **Mindset:** Having the right mindset is crucial to success. Be patient and optimistic, while also being mindful of the complexity.\n\n_____\n\n##### The Story of User 13479136 and User 13710374 ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\nIn the Kaggle's Otto competition, two particular users stood out:\n- [User 13479136](https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678)\n- [User 13710374](https://www.kaggle.com/competitions/otto-recommender-system/discussion/371678)\n- [Chris Deotte](https://www.kaggle.com/cdeotte) managed to extract item categories using RAPIDS TSNE such different categories of items may be clothing and electronics.\n- He then continue to deconstruct the individual user behaviour in the data.\n\n**Very interesting read!**\n\n_____\n\n##### Yes, We Can Use Test Data Leakage. ([Chris Deotte](https://www.kaggle.com/cdeotte))\n\n[Chris Deotte](https://www.kaggle.com/cdeotte) discussed the possibility of using test data leakage in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939).\n\n- Since we are predicting test events occurring within a one week period, it is possible to use the test data to predict future events.\n- This could be done by using a rolling window approach, where the training data is used to predict the test data in the following steps.\n- The test data can then be used to update the model and predict the next step.\n- This approach is beneficial as it allows for more accurate predictions in the future and could be used to improve the model's performance.\n\n_____\n\n##### Say NO to GPU: Nx Faster Co-Visitation Matrices using single CPU! ([Carno Zhao](https://www.kaggle.com/carnozhao))\n\n- [Carno Zhao](https://www.kaggle.com/carnozhao) published a [notebook](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365873) which uses `numba.jit` to compute Nx faster co-visitation matrices.\n- This is an incredibly useful tool to make computations faster and more efficient.\n\n_____\n\n##### 💡 What is the co-visitation matrix, really? ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) discussed a concept known as the co-visitation matrix in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358) and how it relates to modern techniques. \n\n**Interesting starter reading for this competition**\n\n_____\n\n##### A top-down perspective on the current metric values ([narsil](https://www.kaggle.com/narsil))\n\n[narsil](https://www.kaggle.com/narsil) shared a useful post discussing the current metric values from a top-down perspective.\n\n**Sessions in the test set are truncated in a random moment of the session life.**\n- For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart & purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.\n- For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat (see the winning scores for H&M competition - they are close to 0). This is different metric, but shows how difficult such a problem is.\n\n**Some intuitions:**\n- If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard\n- It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.\n_____\n\n\n##### How to train a Word2Vec model for item embeddings - a simple code example 📖 ([Sinan Calisir](https://www.kaggle.com/snnclsr))\n\n[Sinan Calisir](https://www.kaggle.com/snnclsr) shared a useful post on how to train a Word2Vec model to create item embeddings.\nThis includes a simple code example and detailed explanation of the process.\n\n```python\nimport pandas as pd\nfrom gensim.models import Word2Vec\ncs = 100000\ncount = 0\ntotal = 12899779\nwith open(\"w2v_input.txt\", \"w\") as f:\n    for df in pd.read_json(\"../input/otto-recommender-system/train.jsonl\", lines=True, chunksize=cs):\n        for val in df[\"events\"].apply(lambda x: [event[\"aid\"] for event in x]).values:\n            f.write(\" \".join(map(str, val)) + \"\\n\")\n        count += cs\n        print(f\"{count}/{total}\\r\", end=\"\")\nmodel = Word2Vec(corpus_file=\"w2v_input.txt\", vector_size=50, window=5, min_count=1, workers=4)\nmodel.save(\"word2vec.model\")\n```\n\nAnd after training the model we can use it like this:\n\n```python\nmodel = Word2Vec.load(\"word2vec.model\")\nmodel.wv.most_similar(\"1460571\", topn=20)\n```\n_____\n\n\n##### [Starter pack] LGBMRanker with polars 🚀🚀🚀 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- In this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194) by [Radek Osmulski](https://www.kaggle.com/radek1), he explain how re-ranking models are the industry standard for dealing with datasets like those in the Otto Recommender System competition, which have high cardinality categories.\n- He also show how to use LightGBM's ranker to create a baseline model that is competitive with the leaderboard. The post also provides a starter pack to help people getstarted, with instructions on how to use the ranker, what the parameters mean and how to tune them, as well as how to use the model for recommendations. Finally, the post also provides a notebook to help users to get their own model up and running.\n\n_____\n\n##### [Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) shared a great [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366893) on improving candidate generation using a matrix factorization model!\n- This post provides a starter pack with a Pytorch implementation and a Merlin Dataloader to get started. With this starter pack, you can load data, define the model, and start training.\n- This will allow you to quickly get up and running with matrix factorization and improve candidate generation.\n\n_____\n\n##### Saving GPU memory when processing features == (Speedup + Efficiency) ([Giba](https://www.kaggle.com/titericz))\n\n[Giba](https://www.kaggle.com/titericz) discussed a new functionality of Cudf that can help save GPU memory when processing features.\n- By setting the default data type to 32-bit instead of 64-bit, the memory taken up by the GPU can be reduced. The code for this can be found below:\n\n```python\nimport cudf\ncudf.set_allocator('managed', default_dtype=np.float32)\n```\n\n_____\n\n##### I think word2vec is a nice choice. ([taku_sid](https://www.kaggle.com/takusid))\n\n- [taku_sid](https://www.kaggle.com/takusid) shared a post discussing the use of word2vec to generate candidates.\n- Managed to get to a score of lb 0.578 which is great for word2vec.\n- Some of the comments discussions includes important information about differt models used on this competition.\n\n**Recommended!**\n_____\n\n##### Processed dataset ([Konrad Banachewicz](https://www.kaggle.com/konradb))\n\n- [Konrad Banachewicz](https://www.kaggle.com/konradb) created a post [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363617) providing a training dataset in both .csv and .parquet formats for those who love pandas.\n- This can be a useful resource for anyone looking to quickly get started with the dataset.\n\n_____\n\n##### 💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition) ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000) by [Radek Osmulski](https://www.kaggle.com/radek1), he discussed several strategies for working on ML projects more efficiently.\n\n**He include:**\n- Breaking projects into smaller tasks and setting up checkpoints\n- Using domain knowledge to reduce the size of the search space\n- Leveraging existing frameworks, libraries, and tools\n- Looking for pre-existing solutions to problems\n- Focusing on one task at a time\n- Streamlining the data collection and labelling process\n- Setting up automated evaluations and feedback loops\n- Taking breaks to reflect and reassess goals\n- Working in collaboration with others\n\n_____\n\n##### Which metric is correct? ([dehokanta](https://www.kaggle.com/dehokanta))\n\n[dehokanta](https://www.kaggle.com/dehokanta) raised a question on the evaluation metric for this competition: Recall@k.\n\n**[dehokanta](https://www.kaggle.com/dehokanta) asked which of the following two definitions is correct:**\n\n- (1) Recall@k is the fraction of the top k pixels of an image that are the same as the target image.\n- (2) Recall@k is the fraction of the target image that is in the top k pixels of an image.\n\nThe discussion that followed suggested that the **second definition** is the correct one.\n_____\n\n##### 💡How to improve the results of your Approximate Nearest Neighbor search! (annoy) ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) gave an insightful post on [how to improve the results of your Approximate Nearest Neighbor search](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368385)!\n\n**Main Idea:** Use an Approximate Nearest Neighbor (ANN) search to improve the results of your search.\n\n**Benefits:**\n\n- ANN searches are much faster than traditional nearest neighbor searches and can be used to quickly find similar items.\n- ANN searches can also be used to find items with similar characteristics, even if they don’t have an exact match.\n- ANNs are also more scalable, as they can be easily parallelized and distributed across multiple nodes.\n\n**Implementation:**\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) suggests using the [Annoy library](https://github.com/spotify/annoy) to implement the ANN search.\n- Annoy is a Python library that allows you to quickly index and search for points in a vector space.\n- Annoy also includes a few useful features such as the ability to easily control the size of the search pool, as well as the ability to specify the number of neighbors to search for.\n\n**Conclusion:**\n\nUsing an Approximate Nearest Neighbor search can provide a number of benefits, such as increased search speed and scalability. [Radek Osmulski](https://www.kaggle.com/radek1) suggests using the Annoy library to implement the ANN search, which provides a number of useful features for controlling the size of the search pool and specifying the number of neighbors to search for.\n\n_____\n\n##### 💡 Training an XGBoost Ranker on the GPU with Merlin Models 🔥🔥🔥 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368848), [Radek Osmulski](https://www.kaggle.com/radek1) discussed the use of Merlin Models to train an XGBoost Ranker on the GPU.\nThis can greatly reduce the training time for XGBoost models, allowing for faster and more efficient training. The main benefits are that it does not require a Tensorflow backend, and allows for faster training. It also allows for better hyperparameter optimization, since it is possible to train multiple models in parallel. Additionally, the model can be deployed to the cloud for scalability and higher performance. Finally, the model can also be used for inference tasks such as ranking and recommendation.\n\n_____\n\n##### Matrix Factorization with GPU: 6.5x faster! ([CPMP](https://www.kaggle.com/cpmpml))\n\n- [CPMP](https://www.kaggle.com/cpmpml) shared a [very interesting notebook](https://www.kaggle.com/competitions/otto-recommender-system/discussion/371166) to compute matrix factorization, using polar, annoy, Merlin data loader and pytorch.\n- [CPMP](https://www.kaggle.com/cpmpml) then created a GPU version of the notebook which is [available here](https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu) and is 6.5x faster than the original version!\n\n_____\n\n##### 🐘 the elephant in the room -- high cardinality of targets and what to do about this ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364722), [Radek Osmulski](https://www.kaggle.com/radek1) discusses an issue that is often overlooked when working with Machine Learning models, high cardinality of targets and how to handle it. The main idea presented is that, in order to build an effective model, it is important to accurately measure the importance of each feature. The author suggests two main approaches: \n\n- **Feature Engineering:** transforming existing features into a different representation that can be used to better understand the underlying data. \n- **Dimensionality Reduction:** reducing the number of features to a more manageable amount.\n\nThe author also outlines some potential pitfalls, such as overfitting and underfitting, that can arise when dealing with high cardinality. Finally, the post provides some tips on how to effectively use the two main approaches to tackle this problem.\n\n_____\n\n##### Co-Visitation Matrices and Matrix Factorization ([Ravi Shah](https://www.kaggle.com/ravishah1))\n\n[Ravi Shah](https://www.kaggle.com/ravishah1) discussed the use of co-visitation matrices and matrix factorization for recommendation systems in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365589). The main idea is to create features based on items which are frequently viewed together. This technique has been shown to be effective for creating better recommendation systems.\n\n_____\n\n##### 💡What is a good initial goal in the competition? How to improve beyond it? 📈 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368685), [Radek Osmulski](https://www.kaggle.com/radek1) suggested a good initial goal for the competition is to get everything up and running on your end. This includes setting up your environment and running the notebooks. Additionally, [Radek Osmulski](https://www.kaggle.com/radek1) suggested that to improve beyond this goal, it is important to understand the underlying techniques and concepts, such as using the right data structure and algorithm for the task, as well as understanding the limitations of the algorithms. Finally, [Radek Osmulski](https://www.kaggle.com/radek1) recommended debugging your code and understanding the underlying algorithms, such as a greedy approach, to improve beyond the initial goal.\n\n_____\n\n##### To the people still hacking away on this competition... ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) has noticed that there are fewer and fewer submissions on the Leaderboard of [this competition](https://www.kaggle.com/competitions/otto-recommender-system/discussion/369789). He encourages competitors to keep working on their solutions and suggests they explore the following ideas to get better results: \n- Make sure you are using algorithms that are tailored to the problem, such as the Minimum Spanning Tree algorithm.\n- Try out different optimization techniques.\n- Take a close look at your code and identify possible improvements.\n- Use the insights and ideas of other Kagglers. \n- Share your progress and ideas with others.\n\n_____\n\n##### recommenders - Best Practices on Recommendation Systems by Microsoft ([Sinan Calisir](https://www.kaggle.com/snnclsr))\n\n[Sinan Calisir](https://www.kaggle.com/snnclsr) from Microsoft shared a great post discussing best practices when building Recommender Systems. The main points covered include:\n- Data Preprocessing: including data cleaning, normalization, and feature engineering.\n- Model Selection: including selecting the right model for the task, hyperparameter optimization and model validation.\n- Evaluation Metrics: including the use of standard metrics (such as RMSE) as well as business metrics.\n- Deployment & Maintenance: including integrating the model into applications, model retraining and monitoring.\n- User Experience: including recommendations personalization and explainability. \nCheck out the full [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363730) for further details.\n\n_____\n\n##### Why best model might not win RecSys competition ([Roman](https://www.kaggle.com/nroman))\n\n- [Roman](https://www.kaggle.com/nroman) discussed the difference between offline and online metrics in Recommender System (RecSys) competitions.\n- The ground truth labels for RecSys competitions are derived from the organizers' current model. This may mean that the best model might not win the competition, as the model which scores the highest on the offline metric might not be the same as the one which scores the highest on the online metric.\n\n_____\n\n##### Important information regarding test data from competition repository ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) has shared a helpful repository containing important information regarding the test data from the competition.\n- The repository includes a comprehensive README as well as some related code. Check it out [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965)!\n\n_____\n\n##### Otto's Talk on Transformer Recommendation Systems ([BenediktSchifferer](https://www.kaggle.com/benediktschifferer))\n\n[BenediktSchifferer](https://www.kaggle.com/benediktschifferer) presented on the topic of Transformer Recommendation Systems in a [talk](https://www.kaggle.com/competitions/otto-recommender-system/discussion/373224).\n- The main idea of the talk was to explain the use of transformers in recommendation systems and to discuss their potential applications.\n- The talk discussed how transformers can be used to learn user preference and then use them to generate recommendations. It also discussed how transformers can be used to build an end-to-end recommendation pipeline, as well as how they can be used to improve existing models.\n- The talk discussed how transformers can be used to solve cold-start and long-tail problems. The talk highlighted the advantages of using transformers, such as their capability to capture long-term and short-term user preferences, and their ability to handle both structured and unstructured data. It also discussed the challenges of using transformers, such as their large memory requirement and computational cost.\n\n_____\n\n##### RePlay - opensource RecSys constructor ([Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov))\n\n[Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov) has created an opensource RecSys constructor called [RePlay](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365917). It is a library that provides tools for all stages of creating a recommendation system, such as data preprocessing, model evaluation, and comparison. RePlay uses PySpark to handle big data.\n\n_____\n\n##### 💡 ANN -&gt; NN: get better results and run faster with NN search on the GPU 🔥🔥🔥 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) recently shared an amazing post on how to get better results and run faster with NN search on the GPU.\n- [Radek Osmulski](https://www.kaggle.com/radek1) suggested using Matrix Factorization with GPU which is 6.5x faster. Furthermore, [Radek Osmulski](https://www.kaggle.com/radek1) mentioned that this solution was already implemented by their colleagues.\n\n_____\n\n##### Ground-truth? ([Gunes Evitan](https://www.kaggle.com/gunesevitan))\n\n- [Gunes Evitan](https://www.kaggle.com/gunesevitan) asked a question about what \"test set is truncated\" means and if it is possible to create ground-truth for the training set by only one event.\n- The post explains that it is not truncated by a single event, instead it is truncated by the total number of frames in the test set.\n- Therefore, it is not possible to create the ground-truth by a single event, as the test set contains more frames than the training set.\n\n_____\n\n##### The time zone in Germany is UTC+2 in August ([aldparis](https://www.kaggle.com/adaubas))\n\n- [aldparis](https://www.kaggle.com/adaubas) discussed the time zone in Germany being UTC+2 in August.\n- This is due to Daylight Saving Time (DST) which is observed in Germany from the last Sunday in March until the last Sunday in October.\n- During DST, the clocks are moved forward by one hour, resulting in the time zone shifting from UTC+1 to UTC+2. This means that the daylight hours will be longer in the summer months, with the sun setting later in the evening.\n\n_____\n\n##### Ranker models vs. Binary classification models ([Andrej Zubaľ](https://www.kaggle.com/andrejzuba))\n\n[Andrej Zubaľ](https://www.kaggle.com/andrejzuba) discussed the differences between ranker models and binary classification models in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370502).\n- **Ranker models** were found to be more accurate for predicting the probability of a certain event or outcome, as they can rank a list of outcomes from most to least likely.\n- **Binary classification** models are simpler and can be used to classify events into two categories, such as yes or no, true or false.\n- **Ranker models** are better suited for prediction tasks such as predicting the relevance of a search query, while binary classification models are better suited for categorizing events.\n_____\n\n##### How do you train Ranking model? ([dehokanta](https://www.kaggle.com/dehokanta))\n\n[dehokanta](https://www.kaggle.com/dehokanta) asked a question about how to build a ranking model. In the post, they mentioned that they have tried various methods but with not much success.\n\n**This post provides useful information on how to train a ranking model, such as:**\n- Understanding how Ranking works\n- Normalising features\n- Feature selection\n- Data Augmentation\n- Training and Tuning\n- Evaluation Metric Selection\n- Deployment and Monitoring\n\n_____\n\n##### 💡How to ensemble predictions -- a key component to every strong solution 🏅 ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n[Radek Osmulski](https://www.kaggle.com/radek1) in this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368747) discussed the importance of ensembling for a strong solution in Kaggle competitions. Ensembling is combining multiple models to get a better result than the individual models. \n\nThe main idea is to create multiple models that are independent from each other and then combine them to get better results.\nThe author suggests combining models that have different architectures, different pre-processing techniques and different hyperparameters.\n\n**The author also outlines some techniques for ensembling such as:**\n- Averaging\n- Weighted Averaging\n- Model Stacking\n- Bagging and Boosting\n\n**The author also provides some useful tips to keep in mind while ensembling such as:**\n- Identify the best models and ensemble them\n- Try using different weights for each model\n- Use different seed values for each model\n- Use different input data to train each model\n- Use regularization to reduce overfitting\n- Use cross-validation to evaluate the models.\n\nEnsembling is an important technique to improve the accuracy of models and can be used to achieve higher scores in Kaggle competitions.\n\n_____\n\n##### Guidance on what algos to try from one of the authors of Microsoft Recommenders repo ([Miguel Fierro](https://www.kaggle.com/hoaphumanoid))\n\n[Miguel Fierro](https://www.kaggle.com/hoaphumanoid) is one of the authors of the [Microsoft Recommenders repo](https://github.com/microsoft/recommenders/).\n- In this post they provide guidance on which algorithms to try depending on the problem.\n- They mention to start with simple algorithms like popularity, content-based filtering, and matrix factorization.\n- They also suggest to look into user-based collaborative filtering and embeddings. Finally, they suggest to try deep learning methods such as deep matrix factorization, deep autoencoders, and convolutional neural networks.\n\n_____\n\n##### Graph Neural Network, is complex to model but efficient for this task ([Youness EL BRAG](https://www.kaggle.com/younesselbrag))\n\n- [Youness EL BRAG](https://www.kaggle.com/younesselbrag) presented a new approach, [Session-based Recommendation with Graph Neural Networks](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368355), to the task of session-based recommendation.\n- This method models a session as a complex transition of items and estimates user representations in addition to item representations. An accompanying [code](https://github.com/CRIPAC-DIG/SR-GNN?utm_source=catalyzex.com) was also released which enables users to implement the method in their own projects.\n\n_____\n\n##### cuDF vs Pandas Vs Modin for faster data processing and there comparison ([Satya](https://www.kaggle.com/satyaprakashshukl))\n\n- [Satya](https://www.kaggle.com/satyaprakashshukl) discussed different methods for faster data processing and their comparison and speed on this competition.\n- Really useful post for those who are looking for faster data processing methods.\n- Check out the [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366258) for more information.\n\n_____\n\n##### Some Interesting Times Series on Products ([aldparis](https://www.kaggle.com/adaubas))\n\n[aldparis](https://www.kaggle.com/adaubas) recently [shared](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368373) an interesting discussion topic about product times series.\n**The main idea behind it is to use the sales data of products to make predictions and forecast future sales.**\n\n**The post contains a few insights such as:**\n- Times series can be used to forecast future sales more accurately. \n- Seasonality can be taken into account to make more precise predictions.\n- The need to use the most recent data available to make more reliable predictions.\n- Machine Learning algorithms can be used to model times series data.\n- Using multiple linear models can improve accuracy.\n- Different methods of smoothing can be applied to times series data.\n\n_____\n\n##### How can a session start with an order? ([CPMP](https://www.kaggle.com/cpmpml))\n\n- [CPMP](https://www.kaggle.com/cpmpml) posed a question about how to start a session with an order.\n- He suggested that it could be done by creating a queue of tasks and using a scheduler to set the order of tasks and limit the number of tasks run in parallel.\n- Additionally, they suggested the use of a job server that can manage the scheduling. Also, he discussed the use of a separate service for handling long-running tasks, such as background tasks.\n_____\n\n##### Difference between Co-visitation Matrix v.s. Matrix Factorization ([bilzard](https://www.kaggle.com/tatamikenn))\n- [bilzard](https://www.kaggle.com/tatamikenn) discussed the difference between the co-visitation matrix approach, which can be found [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725851), and the matrix decomposition approach, which can be found [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/3725852).\n- The decisive difference between them is the dependence of events, which could be called the context in consideration. This could explain why the co-visitation matrix approach scores higher than the matrix decomposition approach.\n\n_____\n\n##### Some concerns about validation ([Gunes Evitan](https://www.kaggle.com/gunesevitan))\n\n- [Gunes Evitan](https://www.kaggle.com/gunesevitan) noticed that most people were using the test set split code from the organizers' [repository](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370384)\n- But has some concerns that it might have flaws.\n-  He suggests using an alternative method to split the data, such as using a train/validation/test split, which would allow for better validation. This could help avoid overfitting and could potentially lead to better results. He also encourages people to think about the validation method and consider different approaches.\n\n_____\n\n##### We can try more NLP methods in sequence recomendation, like .... ([KKY](https://www.kaggle.com/evilpsycho42))\n\n- In this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367870), [KKY](https://www.kaggle.com/evilpsycho42) discussed the possibility of using NLP methods to improve sequence recommendation in this user-anonymized dataset.\n- They suggested trying out various natural language processing (NLP) techniques to get a better representation of the action/item sequence.\n- Some ideas they proposed include using Long Short Term Memory (LSTM) networks, convolutional neural networks (CNNs), and transformer networks.\n- Additionally, they suggested using methods such as attention mechanisms, bidirectional networks, and sequence-to-sequence learning.\n\n_____\n\n##### Summary about Loading and Preprocessing Big Jsonl Data File ([Lei Wang](https://www.kaggle.com/leiwong))\n\n- [Lei Wang](https://www.kaggle.com/leiwong) shared two notebooks to help with loading and preprocessing data for this competition.\n- **Loading:** The notebook explains how to load big jsonl data file into a pandas dataframe with different methods and also provides a comparison between them. \n- **Preprocessing:** This notebook explains how to convert the data into csv, parquet or a dataframe, as well as the advantages and disadvantages of each.\n\n_____\n\n##### Since there are only few features(only time info), any chance to use ML algorithm? ([cocoshe](https://www.kaggle.com/cocoshe))\n\n- [cocoshe](https://www.kaggle.com/cocoshe) recently posted a discussion about the possibility of using a Machine Learning (ML) algorithm for this competition.\n- They observed that all the recall methods currently used involve \"top-k frequency aids of different session\" and \"top-k frequency aids of different type\", or a \"top-k frequency of all aids\" counting method (co-occurrence matrix). Additionally, padding is used to fill the recall@20.\n- [cocoshe](https://www.kaggle.com/cocoshe) then asked if any chance to use ML algorithm existed, despite the few features that are available.\n\n_____\n\n##### Q about the Data ([Aviel Danin](https://www.kaggle.com/avieldanin))\n\n[Aviel Danin](https://www.kaggle.com/avieldanin) asked a question about the data in [this post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366478):\n\n- In some cases there is a \"cart event\" of a certain aid followed by a \"click\", and then another \"cart event\" of the same aid followed by another \"click\". \n- In other cases, there is a \"cart event\" without a click before it. \n\nThese scenarios could indicate that there are multiple processes related to the data, and it's important to understand them in order to get the most accurate results.\n\n**Answer (By ([roncadr](https://www.kaggle.com/danieleroncaglioni)):**\n```\nYeah I noticed the same, also in the full train dataset there are 16.896.191 carts, while there are less (12.142.933) instances of an aid of some type being followed by the same aid and type \"cart\", so definetely, apparently it is possible to add an item to the cart without having to click that item immediately beforehand. Perhaps when you are browsing on multiple browser tabs and switch between them….?\n```\n_____\n\n##### 💡 How to split the data for training a two-stage recommender? ([Radek Osmulski](https://www.kaggle.com/radek1))\n\n- [Radek Osmulski](https://www.kaggle.com/radek1) discussed the best way to split data when training a two-stage recommender system.\n- He shared his work on [this](https://www.kaggle.com/code/radek1/eda-a-look-at-the-data-training-splits) notebook for anyone interested in learning more about the topic.\n_____\n\n##### How did you read the data? ([takuma0306](https://www.kaggle.com/takuma0306))\n\nIn this [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/372990) [takuma0306](https://www.kaggle.com/takuma0306) asked how others were loading the training data for the competition.\nThey mentioned that when using VS Code, their application was terminating halfway through and when using Google Colab they were running out of system RAM and losing the session. They were wondering how others were successfully loading the data and asked for help.\n\n**Answer (By ([Giannis](https://www.kaggle.com/ikogias))):**\n```\nAs an alternative approach I have created the following PySpark script, which reads-in the JSON files, converts them to a tabular PySpark dataframe and outputs the files as parquet. Using Spark, this code is able to run with a simple Kaggle kernel (in about an hour) w/o running into memory issues.\n- https://www.kaggle.com/code/ikogias/pre-process-with-pyspark-into-parquet\n```\n_____\n\n##### Test Set Clarification Questions ([dpalbrecht](https://www.kaggle.com/dpalbrecht))\n\n- [dpalbrecht](https://www.kaggle.com/dpalbrecht) raised some questions regarding the test set on the [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365554) in the repo.\n- The sentence that confused them was: \"The test set contains 1000 user queries and the ground truth is not known.\"\n\n**The questions raised were:**\n- How the queries are produced?\n- What does \"ground truth\" refer to?\n- Does it mean that for each user query there is a corresponding ground truth image?\n\n**Answer (By [Pietro Maldini](https://www.kaggle.com/pietromaldini1)):**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6043604%2F957ed2d68815e3db8abef7b03304a135%2Fground_truth.png?generation=1668199923227677&alt=media)\n\n_____\n\n\n##### The relatively less data for each test session ([samson fha](https://www.kaggle.com/samsonfha))\n\n- [samson fha](https://www.kaggle.com/samsonfha) discussed the issue of the relatively low amount of data for each test session compared to training sessions, noting that it is difficult to identify user features and generate labels for them.\n- They suggested focusing more on generating labels for goods/aids instead.\n\n**Answer (By [Radek Osmulski](https://www.kaggle.com/radek1))**\nThe reason for the shorter sessions is explained [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363554#2015486) by the organizer.\n\n_____\n\n##### How many candidates you used? ([L0Z1K](https://www.kaggle.com/baekseungyun))\n\n- [L0Z1K](https://www.kaggle.com/baekseungyun) suggests that sharing the number of candidates and recall score can help the community in their analysis of the Santa 2020 competition.\n- This could provide valuable insights into the effectiveness of certain strategies and help to identify potential improvements.\n\n_____\n\n##### Unseen Aids in Test Set ([moth](https://www.kaggle.com/alejopaullier))\n\n- [moth](https://www.kaggle.com/alejopaullier) made an interesting observation in their [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156) about the test set for the competition.\n- They noticed that many items in the test set do not appear in the train set, which can make it difficult to accurately predict the outcomes.\n- **This is an important reminder to be mindful of the unseen data that may exist in the test set when creating models.**\n\n_____\n\n##### Otto recommender Kernel Stats ([Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov))\n\n- [Alexander Ryzhkov](https://www.kaggle.com/alexryzhkov) shared an amazing post about notebook analytics!\n- Such as the number of upvotes or a fork graph.\n\n**Really cool idea!**\nmount of data and also reduces the amount of time needed to train a model. Additionally, it can also reduce the likelihood of overfitting due to the smaller data size.",
    "2076639": "Amazing! Thanks for sharing this with us!",
    "2076928": "Great, so detailed !! Thanks for sharing!",
    "2086845": "This is amazing, Thanks a lot for sharing!",
    "2087467": "So detailed and helpful!!! Thanks for your sharing!",
    "2088258": "Great information! Thanks for your summary!",
    "2093188": "This is amazing and helpful, thanks a lot for sharing!",
    "2095636": "This is amazing and helpful! Thanks for sharing!",
    "2105455": "thedevastator Wow, super cool! Thank you so much for sharing this wonderful post with summarizing all these wonderful resources! I was intended to a similar summary like this in the beginning but failed. It's wonderful you are sharing this with all of us!!! 🙏🙏🙏"
  },
  "source": "meta"
}