{
  "id": 379291,
  "title": "A growing reflections on otto from a beginner's perspective",
  "url": "/competitions/otto-recommender-system/discussion/379291",
  "author_name": "",
  "post_date": "2023-01-19T02:20:12.656418Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<h4></h4>\n<ul>\n<li>writing down my reflections of learning is good for myself and others who share similar experiences</li>\n<li>I am a beginner who learns slowly and can't keep up with all the good and new posts and notebooks, so I will take my own pace</li>\n<li>otto is a worthwhile comp which I will keep learning even after the deadline</li>\n<li>So, this reflection will grow as I keep learning even after the deadline.</li>\n<li>for a quick sum for most if not all amazing posts/notebooks, please check out <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> 's \"One Month Left - Here is what you need to know!\" <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/374229#2105455\" target=\"_blank\">post</a> 🔥🔥🔥🔥</li>\n</ul>\n<h4></h4>\n<p><strong>How the official OTTO dataset repo page describe it</strong></p>\n<blockquote>\n  <ul>\n  <li>12M real-world anonymized user sessions</li>\n  <li>220M events, consiting of&nbsp;<code>clicks</code>,&nbsp;<code>carts</code>&nbsp;and&nbsp;<code>orders</code></li>\n  <li>1.8M unique articles in the catalogue</li>\n  </ul>\n</blockquote>\n<h4></h4>\n<h5>What exactly does this comp want us to predict? 💡💡💡</h5>\n<p>thanks 🙏 to <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> for pointing out the misleading part of my previous description here, below is my second attempt.</p>\n<h5>The official OTTO dataset repo <a href=\"https://github.com/otto-de/recsys-dataset#evaluation\" target=\"_blank\">page</a> describe the prediction task very precisely actually</h5>\n<blockquote>\n  <p>For each&nbsp;<code>session</code>&nbsp;in the test data, your task it to predict the&nbsp;<code>aid</code>&nbsp;values for each&nbsp;<code>type</code>&nbsp;that occur after the last timestamp&nbsp;<code>ts</code>&nbsp;the test session. In other words, the test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation.</p>\n  <p>For&nbsp;<code>clicks</code>&nbsp;there is only a single ground truth value for each session, which is the next&nbsp;<code>aid</code>&nbsp;clicked during the session (although you can still predict up to 20&nbsp;<code>aid</code>&nbsp;values). The ground truth for&nbsp;<code>carts</code>&nbsp;and&nbsp;<code>orders</code>&nbsp;contains all&nbsp;<code>aid</code>&nbsp;values that were added to a cart and ordered respectively during the session.</p>\n</blockquote>\n<h5>Another rephrase to help understanding hopefully</h5>\n<ul>\n<li>I am given a test session which has been truncated at certain timestamp</li>\n<li>The ground truth includes only the 1st clicked aid for each test session after the timestamp above, but I am given 20 chances to get it right</li>\n<li>The ground truth has all the carted aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth carted aids are more than 20, I can still score full as long as I can get 20 of those carted aids right.</li>\n<li>The ground truth has all the ordered aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth ordered aids are more than 20, I can still score full as long as I can get 20 of those ordered aids right.</li>\n</ul>\n<p>The detailed rephase above is derived from the <code>Recall@20</code> metric below</p>\n<h4></h4>\n<h5>How calc the <code>Recall@20</code> metric to score a type of all test sessions</h5>\n<p>$$<br>\nR_{type} = \\frac{ \\sum\\limits_{i=1}^N | \\{ \\text{predicted aids} \\}_{i, type} \\cap \\{ \\text{ground truth aids} \\}_{i, type} | }{ \\sum\\limits_{i=1}^N \\min{( 20, | \\{ \\text{ground truth aids} \\}_{i, type} | )}}<br>\n$$</p>\n<h5>How to add up 3 different type scores above to get the total score</h5>\n<p>$$<br>\nscore = 0.10 \\cdot R_{clicks} + 0.30 \\cdot R_{carts} + 0.60 \\cdot R_{orders}<br>\n$$</p>\n<h4></h4>\n<h5><strong>The official OTTO dataset repo has provided 3 tables for us</strong></h5>\n<p>The formula and meaning of <code>density</code> is answered by the organizer <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>  <a href=\"https://github.com/otto-de/recsys-dataset/issues/2#issuecomment-1313697705\" target=\"_blank\">here</a> </p>\n<table>\n<thead>\n<tr>\n<th>dataset</th>\n<th>num_sessions</th>\n<th>num_items</th>\n<th>num_events</th>\n<th>num_clicks</th>\n<th>num_carts</th>\n<th>num_orders</th>\n<th>Density</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train</td>\n<td>12_899_779</td>\n<td>1_855_603</td>\n<td>216_716_096</td>\n<td>194_720_954</td>\n<td>16_896_191</td>\n<td>5_098_951</td>\n<td>0.0005</td>\n</tr>\n<tr>\n<td>Test</td>\n<td>1_671_803</td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mean</th>\n<th>std</th>\n<th>min</th>\n<th>50%</th>\n<th>75%</th>\n<th>90%</th>\n<th>95%</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train num_events per session</td>\n<td>16.80</td>\n<td>33.58</td>\n<td>2</td>\n<td>6</td>\n<td>15</td>\n<td>39</td>\n<td>68</td>\n<td>500</td>\n</tr>\n<tr>\n<td>Test num_events per session</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mean</th>\n<th>std</th>\n<th>min</th>\n<th>50%</th>\n<th>75%</th>\n<th>90%</th>\n<th>95%</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train num_events per item</td>\n<td>116.79</td>\n<td>728.85</td>\n<td>3</td>\n<td>20</td>\n<td>56</td>\n<td>183</td>\n<td>398</td>\n<td>129004</td>\n</tr>\n<tr>\n<td>Test num_events per item</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n</tr>\n</tbody>\n</table>\n<h5>What insights can we derive?</h5>\n<h4></h4>\n<h5>What is Candidate ReRank model? 💡💡💡</h5>\n<ul>\n<li>first we use a model to select hundreds of aid candidates, then we use another model to select the final 20 aids for predictions</li>\n</ul>\n<h5>Why Chris Deotte said Candidate ReRank model will most likely to win this comp? 💡💡💡</h5>\n<ul>\n<li>maybe it is the most obvious approach and it works in other RecSys comp, according to Chris Deotte</li>\n<li>there are 1.8 milliion aids to choose from, but only need 20 aids for each session_type</li>\n<li>It's kind of making sense to break a large problem into 2 smaller problems: choose hundreds from millions in one model, and then choose 20 from hundreds in another</li>\n<li>but how to select hundreds from millions? through similarities? <ul>\n<li>great kagglers in otto have shared approaches like co-visitation matrix, word2vec, matrix factorization, and maybe more I don't know</li></ul></li>\n</ul>\n<h4></h4>\n<h5>Radek explains it way better than my own reflection below</h5>\n<ul>\n<li>\"A co-visitation matrix counts the co-occurrence of two actions in close proximity.\"</li>\n<li>\"If a user bought A and shortly after bought B, we store these values together.\"</li>\n<li>\"We calculate counts and use them to estimate the probability of future actions based on recent history.\"</li>\n<li>\"It is quite important to understand what is happening in the co-visitation matrix approach…\"</li>\n<li>\"Since it suffers from the same issues as our trigram example!\"</li>\n<li>\"Plus what does the co-visitation matrix resemble?\"</li>\n<li>\"You are right, it is akin to doing Matrix Factorization by counting!\"</li>\n<li>\"It is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏\"</li>\n</ul>\n<h5>How to understand covisitation matrix intuitively? 💡💡💡</h5>\n<ul>\n<li>it’s a way to take any aid and find any number of aids which are most similar to it</li>\n<li>co-visitation matrices differentiate from each other based on how they define similarity or how they select aids to be paired together</li>\n</ul>\n<h5>if you were to play the role of the inventor of co-visitation matrix, what series of ideas/questions could trigger the creation of it?</h5>\n<ul>\n<li>Is there any relationship between one aid/product with other aids/products in the same session or across all sessions?</li>\n<li>Are there some aids more similar to some and more different to others?</li>\n<li>Could it be possible when this aid is viewed, some aids are more likely to be clicked/carted/ordered than other aids?</li>\n<li>Could we pair aids together for each and every session and count the occurrences of pairs?</li>\n<li>Since one aid (eg., '122') could have many pair-partners, by counting the occurrences of the pairs ('122', pair-partner), could we find the most common pair-partners of aid '122'?</li>\n<li>Could the next clicks or carts or orders be the most common pair-partners of the last aid (or all aids) of a test session?</li>\n</ul>\n<h5>what does pairing logic or a definition of simiarlity look like</h5>\n<ul>\n<li>In Radek's notebook, the pairing logic is the following</li>\n<li>use only the last 30 aids of each session to pair on each other with pandas <code>merge</code> or polars <code>join</code> on <code>session</code></li>\n<li>remove the pairs of same partners</li>\n<li>keep pairs whose right-partner is after left-partner within a day</li>\n</ul>\n<h5>We can tweak the pairing logic to change our co-visitation matrix</h5>\n<h4></h4>\n<h5>What shortcomings does co-visitation matrix have? Could Word2Vec be a better model?</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358#2105430\" target=\"_blank\">asked &amp; answered</a> Thank you <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for your insightful reply again!</li>\n</ul>",
  "messages": [
    {
      "id": "2106180",
      "postDate": "01/19/2023 02:20:12",
      "content": "<h4></h4>\n<ul>\n<li>writing down my reflections of learning is good for myself and others who share similar experiences</li>\n<li>I am a beginner who learns slowly and can't keep up with all the good and new posts and notebooks, so I will take my own pace</li>\n<li>otto is a worthwhile comp which I will keep learning even after the deadline</li>\n<li>So, this reflection will grow as I keep learning even after the deadline.</li>\n<li>for a quick sum for most if not all amazing posts/notebooks, please check out <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> 's \"One Month Left - Here is what you need to know!\" <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/374229#2105455\" target=\"_blank\">post</a> 🔥🔥🔥🔥</li>\n</ul>\n<h4></h4>\n<p><strong>How the official OTTO dataset repo page describe it</strong></p>\n<blockquote>\n  <ul>\n  <li>12M real-world anonymized user sessions</li>\n  <li>220M events, consiting of&nbsp;<code>clicks</code>,&nbsp;<code>carts</code>&nbsp;and&nbsp;<code>orders</code></li>\n  <li>1.8M unique articles in the catalogue</li>\n  </ul>\n</blockquote>\n<h4></h4>\n<h5>What exactly does this comp want us to predict? 💡💡💡</h5>\n<p>thanks 🙏 to <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> for pointing out the misleading part of my previous description here, below is my second attempt.</p>\n<h5>The official OTTO dataset repo <a href=\"https://github.com/otto-de/recsys-dataset#evaluation\" target=\"_blank\">page</a> describe the prediction task very precisely actually</h5>\n<blockquote>\n  <p>For each&nbsp;<code>session</code>&nbsp;in the test data, your task it to predict the&nbsp;<code>aid</code>&nbsp;values for each&nbsp;<code>type</code>&nbsp;that occur after the last timestamp&nbsp;<code>ts</code>&nbsp;the test session. In other words, the test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation.</p>\n  <p>For&nbsp;<code>clicks</code>&nbsp;there is only a single ground truth value for each session, which is the next&nbsp;<code>aid</code>&nbsp;clicked during the session (although you can still predict up to 20&nbsp;<code>aid</code>&nbsp;values). The ground truth for&nbsp;<code>carts</code>&nbsp;and&nbsp;<code>orders</code>&nbsp;contains all&nbsp;<code>aid</code>&nbsp;values that were added to a cart and ordered respectively during the session.</p>\n</blockquote>\n<h5>Another rephrase to help understanding hopefully</h5>\n<ul>\n<li>I am given a test session which has been truncated at certain timestamp</li>\n<li>The ground truth includes only the 1st clicked aid for each test session after the timestamp above, but I am given 20 chances to get it right</li>\n<li>The ground truth has all the carted aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth carted aids are more than 20, I can still score full as long as I can get 20 of those carted aids right.</li>\n<li>The ground truth has all the ordered aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth ordered aids are more than 20, I can still score full as long as I can get 20 of those ordered aids right.</li>\n</ul>\n<p>The detailed rephase above is derived from the <code>Recall@20</code> metric below</p>\n<h4></h4>\n<h5>How calc the <code>Recall@20</code> metric to score a type of all test sessions</h5>\n<p>$$<br>\nR_{type} = \\frac{ \\sum\\limits_{i=1}^N | \\{ \\text{predicted aids} \\}_{i, type} \\cap \\{ \\text{ground truth aids} \\}_{i, type} | }{ \\sum\\limits_{i=1}^N \\min{( 20, | \\{ \\text{ground truth aids} \\}_{i, type} | )}}<br>\n$$</p>\n<h5>How to add up 3 different type scores above to get the total score</h5>\n<p>$$<br>\nscore = 0.10 \\cdot R_{clicks} + 0.30 \\cdot R_{carts} + 0.60 \\cdot R_{orders}<br>\n$$</p>\n<h4></h4>\n<h5><strong>The official OTTO dataset repo has provided 3 tables for us</strong></h5>\n<p>The formula and meaning of <code>density</code> is answered by the organizer <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>  <a href=\"https://github.com/otto-de/recsys-dataset/issues/2#issuecomment-1313697705\" target=\"_blank\">here</a> </p>\n<table>\n<thead>\n<tr>\n<th>dataset</th>\n<th>num_sessions</th>\n<th>num_items</th>\n<th>num_events</th>\n<th>num_clicks</th>\n<th>num_carts</th>\n<th>num_orders</th>\n<th>Density</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train</td>\n<td>12_899_779</td>\n<td>1_855_603</td>\n<td>216_716_096</td>\n<td>194_720_954</td>\n<td>16_896_191</td>\n<td>5_098_951</td>\n<td>0.0005</td>\n</tr>\n<tr>\n<td>Test</td>\n<td>1_671_803</td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mean</th>\n<th>std</th>\n<th>min</th>\n<th>50%</th>\n<th>75%</th>\n<th>90%</th>\n<th>95%</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train num_events per session</td>\n<td>16.80</td>\n<td>33.58</td>\n<td>2</td>\n<td>6</td>\n<td>15</td>\n<td>39</td>\n<td>68</td>\n<td>500</td>\n</tr>\n<tr>\n<td>Test num_events per session</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mean</th>\n<th>std</th>\n<th>min</th>\n<th>50%</th>\n<th>75%</th>\n<th>90%</th>\n<th>95%</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Train num_events per item</td>\n<td>116.79</td>\n<td>728.85</td>\n<td>3</td>\n<td>20</td>\n<td>56</td>\n<td>183</td>\n<td>398</td>\n<td>129004</td>\n</tr>\n<tr>\n<td>Test num_events per item</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n<td>TBA</td>\n</tr>\n</tbody>\n</table>\n<h5>What insights can we derive?</h5>\n<h4></h4>\n<h5>What is Candidate ReRank model? 💡💡💡</h5>\n<ul>\n<li>first we use a model to select hundreds of aid candidates, then we use another model to select the final 20 aids for predictions</li>\n</ul>\n<h5>Why Chris Deotte said Candidate ReRank model will most likely to win this comp? 💡💡💡</h5>\n<ul>\n<li>maybe it is the most obvious approach and it works in other RecSys comp, according to Chris Deotte</li>\n<li>there are 1.8 milliion aids to choose from, but only need 20 aids for each session_type</li>\n<li>It's kind of making sense to break a large problem into 2 smaller problems: choose hundreds from millions in one model, and then choose 20 from hundreds in another</li>\n<li>but how to select hundreds from millions? through similarities? <ul>\n<li>great kagglers in otto have shared approaches like co-visitation matrix, word2vec, matrix factorization, and maybe more I don't know</li></ul></li>\n</ul>\n<h4></h4>\n<h5>Radek explains it way better than my own reflection below</h5>\n<ul>\n<li>\"A co-visitation matrix counts the co-occurrence of two actions in close proximity.\"</li>\n<li>\"If a user bought A and shortly after bought B, we store these values together.\"</li>\n<li>\"We calculate counts and use them to estimate the probability of future actions based on recent history.\"</li>\n<li>\"It is quite important to understand what is happening in the co-visitation matrix approach…\"</li>\n<li>\"Since it suffers from the same issues as our trigram example!\"</li>\n<li>\"Plus what does the co-visitation matrix resemble?\"</li>\n<li>\"You are right, it is akin to doing Matrix Factorization by counting!\"</li>\n<li>\"It is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏\"</li>\n</ul>\n<h5>How to understand covisitation matrix intuitively? 💡💡💡</h5>\n<ul>\n<li>it’s a way to take any aid and find any number of aids which are most similar to it</li>\n<li>co-visitation matrices differentiate from each other based on how they define similarity or how they select aids to be paired together</li>\n</ul>\n<h5>if you were to play the role of the inventor of co-visitation matrix, what series of ideas/questions could trigger the creation of it?</h5>\n<ul>\n<li>Is there any relationship between one aid/product with other aids/products in the same session or across all sessions?</li>\n<li>Are there some aids more similar to some and more different to others?</li>\n<li>Could it be possible when this aid is viewed, some aids are more likely to be clicked/carted/ordered than other aids?</li>\n<li>Could we pair aids together for each and every session and count the occurrences of pairs?</li>\n<li>Since one aid (eg., '122') could have many pair-partners, by counting the occurrences of the pairs ('122', pair-partner), could we find the most common pair-partners of aid '122'?</li>\n<li>Could the next clicks or carts or orders be the most common pair-partners of the last aid (or all aids) of a test session?</li>\n</ul>\n<h5>what does pairing logic or a definition of simiarlity look like</h5>\n<ul>\n<li>In Radek's notebook, the pairing logic is the following</li>\n<li>use only the last 30 aids of each session to pair on each other with pandas <code>merge</code> or polars <code>join</code> on <code>session</code></li>\n<li>remove the pairs of same partners</li>\n<li>keep pairs whose right-partner is after left-partner within a day</li>\n</ul>\n<h5>We can tweak the pairing logic to change our co-visitation matrix</h5>\n<h4></h4>\n<h5>What shortcomings does co-visitation matrix have? Could Word2Vec be a better model?</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358#2105430\" target=\"_blank\">asked &amp; answered</a> Thank you <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for your insightful reply again!</li>\n</ul>",
      "rawMarkdown": "#### <mark style=\"background: #FFB86CA6;\">Why?</mark> \n\n- writing down my reflections of learning is good for myself and others who share similar experiences\n- I am a beginner who learns slowly and can't keep up with all the good and new posts and notebooks, so I will take my own pace\n- otto is a worthwhile comp which I will keep learning even after the deadline\n- So, this reflection will grow as I keep learning even after the deadline.\n- for a quick sum for most if not all amazing posts/notebooks, please check out @thedevastator 's \"One Month Left - Here is what you need to know!\" [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/374229#2105455) 🔥🔥🔥🔥\n\n\n#### <mark style=\"background: #FFB86CA6;\">What inside the dataset</mark> \n\n**How the official OTTO dataset repo page describe it**\n\n> -   12M real-world anonymized user sessions\n> -   220M events, consiting of `clicks`, `carts` and `orders`\n> -   1.8M unique articles in the catalogue\n\n\n\n#### <mark style=\"background: #FFB86CA6;\">What to predict</mark> \n\n##### What exactly does this comp want us to predict? 💡💡💡\n\nthanks 🙏 to [@artemfedorov](https://www.kaggle.com/artemfedorov) for pointing out the misleading part of my previous description here, below is my second attempt.\n\n##### The official OTTO dataset repo [page](https://github.com/otto-de/recsys-dataset#evaluation) describe the prediction task very precisely actually\n\n> For each `session` in the test data, your task it to predict the `aid` values for each `type` that occur after the last timestamp `ts` the test session. In other words, the test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation.\n\n> For `clicks` there is only a single ground truth value for each session, which is the next `aid` clicked during the session (although you can still predict up to 20 `aid` values). The ground truth for `carts` and `orders` contains all `aid` values that were added to a cart and ordered respectively during the session.\n\n##### Another rephrase to help understanding hopefully\n- I am given a test session which has been truncated at certain timestamp\n- The ground truth includes only the 1st clicked aid for each test session after the timestamp above, but I am given 20 chances to get it right\n- The ground truth has all the carted aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth carted aids are more than 20, I can still score full as long as I can get 20 of those carted aids right.\n- The ground truth has all the ordered aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth ordered aids are more than 20, I can still score full as long as I can get 20 of those ordered aids right.\n\nThe detailed rephase above is derived from the `Recall@20` metric below\n\n\n#### <mark style=\"background: #FFB86CA6;\">How to evaluate predictions</mark> \n\n##### How calc the `Recall@20` metric to score a type of all test sessions\n\n$$\nR_{type} = \\frac{ \\sum\\limits_{i=1}^N | \\\\{ \\text{predicted aids} \\\\}\\_{i, type} \\cap \\\\{ \\text{ground truth aids} \\\\}\\_{i, type} | }{ \\sum\\limits_{i=1}^N \\min{( 20, | \\\\{ \\text{ground truth aids} \\\\}_{i, type} | )}}\n$$\n\n##### How to add up 3 different type scores above to get the total score\n\n$$\nscore = 0.10 \\cdot R_{clicks} + 0.30 \\cdot R_{carts} + 0.60 \\cdot R_{orders}\n$$\n\n\n\n\n#### <mark style=\"background: #FFB86CA6;\">What insights can be derived from EDA</mark> \n\n##### **The official OTTO dataset repo has provided 3 tables for us**\n\nThe formula and meaning of `density` is answered by the organizer @pnormann  [here](https://github.com/otto-de/recsys-dataset/issues/2#issuecomment-1313697705) \n\n| dataset | num_sessions | num_items | num_events  | num_clicks  | num_carts  | num_orders | Density |\n| ------- | ------------ | --------- | ----------- | ----------- | ---------- | ---------- | ------- |\n| Train   | 12_899_779   | 1_855_603 | 216_716_096 | 194_720_954 | 16_896_191 | 5_098_951  | 0.0005  | \n| Test    |  1_671_803            |           |             |             |            |            |         |\n\n|                           |  mean |   std |  min |  50% |  75% |  90% |  95% |  max |\n| :------------------------ | ----: | ----: | ---: | ---: | ---: | ---: | ---: | ---: |\n| Train num_events per session | 16.80 | 33.58 |    2 |    6 |   15 |   39 |   68 |  500 |\n| Test num_events per session | TBA | TBA | TBA | TBA | TBA | TBA | TBA | TBA |\n\n\n|                        |   mean |    std |  min |  50% |  75% |  90% |  95% |    max |\n| :--------------------- | -----: | -----: | ---: | ---: | ---: | ---: | ---: | -----: |\n| Train num_events per item | 116.79 | 728.85 |    3 |   20 |   56 |  183 |  398 | 129004 |\n| Test num_events per item  |    TBA |    TBA |  TBA |  TBA |  TBA |  TBA |  TBA |    TBA |\n\n##### What insights can we derive?\n\n\n\n#### <mark style=\"background: #FFB86CA6;\">Candidate ReRank Model</mark> \n\n\n##### What is Candidate ReRank model? 💡💡💡\n- first we use a model to select hundreds of aid candidates, then we use another model to select the final 20 aids for predictions\n\n##### Why Chris Deotte said Candidate ReRank model will most likely to win this comp? 💡💡💡\n- maybe it is the most obvious approach and it works in other RecSys comp, according to Chris Deotte\n- there are 1.8 milliion aids to choose from, but only need 20 aids for each session_type\n- It's kind of making sense to break a large problem into 2 smaller problems: choose hundreds from millions in one model, and then choose 20 from hundreds in another\n- but how to select hundreds from millions? through similarities? \n\t- great kagglers in otto have shared approaches like co-visitation matrix, word2vec, matrix factorization, and maybe more I don't know\n\n\n\n####  <mark style=\"background: #FFB86CA6;\">Co-visitation Matrix</mark> \n \n\n##### Radek explains it way better than my own reflection below\n- \"A co-visitation matrix counts the co-occurrence of two actions in close proximity.\"\n- \"If a user bought A and shortly after bought B, we store these values together.\"\n- \"We calculate counts and use them to estimate the probability of future actions based on recent history.\"\n- \"It is quite important to understand what is happening in the co-visitation matrix approach…\"\n- \"Since it suffers from the same issues as our trigram example!\"\n- \"Plus what does the co-visitation matrix resemble?\"\n- \"You are right, it is akin to doing Matrix Factorization by counting!\"\n- \"It is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏\"\n\n##### How to understand covisitation matrix intuitively? 💡💡💡\n-   it’s a way to take any aid and find any number of aids which are most similar to it\n-   co-visitation matrices differentiate from each other based on how they define similarity or how they select aids to be paired together\n\n##### if you were to play the role of the inventor of co-visitation matrix, what series of ideas/questions could trigger the creation of it?\n-   Is there any relationship between one aid/product with other aids/products in the same session or across all sessions?\n-   Are there some aids more similar to some and more different to others?\n-   Could it be possible when this aid is viewed, some aids are more likely to be clicked/carted/ordered than other aids?\n-   Could we pair aids together for each and every session and count the occurrences of pairs?\n-   Since one aid (eg., '122') could have many pair-partners, by counting the occurrences of the pairs ('122', pair-partner), could we find the most common pair-partners of aid '122'?\n-   Could the next clicks or carts or orders be the most common pair-partners of the last aid (or all aids) of a test session?\n\n##### what does pairing logic or a definition of simiarlity look like\n-   In Radek's notebook, the pairing logic is the following\n-   use only the last 30 aids of each session to pair on each other with pandas `merge` or polars `join` on `session`\n-   remove the pairs of same partners\n-   keep pairs whose right-partner is after left-partner within a day\n\n##### We can tweak the pairing logic to change our co-visitation matrix\n\n#### <mark style=\"background: #FFB86CA6;\">Word2Vect</mark> \n\n##### What shortcomings does co-visitation matrix have? Could Word2Vec be a better model?\n- [asked & answered](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358#2105430) Thank you @radek1 for your insightful reply again!",
      "votes": null
    },
    {
      "id": "2107372",
      "postDate": "01/19/2023 18:21:49",
      "content": "<p>\"to predict the next 20 aids to be clicked, 20 to be carted and 20 to be ordered\"(c)<br>\nThat's wrong. You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.<br>\nAs for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.</p>",
      "rawMarkdown": "\"to predict the next 20 aids to be clicked, 20 to be carted and 20 to be ordered\"(c)\nThat's wrong. You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.\nAs for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.",
      "votes": null
    },
    {
      "id": "2107656",
      "postDate": "01/19/2023 23:41:00",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> for your comment and correction! </p>\n<blockquote>\n  <p>You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.</p>\n</blockquote>\n<p>You are very right about it and I have been careless when writing about it above. I will correct it right now.</p>\n<blockquote>\n  <p>As for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.</p>\n</blockquote>\n<p>This is great insight too! I have observed the same but failed to hunt the meaning of it. Thanks again Artem Fedorov!</p>\n<p>It reminds how important it is to keep doing EDA on this dataset.</p>",
      "rawMarkdown": "Thank you very much @artemfedorov for your comment and correction! \n\n> You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.\n\nYou are very right about it and I have been careless when writing about it above. I will correct it right now.\n\n> As for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.\n\nThis is great insight too! I have observed the same but failed to hunt the meaning of it. Thanks again Artem Fedorov!\n\nIt reminds how important it is to keep doing EDA on this dataset.",
      "votes": null
    },
    {
      "id": "2111201",
      "postDate": "01/22/2023 17:15:54",
      "content": "<p>Glad to read my comment was helpfull.</p>\n<blockquote>\n  <p>It reminds how important it is to keep doing EDA on this dataset.</p>\n</blockquote>\n<p>Well, this applies to most datasets, probably to any dataset new for someone.</p>\n<p>P.S. Keep pushing and good luck!</p>",
      "rawMarkdown": "Glad to read my comment was helpfull.\n\n>It reminds how important it is to keep doing EDA on this dataset.\n\nWell, this applies to most datasets, probably to any dataset new for someone.\n\nP.S. Keep pushing and good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2107372,
      "author_name": "artemfedorov",
      "author_url": "",
      "post_date": "01/19/2023 18:21:49",
      "content": "<p>\"to predict the next 20 aids to be clicked, 20 to be carted and 20 to be ordered\"(c)<br>\nThat's wrong. You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.<br>\nAs for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2107656,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "01/19/2023 23:41:00",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> for your comment and correction! </p>\n<blockquote>\n  <p>You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.</p>\n</blockquote>\n<p>You are very right about it and I have been careless when writing about it above. I will correct it right now.</p>\n<blockquote>\n  <p>As for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.</p>\n</blockquote>\n<p>This is great insight too! I have observed the same but failed to hunt the meaning of it. Thanks again Artem Fedorov!</p>\n<p>It reminds how important it is to keep doing EDA on this dataset.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2111201,
              "author_name": "artemfedorov",
              "author_url": "",
              "post_date": "01/22/2023 17:15:54",
              "content": "<p>Glad to read my comment was helpfull.</p>\n<blockquote>\n  <p>It reminds how important it is to keep doing EDA on this dataset.</p>\n</blockquote>\n<p>Well, this applies to most datasets, probably to any dataset new for someone.</p>\n<p>P.S. Keep pushing and good luck!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2106180": "#### <mark style=\"background: #FFB86CA6;\">Why?</mark> \n\n- writing down my reflections of learning is good for myself and others who share similar experiences\n- I am a beginner who learns slowly and can't keep up with all the good and new posts and notebooks, so I will take my own pace\n- otto is a worthwhile comp which I will keep learning even after the deadline\n- So, this reflection will grow as I keep learning even after the deadline.\n- for a quick sum for most if not all amazing posts/notebooks, please check out @thedevastator 's \"One Month Left - Here is what you need to know!\" [post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/374229#2105455) 🔥🔥🔥🔥\n\n\n#### <mark style=\"background: #FFB86CA6;\">What inside the dataset</mark> \n\n**How the official OTTO dataset repo page describe it**\n\n> -   12M real-world anonymized user sessions\n> -   220M events, consiting of `clicks`, `carts` and `orders`\n> -   1.8M unique articles in the catalogue\n\n\n\n#### <mark style=\"background: #FFB86CA6;\">What to predict</mark> \n\n##### What exactly does this comp want us to predict? 💡💡💡\n\nthanks 🙏 to [@artemfedorov](https://www.kaggle.com/artemfedorov) for pointing out the misleading part of my previous description here, below is my second attempt.\n\n##### The official OTTO dataset repo [page](https://github.com/otto-de/recsys-dataset#evaluation) describe the prediction task very precisely actually\n\n> For each `session` in the test data, your task it to predict the `aid` values for each `type` that occur after the last timestamp `ts` the test session. In other words, the test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation.\n\n> For `clicks` there is only a single ground truth value for each session, which is the next `aid` clicked during the session (although you can still predict up to 20 `aid` values). The ground truth for `carts` and `orders` contains all `aid` values that were added to a cart and ordered respectively during the session.\n\n##### Another rephrase to help understanding hopefully\n- I am given a test session which has been truncated at certain timestamp\n- The ground truth includes only the 1st clicked aid for each test session after the timestamp above, but I am given 20 chances to get it right\n- The ground truth has all the carted aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth carted aids are more than 20, I can still score full as long as I can get 20 of those carted aids right.\n- The ground truth has all the ordered aids for each session after the timestamp above, I am given 20 chances to get them all right when they are less than 20; but don't worry even when the ground truth ordered aids are more than 20, I can still score full as long as I can get 20 of those ordered aids right.\n\nThe detailed rephase above is derived from the `Recall@20` metric below\n\n\n#### <mark style=\"background: #FFB86CA6;\">How to evaluate predictions</mark> \n\n##### How calc the `Recall@20` metric to score a type of all test sessions\n\n$$\nR_{type} = \\frac{ \\sum\\limits_{i=1}^N | \\\\{ \\text{predicted aids} \\\\}\\_{i, type} \\cap \\\\{ \\text{ground truth aids} \\\\}\\_{i, type} | }{ \\sum\\limits_{i=1}^N \\min{( 20, | \\\\{ \\text{ground truth aids} \\\\}_{i, type} | )}}\n$$\n\n##### How to add up 3 different type scores above to get the total score\n\n$$\nscore = 0.10 \\cdot R_{clicks} + 0.30 \\cdot R_{carts} + 0.60 \\cdot R_{orders}\n$$\n\n\n\n\n#### <mark style=\"background: #FFB86CA6;\">What insights can be derived from EDA</mark> \n\n##### **The official OTTO dataset repo has provided 3 tables for us**\n\nThe formula and meaning of `density` is answered by the organizer @pnormann  [here](https://github.com/otto-de/recsys-dataset/issues/2#issuecomment-1313697705) \n\n| dataset | num_sessions | num_items | num_events  | num_clicks  | num_carts  | num_orders | Density |\n| ------- | ------------ | --------- | ----------- | ----------- | ---------- | ---------- | ------- |\n| Train   | 12_899_779   | 1_855_603 | 216_716_096 | 194_720_954 | 16_896_191 | 5_098_951  | 0.0005  | \n| Test    |  1_671_803            |           |             |             |            |            |         |\n\n|                           |  mean |   std |  min |  50% |  75% |  90% |  95% |  max |\n| :------------------------ | ----: | ----: | ---: | ---: | ---: | ---: | ---: | ---: |\n| Train num_events per session | 16.80 | 33.58 |    2 |    6 |   15 |   39 |   68 |  500 |\n| Test num_events per session | TBA | TBA | TBA | TBA | TBA | TBA | TBA | TBA |\n\n\n|                        |   mean |    std |  min |  50% |  75% |  90% |  95% |    max |\n| :--------------------- | -----: | -----: | ---: | ---: | ---: | ---: | ---: | -----: |\n| Train num_events per item | 116.79 | 728.85 |    3 |   20 |   56 |  183 |  398 | 129004 |\n| Test num_events per item  |    TBA |    TBA |  TBA |  TBA |  TBA |  TBA |  TBA |    TBA |\n\n##### What insights can we derive?\n\n\n\n#### <mark style=\"background: #FFB86CA6;\">Candidate ReRank Model</mark> \n\n\n##### What is Candidate ReRank model? 💡💡💡\n- first we use a model to select hundreds of aid candidates, then we use another model to select the final 20 aids for predictions\n\n##### Why Chris Deotte said Candidate ReRank model will most likely to win this comp? 💡💡💡\n- maybe it is the most obvious approach and it works in other RecSys comp, according to Chris Deotte\n- there are 1.8 milliion aids to choose from, but only need 20 aids for each session_type\n- It's kind of making sense to break a large problem into 2 smaller problems: choose hundreds from millions in one model, and then choose 20 from hundreds in another\n- but how to select hundreds from millions? through similarities? \n\t- great kagglers in otto have shared approaches like co-visitation matrix, word2vec, matrix factorization, and maybe more I don't know\n\n\n\n####  <mark style=\"background: #FFB86CA6;\">Co-visitation Matrix</mark> \n \n\n##### Radek explains it way better than my own reflection below\n- \"A co-visitation matrix counts the co-occurrence of two actions in close proximity.\"\n- \"If a user bought A and shortly after bought B, we store these values together.\"\n- \"We calculate counts and use them to estimate the probability of future actions based on recent history.\"\n- \"It is quite important to understand what is happening in the co-visitation matrix approach…\"\n- \"Since it suffers from the same issues as our trigram example!\"\n- \"Plus what does the co-visitation matrix resemble?\"\n- \"You are right, it is akin to doing Matrix Factorization by counting!\"\n- \"It is really fun that this competition exposed this heuristic (the co-visitation matrix) that I have not been aware of before! 🙏\"\n\n##### How to understand covisitation matrix intuitively? 💡💡💡\n-   it’s a way to take any aid and find any number of aids which are most similar to it\n-   co-visitation matrices differentiate from each other based on how they define similarity or how they select aids to be paired together\n\n##### if you were to play the role of the inventor of co-visitation matrix, what series of ideas/questions could trigger the creation of it?\n-   Is there any relationship between one aid/product with other aids/products in the same session or across all sessions?\n-   Are there some aids more similar to some and more different to others?\n-   Could it be possible when this aid is viewed, some aids are more likely to be clicked/carted/ordered than other aids?\n-   Could we pair aids together for each and every session and count the occurrences of pairs?\n-   Since one aid (eg., '122') could have many pair-partners, by counting the occurrences of the pairs ('122', pair-partner), could we find the most common pair-partners of aid '122'?\n-   Could the next clicks or carts or orders be the most common pair-partners of the last aid (or all aids) of a test session?\n\n##### what does pairing logic or a definition of simiarlity look like\n-   In Radek's notebook, the pairing logic is the following\n-   use only the last 30 aids of each session to pair on each other with pandas `merge` or polars `join` on `session`\n-   remove the pairs of same partners\n-   keep pairs whose right-partner is after left-partner within a day\n\n##### We can tweak the pairing logic to change our co-visitation matrix\n\n#### <mark style=\"background: #FFB86CA6;\">Word2Vect</mark> \n\n##### What shortcomings does co-visitation matrix have? Could Word2Vec be a better model?\n- [asked & answered](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358#2105430) Thank you @radek1 for your insightful reply again!",
    "2107372": "\"to predict the next 20 aids to be clicked, 20 to be carted and 20 to be ordered\"(c)\nThat's wrong. You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.\nAs for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.",
    "2107656": "Thank you very much @artemfedorov for your comment and correction! \n\n> You need to predict the only one aid clicked. You have 20 aids and to get points you need that only aid clicked to be among those 20.\n\nYou are very right about it and I have been careless when writing about it above. I will correct it right now.\n\n> As for carts/orders, most sessions end without any aid put in cart and without any aid ordered. Yes, we provide predictions for all the sessions - but most of those predictions are meaningless, and would not count whatever aids are listed among those 20, as most sessions have 0 carts and 0 orders.\n\nThis is great insight too! I have observed the same but failed to hunt the meaning of it. Thanks again Artem Fedorov!\n\nIt reminds how important it is to keep doing EDA on this dataset.",
    "2111201": "Glad to read my comment was helpfull.\n\n>It reminds how important it is to keep doing EDA on this dataset.\n\nWell, this applies to most datasets, probably to any dataset new for someone.\n\nP.S. Keep pushing and good luck!"
  },
  "source": "meta"
}