{
  "id": 383657,
  "title": "226th (?!) Place Solution & Two-cents from a First-timer",
  "url": "/competitions/otto-recommender-system/writeups/hoang-nguyen-226th-place-solution-two-cents-from-a",
  "author_name": "",
  "post_date": "2023-02-10T16:36:39.910Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>FOREWORD</h2>\n<p>My work in this competition is nowhere near excellent compared to other more comprehensive approaches that have been shared. In addition, it is largely based on publicly available notebooks and ideas in the competition. Therefore, rather than solution-sharing, this document serves two other main purposes:</p>\n<ul>\n<li>Providing other newbies (like me) with a few tips on getting started with Kaggle real competitions</li>\n<li>Memorializing my personal journey to the first medal</li>\n</ul>\n<p>That said, my sharing consists of two parts</p>\n<ul>\n<li>My workflow in this competition</li>\n<li>What I have learned from it</li>\n</ul>\n<h2>CREDITS</h2>\n<p>First, big thanks for Kaggle and OTTO team for this learning opportunity - I have learned so much!</p>\n<p>Secondly, my achievement and learning in this competition owe primarily to many other participants who have shared their approaches, ideas and feedbacks, especially</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for sharing his <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">co-visitation matrix &amp; rule-based ranker notebook</a> and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">suggestions on building a model-based ranker</a></li>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for the <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">train-validation split dataset</a> and valuable EDA</li>\n<li><a href=\"https://www.kaggle.com/vbmokin\" target=\"_blank\">@vbmokin</a> and <a href=\"https://www.kaggle.com/karakasatarik\" target=\"_blank\">@karakasatarik</a> on their awesome solution ensemble notebooks (<a href=\"https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks\" target=\"_blank\">vbmokin's</a>, <a href=\"https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334\" target=\"_blank\">karakasatarik's</a>).</li>\n</ul>\n<h2>I. Solution Workflow</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F982488%2F8b2e1abf8cff557f28eca83e99654f7d%2FOTTO%20Workflow.png?generation=1675527772718204&amp;alt=media\" alt=\"\"></p>\n<h3>1. Item Co-visitation Matrices</h3>\n<p>This part is entirely based on <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Chris' notebook</a> above.</p>\n<p>There exist products that are frequently clicked/carted/ordered together. A co-visitation matrix, using a pre-defined rule, gives a weight <code>W</code> to a pair of products <code>A</code> &amp; <code>B</code> to signify such relationship between the products.<br>\nWith this notion, my solution includes below three co-visitation matrices mentioned in Chris' notebook</p>\n<ol>\n<li>Order matrix: Click/cart/order to click/cart/order with type weighting</li>\n<li>Buy2buy matrix: Cart/order to cart/order</li>\n<li>Click matrix: click/cart/order to clicks with time weighting</li>\n</ol>\n<p><strong>Tests that did not work</strong><br>\nThe original notebook truncates session to the last 30 events (<code>tail=30</code>) and select for the three matrices top 15, 15 and 20 items most associated with each <code>aid</code>. I additional tested <code>tail</code> of 35 and 40, and top 40-40-50, and finally used <code>tail=40</code> and top 40-40-50. However, this turns out not helpful for final recall scores. The extra information seems to be more noisy than helpful in this case.</p>\n<p><strong>Code</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-matrixv2-tail40-top404050-w136\" target=\"_blank\">Training set's matrix notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-te-matrixv2-tail40-top404050-w136\" target=\"_blank\">Test set's matrix notebook</a></li>\n<li>I did try recreating Chris' co-visitation matrices (and candidate selection) in <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-matrixv1-tail40-top40-40-50\" target=\"_blank\">this notebook</a>, which better utilizes <code>cuDF</code> and therefore shortens running time by more than half.</li>\n</ul>\n<h3>2. Feature Generation</h3>\n<p>Chris in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">his discussion</a> suggests three sets of feature to be created</p>\n<ul>\n<li>Item features</li>\n<li>Session features</li>\n<li>Item-session interaction features</li>\n</ul>\n<p>Based on this idea, I have created the following features</p>\n<ul>\n<li>Item features (for each <code>aid</code>)<ul>\n<li>Count of events (click/cart/order)</li>\n<li>Sum of event weight</li>\n<li>Quarter of day (QoD) with most events (0-3)</li>\n<li>Day of week (DoW) with most events (0-6)</li></ul></li>\n<li>User features (for each <code>session</code>)<ul>\n<li>Count of events (click/cart/order) and interacted items (<code>aid</code>)</li>\n<li>Sum of event weight</li>\n<li>QoD with most events (0-3)</li>\n<li>DoW with most events (0-6)</li>\n<li>Number of days with events</li>\n<li>Days from first to last events</li></ul></li>\n<li>User-item features (for each <code>session</code>-<code>aid</code> pair)<ul>\n<li>Count of events (click/cart/order) and interacted items (<code>aid</code>)</li>\n<li>Sum of event weight</li>\n<li>QoD with most events in both categorical (0-3) and one-hot encoded (0/1 for each) format</li>\n<li>DoW with most events in both categorical (0-6) and one-hot encoded (0/1 for each) format</li>\n<li><code>last_n</code> = <code>item_chronological_rank / user_total_event_count</code></li>\n<li><code>last_ts</code> = <code>(user_item_last_timestamp  - start_week_timestamp) / (end_week_timestamp - start_week_timestamp)</code></li></ul></li>\n</ul>\n<p>However, due to Kaggle's limited computational resources, only a subset of the features was finally selected.</p>\n<p><strong>Code</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Training set's candidate selection &amp; feature generation notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Test set's candidate selection &amp; feature generation notebook</a></li>\n</ul>\n<h3>3. Candidate Selection</h3>\n<p>This section's logic is partly based on that of <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Chris' notebook</a> above.</p>\n<p>For every session, I select top <code>X</code> most relevant items in each event type (click, cart and order). \"Relevancy\" is scored using a number of rules:</p>\n<ul>\n<li>Number of times the session has clicked/carted/ordered the items</li>\n<li>Sum of co-visitation weight</li>\n<li>Whether the item is a top-clicked/bought item of the week</li>\n</ul>\n<p>Candidate selection is a bottleneck of this solution and needs some balancing; having too few candidates results in low recall no matter how good our ranker is, but having too many will exceed computational limit. Therefore, I tested <code>X</code> for three different values 30, 35 and 40. <code>X</code> is finally set at 40.</p>\n<p><strong>Code</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Training set's candidate selection &amp; feature generation notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Test set's candidate selection &amp; feature generation notebook</a></li>\n</ul>\n<h3>4. Ranker</h3>\n<p>Two different ranking methods are used for rankers<br>\n<strong>(1) Rule-based ranker in Chris' notebook</strong><br>\nScores (CV score not computed due to limited time)</p>\n<table>\n<thead>\n<tr>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.54777</td>\n<td>0.54769</td>\n</tr>\n</tbody>\n</table>\n<p>Notebooks</p>\n<ul>\n<li>Inference: <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v2-can40-v2-tail40-top40405\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>(2) <code>XGBRanker</code> with <code>rank:pairwise</code> objective; one single model for each event type</strong><br>\nScores</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.512189</td>\n<td>0.56940</td>\n<td>0.56884</td>\n</tr>\n</tbody>\n</table>\n<p>Notebooks</p>\n<ul>\n<li>Training: <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-trainingclk-can40-matrixv2-tail40-top404050\" target=\"_blank\">click training notebook</a>, <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-trainingcart-can40-matrixv2-tail40-top404050\" target=\"_blank\">cart training notebook</a>, <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-trainingord-can40-matrixv2-tail40-top404050\" target=\"_blank\">order training notebook</a></li>\n<li>Inference: <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v1-can40-v2-tail40-top404050\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>Notes</strong></p>\n<ul>\n<li>Method (2) was overfitting and requires a lot of hyper-parameter tuning.</li>\n<li>In LB score, method (2) managed to beat (1) by 0.0216, so a model-based ranker is, understandably, better than a rule-based approach.</li>\n<li>Method (1) underperforms in LB compared to Chris' original rule-based reranker notebook. The three differences between method (1) and Chris' are<ul>\n<li>Session truncation: Method (1)'s co-visitation matrices use last 40 events of each session, while Chris' use last 30.</li>\n<li>Number of items in co-visitation matrice: method (1)'s matrices use top 40, 40 and 50 items for the three matrices mentioned above, while Chris' use top 15, 15 and 20.</li>\n<li>Event type weight: method (1) uses <code>type_weight = {0:1, 1:3, 2:6}</code> while Chris uses <code>type_weight = {0:1, 1:6, 2:3}</code>. However, in a separate test of mine this disparity proves not to affect the score by much.</li></ul></li>\n</ul>\n<h3>5. Solution Ensemble</h3>\n<p>Due to the poor performance of my individual rankers, I ensemble them and other publicly available submissions (credit to <a href=\"https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks\" target=\"_blank\">@karakasatarik's notebook</a> and <a href=\"https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334\" target=\"_blank\">@vbmokin's notebook</a>), for better score. I weight each submission using their public LB score. My final two submissions are:<br>\n<strong>(A) Ensemble of method (2) above and public submissions</strong><br>\n<strong>(B) Ensemble of above two owned methods and public submissions</strong></p>\n<p>Below are the scores of all public submissions and ensembles</p>\n<p></p><ul><br>\n<li><p>Public LB</p><p></p>\n<table>\n<thead>\n<tr>\n<th><a href=\"https://www.kaggle.com/karakasatarik\" target=\"_blank\">@karakasatarik</a>'s ensemble</th>\n<th><a href=\"https://www.kaggle.com/vbmokin\" target=\"_blank\">@vbmokin</a>'s ensemble</th>\n<th>Submission (A)</th>\n<th>Submission (B)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.57843</td>\n<td>0.57821</td>\n<td>0.57884</td>\n<td>0.57844</td>\n</tr>\n</tbody>\n</table>\n<p></p></li><br>\n<li><p>Private LB<br></p>\n<table>\n<thead>\n<tr>\n<th><a href=\"https://www.kaggle.com/karakasatarik\" target=\"_blank\">@karakasatarik</a>'s ensemble</th>\n<th><a href=\"https://www.kaggle.com/vbmokin\" target=\"_blank\">@vbmokin</a>'s ensemble</th>\n<th>Submission (A)</th>\n<th>Submission (B)</th>\n<th><br></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.57808</td>\n<td>0.57787</td>\n<td>0.57855</td>\n<td>0.57811</td>\n<td><p></p></td></tr></tbody></table></li>\n\n\n</ul>\n\n\n\n\n\n\nMy submission (B) has higher scores than did the public ensembles for private LB, while submission (A) outperforms all other ensembles in both public and private LBs. This means that both methods, though achieving relatively poor scores, did add some valuable info to the final solution (the 0.57808 -&gt; 0.57855 improvement is equal to a boost from rank 653th to rank 255th!).\n\n\n\n\n\n\n\n<h2>II. What I have learned as a Kaggle competition first-timer</h2>\n<p>It can be seen from the above sections that my work depended largely on other participants' help and support - so thank you! Below are what I have learned from this valuable experience - hope it'd be helpful for others too!</p>\n<ul>\n<li><strong>Read the discussion and notebook forums</strong> - especially if you're new to Kaggle/ML. There are always experienced participants sharing their ideas, suggestions and feedbacks, so if you are looking for a place to start the race, this is it! And check back once in a while on notebook/discussion that is helpful for you - the comment section may give you additional insights or unexpected bug-fixings that are no less valuable than the notebook.<br>\nI myself in this competition would definitely not have got the bronze if it was not thanks to the knowledge shared by others.</li>\n<li><strong>Try as many things as you can</strong>. I read through solutions of some of the top achievers (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013\" target=\"_blank\">Top 4</a>, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382802\" target=\"_blank\">Top 6</a>), and realized that they all tried more approaches (Matrix Factorization, Word2Vec) in more depth (hundred of co-visitation matrices). Of course they have done better in other aspects too (preprocessing, train-val splitting, validation, etc.), but such multi in-depth methods alone already lead to more results in better robustness. When ensembled, their solutions understandably far outperform mine.</li>\n<li><strong>Know every line of code</strong> you write. I lost about a month getting discouraged by CV scores going up and down unexpectedly, only to later found out a line of code misplaced in my validation step. <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382277\" target=\"_blank\">This can happen to anyone</a>, so make sure you understand every line of code in your work.</li>\n</ul>",
  "messages": [
    {
      "id": "2129379",
      "postDate": "02/04/2023 15:25:39",
      "content": "<h2>FOREWORD</h2>\n<p>My work in this competition is nowhere near excellent compared to other more comprehensive approaches that have been shared. In addition, it is largely based on publicly available notebooks and ideas in the competition. Therefore, rather than solution-sharing, this document serves two other main purposes:</p>\n<ul>\n<li>Providing other newbies (like me) with a few tips on getting started with Kaggle real competitions</li>\n<li>Memorializing my personal journey to the first medal</li>\n</ul>\n<p>That said, my sharing consists of two parts</p>\n<ul>\n<li>My workflow in this competition</li>\n<li>What I have learned from it</li>\n</ul>\n<h2>CREDITS</h2>\n<p>First, big thanks for Kaggle and OTTO team for this learning opportunity - I have learned so much!</p>\n<p>Secondly, my achievement and learning in this competition owe primarily to many other participants who have shared their approaches, ideas and feedbacks, especially</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for sharing his <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">co-visitation matrix &amp; rule-based ranker notebook</a> and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">suggestions on building a model-based ranker</a></li>\n<li><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for the <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">train-validation split dataset</a> and valuable EDA</li>\n<li><a href=\"https://www.kaggle.com/vbmokin\" target=\"_blank\">@vbmokin</a> and <a href=\"https://www.kaggle.com/karakasatarik\" target=\"_blank\">@karakasatarik</a> on their awesome solution ensemble notebooks (<a href=\"https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks\" target=\"_blank\">vbmokin's</a>, <a href=\"https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334\" target=\"_blank\">karakasatarik's</a>).</li>\n</ul>\n<h2>I. Solution Workflow</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F982488%2F8b2e1abf8cff557f28eca83e99654f7d%2FOTTO%20Workflow.png?generation=1675527772718204&amp;alt=media\" alt=\"\"></p>\n<h3>1. Item Co-visitation Matrices</h3>\n<p>This part is entirely based on <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Chris' notebook</a> above.</p>\n<p>There exist products that are frequently clicked/carted/ordered together. A co-visitation matrix, using a pre-defined rule, gives a weight <code>W</code> to a pair of products <code>A</code> &amp; <code>B</code> to signify such relationship between the products.<br>\nWith this notion, my solution includes below three co-visitation matrices mentioned in Chris' notebook</p>\n<ol>\n<li>Order matrix: Click/cart/order to click/cart/order with type weighting</li>\n<li>Buy2buy matrix: Cart/order to cart/order</li>\n<li>Click matrix: click/cart/order to clicks with time weighting</li>\n</ol>\n<p><strong>Tests that did not work</strong><br>\nThe original notebook truncates session to the last 30 events (<code>tail=30</code>) and select for the three matrices top 15, 15 and 20 items most associated with each <code>aid</code>. I additional tested <code>tail</code> of 35 and 40, and top 40-40-50, and finally used <code>tail=40</code> and top 40-40-50. However, this turns out not helpful for final recall scores. The extra information seems to be more noisy than helpful in this case.</p>\n<p><strong>Code</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-matrixv2-tail40-top404050-w136\" target=\"_blank\">Training set's matrix notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-te-matrixv2-tail40-top404050-w136\" target=\"_blank\">Test set's matrix notebook</a></li>\n<li>I did try recreating Chris' co-visitation matrices (and candidate selection) in <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-matrixv1-tail40-top40-40-50\" target=\"_blank\">this notebook</a>, which better utilizes <code>cuDF</code> and therefore shortens running time by more than half.</li>\n</ul>\n<h3>2. Feature Generation</h3>\n<p>Chris in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">his discussion</a> suggests three sets of feature to be created</p>\n<ul>\n<li>Item features</li>\n<li>Session features</li>\n<li>Item-session interaction features</li>\n</ul>\n<p>Based on this idea, I have created the following features</p>\n<ul>\n<li>Item features (for each <code>aid</code>)<ul>\n<li>Count of events (click/cart/order)</li>\n<li>Sum of event weight</li>\n<li>Quarter of day (QoD) with most events (0-3)</li>\n<li>Day of week (DoW) with most events (0-6)</li></ul></li>\n<li>User features (for each <code>session</code>)<ul>\n<li>Count of events (click/cart/order) and interacted items (<code>aid</code>)</li>\n<li>Sum of event weight</li>\n<li>QoD with most events (0-3)</li>\n<li>DoW with most events (0-6)</li>\n<li>Number of days with events</li>\n<li>Days from first to last events</li></ul></li>\n<li>User-item features (for each <code>session</code>-<code>aid</code> pair)<ul>\n<li>Count of events (click/cart/order) and interacted items (<code>aid</code>)</li>\n<li>Sum of event weight</li>\n<li>QoD with most events in both categorical (0-3) and one-hot encoded (0/1 for each) format</li>\n<li>DoW with most events in both categorical (0-6) and one-hot encoded (0/1 for each) format</li>\n<li><code>last_n</code> = <code>item_chronological_rank / user_total_event_count</code></li>\n<li><code>last_ts</code> = <code>(user_item_last_timestamp  - start_week_timestamp) / (end_week_timestamp - start_week_timestamp)</code></li></ul></li>\n</ul>\n<p>However, due to Kaggle's limited computational resources, only a subset of the features was finally selected.</p>\n<p><strong>Code</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Training set's candidate selection &amp; feature generation notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Test set's candidate selection &amp; feature generation notebook</a></li>\n</ul>\n<h3>3. Candidate Selection</h3>\n<p>This section's logic is partly based on that of <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Chris' notebook</a> above.</p>\n<p>For every session, I select top <code>X</code> most relevant items in each event type (click, cart and order). \"Relevancy\" is scored using a number of rules:</p>\n<ul>\n<li>Number of times the session has clicked/carted/ordered the items</li>\n<li>Sum of co-visitation weight</li>\n<li>Whether the item is a top-clicked/bought item of the week</li>\n</ul>\n<p>Candidate selection is a bottleneck of this solution and needs some balancing; having too few candidates results in low recall no matter how good our ranker is, but having too many will exceed computational limit. Therefore, I tested <code>X</code> for three different values 30, 35 and 40. <code>X</code> is finally set at 40.</p>\n<p><strong>Code</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Training set's candidate selection &amp; feature generation notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50\" target=\"_blank\">Test set's candidate selection &amp; feature generation notebook</a></li>\n</ul>\n<h3>4. Ranker</h3>\n<p>Two different ranking methods are used for rankers<br>\n<strong>(1) Rule-based ranker in Chris' notebook</strong><br>\nScores (CV score not computed due to limited time)</p>\n<table>\n<thead>\n<tr>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.54777</td>\n<td>0.54769</td>\n</tr>\n</tbody>\n</table>\n<p>Notebooks</p>\n<ul>\n<li>Inference: <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v2-can40-v2-tail40-top40405\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>(2) <code>XGBRanker</code> with <code>rank:pairwise</code> objective; one single model for each event type</strong><br>\nScores</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.512189</td>\n<td>0.56940</td>\n<td>0.56884</td>\n</tr>\n</tbody>\n</table>\n<p>Notebooks</p>\n<ul>\n<li>Training: <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-trainingclk-can40-matrixv2-tail40-top404050\" target=\"_blank\">click training notebook</a>, <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-trainingcart-can40-matrixv2-tail40-top404050\" target=\"_blank\">cart training notebook</a>, <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-trainingord-can40-matrixv2-tail40-top404050\" target=\"_blank\">order training notebook</a></li>\n<li>Inference: <a href=\"https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v1-can40-v2-tail40-top404050\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>Notes</strong></p>\n<ul>\n<li>Method (2) was overfitting and requires a lot of hyper-parameter tuning.</li>\n<li>In LB score, method (2) managed to beat (1) by 0.0216, so a model-based ranker is, understandably, better than a rule-based approach.</li>\n<li>Method (1) underperforms in LB compared to Chris' original rule-based reranker notebook. The three differences between method (1) and Chris' are<ul>\n<li>Session truncation: Method (1)'s co-visitation matrices use last 40 events of each session, while Chris' use last 30.</li>\n<li>Number of items in co-visitation matrice: method (1)'s matrices use top 40, 40 and 50 items for the three matrices mentioned above, while Chris' use top 15, 15 and 20.</li>\n<li>Event type weight: method (1) uses <code>type_weight = {0:1, 1:3, 2:6}</code> while Chris uses <code>type_weight = {0:1, 1:6, 2:3}</code>. However, in a separate test of mine this disparity proves not to affect the score by much.</li></ul></li>\n</ul>\n<h3>5. Solution Ensemble</h3>\n<p>Due to the poor performance of my individual rankers, I ensemble them and other publicly available submissions (credit to <a href=\"https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks\" target=\"_blank\">@karakasatarik's notebook</a> and <a href=\"https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334\" target=\"_blank\">@vbmokin's notebook</a>), for better score. I weight each submission using their public LB score. My final two submissions are:<br>\n<strong>(A) Ensemble of method (2) above and public submissions</strong><br>\n<strong>(B) Ensemble of above two owned methods and public submissions</strong></p>\n<p>Below are the scores of all public submissions and ensembles</p>\n<p></p><ul><br>\n<li><p>Public LB</p><p></p>\n<table>\n<thead>\n<tr>\n<th><a href=\"https://www.kaggle.com/karakasatarik\" target=\"_blank\">@karakasatarik</a>'s ensemble</th>\n<th><a href=\"https://www.kaggle.com/vbmokin\" target=\"_blank\">@vbmokin</a>'s ensemble</th>\n<th>Submission (A)</th>\n<th>Submission (B)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.57843</td>\n<td>0.57821</td>\n<td>0.57884</td>\n<td>0.57844</td>\n</tr>\n</tbody>\n</table>\n<p></p></li><br>\n<li><p>Private LB<br></p>\n<table>\n<thead>\n<tr>\n<th><a href=\"https://www.kaggle.com/karakasatarik\" target=\"_blank\">@karakasatarik</a>'s ensemble</th>\n<th><a href=\"https://www.kaggle.com/vbmokin\" target=\"_blank\">@vbmokin</a>'s ensemble</th>\n<th>Submission (A)</th>\n<th>Submission (B)</th>\n<th><br></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.57808</td>\n<td>0.57787</td>\n<td>0.57855</td>\n<td>0.57811</td>\n<td><p></p></td></tr></tbody></table></li>\n\n\n</ul>\n\n\n\n\n\n\nMy submission (B) has higher scores than did the public ensembles for private LB, while submission (A) outperforms all other ensembles in both public and private LBs. This means that both methods, though achieving relatively poor scores, did add some valuable info to the final solution (the 0.57808 -&gt; 0.57855 improvement is equal to a boost from rank 653th to rank 255th!).\n\n\n\n\n\n\n\n<h2>II. What I have learned as a Kaggle competition first-timer</h2>\n<p>It can be seen from the above sections that my work depended largely on other participants' help and support - so thank you! Below are what I have learned from this valuable experience - hope it'd be helpful for others too!</p>\n<ul>\n<li><strong>Read the discussion and notebook forums</strong> - especially if you're new to Kaggle/ML. There are always experienced participants sharing their ideas, suggestions and feedbacks, so if you are looking for a place to start the race, this is it! And check back once in a while on notebook/discussion that is helpful for you - the comment section may give you additional insights or unexpected bug-fixings that are no less valuable than the notebook.<br>\nI myself in this competition would definitely not have got the bronze if it was not thanks to the knowledge shared by others.</li>\n<li><strong>Try as many things as you can</strong>. I read through solutions of some of the top achievers (<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013\" target=\"_blank\">Top 4</a>, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382802\" target=\"_blank\">Top 6</a>), and realized that they all tried more approaches (Matrix Factorization, Word2Vec) in more depth (hundred of co-visitation matrices). Of course they have done better in other aspects too (preprocessing, train-val splitting, validation, etc.), but such multi in-depth methods alone already lead to more results in better robustness. When ensembled, their solutions understandably far outperform mine.</li>\n<li><strong>Know every line of code</strong> you write. I lost about a month getting discouraged by CV scores going up and down unexpectedly, only to later found out a line of code misplaced in my validation step. <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382277\" target=\"_blank\">This can happen to anyone</a>, so make sure you understand every line of code in your work.</li>\n</ul>",
      "rawMarkdown": "## FOREWORD\nMy work in this competition is nowhere near excellent compared to other more comprehensive approaches that have been shared. In addition, it is largely based on publicly available notebooks and ideas in the competition. Therefore, rather than solution-sharing, this document serves two other main purposes:\n- Providing other newbies (like me) with a few tips on getting started with Kaggle real competitions\n- Memorializing my personal journey to the first medal\n\nThat said, my sharing consists of two parts\n- My workflow in this competition\n- What I have learned from it\n\n## CREDITS\nFirst, big thanks for Kaggle and OTTO team for this learning opportunity - I have learned so much!\n\nSecondly, my achievement and learning in this competition owe primarily to many other participants who have shared their approaches, ideas and feedbacks, especially\n- @cdeotte for sharing his [co-visitation matrix & rule-based ranker notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) and [suggestions on building a model-based ranker](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210)\n- @radek1 for the [train-validation split dataset](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation) and valuable EDA\n- @vbmokin and @karakasatarik on their awesome solution ensemble notebooks ([vbmokin's](https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks), [karakasatarik's](https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334)).\n\n## I. Solution Workflow\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F982488%2F8b2e1abf8cff557f28eca83e99654f7d%2FOTTO%20Workflow.png?generation=1675527772718204&alt=media)\n\n### 1. Item Co-visitation Matrices\nThis part is entirely based on [Chris' notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) above.\n\nThere exist products that are frequently clicked/carted/ordered together. A co-visitation matrix, using a pre-defined rule, gives a weight `W` to a pair of products `A` & `B` to signify such relationship between the products.\nWith this notion, my solution includes below three co-visitation matrices mentioned in Chris' notebook\n1. Order matrix: Click/cart/order to click/cart/order with type weighting\n2. Buy2buy matrix: Cart/order to cart/order\n3. Click matrix: click/cart/order to clicks with time weighting\n\n**Tests that did not work**\nThe original notebook truncates session to the last 30 events (`tail=30`) and select for the three matrices top 15, 15 and 20 items most associated with each `aid`. I additional tested `tail` of 35 and 40, and top 40-40-50, and finally used `tail=40` and top 40-40-50. However, this turns out not helpful for final recall scores. The extra information seems to be more noisy than helpful in this case.\n\n**Code**\n- [Training set's matrix notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-matrixv2-tail40-top404050-w136)\n- [Test set's matrix notebook](https://www.kaggle.com/code/hoangnguyen719/otto-te-matrixv2-tail40-top404050-w136)\n- I did try recreating Chris' co-visitation matrices (and candidate selection) in [this notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-matrixv1-tail40-top40-40-50), which better utilizes `cuDF` and therefore shortens running time by more than half.\n\n### 2. Feature Generation\nChris in [his discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) suggests three sets of feature to be created\n- Item features\n- Session features\n- Item-session interaction features\n\nBased on this idea, I have created the following features\n- Item features (for each `aid`)\n    - Count of events (click/cart/order)\n    - Sum of event weight\n    - Quarter of day (QoD) with most events (0-3)\n    - Day of week (DoW) with most events (0-6)\n- User features (for each `session`)\n    - Count of events (click/cart/order) and interacted items (`aid`)\n    - Sum of event weight\n    - QoD with most events (0-3)\n    - DoW with most events (0-6)\n    - Number of days with events\n    - Days from first to last events\n- User-item features (for each `session`-`aid` pair)\n    - Count of events (click/cart/order) and interacted items (`aid`)\n    - Sum of event weight\n    - QoD with most events in both categorical (0-3) and one-hot encoded (0/1 for each) format\n    - DoW with most events in both categorical (0-6) and one-hot encoded (0/1 for each) format\n    - `last_n` = `item_chronological_rank / user_total_event_count`\n    - `last_ts` = `(user_item_last_timestamp  - start_week_timestamp) / (end_week_timestamp - start_week_timestamp)`\n\nHowever, due to Kaggle's limited computational resources, only a subset of the features was finally selected.\n\n**Code**\n- [Training set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50)\n- [Test set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50)\n\n### 3. Candidate Selection\nThis section's logic is partly based on that of [Chris' notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) above.\n\nFor every session, I select top `X` most relevant items in each event type (click, cart and order). \"Relevancy\" is scored using a number of rules:\n- Number of times the session has clicked/carted/ordered the items\n- Sum of co-visitation weight\n- Whether the item is a top-clicked/bought item of the week\n\nCandidate selection is a bottleneck of this solution and needs some balancing; having too few candidates results in low recall no matter how good our ranker is, but having too many will exceed computational limit. Therefore, I tested `X` for three different values 30, 35 and 40. `X` is finally set at 40.\n\n**Code**\n- [Training set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50)\n- [Test set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50)\n\n### 4. Ranker\nTwo different ranking methods are used for rankers\n**(1) Rule-based ranker in Chris' notebook**\nScores (CV score not computed due to limited time)\n| Public LB | Private LB |\n| --- | --- |\n| 0.54777 | 0.54769 |\n\nNotebooks\n- Inference: [here](https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v2-can40-v2-tail40-top40405)\n\n**(2) `XGBRanker` with `rank:pairwise` objective; one single model for each event type**\nScores\n| CV | Public LB | Private LB |\n| --- | --- | --- |\n| 0.512189 | 0.56940 | 0.56884 |\n\nNotebooks\n- Training: [click training notebook](https://www.kaggle.com/code/hoangnguyen719/otto-trainingclk-can40-matrixv2-tail40-top404050), [cart training notebook](https://www.kaggle.com/code/hoangnguyen719/otto-trainingcart-can40-matrixv2-tail40-top404050), [order training notebook](https://www.kaggle.com/code/hoangnguyen719/otto-trainingord-can40-matrixv2-tail40-top404050)\n- Inference: [here](https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v1-can40-v2-tail40-top404050)\n\n**Notes**\n- Method (2) was overfitting and requires a lot of hyper-parameter tuning.\n- In LB score, method (2) managed to beat (1) by 0.0216, so a model-based ranker is, understandably, better than a rule-based approach.\n- Method (1) underperforms in LB compared to Chris' original rule-based reranker notebook. The three differences between method (1) and Chris' are\n    - Session truncation: Method (1)'s co-visitation matrices use last 40 events of each session, while Chris' use last 30.\n    - Number of items in co-visitation matrice: method (1)'s matrices use top 40, 40 and 50 items for the three matrices mentioned above, while Chris' use top 15, 15 and 20.\n    - Event type weight: method (1) uses `type_weight = {0:1, 1:3, 2:6}` while Chris uses `type_weight = {0:1, 1:6, 2:3}`. However, in a separate test of mine this disparity proves not to affect the score by much.\n\n### 5. Solution Ensemble\nDue to the poor performance of my individual rankers, I ensemble them and other publicly available submissions (credit to [@karakasatarik's notebook](https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks) and [@vbmokin's notebook](https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334)), for better score. I weight each submission using their public LB score. My final two submissions are:\n**(A) Ensemble of method (2) above and public submissions**\n**(B) Ensemble of above two owned methods and public submissions**\n\nBelow are the scores of all public submissions and ensembles\n- Public LB\n| @karakasatarik's ensemble | @vbmokin's ensemble | Submission (A) | Submission (B) |\n| --- | --- | --- | --- |\n| 0.57843 | 0.57821 | 0.57884 | 0.57844 |\n\n- Private LB\n| @karakasatarik's ensemble | @vbmokin's ensemble | Submission (A) | Submission (B) |\n| --- | --- | --- | --- |\n| 0.57808 | 0.57787 | 0.57855 | 0.57811 |\n\nMy submission (B) has higher scores than did the public ensembles for private LB, while submission (A) outperforms all other ensembles in both public and private LBs. This means that both methods, though achieving relatively poor scores, did add some valuable info to the final solution (the 0.57808 -> 0.57855 improvement is equal to a boost from rank 653th to rank 255th!).\n\n## II. What I have learned as a Kaggle competition first-timer\nIt can be seen from the above sections that my work depended largely on other participants' help and support - so thank you! Below are what I have learned from this valuable experience - hope it'd be helpful for others too!\n- **Read the discussion and notebook forums** - especially if you're new to Kaggle/ML. There are always experienced participants sharing their ideas, suggestions and feedbacks, so if you are looking for a place to start the race, this is it! And check back once in a while on notebook/discussion that is helpful for you - the comment section may give you additional insights or unexpected bug-fixings that are no less valuable than the notebook.\nI myself in this competition would definitely not have got the bronze if it was not thanks to the knowledge shared by others.\n- **Try as many things as you can**. I read through solutions of some of the top achievers ([Top 4](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013), [Top 6](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382802)), and realized that they all tried more approaches (Matrix Factorization, Word2Vec) in more depth (hundred of co-visitation matrices). Of course they have done better in other aspects too (preprocessing, train-val splitting, validation, etc.), but such multi in-depth methods alone already lead to more results in better robustness. When ensembled, their solutions understandably far outperform mine.\n- **Know every line of code** you write. I lost about a month getting discouraged by CV scores going up and down unexpectedly, only to later found out a line of code misplaced in my validation step. [This can happen to anyone](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382277), so make sure you understand every line of code in your work.",
      "votes": null
    },
    {
      "id": "2129537",
      "postDate": "02/04/2023 17:30:26",
      "content": "<p>Great post thanks for presenting your pipeline ! Thanks for mentioning my  topic !</p>",
      "rawMarkdown": "Great post thanks for presenting your pipeline ! Thanks for mentioning my  topic !",
      "votes": null
    },
    {
      "id": "2129707",
      "postDate": "02/04/2023 20:02:29",
      "content": "<p>Thanks Rayane! Good to see I'm not alone haha. And congratulations on your first medal!</p>",
      "rawMarkdown": "Thanks Rayane! Good to see I'm not alone haha. And congratulations on your first medal!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2129537,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "02/04/2023 17:30:26",
      "content": "<p>Great post thanks for presenting your pipeline ! Thanks for mentioning my  topic !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2129707,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "02/04/2023 20:02:29",
          "content": "<p>Thanks Rayane! Good to see I'm not alone haha. And congratulations on your first medal!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2129379": "## FOREWORD\nMy work in this competition is nowhere near excellent compared to other more comprehensive approaches that have been shared. In addition, it is largely based on publicly available notebooks and ideas in the competition. Therefore, rather than solution-sharing, this document serves two other main purposes:\n- Providing other newbies (like me) with a few tips on getting started with Kaggle real competitions\n- Memorializing my personal journey to the first medal\n\nThat said, my sharing consists of two parts\n- My workflow in this competition\n- What I have learned from it\n\n## CREDITS\nFirst, big thanks for Kaggle and OTTO team for this learning opportunity - I have learned so much!\n\nSecondly, my achievement and learning in this competition owe primarily to many other participants who have shared their approaches, ideas and feedbacks, especially\n- @cdeotte for sharing his [co-visitation matrix & rule-based ranker notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) and [suggestions on building a model-based ranker](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210)\n- @radek1 for the [train-validation split dataset](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation) and valuable EDA\n- @vbmokin and @karakasatarik on their awesome solution ensemble notebooks ([vbmokin's](https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks), [karakasatarik's](https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334)).\n\n## I. Solution Workflow\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F982488%2F8b2e1abf8cff557f28eca83e99654f7d%2FOTTO%20Workflow.png?generation=1675527772718204&alt=media)\n\n### 1. Item Co-visitation Matrices\nThis part is entirely based on [Chris' notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) above.\n\nThere exist products that are frequently clicked/carted/ordered together. A co-visitation matrix, using a pre-defined rule, gives a weight `W` to a pair of products `A` & `B` to signify such relationship between the products.\nWith this notion, my solution includes below three co-visitation matrices mentioned in Chris' notebook\n1. Order matrix: Click/cart/order to click/cart/order with type weighting\n2. Buy2buy matrix: Cart/order to cart/order\n3. Click matrix: click/cart/order to clicks with time weighting\n\n**Tests that did not work**\nThe original notebook truncates session to the last 30 events (`tail=30`) and select for the three matrices top 15, 15 and 20 items most associated with each `aid`. I additional tested `tail` of 35 and 40, and top 40-40-50, and finally used `tail=40` and top 40-40-50. However, this turns out not helpful for final recall scores. The extra information seems to be more noisy than helpful in this case.\n\n**Code**\n- [Training set's matrix notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-matrixv2-tail40-top404050-w136)\n- [Test set's matrix notebook](https://www.kaggle.com/code/hoangnguyen719/otto-te-matrixv2-tail40-top404050-w136)\n- I did try recreating Chris' co-visitation matrices (and candidate selection) in [this notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-matrixv1-tail40-top40-40-50), which better utilizes `cuDF` and therefore shortens running time by more than half.\n\n### 2. Feature Generation\nChris in [his discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) suggests three sets of feature to be created\n- Item features\n- Session features\n- Item-session interaction features\n\nBased on this idea, I have created the following features\n- Item features (for each `aid`)\n    - Count of events (click/cart/order)\n    - Sum of event weight\n    - Quarter of day (QoD) with most events (0-3)\n    - Day of week (DoW) with most events (0-6)\n- User features (for each `session`)\n    - Count of events (click/cart/order) and interacted items (`aid`)\n    - Sum of event weight\n    - QoD with most events (0-3)\n    - DoW with most events (0-6)\n    - Number of days with events\n    - Days from first to last events\n- User-item features (for each `session`-`aid` pair)\n    - Count of events (click/cart/order) and interacted items (`aid`)\n    - Sum of event weight\n    - QoD with most events in both categorical (0-3) and one-hot encoded (0/1 for each) format\n    - DoW with most events in both categorical (0-6) and one-hot encoded (0/1 for each) format\n    - `last_n` = `item_chronological_rank / user_total_event_count`\n    - `last_ts` = `(user_item_last_timestamp  - start_week_timestamp) / (end_week_timestamp - start_week_timestamp)`\n\nHowever, due to Kaggle's limited computational resources, only a subset of the features was finally selected.\n\n**Code**\n- [Training set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50)\n- [Test set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50)\n\n### 3. Candidate Selection\nThis section's logic is partly based on that of [Chris' notebook](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) above.\n\nFor every session, I select top `X` most relevant items in each event type (click, cart and order). \"Relevancy\" is scored using a number of rules:\n- Number of times the session has clicked/carted/ordered the items\n- Sum of co-visitation weight\n- Whether the item is a top-clicked/bought item of the week\n\nCandidate selection is a bottleneck of this solution and needs some balancing; having too few candidates results in low recall no matter how good our ranker is, but having too many will exceed computational limit. Therefore, I tested `X` for three different values 30, 35 and 40. `X` is finally set at 40.\n\n**Code**\n- [Training set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-tr-cand40-v2-tail40-top40-40-50)\n- [Test set's candidate selection & feature generation notebook](https://www.kaggle.com/code/hoangnguyen719/otto-te-cand40-v2-tail40-top40-40-50)\n\n### 4. Ranker\nTwo different ranking methods are used for rankers\n**(1) Rule-based ranker in Chris' notebook**\nScores (CV score not computed due to limited time)\n| Public LB | Private LB |\n| --- | --- |\n| 0.54777 | 0.54769 |\n\nNotebooks\n- Inference: [here](https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v2-can40-v2-tail40-top40405)\n\n**(2) `XGBRanker` with `rank:pairwise` objective; one single model for each event type**\nScores\n| CV | Public LB | Private LB |\n| --- | --- | --- |\n| 0.512189 | 0.56940 | 0.56884 |\n\nNotebooks\n- Training: [click training notebook](https://www.kaggle.com/code/hoangnguyen719/otto-trainingclk-can40-matrixv2-tail40-top404050), [cart training notebook](https://www.kaggle.com/code/hoangnguyen719/otto-trainingcart-can40-matrixv2-tail40-top404050), [order training notebook](https://www.kaggle.com/code/hoangnguyen719/otto-trainingord-can40-matrixv2-tail40-top404050)\n- Inference: [here](https://www.kaggle.com/code/hoangnguyen719/otto-infer-all-v1-can40-v2-tail40-top404050)\n\n**Notes**\n- Method (2) was overfitting and requires a lot of hyper-parameter tuning.\n- In LB score, method (2) managed to beat (1) by 0.0216, so a model-based ranker is, understandably, better than a rule-based approach.\n- Method (1) underperforms in LB compared to Chris' original rule-based reranker notebook. The three differences between method (1) and Chris' are\n    - Session truncation: Method (1)'s co-visitation matrices use last 40 events of each session, while Chris' use last 30.\n    - Number of items in co-visitation matrice: method (1)'s matrices use top 40, 40 and 50 items for the three matrices mentioned above, while Chris' use top 15, 15 and 20.\n    - Event type weight: method (1) uses `type_weight = {0:1, 1:3, 2:6}` while Chris uses `type_weight = {0:1, 1:6, 2:3}`. However, in a separate test of mine this disparity proves not to affect the score by much.\n\n### 5. Solution Ensemble\nDue to the poor performance of my individual rankers, I ensemble them and other publicly available submissions (credit to [@karakasatarik's notebook](https://www.kaggle.com/code/karakasatarik/0-578-ensemble-of-public-notebooks) and [@vbmokin's notebook](https://www.kaggle.com/code/vbmokin/0-578-ensemble-of-public-notebooks-upgrade?scriptVersionId=116497334)), for better score. I weight each submission using their public LB score. My final two submissions are:\n**(A) Ensemble of method (2) above and public submissions**\n**(B) Ensemble of above two owned methods and public submissions**\n\nBelow are the scores of all public submissions and ensembles\n- Public LB\n| @karakasatarik's ensemble | @vbmokin's ensemble | Submission (A) | Submission (B) |\n| --- | --- | --- | --- |\n| 0.57843 | 0.57821 | 0.57884 | 0.57844 |\n\n- Private LB\n| @karakasatarik's ensemble | @vbmokin's ensemble | Submission (A) | Submission (B) |\n| --- | --- | --- | --- |\n| 0.57808 | 0.57787 | 0.57855 | 0.57811 |\n\nMy submission (B) has higher scores than did the public ensembles for private LB, while submission (A) outperforms all other ensembles in both public and private LBs. This means that both methods, though achieving relatively poor scores, did add some valuable info to the final solution (the 0.57808 -> 0.57855 improvement is equal to a boost from rank 653th to rank 255th!).\n\n## II. What I have learned as a Kaggle competition first-timer\nIt can be seen from the above sections that my work depended largely on other participants' help and support - so thank you! Below are what I have learned from this valuable experience - hope it'd be helpful for others too!\n- **Read the discussion and notebook forums** - especially if you're new to Kaggle/ML. There are always experienced participants sharing their ideas, suggestions and feedbacks, so if you are looking for a place to start the race, this is it! And check back once in a while on notebook/discussion that is helpful for you - the comment section may give you additional insights or unexpected bug-fixings that are no less valuable than the notebook.\nI myself in this competition would definitely not have got the bronze if it was not thanks to the knowledge shared by others.\n- **Try as many things as you can**. I read through solutions of some of the top achievers ([Top 4](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013), [Top 6](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382802)), and realized that they all tried more approaches (Matrix Factorization, Word2Vec) in more depth (hundred of co-visitation matrices). Of course they have done better in other aspects too (preprocessing, train-val splitting, validation, etc.), but such multi in-depth methods alone already lead to more results in better robustness. When ensembled, their solutions understandably far outperform mine.\n- **Know every line of code** you write. I lost about a month getting discouraged by CV scores going up and down unexpectedly, only to later found out a line of code misplaced in my validation step. [This can happen to anyone](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382277), so make sure you understand every line of code in your work.",
    "2129537": "Great post thanks for presenting your pipeline ! Thanks for mentioning my  topic !",
    "2129707": "Thanks Rayane! Good to see I'm not alone haha. And congratulations on your first medal!"
  },
  "source": "meta"
}