{
  "id": 366474,
  "title": "💡 Can you beat static rules with a ranker model without additional features?",
  "url": "/competitions/otto-recommender-system/discussion/366474",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-16T09:52:33.641000",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I have encountered the same situation in the H&amp;M competition as here and it is driving me borderline crazy 🙂</p>\n<p>In my <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194\" target=\"_blank\">[Starter pack] LGBMRanker with polars 🚀🚀🚀</a> I created features that are a superset of features that are used in co-visitation matrices.</p>\n<p>But try as hard as I might, I cannot beat the score I attained with the co-visitation matrix 🤯.</p>\n<p>I experienced the same in the H&amp;M competition. There was a popularity-based heuristic (based on bestsellers from the previous week). Despite feeding this data to my ranking model, I could edge closer to the performance of the heuristic, but never quite get there.</p>\n<p>I realize that \"the data is noisy\". One issue I observed in the H&amp;M competition was that certain weeks would have more purchases and that it was one of the things that were throwing off the model (while for any week we could rank bestsellers 1, 2, 3 etc but if we summed the bestsellers for all the week, the count of bestsellers with rank 2 per week might have been &gt; count of bestsellers with position 1)</p>\n<p>Here, again, we have a similar situation. Seems static rules are a distillation, they represent the essence, of what a model could possibly learn from the data. Like in Plato's cave, the model only has access to the instantiation of the ideal, the shadows, while with rules we can pierce through the noisy realization and get at what generates the shadows, the data.</p>\n<p>Of course, this only works if we are operating on a low number of features and we tweak our heuristic approach quite a bit.</p>\n<p>So here is my question. Does my reasoning sound correct? Is this to be expected? That we can hope for our ranking model to outperform competitive static rules, but in order to do so, we need to provide it with additional information?</p>\n<p>But on the same information, with just a handful of features, handcrafted rules, if tweaked appropriately, are likely to outperform a model and that is all good?</p>\n<p>Has that also been your experience?</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2031901,
      "postDate": "2022-11-16T09:52:33.643Z",
      "content": "<p>I have encountered the same situation in the H&amp;M competition as here and it is driving me borderline crazy 🙂</p>\n<p>In my <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194\" target=\"_blank\">[Starter pack] LGBMRanker with polars 🚀🚀🚀</a> I created features that are a superset of features that are used in co-visitation matrices.</p>\n<p>But try as hard as I might, I cannot beat the score I attained with the co-visitation matrix 🤯.</p>\n<p>I experienced the same in the H&amp;M competition. There was a popularity-based heuristic (based on bestsellers from the previous week). Despite feeding this data to my ranking model, I could edge closer to the performance of the heuristic, but never quite get there.</p>\n<p>I realize that \"the data is noisy\". One issue I observed in the H&amp;M competition was that certain weeks would have more purchases and that it was one of the things that were throwing off the model (while for any week we could rank bestsellers 1, 2, 3 etc but if we summed the bestsellers for all the week, the count of bestsellers with rank 2 per week might have been &gt; count of bestsellers with position 1)</p>\n<p>Here, again, we have a similar situation. Seems static rules are a distillation, they represent the essence, of what a model could possibly learn from the data. Like in Plato's cave, the model only has access to the instantiation of the ideal, the shadows, while with rules we can pierce through the noisy realization and get at what generates the shadows, the data.</p>\n<p>Of course, this only works if we are operating on a low number of features and we tweak our heuristic approach quite a bit.</p>\n<p>So here is my question. Does my reasoning sound correct? Is this to be expected? That we can hope for our ranking model to outperform competitive static rules, but in order to do so, we need to provide it with additional information?</p>\n<p>But on the same information, with just a handful of features, handcrafted rules, if tweaked appropriately, are likely to outperform a model and that is all good?</p>\n<p>Has that also been your experience?</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "I have encountered the same situation in the H&M competition as here and it is driving me borderline crazy 🙂\n\nIn my [[Starter pack] LGBMRanker with polars 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194) I created features that are a superset of features that are used in co-visitation matrices.\n\nBut try as hard as I might, I cannot beat the score I attained with the co-visitation matrix 🤯.\n\nI experienced the same in the H&M competition. There was a popularity-based heuristic (based on bestsellers from the previous week). Despite feeding this data to my ranking model, I could edge closer to the performance of the heuristic, but never quite get there.\n\nI realize that \"the data is noisy\". One issue I observed in the H&M competition was that certain weeks would have more purchases and that it was one of the things that were throwing off the model (while for any week we could rank bestsellers 1, 2, 3 etc but if we summed the bestsellers for all the week, the count of bestsellers with rank 2 per week might have been > count of bestsellers with position 1)\n\nHere, again, we have a similar situation. Seems static rules are a distillation, they represent the essence, of what a model could possibly learn from the data. Like in Plato's cave, the model only has access to the instantiation of the ideal, the shadows, while with rules we can pierce through the noisy realization and get at what generates the shadows, the data.\n\nOf course, this only works if we are operating on a low number of features and we tweak our heuristic approach quite a bit.\n\nSo here is my question. Does my reasoning sound correct? Is this to be expected? That we can hope for our ranking model to outperform competitive static rules, but in order to do so, we need to provide it with additional information?\n\nBut on the same information, with just a handful of features, handcrafted rules, if tweaked appropriately, are likely to outperform a model and that is all good?\n\nHas that also been your experience?\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": 24
    },
    {
      "id": 2032493,
      "postDate": "2022-11-16T16:28:42.950Z",
      "content": "<p>Have you tried adding the co-visitation items to (an offline version of your) LGBM reranker notebook? Currently I notice that it only trains with <code>otto-train-and-test-data-for-local-validation/test.parquet</code>. But we need a new row for each co-visitation item. And we need to add them during inference too. If we wish to beat the CV LB of co-visitation notebook.</p>\n<p>Then after adding the co-visitation items, we need to engineer some new features like the boolean <code>is_item_in_user_history</code> and some others. Also i don't think that your current feature set captures which history items are repeats. I think we need to add an item history count variable (or better yet, time weighted item history count feature).</p>",
      "rawMarkdown": "Have you tried adding the co-visitation items to (an offline version of your) LGBM reranker notebook? Currently I notice that it only trains with `otto-train-and-test-data-for-local-validation/test.parquet`. But we need a new row for each co-visitation item. And we need to add them during inference too. If we wish to beat the CV LB of co-visitation notebook.\n\nThen after adding the co-visitation items, we need to engineer some new features like the boolean `is_item_in_user_history` and some others. Also i don't think that your current feature set captures which history items are repeats. I think we need to add an item history count variable (or better yet, time weighted item history count feature).",
      "votes": 12,
      "replies": [
        {
          "id": 2032932,
          "postDate": "2022-11-16T22:24:32.307Z",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙏 You are like a good spirit of the Kaggle forums guiding us newbs! 🙂 Really appreciate your advice! 🙏🙏🙏</p>\n<p>I haven't been very clear in my post, but what I did is I took just the weighted processing of user history from the covisitation matrix approach. Where sessions longer than 20 have aids weighted by type and recency. With this approach, on <code>clicks</code>, I get a local recall@20 of <code>0.3235</code>.</p>\n<p>It is interesting how far one gets with just the last 20 aids strategy, without any reranking: <code>0.3204</code>.</p>\n<p>Using an LGBM reranker, seemingly providing it a superset of features used in the static ruleset that gets us <code>0.3235</code> above, the best I was able to get to was <code>0.3230</code>.</p>\n<p>That is what I found so frustrating that prompted me to writhe the OP 🙂 That no matter what I do, I cannot beat the static rule on this simplified example.</p>\n<p>But thank you very much for your suggestions on using additional features (and a recommendation of what features to use!).</p>\n<p>I am trying to wrap my head around what is happening with these re-ranking models and so far my understanding seems to be that these static rules work extremely well with small amount of features that carry a lot of signal. That it might be hard or even impossible to beat these with a ranking model at times.</p>\n<p>But once we move into a larger feature set regime (which is a threshold one crosses very quickly 🙂), we are unable to make use of all that signal via static rules and that is where those trainable ranking models really shine.</p>\n<p>Hoping my intuition is not too much off as to what might be happening (though I might also be splitting the straw unnecessarily 😉 on the other hand such musings help me understand better what the models are doing and what are their strengths nd weaknesses).</p>\n<p>Thank you so much again for your input 🙏It helps my understanding a lot.</p>",
          "rawMarkdown": "Thank you so much @cdeotte! 🙏 You are like a good spirit of the Kaggle forums guiding us newbs! 🙂 Really appreciate your advice! 🙏🙏🙏\n\nI haven't been very clear in my post, but what I did is I took just the weighted processing of user history from the covisitation matrix approach. Where sessions longer than 20 have aids weighted by type and recency. With this approach, on `clicks`, I get a local recall@20 of `0.3235`.\n\nIt is interesting how far one gets with just the last 20 aids strategy, without any reranking: `0.3204`.\n\nUsing an LGBM reranker, seemingly providing it a superset of features used in the static ruleset that gets us `0.3235` above, the best I was able to get to was `0.3230`.\n\nThat is what I found so frustrating that prompted me to writhe the OP 🙂 That no matter what I do, I cannot beat the static rule on this simplified example.\n\nBut thank you very much for your suggestions on using additional features (and a recommendation of what features to use!).\n\nI am trying to wrap my head around what is happening with these re-ranking models and so far my understanding seems to be that these static rules work extremely well with small amount of features that carry a lot of signal. That it might be hard or even impossible to beat these with a ranking model at times.\n\nBut once we move into a larger feature set regime (which is a threshold one crosses very quickly 🙂), we are unable to make use of all that signal via static rules and that is where those trainable ranking models really shine.\n\nHoping my intuition is not too much off as to what might be happening (though I might also be splitting the straw unnecessarily 😉 on the other hand such musings help me understand better what the models are doing and what are their strengths nd weaknesses).\n\nThank you so much again for your input 🙏It helps my understanding a lot.",
          "votes": 4
        },
        {
          "id": 2032968,
          "postDate": "2022-11-16T23:59:17.087Z",
          "content": "<p>For what it's worth, I was able to nearly match the score using the ranker 🙂</p>\n<p>Ranker: <code>0.3234616931372449</code><br>\nHeuristic: <code>0.323488</code></p>\n<p>Of course, it is such a tiny difference it completely doesn't matter. Had to add a couple more columns 🙂 Here they are along with feature importance coming from the model:</p>\n<pre><code>action_num_reverse_chrono 0.658588344124041\naid 0.11817917885568481\nsec_since_session_start 0.07308456404382999\nsession_length 0.0486059102181944\nrelative_position_in_session 0.04306810550235335\ntype_weighted_log_recency_score 0.0403055170220625\ntype 0.011932710447021396\nlog_recency_score 0.0028580939171428945\naid_clicked_count 0.002075661636954335\naid_carted_count 0.0013019142327152986\naid_ordered_count 0.0\naid_seen_by_type_count 0.0\naid_seen_count 0.0\n</code></pre>\n<p>Let's see how it goes with data from the co-visitation matrix 🙂</p>",
          "rawMarkdown": "For what it's worth, I was able to nearly match the score using the ranker 🙂\n\nRanker: `0.3234616931372449`\nHeuristic: `0.323488`\n\nOf course, it is such a tiny difference it completely doesn't matter. Had to add a couple more columns 🙂 Here they are along with feature importance coming from the model:\n\n```\naction_num_reverse_chrono 0.658588344124041\naid 0.11817917885568481\nsec_since_session_start 0.07308456404382999\nsession_length 0.0486059102181944\nrelative_position_in_session 0.04306810550235335\ntype_weighted_log_recency_score 0.0403055170220625\ntype 0.011932710447021396\nlog_recency_score 0.0028580939171428945\naid_clicked_count 0.002075661636954335\naid_carted_count 0.0013019142327152986\naid_ordered_count 0.0\naid_seen_by_type_count 0.0\naid_seen_count 0.0\n```\n\nLet's see how it goes with data from the co-visitation matrix 🙂",
          "votes": 3
        },
        {
          "id": 2032979,
          "postDate": "2022-11-17T00:15:50.177Z",
          "content": "<p>When a user clicks an item multiple times in their history, do you add that item multiple times to your reranker train data?</p>\n<p>Let's say a user clicked the item <code>12345</code> four times previously. I think your pipeline will add 4 new rows one with weight 0.1 one with weight 0.15 one with weight 0.2 etc. We do not want to add the same item 4 times to our reranker train data.</p>\n<p>To make your notebook like the heuristic reranker, we need to instead make 1 new row with the sum of the weights. So we add one new row for item <code>12345</code> and give it weight <code>0.1+0.15+0.2+0.25</code>. Or we can make our own new feature like <code>history_count = 4</code>. </p>\n<p>Likewise when adding more rows from co-visitation matrices, we must aggregate results and make sure reranker train data only has 1 row per item. An XGB model cannot aggregate multiple rows for the same item in its training. We must do this in our preprocess.</p>",
          "rawMarkdown": "When a user clicks an item multiple times in their history, do you add that item multiple times to your reranker train data?\n\nLet's say a user clicked the item `12345` four times previously. I think your pipeline will add 4 new rows one with weight 0.1 one with weight 0.15 one with weight 0.2 etc. We do not want to add the same item 4 times to our reranker train data.\n\nTo make your notebook like the heuristic reranker, we need to instead make 1 new row with the sum of the weights. So we add one new row for item `12345` and give it weight `0.1+0.15+0.2+0.25`. Or we can make our own new feature like `history_count = 4`. \n\nLikewise when adding more rows from co-visitation matrices, we must aggregate results and make sure reranker train data only has 1 row per item. An XGB model cannot aggregate multiple rows for the same item in its training. We must do this in our preprocess.",
          "votes": 6
        },
        {
          "id": 2032984,
          "postDate": "2022-11-17T00:18:26.050Z",
          "content": "<p>The magic of the reranker is defining feature columns for the item and user that the XGB can understand and use to properly rank the items. If the features are designed correctly, the reranker should always beat heuristics.</p>",
          "rawMarkdown": "The magic of the reranker is defining feature columns for the item and user that the XGB can understand and use to properly rank the items. If the features are designed correctly, the reranker should always beat heuristics.",
          "votes": 5
        },
        {
          "id": 2033087,
          "postDate": "2022-11-17T02:59:15.943Z",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 😊 It was exactly what you mentioned, I wasn't aggregating the history properly (I dropped duplicates keeping last but I didn't sum the recency scores).</p>\n<p>With that fixed, the LGBM reranker beats the heuristic reranker and attains a score of <code>0.3235123899622565</code>! 🙂</p>\n<p>(That is a very small difference in score, but on this size of the validation set the results seem to be quite consistent over multiple runs).</p>\n<blockquote>\n  <p>If the features are designed correctly, the reranker should always beat heuristics.</p>\n</blockquote>\n<p>🙏</p>",
          "rawMarkdown": "Thank you very much, @cdeotte! 😊 It was exactly what you mentioned, I wasn't aggregating the history properly (I dropped duplicates keeping last but I didn't sum the recency scores).\n\nWith that fixed, the LGBM reranker beats the heuristic reranker and attains a score of `0.3235123899622565`! 🙂\n\n(That is a very small difference in score, but on this size of the validation set the results seem to be quite consistent over multiple runs).\n\n> If the features are designed correctly, the reranker should always beat heuristics.\n\n🙏",
          "votes": 3
        },
        {
          "id": 2033398,
          "postDate": "2022-11-17T08:33:23.887Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2042728,
          "postDate": "2022-11-24T23:15:43.790Z",
          "content": "<p>Hey radek. One question: could you point to the part of the code where you are aggregating the User history ? It's not completly clear to me where this happens :) </p>",
          "rawMarkdown": "Hey radek. One question: could you point to the part of the code where you are aggregating the User history ? It's not completly clear to me where this happens :) "
        },
        {
          "id": 2042751,
          "postDate": "2022-11-25T00:18:16.813Z",
          "content": "<p>That is the thing! I am not sure I realized all this while putting together the notebook! I didn't aggregate the user history there at all I think (might be worth checking by seeing if there re duplicate aid entries per session).</p>\n<p>This is from my code I have locally, I think it holds the key to how to do it 🙂</p>\n<pre><code>def sum_recency_scores(df):\n    return df.with_columns([\n        pl.col('^log_recency_score|type_weighted_log_recency_score$').cumsum().over(['session', 'aid'])\n    ])\n</code></pre>\n<p>Essentially, you can stick a column name in <code>pl.col(&lt;col_name&gt;)</code> and then you just need to keep the last entry for an <code>aid</code> in a session using <code>unique</code>. I believe that is all there is to it 🙂 But it is only once I've made the mental leap that it was needed thx to the comment from Chris that it all clicked in 😄</p>",
          "rawMarkdown": "That is the thing! I am not sure I realized all this while putting together the notebook! I didn't aggregate the user history there at all I think (might be worth checking by seeing if there re duplicate aid entries per session).\n\nThis is from my code I have locally, I think it holds the key to how to do it 🙂\n\n```\ndef sum_recency_scores(df):\n    return df.with_columns([\n        pl.col('^log_recency_score|type_weighted_log_recency_score$').cumsum().over(['session', 'aid'])\n    ])\n```\n\nEssentially, you can stick a column name in `pl.col(<col_name>)` and then you just need to keep the last entry for an `aid` in a session using `unique`. I believe that is all there is to it 🙂 But it is only once I've made the mental leap that it was needed thx to the comment from Chris that it all clicked in 😄\n",
          "votes": 3
        },
        {
          "id": 2042965,
          "postDate": "2022-11-25T06:23:19.427Z",
          "content": "<p>Yea makes sense. The \"over\" takes care that we accumulate for each Session for every item and cumsum is exactly what what Chris described!<br>\nBut I will double check myself later.<br>\nTnx :))</p>",
          "rawMarkdown": "Yea makes sense. The \"over\" takes care that we accumulate for each Session for every item and cumsum is exactly what what Chris described!\nBut I will double check myself later.\nTnx :))",
          "votes": 1
        }
      ]
    },
    {
      "id": 2038547,
      "postDate": "2022-11-21T13:30:09.730Z",
      "content": "<p>Another amazing and super beneficial discussions between <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! You guys make this competition a super learning environment!!! Endless thanks to you both!</p>",
      "rawMarkdown": "Another amazing and super beneficial discussions between @radek1 and @cdeotte! You guys make this competition a super learning environment!!! Endless thanks to you both!",
      "votes": 6
    }
  ],
  "comments": [
    {
      "id": 2032493,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-11-16T16:28:42.950000",
      "content": "<p>Have you tried adding the co-visitation items to (an offline version of your) LGBM reranker notebook? Currently I notice that it only trains with <code>otto-train-and-test-data-for-local-validation/test.parquet</code>. But we need a new row for each co-visitation item. And we need to add them during inference too. If we wish to beat the CV LB of co-visitation notebook.</p>\n<p>Then after adding the co-visitation items, we need to engineer some new features like the boolean <code>is_item_in_user_history</code> and some others. Also i don't think that your current feature set captures which history items are repeats. I think we need to add an item history count variable (or better yet, time weighted item history count feature).</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2032932,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-16T22:24:32.307000",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙏 You are like a good spirit of the Kaggle forums guiding us newbs! 🙂 Really appreciate your advice! 🙏🙏🙏</p>\n<p>I haven't been very clear in my post, but what I did is I took just the weighted processing of user history from the covisitation matrix approach. Where sessions longer than 20 have aids weighted by type and recency. With this approach, on <code>clicks</code>, I get a local recall@20 of <code>0.3235</code>.</p>\n<p>It is interesting how far one gets with just the last 20 aids strategy, without any reranking: <code>0.3204</code>.</p>\n<p>Using an LGBM reranker, seemingly providing it a superset of features used in the static ruleset that gets us <code>0.3235</code> above, the best I was able to get to was <code>0.3230</code>.</p>\n<p>That is what I found so frustrating that prompted me to writhe the OP 🙂 That no matter what I do, I cannot beat the static rule on this simplified example.</p>\n<p>But thank you very much for your suggestions on using additional features (and a recommendation of what features to use!).</p>\n<p>I am trying to wrap my head around what is happening with these re-ranking models and so far my understanding seems to be that these static rules work extremely well with small amount of features that carry a lot of signal. That it might be hard or even impossible to beat these with a ranking model at times.</p>\n<p>But once we move into a larger feature set regime (which is a threshold one crosses very quickly 🙂), we are unable to make use of all that signal via static rules and that is where those trainable ranking models really shine.</p>\n<p>Hoping my intuition is not too much off as to what might be happening (though I might also be splitting the straw unnecessarily 😉 on the other hand such musings help me understand better what the models are doing and what are their strengths nd weaknesses).</p>\n<p>Thank you so much again for your input 🙏It helps my understanding a lot.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2032968,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-16T23:59:17.087000",
          "content": "<p>For what it's worth, I was able to nearly match the score using the ranker 🙂</p>\n<p>Ranker: <code>0.3234616931372449</code><br>\nHeuristic: <code>0.323488</code></p>\n<p>Of course, it is such a tiny difference it completely doesn't matter. Had to add a couple more columns 🙂 Here they are along with feature importance coming from the model:</p>\n<pre><code>action_num_reverse_chrono 0.658588344124041\naid 0.11817917885568481\nsec_since_session_start 0.07308456404382999\nsession_length 0.0486059102181944\nrelative_position_in_session 0.04306810550235335\ntype_weighted_log_recency_score 0.0403055170220625\ntype 0.011932710447021396\nlog_recency_score 0.0028580939171428945\naid_clicked_count 0.002075661636954335\naid_carted_count 0.0013019142327152986\naid_ordered_count 0.0\naid_seen_by_type_count 0.0\naid_seen_count 0.0\n</code></pre>\n<p>Let's see how it goes with data from the co-visitation matrix 🙂</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2032979,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-17T00:15:50.177000",
          "content": "<p>When a user clicks an item multiple times in their history, do you add that item multiple times to your reranker train data?</p>\n<p>Let's say a user clicked the item <code>12345</code> four times previously. I think your pipeline will add 4 new rows one with weight 0.1 one with weight 0.15 one with weight 0.2 etc. We do not want to add the same item 4 times to our reranker train data.</p>\n<p>To make your notebook like the heuristic reranker, we need to instead make 1 new row with the sum of the weights. So we add one new row for item <code>12345</code> and give it weight <code>0.1+0.15+0.2+0.25</code>. Or we can make our own new feature like <code>history_count = 4</code>. </p>\n<p>Likewise when adding more rows from co-visitation matrices, we must aggregate results and make sure reranker train data only has 1 row per item. An XGB model cannot aggregate multiple rows for the same item in its training. We must do this in our preprocess.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2032984,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-17T00:18:26.050000",
          "content": "<p>The magic of the reranker is defining feature columns for the item and user that the XGB can understand and use to properly rank the items. If the features are designed correctly, the reranker should always beat heuristics.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2033087,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-17T02:59:15.943000",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 😊 It was exactly what you mentioned, I wasn't aggregating the history properly (I dropped duplicates keeping last but I didn't sum the recency scores).</p>\n<p>With that fixed, the LGBM reranker beats the heuristic reranker and attains a score of <code>0.3235123899622565</code>! 🙂</p>\n<p>(That is a very small difference in score, but on this size of the validation set the results seem to be quite consistent over multiple runs).</p>\n<blockquote>\n  <p>If the features are designed correctly, the reranker should always beat heuristics.</p>\n</blockquote>\n<p>🙏</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2033398,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-17T08:33:23.887000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2042728,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-11-24T23:15:43.790000",
          "content": "<p>Hey radek. One question: could you point to the part of the code where you are aggregating the User history ? It's not completly clear to me where this happens :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2042751,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-25T00:18:16.813000",
          "content": "<p>That is the thing! I am not sure I realized all this while putting together the notebook! I didn't aggregate the user history there at all I think (might be worth checking by seeing if there re duplicate aid entries per session).</p>\n<p>This is from my code I have locally, I think it holds the key to how to do it 🙂</p>\n<pre><code>def sum_recency_scores(df):\n    return df.with_columns([\n        pl.col('^log_recency_score|type_weighted_log_recency_score$').cumsum().over(['session', 'aid'])\n    ])\n</code></pre>\n<p>Essentially, you can stick a column name in <code>pl.col(&lt;col_name&gt;)</code> and then you just need to keep the last entry for an <code>aid</code> in a session using <code>unique</code>. I believe that is all there is to it 🙂 But it is only once I've made the mental leap that it was needed thx to the comment from Chris that it all clicked in 😄</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2042965,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-11-25T06:23:19.427000",
          "content": "<p>Yea makes sense. The \"over\" takes care that we accumulate for each Session for every item and cumsum is exactly what what Chris described!<br>\nBut I will double check myself later.<br>\nTnx :))</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2038547,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "2022-11-21T13:30:09.730000",
      "content": "<p>Another amazing and super beneficial discussions between <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! You guys make this competition a super learning environment!!! Endless thanks to you both!</p>",
      "votes": 6,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031901": "I have encountered the same situation in the H&M competition as here and it is driving me borderline crazy 🙂\n\nIn my [[Starter pack] LGBMRanker with polars 🚀🚀🚀](https://www.kaggle.com/competitions/otto-recommender-system/discussion/366194) I created features that are a superset of features that are used in co-visitation matrices.\n\nBut try as hard as I might, I cannot beat the score I attained with the co-visitation matrix 🤯.\n\nI experienced the same in the H&M competition. There was a popularity-based heuristic (based on bestsellers from the previous week). Despite feeding this data to my ranking model, I could edge closer to the performance of the heuristic, but never quite get there.\n\nI realize that \"the data is noisy\". One issue I observed in the H&M competition was that certain weeks would have more purchases and that it was one of the things that were throwing off the model (while for any week we could rank bestsellers 1, 2, 3 etc but if we summed the bestsellers for all the week, the count of bestsellers with rank 2 per week might have been > count of bestsellers with position 1)\n\nHere, again, we have a similar situation. Seems static rules are a distillation, they represent the essence, of what a model could possibly learn from the data. Like in Plato's cave, the model only has access to the instantiation of the ideal, the shadows, while with rules we can pierce through the noisy realization and get at what generates the shadows, the data.\n\nOf course, this only works if we are operating on a low number of features and we tweak our heuristic approach quite a bit.\n\nSo here is my question. Does my reasoning sound correct? Is this to be expected? That we can hope for our ranking model to outperform competitive static rules, but in order to do so, we need to provide it with additional information?\n\nBut on the same information, with just a handful of features, handcrafted rules, if tweaked appropriately, are likely to outperform a model and that is all good?\n\nHas that also been your experience?\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2032493": "Have you tried adding the co-visitation items to (an offline version of your) LGBM reranker notebook? Currently I notice that it only trains with `otto-train-and-test-data-for-local-validation/test.parquet`. But we need a new row for each co-visitation item. And we need to add them during inference too. If we wish to beat the CV LB of co-visitation notebook.\n\nThen after adding the co-visitation items, we need to engineer some new features like the boolean `is_item_in_user_history` and some others. Also i don't think that your current feature set captures which history items are repeats. I think we need to add an item history count variable (or better yet, time weighted item history count feature).",
    "2038547": "Another amazing and super beneficial discussions between @radek1 and @cdeotte! You guys make this competition a super learning environment!!! Endless thanks to you both!"
  }
}