{
  "id": 366477,
  "title": "📖 What are some good resources to learn about how gradient boosted tree ranking models work?",
  "url": "/competitions/otto-recommender-system/discussion/366477",
  "author_name": "",
  "post_date": "2022-11-16T09:57:27.379047800Z",
  "votes": 22,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Do you know of any good resources on how ranking models work?</p>\n<p>For instance, how does breaking the data into sessions help the model train? How is this information used?</p>\n<p>Why the model seems to improve if I feed it more sessions EVEN if some of the sessions don't have any ground truth associated with them? (they are actual sessions but where nothing was clicked, for instance). I understand that negative examples can also be helpful, but how does the model make use of this information?</p>\n<p>I am looking for resources that would build intuition on how these models work or alternatively that speak about how to improve the performance of these models.</p>\n<p>Thank you very much for your help 🙏</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2031909",
      "postDate": "11/16/2022 09:57:27",
      "content": "<p>Do you know of any good resources on how ranking models work?</p>\n<p>For instance, how does breaking the data into sessions help the model train? How is this information used?</p>\n<p>Why the model seems to improve if I feed it more sessions EVEN if some of the sessions don't have any ground truth associated with them? (they are actual sessions but where nothing was clicked, for instance). I understand that negative examples can also be helpful, but how does the model make use of this information?</p>\n<p>I am looking for resources that would build intuition on how these models work or alternatively that speak about how to improve the performance of these models.</p>\n<p>Thank you very much for your help 🙏</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Do you know of any good resources on how ranking models work?\n\nFor instance, how does breaking the data into sessions help the model train? How is this information used?\n\nWhy the model seems to improve if I feed it more sessions EVEN if some of the sessions don't have any ground truth associated with them? (they are actual sessions but where nothing was clicked, for instance). I understand that negative examples can also be helpful, but how does the model make use of this information?\n\nI am looking for resources that would build intuition on how these models work or alternatively that speak about how to improve the performance of these models.\n\nThank you very much for your help 🙏\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2034655",
      "postDate": "11/18/2022 10:22:56",
      "content": "<p>I 'm waiting for the experienced kagglers to answer your questions on this topic, so that I'll learn too.</p>",
      "rawMarkdown": "I 'm waiting for the experienced kagglers to answer your questions on this topic, so that I'll learn too.",
      "votes": null
    },
    {
      "id": "2035208",
      "postDate": "11/18/2022 17:42:34",
      "content": "<p>It is my understanding that there are 3 ways that GBT can rank</p>\n<ul>\n<li>pointwise - learns regression on each item individually. Then sorts predicted scores</li>\n<li>pairwise - learns binary classification on each pair of items. Then converts predictions into ranked list</li>\n<li>listwise - uses complex algorithms that analyze entire list at once. There are many different algorithms and research papers.</li>\n</ul>\n<p>When choosing parameters for GBT, we can choose which method <code>objective</code> we want to use. Here is Wikipedia page about ranking <a href=\"https://en.wikipedia.org/wiki/Learning_to_rank\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "It is my understanding that there are 3 ways that GBT can rank\n* pointwise - learns regression on each item individually. Then sorts predicted scores\n* pairwise - learns binary classification on each pair of items. Then converts predictions into ranked list\n* listwise - uses complex algorithms that analyze entire list at once. There are many different algorithms and research papers.\n\nWhen choosing parameters for GBT, we can choose which method `objective` we want to use. Here is Wikipedia page about ranking [here][1]\n\n[1]: https://en.wikipedia.org/wiki/Learning_to_rank",
      "votes": null
    },
    {
      "id": "2035360",
      "postDate": "11/18/2022 19:35:25",
      "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙂 Very valuable information, as always 🙂</p>\n<p>Now the question is how do various frameworks expose these options and whether they do 🤔 I only played around very briefly with lgbm in this competition so far, I believe it only supports <code>lambdarank</code> as the objective.</p>\n<p>Not too much information on that and the documentation of the library is not great, but I found <a href=\"https://ffineis.github.io/blog/2021/05/01/lambdarank-lightgbm.html\" target=\"_blank\">this blog post quite informative</a>.</p>\n<p>Next steps would probably be <code>xgboost</code> and possibly even <code>catboost</code>! 🙂</p>",
      "rawMarkdown": "Thank you very much, @cdeotte! 🙂 Very valuable information, as always 🙂\n\nNow the question is how do various frameworks expose these options and whether they do 🤔 I only played around very briefly with lgbm in this competition so far, I believe it only supports `lambdarank` as the objective.\n\nNot too much information on that and the documentation of the library is not great, but I found [this blog post quite informative](https://ffineis.github.io/blog/2021/05/01/lambdarank-lightgbm.html).\n\nNext steps would probably be `xgboost` and possibly even `catboost`! 🙂",
      "votes": null
    },
    {
      "id": "2035696",
      "postDate": "11/19/2022 05:59:33",
      "content": "<p>You can go for voting or bagging<br>\nalso do check my profile</p>",
      "rawMarkdown": "You can go for voting or bagging\nalso do check my profile",
      "votes": null
    },
    {
      "id": "2041941",
      "postDate": "11/24/2022 10:13:38",
      "content": "<p>try this paper <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ,<a href=\"https://www.microsoft.com/en-us/research/uploads/prod/2016/02/MSR-TR-2010-82.pdf\" target=\"_blank\">https://www.microsoft.com/en-us/research/uploads/prod/2016/02/MSR-TR-2010-82.pdf</a> this gives clear info</p>",
      "rawMarkdown": "try this paper @radek1 ,https://www.microsoft.com/en-us/research/uploads/prod/2016/02/MSR-TR-2010-82.pdf this gives clear info",
      "votes": null
    },
    {
      "id": "2041945",
      "postDate": "11/24/2022 10:16:21",
      "content": "<p>Thank you for the recommendation, <a href=\"https://www.kaggle.com/satyanarayanam\" target=\"_blank\">@satyanarayanam</a>! 🙌🙂 </p>",
      "rawMarkdown": "Thank you for the recommendation, @satyanarayanam! 🙌🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2034655,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "11/18/2022 10:22:56",
      "content": "<p>I 'm waiting for the experienced kagglers to answer your questions on this topic, so that I'll learn too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2035208,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/18/2022 17:42:34",
      "content": "<p>It is my understanding that there are 3 ways that GBT can rank</p>\n<ul>\n<li>pointwise - learns regression on each item individually. Then sorts predicted scores</li>\n<li>pairwise - learns binary classification on each pair of items. Then converts predictions into ranked list</li>\n<li>listwise - uses complex algorithms that analyze entire list at once. There are many different algorithms and research papers.</li>\n</ul>\n<p>When choosing parameters for GBT, we can choose which method <code>objective</code> we want to use. Here is Wikipedia page about ranking <a href=\"https://en.wikipedia.org/wiki/Learning_to_rank\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2035360,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/18/2022 19:35:25",
          "content": "<p>Thank you very much, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙂 Very valuable information, as always 🙂</p>\n<p>Now the question is how do various frameworks expose these options and whether they do 🤔 I only played around very briefly with lgbm in this competition so far, I believe it only supports <code>lambdarank</code> as the objective.</p>\n<p>Not too much information on that and the documentation of the library is not great, but I found <a href=\"https://ffineis.github.io/blog/2021/05/01/lambdarank-lightgbm.html\" target=\"_blank\">this blog post quite informative</a>.</p>\n<p>Next steps would probably be <code>xgboost</code> and possibly even <code>catboost</code>! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2035696,
      "author_name": "atharvakadam333",
      "author_url": "",
      "post_date": "11/19/2022 05:59:33",
      "content": "<p>You can go for voting or bagging<br>\nalso do check my profile</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2041941,
      "author_name": "satyanarayanam",
      "author_url": "",
      "post_date": "11/24/2022 10:13:38",
      "content": "<p>try this paper <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ,<a href=\"https://www.microsoft.com/en-us/research/uploads/prod/2016/02/MSR-TR-2010-82.pdf\" target=\"_blank\">https://www.microsoft.com/en-us/research/uploads/prod/2016/02/MSR-TR-2010-82.pdf</a> this gives clear info</p>",
      "votes": null,
      "replies": [
        {
          "id": 2041945,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/24/2022 10:16:21",
          "content": "<p>Thank you for the recommendation, <a href=\"https://www.kaggle.com/satyanarayanam\" target=\"_blank\">@satyanarayanam</a>! 🙌🙂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2031909": "Do you know of any good resources on how ranking models work?\n\nFor instance, how does breaking the data into sessions help the model train? How is this information used?\n\nWhy the model seems to improve if I feed it more sessions EVEN if some of the sessions don't have any ground truth associated with them? (they are actual sessions but where nothing was clicked, for instance). I understand that negative examples can also be helpful, but how does the model make use of this information?\n\nI am looking for resources that would build intuition on how these models work or alternatively that speak about how to improve the performance of these models.\n\nThank you very much for your help 🙏\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2034655": "I 'm waiting for the experienced kagglers to answer your questions on this topic, so that I'll learn too.",
    "2035208": "It is my understanding that there are 3 ways that GBT can rank\n* pointwise - learns regression on each item individually. Then sorts predicted scores\n* pairwise - learns binary classification on each pair of items. Then converts predictions into ranked list\n* listwise - uses complex algorithms that analyze entire list at once. There are many different algorithms and research papers.\n\nWhen choosing parameters for GBT, we can choose which method `objective` we want to use. Here is Wikipedia page about ranking [here][1]\n\n[1]: https://en.wikipedia.org/wiki/Learning_to_rank",
    "2035360": "Thank you very much, @cdeotte! 🙂 Very valuable information, as always 🙂\n\nNow the question is how do various frameworks expose these options and whether they do 🤔 I only played around very briefly with lgbm in this competition so far, I believe it only supports `lambdarank` as the objective.\n\nNot too much information on that and the documentation of the library is not great, but I found [this blog post quite informative](https://ffineis.github.io/blog/2021/05/01/lambdarank-lightgbm.html).\n\nNext steps would probably be `xgboost` and possibly even `catboost`! 🙂",
    "2035696": "You can go for voting or bagging\nalso do check my profile",
    "2041941": "try this paper @radek1 ,https://www.microsoft.com/en-us/research/uploads/prod/2016/02/MSR-TR-2010-82.pdf this gives clear info",
    "2041945": "Thank you for the recommendation, @satyanarayanam! 🙌🙂"
  },
  "source": "meta"
}