{
  "id": 366194,
  "title": "[Starter pack] LGBMRanker with polars 🚀🚀🚀",
  "url": "/competitions/otto-recommender-system/discussion/366194",
  "author_name": "",
  "post_date": "2022-11-15T10:00:31.373179300Z",
  "votes": 34,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In his very informative post, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a> <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> explains how re-ranking models are the industry standard for dealing with datasets like we are presented with in this competition, that is ones with high cardinality categories!</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook\" target=\"_blank\">This starter pack</a> builds exactly such a model!</p>\n<p>I preprocess the data using <a href=\"https://www.pola.rs/\" target=\"_blank\">polars</a> (mostly to try out the library and to decrease memory footprint vs <code>pandas</code>, the added speed is also nice 🙂).</p>\n<p>I then proceed to train an <code>LGBMRanker</code>, output predictions and submit to the competition.</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook\" target=\"_blank\">The notebook</a> provides all the low-level plumbing that you can build on 🙂 The next steps here might include using co-visitation matrices for candidate generation or scoring. One can also generate additional features (such as times an item was clicked or put in cart, popularity scores, temporal features, etc).</p>\n<p>Happy hacking! 🙂 </p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2030220",
      "postDate": "11/15/2022 10:00:31",
      "content": "<p>In his very informative post, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a> <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> explains how re-ranking models are the industry standard for dealing with datasets like we are presented with in this competition, that is ones with high cardinality categories!</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook\" target=\"_blank\">This starter pack</a> builds exactly such a model!</p>\n<p>I preprocess the data using <a href=\"https://www.pola.rs/\" target=\"_blank\">polars</a> (mostly to try out the library and to decrease memory footprint vs <code>pandas</code>, the added speed is also nice 🙂).</p>\n<p>I then proceed to train an <code>LGBMRanker</code>, output predictions and submit to the competition.</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook\" target=\"_blank\">The notebook</a> provides all the low-level plumbing that you can build on 🙂 The next steps here might include using co-visitation matrices for candidate generation or scoring. One can also generate additional features (such as times an item was clicked or put in cart, popularity scores, temporal features, etc).</p>\n<p>Happy hacking! 🙂 </p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "In his very informative post, [Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) @ravishah1 explains how re-ranking models are the industry standard for dealing with datasets like we are presented with in this competition, that is ones with high cardinality categories!\n\n[This starter pack](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook) builds exactly such a model!\n\nI preprocess the data using [polars](https://www.pola.rs/) (mostly to try out the library and to decrease memory footprint vs `pandas`, the added speed is also nice 🙂).\n\nI then proceed to train an `LGBMRanker`, output predictions and submit to the competition.\n\n[The notebook](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook) provides all the low-level plumbing that you can build on 🙂 The next steps here might include using co-visitation matrices for candidate generation or scoring. One can also generate additional features (such as times an item was clicked or put in cart, popularity scores, temporal features, etc).\n\nHappy hacking! 🙂 \n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2032531",
      "postDate": "11/16/2022 16:52:35",
      "content": "<p>Thanks for the starter Radek. In the last Kaggle recsys comp, ML reranker beat heuristic reranker. So i think the winning teams will be using an approach like this.</p>",
      "rawMarkdown": "Thanks for the starter Radek. In the last Kaggle recsys comp, ML reranker beat heuristic reranker. So i think the winning teams will be using an approach like this.",
      "votes": null
    },
    {
      "id": "2032933",
      "postDate": "11/16/2022 22:27:29",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙌 That is a very valuable piece of information 🙂</p>",
      "rawMarkdown": "Thank you, @cdeotte! 🙌 That is a very valuable piece of information 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2032531,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/16/2022 16:52:35",
      "content": "<p>Thanks for the starter Radek. In the last Kaggle recsys comp, ML reranker beat heuristic reranker. So i think the winning teams will be using an approach like this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2032933,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/16/2022 22:27:29",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙌 That is a very valuable piece of information 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2030220": "In his very informative post, [Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) @ravishah1 explains how re-ranking models are the industry standard for dealing with datasets like we are presented with in this competition, that is ones with high cardinality categories!\n\n[This starter pack](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook) builds exactly such a model!\n\nI preprocess the data using [polars](https://www.pola.rs/) (mostly to try out the library and to decrease memory footprint vs `pandas`, the added speed is also nice 🙂).\n\nI then proceed to train an `LGBMRanker`, output predictions and submit to the competition.\n\n[The notebook](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker/notebook) provides all the low-level plumbing that you can build on 🙂 The next steps here might include using co-visitation matrices for candidate generation or scoring. One can also generate additional features (such as times an item was clicked or put in cart, popularity scores, temporal features, etc).\n\nHappy hacking! 🙂 \n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2032531": "Thanks for the starter Radek. In the last Kaggle recsys comp, ML reranker beat heuristic reranker. So i think the winning teams will be using an approach like this.",
    "2032933": "Thank you, @cdeotte! 🙌 That is a very valuable piece of information 🙂"
  },
  "source": "meta"
}