{
  "id": 369464,
  "title": "💡 How to add generated candidates to your train data",
  "url": "/competitions/otto-recommender-system/discussion/369464",
  "author_name": "",
  "post_date": "2022-11-30T08:03:06.712907Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>There are a bunch of ways we can generate candidates for training our ranking models:</p>\n<ul>\n<li>popularity based features (most popular yesterday, most popular ever, most popular this week, etc)</li>\n<li>similarity-based (most similar to last aid, to mean embedding of all aids in a session, etc)</li>\n</ul>\n<p>For the last approach <code>word2vec</code> can be very powerful 🙂 But I know quite a few of us have been struggling with the technicalities of adding candidates to our train data.</p>\n<p>I have also struggled with this myself. First of all, it is not very clear how to approach this conceptually. Secondly, it becomes pretty challenging from a technical perspective in the context of limited RAM. I don't think one can pull this off with <code>pandas</code> on Kaggle. But locally the situation doesn't look much better - merges are extremely RAM hungry.</p>\n<p>I devised a solution using <code>polars</code> and would like to share it with you. I added a section to my <code>word2vec</code> notebook where I explain how all this can be done.</p>\n<p>👉 <a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F90a2d84d6e833f3884b0f56278810dde%2Fadding_candidates.png?generation=1669792938876085&amp;alt=media\" alt=\"\"></p>\n<p>I know a bunch of people asked me about this so I hope this can be useful 🙏</p>\n<p>If you have any questions please let me know! Thanks! 🙌 </p>",
  "messages": [
    {
      "id": "2049560",
      "postDate": "11/30/2022 08:03:06",
      "content": "<p>Hey!</p>\n<p>There are a bunch of ways we can generate candidates for training our ranking models:</p>\n<ul>\n<li>popularity based features (most popular yesterday, most popular ever, most popular this week, etc)</li>\n<li>similarity-based (most similar to last aid, to mean embedding of all aids in a session, etc)</li>\n</ul>\n<p>For the last approach <code>word2vec</code> can be very powerful 🙂 But I know quite a few of us have been struggling with the technicalities of adding candidates to our train data.</p>\n<p>I have also struggled with this myself. First of all, it is not very clear how to approach this conceptually. Secondly, it becomes pretty challenging from a technical perspective in the context of limited RAM. I don't think one can pull this off with <code>pandas</code> on Kaggle. But locally the situation doesn't look much better - merges are extremely RAM hungry.</p>\n<p>I devised a solution using <code>polars</code> and would like to share it with you. I added a section to my <code>word2vec</code> notebook where I explain how all this can be done.</p>\n<p>👉 <a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F90a2d84d6e833f3884b0f56278810dde%2Fadding_candidates.png?generation=1669792938876085&amp;alt=media\" alt=\"\"></p>\n<p>I know a bunch of people asked me about this so I hope this can be useful 🙏</p>\n<p>If you have any questions please let me know! Thanks! 🙌 </p>",
      "rawMarkdown": "Hey!\n\nThere are a bunch of ways we can generate candidates for training our ranking models:\n* popularity based features (most popular yesterday, most popular ever, most popular this week, etc)\n* similarity-based (most similar to last aid, to mean embedding of all aids in a session, etc)\n\nFor the last approach `word2vec` can be very powerful 🙂 But I know quite a few of us have been struggling with the technicalities of adding candidates to our train data.\n\nI have also struggled with this myself. First of all, it is not very clear how to approach this conceptually. Secondly, it becomes pretty challenging from a technical perspective in the context of limited RAM. I don't think one can pull this off with `pandas` on Kaggle. But locally the situation doesn't look much better - merges are extremely RAM hungry.\n\nI devised a solution using `polars` and would like to share it with you. I added a section to my `word2vec` notebook where I explain how all this can be done.\n\n👉 [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F90a2d84d6e833f3884b0f56278810dde%2Fadding_candidates.png?generation=1669792938876085&alt=media)\n\nI know a bunch of people asked me about this so I hope this can be useful 🙏\n\nIf you have any questions please let me know! Thanks! 🙌",
      "votes": null
    },
    {
      "id": "2052894",
      "postDate": "12/02/2022 16:11:29",
      "content": "<p>tnx radek!</p>",
      "rawMarkdown": "tnx radek!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2052894,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "12/02/2022 16:11:29",
      "content": "<p>tnx radek!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2049560": "Hey!\n\nThere are a bunch of ways we can generate candidates for training our ranking models:\n* popularity based features (most popular yesterday, most popular ever, most popular this week, etc)\n* similarity-based (most similar to last aid, to mean embedding of all aids in a session, etc)\n\nFor the last approach `word2vec` can be very powerful 🙂 But I know quite a few of us have been struggling with the technicalities of adding candidates to our train data.\n\nI have also struggled with this myself. First of all, it is not very clear how to approach this conceptually. Secondly, it becomes pretty challenging from a technical perspective in the context of limited RAM. I don't think one can pull this off with `pandas` on Kaggle. But locally the situation doesn't look much better - merges are extremely RAM hungry.\n\nI devised a solution using `polars` and would like to share it with you. I added a section to my `word2vec` notebook where I explain how all this can be done.\n\n👉 [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F90a2d84d6e833f3884b0f56278810dde%2Fadding_candidates.png?generation=1669792938876085&alt=media)\n\nI know a bunch of people asked me about this so I hope this can be useful 🙏\n\nIf you have any questions please let me know! Thanks! 🙌",
    "2052894": "tnx radek!"
  },
  "source": "meta"
}