{
  "id": 368384,
  "title": "How to train a Word2Vec model 🚀🚀🚀",
  "url": "/competitions/otto-recommender-system/discussion/368384",
  "author_name": "",
  "post_date": "2022-11-25T01:33:35.731601400Z",
  "votes": 27,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>I put together a new notebook on training and submitting using Word2Vec 🙂</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></p>\n<p>First of all, I am super surprised how well it performed! With the simplest of logic, without any tweaking, it scores <code>0.521</code> on the LB!</p>\n<p>That is a strong indication it is worth adding this to your solution. For instance, if you are training a ranking model (similarly to how I describe here: 📑 <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278\" target=\"_blank\">[Step-by-step guide] How I got to my current standing on the LB and how to improve going forward)</a>, you probably should absolutely consider adding a <code>Word2Vec</code> model to the mix of candidate generation!</p>\n<p>A couple of other fun things to try:</p>\n<ul>\n<li>generate candidates like in the co-visitation matrix but replace using the dictionaries and counters with <code>Word2Vec</code> similarity score (I have a feeling that would be super powerful if one can get it to run adequately fast 🙂)</li>\n<li>generate candidates but based on the mean of embeddings of the <code>aids</code> in a session (or say, last 3 embeddings) vs just the past one as I am doing in the NB (recency is a super powerful force when it come to recommendations!)</li>\n<li>create some \"similarity\" score based on an embedding for a given candidate aid and some representation of the session, say mean of the embeddings in that session</li>\n</ul>\n<p>So many ideas and things to try 🙂 This competition is amazing in the richness it exposes, how many things one can do, on such seemingly simple data 🥰</p>\n<p>I can't wait to read competition recaps on all the amazing stuff people will be able to build in this competition! 🙂</p>\n<p>Thanks for reading and happy Kaggling! 🥳</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2042761",
      "postDate": "11/25/2022 01:33:35",
      "content": "<p>Hey!</p>\n<p>I put together a new notebook on training and submitting using Word2Vec 🙂</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></p>\n<p>First of all, I am super surprised how well it performed! With the simplest of logic, without any tweaking, it scores <code>0.521</code> on the LB!</p>\n<p>That is a strong indication it is worth adding this to your solution. For instance, if you are training a ranking model (similarly to how I describe here: 📑 <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278\" target=\"_blank\">[Step-by-step guide] How I got to my current standing on the LB and how to improve going forward)</a>, you probably should absolutely consider adding a <code>Word2Vec</code> model to the mix of candidate generation!</p>\n<p>A couple of other fun things to try:</p>\n<ul>\n<li>generate candidates like in the co-visitation matrix but replace using the dictionaries and counters with <code>Word2Vec</code> similarity score (I have a feeling that would be super powerful if one can get it to run adequately fast 🙂)</li>\n<li>generate candidates but based on the mean of embeddings of the <code>aids</code> in a session (or say, last 3 embeddings) vs just the past one as I am doing in the NB (recency is a super powerful force when it come to recommendations!)</li>\n<li>create some \"similarity\" score based on an embedding for a given candidate aid and some representation of the session, say mean of the embeddings in that session</li>\n</ul>\n<p>So many ideas and things to try 🙂 This competition is amazing in the richness it exposes, how many things one can do, on such seemingly simple data 🥰</p>\n<p>I can't wait to read competition recaps on all the amazing stuff people will be able to build in this competition! 🙂</p>\n<p>Thanks for reading and happy Kaggling! 🥳</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nI put together a new notebook on training and submitting using Word2Vec 🙂\n\n[💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\nFirst of all, I am super surprised how well it performed! With the simplest of logic, without any tweaking, it scores `0.521` on the LB!\n\nThat is a strong indication it is worth adding this to your solution. For instance, if you are training a ranking model (similarly to how I describe here: 📑 [[Step-by-step guide] How I got to my current standing on the LB and how to improve going forward)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278), you probably should absolutely consider adding a `Word2Vec` model to the mix of candidate generation!\n\nA couple of other fun things to try:\n* generate candidates like in the co-visitation matrix but replace using the dictionaries and counters with `Word2Vec` similarity score (I have a feeling that would be super powerful if one can get it to run adequately fast 🙂)\n* generate candidates but based on the mean of embeddings of the `aids` in a session (or say, last 3 embeddings) vs just the past one as I am doing in the NB (recency is a super powerful force when it come to recommendations!)\n* create some \"similarity\" score based on an embedding for a given candidate aid and some representation of the session, say mean of the embeddings in that session\n\nSo many ideas and things to try 🙂 This competition is amazing in the richness it exposes, how many things one can do, on such seemingly simple data 🥰\n\nI can't wait to read competition recaps on all the amazing stuff people will be able to build in this competition! 🙂\n\nThanks for reading and happy Kaggling! 🥳\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2043910",
      "postDate": "11/26/2022 04:35:24",
      "content": "<p>thx for sharing.</p>",
      "rawMarkdown": "thx for sharing.",
      "votes": null
    },
    {
      "id": "2043918",
      "postDate": "11/26/2022 04:44:58",
      "content": "<p>my pleasure, <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a>!  Very glad you are finding this useful!!! 🙌🙂</p>",
      "rawMarkdown": "my pleasure, @dragonzhang!  Very glad you are finding this useful!!! 🙌🙂",
      "votes": null
    },
    {
      "id": "2044038",
      "postDate": "11/26/2022 07:10:21",
      "content": "<p>It is always good to learn for me.</p>",
      "rawMarkdown": "It is always good to learn for me.",
      "votes": null
    },
    {
      "id": "2044041",
      "postDate": "11/26/2022 07:14:56",
      "content": "<p>Absolutely!!!</p>",
      "rawMarkdown": "Absolutely!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2043910,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "11/26/2022 04:35:24",
      "content": "<p>thx for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2043918,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/26/2022 04:44:58",
          "content": "<p>my pleasure, <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a>!  Very glad you are finding this useful!!! 🙌🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2044038,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "11/26/2022 07:10:21",
          "content": "<p>It is always good to learn for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2044041,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/26/2022 07:14:56",
          "content": "<p>Absolutely!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2042761": "Hey!\n\nI put together a new notebook on training and submitting using Word2Vec 🙂\n\n[💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\nFirst of all, I am super surprised how well it performed! With the simplest of logic, without any tweaking, it scores `0.521` on the LB!\n\nThat is a strong indication it is worth adding this to your solution. For instance, if you are training a ranking model (similarly to how I describe here: 📑 [[Step-by-step guide] How I got to my current standing on the LB and how to improve going forward)](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368278), you probably should absolutely consider adding a `Word2Vec` model to the mix of candidate generation!\n\nA couple of other fun things to try:\n* generate candidates like in the co-visitation matrix but replace using the dictionaries and counters with `Word2Vec` similarity score (I have a feeling that would be super powerful if one can get it to run adequately fast 🙂)\n* generate candidates but based on the mean of embeddings of the `aids` in a session (or say, last 3 embeddings) vs just the past one as I am doing in the NB (recency is a super powerful force when it come to recommendations!)\n* create some \"similarity\" score based on an embedding for a given candidate aid and some representation of the session, say mean of the embeddings in that session\n\nSo many ideas and things to try 🙂 This competition is amazing in the richness it exposes, how many things one can do, on such seemingly simple data 🥰\n\nI can't wait to read competition recaps on all the amazing stuff people will be able to build in this competition! 🙂\n\nThanks for reading and happy Kaggling! 🥳\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2043910": "thx for sharing.",
    "2043918": "my pleasure, @dragonzhang!  Very glad you are finding this useful!!! 🙌🙂",
    "2044038": "It is always good to learn for me.",
    "2044041": "Absolutely!!!"
  },
  "source": "meta"
}