{
  "id": 367234,
  "title": "How to train a Word2Vec model for item embeddings - a simple code example 📖",
  "url": "/competitions/otto-recommender-system/discussion/367234",
  "author_name": "",
  "post_date": "2022-11-19T17:46:24.771030400Z",
  "votes": 43,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>Here is a small code example of how to train a Word2Vec model from the given user sessions. Since each session is a sequence of clicks, carts, or orders, we can treat it as a sequence of words/tokens like this text. </p>\n<pre><code> pandas  pd\n gensim.models  Word2Vec\n\n\ncs = \ncount = \ntotal = \n (, )  f:\n     df  pd.read_json(, lines=, chunksize=cs):\n         val  df[].apply( x: [event[]  event  x]).values:\n            f.write(.join((, val)) + )\n        count += cs\n        (, end=)\n\n\n\n\nmodel = Word2Vec(corpus_file=, vector_size=, window=, min_count=, workers=)\nmodel.save()\n</code></pre>\n<p>The full list of parameters for the model can be found here. <br>\n<a href=\"https://radimrehurek.com/gensim/models/word2vec.html\" target=\"_blank\">https://radimrehurek.com/gensim/models/word2vec.html</a></p>\n<p>And after training the model we can query for similar items like this:</p>\n<pre><code>model = Word2Vec.load()\n\nmodel.wv.most_similar(, topn=)\n</code></pre>\n<p>But this way generating similar items for the whole item set is a bit painful. To increase the speed we can use <a href=\"https://github.com/spotify/annoy\" target=\"_blank\">annoy indexer</a>.</p>\n<p><a href=\"https://radimrehurek.com/gensim/similarities/annoy.html\" target=\"_blank\">https://radimrehurek.com/gensim/similarities/annoy.html</a></p>\n<pre><code> gensim.similarities.annoy  AnnoyIndexer\n\nannoy_index = AnnoyIndexer(model, )\nmodel.wv.most_similar(, topn=, indexer=annoy_index)\n</code></pre>",
  "messages": [
    {
      "id": "2036416",
      "postDate": "11/19/2022 17:46:24",
      "content": "<p>Hey everyone,</p>\n<p>Here is a small code example of how to train a Word2Vec model from the given user sessions. Since each session is a sequence of clicks, carts, or orders, we can treat it as a sequence of words/tokens like this text. </p>\n<pre><code> pandas  pd\n gensim.models  Word2Vec\n\n\ncs = \ncount = \ntotal = \n (, )  f:\n     df  pd.read_json(, lines=, chunksize=cs):\n         val  df[].apply( x: [event[]  event  x]).values:\n            f.write(.join((, val)) + )\n        count += cs\n        (, end=)\n\n\n\n\nmodel = Word2Vec(corpus_file=, vector_size=, window=, min_count=, workers=)\nmodel.save()\n</code></pre>\n<p>The full list of parameters for the model can be found here. <br>\n<a href=\"https://radimrehurek.com/gensim/models/word2vec.html\" target=\"_blank\">https://radimrehurek.com/gensim/models/word2vec.html</a></p>\n<p>And after training the model we can query for similar items like this:</p>\n<pre><code>model = Word2Vec.load()\n\nmodel.wv.most_similar(, topn=)\n</code></pre>\n<p>But this way generating similar items for the whole item set is a bit painful. To increase the speed we can use <a href=\"https://github.com/spotify/annoy\" target=\"_blank\">annoy indexer</a>.</p>\n<p><a href=\"https://radimrehurek.com/gensim/similarities/annoy.html\" target=\"_blank\">https://radimrehurek.com/gensim/similarities/annoy.html</a></p>\n<pre><code> gensim.similarities.annoy  AnnoyIndexer\n\nannoy_index = AnnoyIndexer(model, )\nmodel.wv.most_similar(, topn=, indexer=annoy_index)\n</code></pre>",
      "rawMarkdown": "Hey everyone,\n\nHere is a small code example of how to train a Word2Vec model from the given user sessions. Since each session is a sequence of clicks, carts, or orders, we can treat it as a sequence of words/tokens like this text. \n\n ```python\nimport pandas as pd\nfrom gensim.models import Word2Vec\n\n# We write the sessions to the text file to make things faster for training the Word2Vec model (could be optimized further).\ncs = 100000\ncount = 0\ntotal = 12899779\nwith open(\"w2v_input.txt\", \"w\") as f:\n    for df in pd.read_json(\"../input/otto-recommender-system/train.jsonl\", lines=True, chunksize=cs):\n        for val in df[\"events\"].apply(lambda x: [event[\"aid\"] for event in x]).values:\n            f.write(\" \".join(map(str, val)) + \"\\n\")\n        count += cs\n        print(f\"{count}/{total}\\r\", end=\"\")\n\n# The lines in the input file look like this at the end.\n\"\"\"\n...\n424964 1492293 1492293 910862 910862 1491172 1491172 424964 1515526 440486 109488 1507622 1734061 854637 854637 718983 215311 215311 718983 711125 711125 50049 105393 105393 959544 1734061 1842593 1464360 207905 1628317 376932 497868 \n...\n\"\"\"\n# Training the word2vec model.\nmodel = Word2Vec(corpus_file=\"w2v_input.txt\", vector_size=50, window=5, min_count=1, workers=4)\nmodel.save(\"word2vec.model\")\n```\nThe full list of parameters for the model can be found here. \nhttps://radimrehurek.com/gensim/models/word2vec.html\n\nAnd after training the model we can query for similar items like this:\n\n```python\nmodel = Word2Vec.load(\"word2vec.model\")\n# It's important to note that item ids (aid) are string instead of int. \nmodel.wv.most_similar(\"1460571\", topn=20)\n```\nBut this way generating similar items for the whole item set is a bit painful. To increase the speed we can use [annoy indexer](https://github.com/spotify/annoy).\n\nhttps://radimrehurek.com/gensim/similarities/annoy.html\n```python\nfrom gensim.similarities.annoy import AnnoyIndexer\n\nannoy_index = AnnoyIndexer(model, 300)\nmodel.wv.most_similar(\"1460571\", topn=20, indexer=annoy_index)\n```",
      "votes": null
    },
    {
      "id": "2037532",
      "postDate": "11/20/2022 18:04:28",
      "content": "<p>Hey, I came across the Word2Vec model for session based systems from this website: <a href=\"https://session-based-recommenders.fastforwardlabs.com/\" target=\"_blank\">https://session-based-recommenders.fastforwardlabs.com/</a></p>\n<p>and I implemented it for the train set.. it was incredibly slow! Thanks for putting this post out, I will try these changes in my notebook.</p>",
      "rawMarkdown": "Hey, I came across the Word2Vec model for session based systems from this website: https://session-based-recommenders.fastforwardlabs.com/\n\nand I implemented it for the train set.. it was incredibly slow! Thanks for putting this post out, I will try these changes in my notebook.",
      "votes": null
    },
    {
      "id": "2037543",
      "postDate": "11/20/2022 18:18:13",
      "content": "<p>Hi, </p>\n<p>This code trains in 45~ minutes in the Kaggle CPU environment. But, I do not remember the inference time (similarity computations) exactly. </p>",
      "rawMarkdown": "Hi, \n\nThis code trains in 45~ minutes in the Kaggle CPU environment. But, I do not remember the inference time (similarity computations) exactly.",
      "votes": null
    },
    {
      "id": "2037596",
      "postDate": "11/20/2022 19:03:10",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> </p>\n<p>Have you evaluated the model performance (recall) and / or LB score?</p>",
      "rawMarkdown": "Thanks @snnclsr \n\nHave you evaluated the model performance (recall) and / or LB score?",
      "votes": null
    },
    {
      "id": "2037627",
      "postDate": "11/20/2022 19:16:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jonimatix\" target=\"_blank\">@jonimatix</a> </p>\n<p>Unfortunately, still no. I am planning to do it in the next week. I will use it only for the candidate generation though.</p>",
      "rawMarkdown": "Hi @jonimatix \n\nUnfortunately, still no. I am planning to do it in the next week. I will use it only for the candidate generation though.",
      "votes": null
    },
    {
      "id": "2037673",
      "postDate": "11/20/2022 20:07:59",
      "content": "<p>Thank, I looked for one person who will be have think, if the classification of product or session is important. I questioned myself, if this person will be developped this solution with word2vec or kmean in td-idf. Thank for you solution !!! </p>",
      "rawMarkdown": "Thank, I looked for one person who will be have think, if the classification of product or session is important. I questioned myself, if this person will be developped this solution with word2vec or kmean in td-idf. Thank for you solution !!!",
      "votes": null
    },
    {
      "id": "2037773",
      "postDate": "11/20/2022 23:34:29",
      "content": "<p>Hey, that was a great read. Thanks for posting it.</p>",
      "rawMarkdown": "Hey, that was a great read. Thanks for posting it.",
      "votes": null
    },
    {
      "id": "2041192",
      "postDate": "11/23/2022 18:21:13",
      "content": "<p>i realized a pattern relate to Event [Aid] , if the task was to predict the next aid corresponding to labels [ click , add , order] , I purpose an idea that instead for  looking at which the Predicted Aid will be next if we treat this problem as Time series Task and check the correlation between labels and Aid given timestep , will be good to use Forecasting sessions and predict sequences of aid in time Axis after apply Fourier-Transformer to convert timestep to Frequency dimension ,</p>",
      "rawMarkdown": "i realized a pattern relate to Event [Aid] , if the task was to predict the next aid corresponding to labels [ click , add , order] , I purpose an idea that instead for  looking at which the Predicted Aid will be next if we treat this problem as Time series Task and check the correlation between labels and Aid given timestep , will be good to use Forecasting sessions and predict sequences of aid in time Axis after apply Fourier-Transformer to convert timestep to Frequency dimension ,",
      "votes": null
    },
    {
      "id": "2053558",
      "postDate": "12/03/2022 11:53:51",
      "content": "<p>Kaggle GPUs are decent but CPUs aren't that good. It takes 2 minutes per epoch on my local machine which has AMD Ryzen 9 5950X CPU.</p>",
      "rawMarkdown": "Kaggle GPUs are decent but CPUs aren't that good. It takes 2 minutes per epoch on my local machine which has AMD Ryzen 9 5950X CPU.",
      "votes": null
    },
    {
      "id": "2069793",
      "postDate": "12/19/2022 10:30:45",
      "content": "<p>thanks a lot</p>",
      "rawMarkdown": "thanks a lot",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2037532,
      "author_name": "rajatrc1705",
      "author_url": "",
      "post_date": "11/20/2022 18:04:28",
      "content": "<p>Hey, I came across the Word2Vec model for session based systems from this website: <a href=\"https://session-based-recommenders.fastforwardlabs.com/\" target=\"_blank\">https://session-based-recommenders.fastforwardlabs.com/</a></p>\n<p>and I implemented it for the train set.. it was incredibly slow! Thanks for putting this post out, I will try these changes in my notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2037543,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "11/20/2022 18:18:13",
          "content": "<p>Hi, </p>\n<p>This code trains in 45~ minutes in the Kaggle CPU environment. But, I do not remember the inference time (similarity computations) exactly. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2037773,
          "author_name": "megan3",
          "author_url": "",
          "post_date": "11/20/2022 23:34:29",
          "content": "<p>Hey, that was a great read. Thanks for posting it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2053558,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/03/2022 11:53:51",
          "content": "<p>Kaggle GPUs are decent but CPUs aren't that good. It takes 2 minutes per epoch on my local machine which has AMD Ryzen 9 5950X CPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2037596,
      "author_name": "jonimatix",
      "author_url": "",
      "post_date": "11/20/2022 19:03:10",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> </p>\n<p>Have you evaluated the model performance (recall) and / or LB score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2037627,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "11/20/2022 19:16:13",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jonimatix\" target=\"_blank\">@jonimatix</a> </p>\n<p>Unfortunately, still no. I am planning to do it in the next week. I will use it only for the candidate generation though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2037673,
      "author_name": "lgregory",
      "author_url": "",
      "post_date": "11/20/2022 20:07:59",
      "content": "<p>Thank, I looked for one person who will be have think, if the classification of product or session is important. I questioned myself, if this person will be developped this solution with word2vec or kmean in td-idf. Thank for you solution !!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2041192,
      "author_name": "younesselbrag",
      "author_url": "",
      "post_date": "11/23/2022 18:21:13",
      "content": "<p>i realized a pattern relate to Event [Aid] , if the task was to predict the next aid corresponding to labels [ click , add , order] , I purpose an idea that instead for  looking at which the Predicted Aid will be next if we treat this problem as Time series Task and check the correlation between labels and Aid given timestep , will be good to use Forecasting sessions and predict sequences of aid in time Axis after apply Fourier-Transformer to convert timestep to Frequency dimension ,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2069793,
      "author_name": "shanggangli",
      "author_url": "",
      "post_date": "12/19/2022 10:30:45",
      "content": "<p>thanks a lot</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2036416": "Hey everyone,\n\nHere is a small code example of how to train a Word2Vec model from the given user sessions. Since each session is a sequence of clicks, carts, or orders, we can treat it as a sequence of words/tokens like this text. \n\n ```python\nimport pandas as pd\nfrom gensim.models import Word2Vec\n\n# We write the sessions to the text file to make things faster for training the Word2Vec model (could be optimized further).\ncs = 100000\ncount = 0\ntotal = 12899779\nwith open(\"w2v_input.txt\", \"w\") as f:\n    for df in pd.read_json(\"../input/otto-recommender-system/train.jsonl\", lines=True, chunksize=cs):\n        for val in df[\"events\"].apply(lambda x: [event[\"aid\"] for event in x]).values:\n            f.write(\" \".join(map(str, val)) + \"\\n\")\n        count += cs\n        print(f\"{count}/{total}\\r\", end=\"\")\n\n# The lines in the input file look like this at the end.\n\"\"\"\n...\n424964 1492293 1492293 910862 910862 1491172 1491172 424964 1515526 440486 109488 1507622 1734061 854637 854637 718983 215311 215311 718983 711125 711125 50049 105393 105393 959544 1734061 1842593 1464360 207905 1628317 376932 497868 \n...\n\"\"\"\n# Training the word2vec model.\nmodel = Word2Vec(corpus_file=\"w2v_input.txt\", vector_size=50, window=5, min_count=1, workers=4)\nmodel.save(\"word2vec.model\")\n```\nThe full list of parameters for the model can be found here. \nhttps://radimrehurek.com/gensim/models/word2vec.html\n\nAnd after training the model we can query for similar items like this:\n\n```python\nmodel = Word2Vec.load(\"word2vec.model\")\n# It's important to note that item ids (aid) are string instead of int. \nmodel.wv.most_similar(\"1460571\", topn=20)\n```\nBut this way generating similar items for the whole item set is a bit painful. To increase the speed we can use [annoy indexer](https://github.com/spotify/annoy).\n\nhttps://radimrehurek.com/gensim/similarities/annoy.html\n```python\nfrom gensim.similarities.annoy import AnnoyIndexer\n\nannoy_index = AnnoyIndexer(model, 300)\nmodel.wv.most_similar(\"1460571\", topn=20, indexer=annoy_index)\n```",
    "2037532": "Hey, I came across the Word2Vec model for session based systems from this website: https://session-based-recommenders.fastforwardlabs.com/\n\nand I implemented it for the train set.. it was incredibly slow! Thanks for putting this post out, I will try these changes in my notebook.",
    "2037543": "Hi, \n\nThis code trains in 45~ minutes in the Kaggle CPU environment. But, I do not remember the inference time (similarity computations) exactly.",
    "2037596": "Thanks @snnclsr \n\nHave you evaluated the model performance (recall) and / or LB score?",
    "2037627": "Hi @jonimatix \n\nUnfortunately, still no. I am planning to do it in the next week. I will use it only for the candidate generation though.",
    "2037673": "Thank, I looked for one person who will be have think, if the classification of product or session is important. I questioned myself, if this person will be developped this solution with word2vec or kmean in td-idf. Thank for you solution !!!",
    "2037773": "Hey, that was a great read. Thanks for posting it.",
    "2041192": "i realized a pattern relate to Event [Aid] , if the task was to predict the next aid corresponding to labels [ click , add , order] , I purpose an idea that instead for  looking at which the Predicted Aid will be next if we treat this problem as Time series Task and check the correlation between labels and Aid given timestep , will be good to use Forecasting sessions and predict sequences of aid in time Axis after apply Fourier-Transformer to convert timestep to Frequency dimension ,",
    "2053558": "Kaggle GPUs are decent but CPUs aren't that good. It takes 2 minutes per epoch on my local machine which has AMD Ryzen 9 5950X CPU.",
    "2069793": "thanks a lot"
  },
  "source": "meta"
}