{
  "id": 370751,
  "title": "I think word2vec is a nice choice.",
  "url": "/competitions/otto-recommender-system/discussion/370751",
  "author_name": "",
  "post_date": "2022-12-06T09:37:47.647723800Z",
  "votes": 40,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I made a word2vec-like algorithm that generates candidates.</p>\n<p>It feels so good：D</p>\n<p>onle use word2vec-like algorithm generates candidates</p>\n<ul>\n<li>lb 0.578</li>\n</ul>",
  "messages": [
    {
      "id": "2056611",
      "postDate": "12/06/2022 09:37:47",
      "content": "<p>I made a word2vec-like algorithm that generates candidates.</p>\n<p>It feels so good：D</p>\n<p>onle use word2vec-like algorithm generates candidates</p>\n<ul>\n<li>lb 0.578</li>\n</ul>",
      "rawMarkdown": "I made a word2vec-like algorithm that generates candidates.\n\nIt feels so good：D\n\nonle use word2vec-like algorithm generates candidates\n- lb 0.578",
      "votes": null
    },
    {
      "id": "2056697",
      "postDate": "12/06/2022 11:19:39",
      "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> Do you use Nearest Neighbour to decide 20 predictions ?</p>",
      "rawMarkdown": "takusid Do you use Nearest Neighbour to decide 20 predictions ?",
      "votes": null
    },
    {
      "id": "2056815",
      "postDate": "12/06/2022 13:12:32",
      "content": "<p><a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a><br>\nNo, I do not use the K neighborhood algorithm.<br>\nI was inspired by word2vec's window size</p>",
      "rawMarkdown": "toshik\nNo, I do not use the K neighborhood algorithm.\nI was inspired by word2vec's window size",
      "votes": null
    },
    {
      "id": "2056881",
      "postDate": "12/06/2022 14:15:37",
      "content": "<p>Nice job. If you generate more than 20 candidates, then add a ranker model <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">here</a>, then select 20 you should get a nice CV LB boost.</p>",
      "rawMarkdown": "Nice job. If you generate more than 20 candidates, then add a ranker model [here][1], then select 20 you should get a nice CV LB boost.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210",
      "votes": null
    },
    {
      "id": "2056897",
      "postDate": "12/06/2022 14:44:20",
      "content": "<p>I also tried train word2vec and then selecting most similar items for users session,  but did not get a good score,  your result is wonderful,  thanks for sharing!</p>",
      "rawMarkdown": "I also tried train word2vec and then selecting most similar items for users session,  but did not get a good score,  your result is wonderful,  thanks for sharing!",
      "votes": null
    },
    {
      "id": "2057793",
      "postDate": "12/07/2022 11:01:27",
      "content": "<p>I think window size is keypoint</p>",
      "rawMarkdown": "I think window size is keypoint",
      "votes": null
    },
    {
      "id": "2057800",
      "postDate": "12/07/2022 11:03:09",
      "content": "<p>thank u for share , I will do my best to catch up with you.</p>",
      "rawMarkdown": "thank u for share , I will do my best to catch up with you.",
      "votes": null
    },
    {
      "id": "2058625",
      "postDate": "12/08/2022 04:54:53",
      "content": "<p>Thanks for your tips and sharing！</p>",
      "rawMarkdown": "Thanks for your tips and sharing！",
      "votes": null
    },
    {
      "id": "2058828",
      "postDate": "12/08/2022 08:42:43",
      "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> that sounds interesting. Could you point to a reference which explains more on the point you made with window size? I am not Familar with word2vec and would like to try out sth like your approach to generate canditates!</p>",
      "rawMarkdown": "takusid that sounds interesting. Could you point to a reference which explains more on the point you made with window size? I am not Familar with word2vec and would like to try out sth like your approach to generate canditates!",
      "votes": null
    },
    {
      "id": "2058834",
      "postDate": "12/08/2022 08:50:03",
      "content": "<p>This is a great explanation of Word2Vec <a href=\"https://jalammar.github.io/illustrated-word2vec/\" target=\"_blank\">https://jalammar.github.io/illustrated-word2vec/</a> . Just think about sessions and aids as sentences and words.</p>",
      "rawMarkdown": "This is a great explanation of Word2Vec https://jalammar.github.io/illustrated-word2vec/ . Just think about sessions and aids as sentences and words.",
      "votes": null
    },
    {
      "id": "2059205",
      "postDate": "12/08/2022 15:16:12",
      "content": "<p>I've noticed that two of the word2vec solutions I've encountered here so far are using nearest neighbor in the embeddings to then recommend the item that looks closest to the last item.</p>\n<p>Is there a reason people aren't instead using predict_output_word (from word2vec) itself to get the proposed next item in context/order?</p>",
      "rawMarkdown": "I've noticed that two of the word2vec solutions I've encountered here so far are using nearest neighbor in the embeddings to then recommend the item that looks closest to the last item.\n\nIs there a reason people aren't instead using predict_output_word (from word2vec) itself to get the proposed next item in context/order?",
      "votes": null
    },
    {
      "id": "2059560",
      "postDate": "12/09/2022 02:51:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> </p>\n<p>How long did it take to learn word2vec model?</p>",
      "rawMarkdown": "Hi @takusid \n\nHow long did it take to learn word2vec model?",
      "votes": null
    },
    {
      "id": "2059748",
      "postDate": "12/09/2022 07:27:23",
      "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> <br>\nWhat is the theoretical maximum score when the selected candidate is perfectly reranked?</p>",
      "rawMarkdown": "takusid \nWhat is the theoretical maximum score when the selected candidate is perfectly reranked?",
      "votes": null
    },
    {
      "id": "2060683",
      "postDate": "12/10/2022 09:49:37",
      "content": "<p>recall@200<br>\nclicks recall = 0.695143472014783    <br>\ncarts recall = 0.5445477916049417    <br>\norders recall = 0.7274427630759999    </p>\n<p>Overall Recall = 0.6693443425285608    </p>",
      "rawMarkdown": "recall@200\nclicks recall = 0.695143472014783    \ncarts recall = 0.5445477916049417    \norders recall = 0.7274427630759999    \n\nOverall Recall = 0.6693443425285608",
      "votes": null
    },
    {
      "id": "2065190",
      "postDate": "12/14/2022 11:37:07",
      "content": "<p>Thanks for sharing!<br>\nOn second thought, I couldn't simply compare the results because it depends on seed when dividing test and ground_truth.</p>",
      "rawMarkdown": "Thanks for sharing!\nOn second thought, I couldn't simply compare the results because it depends on seed when dividing test and ground_truth.",
      "votes": null
    },
    {
      "id": "2071629",
      "postDate": "12/21/2022 08:24:49",
      "content": "<p>I tried FastText unsupervised training and it is meh. Maybe it could be useful for candidate generation.</p>\n<p>I'm using annoy for nn search and I recommend nearest 20 neighbors of the last session aid. My validation scores: <br>\n<code>Scores: Clicks 0.4782 - Carts: 0.4950 - Orders: 0.6586 - Weighted: 0.5620</code> and my lb score is <code>0.547</code>. I also tried recommending nearest neighbor recursively but it didn't work.</p>",
      "rawMarkdown": "I tried FastText unsupervised training and it is meh. Maybe it could be useful for candidate generation.\n\nI'm using annoy for nn search and I recommend nearest 20 neighbors of the last session aid. My validation scores: \n`Scores: Clicks 0.4782 - Carts: 0.4950 - Orders: 0.6586 - Weighted: 0.5620` and my lb score is `0.547`. I also tried recommending nearest neighbor recursively but it didn't work.",
      "votes": null
    },
    {
      "id": "2071770",
      "postDate": "12/21/2022 11:30:06",
      "content": "<p>I tried that. It takes 14 hours to predict my validation set.</p>",
      "rawMarkdown": "I tried that. It takes 14 hours to predict my validation set.",
      "votes": null
    },
    {
      "id": "2082779",
      "postDate": "01/01/2023 23:19:29",
      "content": "<p>Any clue about it ? :P</p>",
      "rawMarkdown": "Any clue about it ? :P",
      "votes": null
    },
    {
      "id": "2124451",
      "postDate": "02/01/2023 02:29:38",
      "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a>, will you share your word2vec solution in a bit detail? Quite interesting on it. Thanks.</p>",
      "rawMarkdown": "takusid, will you share your word2vec solution in a bit detail? Quite interesting on it. Thanks.",
      "votes": null
    },
    {
      "id": "2124526",
      "postDate": "02/01/2023 04:08:02",
      "content": "<p>I was able to reach 0.582 with a combination of covisitation and fasttext skipgram model. I think that's what he did.</p>",
      "rawMarkdown": "I was able to reach 0.582 with a combination of covisitation and fasttext skipgram model. I think that's what he did.",
      "votes": null
    },
    {
      "id": "2124985",
      "postDate": "02/01/2023 10:55:14",
      "content": "<p>yes, I just set the skipGram parameter to 2.<br>\nand train 2 epoch</p>",
      "rawMarkdown": "yes, I just set the skipGram parameter to 2.\nand train 2 epoch",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2056697,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "12/06/2022 11:19:39",
      "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> Do you use Nearest Neighbour to decide 20 predictions ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2056815,
      "author_name": "takusid",
      "author_url": "",
      "post_date": "12/06/2022 13:12:32",
      "content": "<p><a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a><br>\nNo, I do not use the K neighborhood algorithm.<br>\nI was inspired by word2vec's window size</p>",
      "votes": null,
      "replies": [
        {
          "id": 2082779,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/01/2023 23:19:29",
          "content": "<p>Any clue about it ? :P</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2056881,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/06/2022 14:15:37",
      "content": "<p>Nice job. If you generate more than 20 candidates, then add a ranker model <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">here</a>, then select 20 you should get a nice CV LB boost.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2057800,
          "author_name": "takusid",
          "author_url": "",
          "post_date": "12/07/2022 11:03:09",
          "content": "<p>thank u for share , I will do my best to catch up with you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2056897,
      "author_name": "evilpsycho42",
      "author_url": "",
      "post_date": "12/06/2022 14:44:20",
      "content": "<p>I also tried train word2vec and then selecting most similar items for users session,  but did not get a good score,  your result is wonderful,  thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2057793,
          "author_name": "takusid",
          "author_url": "",
          "post_date": "12/07/2022 11:01:27",
          "content": "<p>I think window size is keypoint</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2058625,
          "author_name": "evilpsycho42",
          "author_url": "",
          "post_date": "12/08/2022 04:54:53",
          "content": "<p>Thanks for your tips and sharing！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2058828,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "12/08/2022 08:42:43",
          "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> that sounds interesting. Could you point to a reference which explains more on the point you made with window size? I am not Familar with word2vec and would like to try out sth like your approach to generate canditates!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2058834,
          "author_name": "piotrekga",
          "author_url": "",
          "post_date": "12/08/2022 08:50:03",
          "content": "<p>This is a great explanation of Word2Vec <a href=\"https://jalammar.github.io/illustrated-word2vec/\" target=\"_blank\">https://jalammar.github.io/illustrated-word2vec/</a> . Just think about sessions and aids as sentences and words.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2059205,
      "author_name": "adamcohon",
      "author_url": "",
      "post_date": "12/08/2022 15:16:12",
      "content": "<p>I've noticed that two of the word2vec solutions I've encountered here so far are using nearest neighbor in the embeddings to then recommend the item that looks closest to the last item.</p>\n<p>Is there a reason people aren't instead using predict_output_word (from word2vec) itself to get the proposed next item in context/order?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2071770,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/21/2022 11:30:06",
          "content": "<p>I tried that. It takes 14 hours to predict my validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2059560,
      "author_name": "dehokanta",
      "author_url": "",
      "post_date": "12/09/2022 02:51:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> </p>\n<p>How long did it take to learn word2vec model?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2059748,
      "author_name": "shkanda",
      "author_url": "",
      "post_date": "12/09/2022 07:27:23",
      "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a> <br>\nWhat is the theoretical maximum score when the selected candidate is perfectly reranked?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2060683,
          "author_name": "takusid",
          "author_url": "",
          "post_date": "12/10/2022 09:49:37",
          "content": "<p>recall@200<br>\nclicks recall = 0.695143472014783    <br>\ncarts recall = 0.5445477916049417    <br>\norders recall = 0.7274427630759999    </p>\n<p>Overall Recall = 0.6693443425285608    </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2065190,
          "author_name": "shkanda",
          "author_url": "",
          "post_date": "12/14/2022 11:37:07",
          "content": "<p>Thanks for sharing!<br>\nOn second thought, I couldn't simply compare the results because it depends on seed when dividing test and ground_truth.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2071629,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "12/21/2022 08:24:49",
      "content": "<p>I tried FastText unsupervised training and it is meh. Maybe it could be useful for candidate generation.</p>\n<p>I'm using annoy for nn search and I recommend nearest 20 neighbors of the last session aid. My validation scores: <br>\n<code>Scores: Clicks 0.4782 - Carts: 0.4950 - Orders: 0.6586 - Weighted: 0.5620</code> and my lb score is <code>0.547</code>. I also tried recommending nearest neighbor recursively but it didn't work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124451,
      "author_name": "gongbi",
      "author_url": "",
      "post_date": "02/01/2023 02:29:38",
      "content": "<p><a href=\"https://www.kaggle.com/takusid\" target=\"_blank\">@takusid</a>, will you share your word2vec solution in a bit detail? Quite interesting on it. Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124526,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "02/01/2023 04:08:02",
          "content": "<p>I was able to reach 0.582 with a combination of covisitation and fasttext skipgram model. I think that's what he did.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2124985,
              "author_name": "takusid",
              "author_url": "",
              "post_date": "02/01/2023 10:55:14",
              "content": "<p>yes, I just set the skipGram parameter to 2.<br>\nand train 2 epoch</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2056611": "I made a word2vec-like algorithm that generates candidates.\n\nIt feels so good：D\n\nonle use word2vec-like algorithm generates candidates\n- lb 0.578",
    "2056697": "takusid Do you use Nearest Neighbour to decide 20 predictions ?",
    "2056815": "toshik\nNo, I do not use the K neighborhood algorithm.\nI was inspired by word2vec's window size",
    "2056881": "Nice job. If you generate more than 20 candidates, then add a ranker model [here][1], then select 20 you should get a nice CV LB boost.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210",
    "2056897": "I also tried train word2vec and then selecting most similar items for users session,  but did not get a good score,  your result is wonderful,  thanks for sharing!",
    "2057793": "I think window size is keypoint",
    "2057800": "thank u for share , I will do my best to catch up with you.",
    "2058625": "Thanks for your tips and sharing！",
    "2058828": "takusid that sounds interesting. Could you point to a reference which explains more on the point you made with window size? I am not Familar with word2vec and would like to try out sth like your approach to generate canditates!",
    "2058834": "This is a great explanation of Word2Vec https://jalammar.github.io/illustrated-word2vec/ . Just think about sessions and aids as sentences and words.",
    "2059205": "I've noticed that two of the word2vec solutions I've encountered here so far are using nearest neighbor in the embeddings to then recommend the item that looks closest to the last item.\n\nIs there a reason people aren't instead using predict_output_word (from word2vec) itself to get the proposed next item in context/order?",
    "2059560": "Hi @takusid \n\nHow long did it take to learn word2vec model?",
    "2059748": "takusid \nWhat is the theoretical maximum score when the selected candidate is perfectly reranked?",
    "2060683": "recall@200\nclicks recall = 0.695143472014783    \ncarts recall = 0.5445477916049417    \norders recall = 0.7274427630759999    \n\nOverall Recall = 0.6693443425285608",
    "2065190": "Thanks for sharing!\nOn second thought, I couldn't simply compare the results because it depends on seed when dividing test and ground_truth.",
    "2071629": "I tried FastText unsupervised training and it is meh. Maybe it could be useful for candidate generation.\n\nI'm using annoy for nn search and I recommend nearest 20 neighbors of the last session aid. My validation scores: \n`Scores: Clicks 0.4782 - Carts: 0.4950 - Orders: 0.6586 - Weighted: 0.5620` and my lb score is `0.547`. I also tried recommending nearest neighbor recursively but it didn't work.",
    "2071770": "I tried that. It takes 14 hours to predict my validation set.",
    "2082779": "Any clue about it ? :P",
    "2124451": "takusid, will you share your word2vec solution in a bit detail? Quite interesting on it. Thanks.",
    "2124526": "I was able to reach 0.582 with a combination of covisitation and fasttext skipgram model. I think that's what he did.",
    "2124985": "yes, I just set the skipGram parameter to 2.\nand train 2 epoch"
  },
  "source": "meta"
}