{
  "id": 374677,
  "title": "Some experimental approaches",
  "url": "/competitions/otto-recommender-system/discussion/374677",
  "author_name": "",
  "post_date": "2022-12-28T09:59:37.238226100Z",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This was my first recommender system project so I wanted to start from basics and move on. These are the things I tried but didn't work for me. Maybe you can make them work.</p>\n<ul>\n<li>Collaborative filtering: Either the scores aren't good or the embeddings become too large and inference is too slow</li>\n<li>Matrix factorization: Same as collaborative filtering</li>\n<li>Word2vec: It is pretty fast but not competitive</li>\n<li>FastText: Slightly better and faster training compared to word2vec from gensim because of c++ bindings but the models are same</li>\n<li>Doc2vec: Model doesn't learn anything</li>\n<li>Sequential models from recbole library (GRU4Rec, BERT4Rec and etc.): Training is too slow and they are not competitive</li>\n<li>General models from recbole library (BPR, CF and MF models): Training is fast but inference is too slow because models don't scale</li>\n<li>Models from surprise library: Same as general models from recbole library</li>\n<li>TF-IDF + pairwise similarity: Very slow inference time since I was using argsort to get top 20 most similar aids</li>\n</ul>",
  "messages": [
    {
      "id": "2078404",
      "postDate": "12/28/2022 09:59:37",
      "content": "<p>This was my first recommender system project so I wanted to start from basics and move on. These are the things I tried but didn't work for me. Maybe you can make them work.</p>\n<ul>\n<li>Collaborative filtering: Either the scores aren't good or the embeddings become too large and inference is too slow</li>\n<li>Matrix factorization: Same as collaborative filtering</li>\n<li>Word2vec: It is pretty fast but not competitive</li>\n<li>FastText: Slightly better and faster training compared to word2vec from gensim because of c++ bindings but the models are same</li>\n<li>Doc2vec: Model doesn't learn anything</li>\n<li>Sequential models from recbole library (GRU4Rec, BERT4Rec and etc.): Training is too slow and they are not competitive</li>\n<li>General models from recbole library (BPR, CF and MF models): Training is fast but inference is too slow because models don't scale</li>\n<li>Models from surprise library: Same as general models from recbole library</li>\n<li>TF-IDF + pairwise similarity: Very slow inference time since I was using argsort to get top 20 most similar aids</li>\n</ul>",
      "rawMarkdown": "This was my first recommender system project so I wanted to start from basics and move on. These are the things I tried but didn't work for me. Maybe you can make them work.\n\n* Collaborative filtering: Either the scores aren't good or the embeddings become too large and inference is too slow\n* Matrix factorization: Same as collaborative filtering\n* Word2vec: It is pretty fast but not competitive\n* FastText: Slightly better and faster training compared to word2vec from gensim because of c++ bindings but the models are same\n* Doc2vec: Model doesn't learn anything\n* Sequential models from recbole library (GRU4Rec, BERT4Rec and etc.): Training is too slow and they are not competitive\n* General models from recbole library (BPR, CF and MF models): Training is fast but inference is too slow because models don't scale\n* Models from surprise library: Same as general models from recbole library\n* TF-IDF + pairwise similarity: Very slow inference time since I was using argsort to get top 20 most similar aids",
      "votes": null
    },
    {
      "id": "2078652",
      "postDate": "12/28/2022 13:49:37",
      "content": "<p>The published candidation kernels are pretty good, so seems like that competition is mostly about the reranking.</p>",
      "rawMarkdown": "The published candidation kernels are pretty good, so seems like that competition is mostly about the reranking.",
      "votes": null
    },
    {
      "id": "2078740",
      "postDate": "12/28/2022 15:44:45",
      "content": "<p>I believe this competition will be all about good candidates generation and then feature engineering for the re-ranker model. Training good re-ranker model seems to be a challenge at this point - model overfits quite easily. I am trying to get the whole pipeline running and then will focus primarily on candidate generation/feature engineering. </p>",
      "rawMarkdown": "I believe this competition will be all about good candidates generation and then feature engineering for the re-ranker model. Training good re-ranker model seems to be a challenge at this point - model overfits quite easily. I am trying to get the whole pipeline running and then will focus primarily on candidate generation/feature engineering.",
      "votes": null
    },
    {
      "id": "2079057",
      "postDate": "12/28/2022 23:28:29",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>! I think you really need to use a ranker or combine the scores generated via the models that you mention in some way. </p>\n<p>Don't think using any of the models that you mention can get you far.</p>\n<p>Essentially, candidate generation + scoring them (and generating additional features) and then feeding it all to the reranker is the name of the game here 🙂</p>\n<p>Best of luck!</p>",
      "rawMarkdown": "Hey @gunesevitan! I think you really need to use a ranker or combine the scores generated via the models that you mention in some way. \n\nDon't think using any of the models that you mention can get you far.\n\nEssentially, candidate generation + scoring them (and generating additional features) and then feeding it all to the reranker is the name of the game here 🙂\n\nBest of luck!",
      "votes": null
    },
    {
      "id": "2079165",
      "postDate": "12/29/2022 04:27:00",
      "content": "<p>Yep, that's my plan. My intuition was generating those candidates via models would be more robust compared to hand crafted candidates. </p>",
      "rawMarkdown": "Yep, that's my plan. My intuition was generating those candidates via models would be more robust compared to hand crafted candidates.",
      "votes": null
    },
    {
      "id": "2090251",
      "postDate": "01/07/2023 06:50:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>, Thanks for sharing! I was wondering how you used models like word2vec? Did you concatenate each session's sequence of item aids to make a text document? (treating each aid as a word?)</p>",
      "rawMarkdown": "Hi @gunesevitan, Thanks for sharing! I was wondering how you used models like word2vec? Did you concatenate each session's sequence of item aids to make a text document? (treating each aid as a word?)",
      "votes": null
    },
    {
      "id": "2090311",
      "postDate": "01/07/2023 08:07:47",
      "content": "<p>Yes, that's what I exactly did. All sessions are treated as sentences and aids as words. I didn't use click-cart-order information in those models.</p>",
      "rawMarkdown": "Yes, that's what I exactly did. All sessions are treated as sentences and aids as words. I didn't use click-cart-order information in those models.",
      "votes": null
    },
    {
      "id": "2090552",
      "postDate": "01/07/2023 12:53:07",
      "content": "<p>I see, thank you!</p>",
      "rawMarkdown": "I see, thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2078652,
      "author_name": "cabbage972",
      "author_url": "",
      "post_date": "12/28/2022 13:49:37",
      "content": "<p>The published candidation kernels are pretty good, so seems like that competition is mostly about the reranking.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2078740,
      "author_name": "parthpankajtiwary",
      "author_url": "",
      "post_date": "12/28/2022 15:44:45",
      "content": "<p>I believe this competition will be all about good candidates generation and then feature engineering for the re-ranker model. Training good re-ranker model seems to be a challenge at this point - model overfits quite easily. I am trying to get the whole pipeline running and then will focus primarily on candidate generation/feature engineering. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2079057,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "12/28/2022 23:28:29",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>! I think you really need to use a ranker or combine the scores generated via the models that you mention in some way. </p>\n<p>Don't think using any of the models that you mention can get you far.</p>\n<p>Essentially, candidate generation + scoring them (and generating additional features) and then feeding it all to the reranker is the name of the game here 🙂</p>\n<p>Best of luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2079165,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/29/2022 04:27:00",
          "content": "<p>Yep, that's my plan. My intuition was generating those candidates via models would be more robust compared to hand crafted candidates. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2090251,
      "author_name": "cherrytomatech",
      "author_url": "",
      "post_date": "01/07/2023 06:50:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>, Thanks for sharing! I was wondering how you used models like word2vec? Did you concatenate each session's sequence of item aids to make a text document? (treating each aid as a word?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2090311,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "01/07/2023 08:07:47",
          "content": "<p>Yes, that's what I exactly did. All sessions are treated as sentences and aids as words. I didn't use click-cart-order information in those models.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2090552,
              "author_name": "cherrytomatech",
              "author_url": "",
              "post_date": "01/07/2023 12:53:07",
              "content": "<p>I see, thank you!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2078404": "This was my first recommender system project so I wanted to start from basics and move on. These are the things I tried but didn't work for me. Maybe you can make them work.\n\n* Collaborative filtering: Either the scores aren't good or the embeddings become too large and inference is too slow\n* Matrix factorization: Same as collaborative filtering\n* Word2vec: It is pretty fast but not competitive\n* FastText: Slightly better and faster training compared to word2vec from gensim because of c++ bindings but the models are same\n* Doc2vec: Model doesn't learn anything\n* Sequential models from recbole library (GRU4Rec, BERT4Rec and etc.): Training is too slow and they are not competitive\n* General models from recbole library (BPR, CF and MF models): Training is fast but inference is too slow because models don't scale\n* Models from surprise library: Same as general models from recbole library\n* TF-IDF + pairwise similarity: Very slow inference time since I was using argsort to get top 20 most similar aids",
    "2078652": "The published candidation kernels are pretty good, so seems like that competition is mostly about the reranking.",
    "2078740": "I believe this competition will be all about good candidates generation and then feature engineering for the re-ranker model. Training good re-ranker model seems to be a challenge at this point - model overfits quite easily. I am trying to get the whole pipeline running and then will focus primarily on candidate generation/feature engineering.",
    "2079057": "Hey @gunesevitan! I think you really need to use a ranker or combine the scores generated via the models that you mention in some way. \n\nDon't think using any of the models that you mention can get you far.\n\nEssentially, candidate generation + scoring them (and generating additional features) and then feeding it all to the reranker is the name of the game here 🙂\n\nBest of luck!",
    "2079165": "Yep, that's my plan. My intuition was generating those candidates via models would be more robust compared to hand crafted candidates.",
    "2090251": "Hi @gunesevitan, Thanks for sharing! I was wondering how you used models like word2vec? Did you concatenate each session's sequence of item aids to make a text document? (treating each aid as a word?)",
    "2090311": "Yes, that's what I exactly did. All sessions are treated as sentences and aids as words. I didn't use click-cart-order information in those models.",
    "2090552": "I see, thank you!"
  },
  "source": "meta"
}