{
  "id": 317193,
  "title": "Issue with Pure Pytorch Baseline Notebooks",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/317193",
  "author_name": "",
  "post_date": "2022-04-05T21:54:04.404156400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I've been trying to create a Deep Learning model for the competition and I found the notebook by the user AHMET ERDEM really useful. However, I just noticed something and I believe it might be causing some issues while training the model and also during the prediction time.</p>\n<p>In the notebook: <a href=\"https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline\" target=\"_blank\">https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline</a><br>\nAhmet uses sklearn's LabelEncoder to transform the article_id to integer indices. This causes some article to be encoded as index 0. I believe that the problem with this approach comes afterwards in the HMDataset part of the code, because the sequences are zero-padded. This is a problem because one of the articles is being represented as a 0.  I believe this can make the algorithm predict that particular article a lot. Even when the score is somewhere around 0.018, I also believe that if this issue is addressed, it can be improved.</p>\n<p>Has some one else noticed this? Am I right in this line of reasoning? I'm trying to figure out how to go around this problem, because the calculation of the error and the obtention of predictions depend on the indices given by the encoded articles, so I just can't make all the encodings be +1.</p>",
  "messages": [
    {
      "id": "1746508",
      "postDate": "04/05/2022 21:54:04",
      "content": "<p>I've been trying to create a Deep Learning model for the competition and I found the notebook by the user AHMET ERDEM really useful. However, I just noticed something and I believe it might be causing some issues while training the model and also during the prediction time.</p>\n<p>In the notebook: <a href=\"https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline\" target=\"_blank\">https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline</a><br>\nAhmet uses sklearn's LabelEncoder to transform the article_id to integer indices. This causes some article to be encoded as index 0. I believe that the problem with this approach comes afterwards in the HMDataset part of the code, because the sequences are zero-padded. This is a problem because one of the articles is being represented as a 0.  I believe this can make the algorithm predict that particular article a lot. Even when the score is somewhere around 0.018, I also believe that if this issue is addressed, it can be improved.</p>\n<p>Has some one else noticed this? Am I right in this line of reasoning? I'm trying to figure out how to go around this problem, because the calculation of the error and the obtention of predictions depend on the indices given by the encoded articles, so I just can't make all the encodings be +1.</p>",
      "rawMarkdown": "I've been trying to create a Deep Learning model for the competition and I found the notebook by the user AHMET ERDEM really useful. However, I just noticed something and I believe it might be causing some issues while training the model and also during the prediction time.\n\nIn the notebook: https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline\nAhmet uses sklearn's LabelEncoder to transform the article_id to integer indices. This causes some article to be encoded as index 0. I believe that the problem with this approach comes afterwards in the HMDataset part of the code, because the sequences are zero-padded. This is a problem because one of the articles is being represented as a 0.  I believe this can make the algorithm predict that particular article a lot. Even when the score is somewhere around 0.018, I also believe that if this issue is addressed, it can be improved.\n\nHas some one else noticed this? Am I right in this line of reasoning? I'm trying to figure out how to go around this problem, because the calculation of the error and the obtention of predictions depend on the indices given by the encoded articles, so I just can't make all the encodings be +1.",
      "votes": null
    },
    {
      "id": "1758389",
      "postDate": "04/17/2022 16:22:56",
      "content": "<p>This sounds like a really serious problem, may I ask have you resolved this problem yet? If so, could you please share your resolution? Thank you so much for pointing this out!</p>",
      "rawMarkdown": "This sounds like a really serious problem, may I ask have you resolved this problem yet? If so, could you please share your resolution? Thank you so much for pointing this out!",
      "votes": null
    },
    {
      "id": "1770135",
      "postDate": "04/28/2022 02:31:33",
      "content": "<p>article_hist = torch.ones(self.seq_len).long() * (self.article_len-1)</p>",
      "rawMarkdown": "article_hist = torch.ones(self.seq_len).long() * (self.article_len-1)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1758389,
      "author_name": "skybookreader",
      "author_url": "",
      "post_date": "04/17/2022 16:22:56",
      "content": "<p>This sounds like a really serious problem, may I ask have you resolved this problem yet? If so, could you please share your resolution? Thank you so much for pointing this out!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1770135,
      "author_name": "onemanonearmy",
      "author_url": "",
      "post_date": "04/28/2022 02:31:33",
      "content": "<p>article_hist = torch.ones(self.seq_len).long() * (self.article_len-1)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1746508": "I've been trying to create a Deep Learning model for the competition and I found the notebook by the user AHMET ERDEM really useful. However, I just noticed something and I believe it might be causing some issues while training the model and also during the prediction time.\n\nIn the notebook: https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline\nAhmet uses sklearn's LabelEncoder to transform the article_id to integer indices. This causes some article to be encoded as index 0. I believe that the problem with this approach comes afterwards in the HMDataset part of the code, because the sequences are zero-padded. This is a problem because one of the articles is being represented as a 0.  I believe this can make the algorithm predict that particular article a lot. Even when the score is somewhere around 0.018, I also believe that if this issue is addressed, it can be improved.\n\nHas some one else noticed this? Am I right in this line of reasoning? I'm trying to figure out how to go around this problem, because the calculation of the error and the obtention of predictions depend on the indices given by the encoded articles, so I just can't make all the encodings be +1.",
    "1758389": "This sounds like a really serious problem, may I ask have you resolved this problem yet? If so, could you please share your resolution? Thank you so much for pointing this out!",
    "1770135": "article_hist = torch.ones(self.seq_len).long() * (self.article_len-1)"
  },
  "source": "meta"
}