{
  "id": 324494,
  "title": "A good dataset for learning & doing research in RS?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/324494",
  "author_name": "",
  "post_date": "2022-05-12T03:17:27.649271Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>It seems that there aren't that much real-world datasets in recommendation(there're surely reasons for companys to avoid sharing them). I think the dataset from HM has the following advantages:</p>\n<ul>\n<li>Dataset size is big enough for doing many researches.</li>\n<li>Real-world dataset that with genuine commerical utility. So the skills learned are likely to be useful in real world.</li>\n<li>Rich customer&amp;article features: user&amp;article attributes, numerical and categorical, texts, images. </li>\n<li>Most features are not anonymized. </li>\n</ul>\n<p>I've compared hm dataset with other reco competition dataset, the last 2 point are clear advantages of hm dataset. I'm new to RS and haven't finished yet, so I think it's plausible to continue to work with this dataset.</p>",
  "messages": [
    {
      "id": "1785327",
      "postDate": "05/12/2022 03:17:27",
      "content": "<p>It seems that there aren't that much real-world datasets in recommendation(there're surely reasons for companys to avoid sharing them). I think the dataset from HM has the following advantages:</p>\n<ul>\n<li>Dataset size is big enough for doing many researches.</li>\n<li>Real-world dataset that with genuine commerical utility. So the skills learned are likely to be useful in real world.</li>\n<li>Rich customer&amp;article features: user&amp;article attributes, numerical and categorical, texts, images. </li>\n<li>Most features are not anonymized. </li>\n</ul>\n<p>I've compared hm dataset with other reco competition dataset, the last 2 point are clear advantages of hm dataset. I'm new to RS and haven't finished yet, so I think it's plausible to continue to work with this dataset.</p>",
      "rawMarkdown": "It seems that there aren't that much real-world datasets in recommendation(there're surely reasons for companys to avoid sharing them). I think the dataset from HM has the following advantages:\n- Dataset size is big enough for doing many researches.\n- Real-world dataset that with genuine commerical utility. So the skills learned are likely to be useful in real world.\n- Rich customer&article features: user&article attributes, numerical and categorical, texts, images. \n- Most features are not anonymized. \n\nI've compared hm dataset with other reco competition dataset, the last 2 point are clear advantages of hm dataset. I'm new to RS and haven't finished yet, so I think it's plausible to continue to work with this dataset.",
      "votes": null
    },
    {
      "id": "1785328",
      "postDate": "05/12/2022 03:22:22",
      "content": "<p>But there are indeed disadvantages as well. As GM <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> mentioned, \"A large part of the competition was about predicting the availability of the products for different postal_codes\". Lack of information about availability and product exposure is a problem, the model may need to focus on lots of things unrelated to actual match, which may be part of the reasons why nns don't perform well in this dataset.</p>",
      "rawMarkdown": "But there are indeed disadvantages as well. As GM @paweljankiewicz mentioned, \"A large part of the competition was about predicting the availability of the products for different postal_codes\". Lack of information about availability and product exposure is a problem, the model may need to focus on lots of things unrelated to actual match, which may be part of the reasons why nns don't perform well in this dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1785328,
      "author_name": "homoalways",
      "author_url": "",
      "post_date": "05/12/2022 03:22:22",
      "content": "<p>But there are indeed disadvantages as well. As GM <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> mentioned, \"A large part of the competition was about predicting the availability of the products for different postal_codes\". Lack of information about availability and product exposure is a problem, the model may need to focus on lots of things unrelated to actual match, which may be part of the reasons why nns don't perform well in this dataset.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1785327": "It seems that there aren't that much real-world datasets in recommendation(there're surely reasons for companys to avoid sharing them). I think the dataset from HM has the following advantages:\n- Dataset size is big enough for doing many researches.\n- Real-world dataset that with genuine commerical utility. So the skills learned are likely to be useful in real world.\n- Rich customer&article features: user&article attributes, numerical and categorical, texts, images. \n- Most features are not anonymized. \n\nI've compared hm dataset with other reco competition dataset, the last 2 point are clear advantages of hm dataset. I'm new to RS and haven't finished yet, so I think it's plausible to continue to work with this dataset.",
    "1785328": "But there are indeed disadvantages as well. As GM @paweljankiewicz mentioned, \"A large part of the competition was about predicting the availability of the products for different postal_codes\". Lack of information about availability and product exposure is a problem, the model may need to focus on lots of things unrelated to actual match, which may be part of the reasons why nns don't perform well in this dataset."
  },
  "source": "meta"
}