{
  "id": 320460,
  "title": "Information from H&M website?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/320460",
  "author_name": "",
  "post_date": "2022-04-21T17:46:02.679207100Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I found this competition a couple days ago and have been learning from other participant's notebooks and discussion.  Before diving into the competition and pouring in hours of work, I want to make sure that web-scrapping is not playing a role at this competition.</p>\n<p>Here is the question: <br>\nWill the use of information from H&amp;M's website give significant advantage?</p>\n<p>I took a sample of <code>articles.csv</code> and found that most of the articles still have a product page on H&amp;M's website. The <a href=\"https://www2.hm.com/en_us/productpage.0762558201.html\" target=\"_blank\">product page</a> contains two sections that can be informative: <img src=\"https://i.imgur.com/d7dE2fC.jpg\" alt=\"product_webpage\"><br>\n(1) <code>Others also bought</code>: it shows 12 items, just like the competition's prediction.  <br>\n(2) <code>Style with</code>: also recommends 12 items.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">notebook</a> shows that recommending items frequent purchased together is effective, where the items (that purchased together) are computed from the competition dataset. What if we use the items from the <code>others also bought</code> session from the product webpage instead? Those could contain information about what users bought <strong>AFTER</strong> the training period.</p>\n<p>Please share your thought. Thanks!</p>",
  "messages": [
    {
      "id": "1763677",
      "postDate": "04/21/2022 17:46:02",
      "content": "<p>I found this competition a couple days ago and have been learning from other participant's notebooks and discussion.  Before diving into the competition and pouring in hours of work, I want to make sure that web-scrapping is not playing a role at this competition.</p>\n<p>Here is the question: <br>\nWill the use of information from H&amp;M's website give significant advantage?</p>\n<p>I took a sample of <code>articles.csv</code> and found that most of the articles still have a product page on H&amp;M's website. The <a href=\"https://www2.hm.com/en_us/productpage.0762558201.html\" target=\"_blank\">product page</a> contains two sections that can be informative: <img src=\"https://i.imgur.com/d7dE2fC.jpg\" alt=\"product_webpage\"><br>\n(1) <code>Others also bought</code>: it shows 12 items, just like the competition's prediction.  <br>\n(2) <code>Style with</code>: also recommends 12 items.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">notebook</a> shows that recommending items frequent purchased together is effective, where the items (that purchased together) are computed from the competition dataset. What if we use the items from the <code>others also bought</code> session from the product webpage instead? Those could contain information about what users bought <strong>AFTER</strong> the training period.</p>\n<p>Please share your thought. Thanks!</p>",
      "rawMarkdown": "I found this competition a couple days ago and have been learning from other participant's notebooks and discussion.  Before diving into the competition and pouring in hours of work, I want to make sure that web-scrapping is not playing a role at this competition.\n\nHere is the question: \nWill the use of information from H&M's website give significant advantage?\n\nI took a sample of `articles.csv` and found that most of the articles still have a product page on H&M's website. The [product page](https://www2.hm.com/en_us/productpage.0762558201.html) contains two sections that can be informative: ![product_webpage](https://i.imgur.com/d7dE2fC.jpg)\n(1) `Others also bought`: it shows 12 items, just like the competition's prediction.  \n(2) `Style with`: also recommends 12 items.\n\n\n\n@cdeotte's [notebook] (https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021) shows that recommending items frequent purchased together is effective, where the items (that purchased together) are computed from the competition dataset. What if we use the items from the `others also bought` session from the product webpage instead? Those could contain information about what users bought **AFTER** the training period.\n\nPlease share your thought. Thanks!",
      "votes": null
    },
    {
      "id": "1763935",
      "postDate": "04/22/2022 01:40:20",
      "content": "<p>Thank you for pointing it out.<br>\nI agree that web-scraping should not play any role at this competition. <br>\nI hope the host clearly prohibit this approach.</p>",
      "rawMarkdown": "Thank you for pointing it out.\nI agree that web-scraping should not play any role at this competition. \nI hope the host clearly prohibit this approach.",
      "votes": null
    },
    {
      "id": "1763943",
      "postDate": "04/22/2022 02:00:49",
      "content": "<p>Thanks for sharing.<br>\nI think scrapping from the website is an interesting approach in competition, just like the easter eggs in the game.</p>",
      "rawMarkdown": "Thanks for sharing.\nI think scrapping from the website is an interesting approach in competition, just like the easter eggs in the game.",
      "votes": null
    },
    {
      "id": "1764056",
      "postDate": "04/22/2022 05:39:01",
      "content": "<p>Personally I don't want to participate in a competition where web-scrapping gives advantages, but it's interesting to think about how one can utilize such information.</p>\n<p>Another potential usage of H&amp;M's websites is to identify the sell country of a product (and hence the customer who purchase the product): One can use article_id to find its product webpage. If the webpage is only available at hm.com.cn and not at other countries' websites, we know the product is only sold in China.    </p>",
      "rawMarkdown": "Personally I don't want to participate in a competition where web-scrapping gives advantages, but it's interesting to think about how one can utilize such information.\n\nAnother potential usage of H&M's websites is to identify the sell country of a product (and hence the customer who purchase the product): One can use article_id to find its product webpage. If the webpage is only available at hm.com.cn and not at other countries' websites, we know the product is only sold in China.",
      "votes": null
    },
    {
      "id": "1770738",
      "postDate": "04/28/2022 14:41:32",
      "content": "<p>I want to know that  the data if someone scrapped from H&amp;M website should make it public according the competition rule?</p>",
      "rawMarkdown": "I want to know that  the data if someone scrapped from H&M website should make it public according the competition rule?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1763935,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "04/22/2022 01:40:20",
      "content": "<p>Thank you for pointing it out.<br>\nI agree that web-scraping should not play any role at this competition. <br>\nI hope the host clearly prohibit this approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1763943,
      "author_name": "",
      "author_url": "",
      "post_date": "04/22/2022 02:00:49",
      "content": "<p>Thanks for sharing.<br>\nI think scrapping from the website is an interesting approach in competition, just like the easter eggs in the game.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1764056,
          "author_name": "ncchen",
          "author_url": "",
          "post_date": "04/22/2022 05:39:01",
          "content": "<p>Personally I don't want to participate in a competition where web-scrapping gives advantages, but it's interesting to think about how one can utilize such information.</p>\n<p>Another potential usage of H&amp;M's websites is to identify the sell country of a product (and hence the customer who purchase the product): One can use article_id to find its product webpage. If the webpage is only available at hm.com.cn and not at other countries' websites, we know the product is only sold in China.    </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1770738,
      "author_name": "juzqyxs",
      "author_url": "",
      "post_date": "04/28/2022 14:41:32",
      "content": "<p>I want to know that  the data if someone scrapped from H&amp;M website should make it public according the competition rule?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1763677": "I found this competition a couple days ago and have been learning from other participant's notebooks and discussion.  Before diving into the competition and pouring in hours of work, I want to make sure that web-scrapping is not playing a role at this competition.\n\nHere is the question: \nWill the use of information from H&M's website give significant advantage?\n\nI took a sample of `articles.csv` and found that most of the articles still have a product page on H&M's website. The [product page](https://www2.hm.com/en_us/productpage.0762558201.html) contains two sections that can be informative: ![product_webpage](https://i.imgur.com/d7dE2fC.jpg)\n(1) `Others also bought`: it shows 12 items, just like the competition's prediction.  \n(2) `Style with`: also recommends 12 items.\n\n\n\n@cdeotte's [notebook] (https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021) shows that recommending items frequent purchased together is effective, where the items (that purchased together) are computed from the competition dataset. What if we use the items from the `others also bought` session from the product webpage instead? Those could contain information about what users bought **AFTER** the training period.\n\nPlease share your thought. Thanks!",
    "1763935": "Thank you for pointing it out.\nI agree that web-scraping should not play any role at this competition. \nI hope the host clearly prohibit this approach.",
    "1763943": "Thanks for sharing.\nI think scrapping from the website is an interesting approach in competition, just like the easter eggs in the game.",
    "1764056": "Personally I don't want to participate in a competition where web-scrapping gives advantages, but it's interesting to think about how one can utilize such information.\n\nAnother potential usage of H&M's websites is to identify the sell country of a product (and hence the customer who purchase the product): One can use article_id to find its product webpage. If the webpage is only available at hm.com.cn and not at other countries' websites, we know the product is only sold in China.",
    "1770738": "I want to know that  the data if someone scrapped from H&M website should make it public according the competition rule?"
  },
  "source": "meta"
}