{
  "id": 305952,
  "title": "Greetings!",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/305952",
  "author_name": "FridaRim",
  "post_date": "2022-02-07T16:01:21.065000",
  "votes": 92,
  "comment_count": 68,
  "views": 0,
  "content": "<p>Dear Kagglers,</p>\n<p>On behalf of H&amp;M Group and AI Recommendation Engine team I would like to officially welcome you to the “H&amp;M Personalized Fashion Recommendations” competition!</p>\n<p>In this competition we invite you to produce product recommendations based on data from historical transactions, as well as from customer and product meta data. The available meta data spans from simple data, such as garment type and customer age, to text data from product descriptions, to image data from garment images.</p>\n<p>We can’t wait to see what model approaches or feature engineering techniques the community will use to arrive at the best possible solution. Along with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you develop the best models possible!</p>\n<p>Good luck, and have fun!</p>",
  "messages": [
    {
      "id": 1680156,
      "postDate": "2022-02-07T16:01:21.067Z",
      "content": "<p>Dear Kagglers,</p>\n<p>On behalf of H&amp;M Group and AI Recommendation Engine team I would like to officially welcome you to the “H&amp;M Personalized Fashion Recommendations” competition!</p>\n<p>In this competition we invite you to produce product recommendations based on data from historical transactions, as well as from customer and product meta data. The available meta data spans from simple data, such as garment type and customer age, to text data from product descriptions, to image data from garment images.</p>\n<p>We can’t wait to see what model approaches or feature engineering techniques the community will use to arrive at the best possible solution. Along with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you develop the best models possible!</p>\n<p>Good luck, and have fun!</p>",
      "rawMarkdown": "Dear Kagglers,\n\nOn behalf of H&M Group and AI Recommendation Engine team I would like to officially welcome you to the “H&M Personalized Fashion Recommendations” competition!\n\nIn this competition we invite you to produce product recommendations based on data from historical transactions, as well as from customer and product meta data. The available meta data spans from simple data, such as garment type and customer age, to text data from product descriptions, to image data from garment images.\n\nWe can’t wait to see what model approaches or feature engineering techniques the community will use to arrive at the best possible solution. Along with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you develop the best models possible!\n\nGood luck, and have fun!",
      "votes": 91
    },
    {
      "id": 1711192,
      "postDate": "2022-03-03T17:39:02.117Z",
      "content": "<p>Hi, thanks for a great comp <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a>  FYI, your metric formula is wrong. It should say</p>\n<p>$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U} \\left( \\frac{1}{min(m,12)} \\sum_{k=1}^{min(n,12)}{P(k) \\times rel(k)} \\right) $$</p>\n<p>where <code>m</code> is the number of ground truths per customer, <code>n</code> is the number of predictions per customer, and <code>U</code> is the number of customers. <code>P(k)</code> is the precision at cutoff <code>k</code>, and <code>rel(k)</code> is an indicator function equaling 1 if the item at rank <code>k</code> is a relevant (correct) label, zero otherwise.</p>",
      "rawMarkdown": "Hi, thanks for a great comp @fridarimark @maggiemd  FYI, your metric formula is wrong. It should say\n\n$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U} \\left( \\frac{1}{min(m,12)} \\sum_{k=1}^{min(n,12)}{P(k) \\times rel(k)} \\right) $$\n\nwhere `m` is the number of ground truths per customer, `n` is the number of predictions per customer, and `U` is the number of customers. `P(k)` is the precision at cutoff `k`, and `rel(k)` is an indicator function equaling 1 if the item at rank `k` is a relevant (correct) label, zero otherwise.",
      "votes": 14,
      "replies": [
        {
          "id": 1747078,
          "postDate": "2022-04-06T10:21:02.670Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I have a small doubt. </p>\n<ol>\n<li>should Articles sub be made in an ordered way ?</li>\n<li>If truth value is 2 0 1 and i make the sub 3 1 , what would be the score?</li>\n</ol>",
          "rawMarkdown": "@cdeotte I have a small doubt. \n1. should Articles sub be made in an ordered way ?\n2. If truth value is 2 0 1 and i make the sub 3 1 , what would be the score?"
        }
      ]
    },
    {
      "id": 1683754,
      "postDate": "2022-02-10T03:19:15.497Z",
      "content": "<p>Question about the variables:<br>\n FN, Active - What are these?<br>\nSales channel ID - is is right to assume 2 is in person and 1 is online?<br>\nPostal code - 1.2 mil is a lot of postal codes (approx every country you sell in) is there some else in these hashes?<br>\nwhat percent of the customers are new in the validation set (any idea helps)?<br>\nThanks</p>",
      "rawMarkdown": "Question about the variables:\n FN, Active - What are these?\nSales channel ID - is is right to assume 2 is in person and 1 is online?\nPostal code - 1.2 mil is a lot of postal codes (approx every country you sell in) is there some else in these hashes?\nwhat percent of the customers are new in the validation set (any idea helps)?\nThanks",
      "votes": 10,
      "replies": [
        {
          "id": 1684481,
          "postDate": "2022-02-10T14:21:20.813Z",
          "content": "<p>FN is if a customer get Fashion News newsletter, Active is if the customer is active for communication, sales channel id, 2 is online and 1 store.</p>",
          "rawMarkdown": "FN is if a customer get Fashion News newsletter, Active is if the customer is active for communication, sales channel id, 2 is online and 1 store.",
          "votes": 26
        }
      ]
    },
    {
      "id": 1690119,
      "postDate": "2022-02-14T17:53:06.300Z",
      "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> </p>\n<blockquote>\n  <p>Customer that did not make any purchase during test period are excluded from the scoring</p>\n</blockquote>\n<p>So, we don't need to predict those who didn't purchase any items. <strong>BUT</strong> sample_submission has 1371980 customers, and number of customers who bought items as of last one week is 68984. This means almost 95% of our prediction won't be scored during test period.<br>\nIMHO this is totally nonsense, and we should skip those for saving our time. What do you think? If yes, please update sample_submission or give us the customer_id who are being scored.</p>",
      "rawMarkdown": "@fridarimark \n> Customer that did not make any purchase during test period are excluded from the scoring\n\nSo, we don't need to predict those who didn't purchase any items. **BUT** sample_submission has 1371980 customers, and number of customers who bought items as of last one week is 68984. This means almost 95% of our prediction won't be scored during test period.\nIMHO this is totally nonsense, and we should skip those for saving our time. What do you think? If yes, please update sample_submission or give us the customer_id who are being scored.",
      "votes": 9,
      "replies": [
        {
          "id": 1691807,
          "postDate": "2022-02-15T16:09:09.410Z",
          "content": "<p>I have to disagree on several points. </p>\n<p>First, everything is scored in the sense that the MAP metric will take into account all your predictions. It just happens that some terms in that expression will be weighted 0. The whole challenge here is to design a system that will perform well in the circumstance where we don't know which weights will be 1s.</p>\n<p>Second, if we knew which people made purchases in the test week it would probably amount to considerable data leakage. I'm sure kagglers would find a way to exploit this to optimize their models, but it would make those models much less useful in the realistic scenario where H&amp;M doesn't know who will make purchases next week. Again, that would defeat the purpose of the challenge.</p>\n<p>Third, the use-case for this sort of thing is to generate recommendations for articles to users browsing the website or receiving emails. So you need to be able to generate recommendations for everyone, even if you don't know who will actually follow through and make a purchase.</p>\n<p>Finally, there really aren't that many customers (relatively speaking). I'm not sure what sort of time you are concerned about, but once you have an algorithm that can make predictions for 70k customers, surely it's not that much more costly in terms of time to apply it to all 1.3 M customers…</p>\n<p>Personally, I really like this challenge because the number of publicly available, realistic, managebly sized datasets for training recommender systems with a time series component (such as this one) is quite small, at least to my knowledge.</p>",
          "rawMarkdown": "I have to disagree on several points. \n\nFirst, everything is scored in the sense that the MAP metric will take into account all your predictions. It just happens that some terms in that expression will be weighted 0. The whole challenge here is to design a system that will perform well in the circumstance where we don't know which weights will be 1s.\n\nSecond, if we knew which people made purchases in the test week it would probably amount to considerable data leakage. I'm sure kagglers would find a way to exploit this to optimize their models, but it would make those models much less useful in the realistic scenario where H&M doesn't know who will make purchases next week. Again, that would defeat the purpose of the challenge.\n\nThird, the use-case for this sort of thing is to generate recommendations for articles to users browsing the website or receiving emails. So you need to be able to generate recommendations for everyone, even if you don't know who will actually follow through and make a purchase.\n\nFinally, there really aren't that many customers (relatively speaking). I'm not sure what sort of time you are concerned about, but once you have an algorithm that can make predictions for 70k customers, surely it's not that much more costly in terms of time to apply it to all 1.3 M customers...\n\nPersonally, I really like this challenge because the number of publicly available, realistic, managebly sized datasets for training recommender systems with a time series component (such as this one) is quite small, at least to my knowledge.",
          "votes": 9
        },
        {
          "id": 1691875,
          "postDate": "2022-02-15T16:55:45.660Z",
          "content": "<blockquote>\n  <p>The whole challenge here is to design a system that will perform well in the circumstance where we don't know which weights will be 1s</p>\n</blockquote>\n<p>Yes, I think so. But this is not the objection for unnecessary of dummy data.</p>\n<blockquote>\n  <p>if we knew which people made purchases in the test week it would probably amount to considerable data leakage</p>\n</blockquote>\n<p>What kind of leakage? We are making predictions based on the assumption that all customers will buy something. So, it's totally same thing and our prediction won't change if we know which customer would buy next week.</p>\n<blockquote>\n  <p>you need to be able to generate recommendations for everyone, even if you don't know who will actually follow through and make a purchase.</p>\n</blockquote>\n<p>Yes, I think so. But this is not the objection for unnecessary of dummy data.</p>\n<blockquote>\n  <p>but once you have an algorithm that can make predictions for 70k customers, surely it's not that much more costly in terms of time to apply it to all 1.3 M customers</p>\n</blockquote>\n<p>At least, my DGX which has 250G memory can't inference 1.3 M customers at once.</p>",
          "rawMarkdown": "> The whole challenge here is to design a system that will perform well in the circumstance where we don't know which weights will be 1s\n\nYes, I think so. But this is not the objection for unnecessary of dummy data.\n\n\n> if we knew which people made purchases in the test week it would probably amount to considerable data leakage\n\nWhat kind of leakage? We are making predictions based on the assumption that all customers will buy something. So, it's totally same thing and our prediction won't change if we know which customer would buy next week.\n\n\n> you need to be able to generate recommendations for everyone, even if you don't know who will actually follow through and make a purchase.\n\nYes, I think so. But this is not the objection for unnecessary of dummy data.\n\n> but once you have an algorithm that can make predictions for 70k customers, surely it's not that much more costly in terms of time to apply it to all 1.3 M customers\n\nAt least, my DGX which has 250G memory can't inference 1.3 M customers at once.\n",
          "votes": 1
        },
        {
          "id": 1692061,
          "postDate": "2022-02-15T19:20:05.380Z",
          "content": "<p>I suppose then the challenge for you is to design a more scalable algorithm… I would suggest doing inference in batches… Impossible to know what sort of algorithm you are running, but, e.g.,  scaling matrix factorization based CF methods to large numbers of users x items is a pretty well documented thing. In this case there isn't even <em>that many</em> users x items, just enough that making a matrix of that dimension will crash most reasonably sized systems, so getting around it with sparse representations and batched inference is probably the way to go (at least, it works for me).</p>",
          "rawMarkdown": "I suppose then the challenge for you is to design a more scalable algorithm... I would suggest doing inference in batches... Impossible to know what sort of algorithm you are running, but, e.g.,  scaling matrix factorization based CF methods to large numbers of users x items is a pretty well documented thing. In this case there isn't even *that many* users x items, just enough that making a matrix of that dimension will crash most reasonably sized systems, so getting around it with sparse representations and batched inference is probably the way to go (at least, it works for me).",
          "votes": 1
        },
        {
          "id": 1692063,
          "postDate": "2022-02-15T19:25:43.897Z",
          "content": "<blockquote>\n  <p>I suppose then the challenge for you is to design a more scalable algorithm…</p>\n</blockquote>\n<p>I can agree on that, but is that the host's angle as well? If so why not notebook competition?<br>\nIf not, no need to waste our time.</p>",
          "rawMarkdown": "> I suppose then the challenge for you is to design a more scalable algorithm…\n\nI can agree on that, but is that the host's angle as well? If so why not notebook competition?\nIf not, no need to waste our time.",
          "votes": 2
        },
        {
          "id": 1693558,
          "postDate": "2022-02-16T18:31:25.510Z",
          "content": "<p>I don't think so because it will bias the predictions</p>",
          "rawMarkdown": "I don't think so because it will bias the predictions"
        },
        {
          "id": 1693749,
          "postDate": "2022-02-16T22:50:03.600Z",
          "content": "<p><a href=\"https://www.kaggle.com/achahboune\" target=\"_blank\">@achahboune</a> What do you mean? How do we evaluate the dummies which won't be evaluated?</p>",
          "rawMarkdown": "@achahboune What do you mean? How do we evaluate the dummies which won't be evaluated?",
          "votes": 2
        },
        {
          "id": 1705892,
          "postDate": "2022-02-27T00:28:27.423Z",
          "content": "<p><a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>  </p>\n<p>There are 3 parts of the task</p>\n<ol>\n<li>Which Customers will buy?</li>\n<li>Which Articles will be sold?</li>\n<li>Which Customers will buy which articles?</li>\n</ol>\n<p>This is basically building a recommendation system based on historical data. You know who actually purchased what, so it allows you to test whether your recommendations, if they had been made, would have relevant to those customers. Now, nobody really buys everything that is recommended so even if the same recommendations were made, the same purchases may not have been made.</p>\n<p>The premise here is such a recommendation system may work for all customers for all scenarios. It may not. It may need to be further tested and fine tuned.</p>\n<p>If the host gives you No.1, then the problem becomes less complex. It also becomes more of a prediction problem which means that given a set of Customers, you may predict what they may buy which limits the real world use of it.</p>",
          "rawMarkdown": "@onodera  \n\nThere are 3 parts of the task\n1. Which Customers will buy?\n2. Which Articles will be sold?\n3. Which Customers will buy which articles?\n\nThis is basically building a recommendation system based on historical data. You know who actually purchased what, so it allows you to test whether your recommendations, if they had been made, would have relevant to those customers. Now, nobody really buys everything that is recommended so even if the same recommendations were made, the same purchases may not have been made.\n\nThe premise here is such a recommendation system may work for all customers for all scenarios. It may not. It may need to be further tested and fine tuned.\n\nIf the host gives you No.1, then the problem becomes less complex. It also becomes more of a prediction problem which means that given a set of Customers, you may predict what they may buy which limits the real world use of it."
        },
        {
          "id": 1729766,
          "postDate": "2022-03-20T13:40:40.027Z",
          "content": "<p>No 1 is not included in the metric since those customers are excluded from the leaderboard. Onodera is right, including them in the submission file is a waste of compute power (and thus money + co2).</p>",
          "rawMarkdown": "No 1 is not included in the metric since those customers are excluded from the leaderboard. Onodera is right, including them in the submission file is a waste of compute power (and thus money + co2).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1708321,
      "postDate": "2022-03-01T11:11:24.590Z",
      "content": "<p>Does Kaggle have official position about russian citizens' participation? Will they be able to earn medals and ranking points?</p>\n<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> </p>",
      "rawMarkdown": "Does Kaggle have official position about russian citizens' participation? Will they be able to earn medals and ranking points?\n\n@maggiemd @fridarimark ",
      "votes": 8,
      "replies": [
        {
          "id": 1717941,
          "postDate": "2022-03-10T11:20:22.477Z",
          "content": "<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> may you comment it, please?</p>",
          "rawMarkdown": "@maggiemd @fridarimark may you comment it, please?"
        },
        {
          "id": 1723056,
          "postDate": "2022-03-15T04:52:54.833Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/simakov\" target=\"_blank\">@simakov</a> Kaggle is available to Russian citizens for participation in competitions and for medals, ranking points, etc. The OFAC rules have been updated to exclude participants from disputed areas, but Russian citizens who do not live in these areas can continue to participate. </p>\n<p>\"Competitions are open to residents of the United States and worldwide, except that if you are a resident of Crimea, so-called Donetsk People's Republic (DNR) or Luhansk People's Republic (LNR), Cuba, Iran, Syria, or North Korea, or are subject to U.S. export controls or sanctions, you may not enter the Competition.\"</p>",
          "rawMarkdown": "Hello @simakov Kaggle is available to Russian citizens for participation in competitions and for medals, ranking points, etc. The OFAC rules have been updated to exclude participants from disputed areas, but Russian citizens who do not live in these areas can continue to participate. \n\n\"Competitions are open to residents of the United States and worldwide, except that if you are a resident of Crimea, so-called Donetsk People's Republic (DNR) or Luhansk People's Republic (LNR), Cuba, Iran, Syria, or North Korea, or are subject to U.S. export controls or sanctions, you may not enter the Competition.\"",
          "votes": 5
        }
      ]
    },
    {
      "id": 1690329,
      "postDate": "2022-02-14T21:40:52.207Z",
      "content": "<p>Actually what is missing is the 'available stock' since we don't know that stock, its impossible your site will show the stock that is not avaiable, and you suggest we forecast what they want to purchase, you will endup with a wrong forecast of the historical stock. Or you we should make a guess about the available stock as the last 6month purchases</p>",
      "rawMarkdown": "Actually what is missing is the 'available stock' since we don't know that stock, its impossible your site will show the stock that is not avaiable, and you suggest we forecast what they want to purchase, you will endup with a wrong forecast of the historical stock. Or you we should make a guess about the available stock as the last 6month purchases",
      "votes": 6
    },
    {
      "id": 1687970,
      "postDate": "2022-02-13T09:21:16.153Z",
      "content": "<p>Thanks for hosting this interesting competition.</p>\n<p>I would like to ask a question. If I'm not wrong, <b>there are several product names that have several product codes, and also product codes that are shared by various product names.</b> For example, the product name 'Rose dress' has associated product codes: 552212, 673917, 829601, 892278, 920394. At the same time, the product code 600229 has theses product names associated: 'CAGE SLIM EASY IRON US', 'Cage US', 'Cage easy iron',  'Cage Slim Easy Iron'. Could you explain the difference and which one is important?</p>",
      "rawMarkdown": "Thanks for hosting this interesting competition.\n\nI would like to ask a question. If I'm not wrong, <b>there are several product names that have several product codes, and also product codes that are shared by various product names.</b> For example, the product name 'Rose dress' has associated product codes: 552212, 673917, 829601, 892278, 920394. At the same time, the product code 600229 has theses product names associated: 'CAGE SLIM EASY IRON US', 'Cage US', 'Cage easy iron',  'Cage Slim Easy Iron'. Could you explain the difference and which one is important?",
      "votes": 6,
      "replies": [
        {
          "id": 1700309,
          "postDate": "2022-02-21T20:22:58.057Z",
          "content": "<p>One possible explanation would be that 5 different dresses that classify as 'Rose dress' have been sold. Like one with big roses, one with small roses, one just having the color rose…</p>\n<p>For the second case, it might be human error when entering the product name. All product names sound somewhat similar. So maybe someone did not check for already existing product names or different countries selling this item have different naming rules.</p>\n<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a>, can you confirm?</p>",
          "rawMarkdown": "One possible explanation would be that 5 different dresses that classify as 'Rose dress' have been sold. Like one with big roses, one with small roses, one just having the color rose...\n\nFor the second case, it might be human error when entering the product name. All product names sound somewhat similar. So maybe someone did not check for already existing product names or different countries selling this item have different naming rules.\n\n@fridarimark, can you confirm?"
        },
        {
          "id": 1705898,
          "postDate": "2022-02-27T00:34:39.910Z",
          "content": "<p>Overall, for the purpose of this competition, depending on the numbers of such issues and how you use the relevant data in your work, it shouldn't matter much. Of course, ideally, H&amp;M should clean up their data. May be they are listening in 😊</p>",
          "rawMarkdown": "Overall, for the purpose of this competition, depending on the numbers of such issues and how you use the relevant data in your work, it shouldn't matter much. Of course, ideally, H&M should clean up their data. May be they are listening in 😊"
        }
      ]
    },
    {
      "id": 1754746,
      "postDate": "2022-04-14T00:54:49.320Z",
      "content": "<p>Are the articles that had never been sold in transaction_train.csv and are included in article.csv new release in test data?</p>",
      "rawMarkdown": "Are the articles that had never been sold in transaction_train.csv and are included in article.csv new release in test data?",
      "votes": 1,
      "replies": [
        {
          "id": 1755770,
          "postDate": "2022-04-15T00:52:22.817Z",
          "content": "<p>Thank you for the good information, Ryota-san!<br>\nIf I were you, I would investigate the relation between the size of the 10-digit codes and the release date which means the first date appeared in the train.csv.</p>",
          "rawMarkdown": "Thank you for the good information, Ryota-san!\nIf I were you, I would investigate the relation between the size of the 10-digit codes and the release date which means the first date appeared in the train.csv.",
          "votes": 1
        },
        {
          "id": 1755828,
          "postDate": "2022-04-15T03:15:44.230Z",
          "content": "<p>Thank you for advice!!<br>\nAs you said, I analyzed between article_id and the first day each articles were sold. As a result, articles that has large id number tend to be sold for first time lately.<br>\nBut there are exception. I think maybe there are articles that was launched a long time ago but only recently sold. </p>\n<p>I will try to analyze more!</p>",
          "rawMarkdown": "Thank you for advice!!\nAs you said, I analyzed between article_id and the first day each articles were sold. As a result, articles that has large id number tend to be sold for first time lately.\nBut there are exception. I think maybe there are articles that was launched a long time ago but only recently sold. \n\nI will try to analyze more!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1702205,
      "postDate": "2022-02-23T12:25:45.853Z",
      "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> <br>\n  Thank you for organizing this exciting competition!<br>\n  I am very glad to join!<br>\n  I would like to ask one question.<br>\n  I would like to know the criteria for <strong><code>rel(k)</code></strong> to take <strong><code>1</code></strong> value. (<strong><code>rel(k)</code></strong> appears in the Mean Average Precision formula)<br>\n$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k) $$<br>\n<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/overview/evaluation\" target=\"_blank\">Evaluation</a></p>\n<p>You say that <strong><code>rel(k)</code></strong> will be <strong><code>1</code></strong> if the predicted item is a <strong>related label</strong>, but what is the condition for related?</p>\n<ul>\n<li><p><code>article_id</code> matches (most strict?)</p></li>\n<li><p><code>product_code</code> matches</p></li>\n<li><p><code>section_no</code> matches</p></li>\n<li><p>etc…</p>\n<p>Thank you!</p></li>\n</ul>",
      "rawMarkdown": "@fridarimark \n  Thank you for organizing this exciting competition!\n  I am very glad to join!\n  I would like to ask one question.\n  I would like to know the criteria for **`rel(k)`** to take **`1`** value. (**`rel(k)`** appears in the Mean Average Precision formula)\n$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k) $$\n[Evaluation](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/overview/evaluation)\n\n  You say that **`rel(k)`** will be **`1`** if the predicted item is a **related label**, but what is the condition for related?\n  \n  - `article_id` matches (most strict?)\n  - `product_code` matches\n  - `section_no` matches\n  - etc...\n  \n  Thank you!\n",
      "votes": 1,
      "replies": [
        {
          "id": 1704224,
          "postDate": "2022-02-25T10:39:38.530Z",
          "content": "<p>Relevant label means correct label here, so it needs to match the article id that was bought</p>",
          "rawMarkdown": "Relevant label means correct label here, so it needs to match the article id that was bought",
          "votes": 2
        },
        {
          "id": 1704491,
          "postDate": "2022-02-25T15:15:22.860Z",
          "content": "<p>Thank you for your answer !! 😊</p>",
          "rawMarkdown": "Thank you for your answer !! 😊"
        },
        {
          "id": 1705130,
          "postDate": "2022-02-26T06:54:13.170Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1750907,
          "postDate": "2022-04-10T07:57:42.637Z",
          "content": "<p>Great ! thank you for your answer !</p>",
          "rawMarkdown": "Great ! thank you for your answer !"
        }
      ]
    },
    {
      "id": 1697464,
      "postDate": "2022-02-19T16:33:38.547Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> , thanks a lot for hosting and for the explanations!</p>\n<p>I have a question to the data:<br>\nIn the data we have a table of transactions. Those are in principle sold items. However, this article from 2019 shows that that up to 30% - 40% of cloths and shoes bought online are being returned.</p>\n<p>Is the transactions table already cleaned from the returns?</p>",
      "rawMarkdown": "Hello @fridarimark , thanks a lot for hosting and for the explanations!\n\nI have a question to the data:\nIn the data we have a table of transactions. Those are in principle sold items. However, this article from 2019 shows that that up to 30% - 40% of cloths and shoes bought online are being returned.\n\nIs the transactions table already cleaned from the returns?",
      "votes": 1,
      "replies": [
        {
          "id": 1699535,
          "postDate": "2022-02-21T08:50:31.243Z",
          "content": "<p>The transaction table holds all transactions that happened whether returned later or not. </p>",
          "rawMarkdown": "The transaction table holds all transactions that happened whether returned later or not. "
        },
        {
          "id": 1700371,
          "postDate": "2022-02-21T22:20:27.130Z",
          "content": "<p>Thanks a lot for the answer, <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> !</p>\n<p>Could you also explain why in the data one can find 9699 customers which have no matched transaction? I thought first that those can be customers which have returned the articles. But if not, how did those people appeared in the data base of H&amp;M without having made an order?</p>",
          "rawMarkdown": "Thanks a lot for the answer, @fridarimark !\n\nCould you also explain why in the data one can find 9699 customers which have no matched transaction? I thought first that those can be customers which have returned the articles. But if not, how did those people appeared in the data base of H&M without having made an order?"
        },
        {
          "id": 1712244,
          "postDate": "2022-03-04T18:02:13.597Z",
          "content": "<p><a href=\"https://www.kaggle.com/ikorol\" target=\"_blank\">@ikorol</a> I suggest looking up the \"cold start problem\" for recommender systems. Those customers haven't made purchases yet (perhaps they just registered an account on the website). Finding a good recommendation for them is part of the challenge.</p>",
          "rawMarkdown": "@ikorol I suggest looking up the \"cold start problem\" for recommender systems. Those customers haven't made purchases yet (perhaps they just registered an account on the website). Finding a good recommendation for them is part of the challenge.",
          "votes": 2
        },
        {
          "id": 1719498,
          "postDate": "2022-03-11T20:20:31.167Z",
          "content": "<p>How can we distinguish returns from purchases? on the evaluation metric the returns are going to be measured as 1 or 0? Thx</p>",
          "rawMarkdown": "How can we distinguish returns from purchases? on the evaluation metric the returns are going to be measured as 1 or 0? Thx"
        },
        {
          "id": 1719749,
          "postDate": "2022-03-12T04:42:57.477Z",
          "content": "<p>I think what they are saying is that for the target week, there are items which could either be returned later or could be bought as an exchange for an item bought earlier. But as far as evaluation metric goes, both the cases don't make a difference. </p>",
          "rawMarkdown": "I think what they are saying is that for the target week, there are items which could either be returned later or could be bought as an exchange for an item bought earlier. But as far as evaluation metric goes, both the cases don't make a difference. "
        }
      ]
    },
    {
      "id": 1686307,
      "postDate": "2022-02-11T23:41:59.097Z",
      "content": "<p>I'm going to be entering in my first contest! I'm very excited. However I cannot understand from the rules if the data is allowed to be public or not. Can I post my design to Tableau Public, LinkedIn and social media? Thank you.</p>",
      "rawMarkdown": "I'm going to be entering in my first contest! I'm very excited. However I cannot understand from the rules if the data is allowed to be public or not. Can I post my design to Tableau Public, LinkedIn and social media? Thank you.",
      "votes": 1
    },
    {
      "id": 1684967,
      "postDate": "2022-02-10T22:30:53.130Z",
      "content": "<p>Thank you for hosting the competition. We would like to participate!</p>",
      "rawMarkdown": "Thank you for hosting the competition. We would like to participate!",
      "votes": 1
    },
    {
      "id": 1684625,
      "postDate": "2022-02-10T15:57:49.553Z",
      "content": "<p>Thank you for hosting this interesting competition!! Really exciting what can I do with those Data.</p>",
      "rawMarkdown": "Thank you for hosting this interesting competition!! Really exciting what can I do with those Data.",
      "votes": 1
    },
    {
      "id": 1683633,
      "postDate": "2022-02-09T23:54:51.720Z",
      "content": "<p>Thank you for hosting this interesting competition. Is it possible to have some clarifications on the definitions of some provided features, such as FN in customers.csv? I can guess most of the definitions from their names but there are a few that I am not sure of.</p>",
      "rawMarkdown": "Thank you for hosting this interesting competition. Is it possible to have some clarifications on the definitions of some provided features, such as FN in customers.csv? I can guess most of the definitions from their names but there are a few that I am not sure of.",
      "votes": 1
    },
    {
      "id": 1680570,
      "postDate": "2022-02-07T22:47:43.500Z",
      "content": "<p>Hello from CJ. :D<br>\nWell done!</p>",
      "rawMarkdown": "Hello from CJ. :D\nWell done!",
      "votes": 1
    },
    {
      "id": 1680527,
      "postDate": "2022-02-07T21:49:53.213Z",
      "content": "<p>Thank you for hosting this competition. I am looking forward to taking part in it.  </p>",
      "rawMarkdown": "Thank you for hosting this competition. I am looking forward to taking part in it.  ",
      "votes": 1
    },
    {
      "id": 1688999,
      "postDate": "2022-02-14T00:42:27.243Z",
      "content": "<p>Thank you for organizing this super interesting competition, <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a>!</p>\n<p>Could I please ask how the ground truth was generated?</p>\n<p>If a customer buys items in the following order: <code>2, 1, 1, 3</code>, would they appear in the test set in this order?</p>\n<p>Meaning, item <code>1</code> would not be pushed to rank 1 (<code>1, 2, 3</code>) despite multiple purchases of it being made? And it would not get aggregated, we would still have 2 entries for item <code>1</code>?</p>\n<p>Essentially, is the ground truth a list of items that appears in the chronological order in which they were bought? Or is some ranking applied (such as for instance by value?)</p>\n<p>Apologies if I missed this information somewhere else and thank you very much for your time!</p>",
      "rawMarkdown": "Thank you for organizing this super interesting competition, @fridarimark!\n\nCould I please ask how the ground truth was generated?\n\nIf a customer buys items in the following order: `2, 1, 1, 3`, would they appear in the test set in this order?\n\nMeaning, item `1` would not be pushed to rank 1 (`1, 2, 3`) despite multiple purchases of it being made? And it would not get aggregated, we would still have 2 entries for item `1`?\n\nEssentially, is the ground truth a list of items that appears in the chronological order in which they were bought? Or is some ranking applied (such as for instance by value?)\n\nApologies if I missed this information somewhere else and thank you very much for your time!",
      "votes": 2,
      "replies": [
        {
          "id": 1691701,
          "postDate": "2022-02-15T15:01:03.197Z",
          "content": "<p>This question have been answered here: <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307202\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307202</a></p>",
          "rawMarkdown": " This question have been answered here: https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307202",
          "votes": 1
        },
        {
          "id": 1699191,
          "postDate": "2022-02-21T02:17:19.077Z",
          "content": "<p>Thank you very much, I assumed the metric was doing something else, now it all makes sense :)</p>",
          "rawMarkdown": "Thank you very much, I assumed the metric was doing something else, now it all makes sense :)"
        },
        {
          "id": 1699651,
          "postDate": "2022-02-21T10:38:56.597Z",
          "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> could I please ask you if this is the implementation of the metric that will be used in the competition?</p>\n<p><a href=\"https://raw.githubusercontent.com/benhamner/Metrics/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">link to Ben Hamner's repo</a></p>",
          "rawMarkdown": "@fridarimark could I please ask you if this is the implementation of the metric that will be used in the competition?\n\n[link to Ben Hamner's repo](https://raw.githubusercontent.com/benhamner/Metrics/master/Python/ml_metrics/average_precision.py)\n"
        }
      ]
    },
    {
      "id": 1683608,
      "postDate": "2022-02-09T22:54:18.787Z",
      "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> - I have a question I hope I didn't miss it anywhere. Should we consider predicting items that a user has <strong>NOT</strong> purchased before? Or both items purchased and not purchased? I feel typically the idea is to show the user new content but they can certainly re-purchase the same item 🤔</p>",
      "rawMarkdown": "@fridarimark - I have a question I hope I didn't miss it anywhere. Should we consider predicting items that a user has **NOT** purchased before? Or both items purchased and not purchased? I feel typically the idea is to show the user new content but they can certainly re-purchase the same item 🤔",
      "votes": 2,
      "replies": [
        {
          "id": 1684468,
          "postDate": "2022-02-10T14:08:20.433Z",
          "content": "<p>For this competition we don't require that it should be new content that you provide as recommendations. Therefore you may recommend items that the customer already has bought.</p>",
          "rawMarkdown": "For this competition we don't require that it should be new content that you provide as recommendations. Therefore you may recommend items that the customer already has bought.",
          "votes": 7
        }
      ]
    },
    {
      "id": 1680491,
      "postDate": "2022-02-07T21:05:07.060Z",
      "content": "<p>This is an interesting competition.</p>",
      "rawMarkdown": "This is an interesting competition.",
      "votes": 2
    },
    {
      "id": 1798773,
      "postDate": "2022-05-23T11:08:03.623Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> </p>\n<p>Since the competition is closed, would it be ok to use the dataset for a blog post and demo application?</p>\n<p>KR</p>",
      "rawMarkdown": "Hi @fridarimark \n\nSince the competition is closed, would it be ok to use the dataset for a blog post and demo application?\n\nKR"
    },
    {
      "id": 1785450,
      "postDate": "2022-05-12T05:38:27.270Z",
      "content": "<p>Interesting </p>",
      "rawMarkdown": "Interesting "
    },
    {
      "id": 1769471,
      "postDate": "2022-04-27T09:41:50.593Z",
      "content": "<p>Thank you for hosting the competition!<br>\nThat’s very interesting, I can learn a lot.</p>",
      "rawMarkdown": "Thank you for hosting the competition!\nThat’s very interesting, I can learn a lot."
    },
    {
      "id": 1766324,
      "postDate": "2022-04-24T12:46:15.540Z",
      "content": "<p>This is a interesting competition, thanks for your hosting!</p>",
      "rawMarkdown": "This is a interesting competition, thanks for your hosting!"
    },
    {
      "id": 1757886,
      "postDate": "2022-04-17T05:36:47.847Z",
      "content": "<p>great！！！！！</p>",
      "rawMarkdown": "great！！！！！"
    },
    {
      "id": 1739277,
      "postDate": "2022-03-29T23:23:06.787Z",
      "content": "<p>Thanks for hosting interesting competition!</p>",
      "rawMarkdown": "Thanks for hosting interesting competition!"
    },
    {
      "id": 1735433,
      "postDate": "2022-03-26T08:54:08.703Z",
      "content": "<p>Thanks for hosting interesting competition!</p>",
      "rawMarkdown": "Thanks for hosting interesting competition!"
    },
    {
      "id": 1729393,
      "postDate": "2022-03-20T02:39:41.113Z",
      "content": "<p>Thank you very much for an awesome competition.</p>",
      "rawMarkdown": "Thank you very much for an awesome competition."
    },
    {
      "id": 1719401,
      "postDate": "2022-03-11T18:05:53.953Z",
      "content": "<p>Thanks for the competition! Looking forward to seeing all the different approaches!</p>",
      "rawMarkdown": "Thanks for the competition! Looking forward to seeing all the different approaches!"
    },
    {
      "id": 1705857,
      "postDate": "2022-02-26T22:20:49.043Z",
      "content": "<p>Thank you for hosting the competition. I was wondering if we should use the actual image and use deep learning to process the image data? or should we just first focus on the raw dataset? </p>",
      "rawMarkdown": "Thank you for hosting the competition. I was wondering if we should use the actual image and use deep learning to process the image data? or should we just first focus on the raw dataset? "
    },
    {
      "id": 1700427,
      "postDate": "2022-02-22T00:47:05.043Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> , thanks for hosting this interesting competition.</p>\n<p>I have a question to the data:<br>\nThe transaction table record all transaction from Sep2018 to Sep2020. However, those article_id should change/sold-out every season/month/week. So should I limited my recommendation to the article_id that sell in the Sep2020 only? </p>",
      "rawMarkdown": "Hello @fridarimark , thanks for hosting this interesting competition.\n\nI have a question to the data:\nThe transaction table record all transaction from Sep2018 to Sep2020. However, those article_id should change/sold-out every season/month/week. So should I limited my recommendation to the article_id that sell in the Sep2020 only? "
    },
    {
      "id": 1700169,
      "postDate": "2022-02-21T18:03:18.237Z",
      "content": "<p>Thank you for hosting the competition! 😊</p>",
      "rawMarkdown": "Thank you for hosting the competition! 😊"
    },
    {
      "id": 1698067,
      "postDate": "2022-02-20T05:47:12.700Z",
      "content": "<p>Nice Competition!</p>",
      "rawMarkdown": "Nice Competition!"
    },
    {
      "id": 1692519,
      "postDate": "2022-02-16T05:00:13.397Z",
      "content": "<p>Thanks for hosting such a good competition</p>",
      "rawMarkdown": "Thanks for hosting such a good competition"
    },
    {
      "id": 1686980,
      "postDate": "2022-02-12T13:44:56.117Z",
      "content": "<p>Looks interesting!</p>",
      "rawMarkdown": "Looks interesting!"
    },
    {
      "id": 1686482,
      "postDate": "2022-02-12T05:37:56.677Z",
      "content": "<p>Thank you for this wonderful opportunity!</p>",
      "rawMarkdown": "Thank you for this wonderful opportunity!"
    },
    {
      "id": 1685401,
      "postDate": "2022-02-11T08:42:04.177Z",
      "content": "<p>Seems fun , can't wait to start working on it 🤩</p>",
      "rawMarkdown": "Seems fun , can't wait to start working on it 🤩"
    },
    {
      "id": 3480848,
      "postDate": "2026-06-25T09:43:01.647Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1784585,
      "postDate": "2022-05-11T09:32:38.493Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1690157,
      "postDate": "2022-02-14T18:33:22.770Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1686903,
      "postDate": "2022-02-12T12:35:17.900Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    }
  ],
  "comments": [
    {
      "id": 1711192,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-03-03T17:39:02.117000",
      "content": "<p>Hi, thanks for a great comp <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a>  FYI, your metric formula is wrong. It should say</p>\n<p>$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U} \\left( \\frac{1}{min(m,12)} \\sum_{k=1}^{min(n,12)}{P(k) \\times rel(k)} \\right) $$</p>\n<p>where <code>m</code> is the number of ground truths per customer, <code>n</code> is the number of predictions per customer, and <code>U</code> is the number of customers. <code>P(k)</code> is the precision at cutoff <code>k</code>, and <code>rel(k)</code> is an indicator function equaling 1 if the item at rank <code>k</code> is a relevant (correct) label, zero otherwise.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 1747078,
          "author_name": "Devansh Chowdhury",
          "author_url": "",
          "post_date": "2022-04-06T10:21:02.670000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I have a small doubt. </p>\n<ol>\n<li>should Articles sub be made in an ordered way ?</li>\n<li>If truth value is 2 0 1 and i make the sub 3 1 , what would be the score?</li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1683754,
      "author_name": "Tyrell",
      "author_url": "",
      "post_date": "2022-02-10T03:19:15.497000",
      "content": "<p>Question about the variables:<br>\n FN, Active - What are these?<br>\nSales channel ID - is is right to assume 2 is in person and 1 is online?<br>\nPostal code - 1.2 mil is a lot of postal codes (approx every country you sell in) is there some else in these hashes?<br>\nwhat percent of the customers are new in the validation set (any idea helps)?<br>\nThanks</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1684481,
          "author_name": "FridaRim",
          "author_url": "",
          "post_date": "2022-02-10T14:21:20.813000",
          "content": "<p>FN is if a customer get Fashion News newsletter, Active is if the customer is active for communication, sales channel id, 2 is online and 1 store.</p>",
          "votes": 26,
          "replies": []
        }
      ]
    },
    {
      "id": 1690119,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2022-02-14T17:53:06.300000",
      "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> </p>\n<blockquote>\n  <p>Customer that did not make any purchase during test period are excluded from the scoring</p>\n</blockquote>\n<p>So, we don't need to predict those who didn't purchase any items. <strong>BUT</strong> sample_submission has 1371980 customers, and number of customers who bought items as of last one week is 68984. This means almost 95% of our prediction won't be scored during test period.<br>\nIMHO this is totally nonsense, and we should skip those for saving our time. What do you think? If yes, please update sample_submission or give us the customer_id who are being scored.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1691807,
          "author_name": "jl150",
          "author_url": "",
          "post_date": "2022-02-15T16:09:09.410000",
          "content": "<p>I have to disagree on several points. </p>\n<p>First, everything is scored in the sense that the MAP metric will take into account all your predictions. It just happens that some terms in that expression will be weighted 0. The whole challenge here is to design a system that will perform well in the circumstance where we don't know which weights will be 1s.</p>\n<p>Second, if we knew which people made purchases in the test week it would probably amount to considerable data leakage. I'm sure kagglers would find a way to exploit this to optimize their models, but it would make those models much less useful in the realistic scenario where H&amp;M doesn't know who will make purchases next week. Again, that would defeat the purpose of the challenge.</p>\n<p>Third, the use-case for this sort of thing is to generate recommendations for articles to users browsing the website or receiving emails. So you need to be able to generate recommendations for everyone, even if you don't know who will actually follow through and make a purchase.</p>\n<p>Finally, there really aren't that many customers (relatively speaking). I'm not sure what sort of time you are concerned about, but once you have an algorithm that can make predictions for 70k customers, surely it's not that much more costly in terms of time to apply it to all 1.3 M customers…</p>\n<p>Personally, I really like this challenge because the number of publicly available, realistic, managebly sized datasets for training recommender systems with a time series component (such as this one) is quite small, at least to my knowledge.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1691875,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2022-02-15T16:55:45.660000",
          "content": "<blockquote>\n  <p>The whole challenge here is to design a system that will perform well in the circumstance where we don't know which weights will be 1s</p>\n</blockquote>\n<p>Yes, I think so. But this is not the objection for unnecessary of dummy data.</p>\n<blockquote>\n  <p>if we knew which people made purchases in the test week it would probably amount to considerable data leakage</p>\n</blockquote>\n<p>What kind of leakage? We are making predictions based on the assumption that all customers will buy something. So, it's totally same thing and our prediction won't change if we know which customer would buy next week.</p>\n<blockquote>\n  <p>you need to be able to generate recommendations for everyone, even if you don't know who will actually follow through and make a purchase.</p>\n</blockquote>\n<p>Yes, I think so. But this is not the objection for unnecessary of dummy data.</p>\n<blockquote>\n  <p>but once you have an algorithm that can make predictions for 70k customers, surely it's not that much more costly in terms of time to apply it to all 1.3 M customers</p>\n</blockquote>\n<p>At least, my DGX which has 250G memory can't inference 1.3 M customers at once.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1692061,
          "author_name": "jl150",
          "author_url": "",
          "post_date": "2022-02-15T19:20:05.380000",
          "content": "<p>I suppose then the challenge for you is to design a more scalable algorithm… I would suggest doing inference in batches… Impossible to know what sort of algorithm you are running, but, e.g.,  scaling matrix factorization based CF methods to large numbers of users x items is a pretty well documented thing. In this case there isn't even <em>that many</em> users x items, just enough that making a matrix of that dimension will crash most reasonably sized systems, so getting around it with sparse representations and batched inference is probably the way to go (at least, it works for me).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1692063,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2022-02-15T19:25:43.897000",
          "content": "<blockquote>\n  <p>I suppose then the challenge for you is to design a more scalable algorithm…</p>\n</blockquote>\n<p>I can agree on that, but is that the host's angle as well? If so why not notebook competition?<br>\nIf not, no need to waste our time.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1693558,
          "author_name": "Alaa Chahboune",
          "author_url": "",
          "post_date": "2022-02-16T18:31:25.510000",
          "content": "<p>I don't think so because it will bias the predictions</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1693749,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2022-02-16T22:50:03.600000",
          "content": "<p><a href=\"https://www.kaggle.com/achahboune\" target=\"_blank\">@achahboune</a> What do you mean? How do we evaluate the dummies which won't be evaluated?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1705892,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-02-27T00:28:27.423000",
          "content": "<p><a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>  </p>\n<p>There are 3 parts of the task</p>\n<ol>\n<li>Which Customers will buy?</li>\n<li>Which Articles will be sold?</li>\n<li>Which Customers will buy which articles?</li>\n</ol>\n<p>This is basically building a recommendation system based on historical data. You know who actually purchased what, so it allows you to test whether your recommendations, if they had been made, would have relevant to those customers. Now, nobody really buys everything that is recommended so even if the same recommendations were made, the same purchases may not have been made.</p>\n<p>The premise here is such a recommendation system may work for all customers for all scenarios. It may not. It may need to be further tested and fine tuned.</p>\n<p>If the host gives you No.1, then the problem becomes less complex. It also becomes more of a prediction problem which means that given a set of Customers, you may predict what they may buy which limits the real world use of it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1729766,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "2022-03-20T13:40:40.027000",
          "content": "<p>No 1 is not included in the metric since those customers are excluded from the leaderboard. Onodera is right, including them in the submission file is a waste of compute power (and thus money + co2).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1708321,
      "author_name": "DmitryS",
      "author_url": "",
      "post_date": "2022-03-01T11:11:24.590000",
      "content": "<p>Does Kaggle have official position about russian citizens' participation? Will they be able to earn medals and ranking points?</p>\n<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> </p>",
      "votes": 8,
      "replies": [
        {
          "id": 1717941,
          "author_name": "DmitryS",
          "author_url": "",
          "post_date": "2022-03-10T11:20:22.477000",
          "content": "<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> may you comment it, please?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1723056,
          "author_name": "Maggie",
          "author_url": "",
          "post_date": "2022-03-15T04:52:54.833000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/simakov\" target=\"_blank\">@simakov</a> Kaggle is available to Russian citizens for participation in competitions and for medals, ranking points, etc. The OFAC rules have been updated to exclude participants from disputed areas, but Russian citizens who do not live in these areas can continue to participate. </p>\n<p>\"Competitions are open to residents of the United States and worldwide, except that if you are a resident of Crimea, so-called Donetsk People's Republic (DNR) or Luhansk People's Republic (LNR), Cuba, Iran, Syria, or North Korea, or are subject to U.S. export controls or sanctions, you may not enter the Competition.\"</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1690329,
      "author_name": "Paul Larmuseau",
      "author_url": "",
      "post_date": "2022-02-14T21:40:52.207000",
      "content": "<p>Actually what is missing is the 'available stock' since we don't know that stock, its impossible your site will show the stock that is not avaiable, and you suggest we forecast what they want to purchase, you will endup with a wrong forecast of the historical stock. Or you we should make a guess about the available stock as the last 6month purchases</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1687970,
      "author_name": "jcesquivel",
      "author_url": "",
      "post_date": "2022-02-13T09:21:16.153000",
      "content": "<p>Thanks for hosting this interesting competition.</p>\n<p>I would like to ask a question. If I'm not wrong, <b>there are several product names that have several product codes, and also product codes that are shared by various product names.</b> For example, the product name 'Rose dress' has associated product codes: 552212, 673917, 829601, 892278, 920394. At the same time, the product code 600229 has theses product names associated: 'CAGE SLIM EASY IRON US', 'Cage US', 'Cage easy iron',  'Cage Slim Easy Iron'. Could you explain the difference and which one is important?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1700309,
          "author_name": "Melanie7744",
          "author_url": "",
          "post_date": "2022-02-21T20:22:58.057000",
          "content": "<p>One possible explanation would be that 5 different dresses that classify as 'Rose dress' have been sold. Like one with big roses, one with small roses, one just having the color rose…</p>\n<p>For the second case, it might be human error when entering the product name. All product names sound somewhat similar. So maybe someone did not check for already existing product names or different countries selling this item have different naming rules.</p>\n<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a>, can you confirm?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1705898,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-02-27T00:34:39.910000",
          "content": "<p>Overall, for the purpose of this competition, depending on the numbers of such issues and how you use the relevant data in your work, it shouldn't matter much. Of course, ideally, H&amp;M should clean up their data. May be they are listening in 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1754746,
      "author_name": "Ryota",
      "author_url": "",
      "post_date": "2022-04-14T00:54:49.320000",
      "content": "<p>Are the articles that had never been sold in transaction_train.csv and are included in article.csv new release in test data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1755770,
          "author_name": "torukisaki",
          "author_url": "",
          "post_date": "2022-04-15T00:52:22.817000",
          "content": "<p>Thank you for the good information, Ryota-san!<br>\nIf I were you, I would investigate the relation between the size of the 10-digit codes and the release date which means the first date appeared in the train.csv.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1755828,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2022-04-15T03:15:44.230000",
          "content": "<p>Thank you for advice!!<br>\nAs you said, I analyzed between article_id and the first day each articles were sold. As a result, articles that has large id number tend to be sold for first time lately.<br>\nBut there are exception. I think maybe there are articles that was launched a long time ago but only recently sold. </p>\n<p>I will try to analyze more!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1702205,
      "author_name": "Yukou Takahashi",
      "author_url": "",
      "post_date": "2022-02-23T12:25:45.853000",
      "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> <br>\n  Thank you for organizing this exciting competition!<br>\n  I am very glad to join!<br>\n  I would like to ask one question.<br>\n  I would like to know the criteria for <strong><code>rel(k)</code></strong> to take <strong><code>1</code></strong> value. (<strong><code>rel(k)</code></strong> appears in the Mean Average Precision formula)<br>\n$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k) $$<br>\n<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/overview/evaluation\" target=\"_blank\">Evaluation</a></p>\n<p>You say that <strong><code>rel(k)</code></strong> will be <strong><code>1</code></strong> if the predicted item is a <strong>related label</strong>, but what is the condition for related?</p>\n<ul>\n<li><p><code>article_id</code> matches (most strict?)</p></li>\n<li><p><code>product_code</code> matches</p></li>\n<li><p><code>section_no</code> matches</p></li>\n<li><p>etc…</p>\n<p>Thank you!</p></li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1704224,
          "author_name": "FridaRim",
          "author_url": "",
          "post_date": "2022-02-25T10:39:38.530000",
          "content": "<p>Relevant label means correct label here, so it needs to match the article id that was bought</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1704491,
          "author_name": "Yukou Takahashi",
          "author_url": "",
          "post_date": "2022-02-25T15:15:22.860000",
          "content": "<p>Thank you for your answer !! 😊</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1705130,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-26T06:54:13.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1750907,
          "author_name": "Yesmine Makkes",
          "author_url": "",
          "post_date": "2022-04-10T07:57:42.637000",
          "content": "<p>Great ! thank you for your answer !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1697464,
      "author_name": "ikorol",
      "author_url": "",
      "post_date": "2022-02-19T16:33:38.547000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> , thanks a lot for hosting and for the explanations!</p>\n<p>I have a question to the data:<br>\nIn the data we have a table of transactions. Those are in principle sold items. However, this article from 2019 shows that that up to 30% - 40% of cloths and shoes bought online are being returned.</p>\n<p>Is the transactions table already cleaned from the returns?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1699535,
          "author_name": "FridaRim",
          "author_url": "",
          "post_date": "2022-02-21T08:50:31.243000",
          "content": "<p>The transaction table holds all transactions that happened whether returned later or not. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1700371,
          "author_name": "ikorol",
          "author_url": "",
          "post_date": "2022-02-21T22:20:27.130000",
          "content": "<p>Thanks a lot for the answer, <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> !</p>\n<p>Could you also explain why in the data one can find 9699 customers which have no matched transaction? I thought first that those can be customers which have returned the articles. But if not, how did those people appeared in the data base of H&amp;M without having made an order?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1712244,
          "author_name": "jl150",
          "author_url": "",
          "post_date": "2022-03-04T18:02:13.597000",
          "content": "<p><a href=\"https://www.kaggle.com/ikorol\" target=\"_blank\">@ikorol</a> I suggest looking up the \"cold start problem\" for recommender systems. Those customers haven't made purchases yet (perhaps they just registered an account on the website). Finding a good recommendation for them is part of the challenge.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1719498,
          "author_name": "EnricRovira",
          "author_url": "",
          "post_date": "2022-03-11T20:20:31.167000",
          "content": "<p>How can we distinguish returns from purchases? on the evaluation metric the returns are going to be measured as 1 or 0? Thx</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1719749,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-03-12T04:42:57.477000",
          "content": "<p>I think what they are saying is that for the target week, there are items which could either be returned later or could be bought as an exchange for an item bought earlier. But as far as evaluation metric goes, both the cases don't make a difference. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1686307,
      "author_name": "Dan Fulton",
      "author_url": "",
      "post_date": "2022-02-11T23:41:59.097000",
      "content": "<p>I'm going to be entering in my first contest! I'm very excited. However I cannot understand from the rules if the data is allowed to be public or not. Can I post my design to Tableau Public, LinkedIn and social media? Thank you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1684967,
      "author_name": "N. Peker Çelik",
      "author_url": "",
      "post_date": "2022-02-10T22:30:53.130000",
      "content": "<p>Thank you for hosting the competition. We would like to participate!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1684625,
      "author_name": "AVE",
      "author_url": "",
      "post_date": "2022-02-10T15:57:49.553000",
      "content": "<p>Thank you for hosting this interesting competition!! Really exciting what can I do with those Data.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1683633,
      "author_name": "Luca Loo",
      "author_url": "",
      "post_date": "2022-02-09T23:54:51.720000",
      "content": "<p>Thank you for hosting this interesting competition. Is it possible to have some clarifications on the definitions of some provided features, such as FN in customers.csv? I can guess most of the definitions from their names but there are a few that I am not sure of.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1680570,
      "author_name": "Chengjun JIN",
      "author_url": "",
      "post_date": "2022-02-07T22:47:43.500000",
      "content": "<p>Hello from CJ. :D<br>\nWell done!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1680527,
      "author_name": "wti 200",
      "author_url": "",
      "post_date": "2022-02-07T21:49:53.213000",
      "content": "<p>Thank you for hosting this competition. I am looking forward to taking part in it.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1688999,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-02-14T00:42:27.243000",
      "content": "<p>Thank you for organizing this super interesting competition, <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a>!</p>\n<p>Could I please ask how the ground truth was generated?</p>\n<p>If a customer buys items in the following order: <code>2, 1, 1, 3</code>, would they appear in the test set in this order?</p>\n<p>Meaning, item <code>1</code> would not be pushed to rank 1 (<code>1, 2, 3</code>) despite multiple purchases of it being made? And it would not get aggregated, we would still have 2 entries for item <code>1</code>?</p>\n<p>Essentially, is the ground truth a list of items that appears in the chronological order in which they were bought? Or is some ranking applied (such as for instance by value?)</p>\n<p>Apologies if I missed this information somewhere else and thank you very much for your time!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1691701,
          "author_name": "FridaRim",
          "author_url": "",
          "post_date": "2022-02-15T15:01:03.197000",
          "content": "<p>This question have been answered here: <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307202\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307202</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1699191,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-02-21T02:17:19.077000",
          "content": "<p>Thank you very much, I assumed the metric was doing something else, now it all makes sense :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1699651,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-02-21T10:38:56.597000",
          "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> could I please ask you if this is the implementation of the metric that will be used in the competition?</p>\n<p><a href=\"https://raw.githubusercontent.com/benhamner/Metrics/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">link to Ben Hamner's repo</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1683608,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2022-02-09T22:54:18.787000",
      "content": "<p><a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> - I have a question I hope I didn't miss it anywhere. Should we consider predicting items that a user has <strong>NOT</strong> purchased before? Or both items purchased and not purchased? I feel typically the idea is to show the user new content but they can certainly re-purchase the same item 🤔</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1684468,
          "author_name": "FridaRim",
          "author_url": "",
          "post_date": "2022-02-10T14:08:20.433000",
          "content": "<p>For this competition we don't require that it should be new content that you provide as recommendations. Therefore you may recommend items that the customer already has bought.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1680491,
      "author_name": "Jie Wu",
      "author_url": "",
      "post_date": "2022-02-07T21:05:07.060000",
      "content": "<p>This is an interesting competition.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1798773,
      "author_name": "MartinMarenz",
      "author_url": "",
      "post_date": "2022-05-23T11:08:03.623000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> </p>\n<p>Since the competition is closed, would it be ok to use the dataset for a blog post and demo application?</p>\n<p>KR</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1785450,
      "author_name": "Spandan Samal",
      "author_url": "",
      "post_date": "2022-05-12T05:38:27.270000",
      "content": "<p>Interesting </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1769471,
      "author_name": "tatatataky",
      "author_url": "",
      "post_date": "2022-04-27T09:41:50.593000",
      "content": "<p>Thank you for hosting the competition!<br>\nThat’s very interesting, I can learn a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1766324,
      "author_name": "Zhicheng Wang",
      "author_url": "",
      "post_date": "2022-04-24T12:46:15.540000",
      "content": "<p>This is a interesting competition, thanks for your hosting!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1757886,
      "author_name": "NWPUCS",
      "author_url": "",
      "post_date": "2022-04-17T05:36:47.847000",
      "content": "<p>great！！！！！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1739277,
      "author_name": "kazuki morohashi",
      "author_url": "",
      "post_date": "2022-03-29T23:23:06.787000",
      "content": "<p>Thanks for hosting interesting competition!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1735433,
      "author_name": "potatomonkey",
      "author_url": "",
      "post_date": "2022-03-26T08:54:08.703000",
      "content": "<p>Thanks for hosting interesting competition!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1729393,
      "author_name": "torukisaki",
      "author_url": "",
      "post_date": "2022-03-20T02:39:41.113000",
      "content": "<p>Thank you very much for an awesome competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1719401,
      "author_name": "heuristic_morse",
      "author_url": "",
      "post_date": "2022-03-11T18:05:53.953000",
      "content": "<p>Thanks for the competition! Looking forward to seeing all the different approaches!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1705857,
      "author_name": "BM",
      "author_url": "",
      "post_date": "2022-02-26T22:20:49.043000",
      "content": "<p>Thank you for hosting the competition. I was wondering if we should use the actual image and use deep learning to process the image data? or should we just first focus on the raw dataset? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1700427,
      "author_name": "melody",
      "author_url": "",
      "post_date": "2022-02-22T00:47:05.043000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fridarimark\" target=\"_blank\">@fridarimark</a> , thanks for hosting this interesting competition.</p>\n<p>I have a question to the data:<br>\nThe transaction table record all transaction from Sep2018 to Sep2020. However, those article_id should change/sold-out every season/month/week. So should I limited my recommendation to the article_id that sell in the Sep2020 only? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1700169,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-21T18:03:18.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1698067,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-20T05:47:12.700000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1692519,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-16T05:00:13.397000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1686980,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-12T13:44:56.117000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1686482,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-12T05:37:56.677000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1685401,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-11T08:42:04.177000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3480848,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-06-25T09:43:01.647000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1784585,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-11T09:32:38.493000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1690157,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-14T18:33:22.770000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1686903,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-12T12:35:17.900000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1680156": "Dear Kagglers,\n\nOn behalf of H&M Group and AI Recommendation Engine team I would like to officially welcome you to the “H&M Personalized Fashion Recommendations” competition!\n\nIn this competition we invite you to produce product recommendations based on data from historical transactions, as well as from customer and product meta data. The available meta data spans from simple data, such as garment type and customer age, to text data from product descriptions, to image data from garment images.\n\nWe can’t wait to see what model approaches or feature engineering techniques the community will use to arrive at the best possible solution. Along with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you develop the best models possible!\n\nGood luck, and have fun!",
    "1711192": "Hi, thanks for a great comp @fridarimark @maggiemd  FYI, your metric formula is wrong. It should say\n\n$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U} \\left( \\frac{1}{min(m,12)} \\sum_{k=1}^{min(n,12)}{P(k) \\times rel(k)} \\right) $$\n\nwhere `m` is the number of ground truths per customer, `n` is the number of predictions per customer, and `U` is the number of customers. `P(k)` is the precision at cutoff `k`, and `rel(k)` is an indicator function equaling 1 if the item at rank `k` is a relevant (correct) label, zero otherwise.",
    "1683754": "Question about the variables:\n FN, Active - What are these?\nSales channel ID - is is right to assume 2 is in person and 1 is online?\nPostal code - 1.2 mil is a lot of postal codes (approx every country you sell in) is there some else in these hashes?\nwhat percent of the customers are new in the validation set (any idea helps)?\nThanks",
    "1690119": "@fridarimark \n> Customer that did not make any purchase during test period are excluded from the scoring\n\nSo, we don't need to predict those who didn't purchase any items. **BUT** sample_submission has 1371980 customers, and number of customers who bought items as of last one week is 68984. This means almost 95% of our prediction won't be scored during test period.\nIMHO this is totally nonsense, and we should skip those for saving our time. What do you think? If yes, please update sample_submission or give us the customer_id who are being scored.",
    "1708321": "Does Kaggle have official position about russian citizens' participation? Will they be able to earn medals and ranking points?\n\n@maggiemd @fridarimark ",
    "1690329": "Actually what is missing is the 'available stock' since we don't know that stock, its impossible your site will show the stock that is not avaiable, and you suggest we forecast what they want to purchase, you will endup with a wrong forecast of the historical stock. Or you we should make a guess about the available stock as the last 6month purchases",
    "1687970": "Thanks for hosting this interesting competition.\n\nI would like to ask a question. If I'm not wrong, <b>there are several product names that have several product codes, and also product codes that are shared by various product names.</b> For example, the product name 'Rose dress' has associated product codes: 552212, 673917, 829601, 892278, 920394. At the same time, the product code 600229 has theses product names associated: 'CAGE SLIM EASY IRON US', 'Cage US', 'Cage easy iron',  'Cage Slim Easy Iron'. Could you explain the difference and which one is important?",
    "1754746": "Are the articles that had never been sold in transaction_train.csv and are included in article.csv new release in test data?",
    "1702205": "@fridarimark \n  Thank you for organizing this exciting competition!\n  I am very glad to join!\n  I would like to ask one question.\n  I would like to know the criteria for **`rel(k)`** to take **`1`** value. (**`rel(k)`** appears in the Mean Average Precision formula)\n$$ MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k) $$\n[Evaluation](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/overview/evaluation)\n\n  You say that **`rel(k)`** will be **`1`** if the predicted item is a **related label**, but what is the condition for related?\n  \n  - `article_id` matches (most strict?)\n  - `product_code` matches\n  - `section_no` matches\n  - etc...\n  \n  Thank you!\n",
    "1697464": "Hello @fridarimark , thanks a lot for hosting and for the explanations!\n\nI have a question to the data:\nIn the data we have a table of transactions. Those are in principle sold items. However, this article from 2019 shows that that up to 30% - 40% of cloths and shoes bought online are being returned.\n\nIs the transactions table already cleaned from the returns?",
    "1686307": "I'm going to be entering in my first contest! I'm very excited. However I cannot understand from the rules if the data is allowed to be public or not. Can I post my design to Tableau Public, LinkedIn and social media? Thank you.",
    "1684967": "Thank you for hosting the competition. We would like to participate!",
    "1684625": "Thank you for hosting this interesting competition!! Really exciting what can I do with those Data.",
    "1683633": "Thank you for hosting this interesting competition. Is it possible to have some clarifications on the definitions of some provided features, such as FN in customers.csv? I can guess most of the definitions from their names but there are a few that I am not sure of.",
    "1680570": "Hello from CJ. :D\nWell done!",
    "1680527": "Thank you for hosting this competition. I am looking forward to taking part in it.  ",
    "1688999": "Thank you for organizing this super interesting competition, @fridarimark!\n\nCould I please ask how the ground truth was generated?\n\nIf a customer buys items in the following order: `2, 1, 1, 3`, would they appear in the test set in this order?\n\nMeaning, item `1` would not be pushed to rank 1 (`1, 2, 3`) despite multiple purchases of it being made? And it would not get aggregated, we would still have 2 entries for item `1`?\n\nEssentially, is the ground truth a list of items that appears in the chronological order in which they were bought? Or is some ranking applied (such as for instance by value?)\n\nApologies if I missed this information somewhere else and thank you very much for your time!",
    "1683608": "@fridarimark - I have a question I hope I didn't miss it anywhere. Should we consider predicting items that a user has **NOT** purchased before? Or both items purchased and not purchased? I feel typically the idea is to show the user new content but they can certainly re-purchase the same item 🤔",
    "1680491": "This is an interesting competition.",
    "1798773": "Hi @fridarimark \n\nSince the competition is closed, would it be ok to use the dataset for a blog post and demo application?\n\nKR",
    "1785450": "Interesting ",
    "1769471": "Thank you for hosting the competition!\nThat’s very interesting, I can learn a lot.",
    "1766324": "This is a interesting competition, thanks for your hosting!",
    "1757886": "great！！！！！",
    "1739277": "Thanks for hosting interesting competition!",
    "1735433": "Thanks for hosting interesting competition!",
    "1729393": "Thank you very much for an awesome competition.",
    "1719401": "Thanks for the competition! Looking forward to seeing all the different approaches!",
    "1705857": "Thank you for hosting the competition. I was wondering if we should use the actual image and use deep learning to process the image data? or should we just first focus on the raw dataset? ",
    "1700427": "Hello @fridarimark , thanks for hosting this interesting competition.\n\nI have a question to the data:\nThe transaction table record all transaction from Sep2018 to Sep2020. However, those article_id should change/sold-out every season/month/week. So should I limited my recommendation to the article_id that sell in the Sep2020 only? ",
    "1700169": "Thank you for hosting the competition! 😊",
    "1698067": "Nice Competition!",
    "1692519": "Thanks for hosting such a good competition",
    "1686980": "Looks interesting!",
    "1686482": "Thank you for this wonderful opportunity!",
    "1685401": "Seems fun , can't wait to start working on it 🤩",
    "3480848": "",
    "1784585": "",
    "1690157": "",
    "1686903": "Thank you!"
  }
}