{
  "id": 308402,
  "title": "Question about a zero-target prediction approach",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/308402",
  "author_name": "",
  "post_date": "2022-02-18T13:44:25.822233700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone!<br>\nI would like to get you thoughts on a vary naive approach I am trying to use (and probably some help with debugging.<br>\nI want to take the last day in the available dataset, select the 12 most frequently purchased articles and just send those 12 articles as a prediction for every single customer.<br>\nI ran this approach <a href=\"https://www.kaggle.com/juststas/hm-recsys-first-runs\" target=\"_blank\">here</a>, section \"Modelling attempts\", and got some very weird results.</p>\n<p>When I tested it on a period similar to the current one, but in 2019, I got a map@12 of .347, which can't be right. I don't seem to have any leaks or peeking into the future, but when I use the same approach for 2020, I get a 0.000 submission…</p>\n<p>Any help would be appreciated :)</p>",
  "messages": [
    {
      "id": "1695957",
      "postDate": "02/18/2022 13:44:25",
      "content": "<p>Hi everyone!<br>\nI would like to get you thoughts on a vary naive approach I am trying to use (and probably some help with debugging.<br>\nI want to take the last day in the available dataset, select the 12 most frequently purchased articles and just send those 12 articles as a prediction for every single customer.<br>\nI ran this approach <a href=\"https://www.kaggle.com/juststas/hm-recsys-first-runs\" target=\"_blank\">here</a>, section \"Modelling attempts\", and got some very weird results.</p>\n<p>When I tested it on a period similar to the current one, but in 2019, I got a map@12 of .347, which can't be right. I don't seem to have any leaks or peeking into the future, but when I use the same approach for 2020, I get a 0.000 submission…</p>\n<p>Any help would be appreciated :)</p>",
      "rawMarkdown": "Hi everyone!\nI would like to get you thoughts on a vary naive approach I am trying to use (and probably some help with debugging.\nI want to take the last day in the available dataset, select the 12 most frequently purchased articles and just send those 12 articles as a prediction for every single customer.\nI ran this approach [here](https://www.kaggle.com/juststas/hm-recsys-first-runs), section \"Modelling attempts\", and got some very weird results.\n\nWhen I tested it on a period similar to the current one, but in 2019, I got a map@12 of .347, which can't be right. I don't seem to have any leaks or peeking into the future, but when I use the same approach for 2020, I get a 0.000 submission...\n\nAny help would be appreciated :)",
      "votes": null
    },
    {
      "id": "1696470",
      "postDate": "02/18/2022 21:27:33",
      "content": "<p>In your submission csv, the prediction is a string and each article id must begin with a zero. For example in your notebook your first prediction is <code>924243002 751471001 448509014 918522001 866731...</code>. It needs to be <code>0924243002 0751471001 0448509014 0918522001 0866731...</code> with zeros.</p>",
      "rawMarkdown": "In your submission csv, the prediction is a string and each article id must begin with a zero. For example in your notebook your first prediction is `924243002 751471001 448509014 918522001 866731...`. It needs to be `0924243002 0751471001 0448509014 0918522001 0866731...` with zeros.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1696470,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/18/2022 21:27:33",
      "content": "<p>In your submission csv, the prediction is a string and each article id must begin with a zero. For example in your notebook your first prediction is <code>924243002 751471001 448509014 918522001 866731...</code>. It needs to be <code>0924243002 0751471001 0448509014 0918522001 0866731...</code> with zeros.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1695957": "Hi everyone!\nI would like to get you thoughts on a vary naive approach I am trying to use (and probably some help with debugging.\nI want to take the last day in the available dataset, select the 12 most frequently purchased articles and just send those 12 articles as a prediction for every single customer.\nI ran this approach [here](https://www.kaggle.com/juststas/hm-recsys-first-runs), section \"Modelling attempts\", and got some very weird results.\n\nWhen I tested it on a period similar to the current one, but in 2019, I got a map@12 of .347, which can't be right. I don't seem to have any leaks or peeking into the future, but when I use the same approach for 2020, I get a 0.000 submission...\n\nAny help would be appreciated :)",
    "1696470": "In your submission csv, the prediction is a string and each article id must begin with a zero. For example in your notebook your first prediction is `924243002 751471001 448509014 918522001 866731...`. It needs to be `0924243002 0751471001 0448509014 0918522001 0866731...` with zeros."
  },
  "source": "meta"
}