{
  "id": 313725,
  "title": "Sequence matters?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/313725",
  "author_name": "",
  "post_date": "2022-03-18T15:20:20.717350500Z",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>From what I learned, the sequence of items in prediction vs answer matters a lot in mAPk metric.</p>\n<p>So, how does the organizer order the final answer for each customer? Are they sorted by purchase time, or are they sorted alphabetically by article_id, or some other ways?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1728089",
      "postDate": "03/18/2022 15:20:20",
      "content": "<p>From what I learned, the sequence of items in prediction vs answer matters a lot in mAPk metric.</p>\n<p>So, how does the organizer order the final answer for each customer? Are they sorted by purchase time, or are they sorted alphabetically by article_id, or some other ways?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "From what I learned, the sequence of items in prediction vs answer matters a lot in mAPk metric.\n\nSo, how does the organizer order the final answer for each customer? Are they sorted by purchase time, or are they sorted alphabetically by article_id, or some other ways?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1728224",
      "postDate": "03/18/2022 17:21:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/reaverlee\" target=\"_blank\">@reaverlee</a>, my understanding is that the groundtruth for a customer can be viewed as an <strong>unordered set</strong>. The ordering only matters for <strong>prediction sequence</strong> because the precision at cutoff k (<em>i.e.</em>, <code>P(k)</code>) will drop if the relevant item appears <strong>later</strong> in the sequence. Hence, we hope those relevant can be predicted at top in the prediction sequence. Hope this helps.<br><br>\nIf there's any misunderstanding, please correct me, thanks a lot.  </p>",
      "rawMarkdown": "Hi @reaverlee, my understanding is that the groundtruth for a customer can be viewed as an **unordered set**. The ordering only matters for **prediction sequence** because the precision at cutoff k (*i.e.*, `P(k)`) will drop if the relevant item appears **later** in the sequence. Hence, we hope those relevant can be predicted at top in the prediction sequence. Hope this helps.<br>\nIf there's any misunderstanding, please correct me, thanks a lot.",
      "votes": null
    },
    {
      "id": "1728735",
      "postDate": "03/19/2022 07:57:08",
      "content": "<p>Correct. Ground truth sequence <strong>does not</strong> matter. And prediction sequence <strong>does</strong> matter.</p>",
      "rawMarkdown": "Correct. Ground truth sequence **does not** matter. And prediction sequence **does** matter.",
      "votes": null
    },
    {
      "id": "1729055",
      "postDate": "03/19/2022 15:22:43",
      "content": "<p>Thank you! I got it wrong when I tested mapk function with [1,2,3,4,5] vs [5,4,3,2,1]. I should've tested them for apk function since it's only for one sample.</p>",
      "rawMarkdown": "Thank you! I got it wrong when I tested mapk function with [1,2,3,4,5] vs [5,4,3,2,1]. I should've tested them for apk function since it's only for one sample.",
      "votes": null
    },
    {
      "id": "1729075",
      "postDate": "03/19/2022 15:47:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks a lot for the confirmation and explanation.😊</p>",
      "rawMarkdown": "Hi @cdeotte, thanks a lot for the confirmation and explanation.😊",
      "votes": null
    },
    {
      "id": "1736765",
      "postDate": "03/27/2022 17:39:34",
      "content": "<p>A quick follow-up question RE: 'can be viewed as an (unordered) set'. Would be grateful if someone's able to clarify.</p>\n<p>Should I take this to mean that it is enough to be a member of the ground truth set in order to qualify as a true positive in the precision calculation? What about <em>rel(k)</em> - is it enough for the item at rank <em>k</em> to be a member of the ground truth set for <em>rel(k)</em> be 1, or is that truly matching ranks?</p>\n<p>e.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, …], then P(3) = 1, and rel(3) = 1 too?</p>",
      "rawMarkdown": "A quick follow-up question RE: 'can be viewed as an (unordered) set'. Would be grateful if someone's able to clarify.\n\nShould I take this to mean that it is enough to be a member of the ground truth set in order to qualify as a true positive in the precision calculation? What about *rel(k)* - is it enough for the item at rank *k* to be a member of the ground truth set for *rel(k)* be 1, or is that truly matching ranks?\n\ne.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, ...], then P(3) = 1, and rel(3) = 1 too?",
      "votes": null
    },
    {
      "id": "1736870",
      "postDate": "03/27/2022 20:26:06",
      "content": "<blockquote>\n  <p>e.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, …], then P(3) = 1, and rel(3) = 1 too?</p>\n</blockquote>\n<p>Yes, exactly. I made a simple base line prediction notebook that shows the calculation of the MAP@12 in context of a four-fold validation that matches quite well the leaderboard score: 0.0073 vs. 0.0071</p>\n<p><a href=\"https://www.kaggle.com/code/tbierhance/basic-best-seller-prediction-w-cross-validation\" target=\"_blank\">https://www.kaggle.com/code/tbierhance/basic-best-seller-prediction-w-cross-validation</a></p>",
      "rawMarkdown": "> e.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, …], then P(3) = 1, and rel(3) = 1 too?\n\nYes, exactly. I made a simple base line prediction notebook that shows the calculation of the MAP@12 in context of a four-fold validation that matches quite well the leaderboard score: 0.0073 vs. 0.0071\n\nhttps://www.kaggle.com/code/tbierhance/basic-best-seller-prediction-w-cross-validation",
      "votes": null
    },
    {
      "id": "1736887",
      "postDate": "03/27/2022 20:54:27",
      "content": "<p>Nice, thanks for that <a href=\"https://www.kaggle.com/tbierhance\" target=\"_blank\">@tbierhance</a>, I'll check it out.</p>",
      "rawMarkdown": "Nice, thanks for that @tbierhance, I'll check it out.",
      "votes": null
    },
    {
      "id": "1737541",
      "postDate": "03/28/2022 14:11:12",
      "content": "<p>thanks for your share</p>",
      "rawMarkdown": "thanks for your share",
      "votes": null
    },
    {
      "id": "1737608",
      "postDate": "03/28/2022 14:57:56",
      "content": "<p><a href=\"https://www.kaggle.com/jagofc\" target=\"_blank\">@jagofc</a> </p>\n<blockquote>\n  <p>it is enough to be a member of the ground truth </p>\n</blockquote>\n<p>Yes, but it pays to be right in first n predictions</p>\n<p>see this - <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314192#1733955\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314192#1733955</a></p>\n<p>Also, in your example, the mapk would  $$  \\frac{1}{3} ( (1/1) \\times 1 + (2/2) \\times 1 +(3/3)\\times 1 )= 1 $$</p>\n<p>rel(k) is a mask. For this competition, it is 1 or 0 so you are not penalized for making a wrong prediction. </p>\n<p>So if, in your case, the prediction was [D,C,B,A], then the calculation would change.</p>",
      "rawMarkdown": "jagofc \n\n>  it is enough to be a member of the ground truth \n\nYes, but it pays to be right in first n predictions\n\nsee this - https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314192#1733955\n\nAlso, in your example, the mapk would  $$  \\frac{1}{3} ( (1/1) \\times 1 + (2/2) \\times 1 +(3/3)\\times 1 )= 1 $$\n\nrel(k) is a mask. For this competition, it is 1 or 0 so you are not penalized for making a wrong prediction. \n\nSo if, in your case, the prediction was [D,C,B,A], then the calculation would change.",
      "votes": null
    },
    {
      "id": "1737626",
      "postDate": "03/28/2022 15:09:43",
      "content": "<p>Great, got it, thanks. Linking in the <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">code for the competition metric</a> which makes all of this clear (featured in a couple of other discussions).</p>",
      "rawMarkdown": "Great, got it, thanks. Linking in the [code for the competition metric](https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py) which makes all of this clear (featured in a couple of other discussions).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1728224,
      "author_name": "abaojiang",
      "author_url": "",
      "post_date": "03/18/2022 17:21:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/reaverlee\" target=\"_blank\">@reaverlee</a>, my understanding is that the groundtruth for a customer can be viewed as an <strong>unordered set</strong>. The ordering only matters for <strong>prediction sequence</strong> because the precision at cutoff k (<em>i.e.</em>, <code>P(k)</code>) will drop if the relevant item appears <strong>later</strong> in the sequence. Hence, we hope those relevant can be predicted at top in the prediction sequence. Hope this helps.<br><br>\nIf there's any misunderstanding, please correct me, thanks a lot.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1728735,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/19/2022 07:57:08",
          "content": "<p>Correct. Ground truth sequence <strong>does not</strong> matter. And prediction sequence <strong>does</strong> matter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729055,
          "author_name": "reaverlee",
          "author_url": "",
          "post_date": "03/19/2022 15:22:43",
          "content": "<p>Thank you! I got it wrong when I tested mapk function with [1,2,3,4,5] vs [5,4,3,2,1]. I should've tested them for apk function since it's only for one sample.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729075,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "03/19/2022 15:47:34",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks a lot for the confirmation and explanation.😊</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1736765,
          "author_name": "jagofc",
          "author_url": "",
          "post_date": "03/27/2022 17:39:34",
          "content": "<p>A quick follow-up question RE: 'can be viewed as an (unordered) set'. Would be grateful if someone's able to clarify.</p>\n<p>Should I take this to mean that it is enough to be a member of the ground truth set in order to qualify as a true positive in the precision calculation? What about <em>rel(k)</em> - is it enough for the item at rank <em>k</em> to be a member of the ground truth set for <em>rel(k)</em> be 1, or is that truly matching ranks?</p>\n<p>e.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, …], then P(3) = 1, and rel(3) = 1 too?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1736870,
          "author_name": "tbierhance",
          "author_url": "",
          "post_date": "03/27/2022 20:26:06",
          "content": "<blockquote>\n  <p>e.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, …], then P(3) = 1, and rel(3) = 1 too?</p>\n</blockquote>\n<p>Yes, exactly. I made a simple base line prediction notebook that shows the calculation of the MAP@12 in context of a four-fold validation that matches quite well the leaderboard score: 0.0073 vs. 0.0071</p>\n<p><a href=\"https://www.kaggle.com/code/tbierhance/basic-best-seller-prediction-w-cross-validation\" target=\"_blank\">https://www.kaggle.com/code/tbierhance/basic-best-seller-prediction-w-cross-validation</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1736887,
          "author_name": "jagofc",
          "author_url": "",
          "post_date": "03/27/2022 20:54:27",
          "content": "<p>Nice, thanks for that <a href=\"https://www.kaggle.com/tbierhance\" target=\"_blank\">@tbierhance</a>, I'll check it out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1737608,
          "author_name": "atulverma",
          "author_url": "",
          "post_date": "03/28/2022 14:57:56",
          "content": "<p><a href=\"https://www.kaggle.com/jagofc\" target=\"_blank\">@jagofc</a> </p>\n<blockquote>\n  <p>it is enough to be a member of the ground truth </p>\n</blockquote>\n<p>Yes, but it pays to be right in first n predictions</p>\n<p>see this - <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314192#1733955\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314192#1733955</a></p>\n<p>Also, in your example, the mapk would  $$  \\frac{1}{3} ( (1/1) \\times 1 + (2/2) \\times 1 +(3/3)\\times 1 )= 1 $$</p>\n<p>rel(k) is a mask. For this competition, it is 1 or 0 so you are not penalized for making a wrong prediction. </p>\n<p>So if, in your case, the prediction was [D,C,B,A], then the calculation would change.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1737626,
          "author_name": "jagofc",
          "author_url": "",
          "post_date": "03/28/2022 15:09:43",
          "content": "<p>Great, got it, thanks. Linking in the <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">code for the competition metric</a> which makes all of this clear (featured in a couple of other discussions).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1737541,
      "author_name": "deepkun1995",
      "author_url": "",
      "post_date": "03/28/2022 14:11:12",
      "content": "<p>thanks for your share</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1728089": "From what I learned, the sequence of items in prediction vs answer matters a lot in mAPk metric.\n\nSo, how does the organizer order the final answer for each customer? Are they sorted by purchase time, or are they sorted alphabetically by article_id, or some other ways?\n\nThanks!",
    "1728224": "Hi @reaverlee, my understanding is that the groundtruth for a customer can be viewed as an **unordered set**. The ordering only matters for **prediction sequence** because the precision at cutoff k (*i.e.*, `P(k)`) will drop if the relevant item appears **later** in the sequence. Hence, we hope those relevant can be predicted at top in the prediction sequence. Hope this helps.<br>\nIf there's any misunderstanding, please correct me, thanks a lot.",
    "1728735": "Correct. Ground truth sequence **does not** matter. And prediction sequence **does** matter.",
    "1729055": "Thank you! I got it wrong when I tested mapk function with [1,2,3,4,5] vs [5,4,3,2,1]. I should've tested them for apk function since it's only for one sample.",
    "1729075": "Hi @cdeotte, thanks a lot for the confirmation and explanation.😊",
    "1736765": "A quick follow-up question RE: 'can be viewed as an (unordered) set'. Would be grateful if someone's able to clarify.\n\nShould I take this to mean that it is enough to be a member of the ground truth set in order to qualify as a true positive in the precision calculation? What about *rel(k)* - is it enough for the item at rank *k* to be a member of the ground truth set for *rel(k)* be 1, or is that truly matching ranks?\n\ne.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, ...], then P(3) = 1, and rel(3) = 1 too?",
    "1736870": "> e.g. If the ground truth for a customer was [A, B, C] and we were to predict [C, B, A, …], then P(3) = 1, and rel(3) = 1 too?\n\nYes, exactly. I made a simple base line prediction notebook that shows the calculation of the MAP@12 in context of a four-fold validation that matches quite well the leaderboard score: 0.0073 vs. 0.0071\n\nhttps://www.kaggle.com/code/tbierhance/basic-best-seller-prediction-w-cross-validation",
    "1736887": "Nice, thanks for that @tbierhance, I'll check it out.",
    "1737541": "thanks for your share",
    "1737608": "jagofc \n\n>  it is enough to be a member of the ground truth \n\nYes, but it pays to be right in first n predictions\n\nsee this - https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314192#1733955\n\nAlso, in your example, the mapk would  $$  \\frac{1}{3} ( (1/1) \\times 1 + (2/2) \\times 1 +(3/3)\\times 1 )= 1 $$\n\nrel(k) is a mask. For this competition, it is 1 or 0 so you are not penalized for making a wrong prediction. \n\nSo if, in your case, the prediction was [D,C,B,A], then the calculation would change.",
    "1737626": "Great, got it, thanks. Linking in the [code for the competition metric](https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py) which makes all of this clear (featured in a couple of other discussions)."
  },
  "source": "meta"
}