{
  "id": 307828,
  "title": "Competition Metric - in English",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/307828",
  "author_name": "",
  "post_date": "2022-02-15T19:44:15.876248200Z",
  "votes": 16,
  "comment_count": 9,
  "views": 0,
  "content": "<p>The competition metric's exact definition can be a little hard to intuitively follow.<br>\nHere's an easier, simplified way to think about it:</p>\n<h3>\"Scoring\" our predictions for an individual customer</h3>\n<ul>\n<li>Our predictions for each customer get a score between between 0.0 and 1.0 (higher is better)</li>\n<li>To calculate a customer's score, we start with 1.0, and reduce it in two steps:<ul>\n<li>Reduce score to our recall = (# we correctly predicted / # customer purchased)<br>\n(e.g. we predicted 6 of the 10 items customer bought, so highest the score can be is 0.6)</li>\n<li>Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction<br>\n(by how much is the complicated part of the formula)</li></ul></li>\n</ul>\n<h3>How much each customer's \"score\" contributes to metric</h3>\n<ul>\n<li>Comp Metric is average of individual customers’ scores, hence also between 0.0 and 1.0</li>\n<li>Customers without sales during target week do not affect metric.</li>\n<li>Customers with sales during target week affect metric equally, regardless of # they bought.</li>\n</ul>",
  "messages": [
    {
      "id": "1692079",
      "postDate": "02/15/2022 19:44:15",
      "content": "<p>The competition metric's exact definition can be a little hard to intuitively follow.<br>\nHere's an easier, simplified way to think about it:</p>\n<h3>\"Scoring\" our predictions for an individual customer</h3>\n<ul>\n<li>Our predictions for each customer get a score between between 0.0 and 1.0 (higher is better)</li>\n<li>To calculate a customer's score, we start with 1.0, and reduce it in two steps:<ul>\n<li>Reduce score to our recall = (# we correctly predicted / # customer purchased)<br>\n(e.g. we predicted 6 of the 10 items customer bought, so highest the score can be is 0.6)</li>\n<li>Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction<br>\n(by how much is the complicated part of the formula)</li></ul></li>\n</ul>\n<h3>How much each customer's \"score\" contributes to metric</h3>\n<ul>\n<li>Comp Metric is average of individual customers’ scores, hence also between 0.0 and 1.0</li>\n<li>Customers without sales during target week do not affect metric.</li>\n<li>Customers with sales during target week affect metric equally, regardless of # they bought.</li>\n</ul>",
      "rawMarkdown": "The competition metric's exact definition can be a little hard to intuitively follow.\nHere's an easier, simplified way to think about it:\n\n### \"Scoring\" our predictions for an individual customer\n- Our predictions for each customer get a score between between 0.0 and 1.0 (higher is better)\n- To calculate a customer's score, we start with 1.0, and reduce it in two steps:\n    - Reduce score to our recall = (# we correctly predicted / # customer purchased)\n      (e.g. we predicted 6 of the 10 items customer bought, so highest the score can be is 0.6)\n    - Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction\n      (by how much is the complicated part of the formula)\n\n### How much each customer's \"score\" contributes to metric\n\n- Comp Metric is average of individual customers’ scores, hence also between 0.0 and 1.0\n- Customers without sales during target week do not affect metric.\n- Customers with sales during target week affect metric equally, regardless of # they bought.",
      "votes": null
    },
    {
      "id": "1692668",
      "postDate": "02/16/2022 07:03:58",
      "content": "<p>This is helpful, thanks for sharing!!!</p>",
      "rawMarkdown": "This is helpful, thanks for sharing!!!",
      "votes": null
    },
    {
      "id": "1693187",
      "postDate": "02/16/2022 13:56:50",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> <br>\nThanks for the reply😊</p>\n<p>After all, scoring is far from intuitive and difficult to interpret!</p>\n<p>Thanks for the reference.</p>",
      "rawMarkdown": "jacob34 \nThanks for the reply😊\n\nAfter all, scoring is far from intuitive and difficult to interpret!\n\nThanks for the reference.",
      "votes": null
    },
    {
      "id": "1693209",
      "postDate": "02/16/2022 14:13:22",
      "content": "<p>You're welcome, thanks for the feedback!</p>",
      "rawMarkdown": "You're welcome, thanks for the feedback!",
      "votes": null
    },
    {
      "id": "1693223",
      "postDate": "02/16/2022 14:32:20",
      "content": "<p>Very good explanation, thanks</p>",
      "rawMarkdown": "Very good explanation, thanks",
      "votes": null
    },
    {
      "id": "1695081",
      "postDate": "02/18/2022 00:04:30",
      "content": "<p>Maybe I'm just being a bit stupid but after reading through various threads and notebooks I still felt unclear regarding whether the <strong>order of predictions</strong> matters.</p>\n<p>As far as I could see there were some comments saying 'yes it does' and others saying 'no it doesn't'.</p>\n<p>For me if I submit the same file but with reversed() applied to my list of 12 most popular, I get the same public LB score. I know the public LB is only 3 decimal places but I'm sure the score would be lower if the order mattered?</p>\n<p>If the order doesn't matter, then actually the metric seems (relatively) much easier to me to understand.</p>\n<p>Please feel free to correct me…I usually get stuff wrong…!</p>",
      "rawMarkdown": "Maybe I'm just being a bit stupid but after reading through various threads and notebooks I still felt unclear regarding whether the **order of predictions** matters.\n\nAs far as I could see there were some comments saying 'yes it does' and others saying 'no it doesn't'.\n\nFor me if I submit the same file but with reversed() applied to my list of 12 most popular, I get the same public LB score. I know the public LB is only 3 decimal places but I'm sure the score would be lower if the order mattered?\n\nIf the order doesn't matter, then actually the metric seems (relatively) much easier to me to understand.\n\nPlease feel free to correct me...I usually get stuff wrong...!",
      "votes": null
    },
    {
      "id": "1695987",
      "postDate": "02/18/2022 14:23:29",
      "content": "<p>Order of predictions definitely does matter.<br>\nTo quote from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573\" target=\"_blank\">this post</a> (you can trust him ;)):</p>\n<blockquote>\n  <p>The order of your predictions matters but order of ground truth does not matter.</p>\n</blockquote>\n<p>Very strange that you should get same LB score with reversed - perhaps it's just by chance.<br>\nTry ordering the predictions for each customer, so that you have 6 \"real\" predictions, followed by 6 random predictions, and then try submitting to LB both straight and reversed.</p>\n<p>I agree with you that that's the part of the metric that's complicated.</p>\n<p>As I wrote:</p>\n<blockquote>\n  <p>Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction<br>\n  <strong><em>(by how much is the complicated part of the formula)</em></strong></p>\n</blockquote>",
      "rawMarkdown": "Order of predictions definitely does matter.\nTo quote from @cdeotte in [this post](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573) (you can trust him ;)):\n>The order of your predictions matters but order of ground truth does not matter.\n\nVery strange that you should get same LB score with reversed - perhaps it's just by chance.\nTry ordering the predictions for each customer, so that you have 6 \"real\" predictions, followed by 6 random predictions, and then try submitting to LB both straight and reversed.\n\n I agree with you that that's the part of the metric that's complicated.\n\nAs I wrote:\n> Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction\n***(by how much is the complicated part of the formula)***",
      "votes": null
    },
    {
      "id": "1695995",
      "postDate": "02/18/2022 14:28:04",
      "content": "<p>Yes. Furthermore, rearranging order is one way to approach recommender system competitions. For example, take the best public notebook. Then train a model which rearranges the 12 predictions for each user. This will increase the CV LB.</p>\n<p>Another approach, is to train one model or algorithm which picks 12-100 items per user. Then train a second model or algorithm which rearranges all 12-100 items and submits the top 12 (in the order determined by your second model/algorithm).</p>",
      "rawMarkdown": "Yes. Furthermore, rearranging order is one way to approach recommender system competitions. For example, take the best public notebook. Then train a model which rearranges the 12 predictions for each user. This will increase the CV LB.\n\nAnother approach, is to train one model or algorithm which picks 12-100 items per user. Then train a second model or algorithm which rearranges all 12-100 items and submits the top 12 (in the order determined by your second model/algorithm).",
      "votes": null
    },
    {
      "id": "1696489",
      "postDate": "02/18/2022 22:18:50",
      "content": "<p>Thank you for this useful explanation!</p>",
      "rawMarkdown": "Thank you for this useful explanation!",
      "votes": null
    },
    {
      "id": "1696588",
      "postDate": "02/19/2022 00:30:26",
      "content": "<p>Thanks for the replies.</p>\n<p>Yep apologies it was the correct diagnosis that I just hadn't run a robust enough test - as it happened by chance reversing the specific list I had generated the same LB score to 3 decimal places.</p>\n<p>Running a more robust test as suggested does produce a difference in score.</p>\n<p>Thanks again.</p>\n<p>(as a completely unrelated aside - anyone else think it would be helpful if the submission scoring threw an error if one of the predictions is not a valid article ID? I feel like the loss of the leading '0' from the article IDs giving a sub score of 0.0 is going to continue to confuse new entrants to the comp long after the relevant discussion threads have sunk down the list)</p>",
      "rawMarkdown": "Thanks for the replies.\n\nYep apologies it was the correct diagnosis that I just hadn't run a robust enough test - as it happened by chance reversing the specific list I had generated the same LB score to 3 decimal places.\n\nRunning a more robust test as suggested does produce a difference in score.\n\nThanks again.\n\n(as a completely unrelated aside - anyone else think it would be helpful if the submission scoring threw an error if one of the predictions is not a valid article ID? I feel like the loss of the leading '0' from the article IDs giving a sub score of 0.0 is going to continue to confuse new entrants to the comp long after the relevant discussion threads have sunk down the list)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1692668,
      "author_name": "lallucycle",
      "author_url": "",
      "post_date": "02/16/2022 07:03:58",
      "content": "<p>This is helpful, thanks for sharing!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1693187,
      "author_name": "ryosukekonno",
      "author_url": "",
      "post_date": "02/16/2022 13:56:50",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> <br>\nThanks for the reply😊</p>\n<p>After all, scoring is far from intuitive and difficult to interpret!</p>\n<p>Thanks for the reference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1693209,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "02/16/2022 14:13:22",
          "content": "<p>You're welcome, thanks for the feedback!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1693223,
      "author_name": "ambrul",
      "author_url": "",
      "post_date": "02/16/2022 14:32:20",
      "content": "<p>Very good explanation, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1695081,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "02/18/2022 00:04:30",
      "content": "<p>Maybe I'm just being a bit stupid but after reading through various threads and notebooks I still felt unclear regarding whether the <strong>order of predictions</strong> matters.</p>\n<p>As far as I could see there were some comments saying 'yes it does' and others saying 'no it doesn't'.</p>\n<p>For me if I submit the same file but with reversed() applied to my list of 12 most popular, I get the same public LB score. I know the public LB is only 3 decimal places but I'm sure the score would be lower if the order mattered?</p>\n<p>If the order doesn't matter, then actually the metric seems (relatively) much easier to me to understand.</p>\n<p>Please feel free to correct me…I usually get stuff wrong…!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1695987,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "02/18/2022 14:23:29",
          "content": "<p>Order of predictions definitely does matter.<br>\nTo quote from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573\" target=\"_blank\">this post</a> (you can trust him ;)):</p>\n<blockquote>\n  <p>The order of your predictions matters but order of ground truth does not matter.</p>\n</blockquote>\n<p>Very strange that you should get same LB score with reversed - perhaps it's just by chance.<br>\nTry ordering the predictions for each customer, so that you have 6 \"real\" predictions, followed by 6 random predictions, and then try submitting to LB both straight and reversed.</p>\n<p>I agree with you that that's the part of the metric that's complicated.</p>\n<p>As I wrote:</p>\n<blockquote>\n  <p>Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction<br>\n  <strong><em>(by how much is the complicated part of the formula)</em></strong></p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695995,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2022 14:28:04",
          "content": "<p>Yes. Furthermore, rearranging order is one way to approach recommender system competitions. For example, take the best public notebook. Then train a model which rearranges the 12 predictions for each user. This will increase the CV LB.</p>\n<p>Another approach, is to train one model or algorithm which picks 12-100 items per user. Then train a second model or algorithm which rearranges all 12-100 items and submits the top 12 (in the order determined by your second model/algorithm).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696588,
          "author_name": "davidedwards1",
          "author_url": "",
          "post_date": "02/19/2022 00:30:26",
          "content": "<p>Thanks for the replies.</p>\n<p>Yep apologies it was the correct diagnosis that I just hadn't run a robust enough test - as it happened by chance reversing the specific list I had generated the same LB score to 3 decimal places.</p>\n<p>Running a more robust test as suggested does produce a difference in score.</p>\n<p>Thanks again.</p>\n<p>(as a completely unrelated aside - anyone else think it would be helpful if the submission scoring threw an error if one of the predictions is not a valid article ID? I feel like the loss of the leading '0' from the article IDs giving a sub score of 0.0 is going to continue to confuse new entrants to the comp long after the relevant discussion threads have sunk down the list)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1696489,
      "author_name": "loicge",
      "author_url": "",
      "post_date": "02/18/2022 22:18:50",
      "content": "<p>Thank you for this useful explanation!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1692079": "The competition metric's exact definition can be a little hard to intuitively follow.\nHere's an easier, simplified way to think about it:\n\n### \"Scoring\" our predictions for an individual customer\n- Our predictions for each customer get a score between between 0.0 and 1.0 (higher is better)\n- To calculate a customer's score, we start with 1.0, and reduce it in two steps:\n    - Reduce score to our recall = (# we correctly predicted / # customer purchased)\n      (e.g. we predicted 6 of the 10 items customer bought, so highest the score can be is 0.6)\n    - Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction\n      (by how much is the complicated part of the formula)\n\n### How much each customer's \"score\" contributes to metric\n\n- Comp Metric is average of individual customers’ scores, hence also between 0.0 and 1.0\n- Customers without sales during target week do not affect metric.\n- Customers with sales during target week affect metric equally, regardless of # they bought.",
    "1692668": "This is helpful, thanks for sharing!!!",
    "1693187": "jacob34 \nThanks for the reply😊\n\nAfter all, scoring is far from intuitive and difficult to interpret!\n\nThanks for the reference.",
    "1693209": "You're welcome, thanks for the feedback!",
    "1693223": "Very good explanation, thanks",
    "1695081": "Maybe I'm just being a bit stupid but after reading through various threads and notebooks I still felt unclear regarding whether the **order of predictions** matters.\n\nAs far as I could see there were some comments saying 'yes it does' and others saying 'no it doesn't'.\n\nFor me if I submit the same file but with reversed() applied to my list of 12 most popular, I get the same public LB score. I know the public LB is only 3 decimal places but I'm sure the score would be lower if the order mattered?\n\nIf the order doesn't matter, then actually the metric seems (relatively) much easier to me to understand.\n\nPlease feel free to correct me...I usually get stuff wrong...!",
    "1695987": "Order of predictions definitely does matter.\nTo quote from @cdeotte in [this post](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573) (you can trust him ;)):\n>The order of your predictions matters but order of ground truth does not matter.\n\nVery strange that you should get same LB score with reversed - perhaps it's just by chance.\nTry ordering the predictions for each customer, so that you have 6 \"real\" predictions, followed by 6 random predictions, and then try submitting to LB both straight and reversed.\n\n I agree with you that that's the part of the metric that's complicated.\n\nAs I wrote:\n> Order of predictions matter - score gets reduced further for each False prediction we have before a True prediction\n***(by how much is the complicated part of the formula)***",
    "1695995": "Yes. Furthermore, rearranging order is one way to approach recommender system competitions. For example, take the best public notebook. Then train a model which rearranges the 12 predictions for each user. This will increase the CV LB.\n\nAnother approach, is to train one model or algorithm which picks 12-100 items per user. Then train a second model or algorithm which rearranges all 12-100 items and submits the top 12 (in the order determined by your second model/algorithm).",
    "1696489": "Thank you for this useful explanation!",
    "1696588": "Thanks for the replies.\n\nYep apologies it was the correct diagnosis that I just hadn't run a robust enough test - as it happened by chance reversing the specific list I had generated the same LB score to 3 decimal places.\n\nRunning a more robust test as suggested does produce a difference in score.\n\nThanks again.\n\n(as a completely unrelated aside - anyone else think it would be helpful if the submission scoring threw an error if one of the predictions is not a valid article ID? I feel like the loss of the leading '0' from the article IDs giving a sub score of 0.0 is going to continue to confuse new entrants to the comp long after the relevant discussion threads have sunk down the list)"
  },
  "source": "meta"
}