{
  "id": 357171,
  "title": "Can rounding improve your score? ",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357171",
  "author_name": "broccoli beef",
  "post_date": "2022-10-03T11:47:50.608000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>TL;DR</strong>: Yes, if you are confident that rounding your answer gives the ground truth.</p>\n<p>Every now and then this question arises in kaggle competitions. The answer depends on the scoring metric. </p>\n<p>In this competition, kagglers are required to provide for each test sample, estimates \\(p_1,p_2\\) for the probabilities that team A and B would score in the next 10 seconds. No restrictions were given on \\(p_1,p_2\\) other than saying that the probabilities you provide would be clipped between \\(10^{-15}\\) and \\(1-10^{-15}\\) first (for numerical reasons), but if your model is a consistent one, you really should have \\(p_1+p_2\\le1\\). Each test sample contributes a sample score \\(s(p_1,p_2)\\); the overall score is an average of all sample scores. The sample score \\(s(p_1,p_2)\\) depends in turn on the ground truth \\(y_1,y_2\\) of the sample, as follows:<br>\n<a href=\"https://postimg.cc/FfvKq9M0\" target=\"_blank\"><img src=\"https://i.postimg.cc/W4FJdzSY/Capture.jpg\" alt=\"Capture.jpg\"></a><br>\nwhere \\(x^*=\\max(10^{-15},\\min(1-10^{-15},x))\\). You see that \\(s(p_1,p_2)\\) achieves a minimum of about \\(9.992\\times10^{-16}\\) if \\((p_1,p_2)=(y_1,y_2)\\). </p>\n<p><em>Example</em>. Say your model gives \\(p_1=0.95,p_2=0.01\\). If you round them to \\((1,0)\\) and it happens to be the ground truth, then this sample contributes essentially 0 to the overall score. If you leave the probabilities as they are, the contribution is about \\(0.0307\\). Of course, if your model was wrong and the ground truth was \\((0,1)\\) instead, the contribution would be \\(3.800\\) without rounding whereas with rounding it would be a devastating \\(34.54\\). Similar analysis if the ground truth was \\((0,0)\\).</p>\n<p>So rounding can be beneficial to your score if you are confident about the estimates your model gives, possibly if the estimated probabilities are extreme (e.g., close to 0 or 1 within a threshold you feel comfortable).</p>",
  "messages": [
    {
      "id": 1969242,
      "postDate": "2022-10-03T11:47:50.610Z",
      "content": "<p><strong>TL;DR</strong>: Yes, if you are confident that rounding your answer gives the ground truth.</p>\n<p>Every now and then this question arises in kaggle competitions. The answer depends on the scoring metric. </p>\n<p>In this competition, kagglers are required to provide for each test sample, estimates \\(p_1,p_2\\) for the probabilities that team A and B would score in the next 10 seconds. No restrictions were given on \\(p_1,p_2\\) other than saying that the probabilities you provide would be clipped between \\(10^{-15}\\) and \\(1-10^{-15}\\) first (for numerical reasons), but if your model is a consistent one, you really should have \\(p_1+p_2\\le1\\). Each test sample contributes a sample score \\(s(p_1,p_2)\\); the overall score is an average of all sample scores. The sample score \\(s(p_1,p_2)\\) depends in turn on the ground truth \\(y_1,y_2\\) of the sample, as follows:<br>\n<a href=\"https://postimg.cc/FfvKq9M0\" target=\"_blank\"><img src=\"https://i.postimg.cc/W4FJdzSY/Capture.jpg\" alt=\"Capture.jpg\"></a><br>\nwhere \\(x^*=\\max(10^{-15},\\min(1-10^{-15},x))\\). You see that \\(s(p_1,p_2)\\) achieves a minimum of about \\(9.992\\times10^{-16}\\) if \\((p_1,p_2)=(y_1,y_2)\\). </p>\n<p><em>Example</em>. Say your model gives \\(p_1=0.95,p_2=0.01\\). If you round them to \\((1,0)\\) and it happens to be the ground truth, then this sample contributes essentially 0 to the overall score. If you leave the probabilities as they are, the contribution is about \\(0.0307\\). Of course, if your model was wrong and the ground truth was \\((0,1)\\) instead, the contribution would be \\(3.800\\) without rounding whereas with rounding it would be a devastating \\(34.54\\). Similar analysis if the ground truth was \\((0,0)\\).</p>\n<p>So rounding can be beneficial to your score if you are confident about the estimates your model gives, possibly if the estimated probabilities are extreme (e.g., close to 0 or 1 within a threshold you feel comfortable).</p>",
      "rawMarkdown": "**TL;DR**: Yes, if you are confident that rounding your answer gives the ground truth.\n\nEvery now and then this question arises in kaggle competitions. The answer depends on the scoring metric. \n\nIn this competition, kagglers are required to provide for each test sample, estimates \\\\(p_1,p_2\\\\) for the probabilities that team A and B would score in the next 10 seconds. No restrictions were given on \\\\(p_1,p_2\\\\) other than saying that the probabilities you provide would be clipped between \\\\(10^{-15}\\\\) and \\\\(1-10^{-15}\\\\) first (for numerical reasons), but if your model is a consistent one, you really should have \\\\(p_1+p_2\\le1\\\\). Each test sample contributes a sample score \\\\(s(p_1,p_2)\\\\); the overall score is an average of all sample scores. The sample score \\\\(s(p_1,p_2)\\\\) depends in turn on the ground truth \\\\(y_1,y_2\\\\) of the sample, as follows:\n[![Capture.jpg](https://i.postimg.cc/W4FJdzSY/Capture.jpg)](https://postimg.cc/FfvKq9M0)\nwhere \\\\(x^*=\\max(10^{-15},\\min(1-10^{-15},x))\\\\). You see that \\\\(s(p_1,p_2)\\\\) achieves a minimum of about \\\\(9.992\\times10^{-16}\\\\) if \\\\((p_1,p_2)=(y_1,y_2)\\\\). \n\n*Example*. Say your model gives \\\\(p_1=0.95,p_2=0.01\\\\). If you round them to \\\\((1,0)\\\\) and it happens to be the ground truth, then this sample contributes essentially 0 to the overall score. If you leave the probabilities as they are, the contribution is about \\\\(0.0307\\\\). Of course, if your model was wrong and the ground truth was \\\\((0,1)\\\\) instead, the contribution would be \\\\(3.800\\\\) without rounding whereas with rounding it would be a devastating \\\\(34.54\\\\). Similar analysis if the ground truth was \\\\((0,0)\\\\).\n\nSo rounding can be beneficial to your score if you are confident about the estimates your model gives, possibly if the estimated probabilities are extreme (e.g., close to 0 or 1 within a threshold you feel comfortable).",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1969242": "**TL;DR**: Yes, if you are confident that rounding your answer gives the ground truth.\n\nEvery now and then this question arises in kaggle competitions. The answer depends on the scoring metric. \n\nIn this competition, kagglers are required to provide for each test sample, estimates \\\\(p_1,p_2\\\\) for the probabilities that team A and B would score in the next 10 seconds. No restrictions were given on \\\\(p_1,p_2\\\\) other than saying that the probabilities you provide would be clipped between \\\\(10^{-15}\\\\) and \\\\(1-10^{-15}\\\\) first (for numerical reasons), but if your model is a consistent one, you really should have \\\\(p_1+p_2\\le1\\\\). Each test sample contributes a sample score \\\\(s(p_1,p_2)\\\\); the overall score is an average of all sample scores. The sample score \\\\(s(p_1,p_2)\\\\) depends in turn on the ground truth \\\\(y_1,y_2\\\\) of the sample, as follows:\n[![Capture.jpg](https://i.postimg.cc/W4FJdzSY/Capture.jpg)](https://postimg.cc/FfvKq9M0)\nwhere \\\\(x^*=\\max(10^{-15},\\min(1-10^{-15},x))\\\\). You see that \\\\(s(p_1,p_2)\\\\) achieves a minimum of about \\\\(9.992\\times10^{-16}\\\\) if \\\\((p_1,p_2)=(y_1,y_2)\\\\). \n\n*Example*. Say your model gives \\\\(p_1=0.95,p_2=0.01\\\\). If you round them to \\\\((1,0)\\\\) and it happens to be the ground truth, then this sample contributes essentially 0 to the overall score. If you leave the probabilities as they are, the contribution is about \\\\(0.0307\\\\). Of course, if your model was wrong and the ground truth was \\\\((0,1)\\\\) instead, the contribution would be \\\\(3.800\\\\) without rounding whereas with rounding it would be a devastating \\\\(34.54\\\\). Similar analysis if the ground truth was \\\\((0,0)\\\\).\n\nSo rounding can be beneficial to your score if you are confident about the estimates your model gives, possibly if the estimated probabilities are extreme (e.g., close to 0 or 1 within a threshold you feel comfortable)."
  }
}