{
  "id": 335918,
  "title": "Why not 4 digits precision in the public leaderboard?",
  "url": "/competitions/amex-default-prediction/discussion/335918",
  "author_name": "",
  "post_date": "2022-07-08T13:18:32.114623100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
  "messages": [
    {
      "id": "1848244",
      "postDate": "07/08/2022 13:18:32",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
      "rawMarkdown": "addisonhoward",
      "votes": null
    },
    {
      "id": "1848260",
      "postDate": "07/08/2022 13:30:24",
      "content": "<p>hi, perhaps you'll find this discussion of interest: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/330130\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/330130</a></p>\n<p>main point being: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787</a></p>",
      "rawMarkdown": "hi, perhaps you'll find this discussion of interest: https://www.kaggle.com/competitions/amex-default-prediction/discussion/330130\n\nmain point being: https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787",
      "votes": null
    },
    {
      "id": "1848288",
      "postDate": "07/08/2022 13:52:49",
      "content": "<p>Indeed <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>  has shown via several routes that the 4th decimal place is essentially noise. </p>\n<p>Now this leads to the interesting situation in that there are pretty much only 4 meaningful LB scores; 0.800, 0.799, 0.798 and 0.797.  Given that there are a variety of Public notebooks that will provide 0.797, then perhaps the medals should be awarded thus</p>\n<ul>\n<li>0.800 Gold</li>\n<li>0.799 Silver </li>\n<li>0.798 Bronze</li>\n</ul>\n<p>rather than the standard Top 10(ish), Top 5% and Top 10% stratification?</p>\n<p>I think having a competition that may reach up to 4k participants, but the metric gives only three significant digits, is a bit unfair, as <strong>many</strong> people will be left out of the medal zone based on noise alone. <br>\nI see the business sense in using the ½(<em>G+D</em>) metric, but maybe simply using the normalized Gini coefficient (<em>G</em>) alone would have been better from the point of view of providing a leaderboard score? (The <em>D</em> component seems to be more than 5 times noisier than the <em>G</em> component of the metric, so using <em>G</em> alone would indeed permit more decimal places).</p>\n<p>All the best,<br>\ncarl</p>\n<p><strong>Update</strong>: <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>  has just published a notebook with a  LB score of <code>0.799</code>, so that is the end of the Public leaderboard  😄😄😄😄😄</p>",
      "rawMarkdown": "Indeed @ambrosm  has shown via several routes that the 4th decimal place is essentially noise. \n\nNow this leads to the interesting situation in that there are pretty much only 4 meaningful LB scores; 0.800, 0.799, 0.798 and 0.797.  Given that there are a variety of Public notebooks that will provide 0.797, then perhaps the medals should be awarded thus\n\n* 0.800 Gold\n* 0.799 Silver \n* 0.798 Bronze\n\nrather than the standard Top 10(ish), Top 5% and Top 10% stratification?\n\nI think having a competition that may reach up to 4k participants, but the metric gives only three significant digits, is a bit unfair, as **many** people will be left out of the medal zone based on noise alone. \nI see the business sense in using the ½(*G+D*) metric, but maybe simply using the normalized Gini coefficient (*G*) alone would have been better from the point of view of providing a leaderboard score? (The *D* component seems to be more than 5 times noisier than the *G* component of the metric, so using *G* alone would indeed permit more decimal places).\n\nAll the best,\ncarl\n\n**Update**: @ragnar123  has just published a notebook with a  LB score of `0.799`, so that is the end of the Public leaderboard  😄😄😄😄😄",
      "votes": null
    },
    {
      "id": "1848430",
      "postDate": "07/08/2022 16:09:47",
      "content": "<p>That is true, but I think this is a general scheme with some competitions, especially those which are financial ones. The Ubiquant competition had (and still has) crazy shakeups, and I think the noise in that competition is far greater than in this one. I don't think there is an optimal solution for this issue though since noise may also decide between 0.798 and 0.799, if someone is on the rounding threshold. Maybe log loss would have also been a suitable metric?</p>",
      "rawMarkdown": "That is true, but I think this is a general scheme with some competitions, especially those which are financial ones. The Ubiquant competition had (and still has) crazy shakeups, and I think the noise in that competition is far greater than in this one. I don't think there is an optimal solution for this issue though since noise may also decide between 0.798 and 0.799, if someone is on the rounding threshold. Maybe log loss would have also been a suitable metric?",
      "votes": null
    },
    {
      "id": "1849904",
      "postDate": "07/09/2022 23:30:21",
      "content": "<p>It's something that's been on my mind for a while, too. I think the main reason for not having 4 digits precision on the public leaderboard is that it would give too much information away about individual submissions. If people knew exactly how close they were to the top spot, they might be able to reverse engineer the public test more easily. </p>",
      "rawMarkdown": "It's something that's been on my mind for a while, too. I think the main reason for not having 4 digits precision on the public leaderboard is that it would give too much information away about individual submissions. If people knew exactly how close they were to the top spot, they might be able to reverse engineer the public test more easily.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1848260,
      "author_name": "eduus710",
      "author_url": "",
      "post_date": "07/08/2022 13:30:24",
      "content": "<p>hi, perhaps you'll find this discussion of interest: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/330130\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/330130</a></p>\n<p>main point being: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1848288,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "07/08/2022 13:52:49",
          "content": "<p>Indeed <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>  has shown via several routes that the 4th decimal place is essentially noise. </p>\n<p>Now this leads to the interesting situation in that there are pretty much only 4 meaningful LB scores; 0.800, 0.799, 0.798 and 0.797.  Given that there are a variety of Public notebooks that will provide 0.797, then perhaps the medals should be awarded thus</p>\n<ul>\n<li>0.800 Gold</li>\n<li>0.799 Silver </li>\n<li>0.798 Bronze</li>\n</ul>\n<p>rather than the standard Top 10(ish), Top 5% and Top 10% stratification?</p>\n<p>I think having a competition that may reach up to 4k participants, but the metric gives only three significant digits, is a bit unfair, as <strong>many</strong> people will be left out of the medal zone based on noise alone. <br>\nI see the business sense in using the ½(<em>G+D</em>) metric, but maybe simply using the normalized Gini coefficient (<em>G</em>) alone would have been better from the point of view of providing a leaderboard score? (The <em>D</em> component seems to be more than 5 times noisier than the <em>G</em> component of the metric, so using <em>G</em> alone would indeed permit more decimal places).</p>\n<p>All the best,<br>\ncarl</p>\n<p><strong>Update</strong>: <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>  has just published a notebook with a  LB score of <code>0.799</code>, so that is the end of the Public leaderboard  😄😄😄😄😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1848430,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "07/08/2022 16:09:47",
          "content": "<p>That is true, but I think this is a general scheme with some competitions, especially those which are financial ones. The Ubiquant competition had (and still has) crazy shakeups, and I think the noise in that competition is far greater than in this one. I don't think there is an optimal solution for this issue though since noise may also decide between 0.798 and 0.799, if someone is on the rounding threshold. Maybe log loss would have also been a suitable metric?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1849904,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/09/2022 23:30:21",
      "content": "<p>It's something that's been on my mind for a while, too. I think the main reason for not having 4 digits precision on the public leaderboard is that it would give too much information away about individual submissions. If people knew exactly how close they were to the top spot, they might be able to reverse engineer the public test more easily. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1848244": "addisonhoward",
    "1848260": "hi, perhaps you'll find this discussion of interest: https://www.kaggle.com/competitions/amex-default-prediction/discussion/330130\n\nmain point being: https://www.kaggle.com/competitions/amex-default-prediction/discussion/329787",
    "1848288": "Indeed @ambrosm  has shown via several routes that the 4th decimal place is essentially noise. \n\nNow this leads to the interesting situation in that there are pretty much only 4 meaningful LB scores; 0.800, 0.799, 0.798 and 0.797.  Given that there are a variety of Public notebooks that will provide 0.797, then perhaps the medals should be awarded thus\n\n* 0.800 Gold\n* 0.799 Silver \n* 0.798 Bronze\n\nrather than the standard Top 10(ish), Top 5% and Top 10% stratification?\n\nI think having a competition that may reach up to 4k participants, but the metric gives only three significant digits, is a bit unfair, as **many** people will be left out of the medal zone based on noise alone. \nI see the business sense in using the ½(*G+D*) metric, but maybe simply using the normalized Gini coefficient (*G*) alone would have been better from the point of view of providing a leaderboard score? (The *D* component seems to be more than 5 times noisier than the *G* component of the metric, so using *G* alone would indeed permit more decimal places).\n\nAll the best,\ncarl\n\n**Update**: @ragnar123  has just published a notebook with a  LB score of `0.799`, so that is the end of the Public leaderboard  😄😄😄😄😄",
    "1848430": "That is true, but I think this is a general scheme with some competitions, especially those which are financial ones. The Ubiquant competition had (and still has) crazy shakeups, and I think the noise in that competition is far greater than in this one. I don't think there is an optimal solution for this issue though since noise may also decide between 0.798 and 0.799, if someone is on the rounding threshold. Maybe log loss would have also been a suitable metric?",
    "1849904": "It's something that's been on my mind for a while, too. I think the main reason for not having 4 digits precision on the public leaderboard is that it would give too much information away about individual submissions. If people knew exactly how close they were to the top spot, they might be able to reverse engineer the public test more easily."
  },
  "source": "meta"
}