{
  "id": 220602,
  "title": "Choosing Older Submissions Wins Ties on Leaderboard",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220602",
  "author_name": "",
  "post_date": "2021-02-19T00:34:11.697819100Z",
  "votes": 43,
  "comment_count": 29,
  "views": 0,
  "content": "<p>I just realized that many teams will have the exact same score when all decimal places are revealed because there are approximately 10,000 private test images and the metric is accuracy. </p>\n<p>Each incorrect prediction alters your score by approximately 0.0001, therefore there are only 10 unique scores between 0.898 and 0.899 for example. (This is different than other competitions where there are usually more possible scores between 0.898 and 0.899).</p>\n<p>This means the team that submitted first will go higher. As a result when selecting final submissions in this competition, we should have chosen the oldest submissions. I just realized this now.</p>\n<p><strong>NOTE</strong>: this isn't true in most competitions so I'm not saying we should always do this. It happened here because of the combination of metric and number of private test images.</p>",
  "messages": [
    {
      "id": "1209562",
      "postDate": "02/19/2021 00:34:11",
      "content": "<p>I just realized that many teams will have the exact same score when all decimal places are revealed because there are approximately 10,000 private test images and the metric is accuracy. </p>\n<p>Each incorrect prediction alters your score by approximately 0.0001, therefore there are only 10 unique scores between 0.898 and 0.899 for example. (This is different than other competitions where there are usually more possible scores between 0.898 and 0.899).</p>\n<p>This means the team that submitted first will go higher. As a result when selecting final submissions in this competition, we should have chosen the oldest submissions. I just realized this now.</p>\n<p><strong>NOTE</strong>: this isn't true in most competitions so I'm not saying we should always do this. It happened here because of the combination of metric and number of private test images.</p>",
      "rawMarkdown": "I just realized that many teams will have the exact same score when all decimal places are revealed because there are approximately 10,000 private test images and the metric is accuracy. \n\nEach incorrect prediction alters your score by approximately 0.0001, therefore there are only 10 unique scores between 0.898 and 0.899 for example. (This is different than other competitions where there are usually more possible scores between 0.898 and 0.899).\n\nThis means the team that submitted first will go higher. As a result when selecting final submissions in this competition, we should have chosen the oldest submissions. I just realized this now.\n\n**NOTE**: this isn't true in most competitions so I'm not saying we should always do this. It happened here because of the combination of metric and number of private test images.",
      "votes": null
    },
    {
      "id": "1209570",
      "postDate": "02/19/2021 00:36:15",
      "content": "<p>I still don't quite get how this is scaled. So it it also measured by the submission time? Will they change the leaderboard in order to remedy this? Or is this a good thing?</p>",
      "rawMarkdown": "I still don't quite get how this is scaled. So it it also measured by the submission time? Will they change the leaderboard in order to remedy this? Or is this a good thing?",
      "votes": null
    },
    {
      "id": "1209577",
      "postDate": "02/19/2021 00:39:21",
      "content": "<p>It is my understanding that when two teams have the exact same LB score then the team that submitted first goes higher.</p>\n<p>Therefore when selecting final submissions, we should choose high scoring subs that are months old. Then they will climb higher on LB. </p>\n<p><a href=\"https://www.kaggle.com/andyjianzhou\" target=\"_blank\">@andyjianzhou</a> NOTE: that i'm not talking about the visible 3 decimal places, but even after they reveal more decimal places, many teams will have the exact same score as another team. This is because the metric is accuracy and there are only 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001, the 4th decimal place. So between the visible LB scores of 0.898 and 0.899, there are only 10 possibilities, i.e. 0.8980, 0.8981, …, 0.8989.</p>",
      "rawMarkdown": "It is my understanding that when two teams have the exact same LB score then the team that submitted first goes higher.\n\nTherefore when selecting final submissions, we should choose high scoring subs that are months old. Then they will climb higher on LB. \n\n@andyjianzhou NOTE: that i'm not talking about the visible 3 decimal places, but even after they reveal more decimal places, many teams will have the exact same score as another team. This is because the metric is accuracy and there are only 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001, the 4th decimal place. So between the visible LB scores of 0.898 and 0.899, there are only 10 possibilities, i.e. 0.8980, 0.8981, ..., 0.8989.",
      "votes": null
    },
    {
      "id": "1209578",
      "postDate": "02/19/2021 00:39:33",
      "content": "<p>I suspected that too.</p>",
      "rawMarkdown": "I suspected that too.",
      "votes": null
    },
    {
      "id": "1209583",
      "postDate": "02/19/2021 00:41:32",
      "content": "<p>Yes, I think so. I tracked my public leaderboard today. I got 0.907 (public) on my last submission and it was the last place of 0.907 score. </p>",
      "rawMarkdown": "Yes, I think so. I tracked my public leaderboard today. I got 0.907 (public) on my last submission and it was the last place of 0.907 score.",
      "votes": null
    },
    {
      "id": "1209626",
      "postDate": "02/19/2021 01:23:35",
      "content": "<p>Ah, I see thank you! Love your work by the way. Do you think that this system is good? Choosing scores that scale off of submission date? I guess this is a strategy then haha. Thanks!</p>",
      "rawMarkdown": "Ah, I see thank you! Love your work by the way. Do you think that this system is good? Choosing scores that scale off of submission date? I guess this is a strategy then haha. Thanks!",
      "votes": null
    },
    {
      "id": "1209672",
      "postDate": "02/19/2021 02:18:40",
      "content": "<p>Two year ago in Microsoft Malware competition, if we select the submission two month ago, we will end up at green zone.<br>\nNow if we choose the submission two month ago, we may end up at green zone. we do survived this time but not that good.<br>\n\"Yesterday one more\"</p>",
      "rawMarkdown": "Two year ago in Microsoft Malware competition, if we select the submission two month ago, we will end up at green zone.\nNow if we choose the submission two month ago, we may end up at green zone. we do survived this time but not that good.\n\"Yesterday one more\"",
      "votes": null
    },
    {
      "id": "1209726",
      "postDate": "02/19/2021 02:52:45",
      "content": "<p>This is interesting. I guess that because of the evaluation metric there are more ties than usual, which increases the importance of considering the time dimension. We were lucky to get away with a last day sub :)</p>",
      "rawMarkdown": "This is interesting. I guess that because of the evaluation metric there are more ties than usual, which increases the importance of considering the time dimension. We were lucky to get away with a last day sub :)",
      "votes": null
    },
    {
      "id": "1209748",
      "postDate": "02/19/2021 03:03:42",
      "content": "<p>Congratulations Nikita and Lizzzi on gold medal. Your last day sub must have been a high 901 with high 4th decimal place. I'm curious to see everyone's 4th decimal place. Then we will know who was helped by the time dimension and who could have benefited from the time dimension.</p>",
      "rawMarkdown": "Congratulations Nikita and Lizzzi on gold medal. Your last day sub must have been a high 901 with high 4th decimal place. I'm curious to see everyone's 4th decimal place. Then we will know who was helped by the time dimension and who could have benefited from the time dimension.",
      "votes": null
    },
    {
      "id": "1209831",
      "postDate": "02/19/2021 04:05:16",
      "content": "<blockquote>\n  <p>As a result when selecting final submissions, we should have chosen the oldest submissions. I just realized this now.</p>\n</blockquote>\n<p>And that implies one should join early and make subs/exps faster in early days as well</p>",
      "rawMarkdown": ">As a result when selecting final submissions, we should have chosen the oldest submissions. I just realized this now.\n\nAnd that implies one should join early and make subs/exps faster in early days as well",
      "votes": null
    },
    {
      "id": "1209938",
      "postDate": "02/19/2021 05:48:20",
      "content": "<p>Are you shure they round it up to the 4th decimal, not to the 5th or even 6th?</p>",
      "rawMarkdown": "Are you shure they round it up to the 4th decimal, not to the 5th or even 6th?",
      "votes": null
    },
    {
      "id": "1210152",
      "postDate": "02/19/2021 08:25:52",
      "content": "<p>I think the LB is sorted with unrounded scores and then numbers are rounded for visualisation. But I may be wrong</p>",
      "rawMarkdown": "I think the LB is sorted with unrounded scores and then numbers are rounded for visualisation. But I may be wrong",
      "votes": null
    },
    {
      "id": "1210454",
      "postDate": "02/19/2021 12:55:35",
      "content": "<p>Are you sure they take the time of the submission into account and not the following decimal places to resolve the tie? I'd rather expect more decimal places since it is quite often that many teams have a very similar score - just because of the nature of the metric.</p>",
      "rawMarkdown": "Are you sure they take the time of the submission into account and not the following decimal places to resolve the tie? I'd rather expect more decimal places since it is quite often that many teams have a very similar score - just because of the nature of the metric.",
      "votes": null
    },
    {
      "id": "1210625",
      "postDate": "02/19/2021 15:13:37",
      "content": "<p>There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.</p>\n<p>The possibilities are 0.8980, 0.8981, 0.8982, …, 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.</p>",
      "rawMarkdown": "There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.\n\nThe possibilities are 0.8980, 0.8981, 0.8982, ..., 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.",
      "votes": null
    },
    {
      "id": "1210626",
      "postDate": "02/19/2021 15:13:43",
      "content": "<p>There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.</p>\n<p>The possibilities are 0.8980, 0.8981, 0.8982, …, 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.</p>",
      "rawMarkdown": "There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.\n\nThe possibilities are 0.8980, 0.8981, 0.8982, ..., 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.",
      "votes": null
    },
    {
      "id": "1210628",
      "postDate": "02/19/2021 15:13:51",
      "content": "<p>There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.</p>\n<p>The possibilities are 0.8980, 0.8981, 0.8982, …, 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.</p>",
      "rawMarkdown": "There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.\n\nThe possibilities are 0.8980, 0.8981, 0.8982, ..., 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.",
      "votes": null
    },
    {
      "id": "1210643",
      "postDate": "02/19/2021 15:24:12",
      "content": "<p>hmm, that's interesting, thank you. Now I understand why you wrote about the 4th decimal place - that totally makes sense with this data. Then the idea with time sounds a very plausible explanation of the order.</p>",
      "rawMarkdown": "hmm, that's interesting, thank you. Now I understand why you wrote about the 4th decimal place - that totally makes sense with this data. Then the idea with time sounds a very plausible explanation of the order.",
      "votes": null
    },
    {
      "id": "1210661",
      "postDate": "02/19/2021 15:35:10",
      "content": "<p><a href=\"https://www.kaggle.com/andyjianzhou\" target=\"_blank\">@andyjianzhou</a> Thanks.</p>\n<p>In most competitions it doesn't matter. This competition is unique in this regard.</p>\n<p>I wasnt clear in my original explanation. I'm not referring to the visible 3 digit scores. This comp has accuracy metric and approximately 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001. So there are only 10 unique scores between 0.898 and 0.899 when they reveal all the decimal digits.</p>",
      "rawMarkdown": "andyjianzhou Thanks.\n\nIn most competitions it doesn't matter. This competition is unique in this regard.\n\nI wasnt clear in my original explanation. I'm not referring to the visible 3 digit scores. This comp has accuracy metric and approximately 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001. So there are only 10 unique scores between 0.898 and 0.899 when they reveal all the decimal digits.",
      "votes": null
    },
    {
      "id": "1210836",
      "postDate": "02/19/2021 18:18:17",
      "content": "<p>The 4th decimal values are revealed! Looks like there are indeed many ties. Makes sense that we come last in a group of 3 teams with 0.9016, since our submission was made a few hours before the end.</p>",
      "rawMarkdown": "The 4th decimal values are revealed! Looks like there are indeed many ties. Makes sense that we come last in a group of 3 teams with 0.9016, since our submission was made a few hours before the end.",
      "votes": null
    },
    {
      "id": "1210847",
      "postDate": "02/19/2021 18:32:40",
      "content": "<p>After 4th decimal revealing, it seems that i can approve <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> theory, that same scores are ranked by time in ascending order. For example, i am the last in a train of 0.8995, and it seems to be true, as my best final submission was placed ~4 hours before deadline.</p>",
      "rawMarkdown": "After 4th decimal revealing, it seems that i can approve @cdeotte theory, that same scores are ranked by time in ascending order. For example, i am the last in a train of 0.8995, and it seems to be true, as my best final submission was placed ~4 hours before deadline.",
      "votes": null
    },
    {
      "id": "1210955",
      "postDate": "02/19/2021 20:35:35",
      "content": "<p>40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. I didn't realize it until now but Cassava Comp was like <a href=\"https://www.kaggle.com/c/abstraction-and-reasoning-challenge/leaderboard\" target=\"_blank\">Abstract and Reasoning Comp</a> where there were only a limited number of scores possible and it was a race to get a score before another team.</p>\n<p><img src=\"http://playagricola.com/Kaggle/silver.jpg\" alt=\"image\"></p>",
      "rawMarkdown": "40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. I didn't realize it until now but Cassava Comp was like [Abstract and Reasoning Comp][1] where there were only a limited number of scores possible and it was a race to get a score before another team.\n\n![image](http://playagricola.com/Kaggle/silver.jpg)\n\n[1]: https://www.kaggle.com/c/abstraction-and-reasoning-challenge/leaderboard",
      "votes": null
    },
    {
      "id": "1211424",
      "postDate": "02/20/2021 08:02:38",
      "content": "<p>Actually this is a positive thing. Probably would make people join competitions earlier and reduce a number of strong kagglers joining late and destroing a leaderboard.</p>",
      "rawMarkdown": "Actually this is a positive thing. Probably would make people join competitions earlier and reduce a number of strong kagglers joining late and destroing a leaderboard.",
      "votes": null
    },
    {
      "id": "1211495",
      "postDate": "02/20/2021 09:25:30",
      "content": "<p>Oh I see, I didn't know the test dataset was that small</p>",
      "rawMarkdown": "Oh I see, I didn't know the test dataset was that small",
      "votes": null
    },
    {
      "id": "1211567",
      "postDate": "02/20/2021 10:34:14",
      "content": "<p>I feel like people with the same score should be awarded the same medal, e.g. #18 should get a gold as #17 got one.</p>",
      "rawMarkdown": "I feel like people with the same score should be awarded the same medal, e.g. #18 should get a gold as #17 got one.",
      "votes": null
    },
    {
      "id": "1211673",
      "postDate": "02/20/2021 12:38:21",
      "content": "<p>Actually, my submission was done about ~8 hours before the deadline.<br>\nThat makes sense I am the second one with 0.9016.</p>",
      "rawMarkdown": "Actually, my submission was done about ~8 hours before the deadline.\nThat makes sense I am the second one with 0.9016.",
      "votes": null
    },
    {
      "id": "1213142",
      "postDate": "02/21/2021 21:30:12",
      "content": "<p>For those who have a hard time believing Chris, i.e. that time of submission is used to break ties, here is a competition where 46 teams got the top score but were ranked by the time of submission: <a href=\"https://www.kaggle.com/c/santa-workshop-tour-2019/leaderboard\" target=\"_blank\">https://www.kaggle.com/c/santa-workshop-tour-2019/leaderboard</a></p>",
      "rawMarkdown": "For those who have a hard time believing Chris, i.e. that time of submission is used to break ties, here is a competition where 46 teams got the top score but were ranked by the time of submission: https://www.kaggle.com/c/santa-workshop-tour-2019/leaderboard",
      "votes": null
    },
    {
      "id": "1214128",
      "postDate": "02/22/2021 15:59:57",
      "content": "<blockquote>\n  <p>40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. </p>\n</blockquote>\n<p><strong>Ouch</strong>. On the bright side, our last submission made minutes before the deadline was our top score on the private LB 😂</p>",
      "rawMarkdown": "> 40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. \n\n**Ouch**. On the bright side, our last submission made minutes before the deadline was our top score on the private LB 😂",
      "votes": null
    },
    {
      "id": "1214293",
      "postDate": "02/22/2021 18:12:50",
      "content": "<p>Thanks for your remarks!</p>",
      "rawMarkdown": "Thanks for your remarks!",
      "votes": null
    },
    {
      "id": "1214864",
      "postDate": "02/23/2021 06:56:39",
      "content": "<p>how do you estimate there are 10k images in the private dataset or is this an assumption for explaining <br>\nthanks in advance</p>",
      "rawMarkdown": "how do you estimate there are 10k images in the private dataset or is this an assumption for explaining \nthanks in advance",
      "votes": null
    },
    {
      "id": "1214893",
      "postDate": "02/23/2021 07:23:18",
      "content": "<p>It says on the data description page to expect about 15,000 test images, 30% of which used to calculate the public leaderboard and 70%, the private leaderboard. </p>",
      "rawMarkdown": "It says on the data description page to expect about 15,000 test images, 30% of which used to calculate the public leaderboard and 70%, the private leaderboard.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1209570,
      "author_name": "andyjianzhou",
      "author_url": "",
      "post_date": "02/19/2021 00:36:15",
      "content": "<p>I still don't quite get how this is scaled. So it it also measured by the submission time? Will they change the leaderboard in order to remedy this? Or is this a good thing?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209577,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 00:39:21",
          "content": "<p>It is my understanding that when two teams have the exact same LB score then the team that submitted first goes higher.</p>\n<p>Therefore when selecting final submissions, we should choose high scoring subs that are months old. Then they will climb higher on LB. </p>\n<p><a href=\"https://www.kaggle.com/andyjianzhou\" target=\"_blank\">@andyjianzhou</a> NOTE: that i'm not talking about the visible 3 decimal places, but even after they reveal more decimal places, many teams will have the exact same score as another team. This is because the metric is accuracy and there are only 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001, the 4th decimal place. So between the visible LB scores of 0.898 and 0.899, there are only 10 possibilities, i.e. 0.8980, 0.8981, …, 0.8989.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209583,
          "author_name": "tom88jerry",
          "author_url": "",
          "post_date": "02/19/2021 00:41:32",
          "content": "<p>Yes, I think so. I tracked my public leaderboard today. I got 0.907 (public) on my last submission and it was the last place of 0.907 score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209626,
          "author_name": "andyjianzhou",
          "author_url": "",
          "post_date": "02/19/2021 01:23:35",
          "content": "<p>Ah, I see thank you! Love your work by the way. Do you think that this system is good? Choosing scores that scale off of submission date? I guess this is a strategy then haha. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210661,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 15:35:10",
          "content": "<p><a href=\"https://www.kaggle.com/andyjianzhou\" target=\"_blank\">@andyjianzhou</a> Thanks.</p>\n<p>In most competitions it doesn't matter. This competition is unique in this regard.</p>\n<p>I wasnt clear in my original explanation. I'm not referring to the visible 3 digit scores. This comp has accuracy metric and approximately 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001. So there are only 10 unique scores between 0.898 and 0.899 when they reveal all the decimal digits.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209578,
      "author_name": "tom88jerry",
      "author_url": "",
      "post_date": "02/19/2021 00:39:33",
      "content": "<p>I suspected that too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209672,
      "author_name": "steamedsheep",
      "author_url": "",
      "post_date": "02/19/2021 02:18:40",
      "content": "<p>Two year ago in Microsoft Malware competition, if we select the submission two month ago, we will end up at green zone.<br>\nNow if we choose the submission two month ago, we may end up at green zone. we do survived this time but not that good.<br>\n\"Yesterday one more\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209726,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "02/19/2021 02:52:45",
      "content": "<p>This is interesting. I guess that because of the evaluation metric there are more ties than usual, which increases the importance of considering the time dimension. We were lucky to get away with a last day sub :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209748,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 03:03:42",
          "content": "<p>Congratulations Nikita and Lizzzi on gold medal. Your last day sub must have been a high 901 with high 4th decimal place. I'm curious to see everyone's 4th decimal place. Then we will know who was helped by the time dimension and who could have benefited from the time dimension.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209831,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "02/19/2021 04:05:16",
      "content": "<blockquote>\n  <p>As a result when selecting final submissions, we should have chosen the oldest submissions. I just realized this now.</p>\n</blockquote>\n<p>And that implies one should join early and make subs/exps faster in early days as well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209938,
      "author_name": "nroman",
      "author_url": "",
      "post_date": "02/19/2021 05:48:20",
      "content": "<p>Are you shure they round it up to the 4th decimal, not to the 5th or even 6th?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210625,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 15:13:37",
          "content": "<p>There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.</p>\n<p>The possibilities are 0.8980, 0.8981, 0.8982, …, 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210152,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "02/19/2021 08:25:52",
      "content": "<p>I think the LB is sorted with unrounded scores and then numbers are rounded for visualisation. But I may be wrong</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210626,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 15:13:43",
          "content": "<p>There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.</p>\n<p>The possibilities are 0.8980, 0.8981, 0.8982, …, 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211495,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "02/20/2021 09:25:30",
          "content": "<p>Oh I see, I didn't know the test dataset was that small</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210454,
      "author_name": "lizzzi1",
      "author_url": "",
      "post_date": "02/19/2021 12:55:35",
      "content": "<p>Are you sure they take the time of the submission into account and not the following decimal places to resolve the tie? I'd rather expect more decimal places since it is quite often that many teams have a very similar score - just because of the nature of the metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210628,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 15:13:51",
          "content": "<p>There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.</p>\n<p>The possibilities are 0.8980, 0.8981, 0.8982, …, 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210643,
          "author_name": "lizzzi1",
          "author_url": "",
          "post_date": "02/19/2021 15:24:12",
          "content": "<p>hmm, that's interesting, thank you. Now I understand why you wrote about the 4th decimal place - that totally makes sense with this data. Then the idea with time sounds a very plausible explanation of the order.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210836,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "02/19/2021 18:18:17",
      "content": "<p>The 4th decimal values are revealed! Looks like there are indeed many ties. Makes sense that we come last in a group of 3 teams with 0.9016, since our submission was made a few hours before the end.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211673,
          "author_name": "tmhrkt",
          "author_url": "",
          "post_date": "02/20/2021 12:38:21",
          "content": "<p>Actually, my submission was done about ~8 hours before the deadline.<br>\nThat makes sense I am the second one with 0.9016.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210847,
      "author_name": "sergeydvindenko",
      "author_url": "",
      "post_date": "02/19/2021 18:32:40",
      "content": "<p>After 4th decimal revealing, it seems that i can approve <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> theory, that same scores are ranked by time in ascending order. For example, i am the last in a train of 0.8995, and it seems to be true, as my best final submission was placed ~4 hours before deadline.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210955,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/19/2021 20:35:35",
      "content": "<p>40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. I didn't realize it until now but Cassava Comp was like <a href=\"https://www.kaggle.com/c/abstraction-and-reasoning-challenge/leaderboard\" target=\"_blank\">Abstract and Reasoning Comp</a> where there were only a limited number of scores possible and it was a race to get a score before another team.</p>\n<p><img src=\"http://playagricola.com/Kaggle/silver.jpg\" alt=\"image\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1211424,
          "author_name": "nroman",
          "author_url": "",
          "post_date": "02/20/2021 08:02:38",
          "content": "<p>Actually this is a positive thing. Probably would make people join competitions earlier and reduce a number of strong kagglers joining late and destroing a leaderboard.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1214128,
          "author_name": "eschibli",
          "author_url": "",
          "post_date": "02/22/2021 15:59:57",
          "content": "<blockquote>\n  <p>40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. </p>\n</blockquote>\n<p><strong>Ouch</strong>. On the bright side, our last submission made minutes before the deadline was our top score on the private LB 😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1211567,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/20/2021 10:34:14",
      "content": "<p>I feel like people with the same score should be awarded the same medal, e.g. #18 should get a gold as #17 got one.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1213142,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/21/2021 21:30:12",
      "content": "<p>For those who have a hard time believing Chris, i.e. that time of submission is used to break ties, here is a competition where 46 teams got the top score but were ranked by the time of submission: <a href=\"https://www.kaggle.com/c/santa-workshop-tour-2019/leaderboard\" target=\"_blank\">https://www.kaggle.com/c/santa-workshop-tour-2019/leaderboard</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1214293,
      "author_name": "mdsumonhossain",
      "author_url": "",
      "post_date": "02/22/2021 18:12:50",
      "content": "<p>Thanks for your remarks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1214864,
      "author_name": "sj161199",
      "author_url": "",
      "post_date": "02/23/2021 06:56:39",
      "content": "<p>how do you estimate there are 10k images in the private dataset or is this an assumption for explaining <br>\nthanks in advance</p>",
      "votes": null,
      "replies": [
        {
          "id": 1214893,
          "author_name": "eschibli",
          "author_url": "",
          "post_date": "02/23/2021 07:23:18",
          "content": "<p>It says on the data description page to expect about 15,000 test images, 30% of which used to calculate the public leaderboard and 70%, the private leaderboard. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1209562": "I just realized that many teams will have the exact same score when all decimal places are revealed because there are approximately 10,000 private test images and the metric is accuracy. \n\nEach incorrect prediction alters your score by approximately 0.0001, therefore there are only 10 unique scores between 0.898 and 0.899 for example. (This is different than other competitions where there are usually more possible scores between 0.898 and 0.899).\n\nThis means the team that submitted first will go higher. As a result when selecting final submissions in this competition, we should have chosen the oldest submissions. I just realized this now.\n\n**NOTE**: this isn't true in most competitions so I'm not saying we should always do this. It happened here because of the combination of metric and number of private test images.",
    "1209570": "I still don't quite get how this is scaled. So it it also measured by the submission time? Will they change the leaderboard in order to remedy this? Or is this a good thing?",
    "1209577": "It is my understanding that when two teams have the exact same LB score then the team that submitted first goes higher.\n\nTherefore when selecting final submissions, we should choose high scoring subs that are months old. Then they will climb higher on LB. \n\n@andyjianzhou NOTE: that i'm not talking about the visible 3 decimal places, but even after they reveal more decimal places, many teams will have the exact same score as another team. This is because the metric is accuracy and there are only 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001, the 4th decimal place. So between the visible LB scores of 0.898 and 0.899, there are only 10 possibilities, i.e. 0.8980, 0.8981, ..., 0.8989.",
    "1209578": "I suspected that too.",
    "1209583": "Yes, I think so. I tracked my public leaderboard today. I got 0.907 (public) on my last submission and it was the last place of 0.907 score.",
    "1209626": "Ah, I see thank you! Love your work by the way. Do you think that this system is good? Choosing scores that scale off of submission date? I guess this is a strategy then haha. Thanks!",
    "1209672": "Two year ago in Microsoft Malware competition, if we select the submission two month ago, we will end up at green zone.\nNow if we choose the submission two month ago, we may end up at green zone. we do survived this time but not that good.\n\"Yesterday one more\"",
    "1209726": "This is interesting. I guess that because of the evaluation metric there are more ties than usual, which increases the importance of considering the time dimension. We were lucky to get away with a last day sub :)",
    "1209748": "Congratulations Nikita and Lizzzi on gold medal. Your last day sub must have been a high 901 with high 4th decimal place. I'm curious to see everyone's 4th decimal place. Then we will know who was helped by the time dimension and who could have benefited from the time dimension.",
    "1209831": ">As a result when selecting final submissions, we should have chosen the oldest submissions. I just realized this now.\n\nAnd that implies one should join early and make subs/exps faster in early days as well",
    "1209938": "Are you shure they round it up to the 4th decimal, not to the 5th or even 6th?",
    "1210152": "I think the LB is sorted with unrounded scores and then numbers are rounded for visualisation. But I may be wrong",
    "1210454": "Are you sure they take the time of the submission into account and not the following decimal places to resolve the tie? I'd rather expect more decimal places since it is quite often that many teams have a very similar score - just because of the nature of the metric.",
    "1210625": "There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.\n\nThe possibilities are 0.8980, 0.8981, 0.8982, ..., 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.",
    "1210626": "There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.\n\nThe possibilities are 0.8980, 0.8981, 0.8982, ..., 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.",
    "1210628": "There are approximately 10,000 images in the private test dataset. Therefore each wrong prediction changes your score by approximately 0.0001 the 4th decimal place. Therefore, there are only 10 possible scores for a team that shows 0.898 on the LB.\n\nThe possibilities are 0.8980, 0.8981, 0.8982, ..., 0.8989. There are only 10 unique scores but 250 teams have 0.898 on the leaderboard. Therefore at least 240 teams have the exact same score as another team even when we add more decimal places.",
    "1210643": "hmm, that's interesting, thank you. Now I understand why you wrote about the 4th decimal place - that totally makes sense with this data. Then the idea with time sounds a very plausible explanation of the order.",
    "1210661": "andyjianzhou Thanks.\n\nIn most competitions it doesn't matter. This competition is unique in this regard.\n\nI wasnt clear in my original explanation. I'm not referring to the visible 3 digit scores. This comp has accuracy metric and approximately 10,000 private test images. Therefore each incorrect prediction alters your score by 0.0001. So there are only 10 unique scores between 0.898 and 0.899 when they reveal all the decimal digits.",
    "1210836": "The 4th decimal values are revealed! Looks like there are indeed many ties. Makes sense that we come last in a group of 3 teams with 0.9016, since our submission was made a few hours before the end.",
    "1210847": "After 4th decimal revealing, it seems that i can approve @cdeotte theory, that same scores are ranked by time in ascending order. For example, i am the last in a train of 0.8995, and it seems to be true, as my best final submission was placed ~4 hours before deadline.",
    "1210955": "40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. I didn't realize it until now but Cassava Comp was like [Abstract and Reasoning Comp][1] where there were only a limited number of scores possible and it was a race to get a score before another team.\n\n![image](http://playagricola.com/Kaggle/silver.jpg)\n\n[1]: https://www.kaggle.com/c/abstraction-and-reasoning-challenge/leaderboard",
    "1211424": "Actually this is a positive thing. Probably would make people join competitions earlier and reduce a number of strong kagglers joining late and destroing a leaderboard.",
    "1211495": "Oh I see, I didn't know the test dataset was that small",
    "1211567": "I feel like people with the same score should be awarded the same medal, e.g. #18 should get a gold as #17 got one.",
    "1211673": "Actually, my submission was done about ~8 hours before the deadline.\nThat makes sense I am the second one with 0.9016.",
    "1213142": "For those who have a hard time believing Chris, i.e. that time of submission is used to break ties, here is a competition where 46 teams got the top score but were ranked by the time of submission: https://www.kaggle.com/c/santa-workshop-tour-2019/leaderboard",
    "1214128": "> 40 teams have score 0.8992. The oldest 13 are silver and the newest 27 are bronze. \n\n**Ouch**. On the bright side, our last submission made minutes before the deadline was our top score on the private LB 😂",
    "1214293": "Thanks for your remarks!",
    "1214864": "how do you estimate there are 10k images in the private dataset or is this an assumption for explaining \nthanks in advance",
    "1214893": "It says on the data description page to expect about 15,000 test images, 30% of which used to calculate the public leaderboard and 70%, the private leaderboard."
  },
  "source": "meta"
}