{
  "id": 37137,
  "title": "Near-perfect Dice scores are okay",
  "url": "/competitions/carvana-image-masking-challenge/discussion/37137",
  "author_name": "",
  "post_date": "2017-07-27T16:54:55.754571900Z",
  "votes": 41,
  "comment_count": 44,
  "views": 0,
  "content": "<p>We sometimes observe that newcomers are discouraged by high scores in competitions, assuming the problem is already solved. You should not be discouraged by scores close to one here:</p>\n\n<ul>\n<li>The segmentation accuracy necessary to use cutout images for human consumption is very high</li>\n<li>The difference between good and bad segmentations can be a relatively small number of pixels (that nonetheless makes up a relatively large visual difference)</li>\n</ul>\n\n<p>Every metric and problem has a unique scale when it comes to the absolute metric value and the meaning carried in the decimal places. Don't fall into the trap of thinking this problem is solved just because scores seem like they're \"close enough\" to one.</p>",
  "messages": [
    {
      "id": "207852",
      "postDate": "07/27/2017 16:54:55",
      "content": "<p>We sometimes observe that newcomers are discouraged by high scores in competitions, assuming the problem is already solved. You should not be discouraged by scores close to one here:</p>\n\n<ul>\n<li>The segmentation accuracy necessary to use cutout images for human consumption is very high</li>\n<li>The difference between good and bad segmentations can be a relatively small number of pixels (that nonetheless makes up a relatively large visual difference)</li>\n</ul>\n\n<p>Every metric and problem has a unique scale when it comes to the absolute metric value and the meaning carried in the decimal places. Don't fall into the trap of thinking this problem is solved just because scores seem like they're \"close enough\" to one.</p>",
      "rawMarkdown": "We sometimes observe that newcomers are discouraged by high scores in competitions, assuming the problem is already solved. You should not be discouraged by scores close to one here:\n\n- The segmentation accuracy necessary to use cutout images for human consumption is very high\n- The difference between good and bad segmentations can be a relatively small number of pixels (that nonetheless makes up a relatively large visual difference)\n\nEvery metric and problem has a unique scale when it comes to the absolute metric value and the meaning carried in the decimal places. Don't fall into the trap of thinking this problem is solved just because scores seem like they're \"close enough\" to one.",
      "votes": null
    },
    {
      "id": "207871",
      "postDate": "07/27/2017 18:22:58",
      "content": "<p>I'm thinking this might need more than three digits by the end... is it automated these days?  I remember it being two yesterday.</p>",
      "rawMarkdown": "I'm thinking this might need more than three digits by the end... is it automated these days?  I remember it being two yesterday.",
      "votes": null
    },
    {
      "id": "207876",
      "postDate": "07/27/2017 18:49:17",
      "content": "<p>It's manual (we increased it to 3 this morning).</p>\n\n<p>Scores are saved at full precision behind the scenes, so we only \"need\" to show as many public leaderboard digits as it takes to make the competition fun. We think it's okay to have temporary, apparent ties if the alternative is leaderboard probing and overfit scores.</p>",
      "rawMarkdown": "It's manual (we increased it to 3 this morning).\n\nScores are saved at full precision behind the scenes, so we only \"need\" to show as many public leaderboard digits as it takes to make the competition fun. We think it's okay to have temporary, apparent ties if the alternative is leaderboard probing and overfit scores.",
      "votes": null
    },
    {
      "id": "207919",
      "postDate": "07/27/2017 22:22:43",
      "content": "<p>Didn't even realise at first, but I'm happy you are taking steps to combat leaderboard fitting that are effective and super non-intrusive like this one :)</p>",
      "rawMarkdown": "Didn't even realise at first, but I'm happy you are taking steps to combat leaderboard fitting that are effective and super non-intrusive like this one :)",
      "votes": null
    },
    {
      "id": "208192",
      "postDate": "07/28/2017 23:53:25",
      "content": "<p>Is it possible to increase the size of the data? </p>\n\n<p>With the current set up this problem is somewhere between easy and very easy. </p>\n\n<p>I am expecting top 100 places closer to the end of the competition to have 0.9999+ score.</p>\n\n<p>It would be more challenging/interesting for participants and probably more useful for the organizers if we had more data for train and test sets. </p>\n\n<p>Is it possible that organizers provide more data? More cars models, different locations, different angles, different background etc.</p>",
      "rawMarkdown": "Is it possible to increase the size of the data? \n\nWith the current set up this problem is somewhere between easy and very easy. \n\nI am expecting top 100 places closer to the end of the competition to have 0.9999+ score.\n\nIt would be more challenging/interesting for participants and probably more useful for the organizers if we had more data for train and test sets. \n\nIs it possible that organizers provide more data? More cars models, different locations, different angles, different background etc.",
      "votes": null
    },
    {
      "id": "208212",
      "postDate": "07/29/2017 02:54:49",
      "content": "<p>Could the metric use a log-based Dice or any metric that amplifies the differences near to 1.0? For example, logX(dice*X+smooth)  ) [I haven't really thought this carefully though so it might have certain downsides, e.g. LB feedback information at high scores].  </p>\n\n<p>In the case of X=1.01, log1.01(Dice*1.01+1e-15): Dice=0.0 -&gt; -3471; Dice=0.5 -&gt; -68; Dice=0.991 -&gt; 0.091; Dice=0.999 -&gt; 0.899; Dice=0.9999 -&gt; 0.9899</p>\n\n<p>The main issue is that the segmented object is very large and the \"easy\"/central pixels dominate the calculation, while the real goal of this competition is to make visually appealing cutouts so the boundary (\"harder\") pixels are in practice more important.  Alternatively, a modified Dice metric that puts a higher weight on the boundary pixels of the ground truth object may be useful (though it may complicate computation).</p>",
      "rawMarkdown": "Could the metric use a log-based Dice or any metric that amplifies the differences near to 1.0? For example, logX(dice*X+smooth)  ) [I haven't really thought this carefully though so it might have certain downsides, e.g. LB feedback information at high scores].  \n\nIn the case of X=1.01, log1.01(Dice*1.01+1e-15): Dice=0.0 -&gt; -3471; Dice=0.5 -&gt; -68; Dice=0.991 -&gt; 0.091; Dice=0.999 -&gt; 0.899; Dice=0.9999 -&gt; 0.9899\n\nThe main issue is that the segmented object is very large and the \"easy\"/central pixels dominate the calculation, while the real goal of this competition is to make visually appealing cutouts so the boundary (\"harder\") pixels are in practice more important.  Alternatively, a modified Dice metric that puts a higher weight on the boundary pixels of the ground truth object may be useful (though it may complicate computation).",
      "votes": null
    },
    {
      "id": "208362",
      "postDate": "07/29/2017 14:36:28",
      "content": "<p>Agree with you. Another option is Y = - log_10( 1 - dice ). Then Dice=0.9990-&gt;3.000, Dice=0.9999-&gt;4.000</p>",
      "rawMarkdown": "Agree with you. Another option is Y = - log_10( 1 - dice ). Then Dice=0.9990-&gt;3.000, Dice=0.9999-&gt;4.000",
      "votes": null
    },
    {
      "id": "208466",
      "postDate": "07/29/2017 21:24:42",
      "content": "<p>I don't think it will be practical to provide more manual masks, as they are quite expensive and time-consuming to produce.</p>\n\n<p>With the initial 100,000+ images in the train + test sets, we may have exhausted the distinct car models in Carvana's inventory. It seems adding additional training images of existing car models would make it easier, rather than harder. As for locations, angles, and backgrounds, these photos are what we have to work with.</p>\n\n<p>As noted in the pinned Dice score thread (edit: lol, in <em>this</em> thread /facepalm), the competition may not be on the 3rd significant digit, the 5th, or higher. Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. The difference between 0.9[..]97 and 0.9[..]98 may be what finally makes the difference between an unusable result and a usable result.</p>\n\n<p>In the end, I think it's still hard to say if the top result will fall short of being a usable product, or if 100+ teams will have such accurate masks the winner is determined by who is arbitrarily closest to the slightly inconsistent hand-drawn outlines in the test set. If the latter seems apparent early enough in the competition, we can look into a way to up the ante, but I would be concerned about the fairness of moving the goal posts in the middle of the competition.</p>",
      "rawMarkdown": "I don't think it will be practical to provide more manual masks, as they are quite expensive and time-consuming to produce.\n\nWith the initial 100,000+ images in the train + test sets, we may have exhausted the distinct car models in Carvana's inventory. It seems adding additional training images of existing car models would make it easier, rather than harder. As for locations, angles, and backgrounds, these photos are what we have to work with.\n\nAs noted in the pinned Dice score thread (edit: lol, in *this* thread /facepalm), the competition may not be on the 3rd significant digit, the 5th, or higher. Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. The difference between 0.9[..]97 and 0.9[..]98 may be what finally makes the difference between an unusable result and a usable result.\n\nIn the end, I think it's still hard to say if the top result will fall short of being a usable product, or if 100+ teams will have such accurate masks the winner is determined by who is arbitrarily closest to the slightly inconsistent hand-drawn outlines in the test set. If the latter seems apparent early enough in the competition, we can look into a way to up the ante, but I would be concerned about the fairness of moving the goal posts in the middle of the competition.",
      "votes": null
    },
    {
      "id": "208487",
      "postDate": "07/30/2017 00:25:30",
      "content": "<pre><code>Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. \n</code></pre>\n\n<p>You know, that will be about 250 wrong pixels in a whole image. I bet, it will be extremely accurate and  awesome masks. </p>",
      "rawMarkdown": "Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. \n\nYou know, that will be about 250 wrong pixels in a whole image. I bet, it will be extremely accurate and  awesome masks.",
      "votes": null
    },
    {
      "id": "208683",
      "postDate": "07/30/2017 15:24:59",
      "content": "<p>That's a good point! Due to the <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229\">human error and inconsistencies in the masks we're training/scoring against</a>, the ground truth masks could themselves be inaccurate by an average of 100-300 pixels. The scores are already approaching 0.999 (congrats on hitting 0.995, wow!) but I guess it's a matter of whether we'll continue the pace to 0.9999 similarly quickly or scores will plateau with incremental gains becoming much more difficult.</p>",
      "rawMarkdown": "That's a good point! Due to the [human error and inconsistencies in the masks we're training/scoring against][1], the ground truth masks could themselves be inaccurate by an average of 100-300 pixels. The scores are already approaching 0.999 (congrats on hitting 0.995, wow!) but I guess it's a matter of whether we'll continue the pace to 0.9999 similarly quickly or scores will plateau with incremental gains becoming much more difficult.\n\n  [1]: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229",
      "votes": null
    },
    {
      "id": "208686",
      "postDate": "07/30/2017 15:42:27",
      "content": "<p>Thanks, it was really easy, just a good ol' Unet :) And there is still a big bag with tricks to apply. </p>\n\n<p>On a side question: do you have higher resolution images? Even simplest phone cameras output much bigger files and DSLR cameras typically used for professional photo-shoots have up to 50mpix images. And also JPEG compression artifacts are highly visible on current images (basically we could find rough edges of cars by looking at artifacts). It will be interesting to apply segmentation techniques to 50mpix uncompressed RAW file and compare the results :)  </p>",
      "rawMarkdown": "Thanks, it was really easy, just a good ol' Unet :) And there is still a big bag with tricks to apply. \n\nOn a side question: do you have higher resolution images? Even simplest phone cameras output much bigger files and DSLR cameras typically used for professional photo-shoots have up to 50mpix images. And also JPEG compression artifacts are highly visible on current images (basically we could find rough edges of cars by looking at artifacts). It will be interesting to apply segmentation techniques to 50mpix uncompressed RAW file and compare the results :)",
      "votes": null
    },
    {
      "id": "208689",
      "postDate": "07/30/2017 15:58:14",
      "content": "<p>I don't think Carvana retains original RAW files, but I should be able to acquire JPGs at around ~7500 x ~5000. I'll have to sync up with Kaggle about the process for introducing new source images to the competition.</p>\n\n<p>The manual masks are based on the 1918x1080 versions, so we would not be able to provide masks with higher fidelity.</p>",
      "rawMarkdown": "I don't think Carvana retains original RAW files, but I should be able to acquire JPGs at around ~7500 x ~5000. I'll have to sync up with Kaggle about the process for introducing new source images to the competition.\n\nThe manual masks are based on the 1918x1080 versions, so we would not be able to provide masks with higher fidelity.",
      "votes": null
    },
    {
      "id": "208702",
      "postDate": "07/30/2017 17:33:30",
      "content": "<p>It would be great! I don't know immediately how we can use higher resolution images, however, it should contain some useful information even if there won't be higher resolution masks. For example, we could use such images to find if there are some systematic errors in manual segmentation or something like that. </p>",
      "rawMarkdown": "It would be great! I don't know immediately how we can use higher resolution images, however, it should contain some useful information even if there won't be higher resolution masks. For example, we could use such images to find if there are some systematic errors in manual segmentation or something like that.",
      "votes": null
    },
    {
      "id": "209052",
      "postDate": "08/01/2017 02:44:18",
      "content": "<p>Downsampling higher resolution images back to 1918x1080 would effectively eliminate or greatly reduce jpeg artifacts and noise, and could potentially improve segmentation accuracy.</p>",
      "rawMarkdown": "Downsampling higher resolution images back to 1918x1080 would effectively eliminate or greatly reduce jpeg artifacts and noise, and could potentially improve segmentation accuracy.",
      "votes": null
    },
    {
      "id": "209095",
      "postDate": "08/01/2017 07:04:13",
      "content": "<blockquote>\n  <p>we increased it to 3 this morning</p>\n</blockquote>\n\n<p>@William, I do not think it is a good idea. I am sure it is.</p>",
      "rawMarkdown": "&gt;we increased it to 3 this morning\n\n@William, I do not think it is a good idea. I am sure it is.",
      "votes": null
    },
    {
      "id": "210147",
      "postDate": "08/04/2017 12:54:46",
      "content": "<p>Can you give an estimate of the real size of the test set?\nIs it similar to the train set?</p>",
      "rawMarkdown": "Can you give an estimate of the real size of the test set?\nIs it similar to the train set?",
      "votes": null
    },
    {
      "id": "211366",
      "postDate": "08/08/2017 18:40:41",
      "content": "<p>&gt; <strong>William Cukierski wrote</strong>\n&gt; \n&gt; &gt; It's manual (we increased it to 3 this morning).\n&gt; </p>\n\n<p>Would it be possible to increase the number to 4 (or more) in the last few days of the competition? I am not 100% sure, but I think it would be impossible to mine the LB that much in just a few days, and it would give us a better opportunity to make the decision regarding our final submission, especially since the winner(s) will ultimately be determined based on the fourth or fifth digit after the decimal point.</p>",
      "rawMarkdown": "&gt; **William Cukierski wrote**\n&gt; \n&gt; &gt; It's manual (we increased it to 3 this morning).\n&gt; \n\nWould it be possible to increase the number to 4 (or more) in the last few days of the competition? I am not 100% sure, but I think it would be impossible to mine the LB that much in just a few days, and it would give us a better opportunity to make the decision regarding our final submission, especially since the winner(s) will ultimately be determined based on the fourth or fifth digit after the decimal point.",
      "votes": null
    },
    {
      "id": "211403",
      "postDate": "08/08/2017 20:54:32",
      "content": "<p>What is left for us simple humans when the scores are already that high? hahaha </p>",
      "rawMarkdown": "What is left for us simple humans when the scores are already that high? hahaha",
      "votes": null
    },
    {
      "id": "211980",
      "postDate": "08/10/2017 12:09:08",
      "content": "<p>What is the leaderboard ranking strategy for the same \"rounded\" dice score, for example 0.996?\nLooks like submissions are ranked only by the first submission score for 0.996 range (i.e. additional digits are considered), then nobody moves up inside 0.996. The only way to get higher is to get 0.997. </p>",
      "rawMarkdown": "What is the leaderboard ranking strategy for the same \"rounded\" dice score, for example 0.996?\nLooks like submissions are ranked only by the first submission score for 0.996 range (i.e. additional digits are considered), then nobody moves up inside 0.996. The only way to get higher is to get 0.997.",
      "votes": null
    },
    {
      "id": "212034",
      "postDate": "08/10/2017 14:47:55",
      "content": "<p>The leaderboard scores are truncated, not rounded. Behind the scenes we are computing and ranking based on the full precision (i.e. if you see two teams with 0.996, the higher team has either a better score or the exact same score, but submitted it first).</p>",
      "rawMarkdown": "The leaderboard scores are truncated, not rounded. Behind the scenes we are computing and ranking based on the full precision (i.e. if you see two teams with 0.996, the higher team has either a better score or the exact same score, but submitted it first).",
      "votes": null
    },
    {
      "id": "212065",
      "postDate": "08/10/2017 15:55:15",
      "content": "<p>Looking at the leaderboard it doesn't seem to be the case. Rank is not calculated based on the full precision after the first submission that hits  0.99x, unless the score is improved by at least 0.001.</p>",
      "rawMarkdown": "Looking at the leaderboard it doesn't seem to be the case. Rank is not calculated based on the full precision after the first submission that hits  0.99x, unless the score is improved by at least 0.001.",
      "votes": null
    },
    {
      "id": "212940",
      "postDate": "08/13/2017 06:09:09",
      "content": "<p>Why not show the last few digits of the score. One submit and get a same score (0.996), but even cannot know whether it has an improvement.</p>",
      "rawMarkdown": "Why not show the last few digits of the score. One submit and get a same score (0.996), but even cannot know whether it has an improvement.",
      "votes": null
    },
    {
      "id": "213148",
      "postDate": "08/13/2017 21:52:02",
      "content": "<p>It's because the test set is not to be used for tuning your hyperparameters. I know it's only 25% of the data, so you might treat it as some kind of a validation set but this way it helps you not to overfit to the data as probably any improvement less than 0.001 is not statistically significant.</p>",
      "rawMarkdown": "It's because the test set is not to be used for tuning your hyperparameters. I know it's only 25% of the data, so you might treat it as some kind of a validation set but this way it helps you not to overfit to the data as probably any improvement less than 0.001 is not statistically significant.",
      "votes": null
    },
    {
      "id": "213197",
      "postDate": "08/14/2017 02:26:29",
      "content": "<p>in your submission page, you can do a \"sort by public score\". Then you would know if your latest submission is better than the previous.</p>",
      "rawMarkdown": "in your submission page, you can do a \"sort by public score\". Then you would know if your latest submission is better than the previous.",
      "votes": null
    },
    {
      "id": "213199",
      "postDate": "08/14/2017 02:34:46",
      "content": "<p>i think the the ranking is only done \"first time\". I can do a \"sort by public score\" in my submission page to know which submissions are better than others. But it seems that a better score at my submission page doesn't give a better rank in the leaderboard page.  I wonder if other kagglers can confirm this?</p>\n\n<p>[note]:  Also the raw data csv file which can be downloaded from the leaderboard website captures submission that are better then previous ones. If you see that some kagglers have  several 0.996, it means that later 0.996 is better than the preceeding ones. i believe this is not reflected in the ranking page as well. </p>\n\n<p>Hence i think the ranking is done at full precision at the backend. But for some results, the webpage doesn't show it?</p>",
      "rawMarkdown": "i think the the ranking is only done \"first time\". I can do a \"sort by public score\" in my submission page to know which submissions are better than others. But it seems that a better score at my submission page doesn't give a better rank in the leaderboard page.  I wonder if other kagglers can confirm this?\n\n[note]:  Also the raw data csv file which can be downloaded from the leaderboard website captures submission that are better then previous ones. If you see that some kagglers have  several 0.996, it means that later 0.996 is better than the preceeding ones. i believe this is not reflected in the ranking page as well. \n\nHence i think the ranking is done at full precision at the backend. But for some results, the webpage doesn't show it?",
      "votes": null
    },
    {
      "id": "213205",
      "postDate": "08/14/2017 03:16:22",
      "content": "<p>Agreed, there's something fishy when there are no obvious shifts within the (currently very large) 0.996 band - only entrants to the band from a lower score have been recorded.  Maybe one can try to save the public LB daily to confirm this.</p>",
      "rawMarkdown": "Agreed, there's something fishy when there are no obvious shifts within the (currently very large) 0.996 band - only entrants to the band from a lower score have been recorded.  Maybe one can try to save the public LB daily to confirm this.",
      "votes": null
    },
    {
      "id": "213343",
      "postDate": "08/14/2017 14:54:00",
      "content": "<p>Good idea! Thanks for your advice.</p>",
      "rawMarkdown": "Good idea! Thanks for your advice.",
      "votes": null
    },
    {
      "id": "213982",
      "postDate": "08/15/2017 16:23:39",
      "content": "<p>In future, kaggle can consider providing unlabelled data in the training set. I think weak supervised learning is getting better nowadays.</p>",
      "rawMarkdown": "In future, kaggle can consider providing unlabelled data in the training set. I think weak supervised learning is getting better nowadays.",
      "votes": null
    },
    {
      "id": "214009",
      "postDate": "08/15/2017 17:55:53",
      "content": "<p>I think you can use test set to update your model weights as long as you don't touch the data manually.\nThere are some discussion about pseudo labeling - <a href=\"http://forums.fast.ai/t/pseudo-labeling-in-ml/247/6\">http://forums.fast.ai/t/pseudo-labeling-in-ml/247/6</a> \nthough I don't understand how it works. This competition may be a good chance to try it out :)</p>",
      "rawMarkdown": "I think you can use test set to update your model weights as long as you don't touch the data manually.\nThere are some discussion about pseudo labeling - http://forums.fast.ai/t/pseudo-labeling-in-ml/247/6 \nthough I don't understand how it works. This competition may be a good chance to try it out :)",
      "votes": null
    },
    {
      "id": "214026",
      "postDate": "08/15/2017 19:05:24",
      "content": "<p>I also agree. It is very odd that there are no changes inside same score band except from lower band.\nMy local dice score was very lower-end of 0.997 - it is unlikely that I'm keeping the 1st position with this score.</p>",
      "rawMarkdown": "I also agree. It is very odd that there are no changes inside same score band except from lower band.\nMy local dice score was very lower-end of 0.997 - it is unlikely that I'm keeping the 1st position with this score.",
      "votes": null
    },
    {
      "id": "214275",
      "postDate": "08/16/2017 13:27:52",
      "content": "<p>Hi William and Brian</p>\n\n<p>Can you please ensure us that private set is clear from \"human errors\"?</p>\n\n<p>Now my network is about 0.997-0.998 and I clearly observe that there are many small inaccuracies in your masks.</p>\n\n<p>It makes all further improvements meaningless and this competition became some kind of \"Lottery competition\" </p>",
      "rawMarkdown": "Hi William and Brian\n\nCan you please ensure us that private set is clear from \"human errors\"?\n\nNow my network is about 0.997-0.998 and I clearly observe that there are many small inaccuracies in your masks.\n\nIt makes all further improvements meaningless and this competition became some kind of \"Lottery competition\"",
      "votes": null
    },
    {
      "id": "214597",
      "postDate": "08/17/2017 14:34:12",
      "content": "<p>It has traditionally been the moving on the leaders board that gives me the motivation to keep working, in practical terms it offers more valuable motivation than the the motivation to win. There have been times in other competitions where I knew I was behind the team in front of me by a couple 10 thousandths, and I said to self, \"self, you can find 2 ten thousandths somewhere.\" This competition does not offer that for folks at the top. I just started, so can find motivation for my next few submissions, but when I do get above about .98, where do I find the motivation to get a just a little bit better model?</p>",
      "rawMarkdown": "It has traditionally been the moving on the leaders board that gives me the motivation to keep working, in practical terms it offers more valuable motivation than the the motivation to win. There have been times in other competitions where I knew I was behind the team in front of me by a couple 10 thousandths, and I said to self, \"self, you can find 2 ten thousandths somewhere.\" This competition does not offer that for folks at the top. I just started, so can find motivation for my next few submissions, but when I do get above about .98, where do I find the motivation to get a just a little bit better model?",
      "votes": null
    },
    {
      "id": "215282",
      "postDate": "08/20/2017 23:05:57",
      "content": "<p>Confirm, didn't change my position (and constantly going down replaced by other .996ers from 31st to 81st place) on the leaderboard since I've made it to the \".996\". </p>\n\n<p>i.e. my last submission was worth moving from 0.99596 to 0.99638, however, my position didn't change.</p>",
      "rawMarkdown": "Confirm, didn't change my position (and constantly going down replaced by other .996ers from 31st to 81st place) on the leaderboard since I've made it to the \".996\". \n\ni.e. my last submission was worth moving from 0.99596 to 0.99638, however, my position didn't change.",
      "votes": null
    },
    {
      "id": "216129",
      "postDate": "08/24/2017 12:16:20",
      "content": "<p>Agreed, the public leaderboard is broken. I hope that ranking in \"my submissions\" is at least correct, otherwise we really are in a position where we have no idea about what we are doing!</p>",
      "rawMarkdown": "Agreed, the public leaderboard is broken. I hope that ranking in \"my submissions\" is at least correct, otherwise we really are in a position where we have no idea about what we are doing!",
      "votes": null
    },
    {
      "id": "216561",
      "postDate": "08/26/2017 16:55:11",
      "content": "<p>Please, add 2 more digits after dot on LB. As I understand low precision was made to prevent LB probing. But it's not the case here - 100K images + we predict not the classes but masks. So it's almost not possible to probe something useful from LB. In addition LB became useless and it's hard to track your own progress, because currently we tune 4-5 digit after dot. This situation is actually bad for motivation.</p>",
      "rawMarkdown": "Please, add 2 more digits after dot on LB. As I understand low precision was made to prevent LB probing. But it's not the case here - 100K images + we predict not the classes but masks. So it's almost not possible to probe something useful from LB. In addition LB became useless and it's hard to track your own progress, because currently we tune 4-5 digit after dot. This situation is actually bad for motivation.",
      "votes": null
    },
    {
      "id": "223722",
      "postDate": "09/23/2017 06:22:59",
      "content": "<p>Is the split between Public and Private Test Set made randomly?</p>",
      "rawMarkdown": "Is the split between Public and Private Test Set made randomly?",
      "votes": null
    },
    {
      "id": "223936",
      "postDate": "09/24/2017 11:01:10",
      "content": "<p>Yes, random.</p>",
      "rawMarkdown": "Yes, random.",
      "votes": null
    },
    {
      "id": "225132",
      "postDate": "09/28/2017 10:32:02",
      "content": "<p>Can the digits be expanded (say to 6 digits) now that the competition is over? It's easier to see the difference in each private LB band that way.</p>",
      "rawMarkdown": "Can the digits be expanded (say to 6 digits) now that the competition is over? It's easier to see the difference in each private LB band that way.",
      "votes": null
    },
    {
      "id": "225187",
      "postDate": "09/28/2017 13:13:30",
      "content": "<p>Of course! Done.</p>",
      "rawMarkdown": "Of course! Done.",
      "votes": null
    },
    {
      "id": "225195",
      "postDate": "09/28/2017 13:27:47",
      "content": "<p>And, also can you give some insight about how many actual images were in test set? Thanks!</p>",
      "rawMarkdown": "And, also can you give some insight about how many actual images were in test set? Thanks!",
      "votes": null
    },
    {
      "id": "225215",
      "postDate": "09/28/2017 14:34:55",
      "content": "<pre><code>  95200 Ignored  \n   3664 Private  \n   1200 Public  \n</code></pre>",
      "rawMarkdown": "95200 Ignored  \n       3664 Private  \n       1200 Public",
      "votes": null
    },
    {
      "id": "225217",
      "postDate": "09/28/2017 14:41:03",
      "content": "<p>So, it was exactly as we thought about it. Thanks!</p>",
      "rawMarkdown": "So, it was exactly as we thought about it. Thanks!",
      "votes": null
    },
    {
      "id": "226098",
      "postDate": "10/01/2017 02:47:30",
      "content": "<p>An interesting question: what is the best way to combat hand labeling of test records, a particular problem in computer vision contests where it is perhaps most feasible?  One way is to run two-stage contests, which Kaggle seems to be doing more of lately, which require uploading of final models before the official test data is released.   Another way is to supplement the test data with a very large number of dummy (ignored) records, as was done in this contest.  Personally I prefer the dummy records method, but I'm curious as to what other contestants think about this issue.</p>",
      "rawMarkdown": "An interesting question: what is the best way to combat hand labeling of test records, a particular problem in computer vision contests where it is perhaps most feasible?  One way is to run two-stage contests, which Kaggle seems to be doing more of lately, which require uploading of final models before the official test data is released.   Another way is to supplement the test data with a very large number of dummy (ignored) records, as was done in this contest.  Personally I prefer the dummy records method, but I'm curious as to what other contestants think about this issue.",
      "votes": null
    },
    {
      "id": "294543",
      "postDate": "03/12/2018 05:55:31",
      "content": "<p>Dear William Cukierski,</p>\n\n<p>I am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.</p>\n\n<p>Main content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.</p>\n\n<p>I was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives (attention maps on top of the image etc) are possible content candidates.</p>\n\n<p>If using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.</p>\n\n<p>Please let me know, Best regards, Kweonwoo</p>\n\n<p>Same message is posted on a separate discussion\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/51686\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/51686</a></p>",
      "rawMarkdown": "Dear William Cukierski,\n\nI am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.\n\nMain content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.\n\nI was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives (attention maps on top of the image etc) are possible content candidates.\n\nIf using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.\n\nPlease let me know, Best regards, Kweonwoo\n\nSame message is posted on a separate discussion\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/51686",
      "votes": null
    },
    {
      "id": "371216",
      "postDate": "08/16/2018 08:42:03",
      "content": "<p>It was very great and insightful challenge. Hope, you will consider organizing next similar challenge soon!</p>",
      "rawMarkdown": "It was very great and insightful challenge. Hope, you will consider organizing next similar challenge soon!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 207871,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "07/27/2017 18:22:58",
      "content": "<p>I'm thinking this might need more than three digits by the end... is it automated these days?  I remember it being two yesterday.</p>",
      "votes": null,
      "replies": [
        {
          "id": 207876,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/27/2017 18:49:17",
          "content": "<p>It's manual (we increased it to 3 this morning).</p>\n\n<p>Scores are saved at full precision behind the scenes, so we only \"need\" to show as many public leaderboard digits as it takes to make the competition fun. We think it's okay to have temporary, apparent ties if the alternative is leaderboard probing and overfit scores.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 207919,
          "author_name": "ankasor",
          "author_url": "",
          "post_date": "07/27/2017 22:22:43",
          "content": "<p>Didn't even realise at first, but I'm happy you are taking steps to combat leaderboard fitting that are effective and super non-intrusive like this one :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209095,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "08/01/2017 07:04:13",
          "content": "<blockquote>\n  <p>we increased it to 3 this morning</p>\n</blockquote>\n\n<p>@William, I do not think it is a good idea. I am sure it is.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 211366,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "08/08/2017 18:40:41",
          "content": "<p>&gt; <strong>William Cukierski wrote</strong>\n&gt; \n&gt; &gt; It's manual (we increased it to 3 this morning).\n&gt; </p>\n\n<p>Would it be possible to increase the number to 4 (or more) in the last few days of the competition? I am not 100% sure, but I think it would be impossible to mine the LB that much in just a few days, and it would give us a better opportunity to make the decision regarding our final submission, especially since the winner(s) will ultimately be determined based on the fourth or fifth digit after the decimal point.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208192,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "07/28/2017 23:53:25",
      "content": "<p>Is it possible to increase the size of the data? </p>\n\n<p>With the current set up this problem is somewhere between easy and very easy. </p>\n\n<p>I am expecting top 100 places closer to the end of the competition to have 0.9999+ score.</p>\n\n<p>It would be more challenging/interesting for participants and probably more useful for the organizers if we had more data for train and test sets. </p>\n\n<p>Is it possible that organizers provide more data? More cars models, different locations, different angles, different background etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 208466,
          "author_name": "brianshaler",
          "author_url": "",
          "post_date": "07/29/2017 21:24:42",
          "content": "<p>I don't think it will be practical to provide more manual masks, as they are quite expensive and time-consuming to produce.</p>\n\n<p>With the initial 100,000+ images in the train + test sets, we may have exhausted the distinct car models in Carvana's inventory. It seems adding additional training images of existing car models would make it easier, rather than harder. As for locations, angles, and backgrounds, these photos are what we have to work with.</p>\n\n<p>As noted in the pinned Dice score thread (edit: lol, in <em>this</em> thread /facepalm), the competition may not be on the 3rd significant digit, the 5th, or higher. Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. The difference between 0.9[..]97 and 0.9[..]98 may be what finally makes the difference between an unusable result and a usable result.</p>\n\n<p>In the end, I think it's still hard to say if the top result will fall short of being a usable product, or if 100+ teams will have such accurate masks the winner is determined by who is arbitrarily closest to the slightly inconsistent hand-drawn outlines in the test set. If the latter seems apparent early enough in the competition, we can look into a way to up the ante, but I would be concerned about the fairness of moving the goal posts in the middle of the competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208487,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 00:25:30",
          "content": "<pre><code>Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. \n</code></pre>\n\n<p>You know, that will be about 250 wrong pixels in a whole image. I bet, it will be extremely accurate and  awesome masks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208683,
          "author_name": "brianshaler",
          "author_url": "",
          "post_date": "07/30/2017 15:24:59",
          "content": "<p>That's a good point! Due to the <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229\">human error and inconsistencies in the masks we're training/scoring against</a>, the ground truth masks could themselves be inaccurate by an average of 100-300 pixels. The scores are already approaching 0.999 (congrats on hitting 0.995, wow!) but I guess it's a matter of whether we'll continue the pace to 0.9999 similarly quickly or scores will plateau with incremental gains becoming much more difficult.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208686,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 15:42:27",
          "content": "<p>Thanks, it was really easy, just a good ol' Unet :) And there is still a big bag with tricks to apply. </p>\n\n<p>On a side question: do you have higher resolution images? Even simplest phone cameras output much bigger files and DSLR cameras typically used for professional photo-shoots have up to 50mpix images. And also JPEG compression artifacts are highly visible on current images (basically we could find rough edges of cars by looking at artifacts). It will be interesting to apply segmentation techniques to 50mpix uncompressed RAW file and compare the results :)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208689,
          "author_name": "brianshaler",
          "author_url": "",
          "post_date": "07/30/2017 15:58:14",
          "content": "<p>I don't think Carvana retains original RAW files, but I should be able to acquire JPGs at around ~7500 x ~5000. I'll have to sync up with Kaggle about the process for introducing new source images to the competition.</p>\n\n<p>The manual masks are based on the 1918x1080 versions, so we would not be able to provide masks with higher fidelity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 208702,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "07/30/2017 17:33:30",
          "content": "<p>It would be great! I don't know immediately how we can use higher resolution images, however, it should contain some useful information even if there won't be higher resolution masks. For example, we could use such images to find if there are some systematic errors in manual segmentation or something like that. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 209052,
          "author_name": "hackerpoet",
          "author_url": "",
          "post_date": "08/01/2017 02:44:18",
          "content": "<p>Downsampling higher resolution images back to 1918x1080 would effectively eliminate or greatly reduce jpeg artifacts and noise, and could potentially improve segmentation accuracy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213982,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2017 16:23:39",
          "content": "<p>In future, kaggle can consider providing unlabelled data in the training set. I think weak supervised learning is getting better nowadays.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214009,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "08/15/2017 17:55:53",
          "content": "<p>I think you can use test set to update your model weights as long as you don't touch the data manually.\nThere are some discussion about pseudo labeling - <a href=\"http://forums.fast.ai/t/pseudo-labeling-in-ml/247/6\">http://forums.fast.ai/t/pseudo-labeling-in-ml/247/6</a> \nthough I don't understand how it works. This competition may be a good chance to try it out :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371216,
          "author_name": "amashrabov",
          "author_url": "",
          "post_date": "08/16/2018 08:42:03",
          "content": "<p>It was very great and insightful challenge. Hope, you will consider organizing next similar challenge soon!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 208212,
      "author_name": "kylelee",
      "author_url": "",
      "post_date": "07/29/2017 02:54:49",
      "content": "<p>Could the metric use a log-based Dice or any metric that amplifies the differences near to 1.0? For example, logX(dice*X+smooth)  ) [I haven't really thought this carefully though so it might have certain downsides, e.g. LB feedback information at high scores].  </p>\n\n<p>In the case of X=1.01, log1.01(Dice*1.01+1e-15): Dice=0.0 -&gt; -3471; Dice=0.5 -&gt; -68; Dice=0.991 -&gt; 0.091; Dice=0.999 -&gt; 0.899; Dice=0.9999 -&gt; 0.9899</p>\n\n<p>The main issue is that the segmented object is very large and the \"easy\"/central pixels dominate the calculation, while the real goal of this competition is to make visually appealing cutouts so the boundary (\"harder\") pixels are in practice more important.  Alternatively, a modified Dice metric that puts a higher weight on the boundary pixels of the ground truth object may be useful (though it may complicate computation).</p>",
      "votes": null,
      "replies": [
        {
          "id": 208362,
          "author_name": "dawnbreaker",
          "author_url": "",
          "post_date": "07/29/2017 14:36:28",
          "content": "<p>Agree with you. Another option is Y = - log_10( 1 - dice ). Then Dice=0.9990-&gt;3.000, Dice=0.9999-&gt;4.000</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 210147,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "08/04/2017 12:54:46",
      "content": "<p>Can you give an estimate of the real size of the test set?\nIs it similar to the train set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 211403,
      "author_name": "hsanalytics",
      "author_url": "",
      "post_date": "08/08/2017 20:54:32",
      "content": "<p>What is left for us simple humans when the scores are already that high? hahaha </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 211980,
      "author_name": "selimsef",
      "author_url": "",
      "post_date": "08/10/2017 12:09:08",
      "content": "<p>What is the leaderboard ranking strategy for the same \"rounded\" dice score, for example 0.996?\nLooks like submissions are ranked only by the first submission score for 0.996 range (i.e. additional digits are considered), then nobody moves up inside 0.996. The only way to get higher is to get 0.997. </p>",
      "votes": null,
      "replies": [
        {
          "id": 212034,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "08/10/2017 14:47:55",
          "content": "<p>The leaderboard scores are truncated, not rounded. Behind the scenes we are computing and ranking based on the full precision (i.e. if you see two teams with 0.996, the higher team has either a better score or the exact same score, but submitted it first).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212065,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "08/10/2017 15:55:15",
          "content": "<p>Looking at the leaderboard it doesn't seem to be the case. Rank is not calculated based on the full precision after the first submission that hits  0.99x, unless the score is improved by at least 0.001.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213199,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/14/2017 02:34:46",
          "content": "<p>i think the the ranking is only done \"first time\". I can do a \"sort by public score\" in my submission page to know which submissions are better than others. But it seems that a better score at my submission page doesn't give a better rank in the leaderboard page.  I wonder if other kagglers can confirm this?</p>\n\n<p>[note]:  Also the raw data csv file which can be downloaded from the leaderboard website captures submission that are better then previous ones. If you see that some kagglers have  several 0.996, it means that later 0.996 is better than the preceeding ones. i believe this is not reflected in the ranking page as well. </p>\n\n<p>Hence i think the ranking is done at full precision at the backend. But for some results, the webpage doesn't show it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213205,
          "author_name": "kylelee",
          "author_url": "",
          "post_date": "08/14/2017 03:16:22",
          "content": "<p>Agreed, there's something fishy when there are no obvious shifts within the (currently very large) 0.996 band - only entrants to the band from a lower score have been recorded.  Maybe one can try to save the public LB daily to confirm this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214026,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "08/15/2017 19:05:24",
          "content": "<p>I also agree. It is very odd that there are no changes inside same score band except from lower band.\nMy local dice score was very lower-end of 0.997 - it is unlikely that I'm keeping the 1st position with this score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 215282,
          "author_name": "killthekitten",
          "author_url": "",
          "post_date": "08/20/2017 23:05:57",
          "content": "<p>Confirm, didn't change my position (and constantly going down replaced by other .996ers from 31st to 81st place) on the leaderboard since I've made it to the \".996\". </p>\n\n<p>i.e. my last submission was worth moving from 0.99596 to 0.99638, however, my position didn't change.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 216129,
          "author_name": "dvdbos",
          "author_url": "",
          "post_date": "08/24/2017 12:16:20",
          "content": "<p>Agreed, the public leaderboard is broken. I hope that ranking in \"my submissions\" is at least correct, otherwise we really are in a position where we have no idea about what we are doing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 212940,
      "author_name": "zhanglj",
      "author_url": "",
      "post_date": "08/13/2017 06:09:09",
      "content": "<p>Why not show the last few digits of the score. One submit and get a same score (0.996), but even cannot know whether it has an improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 213148,
          "author_name": "harungunaydin",
          "author_url": "",
          "post_date": "08/13/2017 21:52:02",
          "content": "<p>It's because the test set is not to be used for tuning your hyperparameters. I know it's only 25% of the data, so you might treat it as some kind of a validation set but this way it helps you not to overfit to the data as probably any improvement less than 0.001 is not statistically significant.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213197,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/14/2017 02:26:29",
          "content": "<p>in your submission page, you can do a \"sort by public score\". Then you would know if your latest submission is better than the previous.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213343,
          "author_name": "zhanglj",
          "author_url": "",
          "post_date": "08/14/2017 14:54:00",
          "content": "<p>Good idea! Thanks for your advice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214275,
      "author_name": "gadgysaidoff",
      "author_url": "",
      "post_date": "08/16/2017 13:27:52",
      "content": "<p>Hi William and Brian</p>\n\n<p>Can you please ensure us that private set is clear from \"human errors\"?</p>\n\n<p>Now my network is about 0.997-0.998 and I clearly observe that there are many small inaccuracies in your masks.</p>\n\n<p>It makes all further improvements meaningless and this competition became some kind of \"Lottery competition\" </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 214597,
      "author_name": "joexjmmvhm",
      "author_url": "",
      "post_date": "08/17/2017 14:34:12",
      "content": "<p>It has traditionally been the moving on the leaders board that gives me the motivation to keep working, in practical terms it offers more valuable motivation than the the motivation to win. There have been times in other competitions where I knew I was behind the team in front of me by a couple 10 thousandths, and I said to self, \"self, you can find 2 ten thousandths somewhere.\" This competition does not offer that for folks at the top. I just started, so can find motivation for my next few submissions, but when I do get above about .98, where do I find the motivation to get a just a little bit better model?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 216561,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "08/26/2017 16:55:11",
      "content": "<p>Please, add 2 more digits after dot on LB. As I understand low precision was made to prevent LB probing. But it's not the case here - 100K images + we predict not the classes but masks. So it's almost not possible to probe something useful from LB. In addition LB became useless and it's hard to track your own progress, because currently we tune 4-5 digit after dot. This situation is actually bad for motivation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 223722,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "09/23/2017 06:22:59",
      "content": "<p>Is the split between Public and Private Test Set made randomly?</p>",
      "votes": null,
      "replies": [
        {
          "id": 223936,
          "author_name": "brianshaler",
          "author_url": "",
          "post_date": "09/24/2017 11:01:10",
          "content": "<p>Yes, random.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 225132,
      "author_name": "kylelee",
      "author_url": "",
      "post_date": "09/28/2017 10:32:02",
      "content": "<p>Can the digits be expanded (say to 6 digits) now that the competition is over? It's easier to see the difference in each private LB band that way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 225187,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "09/28/2017 13:13:30",
          "content": "<p>Of course! Done.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225195,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "09/28/2017 13:27:47",
          "content": "<p>And, also can you give some insight about how many actual images were in test set? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225215,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "09/28/2017 14:34:55",
          "content": "<pre><code>  95200 Ignored  \n   3664 Private  \n   1200 Public  \n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225217,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "09/28/2017 14:41:03",
          "content": "<p>So, it was exactly as we thought about it. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 226098,
          "author_name": "dslate",
          "author_url": "",
          "post_date": "10/01/2017 02:47:30",
          "content": "<p>An interesting question: what is the best way to combat hand labeling of test records, a particular problem in computer vision contests where it is perhaps most feasible?  One way is to run two-stage contests, which Kaggle seems to be doing more of lately, which require uploading of final models before the official test data is released.   Another way is to supplement the test data with a very large number of dummy (ignored) records, as was done in this contest.  Personally I prefer the dummy records method, but I'm curious as to what other contestants think about this issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 294543,
      "author_name": "kweonwooj",
      "author_url": "",
      "post_date": "03/12/2018 05:55:31",
      "content": "<p>Dear William Cukierski,</p>\n\n<p>I am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.</p>\n\n<p>Main content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.</p>\n\n<p>I was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives (attention maps on top of the image etc) are possible content candidates.</p>\n\n<p>If using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.</p>\n\n<p>Please let me know, Best regards, Kweonwoo</p>\n\n<p>Same message is posted on a separate discussion\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/51686\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/51686</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "207852": "We sometimes observe that newcomers are discouraged by high scores in competitions, assuming the problem is already solved. You should not be discouraged by scores close to one here:\n\n- The segmentation accuracy necessary to use cutout images for human consumption is very high\n- The difference between good and bad segmentations can be a relatively small number of pixels (that nonetheless makes up a relatively large visual difference)\n\nEvery metric and problem has a unique scale when it comes to the absolute metric value and the meaning carried in the decimal places. Don't fall into the trap of thinking this problem is solved just because scores seem like they're \"close enough\" to one.",
    "207871": "I'm thinking this might need more than three digits by the end... is it automated these days?  I remember it being two yesterday.",
    "207876": "It's manual (we increased it to 3 this morning).\n\nScores are saved at full precision behind the scenes, so we only \"need\" to show as many public leaderboard digits as it takes to make the competition fun. We think it's okay to have temporary, apparent ties if the alternative is leaderboard probing and overfit scores.",
    "207919": "Didn't even realise at first, but I'm happy you are taking steps to combat leaderboard fitting that are effective and super non-intrusive like this one :)",
    "208192": "Is it possible to increase the size of the data? \n\nWith the current set up this problem is somewhere between easy and very easy. \n\nI am expecting top 100 places closer to the end of the competition to have 0.9999+ score.\n\nIt would be more challenging/interesting for participants and probably more useful for the organizers if we had more data for train and test sets. \n\nIs it possible that organizers provide more data? More cars models, different locations, different angles, different background etc.",
    "208212": "Could the metric use a log-based Dice or any metric that amplifies the differences near to 1.0? For example, logX(dice*X+smooth)  ) [I haven't really thought this carefully though so it might have certain downsides, e.g. LB feedback information at high scores].  \n\nIn the case of X=1.01, log1.01(Dice*1.01+1e-15): Dice=0.0 -&gt; -3471; Dice=0.5 -&gt; -68; Dice=0.991 -&gt; 0.091; Dice=0.999 -&gt; 0.899; Dice=0.9999 -&gt; 0.9899\n\nThe main issue is that the segmented object is very large and the \"easy\"/central pixels dominate the calculation, while the real goal of this competition is to make visually appealing cutouts so the boundary (\"harder\") pixels are in practice more important.  Alternatively, a modified Dice metric that puts a higher weight on the boundary pixels of the ground truth object may be useful (though it may complicate computation).",
    "208362": "Agree with you. Another option is Y = - log_10( 1 - dice ). Then Dice=0.9990-&gt;3.000, Dice=0.9999-&gt;4.000",
    "208466": "I don't think it will be practical to provide more manual masks, as they are quite expensive and time-consuming to produce.\n\nWith the initial 100,000+ images in the train + test sets, we may have exhausted the distinct car models in Carvana's inventory. It seems adding additional training images of existing car models would make it easier, rather than harder. As for locations, angles, and backgrounds, these photos are what we have to work with.\n\nAs noted in the pinned Dice score thread (edit: lol, in *this* thread /facepalm), the competition may not be on the 3rd significant digit, the 5th, or higher. Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. The difference between 0.9[..]97 and 0.9[..]98 may be what finally makes the difference between an unusable result and a usable result.\n\nIn the end, I think it's still hard to say if the top result will fall short of being a usable product, or if 100+ teams will have such accurate masks the winner is determined by who is arbitrarily closest to the slightly inconsistent hand-drawn outlines in the test set. If the latter seems apparent early enough in the competition, we can look into a way to up the ante, but I would be concerned about the fairness of moving the goal posts in the middle of the competition.",
    "208487": "Having 100 teams break 0.9999 doesn't really mean anything if they are still visually inaccurate. \n\nYou know, that will be about 250 wrong pixels in a whole image. I bet, it will be extremely accurate and  awesome masks.",
    "208683": "That's a good point! Due to the [human error and inconsistencies in the masks we're training/scoring against][1], the ground truth masks could themselves be inaccurate by an average of 100-300 pixels. The scores are already approaching 0.999 (congrats on hitting 0.995, wow!) but I guess it's a matter of whether we'll continue the pace to 0.9999 similarly quickly or scores will plateau with incremental gains becoming much more difficult.\n\n  [1]: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/37229",
    "208686": "Thanks, it was really easy, just a good ol' Unet :) And there is still a big bag with tricks to apply. \n\nOn a side question: do you have higher resolution images? Even simplest phone cameras output much bigger files and DSLR cameras typically used for professional photo-shoots have up to 50mpix images. And also JPEG compression artifacts are highly visible on current images (basically we could find rough edges of cars by looking at artifacts). It will be interesting to apply segmentation techniques to 50mpix uncompressed RAW file and compare the results :)",
    "208689": "I don't think Carvana retains original RAW files, but I should be able to acquire JPGs at around ~7500 x ~5000. I'll have to sync up with Kaggle about the process for introducing new source images to the competition.\n\nThe manual masks are based on the 1918x1080 versions, so we would not be able to provide masks with higher fidelity.",
    "208702": "It would be great! I don't know immediately how we can use higher resolution images, however, it should contain some useful information even if there won't be higher resolution masks. For example, we could use such images to find if there are some systematic errors in manual segmentation or something like that.",
    "209052": "Downsampling higher resolution images back to 1918x1080 would effectively eliminate or greatly reduce jpeg artifacts and noise, and could potentially improve segmentation accuracy.",
    "209095": "&gt;we increased it to 3 this morning\n\n@William, I do not think it is a good idea. I am sure it is.",
    "210147": "Can you give an estimate of the real size of the test set?\nIs it similar to the train set?",
    "211366": "&gt; **William Cukierski wrote**\n&gt; \n&gt; &gt; It's manual (we increased it to 3 this morning).\n&gt; \n\nWould it be possible to increase the number to 4 (or more) in the last few days of the competition? I am not 100% sure, but I think it would be impossible to mine the LB that much in just a few days, and it would give us a better opportunity to make the decision regarding our final submission, especially since the winner(s) will ultimately be determined based on the fourth or fifth digit after the decimal point.",
    "211403": "What is left for us simple humans when the scores are already that high? hahaha",
    "211980": "What is the leaderboard ranking strategy for the same \"rounded\" dice score, for example 0.996?\nLooks like submissions are ranked only by the first submission score for 0.996 range (i.e. additional digits are considered), then nobody moves up inside 0.996. The only way to get higher is to get 0.997.",
    "212034": "The leaderboard scores are truncated, not rounded. Behind the scenes we are computing and ranking based on the full precision (i.e. if you see two teams with 0.996, the higher team has either a better score or the exact same score, but submitted it first).",
    "212065": "Looking at the leaderboard it doesn't seem to be the case. Rank is not calculated based on the full precision after the first submission that hits  0.99x, unless the score is improved by at least 0.001.",
    "212940": "Why not show the last few digits of the score. One submit and get a same score (0.996), but even cannot know whether it has an improvement.",
    "213148": "It's because the test set is not to be used for tuning your hyperparameters. I know it's only 25% of the data, so you might treat it as some kind of a validation set but this way it helps you not to overfit to the data as probably any improvement less than 0.001 is not statistically significant.",
    "213197": "in your submission page, you can do a \"sort by public score\". Then you would know if your latest submission is better than the previous.",
    "213199": "i think the the ranking is only done \"first time\". I can do a \"sort by public score\" in my submission page to know which submissions are better than others. But it seems that a better score at my submission page doesn't give a better rank in the leaderboard page.  I wonder if other kagglers can confirm this?\n\n[note]:  Also the raw data csv file which can be downloaded from the leaderboard website captures submission that are better then previous ones. If you see that some kagglers have  several 0.996, it means that later 0.996 is better than the preceeding ones. i believe this is not reflected in the ranking page as well. \n\nHence i think the ranking is done at full precision at the backend. But for some results, the webpage doesn't show it?",
    "213205": "Agreed, there's something fishy when there are no obvious shifts within the (currently very large) 0.996 band - only entrants to the band from a lower score have been recorded.  Maybe one can try to save the public LB daily to confirm this.",
    "213343": "Good idea! Thanks for your advice.",
    "213982": "In future, kaggle can consider providing unlabelled data in the training set. I think weak supervised learning is getting better nowadays.",
    "214009": "I think you can use test set to update your model weights as long as you don't touch the data manually.\nThere are some discussion about pseudo labeling - http://forums.fast.ai/t/pseudo-labeling-in-ml/247/6 \nthough I don't understand how it works. This competition may be a good chance to try it out :)",
    "214026": "I also agree. It is very odd that there are no changes inside same score band except from lower band.\nMy local dice score was very lower-end of 0.997 - it is unlikely that I'm keeping the 1st position with this score.",
    "214275": "Hi William and Brian\n\nCan you please ensure us that private set is clear from \"human errors\"?\n\nNow my network is about 0.997-0.998 and I clearly observe that there are many small inaccuracies in your masks.\n\nIt makes all further improvements meaningless and this competition became some kind of \"Lottery competition\"",
    "214597": "It has traditionally been the moving on the leaders board that gives me the motivation to keep working, in practical terms it offers more valuable motivation than the the motivation to win. There have been times in other competitions where I knew I was behind the team in front of me by a couple 10 thousandths, and I said to self, \"self, you can find 2 ten thousandths somewhere.\" This competition does not offer that for folks at the top. I just started, so can find motivation for my next few submissions, but when I do get above about .98, where do I find the motivation to get a just a little bit better model?",
    "215282": "Confirm, didn't change my position (and constantly going down replaced by other .996ers from 31st to 81st place) on the leaderboard since I've made it to the \".996\". \n\ni.e. my last submission was worth moving from 0.99596 to 0.99638, however, my position didn't change.",
    "216129": "Agreed, the public leaderboard is broken. I hope that ranking in \"my submissions\" is at least correct, otherwise we really are in a position where we have no idea about what we are doing!",
    "216561": "Please, add 2 more digits after dot on LB. As I understand low precision was made to prevent LB probing. But it's not the case here - 100K images + we predict not the classes but masks. So it's almost not possible to probe something useful from LB. In addition LB became useless and it's hard to track your own progress, because currently we tune 4-5 digit after dot. This situation is actually bad for motivation.",
    "223722": "Is the split between Public and Private Test Set made randomly?",
    "223936": "Yes, random.",
    "225132": "Can the digits be expanded (say to 6 digits) now that the competition is over? It's easier to see the difference in each private LB band that way.",
    "225187": "Of course! Done.",
    "225195": "And, also can you give some insight about how many actual images were in test set? Thanks!",
    "225215": "95200 Ignored  \n       3664 Private  \n       1200 Public",
    "225217": "So, it was exactly as we thought about it. Thanks!",
    "226098": "An interesting question: what is the best way to combat hand labeling of test records, a particular problem in computer vision contests where it is perhaps most feasible?  One way is to run two-stage contests, which Kaggle seems to be doing more of lately, which require uploading of final models before the official test data is released.   Another way is to supplement the test data with a very large number of dummy (ignored) records, as was done in this contest.  Personally I prefer the dummy records method, but I'm curious as to what other contestants think about this issue.",
    "294543": "Dear William Cukierski,\n\nI am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.\n\nMain content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.\n\nI was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives (attention maps on top of the image etc) are possible content candidates.\n\nIf using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.\n\nPlease let me know, Best regards, Kweonwoo\n\nSame message is posted on a separate discussion\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/51686",
    "371216": "It was very great and insightful challenge. Hope, you will consider organizing next similar challenge soon!"
  },
  "source": "meta"
}