{
  "id": 19351,
  "title": "Just to be 100% sure annotating in 2nd round, reproducibility",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19351",
  "author_name": "Julian de Wit",
  "post_date": "2016-03-07T04:43:01.993000",
  "votes": 1,
  "comment_count": 17,
  "views": 1330,
  "content": "<p>Hello I intend to annotate the to be released validation set.<br>\nBased on the following thread I assume this is allowed.<br>\n<a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18599/questions-about-hand-labeling/109809\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18599/questions-about-hand-labeling/109809</a></p>\n\n<p><br>\nNow for reproducibility..\nCudnn (which almost everyone is using I guess) is non-deterministic in train-mode but deterministic in predict mode.<br>\n<br>\nMust I remove cudnn from my solution since training is non-deterministic.\nIf so.. Well... That would mean I might need to change the 3rd party library I use..\nIs that expected of us ? Since NVIDIA is the sponsor here that would be a BIG bummer for them too.\n<br>\n<br>\nIn the end it only means a very little difference in the final score.\n<br>\n<br>\nPredicting from a trained model IS reproducible.\n<br>\n<br>\nSo to be 100% sure..<br>\n1. Can we annotate patient 500-700 after 7 march ?\n<br>\n<br>\n2. Must the training process be 100% reproducible ? (which means we may not use CUDNN which is provided by the sponsor of this competition)</p>",
  "messages": [
    {
      "id": 110702,
      "postDate": "2016-03-07T17:11:23.143Z",
      "content": "<p>Cherry picking cannot be easily observed.  For example, it's possible for one to create one batch of hand annotations, and systematically replicate that set 100 times with different factors of enlargement or shrinking.  One can then train 100 models and picked the best to submit.  The organizer won't be able to find out:</p>\n\n<ul>\n<li>Whether the test set is hand annotated for location evaluation to find the best model.</li>\n<li>Whether the submitted hand annotation has been systematically altered.</li>\n</ul>\n\n<p>The same applies to training with a wall-time limitation instead of number of iterations.  One can easily mimic different number of iterations by using machines of different speed.</p>\n\n<p>I'm just imagining the ways the rules and evaluation system can be attacked.  After all many of the top teams are close, and the difference might only lie in whether a handful of border line corner cases are properly handled or not.  Considering all test images will be made public, tweaking the system to eliminate one or a few corner cases is totally possible.</p>",
      "rawMarkdown": "Cherry picking cannot be easily observed.  For example, it's possible for one to create one batch of hand annotations, and systematically replicate that set 100 times with different factors of enlargement or shrinking.  One can then train 100 models and picked the best to submit.  The organizer won't be able to find out:\r\n\r\n- Whether the test set is hand annotated for location evaluation to find the best model.\r\n- Whether the submitted hand annotation has been systematically altered.\r\n\r\nThe same applies to training with a wall-time limitation instead of number of iterations.  One can easily mimic different number of iterations by using machines of different speed.\r\n\r\nI'm just imagining the ways the rules and evaluation system can be attacked.  After all many of the top teams are close, and the difference might only lie in whether a handful of border line corner cases are properly handled or not.  Considering all test images will be made public, tweaking the system to eliminate one or a few corner cases is totally possible.\r\n\r\n",
      "votes": 1
    },
    {
      "id": 110624,
      "postDate": "2016-03-07T04:43:01.993Z",
      "content": "<p>Hello I intend to annotate the to be released validation set.<br>\nBased on the following thread I assume this is allowed.<br>\n<a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18599/questions-about-hand-labeling/109809\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18599/questions-about-hand-labeling/109809</a></p>\n\n<p><br>\nNow for reproducibility..\nCudnn (which almost everyone is using I guess) is non-deterministic in train-mode but deterministic in predict mode.<br>\n<br>\nMust I remove cudnn from my solution since training is non-deterministic.\nIf so.. Well... That would mean I might need to change the 3rd party library I use..\nIs that expected of us ? Since NVIDIA is the sponsor here that would be a BIG bummer for them too.\n<br>\n<br>\nIn the end it only means a very little difference in the final score.\n<br>\n<br>\nPredicting from a trained model IS reproducible.\n<br>\n<br>\nSo to be 100% sure..<br>\n1. Can we annotate patient 500-700 after 7 march ?\n<br>\n<br>\n2. Must the training process be 100% reproducible ? (which means we may not use CUDNN which is provided by the sponsor of this competition)</p>",
      "rawMarkdown": "Hello I intend to annotate the to be released validation set.<br>\r\nBased on the following thread I assume this is allowed.<br>\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18599/questions-about-hand-labeling/109809\r\n\r\n<br>\r\nNow for reproducibility..\r\nCudnn (which almost everyone is using I guess) is non-deterministic in train-mode but deterministic in predict mode.<br>\r\n<br>\r\nMust I remove cudnn from my solution since training is non-deterministic.\r\nIf so.. Well... That would mean I might need to change the 3rd party library I use..\r\nIs that expected of us ? Since NVIDIA is the sponsor here that would be a BIG bummer for them too.\r\n<br>\r\n<br>\r\nIn the end it only means a very little difference in the final score.\r\n<br>\r\n<br>\r\nPredicting from a trained model IS reproducible.\r\n<br>\r\n<br>\r\nSo to be 100% sure..<br>\r\n1. Can we annotate patient 500-700 after 7 march ?\r\n<br>\r\n<br>\r\n2. Must the training process be 100% reproducible ? (which means we may not use CUDNN which is provided by the sponsor of this competition)\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n",
      "votes": 1
    },
    {
      "id": 110695,
      "postDate": "2016-03-07T16:41:16.350Z",
      "content": "<p>I hope that the Kaggle team will give some credit to those of us who refuse to do any hand-labeling/annotation in this competition.</p>\n\n<p>Personally, I think it would be somewhat disappointing to see the prizes go away to teams who won based on some sort of manual labor.</p>\n\n<p>I see this competition as a great opportunity to put GPU computing and machine learning to the test, and helping out the machine with any hand-labeling seems to defeat that purpose.</p>\n\n<p>In any case, what I'm looking for is to check the ability of deep learning to surpass human-level performance. If hand-labeling wins, then I guess we are still somewhat far from that goal.</p>",
      "rawMarkdown": "I hope that the Kaggle team will give some credit to those of us who refuse to do any hand-labeling/annotation in this competition.\r\n\r\nPersonally, I think it would be somewhat disappointing to see the prizes go away to teams who won based on some sort of manual labor.\r\n\r\nI see this competition as a great opportunity to put GPU computing and machine learning to the test, and helping out the machine with any hand-labeling seems to defeat that purpose.\r\n\r\nIn any case, what I'm looking for is to check the ability of deep learning to surpass human-level performance. If hand-labeling wins, then I guess we are still somewhat far from that goal."
    },
    {
      "id": 110691,
      "postDate": "2016-03-07T15:32:59.540Z",
      "content": "<p>Cherry picking training cases would violate rules regarding hand labeling. If we see such preferential manual selection we will question it.</p>",
      "rawMarkdown": "Cherry picking training cases would violate rules regarding hand labeling. If we see such preferential manual selection we will question it."
    },
    {
      "id": 110686,
      "postDate": "2016-03-07T14:43:25.270Z",
      "content": "<p>@yuanfang guan.That's what I was thinking too, one can pick cases from validate set to essentially change the model.. We are so lazy to hand label many images (we only did few) and given that we are on the top spot we don't bother to waste time to do more hand labeling training images. However, once we are in stage two and there is nothing to do so the only thing left we might do is to add more training images from the validation set....I hope it is disallowed to save everybody's time..:) but I'm okay with it. \nI hope the test data set is large enough (~500 cases?) so the final Test LB is stable.</p>",
      "rawMarkdown": "@yuanfang guan.That's what I was thinking too, one can pick cases from validate set to essentially change the model.. We are so lazy to hand label many images (we only did few) and given that we are on the top spot we don't bother to waste time to do more hand labeling training images. However, once we are in stage two and there is nothing to do so the only thing left we might do is to add more training images from the validation set....I hope it is disallowed to save everybody's time..:) but I'm okay with it. \r\nI hope the test data set is large enough (~500 cases?) so the final Test LB is stable."
    },
    {
      "id": 110681,
      "postDate": "2016-03-07T14:22:27.427Z",
      "content": "<p>Thank you very much. \nhow large is the test dataset? If it is over like 2000 cases then I guess we might need to remove some results from our network for it to be able to finish in a week..</p>",
      "rawMarkdown": "Thank you very much. \r\nhow large is the test dataset? If it is over like 2000 cases then I guess we might need to remove some results from our network for it to be able to finish in a week..\r\n"
    },
    {
      "id": 110679,
      "postDate": "2016-03-07T14:08:09.807Z",
      "content": "<p>Correct. However, you will not have leaderboard feedback on the test data and are not allowed to hand annotate the test set, so you will not have the ability to pick one that is better than another.</p>\n\n<p>The reason we have to allow this is because it's theoretically possible for a team to already have the validation set labels/annotations in hand (by means of using the leaderboard) during stage one. This allowance is to neutralize that potential unfair advantage.</p>",
      "rawMarkdown": "Correct. However, you will not have leaderboard feedback on the test data and are not allowed to hand annotate the test set, so you will not have the ability to pick one that is better than another.\r\n\r\nThe reason we have to allow this is because it's theoretically possible for a team to already have the validation set labels/annotations in hand (by means of using the leaderboard) during stage one. This allowance is to neutralize that potential unfair advantage."
    },
    {
      "id": 110678,
      "postDate": "2016-03-07T13:49:57.993Z",
      "content": "<p>[quote=William Cukierski;110666]</p>\n\n<p>Apologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set.</p>\n\n<p>[/quote]\nI see....I thought we can't...\nThen to be 100% sure too, we are allowed to do the following:\n1) validate set is released.\n2)we submit a result without adding any new images to do training. Just use our old trained networks.\n3) we make a submission and it is not super good that we do not win the competition.\n4) we'll start to add more hand labeled images from the validate set (depends on the time we have, we may hand label 10 images or 100 images) to retrain our networks, to improve the networks performance. (And we can pick which submission to use as the final ranking submission)</p>",
      "rawMarkdown": "[quote=William Cukierski;110666]\r\n\r\nApologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set.\r\n\r\n[/quote]\r\nI see....I thought we can't...\r\nThen to be 100% sure too, we are allowed to do the following:\r\n1) validate set is released.\r\n2)we submit a result without adding any new images to do training. Just use our old trained networks.\r\n3) we make a submission and it is not super good that we do not win the competition.\r\n4) we'll start to add more hand labeled images from the validate set (depends on the time we have, we may hand label 10 images or 100 images) to retrain our networks, to improve the networks performance. (And we can pick which submission to use as the final ranking submission)"
    },
    {
      "id": 110674,
      "postDate": "2016-03-07T13:14:43.573Z",
      "content": "<p>Correct, that would not be allowed unless it was automated.</p>",
      "rawMarkdown": "Correct, that would not be allowed unless it was automated."
    },
    {
      "id": 110672,
      "postDate": "2016-03-07T13:12:12.407Z",
      "content": "<p>Yes but I meant for instance </p>\n\n<pre><code>If (smooth_my_predictions):\n    smooth...\n</code></pre>\n\n<p>Toggling from true (uploaded model) to false (final submission).</p>\n\n<p>My code is full of these kind of options that I tried.</p>\n\n<p>That is NOT allowed is it ?</p>",
      "rawMarkdown": "Yes but I meant for instance \r\n\r\n    If (smooth_my_predictions):\r\n        smooth...\r\n\r\nToggling from true (uploaded model) to false (final submission).\r\n\r\nMy code is full of these kind of options that I tried.\r\n\r\nThat is NOT allowed is it ?"
    },
    {
      "id": 110670,
      "postDate": "2016-03-07T13:08:46.447Z",
      "content": "<p>You got it.</p>\n\n<p>Changing code paths would be fine (it's a change to get the code to run and not a change to the science). </p>",
      "rawMarkdown": "You got it.\r\n\r\nChanging code paths would be fine (it's a change to get the code to run and not a change to the science). "
    },
    {
      "id": 110668,
      "postDate": "2016-03-07T13:05:36.893Z",
      "content": "<p>Just to be 1000% sure.</p>\n\n<p>I intend to:</p>\n\n<ul>\n<li>Adjust the code for patient 1-700 = train and 701-XXX = submission.\n\n<ul><li>Add patient 500-700 handlabels to the trainset </li>\n<li>Train </li>\n<li>Submit</li></ul></li>\n</ul>\n\n<p>No parameter tuning allowed.</p>\n\n<p>No changing of code paths allowed. </p>\n\n<p>Train + submit needs to be 100% reproducible</p>",
      "rawMarkdown": "Just to be 1000% sure.\r\n\r\nI intend to:\r\n\r\n - Adjust the code for patient 1-700 = train and 701-XXX = submission.\r\n-  Add patient 500-700 handlabels to the trainset \r\n- Train \r\n- Submit\r\n\r\nNo parameter tuning allowed.\r\n\r\nNo changing of code paths allowed. \r\n\r\nTrain + submit needs to be 100% reproducible\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n"
    },
    {
      "id": 110666,
      "postDate": "2016-03-07T12:57:50.757Z",
      "content": "<p>Apologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set.</p>",
      "rawMarkdown": "Apologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set."
    },
    {
      "id": 110664,
      "postDate": "2016-03-07T12:45:09.343Z",
      "content": "<p>Quoting William:</p>\n\n<blockquote>\n  <p>During phase one, you can hand label the training set only. During\n  phase two, after we have released the validation set labels, you can\n  hand label the training and validation set only.</p>\n</blockquote>\n\n<p>I barely used the provided volumes :S</p>\n\n<p>As a matter of fact I already annotated patient 500-700 today so I could upload them with my code.\nThen I do not label anymore after today. I guess that would be fair. But that would mean I voilated rule 1.</p>\n\n<p>I don't want to steal your victory of course but I do want to get the most out of my model.</p>\n\n<p>Good luck, you guys deserve the victory but I'm rooting for myself :)</p>",
      "rawMarkdown": "Quoting William:\r\n\r\n> During phase one, you can hand label the training set only. During\r\n> phase two, after we have released the validation set labels, you can\r\n> hand label the training and validation set only.\r\n\r\nI barely used the provided volumes :S\r\n\r\nAs a matter of fact I already annotated patient 500-700 today so I could upload them with my code.\r\nThen I do not label anymore after today. I guess that would be fair. But that would mean I voilated rule 1.\r\n\r\nI don't want to steal your victory of course but I do want to get the most out of my model.\r\n\r\nGood luck, you guys deserve the victory but I'm rooting for myself :)\r\n\r\n\r\n"
    },
    {
      "id": 110662,
      "postDate": "2016-03-07T12:37:50.910Z",
      "content": "<p>from my understanding for 1, I think you can't annotate patient 500-700. You can only use the true volume from 500-700 to train your model (automatically), and you should not manually annotate more images.\nThat will make each's model pretty unpredictable, because you could decide which images to to annotate.</p>",
      "rawMarkdown": "from my understanding for 1, I think you can't annotate patient 500-700. You can only use the true volume from 500-700 to train your model (automatically), and you should not manually annotate more images.\r\nThat will make each's model pretty unpredictable, because you could decide which images to to annotate.\r\n"
    },
    {
      "id": 110656,
      "postDate": "2016-03-07T10:29:07.460Z",
      "content": "<p>I fixed point 2 myself already.\nHowever.. in the future this will be infeasible I think.</p>\n\n<p>The question about the labeling remains</p>",
      "rawMarkdown": "I fixed point 2 myself already.\r\nHowever.. in the future this will be infeasible I think.\r\n\r\nThe question about the labeling remains"
    },
    {
      "id": 110687,
      "postDate": "2016-03-07T14:47:23.977Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 110684,
      "postDate": "2016-03-07T14:27:23.167Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 110702,
      "author_name": "Wei Dong",
      "author_url": "",
      "post_date": "2016-03-07T17:11:23.143000",
      "content": "<p>Cherry picking cannot be easily observed.  For example, it's possible for one to create one batch of hand annotations, and systematically replicate that set 100 times with different factors of enlargement or shrinking.  One can then train 100 models and picked the best to submit.  The organizer won't be able to find out:</p>\n\n<ul>\n<li>Whether the test set is hand annotated for location evaluation to find the best model.</li>\n<li>Whether the submitted hand annotation has been systematically altered.</li>\n</ul>\n\n<p>The same applies to training with a wall-time limitation instead of number of iterations.  One can easily mimic different number of iterations by using machines of different speed.</p>\n\n<p>I'm just imagining the ways the rules and evaluation system can be attacked.  After all many of the top teams are close, and the difference might only lie in whether a handful of border line corner cases are properly handled or not.  Considering all test images will be made public, tweaking the system to eliminate one or a few corner cases is totally possible.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110695,
      "author_name": "Diogo R. Ferreira",
      "author_url": "",
      "post_date": "2016-03-07T16:41:16.350000",
      "content": "<p>I hope that the Kaggle team will give some credit to those of us who refuse to do any hand-labeling/annotation in this competition.</p>\n\n<p>Personally, I think it would be somewhat disappointing to see the prizes go away to teams who won based on some sort of manual labor.</p>\n\n<p>I see this competition as a great opportunity to put GPU computing and machine learning to the test, and helping out the machine with any hand-labeling seems to defeat that purpose.</p>\n\n<p>In any case, what I'm looking for is to check the ability of deep learning to surpass human-level performance. If hand-labeling wins, then I guess we are still somewhat far from that goal.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110691,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2016-03-07T15:32:59.540000",
      "content": "<p>Cherry picking training cases would violate rules regarding hand labeling. If we see such preferential manual selection we will question it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110686,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-03-07T14:43:25.270000",
      "content": "<p>@yuanfang guan.That's what I was thinking too, one can pick cases from validate set to essentially change the model.. We are so lazy to hand label many images (we only did few) and given that we are on the top spot we don't bother to waste time to do more hand labeling training images. However, once we are in stage two and there is nothing to do so the only thing left we might do is to add more training images from the validation set....I hope it is disallowed to save everybody's time..:) but I'm okay with it. \nI hope the test data set is large enough (~500 cases?) so the final Test LB is stable.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110681,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-03-07T14:22:27.427000",
      "content": "<p>Thank you very much. \nhow large is the test dataset? If it is over like 2000 cases then I guess we might need to remove some results from our network for it to be able to finish in a week..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110679,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2016-03-07T14:08:09.807000",
      "content": "<p>Correct. However, you will not have leaderboard feedback on the test data and are not allowed to hand annotate the test set, so you will not have the ability to pick one that is better than another.</p>\n\n<p>The reason we have to allow this is because it's theoretically possible for a team to already have the validation set labels/annotations in hand (by means of using the leaderboard) during stage one. This allowance is to neutralize that potential unfair advantage.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110678,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-03-07T13:49:57.993000",
      "content": "<p>[quote=William Cukierski;110666]</p>\n\n<p>Apologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set.</p>\n\n<p>[/quote]\nI see....I thought we can't...\nThen to be 100% sure too, we are allowed to do the following:\n1) validate set is released.\n2)we submit a result without adding any new images to do training. Just use our old trained networks.\n3) we make a submission and it is not super good that we do not win the competition.\n4) we'll start to add more hand labeled images from the validate set (depends on the time we have, we may hand label 10 images or 100 images) to retrain our networks, to improve the networks performance. (And we can pick which submission to use as the final ranking submission)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110674,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2016-03-07T13:14:43.573000",
      "content": "<p>Correct, that would not be allowed unless it was automated.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110672,
      "author_name": "Julian de Wit",
      "author_url": "",
      "post_date": "2016-03-07T13:12:12.407000",
      "content": "<p>Yes but I meant for instance </p>\n\n<pre><code>If (smooth_my_predictions):\n    smooth...\n</code></pre>\n\n<p>Toggling from true (uploaded model) to false (final submission).</p>\n\n<p>My code is full of these kind of options that I tried.</p>\n\n<p>That is NOT allowed is it ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110670,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2016-03-07T13:08:46.447000",
      "content": "<p>You got it.</p>\n\n<p>Changing code paths would be fine (it's a change to get the code to run and not a change to the science). </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110668,
      "author_name": "Julian de Wit",
      "author_url": "",
      "post_date": "2016-03-07T13:05:36.893000",
      "content": "<p>Just to be 1000% sure.</p>\n\n<p>I intend to:</p>\n\n<ul>\n<li>Adjust the code for patient 1-700 = train and 701-XXX = submission.\n\n<ul><li>Add patient 500-700 handlabels to the trainset </li>\n<li>Train </li>\n<li>Submit</li></ul></li>\n</ul>\n\n<p>No parameter tuning allowed.</p>\n\n<p>No changing of code paths allowed. </p>\n\n<p>Train + submit needs to be 100% reproducible</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110666,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2016-03-07T12:57:50.757000",
      "content": "<p>Apologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110664,
      "author_name": "Julian de Wit",
      "author_url": "",
      "post_date": "2016-03-07T12:45:09.343000",
      "content": "<p>Quoting William:</p>\n\n<blockquote>\n  <p>During phase one, you can hand label the training set only. During\n  phase two, after we have released the validation set labels, you can\n  hand label the training and validation set only.</p>\n</blockquote>\n\n<p>I barely used the provided volumes :S</p>\n\n<p>As a matter of fact I already annotated patient 500-700 today so I could upload them with my code.\nThen I do not label anymore after today. I guess that would be fair. But that would mean I voilated rule 1.</p>\n\n<p>I don't want to steal your victory of course but I do want to get the most out of my model.</p>\n\n<p>Good luck, you guys deserve the victory but I'm rooting for myself :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110662,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-03-07T12:37:50.910000",
      "content": "<p>from my understanding for 1, I think you can't annotate patient 500-700. You can only use the true volume from 500-700 to train your model (automatically), and you should not manually annotate more images.\nThat will make each's model pretty unpredictable, because you could decide which images to to annotate.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110656,
      "author_name": "Julian de Wit",
      "author_url": "",
      "post_date": "2016-03-07T10:29:07.460000",
      "content": "<p>I fixed point 2 myself already.\nHowever.. in the future this will be infeasible I think.</p>\n\n<p>The question about the labeling remains</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110687,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-07T14:47:23.977000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110684,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-07T14:27:23.167000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "110702": "Cherry picking cannot be easily observed.  For example, it's possible for one to create one batch of hand annotations, and systematically replicate that set 100 times with different factors of enlargement or shrinking.  One can then train 100 models and picked the best to submit.  The organizer won't be able to find out:\r\n\r\n- Whether the test set is hand annotated for location evaluation to find the best model.\r\n- Whether the submitted hand annotation has been systematically altered.\r\n\r\nThe same applies to training with a wall-time limitation instead of number of iterations.  One can easily mimic different number of iterations by using machines of different speed.\r\n\r\nI'm just imagining the ways the rules and evaluation system can be attacked.  After all many of the top teams are close, and the difference might only lie in whether a handful of border line corner cases are properly handled or not.  Considering all test images will be made public, tweaking the system to eliminate one or a few corner cases is totally possible.\r\n\r\n",
    "110624": "Hello I intend to annotate the to be released validation set.<br>\r\nBased on the following thread I assume this is allowed.<br>\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18599/questions-about-hand-labeling/109809\r\n\r\n<br>\r\nNow for reproducibility..\r\nCudnn (which almost everyone is using I guess) is non-deterministic in train-mode but deterministic in predict mode.<br>\r\n<br>\r\nMust I remove cudnn from my solution since training is non-deterministic.\r\nIf so.. Well... That would mean I might need to change the 3rd party library I use..\r\nIs that expected of us ? Since NVIDIA is the sponsor here that would be a BIG bummer for them too.\r\n<br>\r\n<br>\r\nIn the end it only means a very little difference in the final score.\r\n<br>\r\n<br>\r\nPredicting from a trained model IS reproducible.\r\n<br>\r\n<br>\r\nSo to be 100% sure..<br>\r\n1. Can we annotate patient 500-700 after 7 march ?\r\n<br>\r\n<br>\r\n2. Must the training process be 100% reproducible ? (which means we may not use CUDNN which is provided by the sponsor of this competition)\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n",
    "110695": "I hope that the Kaggle team will give some credit to those of us who refuse to do any hand-labeling/annotation in this competition.\r\n\r\nPersonally, I think it would be somewhat disappointing to see the prizes go away to teams who won based on some sort of manual labor.\r\n\r\nI see this competition as a great opportunity to put GPU computing and machine learning to the test, and helping out the machine with any hand-labeling seems to defeat that purpose.\r\n\r\nIn any case, what I'm looking for is to check the ability of deep learning to surpass human-level performance. If hand-labeling wins, then I guess we are still somewhat far from that goal.",
    "110691": "Cherry picking training cases would violate rules regarding hand labeling. If we see such preferential manual selection we will question it.",
    "110686": "@yuanfang guan.That's what I was thinking too, one can pick cases from validate set to essentially change the model.. We are so lazy to hand label many images (we only did few) and given that we are on the top spot we don't bother to waste time to do more hand labeling training images. However, once we are in stage two and there is nothing to do so the only thing left we might do is to add more training images from the validation set....I hope it is disallowed to save everybody's time..:) but I'm okay with it. \r\nI hope the test data set is large enough (~500 cases?) so the final Test LB is stable.",
    "110681": "Thank you very much. \r\nhow large is the test dataset? If it is over like 2000 cases then I guess we might need to remove some results from our network for it to be able to finish in a week..\r\n",
    "110679": "Correct. However, you will not have leaderboard feedback on the test data and are not allowed to hand annotate the test set, so you will not have the ability to pick one that is better than another.\r\n\r\nThe reason we have to allow this is because it's theoretically possible for a team to already have the validation set labels/annotations in hand (by means of using the leaderboard) during stage one. This allowance is to neutralize that potential unfair advantage.",
    "110678": "[quote=William Cukierski;110666]\r\n\r\nApologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set.\r\n\r\n[/quote]\r\nI see....I thought we can't...\r\nThen to be 100% sure too, we are allowed to do the following:\r\n1) validate set is released.\r\n2)we submit a result without adding any new images to do training. Just use our old trained networks.\r\n3) we make a submission and it is not super good that we do not win the competition.\r\n4) we'll start to add more hand labeled images from the validate set (depends on the time we have, we may hand label 10 images or 100 images) to retrain our networks, to improve the networks performance. (And we can pick which submission to use as the final ranking submission)",
    "110674": "Correct, that would not be allowed unless it was automated.",
    "110672": "Yes but I meant for instance \r\n\r\n    If (smooth_my_predictions):\r\n        smooth...\r\n\r\nToggling from true (uploaded model) to false (final submission).\r\n\r\nMy code is full of these kind of options that I tried.\r\n\r\nThat is NOT allowed is it ?",
    "110670": "You got it.\r\n\r\nChanging code paths would be fine (it's a change to get the code to run and not a change to the science). ",
    "110668": "Just to be 1000% sure.\r\n\r\nI intend to:\r\n\r\n - Adjust the code for patient 1-700 = train and 701-XXX = submission.\r\n-  Add patient 500-700 handlabels to the trainset \r\n- Train \r\n- Submit\r\n\r\nNo parameter tuning allowed.\r\n\r\nNo changing of code paths allowed. \r\n\r\nTrain + submit needs to be 100% reproducible\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n",
    "110666": "Apologies for any confusion here. These two stage rules get complicated. You may annotate the validation set because it is similar/equivalent to retraining on the validation set.",
    "110664": "Quoting William:\r\n\r\n> During phase one, you can hand label the training set only. During\r\n> phase two, after we have released the validation set labels, you can\r\n> hand label the training and validation set only.\r\n\r\nI barely used the provided volumes :S\r\n\r\nAs a matter of fact I already annotated patient 500-700 today so I could upload them with my code.\r\nThen I do not label anymore after today. I guess that would be fair. But that would mean I voilated rule 1.\r\n\r\nI don't want to steal your victory of course but I do want to get the most out of my model.\r\n\r\nGood luck, you guys deserve the victory but I'm rooting for myself :)\r\n\r\n\r\n",
    "110662": "from my understanding for 1, I think you can't annotate patient 500-700. You can only use the true volume from 500-700 to train your model (automatically), and you should not manually annotate more images.\r\nThat will make each's model pretty unpredictable, because you could decide which images to to annotate.\r\n",
    "110656": "I fixed point 2 myself already.\r\nHowever.. in the future this will be infeasible I think.\r\n\r\nThe question about the labeling remains",
    "110687": "",
    "110684": ""
  }
}