{
  "id": 64736,
  "title": "How to do K-fold cross validation with image segmentation and average the predictions on the test data?",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64736",
  "author_name": "Robin Smits",
  "post_date": "2018-09-01T07:08:22.531000",
  "votes": 5,
  "comment_count": 20,
  "views": 0,
  "content": "<p>For me working with image segmentation is relatively new and I noticed that I have not yet seen or be able to find a complete solution on doing cross validation on the training set and with each model trained making the predictions for the test data. \nSuppose I do a 5 fold CV...that would give me 5 different trained models and 5 different predictions for the test data.</p>\n\n<p>How should I specifically average those 5 different predictions for the test data.</p>\n\n<p>There are some kernels that make averages for the confidence or averages for a bounding box.</p>\n\n<p>How should that be handled for image segmentation? Should the averaging of confidence and bounding boxes be implemented in a custom CV solution .. and how should you then handle the situation where you have 0, 1 or n bounding boxes predicted for a single patient?</p>\n\n<p>I would be very eager to learn what the Kaggle community has to say about this. </p>",
  "messages": [
    {
      "id": 379904,
      "postDate": "2018-09-01T07:08:22.530Z",
      "content": "<p>For me working with image segmentation is relatively new and I noticed that I have not yet seen or be able to find a complete solution on doing cross validation on the training set and with each model trained making the predictions for the test data. \nSuppose I do a 5 fold CV...that would give me 5 different trained models and 5 different predictions for the test data.</p>\n\n<p>How should I specifically average those 5 different predictions for the test data.</p>\n\n<p>There are some kernels that make averages for the confidence or averages for a bounding box.</p>\n\n<p>How should that be handled for image segmentation? Should the averaging of confidence and bounding boxes be implemented in a custom CV solution .. and how should you then handle the situation where you have 0, 1 or n bounding boxes predicted for a single patient?</p>\n\n<p>I would be very eager to learn what the Kaggle community has to say about this. </p>",
      "rawMarkdown": "For me working with image segmentation is relatively new and I noticed that I have not yet seen or be able to find a complete solution on doing cross validation on the training set and with each model trained making the predictions for the test data. \nSuppose I do a 5 fold CV...that would give me 5 different trained models and 5 different predictions for the test data.\n\nHow should I specifically average those 5 different predictions for the test data.\n\nThere are some kernels that make averages for the confidence or averages for a bounding box.\n\nHow should that be handled for image segmentation? Should the averaging of confidence and bounding boxes be implemented in a custom CV solution .. and how should you then handle the situation where you have 0, 1 or n bounding boxes predicted for a single patient?\n\nI would be very eager to learn what the Kaggle community has to say about this. ",
      "votes": 4
    },
    {
      "id": 382133,
      "postDate": "2018-09-05T19:10:12.453Z",
      "content": "<p>@Robin Smits You should also take into account folds that predict no box will be there. Here is one suggestion: If you are using 5 folds, and you have 1 fold that predicts no boxes, you should reduce the confidence of all predictions for that patient by 4/5. If there are 2 folds that predict no boxes, then reduce the confidence by 3/5. If 4 folds predict no box, reduce confidence by 1/5. If all folds predict no box, then your confidence will be 0, i.e. there is no box =)</p>\n\n<p>Then, based on your confidence, begin to shrink your box...If your confidence is 50%, make your bounding box 1/2 the size. Etc.</p>",
      "rawMarkdown": "@Robin Smits You should also take into account folds that predict no box will be there. Here is one suggestion: If you are using 5 folds, and you have 1 fold that predicts no boxes, you should reduce the confidence of all predictions for that patient by 4/5. If there are 2 folds that predict no boxes, then reduce the confidence by 3/5. If 4 folds predict no box, reduce confidence by 1/5. If all folds predict no box, then your confidence will be 0, i.e. there is no box =)\n\nThen, based on your confidence, begin to shrink your box...If your confidence is 50%, make your bounding box 1/2 the size. Etc.",
      "votes": 1
    },
    {
      "id": 380115,
      "postDate": "2018-09-01T18:46:09.763Z",
      "content": "<blockquote>\n  <p>and how should you then handle the situation where you have 0, 1 or n\n  bounding boxes predicted for a single patient?</p>\n</blockquote>\n\n<p>I would suggest, one should divide Sample to train and validation subsamples by patient ID, but not by entries itselfs.</p>",
      "rawMarkdown": "&gt; and how should you then handle the situation where you have 0, 1 or n\n&gt; bounding boxes predicted for a single patient?\n\nI would suggest, one should divide Sample to train and validation subsamples by patient ID, but not by entries itselfs.",
      "votes": 1,
      "replies": [
        {
          "id": 380169,
          "postDate": "2018-09-01T21:45:51.400Z",
          "content": "<p>Hi Ivan,\nThanks for you answer. Yes fully agree to base it on the patientId.</p>\n\n<p>Suppose I would do a 5 fold CV and would get the following predictions for a specific PatientId( only showing the confidence and bounding box ):</p>\n\n<p>Fold1: 0.4730574205545242 577 385 191 189 \nFold2: \nFold3: 0.5730574205545242 577 385 191 289 \nFold4: <br>\nFold5: 0.6730574205545242 577 385 191 389 </p>\n\n<p>Would I just ignore Fold2 and Fold4? And average Folds 1,3 and 5?</p>",
          "rawMarkdown": "Hi Ivan,\nThanks for you answer. Yes fully agree to base it on the patientId.\n\nSuppose I would do a 5 fold CV and would get the following predictions for a specific PatientId( only showing the confidence and bounding box ):\n\nFold1: 0.4730574205545242 577 385 191 189 \nFold2: \nFold3: 0.5730574205545242 577 385 191 289 \nFold4:  \nFold5: 0.6730574205545242 577 385 191 389 \n\nWould I just ignore Fold2 and Fold4? And average Folds 1,3 and 5?"
        },
        {
          "id": 380182,
          "postDate": "2018-09-01T23:23:00.530Z",
          "content": "<p>Sorry, I'm not sure if I understand you correctly. \nWhy you get various predictions for one patientID during CV split performed by patientID?</p>\n\n<p>Do you mean 5 fold CV something just like that - <a href=\"https://machinelearningmastery.com/k-fold-cross-validation/\">https://machinelearningmastery.com/k-fold-cross-validation/</a> ? or something more advanced?</p>",
          "rawMarkdown": "Sorry, I'm not sure if I understand you correctly. \nWhy you get various predictions for one patientID during CV split performed by patientID?\n\nDo you mean 5 fold CV something just like that - https://machinelearningmastery.com/k-fold-cross-validation/ ? or something more advanced?",
          "votes": 1
        },
        {
          "id": 380284,
          "postDate": "2018-09-02T08:15:39.877Z",
          "content": "<p>Hi Ivan,</p>\n\n<p>Yes that's the one. I noticed however that I did not formulate my question detailed enough so likely caused some confusion there. I've updated the question\nIf I do 5 fold CV then the validation part is ok to do. However with 5 trained models I can also make 5 different predictions for the test data.</p>\n\n<p>So the question is specifically about the averaging of these different test data predictions based on patientId.</p>",
          "rawMarkdown": "Hi Ivan,\n\nYes that's the one. I noticed however that I did not formulate my question detailed enough so likely caused some confusion there. I've updated the question\nIf I do 5 fold CV then the validation part is ok to do. However with 5 trained models I can also make 5 different predictions for the test data.\n\nSo the question is specifically about the averaging of these different test data predictions based on patientId.",
          "votes": 1
        },
        {
          "id": 380557,
          "postDate": "2018-09-03T01:41:49.837Z",
          "content": "<p>Hello, Robin. I see your point. If we have to average 5 prediction, we have kind of trade off -  if we take association, we will get big set and low score. If we take intersection, we will get small set (maybe empty set)  and hight probability of low score to. So, we can use aproach from Andrew Ng course (Deep Learning Specialization - &gt; Convolution Neural Net):</p>\n\n<ol>\n<li>For each bounded box of each algorithm  we have probability, that this box equivalent truth box. </li>\n<li>Delete all bounded boxes with confidence &lt; treshhold1. </li>\n<li>Get bounded boxes with highest confidence (best_box). </li>\n<li>Delete all bounded boxes box : IoU(box, best_box)  &gt; treshhold1. </li>\n<li>Go to print 3. Stop, what we don't have boxes to delete. Remained boxes are our prediction. </li>\n</ol>",
          "rawMarkdown": "Hello, Robin. I see your point. If we have to average 5 prediction, we have kind of trade off -  if we take association, we will get big set and low score. If we take intersection, we will get small set (maybe empty set)  and hight probability of low score to. So, we can use aproach from Andrew Ng course (Deep Learning Specialization - &gt; Convolution Neural Net):\n\n 1.  For each bounded box of each algorithm  we have probability, that this box equivalent truth box. \n 2. Delete all bounded boxes with confidence &lt; treshhold1. \n 3. Get bounded boxes with highest confidence (best_box). \n 4. Delete all bounded boxes box : IoU(box, best_box)  &gt; treshhold1. \n 5. Go to print 3. Stop, what we don't have boxes to delete. Remained boxes are our prediction. ",
          "votes": 4
        },
        {
          "id": 380706,
          "postDate": "2018-09-03T09:26:04.927Z",
          "content": "<p>Hi Dmitrij, Thanks! That sounds like a good approach. I will take a further look at it and work it out in my code.</p>",
          "rawMarkdown": "Hi Dmitrij, Thanks! That sounds like a good approach. I will take a further look at it and work it out in my code.",
          "votes": 1
        },
        {
          "id": 380745,
          "postDate": "2018-09-03T11:20:36.563Z",
          "content": "<p>This algorithm called Non-Max Supression (Deep Learning Specialization - &gt; Convolution Neural Net -&gt; 3rd week)</p>",
          "rawMarkdown": "This algorithm called Non-Max Supression (Deep Learning Specialization - &gt; Convolution Neural Net -&gt; 3rd week)",
          "votes": 1
        },
        {
          "id": 380883,
          "postDate": "2018-09-03T16:29:32.453Z",
          "content": "<p><a href=\"/koza4ukdmitrij\">@koza4ukdmitrij</a> In step 4, you mean a different threshold? (Maybe should be \"2\" instead of \"1\"?) Seems strange to use same threshold for confidence and IoU.</p>",
          "rawMarkdown": "@koza4ukdmitrij In step 4, you mean a different threshold? (Maybe should be \"2\" instead of \"1\"?) Seems strange to use same threshold for confidence and IoU.",
          "votes": 2
        },
        {
          "id": 380899,
          "postDate": "2018-09-03T17:22:08.247Z",
          "content": "<p>Yes, I mean treshhold2, thanks) </p>",
          "rawMarkdown": "Yes, I mean treshhold2, thanks) ",
          "votes": 1
        },
        {
          "id": 380912,
          "postDate": "2018-09-03T17:47:02.867Z",
          "content": "<p><a href=\"/koza4ukdmitrij\">@koza4ukdmitrij</a> <a href=\"/aharless\">@aharless</a> Thanks for both your feedback and clarification.</p>",
          "rawMarkdown": "@koza4ukdmitrij @aharless Thanks for both your feedback and clarification."
        },
        {
          "id": 380939,
          "postDate": "2018-09-03T18:51:33.173Z",
          "content": "<p>I found <a href=\"https://www.pyimagesearch.com/2015/02/16/faster-non-maximum-suppression-python/\">an implementation of NMS</a>.  In the example, they start by making an arbitrary choice based on location rather than choosing the box with the highest confidence, but the comments section discusses how confidence would be a better criterion if it is available.</p>",
          "rawMarkdown": "I found [an implementation of NMS][1].  In the example, they start by making an arbitrary choice based on location rather than choosing the box with the highest confidence, but the comments section discusses how confidence would be a better criterion if it is available.\n\n [1]: https://www.pyimagesearch.com/2015/02/16/faster-non-maximum-suppression-python/",
          "votes": 3
        },
        {
          "id": 380994,
          "postDate": "2018-09-03T21:23:17.963Z",
          "content": "<p>Thanks! Looks like good and very usable code.  When I get to using it I will also try it based on confidence.</p>",
          "rawMarkdown": "Thanks! Looks like good and very usable code.  When I get to using it I will also try it based on confidence."
        },
        {
          "id": 381166,
          "postDate": "2018-09-04T07:52:31.193Z",
          "content": "<p>@Andy Harless and @ Robin Smits, the updated version of his NMS implementation that also uses the confidence is in the <a href=\"https://github.com/jrosebr1/imutils/blob/master/imutils/object_detection.py\">imutils</a> tool, a very useful package that I use quite often.</p>",
          "rawMarkdown": "@Andy Harless and @ Robin Smits, the updated version of his NMS implementation that also uses the confidence is in the [imutils][1] tool, a very useful package that I use quite often.\n\n  [1]: https://github.com/jrosebr1/imutils/blob/master/imutils/object_detection.py",
          "votes": 3
        },
        {
          "id": 381839,
          "postDate": "2018-09-05T09:22:01.137Z",
          "content": "<p>Thanks YaGana! I will for sure take a look at it and try it out.</p>",
          "rawMarkdown": "Thanks YaGana! I will for sure take a look at it and try it out."
        }
      ]
    },
    {
      "id": 380881,
      "postDate": "2018-09-03T16:24:15.790Z",
      "content": "<p>It seems the question is really a more general one: how do you combine multiple predictions? As stated in the post, they are predictions from the same model fit on different folds, but the same issues would apply to combining predictions from different models, (Presumably the \"model\" that we are to upload at the end of stage 1 need not use one only method to make predictions, but could use several separate methods and combine the results. Whatever procedure one uses to combine fold predictions, one could also use to combine predictions from different methods.)</p>",
      "rawMarkdown": "It seems the question is really a more general one: how do you combine multiple predictions? As stated in the post, they are predictions from the same model fit on different folds, but the same issues would apply to combining predictions from different models, (Presumably the \"model\" that we are to upload at the end of stage 1 need not use one only method to make predictions, but could use several separate methods and combine the results. Whatever procedure one uses to combine fold predictions, one could also use to combine predictions from different methods.)",
      "votes": 2,
      "replies": [
        {
          "id": 380914,
          "postDate": "2018-09-03T17:50:48.150Z",
          "content": "<p>So the problem of how to blend the results of multiple models is also solved now ;-)\nWill be interresting to see though how combining multiple models will hold up compared to single model results...</p>",
          "rawMarkdown": "So the problem of how to blend the results of multiple models is also solved now ;-)\nWill be interresting to see though how combining multiple models will hold up compared to single model results..."
        }
      ]
    },
    {
      "id": 412990,
      "postDate": "2018-10-31T05:23:39.913Z",
      "content": "<p>Are you guys balancing the (Normal and Not Normal) training inputs to Lung Opacity samples before computing the folds or doing a 'stratified k-fold' where the relative proportion of Normal, Lung Opacity, and Not Normal is preserved in the folds?  I have seen most public kernels doing single models suggest doing a balance but it does not intuitively feel like the right thing to do, even for single models.  My logic is that in real world conditions, diagnosed (train) and new(test) cases will follow the same distribution but the number of positive detections would never be exactly the same as the negative cases.  I would expect the latter to be larger, which is the case in the training data provided.</p>\n\n<p>Thoughts?</p>",
      "rawMarkdown": "Are you guys balancing the (Normal and Not Normal) training inputs to Lung Opacity samples before computing the folds or doing a 'stratified k-fold' where the relative proportion of Normal, Lung Opacity, and Not Normal is preserved in the folds?  I have seen most public kernels doing single models suggest doing a balance but it does not intuitively feel like the right thing to do, even for single models.  My logic is that in real world conditions, diagnosed (train) and new(test) cases will follow the same distribution but the number of positive detections would never be exactly the same as the negative cases.  I would expect the latter to be larger, which is the case in the training data provided.\n\nThoughts?",
      "replies": [
        {
          "id": 413986,
          "postDate": "2018-11-01T22:41:36.360Z",
          "content": "<p>I've used StratifiedKFold mostly when doing CV. Also used the 'normal' K-Fold but I couldn't notice any significant difference.</p>",
          "rawMarkdown": "I've used StratifiedKFold mostly when doing CV. Also used the 'normal' K-Fold but I couldn't notice any significant difference."
        }
      ]
    },
    {
      "id": 412989,
      "postDate": "2018-10-31T05:23:14.740Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 382133,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2018-09-05T19:10:12.453000",
      "content": "<p>@Robin Smits You should also take into account folds that predict no box will be there. Here is one suggestion: If you are using 5 folds, and you have 1 fold that predicts no boxes, you should reduce the confidence of all predictions for that patient by 4/5. If there are 2 folds that predict no boxes, then reduce the confidence by 3/5. If 4 folds predict no box, reduce confidence by 1/5. If all folds predict no box, then your confidence will be 0, i.e. there is no box =)</p>\n\n<p>Then, based on your confidence, begin to shrink your box...If your confidence is 50%, make your bounding box 1/2 the size. Etc.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 380115,
      "author_name": "Ivan Pozdnyakov",
      "author_url": "",
      "post_date": "2018-09-01T18:46:09.763000",
      "content": "<blockquote>\n  <p>and how should you then handle the situation where you have 0, 1 or n\n  bounding boxes predicted for a single patient?</p>\n</blockquote>\n\n<p>I would suggest, one should divide Sample to train and validation subsamples by patient ID, but not by entries itselfs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 380169,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-09-01T21:45:51.400000",
          "content": "<p>Hi Ivan,\nThanks for you answer. Yes fully agree to base it on the patientId.</p>\n\n<p>Suppose I would do a 5 fold CV and would get the following predictions for a specific PatientId( only showing the confidence and bounding box ):</p>\n\n<p>Fold1: 0.4730574205545242 577 385 191 189 \nFold2: \nFold3: 0.5730574205545242 577 385 191 289 \nFold4: <br>\nFold5: 0.6730574205545242 577 385 191 389 </p>\n\n<p>Would I just ignore Fold2 and Fold4? And average Folds 1,3 and 5?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 380182,
          "author_name": "Ivan Pozdnyakov",
          "author_url": "",
          "post_date": "2018-09-01T23:23:00.530000",
          "content": "<p>Sorry, I'm not sure if I understand you correctly. \nWhy you get various predictions for one patientID during CV split performed by patientID?</p>\n\n<p>Do you mean 5 fold CV something just like that - <a href=\"https://machinelearningmastery.com/k-fold-cross-validation/\">https://machinelearningmastery.com/k-fold-cross-validation/</a> ? or something more advanced?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 380284,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-09-02T08:15:39.877000",
          "content": "<p>Hi Ivan,</p>\n\n<p>Yes that's the one. I noticed however that I did not formulate my question detailed enough so likely caused some confusion there. I've updated the question\nIf I do 5 fold CV then the validation part is ok to do. However with 5 trained models I can also make 5 different predictions for the test data.</p>\n\n<p>So the question is specifically about the averaging of these different test data predictions based on patientId.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 380557,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2018-09-03T01:41:49.837000",
          "content": "<p>Hello, Robin. I see your point. If we have to average 5 prediction, we have kind of trade off -  if we take association, we will get big set and low score. If we take intersection, we will get small set (maybe empty set)  and hight probability of low score to. So, we can use aproach from Andrew Ng course (Deep Learning Specialization - &gt; Convolution Neural Net):</p>\n\n<ol>\n<li>For each bounded box of each algorithm  we have probability, that this box equivalent truth box. </li>\n<li>Delete all bounded boxes with confidence &lt; treshhold1. </li>\n<li>Get bounded boxes with highest confidence (best_box). </li>\n<li>Delete all bounded boxes box : IoU(box, best_box)  &gt; treshhold1. </li>\n<li>Go to print 3. Stop, what we don't have boxes to delete. Remained boxes are our prediction. </li>\n</ol>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 380706,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-09-03T09:26:04.927000",
          "content": "<p>Hi Dmitrij, Thanks! That sounds like a good approach. I will take a further look at it and work it out in my code.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 380745,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2018-09-03T11:20:36.563000",
          "content": "<p>This algorithm called Non-Max Supression (Deep Learning Specialization - &gt; Convolution Neural Net -&gt; 3rd week)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 380883,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-09-03T16:29:32.453000",
          "content": "<p><a href=\"/koza4ukdmitrij\">@koza4ukdmitrij</a> In step 4, you mean a different threshold? (Maybe should be \"2\" instead of \"1\"?) Seems strange to use same threshold for confidence and IoU.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 380899,
          "author_name": "Dmitrij Kozachuk",
          "author_url": "",
          "post_date": "2018-09-03T17:22:08.247000",
          "content": "<p>Yes, I mean treshhold2, thanks) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 380912,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-09-03T17:47:02.867000",
          "content": "<p><a href=\"/koza4ukdmitrij\">@koza4ukdmitrij</a> <a href=\"/aharless\">@aharless</a> Thanks for both your feedback and clarification.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 380939,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-09-03T18:51:33.173000",
          "content": "<p>I found <a href=\"https://www.pyimagesearch.com/2015/02/16/faster-non-maximum-suppression-python/\">an implementation of NMS</a>.  In the example, they start by making an arbitrary choice based on location rather than choosing the box with the highest confidence, but the comments section discusses how confidence would be a better criterion if it is available.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 380994,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-09-03T21:23:17.963000",
          "content": "<p>Thanks! Looks like good and very usable code.  When I get to using it I will also try it based on confidence.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 381166,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-09-04T07:52:31.193000",
          "content": "<p>@Andy Harless and @ Robin Smits, the updated version of his NMS implementation that also uses the confidence is in the <a href=\"https://github.com/jrosebr1/imutils/blob/master/imutils/object_detection.py\">imutils</a> tool, a very useful package that I use quite often.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 381839,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-09-05T09:22:01.137000",
          "content": "<p>Thanks YaGana! I will for sure take a look at it and try it out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 380881,
      "author_name": "Andy Harless",
      "author_url": "",
      "post_date": "2018-09-03T16:24:15.790000",
      "content": "<p>It seems the question is really a more general one: how do you combine multiple predictions? As stated in the post, they are predictions from the same model fit on different folds, but the same issues would apply to combining predictions from different models, (Presumably the \"model\" that we are to upload at the end of stage 1 need not use one only method to make predictions, but could use several separate methods and combine the results. Whatever procedure one uses to combine fold predictions, one could also use to combine predictions from different methods.)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 380914,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-09-03T17:50:48.150000",
          "content": "<p>So the problem of how to blend the results of multiple models is also solved now ;-)\nWill be interresting to see though how combining multiple models will hold up compared to single model results...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 412990,
      "author_name": "Kanwalinder Singh",
      "author_url": "",
      "post_date": "2018-10-31T05:23:39.913000",
      "content": "<p>Are you guys balancing the (Normal and Not Normal) training inputs to Lung Opacity samples before computing the folds or doing a 'stratified k-fold' where the relative proportion of Normal, Lung Opacity, and Not Normal is preserved in the folds?  I have seen most public kernels doing single models suggest doing a balance but it does not intuitively feel like the right thing to do, even for single models.  My logic is that in real world conditions, diagnosed (train) and new(test) cases will follow the same distribution but the number of positive detections would never be exactly the same as the negative cases.  I would expect the latter to be larger, which is the case in the training data provided.</p>\n\n<p>Thoughts?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 413986,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2018-11-01T22:41:36.360000",
          "content": "<p>I've used StratifiedKFold mostly when doing CV. Also used the 'normal' K-Fold but I couldn't notice any significant difference.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 412989,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-31T05:23:14.740000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "379904": "For me working with image segmentation is relatively new and I noticed that I have not yet seen or be able to find a complete solution on doing cross validation on the training set and with each model trained making the predictions for the test data. \nSuppose I do a 5 fold CV...that would give me 5 different trained models and 5 different predictions for the test data.\n\nHow should I specifically average those 5 different predictions for the test data.\n\nThere are some kernels that make averages for the confidence or averages for a bounding box.\n\nHow should that be handled for image segmentation? Should the averaging of confidence and bounding boxes be implemented in a custom CV solution .. and how should you then handle the situation where you have 0, 1 or n bounding boxes predicted for a single patient?\n\nI would be very eager to learn what the Kaggle community has to say about this. ",
    "382133": "@Robin Smits You should also take into account folds that predict no box will be there. Here is one suggestion: If you are using 5 folds, and you have 1 fold that predicts no boxes, you should reduce the confidence of all predictions for that patient by 4/5. If there are 2 folds that predict no boxes, then reduce the confidence by 3/5. If 4 folds predict no box, reduce confidence by 1/5. If all folds predict no box, then your confidence will be 0, i.e. there is no box =)\n\nThen, based on your confidence, begin to shrink your box...If your confidence is 50%, make your bounding box 1/2 the size. Etc.",
    "380115": "&gt; and how should you then handle the situation where you have 0, 1 or n\n&gt; bounding boxes predicted for a single patient?\n\nI would suggest, one should divide Sample to train and validation subsamples by patient ID, but not by entries itselfs.",
    "380881": "It seems the question is really a more general one: how do you combine multiple predictions? As stated in the post, they are predictions from the same model fit on different folds, but the same issues would apply to combining predictions from different models, (Presumably the \"model\" that we are to upload at the end of stage 1 need not use one only method to make predictions, but could use several separate methods and combine the results. Whatever procedure one uses to combine fold predictions, one could also use to combine predictions from different methods.)",
    "412990": "Are you guys balancing the (Normal and Not Normal) training inputs to Lung Opacity samples before computing the folds or doing a 'stratified k-fold' where the relative proportion of Normal, Lung Opacity, and Not Normal is preserved in the folds?  I have seen most public kernels doing single models suggest doing a balance but it does not intuitively feel like the right thing to do, even for single models.  My logic is that in real world conditions, diagnosed (train) and new(test) cases will follow the same distribution but the number of positive detections would never be exactly the same as the negative cases.  I would expect the latter to be larger, which is the case in the training data provided.\n\nThoughts?",
    "412989": ""
  }
}