{
  "id": 19524,
  "title": "Summary of our method",
  "url": "/competitions/second-annual-data-science-bowl/writeups/tencia-woshialex-summary-of-our-method",
  "author_name": "",
  "post_date": "2016-03-15T06:37:04.330Z",
  "votes": 16,
  "comment_count": 14,
  "views": 5994,
  "content": "<p>Hello everyone,</p>\n\n<p>We wanted to share a quick summary of our method. We had a great time on this competition and learned a lot by working with this data, which is interesting by itself and we&#8217;re very glad it is also potentially useful as a real life medical problem. We will briefly explain our approach here and post full documentation in a couple of days.</p>\n\n<p>Our method uses an average of 10 different fully convolutional neural networks (different architecture, input image size, or pre-processing transformation on the image), which output segmentation results for individual MRI images. We trained on the Sunnybrook data set which has one MRI image as input and the contour of the LV as output. We manually added to the training image set some additional images (in total ~200 images). Adding these improves the performance of the network significantly, and we believe the result can still be significantly improved if more training images are added. We found ensembling only improves the result slightly (something like 0.0096 for the best single network to 0.0093 in train set).</p>\n\n<p>For the preprocessing, we tried a combination of different things:</p>\n\n<pre><code>1.  Use the time variance of the images to determine a preliminary center of the LV and bounding box and crop from the center. \n\n2.  We rotated the images so all the cases are aligned to the same direction.\n\n3.  For the input image augmentation, we did random rotation, shift, and contrast normalization.\n</code></pre>\n\n<p>Models were trained on two GPUs, NVIDIA GTX 970 and 980Ti. The entire model takes about 4 days to train and evaluate if both GPUs are used. We used Python, Theano, Lasagne, and cuDNN for neural network implementation.</p>\n\n<p>The CNNs can detect the contour amazingly well for high-quality images. Once we have the contours, the volume is calculated basically as V_preliminary = \\sum (area_i*thickness) and maximum and minimum volume can be determined. We hypothesized that much of the error comes from the end slices, where a human can decide to include it or not. With this in mind, we did a final fitting to correct some error, with a function V_pred = V - \\beta * sqrt(V) and it does much better than a simple linear fitting.</p>\n\n<p>To deal with some extreme cases that our CNN model cannot predict good volumes for, we developed </p>\n\n<pre><code>1.  a sex-age model ( score ~ 0.036 ) (basically the same as https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18375/0-036023-score-without-looking-at-the-images)\n\n2.   a model based only on a single SAX slice. Score ~ 0.015\n\n3.  A model based on the 4-chamber view. We hand labeled many of these images to train this model. Score ~ 0.017\n\n4.  We took the average of these three models (which scores ~ 0.013) as the default model, if our SAX-based CNN model fails, we took the result from this model. \n</code></pre>\n\n<p>We also tried a Fourier-based segmentation method which gives a score about 0.016. Since it is quite complicated and does significantly worse than the CNN, after the early stages of the competition we dropped this model entirely, to simplify our work and code.</p>\n\n<p>Some observations:</p>\n\n<ol>\n<li><p>A lot of our effort was spent cleaning up data and dealing with edge cases. As kunsthart found, much of the error came from just a few cases, and in one case fixing a single prediction in the validation set dropped our score from ~ 0.0102 to 0.0098.</p></li>\n<li><p>The CNN architectures we used were relatively small. We found that adding capacity did not improve results, though we did not experiment much with different types of architectures or activation functions.</p></li>\n<li><p>Batch normalization helped enormously, as did using a modification of the Sorenson-Dice Index as the segmentation objective function, rather than binary cross-entropy.</p></li>\n<li><p>We did no test-time augmentation due to time constraints, this might have helped.</p></li>\n</ol>\n\n<p>We&#8217;ll write up a full documentation of our model in a couple of days. Thanks to Kaggle, Booz Allen Hamilton, and the administrators for running such an interesting and engaging competition, and to our fellow competitors :)</p>",
  "messages": [
    {
      "id": "111523",
      "postDate": "03/15/2016 06:37:04",
      "content": "<p>Hello everyone,</p>\n\n<p>We wanted to share a quick summary of our method. We had a great time on this competition and learned a lot by working with this data, which is interesting by itself and we&#8217;re very glad it is also potentially useful as a real life medical problem. We will briefly explain our approach here and post full documentation in a couple of days.</p>\n\n<p>Our method uses an average of 10 different fully convolutional neural networks (different architecture, input image size, or pre-processing transformation on the image), which output segmentation results for individual MRI images. We trained on the Sunnybrook data set which has one MRI image as input and the contour of the LV as output. We manually added to the training image set some additional images (in total ~200 images). Adding these improves the performance of the network significantly, and we believe the result can still be significantly improved if more training images are added. We found ensembling only improves the result slightly (something like 0.0096 for the best single network to 0.0093 in train set).</p>\n\n<p>For the preprocessing, we tried a combination of different things:</p>\n\n<pre><code>1.  Use the time variance of the images to determine a preliminary center of the LV and bounding box and crop from the center. \n\n2.  We rotated the images so all the cases are aligned to the same direction.\n\n3.  For the input image augmentation, we did random rotation, shift, and contrast normalization.\n</code></pre>\n\n<p>Models were trained on two GPUs, NVIDIA GTX 970 and 980Ti. The entire model takes about 4 days to train and evaluate if both GPUs are used. We used Python, Theano, Lasagne, and cuDNN for neural network implementation.</p>\n\n<p>The CNNs can detect the contour amazingly well for high-quality images. Once we have the contours, the volume is calculated basically as V_preliminary = \\sum (area_i*thickness) and maximum and minimum volume can be determined. We hypothesized that much of the error comes from the end slices, where a human can decide to include it or not. With this in mind, we did a final fitting to correct some error, with a function V_pred = V - \\beta * sqrt(V) and it does much better than a simple linear fitting.</p>\n\n<p>To deal with some extreme cases that our CNN model cannot predict good volumes for, we developed </p>\n\n<pre><code>1.  a sex-age model ( score ~ 0.036 ) (basically the same as https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18375/0-036023-score-without-looking-at-the-images)\n\n2.   a model based only on a single SAX slice. Score ~ 0.015\n\n3.  A model based on the 4-chamber view. We hand labeled many of these images to train this model. Score ~ 0.017\n\n4.  We took the average of these three models (which scores ~ 0.013) as the default model, if our SAX-based CNN model fails, we took the result from this model. \n</code></pre>\n\n<p>We also tried a Fourier-based segmentation method which gives a score about 0.016. Since it is quite complicated and does significantly worse than the CNN, after the early stages of the competition we dropped this model entirely, to simplify our work and code.</p>\n\n<p>Some observations:</p>\n\n<ol>\n<li><p>A lot of our effort was spent cleaning up data and dealing with edge cases. As kunsthart found, much of the error came from just a few cases, and in one case fixing a single prediction in the validation set dropped our score from ~ 0.0102 to 0.0098.</p></li>\n<li><p>The CNN architectures we used were relatively small. We found that adding capacity did not improve results, though we did not experiment much with different types of architectures or activation functions.</p></li>\n<li><p>Batch normalization helped enormously, as did using a modification of the Sorenson-Dice Index as the segmentation objective function, rather than binary cross-entropy.</p></li>\n<li><p>We did no test-time augmentation due to time constraints, this might have helped.</p></li>\n</ol>\n\n<p>We&#8217;ll write up a full documentation of our model in a couple of days. Thanks to Kaggle, Booz Allen Hamilton, and the administrators for running such an interesting and engaging competition, and to our fellow competitors :)</p>",
      "rawMarkdown": "Hello everyone,\r\n\r\nWe wanted to share a quick summary of our method. We had a great time on this competition and learned a lot by working with this data, which is interesting by itself and we’re very glad it is also potentially useful as a real life medical problem. We will briefly explain our approach here and post full documentation in a couple of days.\r\n\r\nOur method uses an average of 10 different fully convolutional neural networks (different architecture, input image size, or pre-processing transformation on the image), which output segmentation results for individual MRI images. We trained on the Sunnybrook data set which has one MRI image as input and the contour of the LV as output. We manually added to the training image set some additional images (in total ~200 images). Adding these improves the performance of the network significantly, and we believe the result can still be significantly improved if more training images are added. We found ensembling only improves the result slightly (something like 0.0096 for the best single network to 0.0093 in train set).\r\n\r\nFor the preprocessing, we tried a combination of different things:\r\n\r\n\t1.\tUse the time variance of the images to determine a preliminary center of the LV and bounding box and crop from the center. \r\n\r\n\t2.\tWe rotated the images so all the cases are aligned to the same direction.\r\n\r\n\t3.\tFor the input image augmentation, we did random rotation, shift, and contrast normalization.\r\n\r\nModels were trained on two GPUs, NVIDIA GTX 970 and 980Ti. The entire model takes about 4 days to train and evaluate if both GPUs are used. We used Python, Theano, Lasagne, and cuDNN for neural network implementation.\r\n\r\nThe CNNs can detect the contour amazingly well for high-quality images. Once we have the contours, the volume is calculated basically as V_preliminary = \\sum (area_i*thickness) and maximum and minimum volume can be determined. We hypothesized that much of the error comes from the end slices, where a human can decide to include it or not. With this in mind, we did a final fitting to correct some error, with a function V_pred = V - \\beta * sqrt(V) and it does much better than a simple linear fitting.\r\n\r\nTo deal with some extreme cases that our CNN model cannot predict good volumes for, we developed \r\n\r\n\t1.\ta sex-age model ( score ~ 0.036 ) (basically the same as https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18375/0-036023-score-without-looking-at-the-images)\r\n\r\n\t2.\t a model based only on a single SAX slice. Score ~ 0.015\r\n\r\n\t3.\tA model based on the 4-chamber view. We hand labeled many of these images to train this model. Score ~ 0.017\r\n\r\n\t4.\tWe took the average of these three models (which scores ~ 0.013) as the default model, if our SAX-based CNN model fails, we took the result from this model. \r\n\r\nWe also tried a Fourier-based segmentation method which gives a score about 0.016. Since it is quite complicated and does significantly worse than the CNN, after the early stages of the competition we dropped this model entirely, to simplify our work and code.\r\n\r\nSome observations:\r\n\r\n1. A lot of our effort was spent cleaning up data and dealing with edge cases. As kunsthart found, much of the error came from just a few cases, and in one case fixing a single prediction in the validation set dropped our score from ~ 0.0102 to 0.0098.\r\n\r\n2. The CNN architectures we used were relatively small. We found that adding capacity did not improve results, though we did not experiment much with different types of architectures or activation functions.\r\n\r\n3. Batch normalization helped enormously, as did using a modification of the Sorenson-Dice Index as the segmentation objective function, rather than binary cross-entropy.\r\n\r\n4. We did no test-time augmentation due to time constraints, this might have helped.\r\n\r\nWe’ll write up a full documentation of our model in a couple of days. Thanks to Kaggle, Booz Allen Hamilton, and the administrators for running such an interesting and engaging competition, and to our fellow competitors :)",
      "votes": null
    },
    {
      "id": "111533",
      "postDate": "03/15/2016 07:43:18",
      "content": "<p>@Tencia,</p>\n\n<p>First congrats of course.</p>\n\n<p>I did some espionage on your github account and was fearing and hoping you did something with your recurrent network experience.. Did you try something in that direction ?</p>",
      "rawMarkdown": "Tencia,\r\n\r\nFirst congrats of course.\r\n\r\nI did some espionage on your github account and was fearing and hoping you did something with your recurrent network experience.. Did you try something in that direction ?",
      "votes": null
    },
    {
      "id": "111570",
      "postDate": "03/15/2016 13:55:37",
      "content": "<p>The one case fixing is the case 595 (and 599) which has only 3 slices available. If we still use the method Volume = sum(area*thickness) we'll underestimate it dramatically so we developed the default model (an average of one-slice model and the 4-chamber model, if it still fails, default to the sex-age model) . This change improved our score from 0.0102 to 0.0098 in the first stage (which has 200 cases). But it seems in the final test dataset, there is no such kind of dataset that has # of slices &lt;5, so our default method is not activated and does not help us very much. </p>\n\n<p>One good thing about this competition is that we can use the sunny-brook dataset to train the network, so the real training data itself basically serves as validation and the validation serves as real test data set for us. To me, this is the only competition I felt pretty comfortable that we had a very good chance to win because we did not overfit and still had a solid #1 position.</p>",
      "rawMarkdown": "The one case fixing is the case 595 (and 599) which has only 3 slices available. If we still use the method Volume = sum(area*thickness) we'll underestimate it dramatically so we developed the default model (an average of one-slice model and the 4-chamber model, if it still fails, default to the sex-age model) . This change improved our score from 0.0102 to 0.0098 in the first stage (which has 200 cases). But it seems in the final test dataset, there is no such kind of dataset that has # of slices <5, so our default method is not activated and does not help us very much. \r\n\r\nOne good thing about this competition is that we can use the sunny-brook dataset to train the network, so the real training data itself basically serves as validation and the validation serves as real test data set for us. To me, this is the only competition I felt pretty comfortable that we had a very good chance to win because we did not overfit and still had a solid #1 position.",
      "votes": null
    },
    {
      "id": "111600",
      "postDate": "03/15/2016 17:07:22",
      "content": "<p>[quote=Julian de Wit;111533]</p>\n\n<p>Did you try something in that direction ?</p>\n\n<p>[/quote]</p>\n\n<p>No we didn't try any recurrent networks, for two reasons. There wasn't enough data to use it in volume calculations. Also, we knew the volumes had been calculated from the areas in a predefined way, so it made sense to just reproduce that way instead of giving a model the chance to come up with the wrong conclusion. If we got the segmentation perfect and also knew the formula the volumes had been calculated with, the error would be zero, so that's the target we aiming towards.</p>\n\n<p>For another option, like DRAW or similar recurrent attention model, again there wasn't enough data and also it did not seem necessary. The fully convolutional feed-forward network was doing a good job with segmentation already, and we thought most of our error was coming from other places.</p>",
      "rawMarkdown": "[quote=Julian de Wit;111533]\r\n\r\nDid you try something in that direction ?\r\n\r\n[/quote]\r\n\r\nNo we didn't try any recurrent networks, for two reasons. There wasn't enough data to use it in volume calculations. Also, we knew the volumes had been calculated from the areas in a predefined way, so it made sense to just reproduce that way instead of giving a model the chance to come up with the wrong conclusion. If we got the segmentation perfect and also knew the formula the volumes had been calculated with, the error would be zero, so that's the target we aiming towards.\r\n\r\nFor another option, like DRAW or similar recurrent attention model, again there wasn't enough data and also it did not seem necessary. The fully convolutional feed-forward network was doing a good job with segmentation already, and we thought most of our error was coming from other places.",
      "votes": null
    },
    {
      "id": "111671",
      "postDate": "03/16/2016 01:35:01",
      "content": "<p>Congratulations Tencia &amp; Woshialex !!! You did a great work.</p>",
      "rawMarkdown": "Congratulations Tencia & Woshialex !!! You did a great work.",
      "votes": null
    },
    {
      "id": "111676",
      "postDate": "03/16/2016 02:50:42",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "111677",
      "postDate": "03/16/2016 02:54:56",
      "content": "<p>And how do you build a distribution for submission?</p>",
      "rawMarkdown": "And how do you build a distribution for submission?",
      "votes": null
    },
    {
      "id": "111694",
      "postDate": "03/16/2016 06:29:58",
      "content": "<p>We used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.</p>",
      "rawMarkdown": "We used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.",
      "votes": null
    },
    {
      "id": "111696",
      "postDate": "03/16/2016 06:42:38",
      "content": "<p>Thank you, I see, so you compute a optimal stdev from the training dataset label firstly, and fit this value with the predicted volume? Have you tried other methods or other features to estimate std?</p>\n\n<p>[quote=Tencia;111694]</p>\n\n<p>We used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Thank you, I see, so you compute a optimal stdev from the training dataset label firstly, and fit this value with the predicted volume? Have you tried other methods or other features to estimate std?\r\n\r\n[quote=Tencia;111694]\r\n\r\nWe used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "111720",
      "postDate": "03/16/2016 11:55:43",
      "content": "<p>@FangzouLiao\nWe did try different methods to find the best prediction for the stdev and it turns out a simple linear fit is one of the best. We introduced some other variables such as 1.0*(#_good_slice&lt;7), or (#_contour_completeness&lt;0.6, which is the a measure of the contour completeness, for the end slices, this number might be 0.8 or smaller) etc.. it seems these variables are significant to improve the prediction for stdev, but at the end, the improve on the performance is tiny so we stick to the simple one variable (predicted_volume) linear fit.</p>",
      "rawMarkdown": "FangzouLiao\r\nWe did try different methods to find the best prediction for the stdev and it turns out a simple linear fit is one of the best. We introduced some other variables such as 1.0*(#_good_slice<7), or (#_contour_completeness<0.6, which is the a measure of the contour completeness, for the end slices, this number might be 0.8 or smaller) etc.. it seems these variables are significant to improve the prediction for stdev, but at the end, the improve on the performance is tiny so we stick to the simple one variable (predicted_volume) linear fit.",
      "votes": null
    },
    {
      "id": "112082",
      "postDate": "03/18/2016 01:50:57",
      "content": "<p>[deleted]</p>",
      "rawMarkdown": "[deleted]",
      "votes": null
    },
    {
      "id": "112094",
      "postDate": "03/18/2016 03:50:14",
      "content": "<p>The additional images were manually segmented from within the training set. These were images used to train the fully convolutional segmentation network, not to calculate volumes, so there was no age/sex info used.</p>",
      "rawMarkdown": "The additional images were manually segmented from within the training set. These were images used to train the fully convolutional segmentation network, not to calculate volumes, so there was no age/sex info used.",
      "votes": null
    },
    {
      "id": "113956",
      "postDate": "04/06/2016 13:10:01",
      "content": "<p>What was your network's performance, in Dice index, on the Sunnybrook dataset?</p>",
      "rawMarkdown": "What was your network's performance, in Dice index, on the Sunnybrook dataset?",
      "votes": null
    },
    {
      "id": "113961",
      "postDate": "04/06/2016 13:23:56",
      "content": "<p>I remember it was something like 94%-96% for train and 93%-95% for validation (Cross validation results on the sunny-brook dataset).</p>\n\n<p>[quote=kungvu;113956]</p>\n\n<p>What was your network's performance, in Dice index, on the Sunnybrook dataset?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I remember it was something like 94%-96% for train and 93%-95% for validation (Cross validation results on the sunny-brook dataset).\r\n\r\n[quote=kungvu;113956]\r\n\r\nWhat was your network's performance, in Dice index, on the Sunnybrook dataset?\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "114313",
      "postDate": "04/09/2016 07:34:41",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 111533,
      "author_name": "juliandewit",
      "author_url": "",
      "post_date": "03/15/2016 07:43:18",
      "content": "<p>@Tencia,</p>\n\n<p>First congrats of course.</p>\n\n<p>I did some espionage on your github account and was fearing and hoping you did something with your recurrent network experience.. Did you try something in that direction ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111570,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "03/15/2016 13:55:37",
      "content": "<p>The one case fixing is the case 595 (and 599) which has only 3 slices available. If we still use the method Volume = sum(area*thickness) we'll underestimate it dramatically so we developed the default model (an average of one-slice model and the 4-chamber model, if it still fails, default to the sex-age model) . This change improved our score from 0.0102 to 0.0098 in the first stage (which has 200 cases). But it seems in the final test dataset, there is no such kind of dataset that has # of slices &lt;5, so our default method is not activated and does not help us very much. </p>\n\n<p>One good thing about this competition is that we can use the sunny-brook dataset to train the network, so the real training data itself basically serves as validation and the validation serves as real test data set for us. To me, this is the only competition I felt pretty comfortable that we had a very good chance to win because we did not overfit and still had a solid #1 position.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111600,
      "author_name": "tencia",
      "author_url": "",
      "post_date": "03/15/2016 17:07:22",
      "content": "<p>[quote=Julian de Wit;111533]</p>\n\n<p>Did you try something in that direction ?</p>\n\n<p>[/quote]</p>\n\n<p>No we didn't try any recurrent networks, for two reasons. There wasn't enough data to use it in volume calculations. Also, we knew the volumes had been calculated from the areas in a predefined way, so it made sense to just reproduce that way instead of giving a model the chance to come up with the wrong conclusion. If we got the segmentation perfect and also knew the formula the volumes had been calculated with, the error would be zero, so that's the target we aiming towards.</p>\n\n<p>For another option, like DRAW or similar recurrent attention model, again there wasn't enough data and also it did not seem necessary. The fully convolutional feed-forward network was doing a good job with segmentation already, and we thought most of our error was coming from other places.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111671,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/16/2016 01:35:01",
      "content": "<p>Congratulations Tencia &amp; Woshialex !!! You did a great work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111676,
      "author_name": "liaofz",
      "author_url": "",
      "post_date": "03/16/2016 02:50:42",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 111677,
      "author_name": "liaofz",
      "author_url": "",
      "post_date": "03/16/2016 02:54:56",
      "content": "<p>And how do you build a distribution for submission?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111694,
      "author_name": "tencia",
      "author_url": "",
      "post_date": "03/16/2016 06:29:58",
      "content": "<p>We used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111696,
      "author_name": "liaofz",
      "author_url": "",
      "post_date": "03/16/2016 06:42:38",
      "content": "<p>Thank you, I see, so you compute a optimal stdev from the training dataset label firstly, and fit this value with the predicted volume? Have you tried other methods or other features to estimate std?</p>\n\n<p>[quote=Tencia;111694]</p>\n\n<p>We used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111720,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "03/16/2016 11:55:43",
      "content": "<p>@FangzouLiao\nWe did try different methods to find the best prediction for the stdev and it turns out a simple linear fit is one of the best. We introduced some other variables such as 1.0*(#_good_slice&lt;7), or (#_contour_completeness&lt;0.6, which is the a measure of the contour completeness, for the end slices, this number might be 0.8 or smaller) etc.. it seems these variables are significant to improve the prediction for stdev, but at the end, the improve on the performance is tiny so we stick to the simple one variable (predicted_volume) linear fit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112082,
      "author_name": "howardreuben",
      "author_url": "",
      "post_date": "03/18/2016 01:50:57",
      "content": "<p>[deleted]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112094,
      "author_name": "tencia",
      "author_url": "",
      "post_date": "03/18/2016 03:50:14",
      "content": "<p>The additional images were manually segmented from within the training set. These were images used to train the fully convolutional segmentation network, not to calculate volumes, so there was no age/sex info used.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 113956,
      "author_name": "vuptran",
      "author_url": "",
      "post_date": "04/06/2016 13:10:01",
      "content": "<p>What was your network's performance, in Dice index, on the Sunnybrook dataset?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 113961,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "04/06/2016 13:23:56",
      "content": "<p>I remember it was something like 94%-96% for train and 93%-95% for validation (Cross validation results on the sunny-brook dataset).</p>\n\n<p>[quote=kungvu;113956]</p>\n\n<p>What was your network's performance, in Dice index, on the Sunnybrook dataset?</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114313,
      "author_name": "shaguftaabid",
      "author_url": "",
      "post_date": "04/09/2016 07:34:41",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "111523": "Hello everyone,\r\n\r\nWe wanted to share a quick summary of our method. We had a great time on this competition and learned a lot by working with this data, which is interesting by itself and we’re very glad it is also potentially useful as a real life medical problem. We will briefly explain our approach here and post full documentation in a couple of days.\r\n\r\nOur method uses an average of 10 different fully convolutional neural networks (different architecture, input image size, or pre-processing transformation on the image), which output segmentation results for individual MRI images. We trained on the Sunnybrook data set which has one MRI image as input and the contour of the LV as output. We manually added to the training image set some additional images (in total ~200 images). Adding these improves the performance of the network significantly, and we believe the result can still be significantly improved if more training images are added. We found ensembling only improves the result slightly (something like 0.0096 for the best single network to 0.0093 in train set).\r\n\r\nFor the preprocessing, we tried a combination of different things:\r\n\r\n\t1.\tUse the time variance of the images to determine a preliminary center of the LV and bounding box and crop from the center. \r\n\r\n\t2.\tWe rotated the images so all the cases are aligned to the same direction.\r\n\r\n\t3.\tFor the input image augmentation, we did random rotation, shift, and contrast normalization.\r\n\r\nModels were trained on two GPUs, NVIDIA GTX 970 and 980Ti. The entire model takes about 4 days to train and evaluate if both GPUs are used. We used Python, Theano, Lasagne, and cuDNN for neural network implementation.\r\n\r\nThe CNNs can detect the contour amazingly well for high-quality images. Once we have the contours, the volume is calculated basically as V_preliminary = \\sum (area_i*thickness) and maximum and minimum volume can be determined. We hypothesized that much of the error comes from the end slices, where a human can decide to include it or not. With this in mind, we did a final fitting to correct some error, with a function V_pred = V - \\beta * sqrt(V) and it does much better than a simple linear fitting.\r\n\r\nTo deal with some extreme cases that our CNN model cannot predict good volumes for, we developed \r\n\r\n\t1.\ta sex-age model ( score ~ 0.036 ) (basically the same as https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18375/0-036023-score-without-looking-at-the-images)\r\n\r\n\t2.\t a model based only on a single SAX slice. Score ~ 0.015\r\n\r\n\t3.\tA model based on the 4-chamber view. We hand labeled many of these images to train this model. Score ~ 0.017\r\n\r\n\t4.\tWe took the average of these three models (which scores ~ 0.013) as the default model, if our SAX-based CNN model fails, we took the result from this model. \r\n\r\nWe also tried a Fourier-based segmentation method which gives a score about 0.016. Since it is quite complicated and does significantly worse than the CNN, after the early stages of the competition we dropped this model entirely, to simplify our work and code.\r\n\r\nSome observations:\r\n\r\n1. A lot of our effort was spent cleaning up data and dealing with edge cases. As kunsthart found, much of the error came from just a few cases, and in one case fixing a single prediction in the validation set dropped our score from ~ 0.0102 to 0.0098.\r\n\r\n2. The CNN architectures we used were relatively small. We found that adding capacity did not improve results, though we did not experiment much with different types of architectures or activation functions.\r\n\r\n3. Batch normalization helped enormously, as did using a modification of the Sorenson-Dice Index as the segmentation objective function, rather than binary cross-entropy.\r\n\r\n4. We did no test-time augmentation due to time constraints, this might have helped.\r\n\r\nWe’ll write up a full documentation of our model in a couple of days. Thanks to Kaggle, Booz Allen Hamilton, and the administrators for running such an interesting and engaging competition, and to our fellow competitors :)",
    "111533": "Tencia,\r\n\r\nFirst congrats of course.\r\n\r\nI did some espionage on your github account and was fearing and hoping you did something with your recurrent network experience.. Did you try something in that direction ?",
    "111570": "The one case fixing is the case 595 (and 599) which has only 3 slices available. If we still use the method Volume = sum(area*thickness) we'll underestimate it dramatically so we developed the default model (an average of one-slice model and the 4-chamber model, if it still fails, default to the sex-age model) . This change improved our score from 0.0102 to 0.0098 in the first stage (which has 200 cases). But it seems in the final test dataset, there is no such kind of dataset that has # of slices <5, so our default method is not activated and does not help us very much. \r\n\r\nOne good thing about this competition is that we can use the sunny-brook dataset to train the network, so the real training data itself basically serves as validation and the validation serves as real test data set for us. To me, this is the only competition I felt pretty comfortable that we had a very good chance to win because we did not overfit and still had a solid #1 position.",
    "111600": "[quote=Julian de Wit;111533]\r\n\r\nDid you try something in that direction ?\r\n\r\n[/quote]\r\n\r\nNo we didn't try any recurrent networks, for two reasons. There wasn't enough data to use it in volume calculations. Also, we knew the volumes had been calculated from the areas in a predefined way, so it made sense to just reproduce that way instead of giving a model the chance to come up with the wrong conclusion. If we got the segmentation perfect and also knew the formula the volumes had been calculated with, the error would be zero, so that's the target we aiming towards.\r\n\r\nFor another option, like DRAW or similar recurrent attention model, again there wasn't enough data and also it did not seem necessary. The fully convolutional feed-forward network was doing a good job with segmentation already, and we thought most of our error was coming from other places.",
    "111671": "Congratulations Tencia & Woshialex !!! You did a great work.",
    "111676": "",
    "111677": "And how do you build a distribution for submission?",
    "111694": "We used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.",
    "111696": "Thank you, I see, so you compute a optimal stdev from the training dataset label firstly, and fit this value with the predicted volume? Have you tried other methods or other features to estimate std?\r\n\r\n[quote=Tencia;111694]\r\n\r\nWe used a Gaussian CDF to come up with the probability values, and used linear fit of stdev = a*predicted_volume + b, fit separately for systole and diastole for each model.\r\n\r\n[/quote]",
    "111720": "FangzouLiao\r\nWe did try different methods to find the best prediction for the stdev and it turns out a simple linear fit is one of the best. We introduced some other variables such as 1.0*(#_good_slice<7), or (#_contour_completeness<0.6, which is the a measure of the contour completeness, for the end slices, this number might be 0.8 or smaller) etc.. it seems these variables are significant to improve the prediction for stdev, but at the end, the improve on the performance is tiny so we stick to the simple one variable (predicted_volume) linear fit.",
    "112082": "[deleted]",
    "112094": "The additional images were manually segmented from within the training set. These were images used to train the fully convolutional segmentation network, not to calculate volumes, so there was no age/sex info used.",
    "113956": "What was your network's performance, in Dice index, on the Sunnybrook dataset?",
    "113961": "I remember it was something like 94%-96% for train and 93%-95% for validation (Cross validation results on the sunny-brook dataset).\r\n\r\n[quote=kungvu;113956]\r\n\r\nWhat was your network's performance, in Dice index, on the Sunnybrook dataset?\r\n\r\n[/quote]",
    "114313": ""
  },
  "source": "meta"
}