{
  "id": 19503,
  "title": "Best non-NN results?",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19503",
  "author_name": "",
  "post_date": "2016-03-14T14:50:29.207Z",
  "votes": 5,
  "comment_count": 7,
  "views": 1515,
  "content": "<p>Once final results are revealed, I'd love to know how well approaches NOT based on neural networks did. Did any make it into the top 10?</p>\n\n<p>My own attempt, such as it is (let's just say I'd be very happy if it managed to stay under 0.03) was all &quot;classic&quot; computer vision - it's an orgy in FFTs, clustering, interpolation, contour finding, region labeling and manipulation, with a smattering of plain vanilla classification and regression on top. Most of the development time went into dreaming up and trying features for LV identification and segmentation. The hardest part was reliably separating LV from RV and arteries. I'm pretty sure NN approaches had a huge advantage there.</p>\n\n<p>Biggest lesson learned, unless somebody managed to do significantly better without NNs: I have to get myself one of them newfangled GPUs. :D</p>",
  "messages": [
    {
      "id": "111401",
      "postDate": "03/14/2016 14:50:29",
      "content": "<p>Once final results are revealed, I'd love to know how well approaches NOT based on neural networks did. Did any make it into the top 10?</p>\n\n<p>My own attempt, such as it is (let's just say I'd be very happy if it managed to stay under 0.03) was all &quot;classic&quot; computer vision - it's an orgy in FFTs, clustering, interpolation, contour finding, region labeling and manipulation, with a smattering of plain vanilla classification and regression on top. Most of the development time went into dreaming up and trying features for LV identification and segmentation. The hardest part was reliably separating LV from RV and arteries. I'm pretty sure NN approaches had a huge advantage there.</p>\n\n<p>Biggest lesson learned, unless somebody managed to do significantly better without NNs: I have to get myself one of them newfangled GPUs. :D</p>",
      "rawMarkdown": "Once final results are revealed, I'd love to know how well approaches NOT based on neural networks did. Did any make it into the top 10?\r\n\r\nMy own attempt, such as it is (let's just say I'd be very happy if it managed to stay under 0.03) was all \"classic\" computer vision - it's an orgy in FFTs, clustering, interpolation, contour finding, region labeling and manipulation, with a smattering of plain vanilla classification and regression on top. Most of the development time went into dreaming up and trying features for LV identification and segmentation. The hardest part was reliably separating LV from RV and arteries. I'm pretty sure NN approaches had a huge advantage there.\r\n\r\nBiggest lesson learned, unless somebody managed to do significantly better without NNs: I have to get myself one of them newfangled GPUs. :D",
      "votes": null
    },
    {
      "id": "111408",
      "postDate": "03/14/2016 15:48:56",
      "content": "<p>From our experience you can go atleast below 0.015 without using any NN's.</p>",
      "rawMarkdown": "From our experience you can go atleast below 0.015 without using any NN's.",
      "votes": null
    },
    {
      "id": "111416",
      "postDate": "03/14/2016 16:16:50",
      "content": "<p>0.015? Are you talking about the score at the first stage of the competition? That's a very good result!</p>\n\n<p>I was able to reach 0.020 during the first stage of the competition using simple methods (thresholding, ...). After that I used NN on this first result to improve my score to 0.014. Since I don't have a GPU, it has been a way to simplify a bit the model. Overall, I am able to localize well the left ventricule without NN for the slices around the middle and end of the ventricule. Where I had difficulties is to localize it correctly for the slices at the beginning of the ventricule (close to the left atrium). That's when I decided to try to use some NN.   </p>",
      "rawMarkdown": "0.015? Are you talking about the score at the first stage of the competition? That's a very good result!\r\n\r\nI was able to reach 0.020 during the first stage of the competition using simple methods (thresholding, ...). After that I used NN on this first result to improve my score to 0.014. Since I don't have a GPU, it has been a way to simplify a bit the model. Overall, I am able to localize well the left ventricule without NN for the slices around the middle and end of the ventricule. Where I had difficulties is to localize it correctly for the slices at the beginning of the ventricule (close to the left atrium). That's when I decided to try to use some NN.",
      "votes": null
    },
    {
      "id": "111562",
      "postDate": "03/15/2016 13:30:20",
      "content": "<p>I used Active Appearance Models (and no neural nets) to get to .013433 on the final test set.</p>\n\n<p>My biggest source of error was that I couldn't find a good way to recognize when the base of the heart \nhad been reached, so I frequently ended up segmenting one too many or one too few slices.</p>\n\n<p>Other than that, the segmentation of the axial slices looked quite accurate in most cases.</p>",
      "rawMarkdown": "I used Active Appearance Models (and no neural nets) to get to .013433 on the final test set.\r\n\r\nMy biggest source of error was that I couldn't find a good way to recognize when the base of the heart \r\nhad been reached, so I frequently ended up segmenting one too many or one too few slices.\r\n\r\nOther than that, the segmentation of the axial slices looked quite accurate in most cases.",
      "votes": null
    },
    {
      "id": "111572",
      "postDate": "03/15/2016 14:12:12",
      "content": "<p>We did use neural network to detect the ROI and to find a coarse contour, but our main technique is to convert the ROI to polar coordinate space and use dynamic programming to find the optimal contour.  We found that improving the NN model has less effect in improving the end result, as the final contour is found using color information. </p>\n\n<p>The fourth image is also the output of NN, which is used to determine the left and right bound of the final contour (the two thin lines in the 3rd image).  Dynamic programming is then applied within the space between the two thin lines.</p>\n\n<p>I think NN in our solution can be replace with other traditional methods without changing the final performance.</p>\n\n<p>Having looked at the top solutions, I feel that trying to find a materialized contour might actually hurt accuracy.  We had the same problem as PaulG, and the best method we have found is to always remove one last slice.</p>",
      "rawMarkdown": "We did use neural network to detect the ROI and to find a coarse contour, but our main technique is to convert the ROI to polar coordinate space and use dynamic programming to find the optimal contour.  We found that improving the NN model has less effect in improving the end result, as the final contour is found using color information. \r\n\r\nThe fourth image is also the output of NN, which is used to determine the left and right bound of the final contour (the two thin lines in the 3rd image).  Dynamic programming is then applied within the space between the two thin lines.\r\n\r\nI think NN in our solution can be replace with other traditional methods without changing the final performance.\r\n\r\nHaving looked at the top solutions, I feel that trying to find a materialized contour might actually hurt accuracy.  We had the same problem as PaulG, and the best method we have found is to always remove one last slice.",
      "votes": null
    },
    {
      "id": "111585",
      "postDate": "03/15/2016 15:39:54",
      "content": "<p>My approach was to use OpenCV train cascade with direct volume calculation. I used modified detect method since I knew there are only one ventricle on image. For images with good quality it works very well. Also it was easy to filter incorrect rectangles since for one patient it must share at least one point (with some exceptions). Using these rectangles I obtain direct volume value as well many different area-based features for each sax like area of enclosing and inclosed circles, different thresholds, partial volumes etc. To improve results I later add all found data in one big table with addition of some DICOM data like age, sex, maximum axis difference and some features from 2ch and 4ch direct area analysis. Then I ran XGBoost 20KFold Monte-Carlo for ~1000 iterations for volume prediction, using median. It allows to achieve score around 0.155 on first set. And I think it's possible to improve this result.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/3913/output.gif\" alt=\"1\" title> </p>",
      "rawMarkdown": "My approach was to use OpenCV train cascade with direct volume calculation. I used modified detect method since I knew there are only one ventricle on image. For images with good quality it works very well. Also it was easy to filter incorrect rectangles since for one patient it must share at least one point (with some exceptions). Using these rectangles I obtain direct volume value as well many different area-based features for each sax like area of enclosing and inclosed circles, different thresholds, partial volumes etc. To improve results I later add all found data in one big table with addition of some DICOM data like age, sex, maximum axis difference and some features from 2ch and 4ch direct area analysis. Then I ran XGBoost 20KFold Monte-Carlo for ~1000 iterations for volume prediction, using median. It allows to achieve score around 0.155 on first set. And I think it's possible to improve this result.\r\n\r\n![1][1] \r\n[1]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/3913/output.gif",
      "votes": null
    },
    {
      "id": "111896",
      "postDate": "03/17/2016 06:11:47",
      "content": "<p>Hi, thank you for sharing your results! You guys did great job.</p>\n\n<p>Let me share key points of our image processing approach which gave us CRPS about 0.023 only. :-(</p>\n\n<p>You could look through our code from <a href=\"https://github.com/brotherofken/national_data_science_bowl_2\">Github repository (not completed&amp;polished yet)</a>.</p>\n\n<p>Our approach to LV detection uses intersection of 2ch, 4ch and sax slices and project that point to sax as initial estimation of LV position. Then we seek for nearest light blob using SIFT keypoint detector.\nThen we trained HOG detector and used it only when intersection based detector failed. Unfortunately this approach led to large number of detection errors.</p>\n\n<p>We detected contours by &quot;color&quot; segmentation and Cascaded Pose Regression with gradient boosted trees (ferns). Also we applied cascaded pose regression for 2ch view in order to estimate LV height. This gave us two estimations of end-systolic/end-diastolic volume per person.</p>\n\n<ul>\n<li><a href=\"https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/sax_sample.png\">Example of landmarks found on sax view (red points).</a></li>\n<li><a href=\"https://github.com/brotherofken/national_data_science_bowl_2/blob/master/slides/img/ch2_sample.png\">Example of landmarks found on 2ch view (red points).</a></li>\n</ul>\n\n<p>After that we combined contour based estimations of volume with patient meta data in order to train XGBoost regression on the top.</p>\n\n<p>Also we tried 3D reconstruction, but without great success.\n<img src=\"https://github.com/brotherofken/national_data_science_bowl_2/raw/master/slides/img/lv1.png\" alt=\"3Dreconstruction\" title></p>\n\n<p>Here is plot of our predictions vs. groundtruth on validation subset.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/predictions_vs_groundtruth.png\" alt=\"prediction vs groundtruth\" title></p>\n\n<ul>\n<li>Blue points - end-systolic volumes in mm^3</li>\n<li>Green points - end-systolic volumes in mm^3</li>\n</ul>",
      "rawMarkdown": "Hi, thank you for sharing your results! You guys did great job.\r\n\r\nLet me share key points of our image processing approach which gave us CRPS about 0.023 only. :-(\r\n\r\nYou could look through our code from [Github repository (not completed&polished yet)][5].\r\n\r\nOur approach to LV detection uses intersection of 2ch, 4ch and sax slices and project that point to sax as initial estimation of LV position. Then we seek for nearest light blob using SIFT keypoint detector.\r\nThen we trained HOG detector and used it only when intersection based detector failed. Unfortunately this approach led to large number of detection errors.\r\n\r\nWe detected contours by \"color\" segmentation and Cascaded Pose Regression with gradient boosted trees (ferns). Also we applied cascaded pose regression for 2ch view in order to estimate LV height. This gave us two estimations of end-systolic/end-diastolic volume per person.\r\n\r\n - [Example of landmarks found on sax view (red points).][1]\r\n - [Example of landmarks found on 2ch view (red points).][2]\r\n\r\nAfter that we combined contour based estimations of volume with patient meta data in order to train XGBoost regression on the top.\r\n\r\nAlso we tried 3D reconstruction, but without great success.\r\n![3Dreconstruction][3]\r\n\r\nHere is plot of our predictions vs. groundtruth on validation subset.\r\n\r\n![prediction vs groundtruth][4]\r\n\r\n - Blue points - end-systolic volumes in mm^3\r\n - Green points - end-systolic volumes in mm^3\r\n\r\n  [1]: https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/sax_sample.png\r\n  [2]: https://github.com/brotherofken/national_data_science_bowl_2/blob/master/slides/img/ch2_sample.png\r\n  [3]: https://github.com/brotherofken/national_data_science_bowl_2/raw/master/slides/img/lv1.png\r\n  [4]: https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/predictions_vs_groundtruth.png\r\n  [5]: https://github.com/brotherofken/national_data_science_bowl_2",
      "votes": null
    },
    {
      "id": "111970",
      "postDate": "03/17/2016 16:58:13",
      "content": "<p>Here are some details about my approach. I made a stupid error in the final submission code (hard coded 200 instead of the test set length in one for loop!!!) - so the corrected code has 0.015570 on the private LB which would give #21. It is not great so I'll just summarize the main features.</p>\n\n<p><strong>Preprocessing</strong></p>\n\n<p>Here I used the image orientation metadata to eliminate sax slices which are not parallel. I noticed that inclusion of such slices led to huge errors as the SliceLocation was often quite off.</p>\n\n<p><strong>Baselines</strong> </p>\n\n<p>I used computer vision techniques (thresholding etc. - more on this later) so it was important to eliminate huge errors or cases when my algorithm failed completely. For that I used some baseline models:</p>\n\n<p>1) simple model based on age-sex metadata as a complete emergency fallback</p>\n\n<p>2) I used my earlier attempt as a fallback for the more refined model in two ways: </p>\n\n<pre><code> a) in case the more refined model failed completely I put in the fallback prediction, \n\n b) in the other case I still averaged the cumulative probability distribution of the more refined model with the one of this model\n</code></pre>\n\n<p><strong>Predictions</strong></p>\n\n<p>My approach was to first identify the LV on the middle sax slice in frame 0. </p>\n\n<p>I used adaptive thresholding using skimage and identified the blob using the intersection of the 2CH and 4CH lines on the sax slice. In cases where this did not work or the 4CH views were absent I used my previous approach where I used some custom measure of circularity, the absolute value of the first Fourier coefficient etc. to identify the LV.</p>\n\n<p>In order to fill in the muscles within the LV, I always used the convex hull of the blob identified above as my prediction for the LV. I then picked a threshold value by estimating the median of the immediate surroundings of the blob. I used this value in the following.</p>\n\n<p>Once the LV was identified on frame 0 of the middle slice, I used thresholding with the above value on other frames and picked the blobs closest to the ones identified on the previous frames. I kept track of the distance of centroids of blobs between neighbouring frames in order to know when something went badly wrong.</p>\n\n<p>Once LV was identified on all frames of the middle slice I adopted the same procedure (picking nearest blob) changing the sax slice and keeping the frame number fixed. I had some sanity fallbacks so that the area would not grow too much on the edges - sometimes I used then intersection with the previous slice - in another attempt I tried watershed segmentation of the blob.\nHere I also used, when available, the intersection of 2CH and 4CH to align the neighboring slices - this helped in cases when some sax slices suddenly flipped from portrait to landscape or even once changed resolution. </p>\n\n<p>I also tried to extract the LV by looking at the approximate ellipse formed by the 2CH and 4CH lines on the sax slices. When the volume in this method as a function of time had a similar shape as the volume as a function of time in the other method I used it. I also tried deducing from this data when to neglect the outermost sax slices. This did not work as well as I hoped.</p>\n\n<p>I repeated the same procedure starting from the two neighbouring slices of the middle slice. I could then take the mean of the volumes obtained in these three tries ignoring cases when that centroid distance was too big.</p>\n\n<p>I also noticed that often the volume of one frame was much higher (or lower) than of its neighbors - to eliminate that I used some smoothing of the obtained volumes as a function of time. </p>\n\n<p><strong>Refinements</strong> </p>\n\n<p>In order to improve the predictions I used some tricks like fitting gradient boosted regression on systole, diastole and age and sex for the baseline model and using a NN (in keras optimizing MAE) to extract predictions from the minimal, maximal volume but also from the volumes in a couple of neighboring frames using the method described above, the 2CH/4CH ellipse and the method described above but with possible watershed segmentation for the outermost slices. From the NN model I got my final predictions for the systole and diastole. </p>\n\n<p>Then I evaluated the standard deviation of the errors and used it for predicting the cumulative probability distribution.\nOf course in the above the NN is used only for callibrating the final predictions and not for any image analysis.</p>\n\n<p><strong>Conclusions</strong></p>\n\n<p>My main problem with the CV algorithmic approach was that I am no cardiologist and there were cases where I thought that my algorithm does a great job while the predicted systole/diastole was significantly away from the ground truth. In other cases I had nice agreement. So I had no idea what to improve - especially at the outermost slices. People using CNN approach for segmentation/prediction did not have that problem I guess as the rules/algorithm were in a way implicitly deduced by the CNN.</p>\n\n<p>I did not have much time in the final weeks of the competition so I settled on exploiting the information in the 2CH and 4CH views. If I would have more time I would have definitely tried some CNN based segmentation approach.</p>",
      "rawMarkdown": "Here are some details about my approach. I made a stupid error in the final submission code (hard coded 200 instead of the test set length in one for loop!!!) - so the corrected code has 0.015570 on the private LB which would give #21. It is not great so I'll just summarize the main features.\r\n\r\n**Preprocessing**\r\n\r\nHere I used the image orientation metadata to eliminate sax slices which are not parallel. I noticed that inclusion of such slices led to huge errors as the SliceLocation was often quite off.\r\n\r\n**Baselines** \r\n\r\nI used computer vision techniques (thresholding etc. - more on this later) so it was important to eliminate huge errors or cases when my algorithm failed completely. For that I used some baseline models:\r\n\r\n1) simple model based on age-sex metadata as a complete emergency fallback\r\n\r\n2) I used my earlier attempt as a fallback for the more refined model in two ways: \r\n\r\n     a) in case the more refined model failed completely I put in the fallback prediction, \r\n\r\n     b) in the other case I still averaged the cumulative probability distribution of the more refined model with the one of this model\r\n\r\n**Predictions**\r\n\r\nMy approach was to first identify the LV on the middle sax slice in frame 0. \r\n\r\nI used adaptive thresholding using skimage and identified the blob using the intersection of the 2CH and 4CH lines on the sax slice. In cases where this did not work or the 4CH views were absent I used my previous approach where I used some custom measure of circularity, the absolute value of the first Fourier coefficient etc. to identify the LV.\r\n\r\nIn order to fill in the muscles within the LV, I always used the convex hull of the blob identified above as my prediction for the LV. I then picked a threshold value by estimating the median of the immediate surroundings of the blob. I used this value in the following.\r\n\r\nOnce the LV was identified on frame 0 of the middle slice, I used thresholding with the above value on other frames and picked the blobs closest to the ones identified on the previous frames. I kept track of the distance of centroids of blobs between neighbouring frames in order to know when something went badly wrong.\r\n\r\nOnce LV was identified on all frames of the middle slice I adopted the same procedure (picking nearest blob) changing the sax slice and keeping the frame number fixed. I had some sanity fallbacks so that the area would not grow too much on the edges - sometimes I used then intersection with the previous slice - in another attempt I tried watershed segmentation of the blob.\r\nHere I also used, when available, the intersection of 2CH and 4CH to align the neighboring slices - this helped in cases when some sax slices suddenly flipped from portrait to landscape or even once changed resolution. \r\n\r\nI also tried to extract the LV by looking at the approximate ellipse formed by the 2CH and 4CH lines on the sax slices. When the volume in this method as a function of time had a similar shape as the volume as a function of time in the other method I used it. I also tried deducing from this data when to neglect the outermost sax slices. This did not work as well as I hoped.\r\n\r\nI repeated the same procedure starting from the two neighbouring slices of the middle slice. I could then take the mean of the volumes obtained in these three tries ignoring cases when that centroid distance was too big.\r\n\r\nI also noticed that often the volume of one frame was much higher (or lower) than of its neighbors - to eliminate that I used some smoothing of the obtained volumes as a function of time. \r\n\r\n**Refinements** \r\n\r\nIn order to improve the predictions I used some tricks like fitting gradient boosted regression on systole, diastole and age and sex for the baseline model and using a NN (in keras optimizing MAE) to extract predictions from the minimal, maximal volume but also from the volumes in a couple of neighboring frames using the method described above, the 2CH/4CH ellipse and the method described above but with possible watershed segmentation for the outermost slices. From the NN model I got my final predictions for the systole and diastole. \r\n\r\nThen I evaluated the standard deviation of the errors and used it for predicting the cumulative probability distribution.\r\nOf course in the above the NN is used only for callibrating the final predictions and not for any image analysis.\r\n\r\n**Conclusions**\r\n\r\nMy main problem with the CV algorithmic approach was that I am no cardiologist and there were cases where I thought that my algorithm does a great job while the predicted systole/diastole was significantly away from the ground truth. In other cases I had nice agreement. So I had no idea what to improve - especially at the outermost slices. People using CNN approach for segmentation/prediction did not have that problem I guess as the rules/algorithm were in a way implicitly deduced by the CNN.\r\n\r\nI did not have much time in the final weeks of the competition so I settled on exploiting the information in the 2CH and 4CH views. If I would have more time I would have definitely tried some CNN based segmentation approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 111408,
      "author_name": "niklaskoehler",
      "author_url": "",
      "post_date": "03/14/2016 15:48:56",
      "content": "<p>From our experience you can go atleast below 0.015 without using any NN's.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111416,
      "author_name": "vincentl",
      "author_url": "",
      "post_date": "03/14/2016 16:16:50",
      "content": "<p>0.015? Are you talking about the score at the first stage of the competition? That's a very good result!</p>\n\n<p>I was able to reach 0.020 during the first stage of the competition using simple methods (thresholding, ...). After that I used NN on this first result to improve my score to 0.014. Since I don't have a GPU, it has been a way to simplify a bit the model. Overall, I am able to localize well the left ventricule without NN for the slices around the middle and end of the ventricule. Where I had difficulties is to localize it correctly for the slices at the beginning of the ventricule (close to the left atrium). That's when I decided to try to use some NN.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111562,
      "author_name": "pgeiger",
      "author_url": "",
      "post_date": "03/15/2016 13:30:20",
      "content": "<p>I used Active Appearance Models (and no neural nets) to get to .013433 on the final test set.</p>\n\n<p>My biggest source of error was that I couldn't find a good way to recognize when the base of the heart \nhad been reached, so I frequently ended up segmenting one too many or one too few slices.</p>\n\n<p>Other than that, the segmentation of the axial slices looked quite accurate in most cases.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111572,
      "author_name": "aaalgo",
      "author_url": "",
      "post_date": "03/15/2016 14:12:12",
      "content": "<p>We did use neural network to detect the ROI and to find a coarse contour, but our main technique is to convert the ROI to polar coordinate space and use dynamic programming to find the optimal contour.  We found that improving the NN model has less effect in improving the end result, as the final contour is found using color information. </p>\n\n<p>The fourth image is also the output of NN, which is used to determine the left and right bound of the final contour (the two thin lines in the 3rd image).  Dynamic programming is then applied within the space between the two thin lines.</p>\n\n<p>I think NN in our solution can be replace with other traditional methods without changing the final performance.</p>\n\n<p>Having looked at the top solutions, I feel that trying to find a materialized contour might actually hurt accuracy.  We had the same problem as PaulG, and the best method we have found is to always remove one last slice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111585,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "03/15/2016 15:39:54",
      "content": "<p>My approach was to use OpenCV train cascade with direct volume calculation. I used modified detect method since I knew there are only one ventricle on image. For images with good quality it works very well. Also it was easy to filter incorrect rectangles since for one patient it must share at least one point (with some exceptions). Using these rectangles I obtain direct volume value as well many different area-based features for each sax like area of enclosing and inclosed circles, different thresholds, partial volumes etc. To improve results I later add all found data in one big table with addition of some DICOM data like age, sex, maximum axis difference and some features from 2ch and 4ch direct area analysis. Then I ran XGBoost 20KFold Monte-Carlo for ~1000 iterations for volume prediction, using median. It allows to achieve score around 0.155 on first set. And I think it's possible to improve this result.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/3913/output.gif\" alt=\"1\" title> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111896,
      "author_name": "rasimakhunzyanov",
      "author_url": "",
      "post_date": "03/17/2016 06:11:47",
      "content": "<p>Hi, thank you for sharing your results! You guys did great job.</p>\n\n<p>Let me share key points of our image processing approach which gave us CRPS about 0.023 only. :-(</p>\n\n<p>You could look through our code from <a href=\"https://github.com/brotherofken/national_data_science_bowl_2\">Github repository (not completed&amp;polished yet)</a>.</p>\n\n<p>Our approach to LV detection uses intersection of 2ch, 4ch and sax slices and project that point to sax as initial estimation of LV position. Then we seek for nearest light blob using SIFT keypoint detector.\nThen we trained HOG detector and used it only when intersection based detector failed. Unfortunately this approach led to large number of detection errors.</p>\n\n<p>We detected contours by &quot;color&quot; segmentation and Cascaded Pose Regression with gradient boosted trees (ferns). Also we applied cascaded pose regression for 2ch view in order to estimate LV height. This gave us two estimations of end-systolic/end-diastolic volume per person.</p>\n\n<ul>\n<li><a href=\"https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/sax_sample.png\">Example of landmarks found on sax view (red points).</a></li>\n<li><a href=\"https://github.com/brotherofken/national_data_science_bowl_2/blob/master/slides/img/ch2_sample.png\">Example of landmarks found on 2ch view (red points).</a></li>\n</ul>\n\n<p>After that we combined contour based estimations of volume with patient meta data in order to train XGBoost regression on the top.</p>\n\n<p>Also we tried 3D reconstruction, but without great success.\n<img src=\"https://github.com/brotherofken/national_data_science_bowl_2/raw/master/slides/img/lv1.png\" alt=\"3Dreconstruction\" title></p>\n\n<p>Here is plot of our predictions vs. groundtruth on validation subset.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/predictions_vs_groundtruth.png\" alt=\"prediction vs groundtruth\" title></p>\n\n<ul>\n<li>Blue points - end-systolic volumes in mm^3</li>\n<li>Green points - end-systolic volumes in mm^3</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111970,
      "author_name": "romualdj",
      "author_url": "",
      "post_date": "03/17/2016 16:58:13",
      "content": "<p>Here are some details about my approach. I made a stupid error in the final submission code (hard coded 200 instead of the test set length in one for loop!!!) - so the corrected code has 0.015570 on the private LB which would give #21. It is not great so I'll just summarize the main features.</p>\n\n<p><strong>Preprocessing</strong></p>\n\n<p>Here I used the image orientation metadata to eliminate sax slices which are not parallel. I noticed that inclusion of such slices led to huge errors as the SliceLocation was often quite off.</p>\n\n<p><strong>Baselines</strong> </p>\n\n<p>I used computer vision techniques (thresholding etc. - more on this later) so it was important to eliminate huge errors or cases when my algorithm failed completely. For that I used some baseline models:</p>\n\n<p>1) simple model based on age-sex metadata as a complete emergency fallback</p>\n\n<p>2) I used my earlier attempt as a fallback for the more refined model in two ways: </p>\n\n<pre><code> a) in case the more refined model failed completely I put in the fallback prediction, \n\n b) in the other case I still averaged the cumulative probability distribution of the more refined model with the one of this model\n</code></pre>\n\n<p><strong>Predictions</strong></p>\n\n<p>My approach was to first identify the LV on the middle sax slice in frame 0. </p>\n\n<p>I used adaptive thresholding using skimage and identified the blob using the intersection of the 2CH and 4CH lines on the sax slice. In cases where this did not work or the 4CH views were absent I used my previous approach where I used some custom measure of circularity, the absolute value of the first Fourier coefficient etc. to identify the LV.</p>\n\n<p>In order to fill in the muscles within the LV, I always used the convex hull of the blob identified above as my prediction for the LV. I then picked a threshold value by estimating the median of the immediate surroundings of the blob. I used this value in the following.</p>\n\n<p>Once the LV was identified on frame 0 of the middle slice, I used thresholding with the above value on other frames and picked the blobs closest to the ones identified on the previous frames. I kept track of the distance of centroids of blobs between neighbouring frames in order to know when something went badly wrong.</p>\n\n<p>Once LV was identified on all frames of the middle slice I adopted the same procedure (picking nearest blob) changing the sax slice and keeping the frame number fixed. I had some sanity fallbacks so that the area would not grow too much on the edges - sometimes I used then intersection with the previous slice - in another attempt I tried watershed segmentation of the blob.\nHere I also used, when available, the intersection of 2CH and 4CH to align the neighboring slices - this helped in cases when some sax slices suddenly flipped from portrait to landscape or even once changed resolution. </p>\n\n<p>I also tried to extract the LV by looking at the approximate ellipse formed by the 2CH and 4CH lines on the sax slices. When the volume in this method as a function of time had a similar shape as the volume as a function of time in the other method I used it. I also tried deducing from this data when to neglect the outermost sax slices. This did not work as well as I hoped.</p>\n\n<p>I repeated the same procedure starting from the two neighbouring slices of the middle slice. I could then take the mean of the volumes obtained in these three tries ignoring cases when that centroid distance was too big.</p>\n\n<p>I also noticed that often the volume of one frame was much higher (or lower) than of its neighbors - to eliminate that I used some smoothing of the obtained volumes as a function of time. </p>\n\n<p><strong>Refinements</strong> </p>\n\n<p>In order to improve the predictions I used some tricks like fitting gradient boosted regression on systole, diastole and age and sex for the baseline model and using a NN (in keras optimizing MAE) to extract predictions from the minimal, maximal volume but also from the volumes in a couple of neighboring frames using the method described above, the 2CH/4CH ellipse and the method described above but with possible watershed segmentation for the outermost slices. From the NN model I got my final predictions for the systole and diastole. </p>\n\n<p>Then I evaluated the standard deviation of the errors and used it for predicting the cumulative probability distribution.\nOf course in the above the NN is used only for callibrating the final predictions and not for any image analysis.</p>\n\n<p><strong>Conclusions</strong></p>\n\n<p>My main problem with the CV algorithmic approach was that I am no cardiologist and there were cases where I thought that my algorithm does a great job while the predicted systole/diastole was significantly away from the ground truth. In other cases I had nice agreement. So I had no idea what to improve - especially at the outermost slices. People using CNN approach for segmentation/prediction did not have that problem I guess as the rules/algorithm were in a way implicitly deduced by the CNN.</p>\n\n<p>I did not have much time in the final weeks of the competition so I settled on exploiting the information in the 2CH and 4CH views. If I would have more time I would have definitely tried some CNN based segmentation approach.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "111401": "Once final results are revealed, I'd love to know how well approaches NOT based on neural networks did. Did any make it into the top 10?\r\n\r\nMy own attempt, such as it is (let's just say I'd be very happy if it managed to stay under 0.03) was all \"classic\" computer vision - it's an orgy in FFTs, clustering, interpolation, contour finding, region labeling and manipulation, with a smattering of plain vanilla classification and regression on top. Most of the development time went into dreaming up and trying features for LV identification and segmentation. The hardest part was reliably separating LV from RV and arteries. I'm pretty sure NN approaches had a huge advantage there.\r\n\r\nBiggest lesson learned, unless somebody managed to do significantly better without NNs: I have to get myself one of them newfangled GPUs. :D",
    "111408": "From our experience you can go atleast below 0.015 without using any NN's.",
    "111416": "0.015? Are you talking about the score at the first stage of the competition? That's a very good result!\r\n\r\nI was able to reach 0.020 during the first stage of the competition using simple methods (thresholding, ...). After that I used NN on this first result to improve my score to 0.014. Since I don't have a GPU, it has been a way to simplify a bit the model. Overall, I am able to localize well the left ventricule without NN for the slices around the middle and end of the ventricule. Where I had difficulties is to localize it correctly for the slices at the beginning of the ventricule (close to the left atrium). That's when I decided to try to use some NN.",
    "111562": "I used Active Appearance Models (and no neural nets) to get to .013433 on the final test set.\r\n\r\nMy biggest source of error was that I couldn't find a good way to recognize when the base of the heart \r\nhad been reached, so I frequently ended up segmenting one too many or one too few slices.\r\n\r\nOther than that, the segmentation of the axial slices looked quite accurate in most cases.",
    "111572": "We did use neural network to detect the ROI and to find a coarse contour, but our main technique is to convert the ROI to polar coordinate space and use dynamic programming to find the optimal contour.  We found that improving the NN model has less effect in improving the end result, as the final contour is found using color information. \r\n\r\nThe fourth image is also the output of NN, which is used to determine the left and right bound of the final contour (the two thin lines in the 3rd image).  Dynamic programming is then applied within the space between the two thin lines.\r\n\r\nI think NN in our solution can be replace with other traditional methods without changing the final performance.\r\n\r\nHaving looked at the top solutions, I feel that trying to find a materialized contour might actually hurt accuracy.  We had the same problem as PaulG, and the best method we have found is to always remove one last slice.",
    "111585": "My approach was to use OpenCV train cascade with direct volume calculation. I used modified detect method since I knew there are only one ventricle on image. For images with good quality it works very well. Also it was easy to filter incorrect rectangles since for one patient it must share at least one point (with some exceptions). Using these rectangles I obtain direct volume value as well many different area-based features for each sax like area of enclosing and inclosed circles, different thresholds, partial volumes etc. To improve results I later add all found data in one big table with addition of some DICOM data like age, sex, maximum axis difference and some features from 2ch and 4ch direct area analysis. Then I ran XGBoost 20KFold Monte-Carlo for ~1000 iterations for volume prediction, using median. It allows to achieve score around 0.155 on first set. And I think it's possible to improve this result.\r\n\r\n![1][1] \r\n[1]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/3913/output.gif",
    "111896": "Hi, thank you for sharing your results! You guys did great job.\r\n\r\nLet me share key points of our image processing approach which gave us CRPS about 0.023 only. :-(\r\n\r\nYou could look through our code from [Github repository (not completed&polished yet)][5].\r\n\r\nOur approach to LV detection uses intersection of 2ch, 4ch and sax slices and project that point to sax as initial estimation of LV position. Then we seek for nearest light blob using SIFT keypoint detector.\r\nThen we trained HOG detector and used it only when intersection based detector failed. Unfortunately this approach led to large number of detection errors.\r\n\r\nWe detected contours by \"color\" segmentation and Cascaded Pose Regression with gradient boosted trees (ferns). Also we applied cascaded pose regression for 2ch view in order to estimate LV height. This gave us two estimations of end-systolic/end-diastolic volume per person.\r\n\r\n - [Example of landmarks found on sax view (red points).][1]\r\n - [Example of landmarks found on 2ch view (red points).][2]\r\n\r\nAfter that we combined contour based estimations of volume with patient meta data in order to train XGBoost regression on the top.\r\n\r\nAlso we tried 3D reconstruction, but without great success.\r\n![3Dreconstruction][3]\r\n\r\nHere is plot of our predictions vs. groundtruth on validation subset.\r\n\r\n![prediction vs groundtruth][4]\r\n\r\n - Blue points - end-systolic volumes in mm^3\r\n - Green points - end-systolic volumes in mm^3\r\n\r\n  [1]: https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/sax_sample.png\r\n  [2]: https://github.com/brotherofken/national_data_science_bowl_2/blob/master/slides/img/ch2_sample.png\r\n  [3]: https://github.com/brotherofken/national_data_science_bowl_2/raw/master/slides/img/lv1.png\r\n  [4]: https://raw.githubusercontent.com/brotherofken/national_data_science_bowl_2/master/slides/img/predictions_vs_groundtruth.png\r\n  [5]: https://github.com/brotherofken/national_data_science_bowl_2",
    "111970": "Here are some details about my approach. I made a stupid error in the final submission code (hard coded 200 instead of the test set length in one for loop!!!) - so the corrected code has 0.015570 on the private LB which would give #21. It is not great so I'll just summarize the main features.\r\n\r\n**Preprocessing**\r\n\r\nHere I used the image orientation metadata to eliminate sax slices which are not parallel. I noticed that inclusion of such slices led to huge errors as the SliceLocation was often quite off.\r\n\r\n**Baselines** \r\n\r\nI used computer vision techniques (thresholding etc. - more on this later) so it was important to eliminate huge errors or cases when my algorithm failed completely. For that I used some baseline models:\r\n\r\n1) simple model based on age-sex metadata as a complete emergency fallback\r\n\r\n2) I used my earlier attempt as a fallback for the more refined model in two ways: \r\n\r\n     a) in case the more refined model failed completely I put in the fallback prediction, \r\n\r\n     b) in the other case I still averaged the cumulative probability distribution of the more refined model with the one of this model\r\n\r\n**Predictions**\r\n\r\nMy approach was to first identify the LV on the middle sax slice in frame 0. \r\n\r\nI used adaptive thresholding using skimage and identified the blob using the intersection of the 2CH and 4CH lines on the sax slice. In cases where this did not work or the 4CH views were absent I used my previous approach where I used some custom measure of circularity, the absolute value of the first Fourier coefficient etc. to identify the LV.\r\n\r\nIn order to fill in the muscles within the LV, I always used the convex hull of the blob identified above as my prediction for the LV. I then picked a threshold value by estimating the median of the immediate surroundings of the blob. I used this value in the following.\r\n\r\nOnce the LV was identified on frame 0 of the middle slice, I used thresholding with the above value on other frames and picked the blobs closest to the ones identified on the previous frames. I kept track of the distance of centroids of blobs between neighbouring frames in order to know when something went badly wrong.\r\n\r\nOnce LV was identified on all frames of the middle slice I adopted the same procedure (picking nearest blob) changing the sax slice and keeping the frame number fixed. I had some sanity fallbacks so that the area would not grow too much on the edges - sometimes I used then intersection with the previous slice - in another attempt I tried watershed segmentation of the blob.\r\nHere I also used, when available, the intersection of 2CH and 4CH to align the neighboring slices - this helped in cases when some sax slices suddenly flipped from portrait to landscape or even once changed resolution. \r\n\r\nI also tried to extract the LV by looking at the approximate ellipse formed by the 2CH and 4CH lines on the sax slices. When the volume in this method as a function of time had a similar shape as the volume as a function of time in the other method I used it. I also tried deducing from this data when to neglect the outermost sax slices. This did not work as well as I hoped.\r\n\r\nI repeated the same procedure starting from the two neighbouring slices of the middle slice. I could then take the mean of the volumes obtained in these three tries ignoring cases when that centroid distance was too big.\r\n\r\nI also noticed that often the volume of one frame was much higher (or lower) than of its neighbors - to eliminate that I used some smoothing of the obtained volumes as a function of time. \r\n\r\n**Refinements** \r\n\r\nIn order to improve the predictions I used some tricks like fitting gradient boosted regression on systole, diastole and age and sex for the baseline model and using a NN (in keras optimizing MAE) to extract predictions from the minimal, maximal volume but also from the volumes in a couple of neighboring frames using the method described above, the 2CH/4CH ellipse and the method described above but with possible watershed segmentation for the outermost slices. From the NN model I got my final predictions for the systole and diastole. \r\n\r\nThen I evaluated the standard deviation of the errors and used it for predicting the cumulative probability distribution.\r\nOf course in the above the NN is used only for callibrating the final predictions and not for any image analysis.\r\n\r\n**Conclusions**\r\n\r\nMy main problem with the CV algorithmic approach was that I am no cardiologist and there were cases where I thought that my algorithm does a great job while the predicted systole/diastole was significantly away from the ground truth. In other cases I had nice agreement. So I had no idea what to improve - especially at the outermost slices. People using CNN approach for segmentation/prediction did not have that problem I guess as the rules/algorithm were in a way implicitly deduced by the CNN.\r\n\r\nI did not have much time in the final weeks of the competition so I settled on exploiting the information in the 2CH and 4CH views. If I would have more time I would have definitely tried some CNN based segmentation approach."
  },
  "source": "meta"
}