{
  "id": 19256,
  "title": "Starting from scratch",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19256",
  "author_name": "",
  "post_date": "2016-03-01T12:38:41.903Z",
  "votes": 4,
  "comment_count": 6,
  "views": 951,
  "content": "<p>Hi All,</p>\n\n<p>I gave up on the tutorials because I couldn't figure out how to tune them for better performance. So I decided to develop my own methodology, starting from scratch. What I'm doing is:\n1) clean up the images (aspect ratio, contrast, brightness, deduplicate)\n2) train a convolutional neural network to find the centroid of the left ventricle in the images\n3) crop the image around the left ventricle, then binarise the image and apply image morphology to smooth the blobs and fill any holes inside\n4) train a GBM to recognise which blob is the left ventricle\n5) track the blob through the 30 time slices, to find the times of diastole and systole\n6) track the diastole and systole blobs through the sax location slices to get a set of slice areas\n7) train xgboost to estimate the left ventricle volume, with features derived from the dicom image header and the sax location slices</p>\n\n<p>I learned a lot from doing all of this, but it sure took a lot of time, and I don't think I have time left to improve the model :-(</p>\n\n<p>I will write blog tutorial if enough people are interested.</p>\n\n<p>Colin</p>",
  "messages": [
    {
      "id": "109893",
      "postDate": "03/01/2016 12:38:41",
      "content": "<p>Hi All,</p>\n\n<p>I gave up on the tutorials because I couldn't figure out how to tune them for better performance. So I decided to develop my own methodology, starting from scratch. What I'm doing is:\n1) clean up the images (aspect ratio, contrast, brightness, deduplicate)\n2) train a convolutional neural network to find the centroid of the left ventricle in the images\n3) crop the image around the left ventricle, then binarise the image and apply image morphology to smooth the blobs and fill any holes inside\n4) train a GBM to recognise which blob is the left ventricle\n5) track the blob through the 30 time slices, to find the times of diastole and systole\n6) track the diastole and systole blobs through the sax location slices to get a set of slice areas\n7) train xgboost to estimate the left ventricle volume, with features derived from the dicom image header and the sax location slices</p>\n\n<p>I learned a lot from doing all of this, but it sure took a lot of time, and I don't think I have time left to improve the model :-(</p>\n\n<p>I will write blog tutorial if enough people are interested.</p>\n\n<p>Colin</p>",
      "rawMarkdown": "Hi All,\r\n\r\nI gave up on the tutorials because I couldn't figure out how to tune them for better performance. So I decided to develop my own methodology, starting from scratch. What I'm doing is:\r\n1) clean up the images (aspect ratio, contrast, brightness, deduplicate)\r\n2) train a convolutional neural network to find the centroid of the left ventricle in the images\r\n3) crop the image around the left ventricle, then binarise the image and apply image morphology to smooth the blobs and fill any holes inside\r\n4) train a GBM to recognise which blob is the left ventricle\r\n5) track the blob through the 30 time slices, to find the times of diastole and systole\r\n6) track the diastole and systole blobs through the sax location slices to get a set of slice areas\r\n7) train xgboost to estimate the left ventricle volume, with features derived from the dicom image header and the sax location slices\r\n\r\nI learned a lot from doing all of this, but it sure took a lot of time, and I don't think I have time left to improve the model :-(\r\n\r\nI will write blog tutorial if enough people are interested.\r\n\r\nColin",
      "votes": null
    },
    {
      "id": "110500",
      "postDate": "03/06/2016 03:54:14",
      "content": "<p>Hi Colin, I am very interested in your methodology. And I am eager to see your blogs. Please go on~</p>",
      "rawMarkdown": "Hi Colin, I am very interested in your methodology. And I am eager to see your blogs. Please go on~",
      "votes": null
    },
    {
      "id": "110518",
      "postDate": "03/06/2016 08:18:15",
      "content": "<p>Hi HawkWang,</p>\n\n<p>Thanks for the encouragement. </p>\n\n<p>Since it's quite clear that I don't have the time to improve my models enough to get a placing, I'm going to start publishing my approach. So I've just posted my first blog about this competition <a href=\"http://colinpriest.com/2016/03/06/second-annual-data-science-bowl-part-1/\">http://colinpriest.com/2016/03/06/second-annual-data-science-bowl-part-1/</a> </p>\n\n<p>In it I describe the structure of the image data, and the problems with the image quality. Normalising the pixel brightnesses (to mean 0.5, standard deviation 0.25) doesn't work very well. Instead I use an empirical approach, choosing exemplar images, and then mapping the pixel brightness rankings across to a poor quality image.</p>\n\n<p>The blog provides an R script for doing this.</p>\n\n<p>Colin</p>",
      "rawMarkdown": "Hi HawkWang,\r\n\r\nThanks for the encouragement. \r\n\r\nSince it's quite clear that I don't have the time to improve my models enough to get a placing, I'm going to start publishing my approach. So I've just posted my first blog about this competition http://colinpriest.com/2016/03/06/second-annual-data-science-bowl-part-1/ \r\n\r\nIn it I describe the structure of the image data, and the problems with the image quality. Normalising the pixel brightnesses (to mean 0.5, standard deviation 0.25) doesn't work very well. Instead I use an empirical approach, choosing exemplar images, and then mapping the pixel brightness rankings across to a poor quality image.\r\n\r\nThe blog provides an R script for doing this.\r\n\r\nColin",
      "votes": null
    },
    {
      "id": "110534",
      "postDate": "03/06/2016 10:54:53",
      "content": "<p>Nice. I've also started to write down my approach - it took weeks to develop and days to run, see <a href=\"http://ottopdatascience.blogspot.nl/p/blog-page.html\">http://ottopdatascience.blogspot.nl/p/blog-page.html</a></p>",
      "rawMarkdown": "Nice. I've also started to write down my approach - it took weeks to develop and days to run, see http://ottopdatascience.blogspot.nl/p/blog-page.html",
      "votes": null
    },
    {
      "id": "110537",
      "postDate": "03/06/2016 11:02:03",
      "content": "<p>Hi Otto,</p>\n\n<p>I really like your blob detector. Mine isn't as good as yours - I hadn't realised that EBImage could give me those features!</p>\n\n<p>Colin</p>",
      "rawMarkdown": "Hi Otto,\r\n\r\nI really like your blob detector. Mine isn't as good as yours - I hadn't realised that EBImage could give me those features!\r\n\r\nColin",
      "votes": null
    },
    {
      "id": "110538",
      "postDate": "03/06/2016 11:15:59",
      "content": "<p>Hi Colin - we should have paired up ;). The LV detector for good images is pretty strong indeed, but I believe the main weakness in my code is lacking to take more advantage of the 3-D structure by projecting the LV from the middle slices better into the slices away from the middle.</p>",
      "rawMarkdown": "Hi Colin - we should have paired up ;). The LV detector for good images is pretty strong indeed, but I believe the main weakness in my code is lacking to take more advantage of the 3-D structure by projecting the LV from the middle slices better into the slices away from the middle.",
      "votes": null
    },
    {
      "id": "110613",
      "postDate": "03/07/2016 00:35:13",
      "content": "<p>Hi All,</p>\n\n<p>Here's my second blog, showing how I cleaned up the images and rearranged them, ready for a convolutional neural network. <a href=\"http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/\">http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/</a> </p>\n\n<p>Colin</p>",
      "rawMarkdown": "Hi All,\r\n\r\nHere's my second blog, showing how I cleaned up the images and rearranged them, ready for a convolutional neural network. http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/ \r\n\r\nColin",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 110500,
      "author_name": "yourwanghao",
      "author_url": "",
      "post_date": "03/06/2016 03:54:14",
      "content": "<p>Hi Colin, I am very interested in your methodology. And I am eager to see your blogs. Please go on~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110518,
      "author_name": "colinpriest",
      "author_url": "",
      "post_date": "03/06/2016 08:18:15",
      "content": "<p>Hi HawkWang,</p>\n\n<p>Thanks for the encouragement. </p>\n\n<p>Since it's quite clear that I don't have the time to improve my models enough to get a placing, I'm going to start publishing my approach. So I've just posted my first blog about this competition <a href=\"http://colinpriest.com/2016/03/06/second-annual-data-science-bowl-part-1/\">http://colinpriest.com/2016/03/06/second-annual-data-science-bowl-part-1/</a> </p>\n\n<p>In it I describe the structure of the image data, and the problems with the image quality. Normalising the pixel brightnesses (to mean 0.5, standard deviation 0.25) doesn't work very well. Instead I use an empirical approach, choosing exemplar images, and then mapping the pixel brightness rankings across to a poor quality image.</p>\n\n<p>The blog provides an R script for doing this.</p>\n\n<p>Colin</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110534,
      "author_name": "operdeck",
      "author_url": "",
      "post_date": "03/06/2016 10:54:53",
      "content": "<p>Nice. I've also started to write down my approach - it took weeks to develop and days to run, see <a href=\"http://ottopdatascience.blogspot.nl/p/blog-page.html\">http://ottopdatascience.blogspot.nl/p/blog-page.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110537,
      "author_name": "colinpriest",
      "author_url": "",
      "post_date": "03/06/2016 11:02:03",
      "content": "<p>Hi Otto,</p>\n\n<p>I really like your blob detector. Mine isn't as good as yours - I hadn't realised that EBImage could give me those features!</p>\n\n<p>Colin</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110538,
      "author_name": "operdeck",
      "author_url": "",
      "post_date": "03/06/2016 11:15:59",
      "content": "<p>Hi Colin - we should have paired up ;). The LV detector for good images is pretty strong indeed, but I believe the main weakness in my code is lacking to take more advantage of the 3-D structure by projecting the LV from the middle slices better into the slices away from the middle.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110613,
      "author_name": "colinpriest",
      "author_url": "",
      "post_date": "03/07/2016 00:35:13",
      "content": "<p>Hi All,</p>\n\n<p>Here's my second blog, showing how I cleaned up the images and rearranged them, ready for a convolutional neural network. <a href=\"http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/\">http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/</a> </p>\n\n<p>Colin</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "109893": "Hi All,\r\n\r\nI gave up on the tutorials because I couldn't figure out how to tune them for better performance. So I decided to develop my own methodology, starting from scratch. What I'm doing is:\r\n1) clean up the images (aspect ratio, contrast, brightness, deduplicate)\r\n2) train a convolutional neural network to find the centroid of the left ventricle in the images\r\n3) crop the image around the left ventricle, then binarise the image and apply image morphology to smooth the blobs and fill any holes inside\r\n4) train a GBM to recognise which blob is the left ventricle\r\n5) track the blob through the 30 time slices, to find the times of diastole and systole\r\n6) track the diastole and systole blobs through the sax location slices to get a set of slice areas\r\n7) train xgboost to estimate the left ventricle volume, with features derived from the dicom image header and the sax location slices\r\n\r\nI learned a lot from doing all of this, but it sure took a lot of time, and I don't think I have time left to improve the model :-(\r\n\r\nI will write blog tutorial if enough people are interested.\r\n\r\nColin",
    "110500": "Hi Colin, I am very interested in your methodology. And I am eager to see your blogs. Please go on~",
    "110518": "Hi HawkWang,\r\n\r\nThanks for the encouragement. \r\n\r\nSince it's quite clear that I don't have the time to improve my models enough to get a placing, I'm going to start publishing my approach. So I've just posted my first blog about this competition http://colinpriest.com/2016/03/06/second-annual-data-science-bowl-part-1/ \r\n\r\nIn it I describe the structure of the image data, and the problems with the image quality. Normalising the pixel brightnesses (to mean 0.5, standard deviation 0.25) doesn't work very well. Instead I use an empirical approach, choosing exemplar images, and then mapping the pixel brightness rankings across to a poor quality image.\r\n\r\nThe blog provides an R script for doing this.\r\n\r\nColin",
    "110534": "Nice. I've also started to write down my approach - it took weeks to develop and days to run, see http://ottopdatascience.blogspot.nl/p/blog-page.html",
    "110537": "Hi Otto,\r\n\r\nI really like your blob detector. Mine isn't as good as yours - I hadn't realised that EBImage could give me those features!\r\n\r\nColin",
    "110538": "Hi Colin - we should have paired up ;). The LV detector for good images is pretty strong indeed, but I believe the main weakness in my code is lacking to take more advantage of the 3-D structure by projecting the LV from the middle slices better into the slices away from the middle.",
    "110613": "Hi All,\r\n\r\nHere's my second blog, showing how I cleaned up the images and rearranged them, ready for a convolutional neural network. http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/ \r\n\r\nColin"
  },
  "source": "meta"
}