{
  "id": 19711,
  "title": "11th Place Quick Summary",
  "url": "/competitions/second-annual-data-science-bowl/writeups/tim-hochberg-11th-place-quick-summary",
  "author_name": "",
  "post_date": "2016-03-22T16:57:23.847Z",
  "votes": 4,
  "comment_count": 4,
  "views": 1016,
  "content": "<p>I've been meaning to write up something about my approach, but I haven't been able to find the time, so here's a quick summary. </p>\n\n<p>But first: congratulations to the winning teams. Really impressive work!</p>\n\n<p>Onward, to my approach:</p>\n\n<ol>\n<li><p>Crop and down sample images to a common resolution and size (84x84). I played around with various complicated cropping strategies, but ended up just using center cropping.</p></li>\n<li><p>Divide the images from each heart into 3 sets; one from the bottom third, one from the middle third and one from the top third of the heart.</p></li>\n<li><p>Train a neural net to predict both CDFs directly that took 3 images at a time, one randomly selected from each set.  These images were fed in at either the odd or even time steps, randomly selected. This gave a lot more effective training samples than I would otherwise would have had. I also used translations and flips to augment the data.</p></li>\n<li><p>Compute the prediction for each heart by randomly selecting many different combinations for images from the three regions as well as different translations and flips and combing the resulting CDFs.  I got slightly better results by averaging the inverse CDFs than averaging the CDFs directly.</p></li>\n</ol>\n\n<p>That's pretty much it. This had the advantage of being quick to implement and not too computationally expensive: it could easily be run overnight on 3 year old macbook pro. Unfortunately, I ran into severe overfitting problems when I tried to up the accuracy of this by moving to more layers or higher resolutions.</p>",
  "messages": [
    {
      "id": "112655",
      "postDate": "03/22/2016 16:57:23",
      "content": "<p>I've been meaning to write up something about my approach, but I haven't been able to find the time, so here's a quick summary. </p>\n\n<p>But first: congratulations to the winning teams. Really impressive work!</p>\n\n<p>Onward, to my approach:</p>\n\n<ol>\n<li><p>Crop and down sample images to a common resolution and size (84x84). I played around with various complicated cropping strategies, but ended up just using center cropping.</p></li>\n<li><p>Divide the images from each heart into 3 sets; one from the bottom third, one from the middle third and one from the top third of the heart.</p></li>\n<li><p>Train a neural net to predict both CDFs directly that took 3 images at a time, one randomly selected from each set.  These images were fed in at either the odd or even time steps, randomly selected. This gave a lot more effective training samples than I would otherwise would have had. I also used translations and flips to augment the data.</p></li>\n<li><p>Compute the prediction for each heart by randomly selecting many different combinations for images from the three regions as well as different translations and flips and combing the resulting CDFs.  I got slightly better results by averaging the inverse CDFs than averaging the CDFs directly.</p></li>\n</ol>\n\n<p>That's pretty much it. This had the advantage of being quick to implement and not too computationally expensive: it could easily be run overnight on 3 year old macbook pro. Unfortunately, I ran into severe overfitting problems when I tried to up the accuracy of this by moving to more layers or higher resolutions.</p>",
      "rawMarkdown": "I've been meaning to write up something about my approach, but I haven't been able to find the time, so here's a quick summary. \r\n\r\nBut first: congratulations to the winning teams. Really impressive work!\r\n\r\nOnward, to my approach:\r\n\r\n1. Crop and down sample images to a common resolution and size (84x84). I played around with various complicated cropping strategies, but ended up just using center cropping.\r\n\r\n2. Divide the images from each heart into 3 sets; one from the bottom third, one from the middle third and one from the top third of the heart.\r\n\r\n3. Train a neural net to predict both CDFs directly that took 3 images at a time, one randomly selected from each set.  These images were fed in at either the odd or even time steps, randomly selected. This gave a lot more effective training samples than I would otherwise would have had. I also used translations and flips to augment the data.\r\n\r\n4. Compute the prediction for each heart by randomly selecting many different combinations for images from the three regions as well as different translations and flips and combing the resulting CDFs.  I got slightly better results by averaging the inverse CDFs than averaging the CDFs directly.\r\n\r\nThat's pretty much it. This had the advantage of being quick to implement and not too computationally expensive: it could easily be run overnight on 3 year old macbook pro. Unfortunately, I ran into severe overfitting problems when I tried to up the accuracy of this by moving to more layers or higher resolutions.",
      "votes": null
    },
    {
      "id": "112692",
      "postDate": "03/23/2016 02:09:26",
      "content": "<p>wow.. .. I am amazed you got such a low score with this approach.... what's the intuition behind the 3 regions? and what was your network architecture?  </p>",
      "rawMarkdown": "wow.. .. I am amazed you got such a low score with this approach.... what's the intuition behind the 3 regions? and what was your network architecture?",
      "votes": null
    },
    {
      "id": "112755",
      "postDate": "03/23/2016 16:18:39",
      "content": "<p>@DavidGbodiOdaibo, there wasn't any specific intuition about three regions. In fact, I would have preferred using 8 regions as that would have given decent spatial resolution only the short axis. However, when I went to more than three regions, overfitting overwhelmed any potential gains due to increased resolution.</p>\n\n<p>If you can read Lasagne, the actual net is below.  The 62 input layers are 30 time slices for the central region, 15 for each of the end regions (randomly selected even or odd time points), plus the spacing between the end slices and the central slice. </p>\n\n<p>With enough data,  I suspect you could get respectable results with extensions to this method (more regions etc), since I was limited by overfitting. However, segmenting the individual slices is really a better approach since then you get visualization of the volumes as well.</p>\n\n<pre><code>layers = [\n    LF(InputLayer, shape=(None, 62, SIZE, SIZE), layer_name=&quot;images&quot;),\n    LF(BatchNormLayer, nonlinearity=None),\n    #\n    LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify), \n    #\n    LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify), \n    #\n    LF(Conv2DLayer, num_filters=48, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=48, filter_size=(2,2), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify),\n    LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \n    #\n    LF(Conv2DLayer, num_filters=96, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=96, filter_size=(2,2), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify),   \n    LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \n    #\n    LF(Conv2DLayer, num_filters=192, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=192, filter_size=(2,2), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify),   \n    LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2), layer_name=&quot;jointOut&quot;),\n    ]\n\nfor name in [&quot;systole&quot;, &quot;diastole&quot;]:\n    layers += [\n        LF(Conv2DLayer, incoming=&quot;jointOut&quot;,  layer_name=name+&quot;In&quot;, \n                          num_filters=384, filter_size=(3,3), nonlinearity=leaky_rectify),\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \n        #\n        LF(DropoutLayer, p=0.5),            \n        LF(DenseLayer, num_units=OUTPUT_SIZE, nonlinearity=sigmoid),            \n        LF(ReshapeLayer, shape=([0],1,OUTPUT_SIZE), layer_name=name+&quot;Out&quot;)\n    ]\n\nlayers += [\n    LF(ConcatLayer, incomings=[&quot;systoleOut&quot;, &quot;diastoleOut&quot;], axis=1, layer_name=&quot;output&quot;)\n]\n</code></pre>",
      "rawMarkdown": "@DavidGbodiOdaibo, there wasn't any specific intuition about three regions. In fact, I would have preferred using 8 regions as that would have given decent spatial resolution only the short axis. However, when I went to more than three regions, overfitting overwhelmed any potential gains due to increased resolution.\r\n\r\nIf you can read Lasagne, the actual net is below.  The 62 input layers are 30 time slices for the central region, 15 for each of the end regions (randomly selected even or odd time points), plus the spacing between the end slices and the central slice. \r\n\r\nWith enough data,  I suspect you could get respectable results with extensions to this method (more regions etc), since I was limited by overfitting. However, segmenting the individual slices is really a better approach since then you get visualization of the volumes as well.\r\n\r\n\r\n\r\n    layers = [\r\n        LF(InputLayer, shape=(None, 62, SIZE, SIZE), layer_name=\"images\"),\r\n        LF(BatchNormLayer, nonlinearity=None),\r\n        #\r\n        LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify), \r\n        #\r\n        LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify), \r\n        #\r\n        LF(Conv2DLayer, num_filters=48, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=48, filter_size=(2,2), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify),\r\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \r\n        #\r\n        LF(Conv2DLayer, num_filters=96, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=96, filter_size=(2,2), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify),   \r\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \r\n        #\r\n        LF(Conv2DLayer, num_filters=192, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=192, filter_size=(2,2), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify),   \r\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2), layer_name=\"jointOut\"),\r\n        ]\r\n    \r\n    for name in [\"systole\", \"diastole\"]:\r\n        layers += [\r\n            LF(Conv2DLayer, incoming=\"jointOut\",  layer_name=name+\"In\", \r\n                              num_filters=384, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n            LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \r\n            #\r\n            LF(DropoutLayer, p=0.5),            \r\n            LF(DenseLayer, num_units=OUTPUT_SIZE, nonlinearity=sigmoid),            \r\n            LF(ReshapeLayer, shape=([0],1,OUTPUT_SIZE), layer_name=name+\"Out\")\r\n        ]\r\n        \r\n    layers += [\r\n        LF(ConcatLayer, incomings=[\"systoleOut\", \"diastoleOut\"], axis=1, layer_name=\"output\")\r\n    ]",
      "votes": null
    },
    {
      "id": "112778",
      "postDate": "03/23/2016 18:31:01",
      "content": "<p>wow and you did not use a lot of 3x3 kernels. I am still trying to figure out what I did wrong I think the pixelspacing really messed me up. Did you rescale/normalize the images with pixel spacing in preprocessing, you said you cropped them to 84X84 but did you rescale with pixelspacing.</p>",
      "rawMarkdown": "wow and you did not use a lot of 3x3 kernels. I am still trying to figure out what I did wrong I think the pixelspacing really messed me up. Did you rescale/normalize the images with pixel spacing in preprocessing, you said you cropped them to 84X84 but did you rescale with pixelspacing.",
      "votes": null
    },
    {
      "id": "112779",
      "postDate": "03/23/2016 18:34:34",
      "content": "<p>@DavidGbodiOdaibo, yes the images were scaled so that the pixels in each image where the same size.</p>",
      "rawMarkdown": "DavidGbodiOdaibo, yes the images were scaled so that the pixels in each image where the same size.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 112692,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "03/23/2016 02:09:26",
      "content": "<p>wow.. .. I am amazed you got such a low score with this approach.... what's the intuition behind the 3 regions? and what was your network architecture?  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112755,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "03/23/2016 16:18:39",
      "content": "<p>@DavidGbodiOdaibo, there wasn't any specific intuition about three regions. In fact, I would have preferred using 8 regions as that would have given decent spatial resolution only the short axis. However, when I went to more than three regions, overfitting overwhelmed any potential gains due to increased resolution.</p>\n\n<p>If you can read Lasagne, the actual net is below.  The 62 input layers are 30 time slices for the central region, 15 for each of the end regions (randomly selected even or odd time points), plus the spacing between the end slices and the central slice. </p>\n\n<p>With enough data,  I suspect you could get respectable results with extensions to this method (more regions etc), since I was limited by overfitting. However, segmenting the individual slices is really a better approach since then you get visualization of the volumes as well.</p>\n\n<pre><code>layers = [\n    LF(InputLayer, shape=(None, 62, SIZE, SIZE), layer_name=&quot;images&quot;),\n    LF(BatchNormLayer, nonlinearity=None),\n    #\n    LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify), \n    #\n    LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify), \n    #\n    LF(Conv2DLayer, num_filters=48, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=48, filter_size=(2,2), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify),\n    LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \n    #\n    LF(Conv2DLayer, num_filters=96, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=96, filter_size=(2,2), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify),   \n    LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \n    #\n    LF(Conv2DLayer, num_filters=192, filter_size=(3,3), nonlinearity=leaky_rectify),\n    LF(Conv2DLayer, num_filters=192, filter_size=(2,2), nonlinearity=None, b=None),\n    LF(BatchNormLayer, nonlinearity=leaky_rectify),   \n    LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2), layer_name=&quot;jointOut&quot;),\n    ]\n\nfor name in [&quot;systole&quot;, &quot;diastole&quot;]:\n    layers += [\n        LF(Conv2DLayer, incoming=&quot;jointOut&quot;,  layer_name=name+&quot;In&quot;, \n                          num_filters=384, filter_size=(3,3), nonlinearity=leaky_rectify),\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \n        #\n        LF(DropoutLayer, p=0.5),            \n        LF(DenseLayer, num_units=OUTPUT_SIZE, nonlinearity=sigmoid),            \n        LF(ReshapeLayer, shape=([0],1,OUTPUT_SIZE), layer_name=name+&quot;Out&quot;)\n    ]\n\nlayers += [\n    LF(ConcatLayer, incomings=[&quot;systoleOut&quot;, &quot;diastoleOut&quot;], axis=1, layer_name=&quot;output&quot;)\n]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112778,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "03/23/2016 18:31:01",
      "content": "<p>wow and you did not use a lot of 3x3 kernels. I am still trying to figure out what I did wrong I think the pixelspacing really messed me up. Did you rescale/normalize the images with pixel spacing in preprocessing, you said you cropped them to 84X84 but did you rescale with pixelspacing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112779,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "03/23/2016 18:34:34",
      "content": "<p>@DavidGbodiOdaibo, yes the images were scaled so that the pixels in each image where the same size.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "112655": "I've been meaning to write up something about my approach, but I haven't been able to find the time, so here's a quick summary. \r\n\r\nBut first: congratulations to the winning teams. Really impressive work!\r\n\r\nOnward, to my approach:\r\n\r\n1. Crop and down sample images to a common resolution and size (84x84). I played around with various complicated cropping strategies, but ended up just using center cropping.\r\n\r\n2. Divide the images from each heart into 3 sets; one from the bottom third, one from the middle third and one from the top third of the heart.\r\n\r\n3. Train a neural net to predict both CDFs directly that took 3 images at a time, one randomly selected from each set.  These images were fed in at either the odd or even time steps, randomly selected. This gave a lot more effective training samples than I would otherwise would have had. I also used translations and flips to augment the data.\r\n\r\n4. Compute the prediction for each heart by randomly selecting many different combinations for images from the three regions as well as different translations and flips and combing the resulting CDFs.  I got slightly better results by averaging the inverse CDFs than averaging the CDFs directly.\r\n\r\nThat's pretty much it. This had the advantage of being quick to implement and not too computationally expensive: it could easily be run overnight on 3 year old macbook pro. Unfortunately, I ran into severe overfitting problems when I tried to up the accuracy of this by moving to more layers or higher resolutions.",
    "112692": "wow.. .. I am amazed you got such a low score with this approach.... what's the intuition behind the 3 regions? and what was your network architecture?",
    "112755": "@DavidGbodiOdaibo, there wasn't any specific intuition about three regions. In fact, I would have preferred using 8 regions as that would have given decent spatial resolution only the short axis. However, when I went to more than three regions, overfitting overwhelmed any potential gains due to increased resolution.\r\n\r\nIf you can read Lasagne, the actual net is below.  The 62 input layers are 30 time slices for the central region, 15 for each of the end regions (randomly selected even or odd time points), plus the spacing between the end slices and the central slice. \r\n\r\nWith enough data,  I suspect you could get respectable results with extensions to this method (more regions etc), since I was limited by overfitting. However, segmenting the individual slices is really a better approach since then you get visualization of the volumes as well.\r\n\r\n\r\n\r\n    layers = [\r\n        LF(InputLayer, shape=(None, 62, SIZE, SIZE), layer_name=\"images\"),\r\n        LF(BatchNormLayer, nonlinearity=None),\r\n        #\r\n        LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=16, filter_size=(3,3), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify), \r\n        #\r\n        LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=32, filter_size=(3,3), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify), \r\n        #\r\n        LF(Conv2DLayer, num_filters=48, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=48, filter_size=(2,2), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify),\r\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \r\n        #\r\n        LF(Conv2DLayer, num_filters=96, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=96, filter_size=(2,2), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify),   \r\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \r\n        #\r\n        LF(Conv2DLayer, num_filters=192, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n        LF(Conv2DLayer, num_filters=192, filter_size=(2,2), nonlinearity=None, b=None),\r\n        LF(BatchNormLayer, nonlinearity=leaky_rectify),   \r\n        LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2), layer_name=\"jointOut\"),\r\n        ]\r\n    \r\n    for name in [\"systole\", \"diastole\"]:\r\n        layers += [\r\n            LF(Conv2DLayer, incoming=\"jointOut\",  layer_name=name+\"In\", \r\n                              num_filters=384, filter_size=(3,3), nonlinearity=leaky_rectify),\r\n            LF(MaxPool2DLayer, pool_size=(3,3), stride=(2,2)), \r\n            #\r\n            LF(DropoutLayer, p=0.5),            \r\n            LF(DenseLayer, num_units=OUTPUT_SIZE, nonlinearity=sigmoid),            \r\n            LF(ReshapeLayer, shape=([0],1,OUTPUT_SIZE), layer_name=name+\"Out\")\r\n        ]\r\n        \r\n    layers += [\r\n        LF(ConcatLayer, incomings=[\"systoleOut\", \"diastoleOut\"], axis=1, layer_name=\"output\")\r\n    ]",
    "112778": "wow and you did not use a lot of 3x3 kernels. I am still trying to figure out what I did wrong I think the pixelspacing really messed me up. Did you rescale/normalize the images with pixel spacing in preprocessing, you said you cropped them to 84X84 but did you rescale with pixelspacing.",
    "112779": "DavidGbodiOdaibo, yes the images were scaled so that the pixels in each image where the same size."
  },
  "source": "meta"
}