{
  "id": 18783,
  "title": "BAH-NVIDIA Team Blog Post #2",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18783",
  "author_name": "",
  "post_date": "2016-02-05T22:59:06.310Z",
  "votes": 5,
  "comment_count": 20,
  "views": 4141,
  "content": "<p>Hi!  We are pleased to share some notes on our current deep learning based approach to the contest.  Whilst we haven't given all the model specifics, we're happy to engage in discussion about our methods and what has and hasn't worked!</p>\n\n<p><a href=\"http://www.datasciencebowl.com/first_dl_submission/\">http://www.datasciencebowl.com/first_dl_submission/</a></p>",
  "messages": [
    {
      "id": "107024",
      "postDate": "02/05/2016 22:59:06",
      "content": "<p>Hi!  We are pleased to share some notes on our current deep learning based approach to the contest.  Whilst we haven't given all the model specifics, we're happy to engage in discussion about our methods and what has and hasn't worked!</p>\n\n<p><a href=\"http://www.datasciencebowl.com/first_dl_submission/\">http://www.datasciencebowl.com/first_dl_submission/</a></p>",
      "rawMarkdown": "Hi!  We are pleased to share some notes on our current deep learning based approach to the contest.  Whilst we haven't given all the model specifics, we're happy to engage in discussion about our methods and what has and hasn't worked!\r\n\r\nhttp://www.datasciencebowl.com/first_dl_submission/",
      "votes": null
    },
    {
      "id": "107053",
      "postDate": "02/06/2016 07:00:36",
      "content": "<p>Nice to have a Titan X, How many iterations typically needs to reach around 0.23? I am using GTX 860M which is like a dwarf next to Titan X and needs close to 350 iterations each taking close to 2.5 minutes to reach around 0.31.</p>",
      "rawMarkdown": "Nice to have a Titan X, How many iterations typically needs to reach around 0.23? I am using GTX 860M which is like a dwarf next to Titan X and needs close to 350 iterations each taking close to 2.5 minutes to reach around 0.31.",
      "votes": null
    },
    {
      "id": "107072",
      "postDate": "02/06/2016 12:54:26",
      "content": "<p>Each iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.</p>",
      "rawMarkdown": "Each iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.",
      "votes": null
    },
    {
      "id": "107079",
      "postDate": "02/06/2016 15:24:57",
      "content": "<p>[quote=Senecaur;107072]</p>\n\n<p>Each iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.</p>\n\n<p>[/quote]</p>\n\n<p>@Senecaur Would you please detail the strategy of dealing with overfitting? </p>",
      "rawMarkdown": "[quote=Senecaur;107072]\r\n\r\nEach iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.\r\n\r\n[/quote]\r\n\r\n\r\n@Senecaur Would you please detail the strategy of dealing with overfitting?",
      "votes": null
    },
    {
      "id": "107127",
      "postDate": "02/07/2016 09:53:07",
      "content": "<p>Maybe you should have used something like convex hulls, to get rid of the black points in the circle. ;)</p>",
      "rawMarkdown": "Maybe you should have used something like convex hulls, to get rid of the black points in the circle. ;)",
      "votes": null
    },
    {
      "id": "107133",
      "postDate": "02/07/2016 14:59:54",
      "content": "<p>@PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.</p>",
      "rawMarkdown": "PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.",
      "votes": null
    },
    {
      "id": "107134",
      "postDate": "02/07/2016 15:00:48",
      "content": "<p>@Icedragon - Great suggestion.  To be clear we are not currently using the segmented images as an input to our model, but we are experimenting with segmentation based models too, so I think that some form of smoothing could be very beneficial.</p>",
      "rawMarkdown": "Icedragon - Great suggestion.  To be clear we are not currently using the segmented images as an input to our model, but we are experimenting with segmentation based models too, so I think that some form of smoothing could be very beneficial.",
      "votes": null
    },
    {
      "id": "107213",
      "postDate": "02/08/2016 13:10:33",
      "content": "<p>@Senecaur. Thanks for posting this. I had two questoins: </p>\n\n<ul>\n<li><p>You mention that the model &quot;predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently&quot;. I was wondering - does the model predict 600 logistic values (all between 0 and 1) which are the subsequently &quot;manually&quot; transformed into a CDF, or does the model directly output a CDF that starts at 0 for the value 1 and ends at 1 for value 600? </p>\n\n<ul><li>you are mentioning that &quot;we are currently choosing to standardize the dimensions of our input by randomly sampling a fixed number of slices and a fixed number of timesteps from each patient study before ingesting it into the CNN&quot;. I am trying to wrap my head around this. When you say slice - do you mean the slice location (e.g. &quot;sax42&quot;) or do you mean an invididual image (e.g. image 0004 at sax42)? I am understanding it correctly when i assume that for each patient you take a random subset of slice locations (e.g. sax 34 and sax 35 for patient 1, sax 64 and sax 42 for patient 2, etc.) and only use a random time-period (e.g. seconds 40 to 80) that is similar for each patient?</li></ul></li>\n</ul>\n\n<p>please let me know! thanks! </p>",
      "rawMarkdown": "Senecaur. Thanks for posting this. I had two questoins: \r\n\r\n - You mention that the model \"predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently\". I was wondering - does the model predict 600 logistic values (all between 0 and 1) which are the subsequently \"manually\" transformed into a CDF, or does the model directly output a CDF that starts at 0 for the value 1 and ends at 1 for value 600? \r\n\r\n- you are mentioning that \"we are currently choosing to standardize the dimensions of our input by randomly sampling a fixed number of slices and a fixed number of timesteps from each patient study before ingesting it into the CNN\". I am trying to wrap my head around this. When you say slice - do you mean the slice location (e.g. \"sax42\") or do you mean an invididual image (e.g. image 0004 at sax42)? I am understanding it correctly when i assume that for each patient you take a random subset of slice locations (e.g. sax 34 and sax 35 for patient 1, sax 64 and sax 42 for patient 2, etc.) and only use a random time-period (e.g. seconds 40 to 80) that is similar for each patient?\r\n\r\nplease let me know! thanks!",
      "votes": null
    },
    {
      "id": "107215",
      "postDate": "02/08/2016 13:31:26",
      "content": "<p>@WD</p>\n\n<p>For systolic and diastolic separately we output a 600-dimensional softmax and then apply a cumulative sum operation to get our CDF prediction.  </p>\n\n<p>You are correct, by slice I mean the slice location within a patient study, e.g. sax42, not an individual image.  Each iteration we randomly pick a fixed number of these slices and we pick a fixed number of timesteps from within the slice.  We do not pick these to be similar for each patient, everything is randomized on a per patient per iteration basis.  The fixed numbers of slices and timesteps are quite large though, so actually our sampling is not very sparse at all.</p>",
      "rawMarkdown": "WD\r\n\r\nFor systolic and diastolic separately we output a 600-dimensional softmax and then apply a cumulative sum operation to get our CDF prediction.  \r\n\r\nYou are correct, by slice I mean the slice location within a patient study, e.g. sax42, not an individual image.  Each iteration we randomly pick a fixed number of these slices and we pick a fixed number of timesteps from within the slice.  We do not pick these to be similar for each patient, everything is randomized on a per patient per iteration basis.  The fixed numbers of slices and timesteps are quite large though, so actually our sampling is not very sparse at all.",
      "votes": null
    },
    {
      "id": "107221",
      "postDate": "02/08/2016 14:22:26",
      "content": "<p>@Senecaur. Thanks so much. That makes a lot of sense. I am relatively new to image recognition, and might not know all the tools - especially for preprocessing. What kind of tools / packages did you use for the image scaling and segmentation? I presume that these tools are &quot;outside&quot; lasagne, any links / references would be much appreciated! </p>",
      "rawMarkdown": "Senecaur. Thanks so much. That makes a lot of sense. I am relatively new to image recognition, and might not know all the tools - especially for preprocessing. What kind of tools / packages did you use for the image scaling and segmentation? I presume that these tools are \"outside\" lasagne, any links / references would be much appreciated!",
      "votes": null
    },
    {
      "id": "107222",
      "postDate": "02/08/2016 14:27:14",
      "content": "<p>@WD No problem - our image preprocessing and augmentation is all done using the Python library <a href=\"http://scikit-image.org/\">scikit-image</a>, this means it is easy to integrate these processes inline with Lasagne based models.  If you need to manually label/segment any of the training images then I recommend <a href=\"https://cvhci.anthropomatik.kit.edu/~baeuml/projects/a-universal-labeling-tool-for-computer-vision-sloth/\">Sloth</a></p>",
      "rawMarkdown": "WD No problem - our image preprocessing and augmentation is all done using the Python library [scikit-image][1], this means it is easy to integrate these processes inline with Lasagne based models.  If you need to manually label/segment any of the training images then I recommend [Sloth][2]\r\n\r\n  [1]: http://scikit-image.org/\r\n  [2]: https://cvhci.anthropomatik.kit.edu/~baeuml/projects/a-universal-labeling-tool-for-computer-vision-sloth/",
      "votes": null
    },
    {
      "id": "107261",
      "postDate": "02/08/2016 21:35:22",
      "content": "<p>@Senecaur - thanks so much. This is very interesting. Are you using 64*64 images? Can you provide further explanation / intuition as well why the random sampling of the slices and the the time-intervals works? I understand that it might provide additonal regularization, but from your description it seemed that the original thought came from the limitations in the data (with regards to slice availability, etc.)? </p>",
      "rawMarkdown": "Senecaur - thanks so much. This is very interesting. Are you using 64*64 images? Can you provide further explanation / intuition as well why the random sampling of the slices and the the time-intervals works? I understand that it might provide additonal regularization, but from your description it seemed that the original thought came from the limitations in the data (with regards to slice availability, etc.)?",
      "votes": null
    },
    {
      "id": "107322",
      "postDate": "02/09/2016 05:33:53",
      "content": "<p>[quote=Senecaur;107133]</p>\n\n<p>@PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.</p>\n\n<p>[/quote]\n@Senecaur, I ran into heavy overfitting due to powerful model and limited data.  I used Batch Normalization and Global Averaging Pooling which claimed that Dropout is better not involved.  Thank you for your strategy!</p>",
      "rawMarkdown": "[quote=Senecaur;107133]\r\n\r\n@PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.\r\n\r\n[/quote]\r\n@Senecaur, I ran into heavy overfitting due to powerful model and limited data.  I used Batch Normalization and Global Averaging Pooling which claimed that Dropout is better not involved.  Thank you for your strategy!",
      "votes": null
    },
    {
      "id": "108442",
      "postDate": "02/17/2016 11:09:18",
      "content": "<p>@senecaur - did you use k-means to smoothen out the images? much appreciated! </p>",
      "rawMarkdown": "senecaur - did you use k-means to smoothen out the images? much appreciated!",
      "votes": null
    },
    {
      "id": "108590",
      "postDate": "02/18/2016 14:36:02",
      "content": "<p>@Senecaur, Your idea about standardizing the dimensions of input is brilliant, but I have a question about your test data. </p>\n\n<p>You mention &quot;predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently&quot;, but your train data for each patient study have fixed number images(randomly picked from different slices and timesteps),  let's assume the fixed number is 60 for each patient study, how do you choose 60 images for each patient in test?</p>",
      "rawMarkdown": "Senecaur, Your idea about standardizing the dimensions of input is brilliant, but I have a question about your test data. \r\n\r\nYou mention \"predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently\", but your train data for each patient study have fixed number images(randomly picked from different slices and timesteps),  let's assume the fixed number is 60 for each patient study, how do you choose 60 images for each patient in test?",
      "votes": null
    },
    {
      "id": "108614",
      "postDate": "02/18/2016 17:55:08",
      "content": "<p>@WD I tried both 64x64 and 128x128 images.  For the method I described there was no real improvement from using the larger images but it was much slower and more memory hungry to train the model.  You are right that the random sampling of time-steps and slices was more born out of the need to standardize the input tensor than with the plan that it would improve model performance somehow.  I am taking an educated guess that it helps with regularization - I don't really have proof that that is the case!</p>",
      "rawMarkdown": "WD I tried both 64x64 and 128x128 images.  For the method I described there was no real improvement from using the larger images but it was much slower and more memory hungry to train the model.  You are right that the random sampling of time-steps and slices was more born out of the need to standardize the input tensor than with the plan that it would improve model performance somehow.  I am taking an educated guess that it helps with regularization - I don't really have proof that that is the case!",
      "votes": null
    },
    {
      "id": "108615",
      "postDate": "02/18/2016 17:55:38",
      "content": "<p>@WD Yes, we do use k-means</p>",
      "rawMarkdown": "WD Yes, we do use k-means",
      "votes": null
    },
    {
      "id": "108616",
      "postDate": "02/18/2016 17:57:32",
      "content": "<p>@Gzs_iceberg We choose the slices and time-steps for the test data in the same way - randomly - and then average over multiple samples for the same test patient study.  Test time augmentation has proven in previous competitions (e.g. Galaxy Zoo, DSB1) to be a benefit even when it is not necessary, so it seemed reasonable to do here too.</p>",
      "rawMarkdown": "Gzs_iceberg We choose the slices and time-steps for the test data in the same way - randomly - and then average over multiple samples for the same test patient study.  Test time augmentation has proven in previous competitions (e.g. Galaxy Zoo, DSB1) to be a benefit even when it is not necessary, so it seemed reasonable to do here too.",
      "votes": null
    },
    {
      "id": "108651",
      "postDate": "02/18/2016 23:02:24",
      "content": "<p>@Senecaur / others. I am new to image manipulation and still a beginner to CNNs. I am using an implementation where one glues together a stack of 30 images into one big &quot;mega&quot; image, and one feeds those images into the CNN. In this case, should one rotate all 30 images by a random set of degrees, or just a random number of images before glueing them together in the the larger stack? I would think the former, but let me know. Finally, should one rotate and replace, or should one rotate and copy (and hence create more data)? Many thanks in advance! W</p>",
      "rawMarkdown": "Senecaur / others. I am new to image manipulation and still a beginner to CNNs. I am using an implementation where one glues together a stack of 30 images into one big \"mega\" image, and one feeds those images into the CNN. In this case, should one rotate all 30 images by a random set of degrees, or just a random number of images before glueing them together in the the larger stack? I would think the former, but let me know. Finally, should one rotate and replace, or should one rotate and copy (and hence create more data)? Many thanks in advance! W",
      "votes": null
    },
    {
      "id": "108711",
      "postDate": "02/19/2016 09:36:35",
      "content": "<p>This is probably a naive question. I understand that one can scale and rotate images in the 2D plane or within the axes or orientation of a specific image. However, is it possible and does it make sense to try to reorient all images so that they have the same 3D orientation (e.g. the patient orientation cosines as defined in the DICOM image)? I am trying to wrap my head around on whether that is practically and theoretically possible. I played around with it in Mango Viewer, but without much success. Any thoughts appreciated. <a href=\"http://nipy.org/nibabel/dicom/dicom_orientation.html\">http://nipy.org/nibabel/dicom/dicom_orientation.html</a> has some useful links as well </p>",
      "rawMarkdown": "This is probably a naive question. I understand that one can scale and rotate images in the 2D plane or within the axes or orientation of a specific image. However, is it possible and does it make sense to try to reorient all images so that they have the same 3D orientation (e.g. the patient orientation cosines as defined in the DICOM image)? I am trying to wrap my head around on whether that is practically and theoretically possible. I played around with it in Mango Viewer, but without much success. Any thoughts appreciated. http://nipy.org/nibabel/dicom/dicom_orientation.html has some useful links as well",
      "votes": null
    },
    {
      "id": "113737",
      "postDate": "04/04/2016 17:08:33",
      "content": "<p>@Senecaur</p>\n\n<p>I am looking at ways to label the images and I noticed that you recommended Sloth to do so.  Could you go into more detail in how you and your team used Sloth (I've never worked with Sloth or Dicom images before so they are both new to me)?</p>\n\n<p>Did you first convert to another format (jpeg, etc?) as I cannot get Sloth to recognize dicom images.  If so, were there any information loss?  Is it a complete manual process?  </p>\n\n<p>Thank you in advance.</p>",
      "rawMarkdown": "Senecaur\r\n\r\nI am looking at ways to label the images and I noticed that you recommended Sloth to do so.  Could you go into more detail in how you and your team used Sloth (I've never worked with Sloth or Dicom images before so they are both new to me)?\r\n\r\nDid you first convert to another format (jpeg, etc?) as I cannot get Sloth to recognize dicom images.  If so, were there any information loss?  Is it a complete manual process?  \r\n\r\nThank you in advance.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 107053,
      "author_name": "esorar",
      "author_url": "",
      "post_date": "02/06/2016 07:00:36",
      "content": "<p>Nice to have a Titan X, How many iterations typically needs to reach around 0.23? I am using GTX 860M which is like a dwarf next to Titan X and needs close to 350 iterations each taking close to 2.5 minutes to reach around 0.31.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107072,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/06/2016 12:54:26",
      "content": "<p>Each iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107079,
      "author_name": "pengpai",
      "author_url": "",
      "post_date": "02/06/2016 15:24:57",
      "content": "<p>[quote=Senecaur;107072]</p>\n\n<p>Each iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.</p>\n\n<p>[/quote]</p>\n\n<p>@Senecaur Would you please detail the strategy of dealing with overfitting? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107127,
      "author_name": "icedragon",
      "author_url": "",
      "post_date": "02/07/2016 09:53:07",
      "content": "<p>Maybe you should have used something like convex hulls, to get rid of the black points in the circle. ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107133,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/07/2016 14:59:54",
      "content": "<p>@PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107134,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/07/2016 15:00:48",
      "content": "<p>@Icedragon - Great suggestion.  To be clear we are not currently using the segmented images as an input to our model, but we are experimenting with segmentation based models too, so I think that some form of smoothing could be very beneficial.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107213,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/08/2016 13:10:33",
      "content": "<p>@Senecaur. Thanks for posting this. I had two questoins: </p>\n\n<ul>\n<li><p>You mention that the model &quot;predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently&quot;. I was wondering - does the model predict 600 logistic values (all between 0 and 1) which are the subsequently &quot;manually&quot; transformed into a CDF, or does the model directly output a CDF that starts at 0 for the value 1 and ends at 1 for value 600? </p>\n\n<ul><li>you are mentioning that &quot;we are currently choosing to standardize the dimensions of our input by randomly sampling a fixed number of slices and a fixed number of timesteps from each patient study before ingesting it into the CNN&quot;. I am trying to wrap my head around this. When you say slice - do you mean the slice location (e.g. &quot;sax42&quot;) or do you mean an invididual image (e.g. image 0004 at sax42)? I am understanding it correctly when i assume that for each patient you take a random subset of slice locations (e.g. sax 34 and sax 35 for patient 1, sax 64 and sax 42 for patient 2, etc.) and only use a random time-period (e.g. seconds 40 to 80) that is similar for each patient?</li></ul></li>\n</ul>\n\n<p>please let me know! thanks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107215,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/08/2016 13:31:26",
      "content": "<p>@WD</p>\n\n<p>For systolic and diastolic separately we output a 600-dimensional softmax and then apply a cumulative sum operation to get our CDF prediction.  </p>\n\n<p>You are correct, by slice I mean the slice location within a patient study, e.g. sax42, not an individual image.  Each iteration we randomly pick a fixed number of these slices and we pick a fixed number of timesteps from within the slice.  We do not pick these to be similar for each patient, everything is randomized on a per patient per iteration basis.  The fixed numbers of slices and timesteps are quite large though, so actually our sampling is not very sparse at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107221,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/08/2016 14:22:26",
      "content": "<p>@Senecaur. Thanks so much. That makes a lot of sense. I am relatively new to image recognition, and might not know all the tools - especially for preprocessing. What kind of tools / packages did you use for the image scaling and segmentation? I presume that these tools are &quot;outside&quot; lasagne, any links / references would be much appreciated! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107222,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/08/2016 14:27:14",
      "content": "<p>@WD No problem - our image preprocessing and augmentation is all done using the Python library <a href=\"http://scikit-image.org/\">scikit-image</a>, this means it is easy to integrate these processes inline with Lasagne based models.  If you need to manually label/segment any of the training images then I recommend <a href=\"https://cvhci.anthropomatik.kit.edu/~baeuml/projects/a-universal-labeling-tool-for-computer-vision-sloth/\">Sloth</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107261,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/08/2016 21:35:22",
      "content": "<p>@Senecaur - thanks so much. This is very interesting. Are you using 64*64 images? Can you provide further explanation / intuition as well why the random sampling of the slices and the the time-intervals works? I understand that it might provide additonal regularization, but from your description it seemed that the original thought came from the limitations in the data (with regards to slice availability, etc.)? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107322,
      "author_name": "pengpai",
      "author_url": "",
      "post_date": "02/09/2016 05:33:53",
      "content": "<p>[quote=Senecaur;107133]</p>\n\n<p>@PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.</p>\n\n<p>[/quote]\n@Senecaur, I ran into heavy overfitting due to powerful model and limited data.  I used Batch Normalization and Global Averaging Pooling which claimed that Dropout is better not involved.  Thank you for your strategy!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108442,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/17/2016 11:09:18",
      "content": "<p>@senecaur - did you use k-means to smoothen out the images? much appreciated! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108590,
      "author_name": "gzsiceberg",
      "author_url": "",
      "post_date": "02/18/2016 14:36:02",
      "content": "<p>@Senecaur, Your idea about standardizing the dimensions of input is brilliant, but I have a question about your test data. </p>\n\n<p>You mention &quot;predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently&quot;, but your train data for each patient study have fixed number images(randomly picked from different slices and timesteps),  let's assume the fixed number is 60 for each patient study, how do you choose 60 images for each patient in test?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108614,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/18/2016 17:55:08",
      "content": "<p>@WD I tried both 64x64 and 128x128 images.  For the method I described there was no real improvement from using the larger images but it was much slower and more memory hungry to train the model.  You are right that the random sampling of time-steps and slices was more born out of the need to standardize the input tensor than with the plan that it would improve model performance somehow.  I am taking an educated guess that it helps with regularization - I don't really have proof that that is the case!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108615,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/18/2016 17:55:38",
      "content": "<p>@WD Yes, we do use k-means</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108616,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/18/2016 17:57:32",
      "content": "<p>@Gzs_iceberg We choose the slices and time-steps for the test data in the same way - randomly - and then average over multiple samples for the same test patient study.  Test time augmentation has proven in previous competitions (e.g. Galaxy Zoo, DSB1) to be a benefit even when it is not necessary, so it seemed reasonable to do here too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108651,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/18/2016 23:02:24",
      "content": "<p>@Senecaur / others. I am new to image manipulation and still a beginner to CNNs. I am using an implementation where one glues together a stack of 30 images into one big &quot;mega&quot; image, and one feeds those images into the CNN. In this case, should one rotate all 30 images by a random set of degrees, or just a random number of images before glueing them together in the the larger stack? I would think the former, but let me know. Finally, should one rotate and replace, or should one rotate and copy (and hence create more data)? Many thanks in advance! W</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108711,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/19/2016 09:36:35",
      "content": "<p>This is probably a naive question. I understand that one can scale and rotate images in the 2D plane or within the axes or orientation of a specific image. However, is it possible and does it make sense to try to reorient all images so that they have the same 3D orientation (e.g. the patient orientation cosines as defined in the DICOM image)? I am trying to wrap my head around on whether that is practically and theoretically possible. I played around with it in Mango Viewer, but without much success. Any thoughts appreciated. <a href=\"http://nipy.org/nibabel/dicom/dicom_orientation.html\">http://nipy.org/nibabel/dicom/dicom_orientation.html</a> has some useful links as well </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 113737,
      "author_name": "josephclay",
      "author_url": "",
      "post_date": "04/04/2016 17:08:33",
      "content": "<p>@Senecaur</p>\n\n<p>I am looking at ways to label the images and I noticed that you recommended Sloth to do so.  Could you go into more detail in how you and your team used Sloth (I've never worked with Sloth or Dicom images before so they are both new to me)?</p>\n\n<p>Did you first convert to another format (jpeg, etc?) as I cannot get Sloth to recognize dicom images.  If so, were there any information loss?  Is it a complete manual process?  </p>\n\n<p>Thank you in advance.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "107024": "Hi!  We are pleased to share some notes on our current deep learning based approach to the contest.  Whilst we haven't given all the model specifics, we're happy to engage in discussion about our methods and what has and hasn't worked!\r\n\r\nhttp://www.datasciencebowl.com/first_dl_submission/",
    "107053": "Nice to have a Titan X, How many iterations typically needs to reach around 0.23? I am using GTX 860M which is like a dwarf next to Titan X and needs close to 350 iterations each taking close to 2.5 minutes to reach around 0.31.",
    "107072": "Each iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.",
    "107079": "[quote=Senecaur;107072]\r\n\r\nEach iteration takes about a minute to train, 15 seconds to validate and it takes about 100 iterations for our model to train.  So total training time is about 2 hours.  We are using a relatively small batch size, so we could probably further lower this time per iteration, but that may not speed-up the total training time.\r\n\r\n[/quote]\r\n\r\n\r\n@Senecaur Would you please detail the strategy of dealing with overfitting?",
    "107127": "Maybe you should have used something like convex hulls, to get rid of the black points in the circle. ;)",
    "107133": "PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.",
    "107134": "Icedragon - Great suggestion.  To be clear we are not currently using the segmented images as an input to our model, but we are experimenting with segmentation based models too, so I think that some form of smoothing could be very beneficial.",
    "107213": "Senecaur. Thanks for posting this. I had two questoins: \r\n\r\n - You mention that the model \"predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently\". I was wondering - does the model predict 600 logistic values (all between 0 and 1) which are the subsequently \"manually\" transformed into a CDF, or does the model directly output a CDF that starts at 0 for the value 1 and ends at 1 for value 600? \r\n\r\n- you are mentioning that \"we are currently choosing to standardize the dimensions of our input by randomly sampling a fixed number of slices and a fixed number of timesteps from each patient study before ingesting it into the CNN\". I am trying to wrap my head around this. When you say slice - do you mean the slice location (e.g. \"sax42\") or do you mean an invididual image (e.g. image 0004 at sax42)? I am understanding it correctly when i assume that for each patient you take a random subset of slice locations (e.g. sax 34 and sax 35 for patient 1, sax 64 and sax 42 for patient 2, etc.) and only use a random time-period (e.g. seconds 40 to 80) that is similar for each patient?\r\n\r\nplease let me know! thanks!",
    "107215": "WD\r\n\r\nFor systolic and diastolic separately we output a 600-dimensional softmax and then apply a cumulative sum operation to get our CDF prediction.  \r\n\r\nYou are correct, by slice I mean the slice location within a patient study, e.g. sax42, not an individual image.  Each iteration we randomly pick a fixed number of these slices and we pick a fixed number of timesteps from within the slice.  We do not pick these to be similar for each patient, everything is randomized on a per patient per iteration basis.  The fixed numbers of slices and timesteps are quite large though, so actually our sampling is not very sparse at all.",
    "107221": "Senecaur. Thanks so much. That makes a lot of sense. I am relatively new to image recognition, and might not know all the tools - especially for preprocessing. What kind of tools / packages did you use for the image scaling and segmentation? I presume that these tools are \"outside\" lasagne, any links / references would be much appreciated!",
    "107222": "WD No problem - our image preprocessing and augmentation is all done using the Python library [scikit-image][1], this means it is easy to integrate these processes inline with Lasagne based models.  If you need to manually label/segment any of the training images then I recommend [Sloth][2]\r\n\r\n  [1]: http://scikit-image.org/\r\n  [2]: https://cvhci.anthropomatik.kit.edu/~baeuml/projects/a-universal-labeling-tool-for-computer-vision-sloth/",
    "107261": "Senecaur - thanks so much. This is very interesting. Are you using 64*64 images? Can you provide further explanation / intuition as well why the random sampling of the slices and the the time-intervals works? I understand that it might provide additonal regularization, but from your description it seemed that the original thought came from the limitations in the data (with regards to slice availability, etc.)?",
    "107322": "[quote=Senecaur;107133]\r\n\r\n@PengPai Sure - we use dropout on the inputs to the softmax output layer, we use L2 regularization on the trainable parameters and we use online data augmentation (random rotations, brightness adjustments and translations).  We also randomize the sampling and ordering of the timesteps within a slice and the sampling and ordering of slices per study each iteration - this has a regularization effect too.\r\n\r\n[/quote]\r\n@Senecaur, I ran into heavy overfitting due to powerful model and limited data.  I used Batch Normalization and Global Averaging Pooling which claimed that Dropout is better not involved.  Thank you for your strategy!",
    "108442": "senecaur - did you use k-means to smoothen out the images? much appreciated!",
    "108590": "Senecaur, Your idea about standardizing the dimensions of input is brilliant, but I have a question about your test data. \r\n\r\nYou mention \"predicts 600-dimensional systolic and diastolic CDFs for each timestep and each slice in the input tensor independently\", but your train data for each patient study have fixed number images(randomly picked from different slices and timesteps),  let's assume the fixed number is 60 for each patient study, how do you choose 60 images for each patient in test?",
    "108614": "WD I tried both 64x64 and 128x128 images.  For the method I described there was no real improvement from using the larger images but it was much slower and more memory hungry to train the model.  You are right that the random sampling of time-steps and slices was more born out of the need to standardize the input tensor than with the plan that it would improve model performance somehow.  I am taking an educated guess that it helps with regularization - I don't really have proof that that is the case!",
    "108615": "WD Yes, we do use k-means",
    "108616": "Gzs_iceberg We choose the slices and time-steps for the test data in the same way - randomly - and then average over multiple samples for the same test patient study.  Test time augmentation has proven in previous competitions (e.g. Galaxy Zoo, DSB1) to be a benefit even when it is not necessary, so it seemed reasonable to do here too.",
    "108651": "Senecaur / others. I am new to image manipulation and still a beginner to CNNs. I am using an implementation where one glues together a stack of 30 images into one big \"mega\" image, and one feeds those images into the CNN. In this case, should one rotate all 30 images by a random set of degrees, or just a random number of images before glueing them together in the the larger stack? I would think the former, but let me know. Finally, should one rotate and replace, or should one rotate and copy (and hence create more data)? Many thanks in advance! W",
    "108711": "This is probably a naive question. I understand that one can scale and rotate images in the 2D plane or within the axes or orientation of a specific image. However, is it possible and does it make sense to try to reorient all images so that they have the same 3D orientation (e.g. the patient orientation cosines as defined in the DICOM image)? I am trying to wrap my head around on whether that is practically and theoretically possible. I played around with it in Mango Viewer, but without much success. Any thoughts appreciated. http://nipy.org/nibabel/dicom/dicom_orientation.html has some useful links as well",
    "113737": "Senecaur\r\n\r\nI am looking at ways to label the images and I noticed that you recommended Sloth to do so.  Could you go into more detail in how you and your team used Sloth (I've never worked with Sloth or Dicom images before so they are both new to me)?\r\n\r\nDid you first convert to another format (jpeg, etc?) as I cannot get Sloth to recognize dicom images.  If so, were there any information loss?  Is it a complete manual process?  \r\n\r\nThank you in advance."
  },
  "source": "meta"
}