{
  "id": 19013,
  "title": "Using age and gender data",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19013",
  "author_name": "",
  "post_date": "2016-02-16T22:34:05.437Z",
  "votes": null,
  "comment_count": 7,
  "views": 1113,
  "content": "<p>i was thinking to use gender and age data to furhter improve the outcomes of the CNN. I was thinking of two approaches, but would love the expertise of those with more experience with stacking. </p>\n\n<p>In my first approach, i would build a second regressoin model only using age and gender information to make sys and dias predictions. I would train this on the entire dataset. Subsequently, i would combine the outcomes of the two models (e.g. again with linear regression). </p>\n\n<p>Additionally, i was thinking of training a linear regression model on the residual between the CNN-predicted systole and diastole values, and then add the prediction of this regression model to the CNN prediction. </p>\n\n<p>These approaches might be identical, and/or might be both be invalid or not smart. I would love to get thougths and guidance of those with more experience </p>",
  "messages": [
    {
      "id": "108371",
      "postDate": "02/16/2016 22:34:05",
      "content": "<p>i was thinking to use gender and age data to furhter improve the outcomes of the CNN. I was thinking of two approaches, but would love the expertise of those with more experience with stacking. </p>\n\n<p>In my first approach, i would build a second regressoin model only using age and gender information to make sys and dias predictions. I would train this on the entire dataset. Subsequently, i would combine the outcomes of the two models (e.g. again with linear regression). </p>\n\n<p>Additionally, i was thinking of training a linear regression model on the residual between the CNN-predicted systole and diastole values, and then add the prediction of this regression model to the CNN prediction. </p>\n\n<p>These approaches might be identical, and/or might be both be invalid or not smart. I would love to get thougths and guidance of those with more experience </p>",
      "rawMarkdown": "i was thinking to use gender and age data to furhter improve the outcomes of the CNN. I was thinking of two approaches, but would love the expertise of those with more experience with stacking. \r\n\r\nIn my first approach, i would build a second regressoin model only using age and gender information to make sys and dias predictions. I would train this on the entire dataset. Subsequently, i would combine the outcomes of the two models (e.g. again with linear regression). \r\n\r\nAdditionally, i was thinking of training a linear regression model on the residual between the CNN-predicted systole and diastole values, and then add the prediction of this regression model to the CNN prediction. \r\n\r\nThese approaches might be identical, and/or might be both be invalid or not smart. I would love to get thougths and guidance of those with more experience",
      "votes": null
    },
    {
      "id": "108397",
      "postDate": "02/17/2016 01:47:36",
      "content": "<p>I have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.</p>\n\n<p>Also I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.</p>",
      "rawMarkdown": "I have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.\r\n\r\nAlso I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.",
      "votes": null
    },
    {
      "id": "108399",
      "postDate": "02/17/2016 02:02:22",
      "content": "<p>I am part of the Booze-Allen-Hamilton and NVIDIA team.  I wrote a short blog post earlier this month describing a neural network approach that scored 0.022886 on the leaderboard.  You can read the post <a href=\"http://www.datasciencebowl.com/first_dl_submission/\">here</a>.  It's essentially an extension of the mxnet and keras tutorials in that we are applying a convolutional neural network to the individual slices and ignoring their spatial relationships (and temporal ordering for that matter). </p>",
      "rawMarkdown": "I am part of the Booze-Allen-Hamilton and NVIDIA team.  I wrote a short blog post earlier this month describing a neural network approach that scored 0.022886 on the leaderboard.  You can read the post [here][1].  It's essentially an extension of the mxnet and keras tutorials in that we are applying a convolutional neural network to the individual slices and ignoring their spatial relationships (and temporal ordering for that matter). \r\n\r\n  [1]: http://www.datasciencebowl.com/first_dl_submission/",
      "votes": null
    },
    {
      "id": "108520",
      "postDate": "02/18/2016 02:08:36",
      "content": "<p>I made it to about 0.28 by using the mxnet tutorial and adjusting the cropping and resizing bits before I hit my laptop GPU's ceiling.  If I was on the NVIDIA team, I would feed the new network I put together into a bright shiny Titan X and go to a larger scaled image size (e.g. 96x96).  It baked on my laptop from Friday evening to Sunday midnight before it memory dumped.  This is my first foray with convnets and it has me itching for new hardware.</p>",
      "rawMarkdown": "I made it to about 0.28 by using the mxnet tutorial and adjusting the cropping and resizing bits before I hit my laptop GPU's ceiling.  If I was on the NVIDIA team, I would feed the new network I put together into a bright shiny Titan X and go to a larger scaled image size (e.g. 96x96).  It baked on my laptop from Friday evening to Sunday midnight before it memory dumped.  This is my first foray with convnets and it has me itching for new hardware.",
      "votes": null
    },
    {
      "id": "108582",
      "postDate": "02/18/2016 13:02:32",
      "content": "<p>[quote=EIGSI;108397]</p>\n\n<p>I have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.</p>\n\n<p>Also I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.</p>\n\n<p>[/quote]</p>\n\n<p>You are definitely right. There are so many potential approaches for this match, and for people who have no prior knowledge of the domain like us, it's really hard to find the right approach.  Maybe we should read some relevant papers at first. </p>",
      "rawMarkdown": "[quote=EIGSI;108397]\r\n\r\nI have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.\r\n\r\nAlso I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.\r\n\r\n\r\n[/quote]\r\n\r\nYou are definitely right. There are so many potential approaches for this match, and for people who have no prior knowledge of the domain like us, it's really hard to find the right approach.  Maybe we should read some relevant papers at first.",
      "votes": null
    },
    {
      "id": "108596",
      "postDate": "02/18/2016 15:15:53",
      "content": "<p>@Jiming Ye I think the fourier tutorial is good for understanding of the domain. I ran it with python3 and it estimates volume for directory 1 as 3000 something so it quits after that point due to the assertion that the volume can be 600 max. Must be something to do with python version, it was wriiten for python 2. If anyone figured out how to make this work it would be great</p>",
      "rawMarkdown": "Jiming Ye I think the fourier tutorial is good for understanding of the domain. I ran it with python3 and it estimates volume for directory 1 as 3000 something so it quits after that point due to the assertion that the volume can be 600 max. Must be something to do with python version, it was wriiten for python 2. If anyone figured out how to make this work it would be great",
      "votes": null
    },
    {
      "id": "108597",
      "postDate": "02/18/2016 15:16:07",
      "content": "<p>@senecaur yes I have seen your tutorial, your input data is 4D which is something I wanted to try as well. I think the mxnet ignores those slices with less than 30 frames and that still leaves over 5000 training slices, I wonder why you had to resample temporally as well</p>",
      "rawMarkdown": "@senecaur yes I have seen your tutorial, your input data is 4D which is something I wanted to try as well. I think the mxnet ignores those slices with less than 30 frames and that still leaves over 5000 training slices, I wonder why you had to resample temporally as well",
      "votes": null
    },
    {
      "id": "108617",
      "postDate": "02/18/2016 17:59:11",
      "content": "<p>@EIGSI  We <em>chose</em> to sample so that we could make use of all slices, even if they have &lt;30 time-steps; however, our results seem to suggest that performance of the model doesn't dramatically drop if you use a smaller sample of time-steps per slice but your memory usage and training time go down.</p>",
      "rawMarkdown": "EIGSI  We *chose* to sample so that we could make use of all slices, even if they have <30 time-steps; however, our results seem to suggest that performance of the model doesn't dramatically drop if you use a smaller sample of time-steps per slice but your memory usage and training time go down.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 108397,
      "author_name": "",
      "author_url": "",
      "post_date": "02/17/2016 01:47:36",
      "content": "<p>I have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.</p>\n\n<p>Also I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108399,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/17/2016 02:02:22",
      "content": "<p>I am part of the Booze-Allen-Hamilton and NVIDIA team.  I wrote a short blog post earlier this month describing a neural network approach that scored 0.022886 on the leaderboard.  You can read the post <a href=\"http://www.datasciencebowl.com/first_dl_submission/\">here</a>.  It's essentially an extension of the mxnet and keras tutorials in that we are applying a convolutional neural network to the individual slices and ignoring their spatial relationships (and temporal ordering for that matter). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108520,
      "author_name": "scsmith2",
      "author_url": "",
      "post_date": "02/18/2016 02:08:36",
      "content": "<p>I made it to about 0.28 by using the mxnet tutorial and adjusting the cropping and resizing bits before I hit my laptop GPU's ceiling.  If I was on the NVIDIA team, I would feed the new network I put together into a bright shiny Titan X and go to a larger scaled image size (e.g. 96x96).  It baked on my laptop from Friday evening to Sunday midnight before it memory dumped.  This is my first foray with convnets and it has me itching for new hardware.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108582,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "02/18/2016 13:02:32",
      "content": "<p>[quote=EIGSI;108397]</p>\n\n<p>I have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.</p>\n\n<p>Also I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.</p>\n\n<p>[/quote]</p>\n\n<p>You are definitely right. There are so many potential approaches for this match, and for people who have no prior knowledge of the domain like us, it's really hard to find the right approach.  Maybe we should read some relevant papers at first. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108596,
      "author_name": "",
      "author_url": "",
      "post_date": "02/18/2016 15:15:53",
      "content": "<p>@Jiming Ye I think the fourier tutorial is good for understanding of the domain. I ran it with python3 and it estimates volume for directory 1 as 3000 something so it quits after that point due to the assertion that the volume can be 600 max. Must be something to do with python version, it was wriiten for python 2. If anyone figured out how to make this work it would be great</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108597,
      "author_name": "",
      "author_url": "",
      "post_date": "02/18/2016 15:16:07",
      "content": "<p>@senecaur yes I have seen your tutorial, your input data is 4D which is something I wanted to try as well. I think the mxnet ignores those slices with less than 30 frames and that still leaves over 5000 training slices, I wonder why you had to resample temporally as well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108617,
      "author_name": "senecaur",
      "author_url": "",
      "post_date": "02/18/2016 17:59:11",
      "content": "<p>@EIGSI  We <em>chose</em> to sample so that we could make use of all slices, even if they have &lt;30 time-steps; however, our results seem to suggest that performance of the model doesn't dramatically drop if you use a smaller sample of time-steps per slice but your memory usage and training time go down.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "108371": "i was thinking to use gender and age data to furhter improve the outcomes of the CNN. I was thinking of two approaches, but would love the expertise of those with more experience with stacking. \r\n\r\nIn my first approach, i would build a second regressoin model only using age and gender information to make sys and dias predictions. I would train this on the entire dataset. Subsequently, i would combine the outcomes of the two models (e.g. again with linear regression). \r\n\r\nAdditionally, i was thinking of training a linear regression model on the residual between the CNN-predicted systole and diastole values, and then add the prediction of this regression model to the CNN prediction. \r\n\r\nThese approaches might be identical, and/or might be both be invalid or not smart. I would love to get thougths and guidance of those with more experience",
    "108397": "I have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.\r\n\r\nAlso I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.",
    "108399": "I am part of the Booze-Allen-Hamilton and NVIDIA team.  I wrote a short blog post earlier this month describing a neural network approach that scored 0.022886 on the leaderboard.  You can read the post [here][1].  It's essentially an extension of the mxnet and keras tutorials in that we are applying a convolutional neural network to the individual slices and ignoring their spatial relationships (and temporal ordering for that matter). \r\n\r\n  [1]: http://www.datasciencebowl.com/first_dl_submission/",
    "108520": "I made it to about 0.28 by using the mxnet tutorial and adjusting the cropping and resizing bits before I hit my laptop GPU's ceiling.  If I was on the NVIDIA team, I would feed the new network I put together into a bright shiny Titan X and go to a larger scaled image size (e.g. 96x96).  It baked on my laptop from Friday evening to Sunday midnight before it memory dumped.  This is my first foray with convnets and it has me itching for new hardware.",
    "108582": "[quote=EIGSI;108397]\r\n\r\nI have tried both mxnet and keras tutorials as a learning experience as I have no prior knowledge in this problem domain and limited programming experience on deep learning. However I have come to realize that the tutorials were a bit misleading. The heart volume is made up of the multiple slices yet these tutorials ignore spatial relationships between the slices as well as slice thickness/distance etc.  I saw no point in messing with cnn parameters etc to lower crps. I wonder if anybody was able to tweak these tutorials and get crps around 0.02? It would be surprising to me.\r\n\r\nAlso I am wondering if expert manual segmentations would ever be made available on this dataset. I would be more interested in training neural nets to learn the actual segmentations. Computing volumes after that would be much easier.\r\n\r\n\r\n[/quote]\r\n\r\nYou are definitely right. There are so many potential approaches for this match, and for people who have no prior knowledge of the domain like us, it's really hard to find the right approach.  Maybe we should read some relevant papers at first.",
    "108596": "Jiming Ye I think the fourier tutorial is good for understanding of the domain. I ran it with python3 and it estimates volume for directory 1 as 3000 something so it quits after that point due to the assertion that the volume can be 600 max. Must be something to do with python version, it was wriiten for python 2. If anyone figured out how to make this work it would be great",
    "108597": "@senecaur yes I have seen your tutorial, your input data is 4D which is something I wanted to try as well. I think the mxnet ignores those slices with less than 30 frames and that still leaves over 5000 training slices, I wonder why you had to resample temporally as well",
    "108617": "EIGSI  We *chose* to sample so that we could make use of all slices, even if they have <30 time-steps; however, our results seem to suggest that performance of the model doesn't dramatically drop if you use a smaller sample of time-steps per slice but your memory usage and training time go down."
  },
  "source": "meta"
}