{
  "id": 18375,
  "title": "0.036023 score without looking at the images",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18375",
  "author_name": "",
  "post_date": "2016-01-13T08:37:04.247Z",
  "votes": 31,
  "comment_count": 15,
  "views": 4380,
  "content": "<p>Before starting with image processing, I've built the basic infrastructure for my C++ application: DICOM and CSV file I/O, submission generation and evaluation, machine learning pipeline, etc. Having it all working, I wanted to generate a few preliminary results staying close to the physical reality and far away from deep learning.</p>\n\n<p>I noticed that patient data contains age and sex information in DICOM tags. Common sense supported with a quick web search confirmed that female hearts are smaller than male hearts and normal heart growth stops at adulthood. I created a model for LV volume based on piecewise-linear functions of patient's age, one for each of {male, female}x{systole, diastole} domains. Each function has a positive slope for age arguments in child range and zero slope at median value for arguments in adult range. Here are the parameters of the best scoring model I found with supervised training using a simple brute force minimum value search:</p>\n\n<ul>\n<li>female child systole:        2.41x + 15</li>\n<li>female child diastole:       7.61x + 22</li>\n<li>female adult median systole:     53.6</li>\n<li>female adult median diastole:    144</li>\n<li>female adult age:        16</li>\n<li>male child systole:      4.69x + 0</li>\n<li>male child diastole:         10.8x + 9</li>\n<li>male adult median systole:   75</li>\n<li>male adult median diastole:  181</li>\n<li>male adult age:      16</li>\n</ul>\n\n<p>The next step was generating the CDF (cumulative distribution function). Simple step function evaluates to 0.050397 on the training set. Replacing it with piecewise linear CDF (horizontal-slope-horizontal) with a slope of 0.01 improves the result significantly to 0.038542. The best scoring CDF uses approximation of <a href=\"http://mathworld.wolfram.com/NormalDistributionFunction.html\">Normal Distribution Function</a> given by equation (13) on that page, but the gain is very tiny with final result of 0.036023.</p>",
  "messages": [
    {
      "id": "104492",
      "postDate": "01/13/2016 08:37:04",
      "content": "<p>Before starting with image processing, I've built the basic infrastructure for my C++ application: DICOM and CSV file I/O, submission generation and evaluation, machine learning pipeline, etc. Having it all working, I wanted to generate a few preliminary results staying close to the physical reality and far away from deep learning.</p>\n\n<p>I noticed that patient data contains age and sex information in DICOM tags. Common sense supported with a quick web search confirmed that female hearts are smaller than male hearts and normal heart growth stops at adulthood. I created a model for LV volume based on piecewise-linear functions of patient's age, one for each of {male, female}x{systole, diastole} domains. Each function has a positive slope for age arguments in child range and zero slope at median value for arguments in adult range. Here are the parameters of the best scoring model I found with supervised training using a simple brute force minimum value search:</p>\n\n<ul>\n<li>female child systole:        2.41x + 15</li>\n<li>female child diastole:       7.61x + 22</li>\n<li>female adult median systole:     53.6</li>\n<li>female adult median diastole:    144</li>\n<li>female adult age:        16</li>\n<li>male child systole:      4.69x + 0</li>\n<li>male child diastole:         10.8x + 9</li>\n<li>male adult median systole:   75</li>\n<li>male adult median diastole:  181</li>\n<li>male adult age:      16</li>\n</ul>\n\n<p>The next step was generating the CDF (cumulative distribution function). Simple step function evaluates to 0.050397 on the training set. Replacing it with piecewise linear CDF (horizontal-slope-horizontal) with a slope of 0.01 improves the result significantly to 0.038542. The best scoring CDF uses approximation of <a href=\"http://mathworld.wolfram.com/NormalDistributionFunction.html\">Normal Distribution Function</a> given by equation (13) on that page, but the gain is very tiny with final result of 0.036023.</p>",
      "rawMarkdown": "Before starting with image processing, I've built the basic infrastructure for my C++ application: DICOM and CSV file I/O, submission generation and evaluation, machine learning pipeline, etc. Having it all working, I wanted to generate a few preliminary results staying close to the physical reality and far away from deep learning.\r\n\r\nI noticed that patient data contains age and sex information in DICOM tags. Common sense supported with a quick web search confirmed that female hearts are smaller than male hearts and normal heart growth stops at adulthood. I created a model for LV volume based on piecewise-linear functions of patient's age, one for each of {male, female}x{systole, diastole} domains. Each function has a positive slope for age arguments in child range and zero slope at median value for arguments in adult range. Here are the parameters of the best scoring model I found with supervised training using a simple brute force minimum value search:\r\n\r\n - female child systole: \t\t2.41x + 15\r\n - female child diastole: \t\t7.61x + 22\r\n - female adult median systole: \t53.6\r\n - female adult median diastole: \t144\r\n - female adult age: \t\t16\r\n - male child systole: \t\t4.69x + 0\r\n - male child diastole: \t\t10.8x + 9\r\n - male adult median systole: \t75\r\n - male adult median diastole: \t181\r\n - male adult age: \t\t16\r\n\r\nThe next step was generating the CDF (cumulative distribution function). Simple step function evaluates to 0.050397 on the training set. Replacing it with piecewise linear CDF (horizontal-slope-horizontal) with a slope of 0.01 improves the result significantly to 0.038542. The best scoring CDF uses approximation of [Normal Distribution Function][1] given by equation (13) on that page, but the gain is very tiny with final result of 0.036023.\r\n\r\n\r\n  [1]: http://mathworld.wolfram.com/NormalDistributionFunction.html",
      "votes": null
    },
    {
      "id": "104510",
      "postDate": "01/13/2016 13:39:49",
      "content": "<p>Just a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation. </p>",
      "rawMarkdown": "Just a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation.",
      "votes": null
    },
    {
      "id": "104554",
      "postDate": "01/13/2016 19:34:54",
      "content": "<p>Maybe not too surprising but I still think it is a nice result, 'proving' that it is probably valuable to include this metadata in the learning process.</p>",
      "rawMarkdown": "Maybe not too surprising but I still think it is a nice result, 'proving' that it is probably valuable to include this metadata in the learning process.",
      "votes": null
    },
    {
      "id": "104668",
      "postDate": "01/15/2016 05:01:32",
      "content": "<p>[quote=Michael Hansen;104510]</p>\n\n<p>Just a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation. </p>\n\n<p>[/quote]</p>\n\n<p>If  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.</p>",
      "rawMarkdown": "[quote=Michael Hansen;104510]\r\n\r\nJust a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation. \r\n\r\n[/quote]\r\n\r\nIf  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.",
      "votes": null
    },
    {
      "id": "104695",
      "postDate": "01/15/2016 13:07:36",
      "content": "<p>[quote=Torgos;104668]</p>\n\n<p>If  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.</p>\n\n<p>[/quote]</p>\n\n<p>Using the metadata you have is and will be allowed. If age and gender truly supplement the best image-based models, then they deserve to be part of the algorithm! After all, they're available to the clinician at the MRI workstation, so they're not leakage and are fair game for the computers.</p>",
      "rawMarkdown": "[quote=Torgos;104668]\r\n\r\nIf  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.\r\n\r\n[/quote]\r\n\r\nUsing the metadata you have is and will be allowed. If age and gender truly supplement the best image-based models, then they deserve to be part of the algorithm! After all, they're available to the clinician at the MRI workstation, so they're not leakage and are fair game for the computers.",
      "votes": null
    },
    {
      "id": "107666",
      "postDate": "02/11/2016 23:32:50",
      "content": "<p>I just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.</p>",
      "rawMarkdown": "I just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.",
      "votes": null
    },
    {
      "id": "107988",
      "postDate": "02/14/2016 16:33:13",
      "content": "<p>Paul, how did you identify which age is expressed in months/weeks rather than years?</p>\n\n<p>[quote=Paul Jurczak;107666]</p>\n\n<p>I just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Paul, how did you identify which age is expressed in months/weeks rather than years?\r\n\r\n[quote=Paul Jurczak;107666]\r\n\r\nI just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "108027",
      "postDate": "02/14/2016 22:12:21",
      "content": "<p>The age tag is a 4 character string: 3 digits followed by 1 letter (Y, M or W). </p>",
      "rawMarkdown": "The age tag is a 4 character string: 3 digits followed by 1 letter (Y, M or W).",
      "votes": null
    },
    {
      "id": "108028",
      "postDate": "02/14/2016 23:12:22",
      "content": "<p>Ah, true, you are right, I just found that as well. Thanks for posting!</p>",
      "rawMarkdown": "Ah, true, you are right, I just found that as well. Thanks for posting!",
      "votes": null
    },
    {
      "id": "108270",
      "postDate": "02/16/2016 14:50:26",
      "content": "<p>@Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W</p>",
      "rawMarkdown": "Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W",
      "votes": null
    },
    {
      "id": "108374",
      "postDate": "02/16/2016 23:26:01",
      "content": "<p>[quote=WD;108270]</p>\n\n<p>@Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W</p>\n\n<p>[/quote]</p>\n\n<p>It is just a brute force search coded in C++ with <a href=\"https://imebra.com/\">Imebra</a> dependency for reading DICOM. If you are still interested, I can put it on github.</p>",
      "rawMarkdown": "[quote=WD;108270]\r\n\r\n@Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W\r\n\r\n[/quote]\r\n\r\nIt is just a brute force search coded in C++ with [Imebra][1] dependency for reading DICOM. If you are still interested, I can put it on github.\r\n\r\n\r\n  [1]: https://imebra.com/",
      "votes": null
    },
    {
      "id": "108801",
      "postDate": "02/20/2016 00:00:02",
      "content": "<p>@Paul Jurczak definitely interested, please put it on github</p>",
      "rawMarkdown": "Paul Jurczak definitely interested, please put it on github",
      "votes": null
    },
    {
      "id": "108818",
      "postDate": "02/20/2016 07:11:44",
      "content": "<p>@phunter</p>\n\n<p>I posted it to <a href=\"https://github.com/pauljurczak/Second-Annual-Data-Science-Bowl\">https://github.com/pauljurczak/Second-Annual-Data-Science-Bowl</a>. Original project directory structure is not preserved, because I have only a faint idea how to use github.</p>",
      "rawMarkdown": "phunter\r\n\r\nI posted it to https://github.com/pauljurczak/Second-Annual-Data-Science-Bowl. Original project directory structure is not preserved, because I have only a faint idea how to use github.",
      "votes": null
    },
    {
      "id": "108827",
      "postDate": "02/20/2016 08:27:50",
      "content": "<p>@Paul, @phunter. I am debating and would love your thoughts on the best way to stack models. If one uses a model that has a 600-datapoint &quot;CDF&quot; as output, is the only way to stack this model with a linear regression estimate following Paul's approach to a) transform the CNN output to a point estimate (e.g. the mean of the PDF), then do some form of stacking (e.g. linear regression) on top of both Paul's model and the CNN model, and then subsequently transform the prediction with a sigma back into a CDF? I tried this approach, and performance was not great, maybe because there is so much risk of noise in the transformation from CDF to PDF and back again. Is there any other elegant approach that I am missing? Is there  a way to directly take the 600-datapoint CDF as input in a stacked model? </p>",
      "rawMarkdown": "Paul, @phunter. I am debating and would love your thoughts on the best way to stack models. If one uses a model that has a 600-datapoint \"CDF\" as output, is the only way to stack this model with a linear regression estimate following Paul's approach to a) transform the CNN output to a point estimate (e.g. the mean of the PDF), then do some form of stacking (e.g. linear regression) on top of both Paul's model and the CNN model, and then subsequently transform the prediction with a sigma back into a CDF? I tried this approach, and performance was not great, maybe because there is so much risk of noise in the transformation from CDF to PDF and back again. Is there any other elegant approach that I am missing? Is there  a way to directly take the 600-datapoint CDF as input in a stacked model?",
      "votes": null
    },
    {
      "id": "108828",
      "postDate": "02/20/2016 09:10:52",
      "content": "<p>@WD</p>\n\n<p>I have only a faint understanding of CNN and I'm hoping to get a good score without using them. This model was only an exercise, but I may use it as a sanity check for the next model based on image processing and as a fallback option for studies where image processing fails.</p>\n\n<p>One idea for you to try is to look for the extreme outliers (relative to my model) generated by CNN and replace them by my model's prediction. You may consider augmenting the model with lower and upper volume range to filter out really improbable CNN predictions.</p>",
      "rawMarkdown": "WD\r\n\r\nI have only a faint understanding of CNN and I'm hoping to get a good score without using them. This model was only an exercise, but I may use it as a sanity check for the next model based on image processing and as a fallback option for studies where image processing fails.\r\n\r\nOne idea for you to try is to look for the extreme outliers (relative to my model) generated by CNN and replace them by my model's prediction. You may consider augmenting the model with lower and upper volume range to filter out really improbable CNN predictions.",
      "votes": null
    },
    {
      "id": "110301",
      "postDate": "03/04/2016 14:15:27",
      "content": "<p>This post for some reason made me picture Michael Jordan taking a foul shot with his eyes closed :-)</p>",
      "rawMarkdown": "This post for some reason made me picture Michael Jordan taking a foul shot with his eyes closed :-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 104510,
      "author_name": "michaelhansen",
      "author_url": "",
      "post_date": "01/13/2016 13:39:49",
      "content": "<p>Just a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104554,
      "author_name": "diwiwo",
      "author_url": "",
      "post_date": "01/13/2016 19:34:54",
      "content": "<p>Maybe not too surprising but I still think it is a nice result, 'proving' that it is probably valuable to include this metadata in the learning process.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104668,
      "author_name": "telser",
      "author_url": "",
      "post_date": "01/15/2016 05:01:32",
      "content": "<p>[quote=Michael Hansen;104510]</p>\n\n<p>Just a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation. </p>\n\n<p>[/quote]</p>\n\n<p>If  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104695,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "01/15/2016 13:07:36",
      "content": "<p>[quote=Torgos;104668]</p>\n\n<p>If  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.</p>\n\n<p>[/quote]</p>\n\n<p>Using the metadata you have is and will be allowed. If age and gender truly supplement the best image-based models, then they deserve to be part of the algorithm! After all, they're available to the clinician at the MRI workstation, so they're not leakage and are fair game for the computers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107666,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/11/2016 23:32:50",
      "content": "<p>I just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107988,
      "author_name": "kosinski",
      "author_url": "",
      "post_date": "02/14/2016 16:33:13",
      "content": "<p>Paul, how did you identify which age is expressed in months/weeks rather than years?</p>\n\n<p>[quote=Paul Jurczak;107666]</p>\n\n<p>I just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108027,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/14/2016 22:12:21",
      "content": "<p>The age tag is a 4 character string: 3 digits followed by 1 letter (Y, M or W). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108028,
      "author_name": "kosinski",
      "author_url": "",
      "post_date": "02/14/2016 23:12:22",
      "content": "<p>Ah, true, you are right, I just found that as well. Thanks for posting!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108270,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/16/2016 14:50:26",
      "content": "<p>@Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108374,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/16/2016 23:26:01",
      "content": "<p>[quote=WD;108270]</p>\n\n<p>@Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W</p>\n\n<p>[/quote]</p>\n\n<p>It is just a brute force search coded in C++ with <a href=\"https://imebra.com/\">Imebra</a> dependency for reading DICOM. If you are still interested, I can put it on github.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108801,
      "author_name": "phunter",
      "author_url": "",
      "post_date": "02/20/2016 00:00:02",
      "content": "<p>@Paul Jurczak definitely interested, please put it on github</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108818,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/20/2016 07:11:44",
      "content": "<p>@phunter</p>\n\n<p>I posted it to <a href=\"https://github.com/pauljurczak/Second-Annual-Data-Science-Bowl\">https://github.com/pauljurczak/Second-Annual-Data-Science-Bowl</a>. Original project directory structure is not preserved, because I have only a faint idea how to use github.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108827,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/20/2016 08:27:50",
      "content": "<p>@Paul, @phunter. I am debating and would love your thoughts on the best way to stack models. If one uses a model that has a 600-datapoint &quot;CDF&quot; as output, is the only way to stack this model with a linear regression estimate following Paul's approach to a) transform the CNN output to a point estimate (e.g. the mean of the PDF), then do some form of stacking (e.g. linear regression) on top of both Paul's model and the CNN model, and then subsequently transform the prediction with a sigma back into a CDF? I tried this approach, and performance was not great, maybe because there is so much risk of noise in the transformation from CDF to PDF and back again. Is there any other elegant approach that I am missing? Is there  a way to directly take the 600-datapoint CDF as input in a stacked model? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108828,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/20/2016 09:10:52",
      "content": "<p>@WD</p>\n\n<p>I have only a faint understanding of CNN and I'm hoping to get a good score without using them. This model was only an exercise, but I may use it as a sanity check for the next model based on image processing and as a fallback option for studies where image processing fails.</p>\n\n<p>One idea for you to try is to look for the extreme outliers (relative to my model) generated by CNN and replace them by my model's prediction. You may consider augmenting the model with lower and upper volume range to filter out really improbable CNN predictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110301,
      "author_name": "neuralnetworks",
      "author_url": "",
      "post_date": "03/04/2016 14:15:27",
      "content": "<p>This post for some reason made me picture Michael Jordan taking a foul shot with his eyes closed :-)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "104492": "Before starting with image processing, I've built the basic infrastructure for my C++ application: DICOM and CSV file I/O, submission generation and evaluation, machine learning pipeline, etc. Having it all working, I wanted to generate a few preliminary results staying close to the physical reality and far away from deep learning.\r\n\r\nI noticed that patient data contains age and sex information in DICOM tags. Common sense supported with a quick web search confirmed that female hearts are smaller than male hearts and normal heart growth stops at adulthood. I created a model for LV volume based on piecewise-linear functions of patient's age, one for each of {male, female}x{systole, diastole} domains. Each function has a positive slope for age arguments in child range and zero slope at median value for arguments in adult range. Here are the parameters of the best scoring model I found with supervised training using a simple brute force minimum value search:\r\n\r\n - female child systole: \t\t2.41x + 15\r\n - female child diastole: \t\t7.61x + 22\r\n - female adult median systole: \t53.6\r\n - female adult median diastole: \t144\r\n - female adult age: \t\t16\r\n - male child systole: \t\t4.69x + 0\r\n - male child diastole: \t\t10.8x + 9\r\n - male adult median systole: \t75\r\n - male adult median diastole: \t181\r\n - male adult age: \t\t16\r\n\r\nThe next step was generating the CDF (cumulative distribution function). Simple step function evaluates to 0.050397 on the training set. Replacing it with piecewise linear CDF (horizontal-slope-horizontal) with a slope of 0.01 improves the result significantly to 0.038542. The best scoring CDF uses approximation of [Normal Distribution Function][1] given by equation (13) on that page, but the gain is very tiny with final result of 0.036023.\r\n\r\n\r\n  [1]: http://mathworld.wolfram.com/NormalDistributionFunction.html",
    "104510": "Just a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation.",
    "104554": "Maybe not too surprising but I still think it is a nice result, 'proving' that it is probably valuable to include this metadata in the learning process.",
    "104668": "[quote=Michael Hansen;104510]\r\n\r\nJust a quick comment. It is not too surprising that age and gender are correlated with volume since they would be correlated with patient size. In fact, this is used clinically to evaluate whether a given volume is abnormal given the patient size. But be careful about relying too much on age and gender in the models, since disease cases would go against exactly that correlation. \r\n\r\n[/quote]\r\n\r\nIf  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.",
    "104695": "[quote=Torgos;104668]\r\n\r\nIf  the usage of metadata is helpful in more accurately predicting volume in most cases but defeats the purpose of this exercise (which I interpreted to be the extraction of that information from the images) in a way such as to weaken its clinical use, perhaps you should not allow it for a final solution?  It's still relatively early in the competition, but any such rule change should be done asap.\r\n\r\n[/quote]\r\n\r\nUsing the metadata you have is and will be allowed. If age and gender truly supplement the best image-based models, then they deserve to be part of the algorithm! After all, they're available to the clinician at the MRI workstation, so they're not leakage and are fair game for the computers.",
    "107666": "I just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.",
    "107988": "Paul, how did you identify which age is expressed in months/weeks rather than years?\r\n\r\n[quote=Paul Jurczak;107666]\r\n\r\nI just noticed that patient's age is sometimes expressed in months or weeks instead of years. My report was produced under a wrong assumption that patient's age is always expressed in years. Correcting this error improves the score to 0.034560. Incidentally, it also demonstrates a degree of robustness of this model.\r\n\r\n[/quote]",
    "108027": "The age tag is a 4 character string: 3 digits followed by 1 letter (Y, M or W).",
    "108028": "Ah, true, you are right, I just found that as well. Thanks for posting!",
    "108270": "Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W",
    "108374": "[quote=WD;108270]\r\n\r\n@Paul. this is very interesting. I am less familiar wiht piecewise-linear functions. Would you be willing and able to post your code? would love to take a look! W\r\n\r\n[/quote]\r\n\r\nIt is just a brute force search coded in C++ with [Imebra][1] dependency for reading DICOM. If you are still interested, I can put it on github.\r\n\r\n\r\n  [1]: https://imebra.com/",
    "108801": "Paul Jurczak definitely interested, please put it on github",
    "108818": "phunter\r\n\r\nI posted it to https://github.com/pauljurczak/Second-Annual-Data-Science-Bowl. Original project directory structure is not preserved, because I have only a faint idea how to use github.",
    "108827": "Paul, @phunter. I am debating and would love your thoughts on the best way to stack models. If one uses a model that has a 600-datapoint \"CDF\" as output, is the only way to stack this model with a linear regression estimate following Paul's approach to a) transform the CNN output to a point estimate (e.g. the mean of the PDF), then do some form of stacking (e.g. linear regression) on top of both Paul's model and the CNN model, and then subsequently transform the prediction with a sigma back into a CDF? I tried this approach, and performance was not great, maybe because there is so much risk of noise in the transformation from CDF to PDF and back again. Is there any other elegant approach that I am missing? Is there  a way to directly take the 600-datapoint CDF as input in a stacked model?",
    "108828": "WD\r\n\r\nI have only a faint understanding of CNN and I'm hoping to get a good score without using them. This model was only an exercise, but I may use it as a sanity check for the next model based on image processing and as a fallback option for studies where image processing fails.\r\n\r\nOne idea for you to try is to look for the extreme outliers (relative to my model) generated by CNN and replace them by my model's prediction. You may consider augmenting the model with lower and upper volume range to filter out really improbable CNN predictions.",
    "110301": "This post for some reason made me picture Michael Jordan taking a foul shot with his eyes closed :-)"
  },
  "source": "meta"
}