{
  "id": 180434,
  "title": "Naive (Model-less) Baseline vs Top Models",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/180434",
  "author_name": "Bryan Arnold",
  "post_date": "2020-09-05T03:26:56.971000",
  "votes": 23,
  "comment_count": 15,
  "views": 0,
  "content": "<p>See my <a href=\"https://www.kaggle.com/puremath86/naive-baseline-lb-7-1?scriptVersionId=42010889\" target=\"_blank\">notebook</a>.</p>\n<p>I'm a late starter to this competition, so my apologies if this is old news to anyone. We are able to achieve a -7.1 Public LB score with a completely naive approach. The top score, as of yet, is about a -6.7. The theoretically best score is about a -4.6. So what are everyone's thoughts on this?</p>\n<p>I have seen a lot of public kernels using some really leaky features and information that is not available for the final prediction. It appears we can easily get a -6.9 or so with relatively naive models --<code>FVC</code> as a linear function of <code>Weeks</code> for each of the five <code>Patient</code>s in the test set. But to do much better we are going to have to incorporate information about the lungs.</p>\n<p>I know no one likes to give away their secret sauce --especially if they're on the trail of a potentially good idea-- but I'd like to discuss ideas and/or approaches. At the very least we can talk about what we have tried and hasn't worked… and save newbies and late-comers like myself some time. :P</p>\n<p>I was thinking we should be able to calculate the lung volume directly from a pixel count of the lungs. We would have to adjust for zoom/resolution/angle/etc., but that should be doable. I was also thinking that we can denoise the target in the training set (LOWESS?), remove outliers, and do some post-processing on the predictions --since it is just a time-series. Thoughts?</p>\n<p>There are a myriad of other issues to address too. So what does everyone think?</p>\n<p><strong>EDIT</strong> Ok. I guess I should also ask the question: LB Probing?</p>",
  "messages": [
    {
      "id": 998754,
      "postDate": "2020-09-05T03:26:56.973Z",
      "content": "<p>See my <a href=\"https://www.kaggle.com/puremath86/naive-baseline-lb-7-1?scriptVersionId=42010889\" target=\"_blank\">notebook</a>.</p>\n<p>I'm a late starter to this competition, so my apologies if this is old news to anyone. We are able to achieve a -7.1 Public LB score with a completely naive approach. The top score, as of yet, is about a -6.7. The theoretically best score is about a -4.6. So what are everyone's thoughts on this?</p>\n<p>I have seen a lot of public kernels using some really leaky features and information that is not available for the final prediction. It appears we can easily get a -6.9 or so with relatively naive models --<code>FVC</code> as a linear function of <code>Weeks</code> for each of the five <code>Patient</code>s in the test set. But to do much better we are going to have to incorporate information about the lungs.</p>\n<p>I know no one likes to give away their secret sauce --especially if they're on the trail of a potentially good idea-- but I'd like to discuss ideas and/or approaches. At the very least we can talk about what we have tried and hasn't worked… and save newbies and late-comers like myself some time. :P</p>\n<p>I was thinking we should be able to calculate the lung volume directly from a pixel count of the lungs. We would have to adjust for zoom/resolution/angle/etc., but that should be doable. I was also thinking that we can denoise the target in the training set (LOWESS?), remove outliers, and do some post-processing on the predictions --since it is just a time-series. Thoughts?</p>\n<p>There are a myriad of other issues to address too. So what does everyone think?</p>\n<p><strong>EDIT</strong> Ok. I guess I should also ask the question: LB Probing?</p>",
      "rawMarkdown": "See my [notebook](https://www.kaggle.com/puremath86/naive-baseline-lb-7-1?scriptVersionId=42010889).\n\nI'm a late starter to this competition, so my apologies if this is old news to anyone. We are able to achieve a -7.1 Public LB score with a completely naive approach. The top score, as of yet, is about a -6.7. The theoretically best score is about a -4.6. So what are everyone's thoughts on this?\n\nI have seen a lot of public kernels using some really leaky features and information that is not available for the final prediction. It appears we can easily get a -6.9 or so with relatively naive models --`FVC` as a linear function of `Weeks` for each of the five `Patient`s in the test set. But to do much better we are going to have to incorporate information about the lungs.\n\nI know no one likes to give away their secret sauce --especially if they're on the trail of a potentially good idea-- but I'd like to discuss ideas and/or approaches. At the very least we can talk about what we have tried and hasn't worked... and save newbies and late-comers like myself some time. :P\n\nI was thinking we should be able to calculate the lung volume directly from a pixel count of the lungs. We would have to adjust for zoom/resolution/angle/etc., but that should be doable. I was also thinking that we can denoise the target in the training set (LOWESS?), remove outliers, and do some post-processing on the predictions --since it is just a time-series. Thoughts?\n\nThere are a myriad of other issues to address too. So what does everyone think?\n\n**EDIT** Ok. I guess I should also ask the question: LB Probing?",
      "votes": 23
    },
    {
      "id": 1002191,
      "postDate": "2020-09-07T22:48:41.797Z",
      "content": "<p>At least in my experiments, lung volume information doesn't bring me any improvement yet.</p>",
      "rawMarkdown": "At least in my experiments, lung volume information doesn't bring me any improvement yet.",
      "votes": 5,
      "replies": [
        {
          "id": 1002268,
          "postDate": "2020-09-08T02:12:59.753Z",
          "content": "<p>Thanks for sharing. I'm thinking the CT Scans won't be very useful for this competition due to lack of data and the noise in the FVC measurements.</p>",
          "rawMarkdown": "Thanks for sharing. I'm thinking the CT Scans won't be very useful for this competition due to lack of data and the noise in the FVC measurements.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1002373,
      "postDate": "2020-09-08T04:54:05.500Z",
      "content": "<p>I think this is an amazing competition. Especially because it is hard.<br>\nIMHO there are 2 key important things to create great models for this competition:</p>\n<ol>\n<li>Extracting relevant features from the CT scans</li>\n<li>Properly predicting uncertainty</li>\n</ol>\n<p>Fitting curves through the ~1500 FVC points in the training set is easy. The hard thing to do is (1) and (2). In (1), I've tried manual feature engineering (+21 features) and unsupervised learning techniques (Variational Autoencoder, +16 features). This code is not public yet, but I intend to do it once the competition ends. In (2), there are 2 approaches: quantile models/pinball loss, and Bayesian approach/probabilistic machine learning. Also tried both, and made some of the code already public. </p>",
      "rawMarkdown": "I think this is an amazing competition. Especially because it is hard.\nIMHO there are 2 key important things to create great models for this competition:\n1. Extracting relevant features from the CT scans\n2. Properly predicting uncertainty\n\nFitting curves through the ~1500 FVC points in the training set is easy. The hard thing to do is (1) and (2). In (1), I've tried manual feature engineering (+21 features) and unsupervised learning techniques (Variational Autoencoder, +16 features). This code is not public yet, but I intend to do it once the competition ends. In (2), there are 2 approaches: quantile models/pinball loss, and Bayesian approach/probabilistic machine learning. Also tried both, and made some of the code already public. \n",
      "votes": 3,
      "replies": [
        {
          "id": 1002395,
          "postDate": "2020-09-08T05:23:45.973Z",
          "content": "<p>Thanks for sharing your thoughts. I agree the competition is interesting. I'm just afraid the winning solutions will not be useful to the competition host. If computer vision-based approaches don't add substantially more value than some rudimentary approaches --or just collecting more spirometry measurements directly from patients-- then they aren't worth it. I'd loved to be wrong though. And maybe the feature engineering and approaches themselves will be of value.</p>\n<p>In my preliminary experiments, I'm finding that probing the LB should be able to attain a score of ~ -5.7 --if you can probe enough data for each of the five patients before the competition is over. And I'm not convinced I'll be able to achieve that since I started so late and only 15% of the LB is public. Fortunately, it doesn't look like many competitors are taking this approach. But who knows?</p>",
          "rawMarkdown": "Thanks for sharing your thoughts. I agree the competition is interesting. I'm just afraid the winning solutions will not be useful to the competition host. If computer vision-based approaches don't add substantially more value than some rudimentary approaches --or just collecting more spirometry measurements directly from patients-- then they aren't worth it. I'd loved to be wrong though. And maybe the feature engineering and approaches themselves will be of value.\n\nIn my preliminary experiments, I'm finding that probing the LB should be able to attain a score of ~ -5.7 --if you can probe enough data for each of the five patients before the competition is over. And I'm not convinced I'll be able to achieve that since I started so late and only 15% of the LB is public. Fortunately, it doesn't look like many competitors are taking this approach. But who knows?"
        },
        {
          "id": 1002405,
          "postDate": "2020-09-08T05:29:25.657Z",
          "content": "<p>If I were the host, I wouldn't care about the solutions. I'd care about the people who wrote them ;)</p>",
          "rawMarkdown": "If I were the host, I wouldn't care about the solutions. I'd care about the people who wrote them ;)",
          "votes": 2
        },
        {
          "id": 1008509,
          "postDate": "2020-09-13T07:08:22.723Z",
          "content": "<p>Hi Carlos, </p>\n<p>21 manual features!? 🤪 are those all from the ct scans? I can think only of: Lung volume, airway volume, some measure of texture. I know you are keeping your cards close to your chest on this until later in the competition but please could you give some hints of what other features are promising?</p>",
          "rawMarkdown": "Hi Carlos, \n\n21 manual features!? 🤪 are those all from the ct scans? I can think only of: Lung volume, airway volume, some measure of texture. I know you are keeping your cards close to your chest on this until later in the competition but please could you give some hints of what other features are promising?",
          "votes": 1
        },
        {
          "id": 1009355,
          "postDate": "2020-09-13T22:31:07.470Z",
          "content": "<p>The manual features are no secret… I think I mentioned in a post or notebook:</p>\n<ul>\n<li>Lung volume in L</li>\n<li>Chest circumference in mm</li>\n<li>Lung height in mm</li>\n<li>4 moments of HU distribution (mean, stddev, skew, kurtosis)</li>\n<li>% of voxels in several different bins of different HU intervals (e.g. -1000 to -950 HUs, -950 to -900 HUs, etc)</li>\n</ul>\n<p>There's also the features from the unsupervised approach…<br>\nHope it helps! Cheers!</p>",
          "rawMarkdown": "The manual features are no secret... I think I mentioned in a post or notebook:\n- Lung volume in L\n- Chest circumference in mm\n- Lung height in mm\n- 4 moments of HU distribution (mean, stddev, skew, kurtosis)\n- % of voxels in several different bins of different HU intervals (e.g. -1000 to -950 HUs, -950 to -900 HUs, etc)\n\nThere's also the features from the unsupervised approach...\nHope it helps! Cheers!",
          "votes": 5
        }
      ]
    },
    {
      "id": 1001377,
      "postDate": "2020-09-07T09:45:16.110Z",
      "content": "<p>Doesn't work for me: Include simple image statistics (from segmented lungs) like mean, std, skew, kurtosis = didn't improve CV</p>",
      "rawMarkdown": "Doesn't work for me: Include simple image statistics (from segmented lungs) like mean, std, skew, kurtosis = didn't improve CV",
      "votes": 3,
      "replies": [
        {
          "id": 1001381,
          "postDate": "2020-09-07T09:50:43.430Z",
          "content": "<p>Thanks for sharing. </p>",
          "rawMarkdown": "Thanks for sharing. "
        }
      ]
    },
    {
      "id": 1003123,
      "postDate": "2020-09-08T17:11:57.063Z",
      "content": "<p>Hello Bryan, I too recently joined the competition and have been trying to extract features form the CT scans. I am going to share some of my initial findings. So far I have managed to obtain the following features from images:</p>\n<ol>\n<li><strong>Lung Volume:</strong> This variable is positively correlated with FVC and has a Pearson corr. coeff. of 0.35</li>\n<li><strong>Area of tissue pixels:</strong> With some effort I managed to segment the tissue pixels and obtained the number of tissue pixels by setting some threshold. This variable is negatively correlated with FVC and a Pearson corr. coeff. of -0.02. I am planning to experiment with the threshold values in the future and include multiple features with different thresholds. </li>\n<li><strong>Other variables:</strong> I have also extracted a few other tissue related features which show a good negative correlation with FVC. Might not be able reveal much at this point though.</li>\n</ol>\n<p>Below are a couple of normalized distributions of the variables mentioned above:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2568215%2F04f8262565ef36522a3024495e515d16%2FScreenshot%202020-09-08%20at%209.59.26%20PM.png?generation=1599582732378028&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello Bryan, I too recently joined the competition and have been trying to extract features form the CT scans. I am going to share some of my initial findings. So far I have managed to obtain the following features from images:\n1. **Lung Volume:** This variable is positively correlated with FVC and has a Pearson corr. coeff. of 0.35\n2. **Area of tissue pixels:** With some effort I managed to segment the tissue pixels and obtained the number of tissue pixels by setting some threshold. This variable is negatively correlated with FVC and a Pearson corr. coeff. of -0.02. I am planning to experiment with the threshold values in the future and include multiple features with different thresholds. \n3. **Other variables:** I have also extracted a few other tissue related features which show a good negative correlation with FVC. Might not be able reveal much at this point though.\n\nBelow are a couple of normalized distributions of the variables mentioned above:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2568215%2F04f8262565ef36522a3024495e515d16%2FScreenshot%202020-09-08%20at%209.59.26%20PM.png?generation=1599582732378028&alt=media)\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 1003130,
          "postDate": "2020-09-08T17:20:57.173Z",
          "content": "<p>Thanks, Abhishek. That's interesting. I'm definitely going to look at lung volume features in the next few days.</p>",
          "rawMarkdown": "Thanks, Abhishek. That's interesting. I'm definitely going to look at lung volume features in the next few days."
        }
      ]
    },
    {
      "id": 1002627,
      "postDate": "2020-09-08T09:33:58.670Z",
      "content": "<p>I've found that the most important factor in this competition is \"gaming the metric\". By this I mean that a model that predicts future FVCs can be easily useless if given uncertainties aren't accurate enough (according to metric).</p>",
      "rawMarkdown": "I've found that the most important factor in this competition is \"gaming the metric\". By this I mean that a model that predicts future FVCs can be easily useless if given uncertainties aren't accurate enough (according to metric)."
    },
    {
      "id": 1000771,
      "postDate": "2020-09-06T19:24:27.860Z",
      "content": "<p>No one wants to share their recipe….</p>",
      "rawMarkdown": "No one wants to share their recipe....",
      "replies": [
        {
          "id": 1000920,
          "postDate": "2020-09-06T23:05:08.420Z",
          "content": "<p>No worries. I don't blame them. The lack of data and various other factors make this a bit of a challenge.</p>",
          "rawMarkdown": "No worries. I don't blame them. The lack of data and various other factors make this a bit of a challenge."
        }
      ]
    },
    {
      "id": 1003355,
      "postDate": "2020-09-08T21:32:38.503Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1002191,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2020-09-07T22:48:41.797000",
      "content": "<p>At least in my experiments, lung volume information doesn't bring me any improvement yet.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1002268,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2020-09-08T02:12:59.753000",
          "content": "<p>Thanks for sharing. I'm thinking the CT Scans won't be very useful for this competition due to lack of data and the noise in the FVC measurements.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1002373,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-09-08T04:54:05.500000",
      "content": "<p>I think this is an amazing competition. Especially because it is hard.<br>\nIMHO there are 2 key important things to create great models for this competition:</p>\n<ol>\n<li>Extracting relevant features from the CT scans</li>\n<li>Properly predicting uncertainty</li>\n</ol>\n<p>Fitting curves through the ~1500 FVC points in the training set is easy. The hard thing to do is (1) and (2). In (1), I've tried manual feature engineering (+21 features) and unsupervised learning techniques (Variational Autoencoder, +16 features). This code is not public yet, but I intend to do it once the competition ends. In (2), there are 2 approaches: quantile models/pinball loss, and Bayesian approach/probabilistic machine learning. Also tried both, and made some of the code already public. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1002395,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2020-09-08T05:23:45.973000",
          "content": "<p>Thanks for sharing your thoughts. I agree the competition is interesting. I'm just afraid the winning solutions will not be useful to the competition host. If computer vision-based approaches don't add substantially more value than some rudimentary approaches --or just collecting more spirometry measurements directly from patients-- then they aren't worth it. I'd loved to be wrong though. And maybe the feature engineering and approaches themselves will be of value.</p>\n<p>In my preliminary experiments, I'm finding that probing the LB should be able to attain a score of ~ -5.7 --if you can probe enough data for each of the five patients before the competition is over. And I'm not convinced I'll be able to achieve that since I started so late and only 15% of the LB is public. Fortunately, it doesn't look like many competitors are taking this approach. But who knows?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1002405,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-09-08T05:29:25.657000",
          "content": "<p>If I were the host, I wouldn't care about the solutions. I'd care about the people who wrote them ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1008509,
          "author_name": "Mark P",
          "author_url": "",
          "post_date": "2020-09-13T07:08:22.723000",
          "content": "<p>Hi Carlos, </p>\n<p>21 manual features!? 🤪 are those all from the ct scans? I can think only of: Lung volume, airway volume, some measure of texture. I know you are keeping your cards close to your chest on this until later in the competition but please could you give some hints of what other features are promising?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1009355,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-09-13T22:31:07.470000",
          "content": "<p>The manual features are no secret… I think I mentioned in a post or notebook:</p>\n<ul>\n<li>Lung volume in L</li>\n<li>Chest circumference in mm</li>\n<li>Lung height in mm</li>\n<li>4 moments of HU distribution (mean, stddev, skew, kurtosis)</li>\n<li>% of voxels in several different bins of different HU intervals (e.g. -1000 to -950 HUs, -950 to -900 HUs, etc)</li>\n</ul>\n<p>There's also the features from the unsupervised approach…<br>\nHope it helps! Cheers!</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1001377,
      "author_name": "A.Demyanchuk",
      "author_url": "",
      "post_date": "2020-09-07T09:45:16.110000",
      "content": "<p>Doesn't work for me: Include simple image statistics (from segmented lungs) like mean, std, skew, kurtosis = didn't improve CV</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1001381,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2020-09-07T09:50:43.430000",
          "content": "<p>Thanks for sharing. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1003123,
      "author_name": "Abhishek Bhat",
      "author_url": "",
      "post_date": "2020-09-08T17:11:57.063000",
      "content": "<p>Hello Bryan, I too recently joined the competition and have been trying to extract features form the CT scans. I am going to share some of my initial findings. So far I have managed to obtain the following features from images:</p>\n<ol>\n<li><strong>Lung Volume:</strong> This variable is positively correlated with FVC and has a Pearson corr. coeff. of 0.35</li>\n<li><strong>Area of tissue pixels:</strong> With some effort I managed to segment the tissue pixels and obtained the number of tissue pixels by setting some threshold. This variable is negatively correlated with FVC and a Pearson corr. coeff. of -0.02. I am planning to experiment with the threshold values in the future and include multiple features with different thresholds. </li>\n<li><strong>Other variables:</strong> I have also extracted a few other tissue related features which show a good negative correlation with FVC. Might not be able reveal much at this point though.</li>\n</ol>\n<p>Below are a couple of normalized distributions of the variables mentioned above:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2568215%2F04f8262565ef36522a3024495e515d16%2FScreenshot%202020-09-08%20at%209.59.26%20PM.png?generation=1599582732378028&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1003130,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2020-09-08T17:20:57.173000",
          "content": "<p>Thanks, Abhishek. That's interesting. I'm definitely going to look at lung volume features in the next few days.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1002627,
      "author_name": "David Castillo",
      "author_url": "",
      "post_date": "2020-09-08T09:33:58.670000",
      "content": "<p>I've found that the most important factor in this competition is \"gaming the metric\". By this I mean that a model that predicts future FVCs can be easily useless if given uncertainties aren't accurate enough (according to metric).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1000771,
      "author_name": "Ahmmad Rashid",
      "author_url": "",
      "post_date": "2020-09-06T19:24:27.860000",
      "content": "<p>No one wants to share their recipe….</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1000920,
          "author_name": "Bryan Arnold",
          "author_url": "",
          "post_date": "2020-09-06T23:05:08.420000",
          "content": "<p>No worries. I don't blame them. The lack of data and various other factors make this a bit of a challenge.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1003355,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-08T21:32:38.503000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "998754": "See my [notebook](https://www.kaggle.com/puremath86/naive-baseline-lb-7-1?scriptVersionId=42010889).\n\nI'm a late starter to this competition, so my apologies if this is old news to anyone. We are able to achieve a -7.1 Public LB score with a completely naive approach. The top score, as of yet, is about a -6.7. The theoretically best score is about a -4.6. So what are everyone's thoughts on this?\n\nI have seen a lot of public kernels using some really leaky features and information that is not available for the final prediction. It appears we can easily get a -6.9 or so with relatively naive models --`FVC` as a linear function of `Weeks` for each of the five `Patient`s in the test set. But to do much better we are going to have to incorporate information about the lungs.\n\nI know no one likes to give away their secret sauce --especially if they're on the trail of a potentially good idea-- but I'd like to discuss ideas and/or approaches. At the very least we can talk about what we have tried and hasn't worked... and save newbies and late-comers like myself some time. :P\n\nI was thinking we should be able to calculate the lung volume directly from a pixel count of the lungs. We would have to adjust for zoom/resolution/angle/etc., but that should be doable. I was also thinking that we can denoise the target in the training set (LOWESS?), remove outliers, and do some post-processing on the predictions --since it is just a time-series. Thoughts?\n\nThere are a myriad of other issues to address too. So what does everyone think?\n\n**EDIT** Ok. I guess I should also ask the question: LB Probing?",
    "1002191": "At least in my experiments, lung volume information doesn't bring me any improvement yet.",
    "1002373": "I think this is an amazing competition. Especially because it is hard.\nIMHO there are 2 key important things to create great models for this competition:\n1. Extracting relevant features from the CT scans\n2. Properly predicting uncertainty\n\nFitting curves through the ~1500 FVC points in the training set is easy. The hard thing to do is (1) and (2). In (1), I've tried manual feature engineering (+21 features) and unsupervised learning techniques (Variational Autoencoder, +16 features). This code is not public yet, but I intend to do it once the competition ends. In (2), there are 2 approaches: quantile models/pinball loss, and Bayesian approach/probabilistic machine learning. Also tried both, and made some of the code already public. \n",
    "1001377": "Doesn't work for me: Include simple image statistics (from segmented lungs) like mean, std, skew, kurtosis = didn't improve CV",
    "1003123": "Hello Bryan, I too recently joined the competition and have been trying to extract features form the CT scans. I am going to share some of my initial findings. So far I have managed to obtain the following features from images:\n1. **Lung Volume:** This variable is positively correlated with FVC and has a Pearson corr. coeff. of 0.35\n2. **Area of tissue pixels:** With some effort I managed to segment the tissue pixels and obtained the number of tissue pixels by setting some threshold. This variable is negatively correlated with FVC and a Pearson corr. coeff. of -0.02. I am planning to experiment with the threshold values in the future and include multiple features with different thresholds. \n3. **Other variables:** I have also extracted a few other tissue related features which show a good negative correlation with FVC. Might not be able reveal much at this point though.\n\nBelow are a couple of normalized distributions of the variables mentioned above:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2568215%2F04f8262565ef36522a3024495e515d16%2FScreenshot%202020-09-08%20at%209.59.26%20PM.png?generation=1599582732378028&alt=media)\n\n",
    "1002627": "I've found that the most important factor in this competition is \"gaming the metric\". By this I mean that a model that predicts future FVCs can be easily useless if given uncertainties aren't accurate enough (according to metric).",
    "1000771": "No one wants to share their recipe....",
    "1003355": ""
  }
}