{
  "id": 183235,
  "title": "Test time augmentation for uncertainty estimation",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/183235",
  "author_name": "",
  "post_date": "2020-09-16T02:35:27.274205100Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>One method for estimating the uncertainty in predictions is the use of test time augmentation. This method allows image data to be used to make predictions and from the experiments I have performed it looks like it could perform very well if executed correctly. </p>\n<p>The idea that I had to implement this can be described step by step as follows.<br>\n 1) Train a network to predict the slope and intercept using both image and tabular data<br>\n 2) Make <code>n_tta</code> predictions per slice for the slope and intercept<br>\n 3) Make a prediction for weeks from (-12,133) using all slope and intercept pairs for all slices<br>\n 4) Take the arithmetic mean across all weeks and all slices for the prediction of each patient, and the standard deviation for the confidence</p>\n<p>This approach seemed to work fairly well with an OOF CV score of ~-8 when predicting both intercept and slope with the network, and a score of ~-6.9 when just predicting the slope and inferring the intercept (both results are the best scores seen). The latter approached could be significantly improved by developing a method for predicting the intercept accurately.</p>\n<p>I no longer have the time to work on developing this idea further, and I have made all of the notebooks that I have been working on publicly available.</p>\n<p>The notebook that predicts the slope is <a href=\"https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu\" target=\"_blank\">here</a>. In this notebook you can see that the metric does very well on most examples in the validation set but the few outliers skew everything, also note that this is not the true metric that is used in evaluation when submitting. There are a lot of things that can be improved upon from this notebook, adding a weight to less commonly observed slopes could be a good start as that could reduce the high scores that ruin the average, or predicting the slope at the same time (although this is a little tricky because of the different scales in the slope and intercept). Also the data that is trained on needs to be explored as the results depend heavily on this. Note also that the results from this notebook are slightly misleading because they assume the intercept is known perfectly, but even when predicting the intercept as well the scores don't do much worse.</p>\n<p>Another attempt I made was to predict the slopes and intercept predicted by Bayesian programming as in <a href=\"https://www.kaggle.com/samklein/kfold-osic-efficientnet\" target=\"_blank\">this notebook</a>. Earlier versions of this (ver. 15 for example) has the code to predict the slope and intercept simultaneously although I think it has a few bugs that have been fixed in the first notebook commented in this post.</p>\n<p>Still quite a bit of work to do but hopefully this is helpful to someone!</p>",
  "messages": [
    {
      "id": "1012287",
      "postDate": "09/16/2020 02:35:27",
      "content": "<p>One method for estimating the uncertainty in predictions is the use of test time augmentation. This method allows image data to be used to make predictions and from the experiments I have performed it looks like it could perform very well if executed correctly. </p>\n<p>The idea that I had to implement this can be described step by step as follows.<br>\n 1) Train a network to predict the slope and intercept using both image and tabular data<br>\n 2) Make <code>n_tta</code> predictions per slice for the slope and intercept<br>\n 3) Make a prediction for weeks from (-12,133) using all slope and intercept pairs for all slices<br>\n 4) Take the arithmetic mean across all weeks and all slices for the prediction of each patient, and the standard deviation for the confidence</p>\n<p>This approach seemed to work fairly well with an OOF CV score of ~-8 when predicting both intercept and slope with the network, and a score of ~-6.9 when just predicting the slope and inferring the intercept (both results are the best scores seen). The latter approached could be significantly improved by developing a method for predicting the intercept accurately.</p>\n<p>I no longer have the time to work on developing this idea further, and I have made all of the notebooks that I have been working on publicly available.</p>\n<p>The notebook that predicts the slope is <a href=\"https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu\" target=\"_blank\">here</a>. In this notebook you can see that the metric does very well on most examples in the validation set but the few outliers skew everything, also note that this is not the true metric that is used in evaluation when submitting. There are a lot of things that can be improved upon from this notebook, adding a weight to less commonly observed slopes could be a good start as that could reduce the high scores that ruin the average, or predicting the slope at the same time (although this is a little tricky because of the different scales in the slope and intercept). Also the data that is trained on needs to be explored as the results depend heavily on this. Note also that the results from this notebook are slightly misleading because they assume the intercept is known perfectly, but even when predicting the intercept as well the scores don't do much worse.</p>\n<p>Another attempt I made was to predict the slopes and intercept predicted by Bayesian programming as in <a href=\"https://www.kaggle.com/samklein/kfold-osic-efficientnet\" target=\"_blank\">this notebook</a>. Earlier versions of this (ver. 15 for example) has the code to predict the slope and intercept simultaneously although I think it has a few bugs that have been fixed in the first notebook commented in this post.</p>\n<p>Still quite a bit of work to do but hopefully this is helpful to someone!</p>",
      "rawMarkdown": "One method for estimating the uncertainty in predictions is the use of test time augmentation. This method allows image data to be used to make predictions and from the experiments I have performed it looks like it could perform very well if executed correctly. \n\nThe idea that I had to implement this can be described step by step as follows.\n 1) Train a network to predict the slope and intercept using both image and tabular data\n 2) Make `n_tta` predictions per slice for the slope and intercept\n 3) Make a prediction for weeks from (-12,133) using all slope and intercept pairs for all slices\n 4) Take the arithmetic mean across all weeks and all slices for the prediction of each patient, and the standard deviation for the confidence\n\nThis approach seemed to work fairly well with an OOF CV score of ~-8 when predicting both intercept and slope with the network, and a score of ~-6.9 when just predicting the slope and inferring the intercept (both results are the best scores seen). The latter approached could be significantly improved by developing a method for predicting the intercept accurately.\n\nI no longer have the time to work on developing this idea further, and I have made all of the notebooks that I have been working on publicly available.\n\nThe notebook that predicts the slope is [here](https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu). In this notebook you can see that the metric does very well on most examples in the validation set but the few outliers skew everything, also note that this is not the true metric that is used in evaluation when submitting. There are a lot of things that can be improved upon from this notebook, adding a weight to less commonly observed slopes could be a good start as that could reduce the high scores that ruin the average, or predicting the slope at the same time (although this is a little tricky because of the different scales in the slope and intercept). Also the data that is trained on needs to be explored as the results depend heavily on this. Note also that the results from this notebook are slightly misleading because they assume the intercept is known perfectly, but even when predicting the intercept as well the scores don't do much worse.\n\nAnother attempt I made was to predict the slopes and intercept predicted by Bayesian programming as in [this notebook](https://www.kaggle.com/samklein/kfold-osic-efficientnet). Earlier versions of this (ver. 15 for example) has the code to predict the slope and intercept simultaneously although I think it has a few bugs that have been fixed in the first notebook commented in this post.\n\nStill quite a bit of work to do but hopefully this is helpful to someone!",
      "votes": null
    },
    {
      "id": "1012308",
      "postDate": "09/16/2020 02:50:59",
      "content": "<p>Really helpful! I'll be definitely trying this soon! Thanks for your contribution towards the community. Appreciate it :)</p>",
      "rawMarkdown": "Really helpful! I'll be definitely trying this soon! Thanks for your contribution towards the community. Appreciate it :)",
      "votes": null
    },
    {
      "id": "1012374",
      "postDate": "09/16/2020 03:54:28",
      "content": "<p>No worries, and good luck. It would be great to see how it goes!</p>",
      "rawMarkdown": "No worries, and good luck. It would be great to see how it goes!",
      "votes": null
    },
    {
      "id": "1014555",
      "postDate": "09/17/2020 14:39:09",
      "content": "<p>Sorry, but what do you mean by n_tta prediction and slice? </p>",
      "rawMarkdown": "Sorry, but what do you mean by n_tta prediction and slice?",
      "votes": null
    },
    {
      "id": "1015129",
      "postDate": "09/18/2020 01:20:07",
      "content": "<p>Sorry not to be clear. Each CT scan comes as two dimensional images, and I refer to each of these as a slice. I make <code>n_tta</code> predictions per slice where each prediction is augmented randomly. If that is not clear let me know.</p>",
      "rawMarkdown": "Sorry not to be clear. Each CT scan comes as two dimensional images, and I refer to each of these as a slice. I make `n_tta` predictions per slice where each prediction is augmented randomly. If that is not clear let me know.",
      "votes": null
    },
    {
      "id": "1016156",
      "postDate": "09/18/2020 17:49:42",
      "content": "<p>This sounds like a super interesting idea, will surely give it a try. The only worry is that I might not be able to keep the notebook runtime below 4hrs with these TTA steps. Mainly because I'm also segmenting the lung, extracting tissue area, lung volumn etc. and including them in the tabular input for CNN. I am also using a 5-fold CV and average ensembling my predictions from models built in each fold. All this takes around a 3 hours. Will have to optimize my code further to include TTA. </p>",
      "rawMarkdown": "This sounds like a super interesting idea, will surely give it a try. The only worry is that I might not be able to keep the notebook runtime below 4hrs with these TTA steps. Mainly because I'm also segmenting the lung, extracting tissue area, lung volumn etc. and including them in the tabular input for CNN. I am also using a 5-fold CV and average ensembling my predictions from models built in each fold. All this takes around a 3 hours. Will have to optimize my code further to include TTA.",
      "votes": null
    },
    {
      "id": "1016324",
      "postDate": "09/18/2020 20:33:41",
      "content": "<p>Cool, I hope it works for you. The speed could definitely be an issue, but training models on TPU and loading them for evaluation of GPU works, but with the TTA the inference does take a little while. Good luck!</p>",
      "rawMarkdown": "Cool, I hope it works for you. The speed could definitely be an issue, but training models on TPU and loading them for evaluation of GPU works, but with the TTA the inference does take a little while. Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012308,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "09/16/2020 02:50:59",
      "content": "<p>Really helpful! I'll be definitely trying this soon! Thanks for your contribution towards the community. Appreciate it :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012374,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/16/2020 03:54:28",
          "content": "<p>No worries, and good luck. It would be great to see how it goes!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1014555,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "09/17/2020 14:39:09",
      "content": "<p>Sorry, but what do you mean by n_tta prediction and slice? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1015129,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/18/2020 01:20:07",
          "content": "<p>Sorry not to be clear. Each CT scan comes as two dimensional images, and I refer to each of these as a slice. I make <code>n_tta</code> predictions per slice where each prediction is augmented randomly. If that is not clear let me know.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1016156,
      "author_name": "abhishekgbhat",
      "author_url": "",
      "post_date": "09/18/2020 17:49:42",
      "content": "<p>This sounds like a super interesting idea, will surely give it a try. The only worry is that I might not be able to keep the notebook runtime below 4hrs with these TTA steps. Mainly because I'm also segmenting the lung, extracting tissue area, lung volumn etc. and including them in the tabular input for CNN. I am also using a 5-fold CV and average ensembling my predictions from models built in each fold. All this takes around a 3 hours. Will have to optimize my code further to include TTA. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1016324,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/18/2020 20:33:41",
          "content": "<p>Cool, I hope it works for you. The speed could definitely be an issue, but training models on TPU and loading them for evaluation of GPU works, but with the TTA the inference does take a little while. Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1012287": "One method for estimating the uncertainty in predictions is the use of test time augmentation. This method allows image data to be used to make predictions and from the experiments I have performed it looks like it could perform very well if executed correctly. \n\nThe idea that I had to implement this can be described step by step as follows.\n 1) Train a network to predict the slope and intercept using both image and tabular data\n 2) Make `n_tta` predictions per slice for the slope and intercept\n 3) Make a prediction for weeks from (-12,133) using all slope and intercept pairs for all slices\n 4) Take the arithmetic mean across all weeks and all slices for the prediction of each patient, and the standard deviation for the confidence\n\nThis approach seemed to work fairly well with an OOF CV score of ~-8 when predicting both intercept and slope with the network, and a score of ~-6.9 when just predicting the slope and inferring the intercept (both results are the best scores seen). The latter approached could be significantly improved by developing a method for predicting the intercept accurately.\n\nI no longer have the time to work on developing this idea further, and I have made all of the notebooks that I have been working on publicly available.\n\nThe notebook that predicts the slope is [here](https://www.kaggle.com/samklein/kfold-osic-efficientnet-slope-tta-confidence-tpu). In this notebook you can see that the metric does very well on most examples in the validation set but the few outliers skew everything, also note that this is not the true metric that is used in evaluation when submitting. There are a lot of things that can be improved upon from this notebook, adding a weight to less commonly observed slopes could be a good start as that could reduce the high scores that ruin the average, or predicting the slope at the same time (although this is a little tricky because of the different scales in the slope and intercept). Also the data that is trained on needs to be explored as the results depend heavily on this. Note also that the results from this notebook are slightly misleading because they assume the intercept is known perfectly, but even when predicting the intercept as well the scores don't do much worse.\n\nAnother attempt I made was to predict the slopes and intercept predicted by Bayesian programming as in [this notebook](https://www.kaggle.com/samklein/kfold-osic-efficientnet). Earlier versions of this (ver. 15 for example) has the code to predict the slope and intercept simultaneously although I think it has a few bugs that have been fixed in the first notebook commented in this post.\n\nStill quite a bit of work to do but hopefully this is helpful to someone!",
    "1012308": "Really helpful! I'll be definitely trying this soon! Thanks for your contribution towards the community. Appreciate it :)",
    "1012374": "No worries, and good luck. It would be great to see how it goes!",
    "1014555": "Sorry, but what do you mean by n_tta prediction and slice?",
    "1015129": "Sorry not to be clear. Each CT scan comes as two dimensional images, and I refer to each of these as a slice. I make `n_tta` predictions per slice where each prediction is augmented randomly. If that is not clear let me know.",
    "1016156": "This sounds like a super interesting idea, will surely give it a try. The only worry is that I might not be able to keep the notebook runtime below 4hrs with these TTA steps. Mainly because I'm also segmenting the lung, extracting tissue area, lung volumn etc. and including them in the tabular input for CNN. I am also using a 5-fold CV and average ensembling my predictions from models built in each fold. All this takes around a 3 hours. Will have to optimize my code further to include TTA.",
    "1016324": "Cool, I hope it works for you. The speed could definitely be an issue, but training models on TPU and loading them for evaluation of GPU works, but with the TTA the inference does take a little while. Good luck!"
  },
  "source": "meta"
}