{
  "id": 188861,
  "title": "Data augmentation via generation of synthetic CT scans & tabular data",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/188861",
  "author_name": "",
  "post_date": "2020-10-05T17:24:38.383604Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey guys,<br>\nAs our team is already reflecting on what worked &amp; didn't work, we discussed one avenue unfortunately we didn't have enough time to pursue: <strong>data augmentation via generation of synthetic CT scans &amp; tabular data</strong>. (There is a plethora of articles about this topic, because medical datasets are generally too expensive and too unbalanced.)</p>\n<p>Out of curiosity: has anyone tried anything in this direction? Did it work?</p>\n<p>Cheers!</p>",
  "messages": [
    {
      "id": "1038267",
      "postDate": "10/05/2020 17:24:38",
      "content": "<p>Hey guys,<br>\nAs our team is already reflecting on what worked &amp; didn't work, we discussed one avenue unfortunately we didn't have enough time to pursue: <strong>data augmentation via generation of synthetic CT scans &amp; tabular data</strong>. (There is a plethora of articles about this topic, because medical datasets are generally too expensive and too unbalanced.)</p>\n<p>Out of curiosity: has anyone tried anything in this direction? Did it work?</p>\n<p>Cheers!</p>",
      "rawMarkdown": "Hey guys,\nAs our team is already reflecting on what worked & didn't work, we discussed one avenue unfortunately we didn't have enough time to pursue: **data augmentation via generation of synthetic CT scans & tabular data**. (There is a plethora of articles about this topic, because medical datasets are generally too expensive and too unbalanced.)\n\nOut of curiosity: has anyone tried anything in this direction? Did it work?\n\nCheers!",
      "votes": null
    },
    {
      "id": "1038388",
      "postDate": "10/05/2020 19:06:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> ,</p>\n<p>i have been thinking about this to, but in this case it would only make sense to generate extra slices for the low slice-volume scans (like missing slices inbetween) to reduce a 5 mm  slice distance to 2.5 and so on.</p>\n<p>There are already plenty of CT thorax scans available, and one high slice count scan could be used multiple times with multiple gaps to train a GAN : take n-k-th slice and n+k-th slice to predict n-th slice… the lung and vessels, even mediastinum with heart could be predicted quite good i think, however the problem is with the fibrosis… i am not sure how well they would get.</p>",
      "rawMarkdown": "Hi @carlossouza ,\n\ni have been thinking about this to, but in this case it would only make sense to generate extra slices for the low slice-volume scans (like missing slices inbetween) to reduce a 5 mm  slice distance to 2.5 and so on.\n\nThere are already plenty of CT thorax scans available, and one high slice count scan could be used multiple times with multiple gaps to train a GAN : take n-k-th slice and n+k-th slice to predict n-th slice... the lung and vessels, even mediastinum with heart could be predicted quite good i think, however the problem is with the fibrosis... i am not sure how well they would get.",
      "votes": null
    },
    {
      "id": "1038401",
      "postDate": "10/05/2020 19:15:32",
      "content": "<p>Yes… but I think generating extra slices would produce marginal gains (if any)…</p>\n<p>My thinking was to completely synthesize new patients. The problem is not the resolution of the 171 CT scans - I think those are fine. The problem is that we have only 171 FVC curves and corresponding CT scans…</p>\n<p>For next year, if I were in the organization committee, I would certainly recommend a better split of seen/unseen data: imho 15/85 significantly limited the community capacity to generate insights from the available data.</p>\n<p>If we had 50/50, for example, I believe the community would have created way more innovative approaches, generating far more insightful solutions…</p>",
      "rawMarkdown": "Yes... but I think generating extra slices would produce marginal gains (if any)...\n\nMy thinking was to completely synthesize new patients. The problem is not the resolution of the 171 CT scans - I think those are fine. The problem is that we have only 171 FVC curves and corresponding CT scans...\n\nFor next year, if I were in the organization committee, I would certainly recommend a better split of seen/unseen data: imho 15/85 significantly limited the community capacity to generate insights from the available data.\n\nIf we had 50/50, for example, I believe the community would have created way more innovative approaches, generating far more insightful solutions...",
      "votes": null
    },
    {
      "id": "1038565",
      "postDate": "10/05/2020 21:46:50",
      "content": "<p>Lost my motivation a with this competition as the gains were so small and was hoping someone might break some ground with the CT scans. In the end have just set up a robust cross-validation and ensemble and crossing my fingers!</p>\n<p>The problem as I see it with synthetic data is all we've got is some limited observed data so the benefit could only really come from imposing better priors on the data generation model. I guess it could be that the inductive bias of a GAN for generating CTs could be that prior.</p>\n<p>Hopefully the winner in the end has a nice novelty it would seem a really bad outcome for everyone if it was just a tabular model </p>",
      "rawMarkdown": "Lost my motivation a with this competition as the gains were so small and was hoping someone might break some ground with the CT scans. In the end have just set up a robust cross-validation and ensemble and crossing my fingers!\n\nThe problem as I see it with synthetic data is all we've got is some limited observed data so the benefit could only really come from imposing better priors on the data generation model. I guess it could be that the inductive bias of a GAN for generating CTs could be that prior.\n\nHopefully the winner in the end has a nice novelty it would seem a really bad outcome for everyone if it was just a tabular model",
      "votes": null
    },
    {
      "id": "1038567",
      "postDate": "10/05/2020 21:50:22",
      "content": "<p><a href=\"https://www.kaggle.com/jameschapman19\" target=\"_blank\">@jameschapman19</a> <br>\ni am pretty sure that at least some information is indeed in the scans =) </p>",
      "rawMarkdown": "jameschapman19 \ni am pretty sure that at least some information is indeed in the scans =)",
      "votes": null
    },
    {
      "id": "1039131",
      "postDate": "10/06/2020 11:00:59",
      "content": "<p>I have the feeling that they made it on purpose. Reducing the dataset by taking a biased/unbalanced dataset would force us to have the best denoising methods and find way to create new data. It could then generalize better. As you say, medical data are usually unbalanced. They might have the same problem even with their complete dataset. By forcing us to search for those methods, they can afterwards use it on the full dataset.<br>\nI'm still a beginner, so I might be completely wrong.</p>\n<p>Btw, I looked at your notebook <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>. They're great !</p>",
      "rawMarkdown": "I have the feeling that they made it on purpose. Reducing the dataset by taking a biased/unbalanced dataset would force us to have the best denoising methods and find way to create new data. It could then generalize better. As you say, medical data are usually unbalanced. They might have the same problem even with their complete dataset. By forcing us to search for those methods, they can afterwards use it on the full dataset.\nI'm still a beginner, so I might be completely wrong.\n\nBtw, I looked at your notebook @carlossouza. They're great !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1038388,
      "author_name": "sandorkonya",
      "author_url": "",
      "post_date": "10/05/2020 19:06:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> ,</p>\n<p>i have been thinking about this to, but in this case it would only make sense to generate extra slices for the low slice-volume scans (like missing slices inbetween) to reduce a 5 mm  slice distance to 2.5 and so on.</p>\n<p>There are already plenty of CT thorax scans available, and one high slice count scan could be used multiple times with multiple gaps to train a GAN : take n-k-th slice and n+k-th slice to predict n-th slice… the lung and vessels, even mediastinum with heart could be predicted quite good i think, however the problem is with the fibrosis… i am not sure how well they would get.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1038401,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "10/05/2020 19:15:32",
          "content": "<p>Yes… but I think generating extra slices would produce marginal gains (if any)…</p>\n<p>My thinking was to completely synthesize new patients. The problem is not the resolution of the 171 CT scans - I think those are fine. The problem is that we have only 171 FVC curves and corresponding CT scans…</p>\n<p>For next year, if I were in the organization committee, I would certainly recommend a better split of seen/unseen data: imho 15/85 significantly limited the community capacity to generate insights from the available data.</p>\n<p>If we had 50/50, for example, I believe the community would have created way more innovative approaches, generating far more insightful solutions…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1039131,
          "author_name": "yohannwattiez",
          "author_url": "",
          "post_date": "10/06/2020 11:00:59",
          "content": "<p>I have the feeling that they made it on purpose. Reducing the dataset by taking a biased/unbalanced dataset would force us to have the best denoising methods and find way to create new data. It could then generalize better. As you say, medical data are usually unbalanced. They might have the same problem even with their complete dataset. By forcing us to search for those methods, they can afterwards use it on the full dataset.<br>\nI'm still a beginner, so I might be completely wrong.</p>\n<p>Btw, I looked at your notebook <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>. They're great !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1038565,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "10/05/2020 21:46:50",
      "content": "<p>Lost my motivation a with this competition as the gains were so small and was hoping someone might break some ground with the CT scans. In the end have just set up a robust cross-validation and ensemble and crossing my fingers!</p>\n<p>The problem as I see it with synthetic data is all we've got is some limited observed data so the benefit could only really come from imposing better priors on the data generation model. I guess it could be that the inductive bias of a GAN for generating CTs could be that prior.</p>\n<p>Hopefully the winner in the end has a nice novelty it would seem a really bad outcome for everyone if it was just a tabular model </p>",
      "votes": null,
      "replies": [
        {
          "id": 1038567,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "10/05/2020 21:50:22",
          "content": "<p><a href=\"https://www.kaggle.com/jameschapman19\" target=\"_blank\">@jameschapman19</a> <br>\ni am pretty sure that at least some information is indeed in the scans =) </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1038267": "Hey guys,\nAs our team is already reflecting on what worked & didn't work, we discussed one avenue unfortunately we didn't have enough time to pursue: **data augmentation via generation of synthetic CT scans & tabular data**. (There is a plethora of articles about this topic, because medical datasets are generally too expensive and too unbalanced.)\n\nOut of curiosity: has anyone tried anything in this direction? Did it work?\n\nCheers!",
    "1038388": "Hi @carlossouza ,\n\ni have been thinking about this to, but in this case it would only make sense to generate extra slices for the low slice-volume scans (like missing slices inbetween) to reduce a 5 mm  slice distance to 2.5 and so on.\n\nThere are already plenty of CT thorax scans available, and one high slice count scan could be used multiple times with multiple gaps to train a GAN : take n-k-th slice and n+k-th slice to predict n-th slice... the lung and vessels, even mediastinum with heart could be predicted quite good i think, however the problem is with the fibrosis... i am not sure how well they would get.",
    "1038401": "Yes... but I think generating extra slices would produce marginal gains (if any)...\n\nMy thinking was to completely synthesize new patients. The problem is not the resolution of the 171 CT scans - I think those are fine. The problem is that we have only 171 FVC curves and corresponding CT scans...\n\nFor next year, if I were in the organization committee, I would certainly recommend a better split of seen/unseen data: imho 15/85 significantly limited the community capacity to generate insights from the available data.\n\nIf we had 50/50, for example, I believe the community would have created way more innovative approaches, generating far more insightful solutions...",
    "1038565": "Lost my motivation a with this competition as the gains were so small and was hoping someone might break some ground with the CT scans. In the end have just set up a robust cross-validation and ensemble and crossing my fingers!\n\nThe problem as I see it with synthetic data is all we've got is some limited observed data so the benefit could only really come from imposing better priors on the data generation model. I guess it could be that the inductive bias of a GAN for generating CTs could be that prior.\n\nHopefully the winner in the end has a nice novelty it would seem a really bad outcome for everyone if it was just a tabular model",
    "1038567": "jameschapman19 \ni am pretty sure that at least some information is indeed in the scans =)",
    "1039131": "I have the feeling that they made it on purpose. Reducing the dataset by taking a biased/unbalanced dataset would force us to have the best denoising methods and find way to create new data. It could then generalize better. As you say, medical data are usually unbalanced. They might have the same problem even with their complete dataset. By forcing us to search for those methods, they can afterwards use it on the full dataset.\nI'm still a beginner, so I might be completely wrong.\n\nBtw, I looked at your notebook @carlossouza. They're great !"
  },
  "source": "meta"
}