{
  "id": 167025,
  "title": "3D Deep Convolutional Auto Encoders",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/167025",
  "author_name": "",
  "post_date": "2020-07-14T23:44:50.482358300Z",
  "votes": 35,
  "comment_count": 16,
  "views": 0,
  "content": "<p>There are lots of great papers describing applications of 3D convolutional auto encoders in solving medical imaging problems. Most of them apply this technology to image segmentation, although there are interesting applications on image generation as well. Whatever the end application is, all auto encoders learn latent features that can represent the data, but at a fraction of its size/dimensions. \nIn this first experiment shown below, my auto encoder is compressing the input to ~4% of its dimensions:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F7e5f708a59341c33cd4fd8d83b3a3f5f%2Fezgif-5-606f3222355e.gif?generation=1594768840412636&amp;alt=media\" alt=\"\"></p>\n\n<p>Unfortunately, it is extremely slow... but it is learning :)</p>\n\n<p>\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that we can train a network on just two volumetric images in order to generalize to a third one.\"</p>\n\n<p>When I first read that in one of the papers, at first I didn't believe... but after seeing it in action, it is actually true. :)</p>\n\n<p>I didn't finish reading them all yet, but there are some good ideas in these articles:\n- <a href=\"https://arxiv.org/abs/1606.06650\">3D U-net: Learning dense volumetric segmentation from sparse annotation</a>\n- <a href=\"https://arxiv.org/abs/1811.07999\">Synthetic Lung Nodule 3D Image Generation Using Autoencoders</a>\n- <a href=\"https://arxiv.org/abs/1702.00288\">Low-Dose CT with a Residual Encoder-Decoder Convolutional Neural Network (RED-CNN)</a>\n- <a href=\"https://arxiv.org/abs/1802.05656\">3-D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2-D Trained Network</a>\n- <a href=\"https://srinjaypaul.github.io/3D_Convolutional_autoencoder_for_brain_volumes/\">3D Convolutional autoencoder for brain volumes – Srinjay Paul – Deep Learning and brain image Analysis</a></p>",
  "messages": [
    {
      "id": "929772",
      "postDate": "07/14/2020 23:44:50",
      "content": "<p>There are lots of great papers describing applications of 3D convolutional auto encoders in solving medical imaging problems. Most of them apply this technology to image segmentation, although there are interesting applications on image generation as well. Whatever the end application is, all auto encoders learn latent features that can represent the data, but at a fraction of its size/dimensions. \nIn this first experiment shown below, my auto encoder is compressing the input to ~4% of its dimensions:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F7e5f708a59341c33cd4fd8d83b3a3f5f%2Fezgif-5-606f3222355e.gif?generation=1594768840412636&amp;alt=media\" alt=\"\"></p>\n\n<p>Unfortunately, it is extremely slow... but it is learning :)</p>\n\n<p>\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that we can train a network on just two volumetric images in order to generalize to a third one.\"</p>\n\n<p>When I first read that in one of the papers, at first I didn't believe... but after seeing it in action, it is actually true. :)</p>\n\n<p>I didn't finish reading them all yet, but there are some good ideas in these articles:\n- <a href=\"https://arxiv.org/abs/1606.06650\">3D U-net: Learning dense volumetric segmentation from sparse annotation</a>\n- <a href=\"https://arxiv.org/abs/1811.07999\">Synthetic Lung Nodule 3D Image Generation Using Autoencoders</a>\n- <a href=\"https://arxiv.org/abs/1702.00288\">Low-Dose CT with a Residual Encoder-Decoder Convolutional Neural Network (RED-CNN)</a>\n- <a href=\"https://arxiv.org/abs/1802.05656\">3-D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2-D Trained Network</a>\n- <a href=\"https://srinjaypaul.github.io/3D_Convolutional_autoencoder_for_brain_volumes/\">3D Convolutional autoencoder for brain volumes – Srinjay Paul – Deep Learning and brain image Analysis</a></p>",
      "rawMarkdown": "There are lots of great papers describing applications of 3D convolutional auto encoders in solving medical imaging problems. Most of them apply this technology to image segmentation, although there are interesting applications on image generation as well. Whatever the end application is, all auto encoders learn latent features that can represent the data, but at a fraction of its size/dimensions. \nIn this first experiment shown below, my auto encoder is compressing the input to ~4% of its dimensions:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F7e5f708a59341c33cd4fd8d83b3a3f5f%2Fezgif-5-606f3222355e.gif?generation=1594768840412636&amp;alt=media)\n\nUnfortunately, it is extremely slow... but it is learning :)\n\n\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that we can train a network on just two volumetric images in order to generalize to a third one.\"\n\nWhen I first read that in one of the papers, at first I didn't believe... but after seeing it in action, it is actually true. :)\n\nI didn't finish reading them all yet, but there are some good ideas in these articles:\n- [3D U-net: Learning dense volumetric segmentation from sparse annotation](https://arxiv.org/abs/1606.06650)\n- [Synthetic Lung Nodule 3D Image Generation Using Autoencoders](https://arxiv.org/abs/1811.07999)\n- [Low-Dose CT with a Residual Encoder-Decoder Convolutional Neural Network (RED-CNN)](https://arxiv.org/abs/1702.00288)\n- [3-D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2-D Trained Network](https://arxiv.org/abs/1802.05656)\n- [3D Convolutional autoencoder for brain volumes – Srinjay Paul – Deep Learning and brain image Analysis](https://srinjaypaul.github.io/3D_Convolutional_autoencoder_for_brain_volumes/)",
      "votes": null
    },
    {
      "id": "929793",
      "postDate": "07/15/2020 00:38:51",
      "content": "<p>I suppose the tricky thing is reducing the dimensionality while keeping the variance associated with FVC</p>",
      "rawMarkdown": "I suppose the tricky thing is reducing the dimensionality while keeping the variance associated with FVC",
      "votes": null
    },
    {
      "id": "930349",
      "postDate": "07/15/2020 11:43:40",
      "content": "<p>Thnks <a href=\"/carlossouza\">@carlossouza</a> . Do you know how to combine slices to get 3D images? I hope a yes,  since you trained that kind of auto-encoder.</p>",
      "rawMarkdown": "Thnks @carlossouza . Do you know how to combine slices to get 3D images? I hope a yes,  since you trained that kind of auto-encoder.",
      "votes": null
    },
    {
      "id": "930426",
      "postDate": "07/15/2020 13:01:39",
      "content": "<p>Just stack them! My tensors had 4 dimensions: C (channel) x D (slice) x H (height) x W (width). To pass them through the network, you have to add the 5th, N (batch): N x C x D x H x W.</p>\n\n<p>The key is how to pre-process these images, to avoid garbage-in-garbage-out. In this first experiment I did very little pre-processing, as you can see in the images (i.e. the chamber &amp; table are visible). Next step is to mask them out. Another area that requires more experimentation is how to ensure all tensors have the same shape. This can be achieved by different methods of resizing, interpolation, cropping, etc.</p>",
      "rawMarkdown": "Just stack them! My tensors had 4 dimensions: C (channel) x D (slice) x H (height) x W (width). To pass them through the network, you have to add the 5th, N (batch): N x C x D x H x W.\n\nThe key is how to pre-process these images, to avoid garbage-in-garbage-out. In this first experiment I did very little pre-processing, as you can see in the images (i.e. the chamber &amp; table are visible). Next step is to mask them out. Another area that requires more experimentation is how to ensure all tensors have the same shape. This can be achieved by different methods of resizing, interpolation, cropping, etc.",
      "votes": null
    },
    {
      "id": "930692",
      "postDate": "07/15/2020 16:42:49",
      "content": "<p>Hello Carlos. How would you use these latent features as input in this competition? Combine them in some way ( I have no idea how at the moment ) into a unique Dataset or use them at some point in a Pipeline. Thanks</p>",
      "rawMarkdown": "Hello Carlos. How would you use these latent features as input in this competition? Combine them in some way ( I have no idea how at the moment ) into a unique Dataset or use them at some point in a Pipeline. Thanks",
      "votes": null
    },
    {
      "id": "931003",
      "postDate": "07/15/2020 22:11:40",
      "content": "<p>Hi <a href=\"/ronaldokun\">@ronaldokun</a> ! The idea is to train a model exactly as all great starter models people already published, which basically uses only the tabular data, but additionally feeding these latent features. </p>",
      "rawMarkdown": "Hi @ronaldokun ! The idea is to train a model exactly as all great starter models people already published, which basically uses only the tabular data, but additionally feeding these latent features.",
      "votes": null
    },
    {
      "id": "931053",
      "postDate": "07/15/2020 23:43:10",
      "content": "<p>Hello Carlos, thank you for your reply. That's great 🤓, the process is actually simpler than what I was thinking ( the logic I mean ). </p>\n\n<p>Great insight!</p>",
      "rawMarkdown": "Hello Carlos, thank you for your reply. That's great 🤓, the process is actually simpler than what I was thinking ( the logic I mean ). \n\nGreat insight!",
      "votes": null
    },
    {
      "id": "932138",
      "postDate": "07/16/2020 18:25:08",
      "content": "<p>Update: now using properly masked lungs:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F08729db2260bf57e489dc3b1eb76da2a%2Fsample.gif?generation=1594923885543748&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F2b9e75fcdc7b3faa713afb8b1b88eb71%2Foutput.gif?generation=1594923900605055&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Update: now using properly masked lungs:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F08729db2260bf57e489dc3b1eb76da2a%2Fsample.gif?generation=1594923885543748&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F2b9e75fcdc7b3faa713afb8b1b88eb71%2Foutput.gif?generation=1594923900605055&amp;alt=media)",
      "votes": null
    },
    {
      "id": "932978",
      "postDate": "07/17/2020 12:04:22",
      "content": "<p>This is pretty cool <a href=\"/carlossouza\">@carlossouza</a> ! A question that comes to my mind is whether this model is able to reconstruct the features that are relevant for pulmonary fibrosis though or if one could somehow make it care stronger for the right features, e.g. by a bimodal (or whatever you call it) model structure:\n1. encoder\n2.1 decoder (reconstruction task)\n2.2 FVC prediction from the latent features (maybe simply for week 0, i.e. the week from that the scan itself is aswell); just to make it focus more on features that are important for FVC prediction</p>\n\n<p>And train this as one model.</p>\n\n<p>Then proceed e.g. as you said below, just including the latent features in a model working with tabular data.</p>",
      "rawMarkdown": "This is pretty cool @carlossouza ! A question that comes to my mind is whether this model is able to reconstruct the features that are relevant for pulmonary fibrosis though or if one could somehow make it care stronger for the right features, e.g. by a bimodal (or whatever you call it) model structure:\n1. encoder\n2.1 decoder (reconstruction task)\n2.2 FVC prediction from the latent features (maybe simply for week 0, i.e. the week from that the scan itself is aswell); just to make it focus more on features that are important for FVC prediction\n\nAnd train this as one model.\n\nThen proceed e.g. as you said below, just including the latent features in a model working with tabular data.",
      "votes": null
    },
    {
      "id": "933274",
      "postDate": "07/17/2020 16:02:08",
      "content": "<p>Hi Carlos, if you get the result vector to concat in the tabular data, it is only going to add data to the week 0. Any idea how to deal with the other weeks?</p>",
      "rawMarkdown": "Hi Carlos, if you get the result vector to concat in the tabular data, it is only going to add data to the week 0. Any idea how to deal with the other weeks?",
      "votes": null
    },
    {
      "id": "944749",
      "postDate": "07/25/2020 10:12:01",
      "content": "<p>Would be interesting to use a variational autoencoder to generate images for the remaining weeks using FVC in the loss function. Although without additional external data the results might not be that great.</p>",
      "rawMarkdown": "Would be interesting to use a variational autoencoder to generate images for the remaining weeks using FVC in the loss function. Although without additional external data the results might not be that great.",
      "votes": null
    },
    {
      "id": "948739",
      "postDate": "07/28/2020 07:20:48",
      "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a> thanks for your work, it could be the ground-breaking solution that everyone's waiting.\nYou mentioned that training is slow, did you try to use JPEG images instead( both training and inference) ? </p>",
      "rawMarkdown": "Hi @carlossouza thanks for your work, it could be the ground-breaking solution that everyone's waiting.\nYou mentioned that training is slow, did you try to use JPEG images instead( both training and inference) ?",
      "votes": null
    },
    {
      "id": "960656",
      "postDate": "08/06/2020 15:19:01",
      "content": "<p>Cool thread and work <a href=\"/carlossouza\">@carlossouza</a> ! </p>\n\n<p>What specific autoencoder (AE) are you implementing within this thread, and do you plan to play with other types of AEs (e.g., denoising, VAE, etc.)? </p>\n\n<p>What do you expect the latent space to learn, and how do you plan to implement these potentially useful features for future analyzes? </p>",
      "rawMarkdown": "Cool thread and work @carlossouza ! \n\nWhat specific autoencoder (AE) are you implementing within this thread, and do you plan to play with other types of AEs (e.g., denoising, VAE, etc.)? \n\nWhat do you expect the latent space to learn, and how do you plan to implement these potentially useful features for future analyzes?",
      "votes": null
    },
    {
      "id": "960970",
      "postDate": "08/06/2020 20:28:31",
      "content": "<p>Hi <a href=\"/niksapraljak\">@niksapraljak</a> ! </p>\n\n<p>My objective with this competition was to learn how to implement auto-encoders. After I implemented a very basic vanilla auto-encoder, I observed the latent space was very irregular... Then, I started reading all seminal papers on auto-encoders, especially variational auto-encoders... Since then, I learned a lot about VAEs. There are some great papers about the theory (the math is not trivial) and applications (I'm especially fond of applications to recommender systems).</p>\n\n<p>I even managed to perfectly reproduce a cool paper: <a href=\"https://www.semanticscholar.org/paper/Partial-VAE-for-Hybrid-Recommender-System-Ma-Gong/e81b4bc889e99ad3ebd588eef12a0216a015d7bf\">Partial VAE for Hybrid Recommender System</a>. The <a href=\"https://www.kaggle.com/carlossouza/partial-vae-for-hybrid-recommender-system\">notebook with the code is here</a>. The application is very cool: predicting movie movie ratings in very sparse datasets (genetics has drosophilas, computer vision has MNIST, variational inference has MovieLens :))</p>\n\n<p>Anyway, I'm experimenting with VAEs in the past 2 weeks... VAEs are hard. More hyperparameters, math is not trivial, if you don't train it properly the model won't learn or won't generalize well... And as it is with ML in general, when there is a bug in yout code, the model will fail silently, leaving you clueless about where is the error and how to debug. In these situations, IMHO it is better to learn on a smaller dataset, in which I can experiment faster, than in a large dataset (such as the one in this competition). Baby steps :)</p>\n\n<p>That wouldn't be the case if I actually had managed to run experiments on TPUs, but running auto-encoders on TPUs didn't work. To my surprise, the code was right, and the problem was with Pytorch XLA library (accordingly to the team developing the library this was the 1st time someone tried using 3D images in a network on TPUs... how crazy was that :)). They are investigating the issue.</p>\n\n<p>In the mean time, I continue learning VAEs, but on smaller datasets such as this MovieLens example, and on a NeurIPS competition: <a href=\"https://www.microsoft.com/en-us/research/academic-program/diagnostic-questions/\">Diagnostic Questions: Predicting Student Responses and Measuring Question Quality</a>. Variational inference is really cool! :)</p>",
      "rawMarkdown": "Hi @niksapraljak ! \n\nMy objective with this competition was to learn how to implement auto-encoders. After I implemented a very basic vanilla auto-encoder, I observed the latent space was very irregular... Then, I started reading all seminal papers on auto-encoders, especially variational auto-encoders... Since then, I learned a lot about VAEs. There are some great papers about the theory (the math is not trivial) and applications (I'm especially fond of applications to recommender systems).\n\nI even managed to perfectly reproduce a cool paper: [Partial VAE for Hybrid Recommender System](https://www.semanticscholar.org/paper/Partial-VAE-for-Hybrid-Recommender-System-Ma-Gong/e81b4bc889e99ad3ebd588eef12a0216a015d7bf). The [notebook with the code is here](https://www.kaggle.com/carlossouza/partial-vae-for-hybrid-recommender-system). The application is very cool: predicting movie movie ratings in very sparse datasets (genetics has drosophilas, computer vision has MNIST, variational inference has MovieLens :))\n\nAnyway, I'm experimenting with VAEs in the past 2 weeks... VAEs are hard. More hyperparameters, math is not trivial, if you don't train it properly the model won't learn or won't generalize well... And as it is with ML in general, when there is a bug in yout code, the model will fail silently, leaving you clueless about where is the error and how to debug. In these situations, IMHO it is better to learn on a smaller dataset, in which I can experiment faster, than in a large dataset (such as the one in this competition). Baby steps :)\n\nThat wouldn't be the case if I actually had managed to run experiments on TPUs, but running auto-encoders on TPUs didn't work. To my surprise, the code was right, and the problem was with Pytorch XLA library (accordingly to the team developing the library this was the 1st time someone tried using 3D images in a network on TPUs... how crazy was that :)). They are investigating the issue.\n\nIn the mean time, I continue learning VAEs, but on smaller datasets such as this MovieLens example, and on a NeurIPS competition: [Diagnostic Questions: Predicting Student Responses and Measuring Question Quality](https://www.microsoft.com/en-us/research/academic-program/diagnostic-questions/). Variational inference is really cool! :)",
      "votes": null
    },
    {
      "id": "961236",
      "postDate": "08/07/2020 03:07:06",
      "content": "<p>Oh cool, I would be very interested also to hear what the Kaggle team has to say about using 3D images in a network on TPUs. </p>\n<p>VAEs are quite a little bit more challenging because the latent space is constructed as a normal distribution, which is assumed to generate the data. Lastly, I hope to stumble upon your future notebook that involves VAE+TPUs! </p>",
      "rawMarkdown": "Oh cool, I would be very interested also to hear what the Kaggle team has to say about using 3D images in a network on TPUs. \n\nVAEs are quite a little bit more challenging because the latent space is constructed as a normal distribution, which is assumed to generate the data. Lastly, I hope to stumble upon your future notebook that involves VAE+TPUs!",
      "votes": null
    },
    {
      "id": "983982",
      "postDate": "08/24/2020 18:30:38",
      "content": "<p>After convolutional operations, a deep learning model should be able to find the degree of fibrosis at Week 0 and correlate it with the time evolution of the FVC.</p>",
      "rawMarkdown": "After convolutional operations, a deep learning model should be able to find the degree of fibrosis at Week 0 and correlate it with the time evolution of the FVC.",
      "votes": null
    },
    {
      "id": "1640007",
      "postDate": "01/06/2022 05:40:27",
      "content": "<p>sorry<br>\nit's a good work</p>",
      "rawMarkdown": "sorry\nit's a good work",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 933274,
      "author_name": "cafalchio",
      "author_url": "",
      "post_date": "07/17/2020 16:02:08",
      "content": "<p>Hi Carlos, if you get the result vector to concat in the tabular data, it is only going to add data to the week 0. Any idea how to deal with the other weeks?</p>",
      "votes": null,
      "replies": [
        {
          "id": 944749,
          "author_name": "metathesis",
          "author_url": "",
          "post_date": "07/25/2020 10:12:01",
          "content": "<p>Would be interesting to use a variational autoencoder to generate images for the remaining weeks using FVC in the loss function. Although without additional external data the results might not be that great.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 983982,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/24/2020 18:30:38",
          "content": "<p>After convolutional operations, a deep learning model should be able to find the degree of fibrosis at Week 0 and correlate it with the time evolution of the FVC.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1640007,
      "author_name": "ilmechaju",
      "author_url": "",
      "post_date": "01/06/2022 05:40:27",
      "content": "<p>sorry<br>\nit's a good work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 929793,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "07/15/2020 00:38:51",
      "content": "<p>I suppose the tricky thing is reducing the dimensionality while keeping the variance associated with FVC</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 930349,
      "author_name": "ulrich07",
      "author_url": "",
      "post_date": "07/15/2020 11:43:40",
      "content": "<p>Thnks <a href=\"/carlossouza\">@carlossouza</a> . Do you know how to combine slices to get 3D images? I hope a yes,  since you trained that kind of auto-encoder.</p>",
      "votes": null,
      "replies": [
        {
          "id": 930426,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "07/15/2020 13:01:39",
          "content": "<p>Just stack them! My tensors had 4 dimensions: C (channel) x D (slice) x H (height) x W (width). To pass them through the network, you have to add the 5th, N (batch): N x C x D x H x W.</p>\n\n<p>The key is how to pre-process these images, to avoid garbage-in-garbage-out. In this first experiment I did very little pre-processing, as you can see in the images (i.e. the chamber &amp; table are visible). Next step is to mask them out. Another area that requires more experimentation is how to ensure all tensors have the same shape. This can be achieved by different methods of resizing, interpolation, cropping, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930692,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "07/15/2020 16:42:49",
          "content": "<p>Hello Carlos. How would you use these latent features as input in this competition? Combine them in some way ( I have no idea how at the moment ) into a unique Dataset or use them at some point in a Pipeline. Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931003,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "07/15/2020 22:11:40",
          "content": "<p>Hi <a href=\"/ronaldokun\">@ronaldokun</a> ! The idea is to train a model exactly as all great starter models people already published, which basically uses only the tabular data, but additionally feeding these latent features. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931053,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "07/15/2020 23:43:10",
          "content": "<p>Hello Carlos, thank you for your reply. That's great 🤓, the process is actually simpler than what I was thinking ( the logic I mean ). </p>\n\n<p>Great insight!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 932138,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "07/16/2020 18:25:08",
      "content": "<p>Update: now using properly masked lungs:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F08729db2260bf57e489dc3b1eb76da2a%2Fsample.gif?generation=1594923885543748&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F2b9e75fcdc7b3faa713afb8b1b88eb71%2Foutput.gif?generation=1594923900605055&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 932978,
      "author_name": "tobiit",
      "author_url": "",
      "post_date": "07/17/2020 12:04:22",
      "content": "<p>This is pretty cool <a href=\"/carlossouza\">@carlossouza</a> ! A question that comes to my mind is whether this model is able to reconstruct the features that are relevant for pulmonary fibrosis though or if one could somehow make it care stronger for the right features, e.g. by a bimodal (or whatever you call it) model structure:\n1. encoder\n2.1 decoder (reconstruction task)\n2.2 FVC prediction from the latent features (maybe simply for week 0, i.e. the week from that the scan itself is aswell); just to make it focus more on features that are important for FVC prediction</p>\n\n<p>And train this as one model.</p>\n\n<p>Then proceed e.g. as you said below, just including the latent features in a model working with tabular data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 948739,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "07/28/2020 07:20:48",
      "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a> thanks for your work, it could be the ground-breaking solution that everyone's waiting.\nYou mentioned that training is slow, did you try to use JPEG images instead( both training and inference) ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 960656,
      "author_name": "niksapraljak",
      "author_url": "",
      "post_date": "08/06/2020 15:19:01",
      "content": "<p>Cool thread and work <a href=\"/carlossouza\">@carlossouza</a> ! </p>\n\n<p>What specific autoencoder (AE) are you implementing within this thread, and do you plan to play with other types of AEs (e.g., denoising, VAE, etc.)? </p>\n\n<p>What do you expect the latent space to learn, and how do you plan to implement these potentially useful features for future analyzes? </p>",
      "votes": null,
      "replies": [
        {
          "id": 960970,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "08/06/2020 20:28:31",
          "content": "<p>Hi <a href=\"/niksapraljak\">@niksapraljak</a> ! </p>\n\n<p>My objective with this competition was to learn how to implement auto-encoders. After I implemented a very basic vanilla auto-encoder, I observed the latent space was very irregular... Then, I started reading all seminal papers on auto-encoders, especially variational auto-encoders... Since then, I learned a lot about VAEs. There are some great papers about the theory (the math is not trivial) and applications (I'm especially fond of applications to recommender systems).</p>\n\n<p>I even managed to perfectly reproduce a cool paper: <a href=\"https://www.semanticscholar.org/paper/Partial-VAE-for-Hybrid-Recommender-System-Ma-Gong/e81b4bc889e99ad3ebd588eef12a0216a015d7bf\">Partial VAE for Hybrid Recommender System</a>. The <a href=\"https://www.kaggle.com/carlossouza/partial-vae-for-hybrid-recommender-system\">notebook with the code is here</a>. The application is very cool: predicting movie movie ratings in very sparse datasets (genetics has drosophilas, computer vision has MNIST, variational inference has MovieLens :))</p>\n\n<p>Anyway, I'm experimenting with VAEs in the past 2 weeks... VAEs are hard. More hyperparameters, math is not trivial, if you don't train it properly the model won't learn or won't generalize well... And as it is with ML in general, when there is a bug in yout code, the model will fail silently, leaving you clueless about where is the error and how to debug. In these situations, IMHO it is better to learn on a smaller dataset, in which I can experiment faster, than in a large dataset (such as the one in this competition). Baby steps :)</p>\n\n<p>That wouldn't be the case if I actually had managed to run experiments on TPUs, but running auto-encoders on TPUs didn't work. To my surprise, the code was right, and the problem was with Pytorch XLA library (accordingly to the team developing the library this was the 1st time someone tried using 3D images in a network on TPUs... how crazy was that :)). They are investigating the issue.</p>\n\n<p>In the mean time, I continue learning VAEs, but on smaller datasets such as this MovieLens example, and on a NeurIPS competition: <a href=\"https://www.microsoft.com/en-us/research/academic-program/diagnostic-questions/\">Diagnostic Questions: Predicting Student Responses and Measuring Question Quality</a>. Variational inference is really cool! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961236,
          "author_name": "niksapraljak",
          "author_url": "",
          "post_date": "08/07/2020 03:07:06",
          "content": "<p>Oh cool, I would be very interested also to hear what the Kaggle team has to say about using 3D images in a network on TPUs. </p>\n<p>VAEs are quite a little bit more challenging because the latent space is constructed as a normal distribution, which is assumed to generate the data. Lastly, I hope to stumble upon your future notebook that involves VAE+TPUs! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "929772": "There are lots of great papers describing applications of 3D convolutional auto encoders in solving medical imaging problems. Most of them apply this technology to image segmentation, although there are interesting applications on image generation as well. Whatever the end application is, all auto encoders learn latent features that can represent the data, but at a fraction of its size/dimensions. \nIn this first experiment shown below, my auto encoder is compressing the input to ~4% of its dimensions:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F7e5f708a59341c33cd4fd8d83b3a3f5f%2Fezgif-5-606f3222355e.gif?generation=1594768840412636&amp;alt=media)\n\nUnfortunately, it is extremely slow... but it is learning :)\n\n\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that we can train a network on just two volumetric images in order to generalize to a third one.\"\n\nWhen I first read that in one of the papers, at first I didn't believe... but after seeing it in action, it is actually true. :)\n\nI didn't finish reading them all yet, but there are some good ideas in these articles:\n- [3D U-net: Learning dense volumetric segmentation from sparse annotation](https://arxiv.org/abs/1606.06650)\n- [Synthetic Lung Nodule 3D Image Generation Using Autoencoders](https://arxiv.org/abs/1811.07999)\n- [Low-Dose CT with a Residual Encoder-Decoder Convolutional Neural Network (RED-CNN)](https://arxiv.org/abs/1702.00288)\n- [3-D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2-D Trained Network](https://arxiv.org/abs/1802.05656)\n- [3D Convolutional autoencoder for brain volumes – Srinjay Paul – Deep Learning and brain image Analysis](https://srinjaypaul.github.io/3D_Convolutional_autoencoder_for_brain_volumes/)",
    "929793": "I suppose the tricky thing is reducing the dimensionality while keeping the variance associated with FVC",
    "930349": "Thnks @carlossouza . Do you know how to combine slices to get 3D images? I hope a yes,  since you trained that kind of auto-encoder.",
    "930426": "Just stack them! My tensors had 4 dimensions: C (channel) x D (slice) x H (height) x W (width). To pass them through the network, you have to add the 5th, N (batch): N x C x D x H x W.\n\nThe key is how to pre-process these images, to avoid garbage-in-garbage-out. In this first experiment I did very little pre-processing, as you can see in the images (i.e. the chamber &amp; table are visible). Next step is to mask them out. Another area that requires more experimentation is how to ensure all tensors have the same shape. This can be achieved by different methods of resizing, interpolation, cropping, etc.",
    "930692": "Hello Carlos. How would you use these latent features as input in this competition? Combine them in some way ( I have no idea how at the moment ) into a unique Dataset or use them at some point in a Pipeline. Thanks",
    "931003": "Hi @ronaldokun ! The idea is to train a model exactly as all great starter models people already published, which basically uses only the tabular data, but additionally feeding these latent features.",
    "931053": "Hello Carlos, thank you for your reply. That's great 🤓, the process is actually simpler than what I was thinking ( the logic I mean ). \n\nGreat insight!",
    "932138": "Update: now using properly masked lungs:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F08729db2260bf57e489dc3b1eb76da2a%2Fsample.gif?generation=1594923885543748&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F2b9e75fcdc7b3faa713afb8b1b88eb71%2Foutput.gif?generation=1594923900605055&amp;alt=media)",
    "932978": "This is pretty cool @carlossouza ! A question that comes to my mind is whether this model is able to reconstruct the features that are relevant for pulmonary fibrosis though or if one could somehow make it care stronger for the right features, e.g. by a bimodal (or whatever you call it) model structure:\n1. encoder\n2.1 decoder (reconstruction task)\n2.2 FVC prediction from the latent features (maybe simply for week 0, i.e. the week from that the scan itself is aswell); just to make it focus more on features that are important for FVC prediction\n\nAnd train this as one model.\n\nThen proceed e.g. as you said below, just including the latent features in a model working with tabular data.",
    "933274": "Hi Carlos, if you get the result vector to concat in the tabular data, it is only going to add data to the week 0. Any idea how to deal with the other weeks?",
    "944749": "Would be interesting to use a variational autoencoder to generate images for the remaining weeks using FVC in the loss function. Although without additional external data the results might not be that great.",
    "948739": "Hi @carlossouza thanks for your work, it could be the ground-breaking solution that everyone's waiting.\nYou mentioned that training is slow, did you try to use JPEG images instead( both training and inference) ?",
    "960656": "Cool thread and work @carlossouza ! \n\nWhat specific autoencoder (AE) are you implementing within this thread, and do you plan to play with other types of AEs (e.g., denoising, VAE, etc.)? \n\nWhat do you expect the latent space to learn, and how do you plan to implement these potentially useful features for future analyzes?",
    "960970": "Hi @niksapraljak ! \n\nMy objective with this competition was to learn how to implement auto-encoders. After I implemented a very basic vanilla auto-encoder, I observed the latent space was very irregular... Then, I started reading all seminal papers on auto-encoders, especially variational auto-encoders... Since then, I learned a lot about VAEs. There are some great papers about the theory (the math is not trivial) and applications (I'm especially fond of applications to recommender systems).\n\nI even managed to perfectly reproduce a cool paper: [Partial VAE for Hybrid Recommender System](https://www.semanticscholar.org/paper/Partial-VAE-for-Hybrid-Recommender-System-Ma-Gong/e81b4bc889e99ad3ebd588eef12a0216a015d7bf). The [notebook with the code is here](https://www.kaggle.com/carlossouza/partial-vae-for-hybrid-recommender-system). The application is very cool: predicting movie movie ratings in very sparse datasets (genetics has drosophilas, computer vision has MNIST, variational inference has MovieLens :))\n\nAnyway, I'm experimenting with VAEs in the past 2 weeks... VAEs are hard. More hyperparameters, math is not trivial, if you don't train it properly the model won't learn or won't generalize well... And as it is with ML in general, when there is a bug in yout code, the model will fail silently, leaving you clueless about where is the error and how to debug. In these situations, IMHO it is better to learn on a smaller dataset, in which I can experiment faster, than in a large dataset (such as the one in this competition). Baby steps :)\n\nThat wouldn't be the case if I actually had managed to run experiments on TPUs, but running auto-encoders on TPUs didn't work. To my surprise, the code was right, and the problem was with Pytorch XLA library (accordingly to the team developing the library this was the 1st time someone tried using 3D images in a network on TPUs... how crazy was that :)). They are investigating the issue.\n\nIn the mean time, I continue learning VAEs, but on smaller datasets such as this MovieLens example, and on a NeurIPS competition: [Diagnostic Questions: Predicting Student Responses and Measuring Question Quality](https://www.microsoft.com/en-us/research/academic-program/diagnostic-questions/). Variational inference is really cool! :)",
    "961236": "Oh cool, I would be very interested also to hear what the Kaggle team has to say about using 3D images in a network on TPUs. \n\nVAEs are quite a little bit more challenging because the latent space is constructed as a normal distribution, which is assumed to generate the data. Lastly, I hope to stumble upon your future notebook that involves VAE+TPUs!",
    "983982": "After convolutional operations, a deep learning model should be able to find the degree of fibrosis at Week 0 and correlate it with the time evolution of the FVC.",
    "1640007": "sorry\nit's a good work"
  },
  "source": "meta"
}