{
  "id": 180091,
  "title": "A Begginer's Question",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/180091",
  "author_name": "",
  "post_date": "2020-09-03T20:07:53.896616400Z",
  "votes": 1,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Can somebody explain me about how image data can be used along with csv to make predictions. Basically how to make use of image data here?</p>",
  "messages": [
    {
      "id": "997197",
      "postDate": "09/03/2020 20:07:53",
      "content": "<p>Can somebody explain me about how image data can be used along with csv to make predictions. Basically how to make use of image data here?</p>",
      "rawMarkdown": "Can somebody explain me about how image data can be used along with csv to make predictions. Basically how to make use of image data here?",
      "votes": null
    },
    {
      "id": "997251",
      "postDate": "09/03/2020 21:11:21",
      "content": "<p>Hello! This competition is unusual in some degree, cause we have a tabular data and some type of image data simultaneously. There are lots of great participants used tabular data only (almost all inside topic <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177045\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177045</a> with high CV and LB scores) and that's good way to start: dataset's length absolutely small so we can check the hypothesis fast enough without any extra resources. As an improvement we can use \"image data\" either conventional way (ResNet, EfficientNet, Inception and so on) or for beneficial features extraction (AE, VAE, middle CNN's feature maps).</p>",
      "rawMarkdown": "Hello! This competition is unusual in some degree, cause we have a tabular data and some type of image data simultaneously. There are lots of great participants used tabular data only (almost all inside topic https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177045 with high CV and LB scores) and that's good way to start: dataset's length absolutely small so we can check the hypothesis fast enough without any extra resources. As an improvement we can use \"image data\" either conventional way (ResNet, EfficientNet, Inception and so on) or for beneficial features extraction (AE, VAE, middle CNN's feature maps).",
      "votes": null
    },
    {
      "id": "997301",
      "postDate": "09/03/2020 23:15:23",
      "content": "<p>Just to add to this, people have used autoencoders to reduce the dimension of the dicom images into say 10 and added that along with the regular values from csv to create a combined model.</p>",
      "rawMarkdown": "Just to add to this, people have used autoencoders to reduce the dimension of the dicom images into say 10 and added that along with the regular values from csv to create a combined model.",
      "votes": null
    },
    {
      "id": "997426",
      "postDate": "09/04/2020 03:36:29",
      "content": "<p>Yeah, right!</p>",
      "rawMarkdown": "Yeah, right!",
      "votes": null
    },
    {
      "id": "997488",
      "postDate": "09/04/2020 04:37:10",
      "content": "<p>this is THE question :)</p>",
      "rawMarkdown": "this is THE question :)",
      "votes": null
    },
    {
      "id": "997515",
      "postDate": "09/04/2020 05:11:24",
      "content": "<p>Thanks for it. Really helpful.</p>",
      "rawMarkdown": "Thanks for it. Really helpful.",
      "votes": null
    },
    {
      "id": "997518",
      "postDate": "09/04/2020 05:12:43",
      "content": "<p><a href=\"https://www.kaggle.com/Jony\" target=\"_blank\">@Jony</a> Karki Thanks for adding on. Good luck for the competition.</p>",
      "rawMarkdown": "Jony Karki Thanks for adding on. Good luck for the competition.",
      "votes": null
    },
    {
      "id": "997842",
      "postDate": "09/04/2020 09:26:50",
      "content": "<p>Also I would add that when you are using the ResNet, EfficientNet approach you can include meta data by concatenating it with the output of the convolutional network of choice</p>",
      "rawMarkdown": "Also I would add that when you are using the ResNet, EfficientNet approach you can include meta data by concatenating it with the output of the convolutional network of choice",
      "votes": null
    },
    {
      "id": "997932",
      "postDate": "09/04/2020 10:51:55",
      "content": "<p>There are tons  of stuff you can do with image data!! I would suggest to go through the openCV docs if you are just starting out with images. Rest, if you need assistance with dicom, as this competition is based on images in that format; visit this kaggle thread : </p>\n<p><strong><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177285\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177285</a></strong></p>\n<p>~May the force be with you!</p>",
      "rawMarkdown": "There are tons  of stuff you can do with image data!! I would suggest to go through the openCV docs if you are just starting out with images. Rest, if you need assistance with dicom, as this competition is based on images in that format; visit this kaggle thread : \n\n**https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177285**\n\n~May the force be with you!",
      "votes": null
    },
    {
      "id": "998068",
      "postDate": "09/04/2020 13:21:36",
      "content": "<p>Thanks for the help :)</p>",
      "rawMarkdown": "Thanks for the help :)",
      "votes": null
    },
    {
      "id": "998280",
      "postDate": "09/04/2020 16:14:29",
      "content": "<p>yeah. I was following your notebook on autoencoders and just wanted to ask does it work or not?.</p>\n<p>I tried using a ConvNet and appending the ConvNet's output with the meta data (similar to what <a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a> is saying above), but it didn't seem to perform well for the validation data.</p>",
      "rawMarkdown": "yeah. I was following your notebook on autoencoders and just wanted to ask does it work or not?.\n\nI tried using a ConvNet and appending the ConvNet's output with the meta data (similar to what @samklein is saying above), but it didn't seem to perform well for the validation data.",
      "votes": null
    },
    {
      "id": "998317",
      "postDate": "09/04/2020 16:44:49",
      "content": "<p>I'm experimenting with VAEs, they are way better than vanilla AEs.. the math is a little harder, that's why I didn't publish anything public yet… but will do soon!</p>",
      "rawMarkdown": "I'm experimenting with VAEs, they are way better than vanilla AEs.. the math is a little harder, that's why I didn't publish anything public yet... but will do soon!",
      "votes": null
    },
    {
      "id": "998622",
      "postDate": "09/04/2020 22:12:57",
      "content": "<p>I would be interested to know about your attempt <a href=\"https://www.kaggle.com/jonykarki\" target=\"_blank\">@jonykarki</a> I haven't had the time to properly set anything up yet but I should today. It sounds like you were overfitting? Or could nothing be learned from the images?</p>",
      "rawMarkdown": "I would be interested to know about your attempt @jonykarki I haven't had the time to properly set anything up yet but I should today. It sounds like you were overfitting? Or could nothing be learned from the images?",
      "votes": null
    },
    {
      "id": "998712",
      "postDate": "09/05/2020 01:58:56",
      "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a> from what I've seen till now I feel like my model isn't really learning anything from the images rather than overfitting. I think it could be because of how I'm loading the images. I'm averaging the images to turn them all into [8, 256, 256] and using convnet on them. So, that could be the problem, but I haven't figured it out yet.<br>\nThis is how I'm averaging them:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1887174%2F0931ae9ef25bc37e4a719c474a4af518%2Finbox_1887174_434cb4a45e2eff9ed40cb13d06871ec9_fsafsed.png?generation=1599271109483087&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "samklein from what I've seen till now I feel like my model isn't really learning anything from the images rather than overfitting. I think it could be because of how I'm loading the images. I'm averaging the images to turn them all into [8, 256, 256] and using convnet on them. So, that could be the problem, but I haven't figured it out yet.\nThis is how I'm averaging them:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1887174%2F0931ae9ef25bc37e4a719c474a4af518%2Finbox_1887174_434cb4a45e2eff9ed40cb13d06871ec9_fsafsed.png?generation=1599271109483087&alt=media)",
      "votes": null
    },
    {
      "id": "998744",
      "postDate": "09/05/2020 03:08:49",
      "content": "<p>I have seen other kernels where people are learning something with this approach, and so I imagine that the averaging destroys the features. Why were you averaging in the first place? Just out of curiosity</p>",
      "rawMarkdown": "I have seen other kernels where people are learning something with this approach, and so I imagine that the averaging destroys the features. Why were you averaging in the first place? Just out of curiosity",
      "votes": null
    },
    {
      "id": "998825",
      "postDate": "09/05/2020 05:54:12",
      "content": "<p>I was trying to avoid having to use very large matrices. And I thought averaging would be easier and faster than padding or something else. Now that you mention though maybe averaging did destroy the features. Thanks for that insight. I'll probably look for something else to try.</p>",
      "rawMarkdown": "I was trying to avoid having to use very large matrices. And I thought averaging would be easier and faster than padding or something else. Now that you mention though maybe averaging did destroy the features. Thanks for that insight. I'll probably look for something else to try.",
      "votes": null
    },
    {
      "id": "999030",
      "postDate": "09/05/2020 09:25:52",
      "content": "<p>No worries, and good luck!</p>",
      "rawMarkdown": "No worries, and good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 997251,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "09/03/2020 21:11:21",
      "content": "<p>Hello! This competition is unusual in some degree, cause we have a tabular data and some type of image data simultaneously. There are lots of great participants used tabular data only (almost all inside topic <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177045\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177045</a> with high CV and LB scores) and that's good way to start: dataset's length absolutely small so we can check the hypothesis fast enough without any extra resources. As an improvement we can use \"image data\" either conventional way (ResNet, EfficientNet, Inception and so on) or for beneficial features extraction (AE, VAE, middle CNN's feature maps).</p>",
      "votes": null,
      "replies": [
        {
          "id": 997301,
          "author_name": "jonykarki",
          "author_url": "",
          "post_date": "09/03/2020 23:15:23",
          "content": "<p>Just to add to this, people have used autoencoders to reduce the dimension of the dicom images into say 10 and added that along with the regular values from csv to create a combined model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997426,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/04/2020 03:36:29",
          "content": "<p>Yeah, right!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997515,
          "author_name": "taralsarvagod13",
          "author_url": "",
          "post_date": "09/04/2020 05:11:24",
          "content": "<p>Thanks for it. Really helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997518,
          "author_name": "taralsarvagod13",
          "author_url": "",
          "post_date": "09/04/2020 05:12:43",
          "content": "<p><a href=\"https://www.kaggle.com/Jony\" target=\"_blank\">@Jony</a> Karki Thanks for adding on. Good luck for the competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997842,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 09:26:50",
          "content": "<p>Also I would add that when you are using the ResNet, EfficientNet approach you can include meta data by concatenating it with the output of the convolutional network of choice</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997488,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "09/04/2020 04:37:10",
      "content": "<p>this is THE question :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 998280,
          "author_name": "jonykarki",
          "author_url": "",
          "post_date": "09/04/2020 16:14:29",
          "content": "<p>yeah. I was following your notebook on autoencoders and just wanted to ask does it work or not?.</p>\n<p>I tried using a ConvNet and appending the ConvNet's output with the meta data (similar to what <a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a> is saying above), but it didn't seem to perform well for the validation data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998317,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "09/04/2020 16:44:49",
          "content": "<p>I'm experimenting with VAEs, they are way better than vanilla AEs.. the math is a little harder, that's why I didn't publish anything public yet… but will do soon!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998622,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/04/2020 22:12:57",
          "content": "<p>I would be interested to know about your attempt <a href=\"https://www.kaggle.com/jonykarki\" target=\"_blank\">@jonykarki</a> I haven't had the time to properly set anything up yet but I should today. It sounds like you were overfitting? Or could nothing be learned from the images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998712,
          "author_name": "jonykarki",
          "author_url": "",
          "post_date": "09/05/2020 01:58:56",
          "content": "<p><a href=\"https://www.kaggle.com/samklein\" target=\"_blank\">@samklein</a> from what I've seen till now I feel like my model isn't really learning anything from the images rather than overfitting. I think it could be because of how I'm loading the images. I'm averaging the images to turn them all into [8, 256, 256] and using convnet on them. So, that could be the problem, but I haven't figured it out yet.<br>\nThis is how I'm averaging them:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1887174%2F0931ae9ef25bc37e4a719c474a4af518%2Finbox_1887174_434cb4a45e2eff9ed40cb13d06871ec9_fsafsed.png?generation=1599271109483087&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998744,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/05/2020 03:08:49",
          "content": "<p>I have seen other kernels where people are learning something with this approach, and so I imagine that the averaging destroys the features. Why were you averaging in the first place? Just out of curiosity</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998825,
          "author_name": "jonykarki",
          "author_url": "",
          "post_date": "09/05/2020 05:54:12",
          "content": "<p>I was trying to avoid having to use very large matrices. And I thought averaging would be easier and faster than padding or something else. Now that you mention though maybe averaging did destroy the features. Thanks for that insight. I'll probably look for something else to try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 999030,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "09/05/2020 09:25:52",
          "content": "<p>No worries, and good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997932,
      "author_name": "fireheart7",
      "author_url": "",
      "post_date": "09/04/2020 10:51:55",
      "content": "<p>There are tons  of stuff you can do with image data!! I would suggest to go through the openCV docs if you are just starting out with images. Rest, if you need assistance with dicom, as this competition is based on images in that format; visit this kaggle thread : </p>\n<p><strong><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177285\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177285</a></strong></p>\n<p>~May the force be with you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 998068,
          "author_name": "taralsarvagod13",
          "author_url": "",
          "post_date": "09/04/2020 13:21:36",
          "content": "<p>Thanks for the help :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "997197": "Can somebody explain me about how image data can be used along with csv to make predictions. Basically how to make use of image data here?",
    "997251": "Hello! This competition is unusual in some degree, cause we have a tabular data and some type of image data simultaneously. There are lots of great participants used tabular data only (almost all inside topic https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177045 with high CV and LB scores) and that's good way to start: dataset's length absolutely small so we can check the hypothesis fast enough without any extra resources. As an improvement we can use \"image data\" either conventional way (ResNet, EfficientNet, Inception and so on) or for beneficial features extraction (AE, VAE, middle CNN's feature maps).",
    "997301": "Just to add to this, people have used autoencoders to reduce the dimension of the dicom images into say 10 and added that along with the regular values from csv to create a combined model.",
    "997426": "Yeah, right!",
    "997488": "this is THE question :)",
    "997515": "Thanks for it. Really helpful.",
    "997518": "Jony Karki Thanks for adding on. Good luck for the competition.",
    "997842": "Also I would add that when you are using the ResNet, EfficientNet approach you can include meta data by concatenating it with the output of the convolutional network of choice",
    "997932": "There are tons  of stuff you can do with image data!! I would suggest to go through the openCV docs if you are just starting out with images. Rest, if you need assistance with dicom, as this competition is based on images in that format; visit this kaggle thread : \n\n**https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177285**\n\n~May the force be with you!",
    "998068": "Thanks for the help :)",
    "998280": "yeah. I was following your notebook on autoencoders and just wanted to ask does it work or not?.\n\nI tried using a ConvNet and appending the ConvNet's output with the meta data (similar to what @samklein is saying above), but it didn't seem to perform well for the validation data.",
    "998317": "I'm experimenting with VAEs, they are way better than vanilla AEs.. the math is a little harder, that's why I didn't publish anything public yet... but will do soon!",
    "998622": "I would be interested to know about your attempt @jonykarki I haven't had the time to properly set anything up yet but I should today. It sounds like you were overfitting? Or could nothing be learned from the images?",
    "998712": "samklein from what I've seen till now I feel like my model isn't really learning anything from the images rather than overfitting. I think it could be because of how I'm loading the images. I'm averaging the images to turn them all into [8, 256, 256] and using convnet on them. So, that could be the problem, but I haven't figured it out yet.\nThis is how I'm averaging them:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1887174%2F0931ae9ef25bc37e4a719c474a4af518%2Finbox_1887174_434cb4a45e2eff9ed40cb13d06871ec9_fsafsed.png?generation=1599271109483087&alt=media)",
    "998744": "I have seen other kernels where people are learning something with this approach, and so I imagine that the averaging destroys the features. Why were you averaging in the first place? Just out of curiosity",
    "998825": "I was trying to avoid having to use very large matrices. And I thought averaging would be easier and faster than padding or something else. Now that you mention though maybe averaging did destroy the features. Thanks for that insight. I'll probably look for something else to try.",
    "999030": "No worries, and good luck!"
  },
  "source": "meta"
}