{
  "id": 187055,
  "title": "Approache to use images ( using siamese to find difference as nummerical difference )",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/187055",
  "author_name": "",
  "post_date": "2020-09-27T09:46:30.664851300Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>How we can use images to model in this competition.</p>\n<p>One way I though of how an image of patient is different from other.<br>\nAs images taken of single breathe in-out.</p>\n<p>Approach 1 :<br>\nUsing siamese network it can be found as distance that could be use to as one of quantitative data with other tabular data to use.<br>\nNot to use triplet but just difference of 2 images as in twin siamese.<br>\nLess number of images is not a issue for siamese.</p>\n<p>Approach 2:<br>\nThere is one notebook of autoencoder to generate feature vectors.<br>\nDifference of these vectors one at full breath-in and one at full breath-out could also provide quantitative data to use with tabular.</p>",
  "messages": [
    {
      "id": "1028890",
      "postDate": "09/27/2020 09:46:30",
      "content": "<p>How we can use images to model in this competition.</p>\n<p>One way I though of how an image of patient is different from other.<br>\nAs images taken of single breathe in-out.</p>\n<p>Approach 1 :<br>\nUsing siamese network it can be found as distance that could be use to as one of quantitative data with other tabular data to use.<br>\nNot to use triplet but just difference of 2 images as in twin siamese.<br>\nLess number of images is not a issue for siamese.</p>\n<p>Approach 2:<br>\nThere is one notebook of autoencoder to generate feature vectors.<br>\nDifference of these vectors one at full breath-in and one at full breath-out could also provide quantitative data to use with tabular.</p>",
      "rawMarkdown": "How we can use images to model in this competition.\n\nOne way I though of how an image of patient is different from other.\nAs images taken of single breathe in-out.\n\nApproach 1 :\nUsing siamese network it can be found as distance that could be use to as one of quantitative data with other tabular data to use.\nNot to use triplet but just difference of 2 images as in twin siamese.\nLess number of images is not a issue for siamese.\n\nApproach 2:\nThere is one notebook of autoencoder to generate feature vectors.\nDifference of these vectors one at full breath-in and one at full breath-out could also provide quantitative data to use with tabular.",
      "votes": null
    },
    {
      "id": "1029461",
      "postDate": "09/27/2020 18:44:56",
      "content": "<p>That's a great idea Rajnish. Although my team and many others failed at extracting any useful information from the image at this point, checkout this thread <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077</a> for more info.</p>\n<p>Let us know if you succeed in those attempts.</p>",
      "rawMarkdown": "That's a great idea Rajnish. Although my team and many others failed at extracting any useful information from the image at this point, checkout this thread https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077 for more info.\n\nLet us know if you succeed in those attempts.",
      "votes": null
    },
    {
      "id": "1030042",
      "postDate": "09/28/2020 11:23:57",
      "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> Could not actively follow this competition ( after 2 continuous one ) due to other priority and focused on learning PyTorch and Fast.AI as new papers coming up has more PyTorch implementation and FAST.AI have some cool stuff to accelerate learning and training.</p>\n<p>I could infer from images that image:1 of patient have no visible air and mid image ( if a patient has 30 image then approximately 14,15  or 16 ) has air visibility.<br>\nNow this is the pattern for each patient's image.  If we can find distance of these images considering image-1 as base image ,we can find kind of lung capacity for each patient by either comparing using siamese or encoded vector coming from autoencoder.</p>\n<p>I created one autoencoder here <a href=\"https://www.kaggle.com/rajnishe/rc-siim-autoencoder-model\" target=\"_blank\">https://www.kaggle.com/rajnishe/rc-siim-autoencoder-model</a></p>\n<p>Take output of encoeder for each image save on disk or compare difference at run time and save it.<br>\nOne hyperparameter could be what distance you will use as in KNN .</p>\n<p>To reduce too many features , Gaussian blur will help , I have put that code here<br>\n<a href=\"https://www.kaggle.com/rajnishe/rc-osic-dcm-to-tif-n-jpg-conversion?scriptVersionId=41671497\" target=\"_blank\">https://www.kaggle.com/rajnishe/rc-osic-dcm-to-tif-n-jpg-conversion?scriptVersionId=41671497</a></p>",
      "rawMarkdown": "ronaldokun Could not actively follow this competition ( after 2 continuous one ) due to other priority and focused on learning PyTorch and Fast.AI as new papers coming up has more PyTorch implementation and FAST.AI have some cool stuff to accelerate learning and training.\n\nI could infer from images that image:1 of patient have no visible air and mid image ( if a patient has 30 image then approximately 14,15  or 16 ) has air visibility.\nNow this is the pattern for each patient's image.  If we can find distance of these images considering image-1 as base image ,we can find kind of lung capacity for each patient by either comparing using siamese or encoded vector coming from autoencoder.\n\nI created one autoencoder here https://www.kaggle.com/rajnishe/rc-siim-autoencoder-model\n\nTake output of encoeder for each image save on disk or compare difference at run time and save it.\nOne hyperparameter could be what distance you will use as in KNN .\n\nTo reduce too many features , Gaussian blur will help , I have put that code here\nhttps://www.kaggle.com/rajnishe/rc-osic-dcm-to-tif-n-jpg-conversion?scriptVersionId=41671497",
      "votes": null
    },
    {
      "id": "1030362",
      "postDate": "09/28/2020 15:51:40",
      "content": "<p>That's definitely a different approach. I'll try to implement it before the end of the competition and let you know.</p>",
      "rawMarkdown": "That's definitely a different approach. I'll try to implement it before the end of the competition and let you know.",
      "votes": null
    },
    {
      "id": "1039079",
      "postDate": "10/06/2020 09:58:50",
      "content": "<p>Gave this a try. I didn't have time for a robust test. My main indicator was the behaviour of the loss - whether it had a decaying curve or was just all over the place. It was the latter, so I considered it probably wouldn't do any better than my other CNN approaches.</p>",
      "rawMarkdown": "Gave this a try. I didn't have time for a robust test. My main indicator was the behaviour of the loss - whether it had a decaying curve or was just all over the place. It was the latter, so I considered it probably wouldn't do any better than my other CNN approaches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1029461,
      "author_name": "ronaldokun",
      "author_url": "",
      "post_date": "09/27/2020 18:44:56",
      "content": "<p>That's a great idea Rajnish. Although my team and many others failed at extracting any useful information from the image at this point, checkout this thread <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077</a> for more info.</p>\n<p>Let us know if you succeed in those attempts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1030042,
      "author_name": "rajnishe",
      "author_url": "",
      "post_date": "09/28/2020 11:23:57",
      "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> Could not actively follow this competition ( after 2 continuous one ) due to other priority and focused on learning PyTorch and Fast.AI as new papers coming up has more PyTorch implementation and FAST.AI have some cool stuff to accelerate learning and training.</p>\n<p>I could infer from images that image:1 of patient have no visible air and mid image ( if a patient has 30 image then approximately 14,15  or 16 ) has air visibility.<br>\nNow this is the pattern for each patient's image.  If we can find distance of these images considering image-1 as base image ,we can find kind of lung capacity for each patient by either comparing using siamese or encoded vector coming from autoencoder.</p>\n<p>I created one autoencoder here <a href=\"https://www.kaggle.com/rajnishe/rc-siim-autoencoder-model\" target=\"_blank\">https://www.kaggle.com/rajnishe/rc-siim-autoencoder-model</a></p>\n<p>Take output of encoeder for each image save on disk or compare difference at run time and save it.<br>\nOne hyperparameter could be what distance you will use as in KNN .</p>\n<p>To reduce too many features , Gaussian blur will help , I have put that code here<br>\n<a href=\"https://www.kaggle.com/rajnishe/rc-osic-dcm-to-tif-n-jpg-conversion?scriptVersionId=41671497\" target=\"_blank\">https://www.kaggle.com/rajnishe/rc-osic-dcm-to-tif-n-jpg-conversion?scriptVersionId=41671497</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1030362,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "09/28/2020 15:51:40",
          "content": "<p>That's definitely a different approach. I'll try to implement it before the end of the competition and let you know.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1039079,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "10/06/2020 09:58:50",
      "content": "<p>Gave this a try. I didn't have time for a robust test. My main indicator was the behaviour of the loss - whether it had a decaying curve or was just all over the place. It was the latter, so I considered it probably wouldn't do any better than my other CNN approaches.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1028890": "How we can use images to model in this competition.\n\nOne way I though of how an image of patient is different from other.\nAs images taken of single breathe in-out.\n\nApproach 1 :\nUsing siamese network it can be found as distance that could be use to as one of quantitative data with other tabular data to use.\nNot to use triplet but just difference of 2 images as in twin siamese.\nLess number of images is not a issue for siamese.\n\nApproach 2:\nThere is one notebook of autoencoder to generate feature vectors.\nDifference of these vectors one at full breath-in and one at full breath-out could also provide quantitative data to use with tabular.",
    "1029461": "That's a great idea Rajnish. Although my team and many others failed at extracting any useful information from the image at this point, checkout this thread https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077 for more info.\n\nLet us know if you succeed in those attempts.",
    "1030042": "ronaldokun Could not actively follow this competition ( after 2 continuous one ) due to other priority and focused on learning PyTorch and Fast.AI as new papers coming up has more PyTorch implementation and FAST.AI have some cool stuff to accelerate learning and training.\n\nI could infer from images that image:1 of patient have no visible air and mid image ( if a patient has 30 image then approximately 14,15  or 16 ) has air visibility.\nNow this is the pattern for each patient's image.  If we can find distance of these images considering image-1 as base image ,we can find kind of lung capacity for each patient by either comparing using siamese or encoded vector coming from autoencoder.\n\nI created one autoencoder here https://www.kaggle.com/rajnishe/rc-siim-autoencoder-model\n\nTake output of encoeder for each image save on disk or compare difference at run time and save it.\nOne hyperparameter could be what distance you will use as in KNN .\n\nTo reduce too many features , Gaussian blur will help , I have put that code here\nhttps://www.kaggle.com/rajnishe/rc-osic-dcm-to-tif-n-jpg-conversion?scriptVersionId=41671497",
    "1030362": "That's definitely a different approach. I'll try to implement it before the end of the competition and let you know.",
    "1039079": "Gave this a try. I didn't have time for a robust test. My main indicator was the behaviour of the loss - whether it had a decaying curve or was just all over the place. It was the latter, so I considered it probably wouldn't do any better than my other CNN approaches."
  },
  "source": "meta"
}