{
  "id": 69371,
  "title": "Are full-sized images any good?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/69371",
  "author_name": "",
  "post_date": "2018-10-23T09:22:55.031737500Z",
  "votes": 11,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I've just joined the competition and I see that most (if not all) of the kernels are using a resized version of the data, usually 299x299. Aren't the 512x512 images worth training? Or was that just due to the time limit?</p>",
  "messages": [
    {
      "id": "408667",
      "postDate": "10/23/2018 09:22:55",
      "content": "<p>I've just joined the competition and I see that most (if not all) of the kernels are using a resized version of the data, usually 299x299. Aren't the 512x512 images worth training? Or was that just due to the time limit?</p>",
      "rawMarkdown": "I've just joined the competition and I see that most (if not all) of the kernels are using a resized version of the data, usually 299x299. Aren't the 512x512 images worth training? Or was that just due to the time limit?",
      "votes": null
    },
    {
      "id": "408727",
      "postDate": "10/23/2018 11:15:10",
      "content": "<p>Another interesting question is: are 250Gb-dataset version any good?</p>",
      "rawMarkdown": "Another interesting question is: are 250Gb-dataset version any good?",
      "votes": null
    },
    {
      "id": "408745",
      "postDate": "10/23/2018 11:40:57",
      "content": "<p>I guess yes. \nFrom this kernel: <a href=\"https://www.kaggle.com/allunia/protein-atlas-exploration-and-baseline\">https://www.kaggle.com/allunia/protein-atlas-exploration-and-baseline</a></p>\n\n<blockquote>\n  <p>You can see that we have lost a lot of information by downscaling the image!  </p>\n</blockquote>\n\n<p>There are some targets (Ex: MICROTUBULE ORGANIZING CENTER) have a small pattern recognition. So, I think image size is important\n<img src=\"https://www.proteinatlas.org/images_dictionary/microtubule-organizing_center__2__52269__1_red_green.jpg\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "I guess yes. \nFrom this kernel: https://www.kaggle.com/allunia/protein-atlas-exploration-and-baseline\n&gt;  You can see that we have lost a lot of information by downscaling the image!  \n\nThere are some targets (Ex: MICROTUBULE ORGANIZING CENTER) have a small pattern recognition. So, I think image size is important\n![enter image description here][1]\n\n[1]: https://www.proteinatlas.org/images_dictionary/microtubule-organizing_center__2__52269__1_red_green.jpg",
      "votes": null
    },
    {
      "id": "408774",
      "postDate": "10/23/2018 12:43:18",
      "content": "<p>Definitely larger images worth training as they capture very low level details. During resizing one might loose few details.  If you fear that you might loose data, do multiple crops and use these to augment your input data, so that the original image will be split into different images of smaller sizes while you will preserve details of original image.</p>",
      "rawMarkdown": "Definitely larger images worth training as they capture very low level details. During resizing one might loose few details.  If you fear that you might loose data, do multiple crops and use these to augment your input data, so that the original image will be split into different images of smaller sizes while you will preserve details of original image.",
      "votes": null
    },
    {
      "id": "408867",
      "postDate": "10/23/2018 14:54:55",
      "content": "<p>How to ensure that the crops are focusing on the area containing the target proteins?</p>",
      "rawMarkdown": "How to ensure that the crops are focusing on the area containing the target proteins?",
      "votes": null
    },
    {
      "id": "409152",
      "postDate": "10/23/2018 21:01:52",
      "content": "<p>They are definitely good. The problem is, it's much harder to use 512x512 images for at least 2 reasons:\n 1/ most pretrained networks have maximum size 299x299 and using them promises quick win without too much effort early on the competition \n 2/ 512x512 take 3 times larger space in memory than 299x299, take longer to load from hdd and longer to converge. And Kaggle resources are limited. So you need either dedicated desktop pc with GPU or your own gcloud instance to train network for this size and it's much harder to train something usuable on Kaggle public kernels.</p>",
      "rawMarkdown": "They are definitely good. The problem is, it's much harder to use 512x512 images for at least 2 reasons:\n 1/ most pretrained networks have maximum size 299x299 and using them promises quick win without too much effort early on the competition \n 2/ 512x512 take 3 times larger space in memory than 299x299, take longer to load from hdd and longer to converge. And Kaggle resources are limited. So you need either dedicated desktop pc with GPU or your own gcloud instance to train network for this size and it's much harder to train something usuable on Kaggle public kernels.",
      "votes": null
    },
    {
      "id": "409209",
      "postDate": "10/24/2018 00:43:48",
      "content": "<p>Interesting insights, thanks for sharing!</p>",
      "rawMarkdown": "Interesting insights, thanks for sharing!",
      "votes": null
    },
    {
      "id": "409210",
      "postDate": "10/24/2018 00:46:20",
      "content": "<p>I'm trying to train one at full size right now and it is definitely taking much longer to converge than at the default size. We'll have to wait a few days and see.</p>",
      "rawMarkdown": "I'm trying to train one at full size right now and it is definitely taking much longer to converge than at the default size. We'll have to wait a few days and see.",
      "votes": null
    },
    {
      "id": "409297",
      "postDate": "10/24/2018 04:36:50",
      "content": "<p>If you are using crops which covers more than 50% of your original image(e.g. create 299x299 crops for 512x512 image) then i think it would be reasonably fair to apply same target of image to its crops. This is an experiment but i think using this way allow you to preserve original details also with augmentation of data. I tried this and found 0.018% improvement from 0.15 to 0.168 though its not significant but its all about experimenting.</p>",
      "rawMarkdown": "If you are using crops which covers more than 50% of your original image(e.g. create 299x299 crops for 512x512 image) then i think it would be reasonably fair to apply same target of image to its crops. This is an experiment but i think using this way allow you to preserve original details also with augmentation of data. I tried this and found 0.018% improvement from 0.15 to 0.168 though its not significant but its all about experimenting.",
      "votes": null
    },
    {
      "id": "410476",
      "postDate": "10/26/2018 05:37:00",
      "content": "<p>My latest networks use the 512x512x3 using RGB only. The most recent design I am trying turns out to work much better with the 512x512 images. It's not a huge model, I can run a batch size of 48 in 10GB ram. Training is from scratch, 20-30 minutes per epoch over 123K input images.</p>",
      "rawMarkdown": "My latest networks use the 512x512x3 using RGB only. The most recent design I am trying turns out to work much better with the 512x512 images. It's not a huge model, I can run a batch size of 48 in 10GB ram. Training is from scratch, 20-30 minutes per epoch over 123K input images.",
      "votes": null
    },
    {
      "id": "410497",
      "postDate": "10/26/2018 06:36:32",
      "content": "<p>Do you think that Y channel literally worsens the performance of your model or that the added precision doesn't weight out the longer training and memory consumption what would be neccessary?</p>",
      "rawMarkdown": "Do you think that Y channel literally worsens the performance of your model or that the added precision doesn't weight out the longer training and memory consumption what would be neccessary?",
      "votes": null
    },
    {
      "id": "410601",
      "postDate": "10/26/2018 10:27:43",
      "content": "<p>123K input images?</p>\n\n<p>Are they in the dataset? (Are Kernel datasets limited also in the amount of images? 30K) </p>",
      "rawMarkdown": "123K input images?\n\nAre they in the dataset? (Are Kernel datasets limited also in the amount of images? 30K)",
      "votes": null
    },
    {
      "id": "410752",
      "postDate": "10/26/2018 15:48:54",
      "content": "<p>@Daniel, I am manipulating and duplicating images to help with training and class imbalance. At 512x512 it works better if augmentation is done before and then the results saved in a folder. Otherwise I become limited by the CPU augmenting the images. I then use Keras ImageDataGenerator to stream from the pre augmented files. If I am running a \"fast\" model I load subsets of the training into memory. </p>\n\n<p>First I split the validation set from the train set, then I augment the train set with a list of images. The final counts look like:</p>\n\n<pre>27964 original training\n68093 augmented data\n124632 training masks\n3108 validation masks\n</pre>\n\n<p>@Michal, Y channel didn't seem to help or hurt model performance. Longer training and memory usage were the main reasons for dropping it. Researching the problem domain I also found that typically the RGB filters are used and Y was rarely mentioned. The earlier papers only used R and G. I am also aiming for the special prize, so memory usage and model size is a large factor.  My current best model is only 3.5M parameters, and half of that is the dense output layers.</p>",
      "rawMarkdown": "Daniel, I am manipulating and duplicating images to help with training and class imbalance. At 512x512 it works better if augmentation is done before and then the results saved in a folder. Otherwise I become limited by the CPU augmenting the images. I then use Keras ImageDataGenerator to stream from the pre augmented files. If I am running a \"fast\" model I load subsets of the training into memory. \n\nFirst I split the validation set from the train set, then I augment the train set with a list of images. The final counts look like:\n<pre>27964 original training\n68093 augmented data\n124632 training masks\n3108 validation masks\n</pre>\n\n@Michal, Y channel didn't seem to help or hurt model performance. Longer training and memory usage were the main reasons for dropping it. Researching the problem domain I also found that typically the RGB filters are used and Y was rarely mentioned. The earlier papers only used R and G. I am also aiming for the special prize, so memory usage and model size is a large factor.  My current best model is only 3.5M parameters, and half of that is the dense output layers.",
      "votes": null
    },
    {
      "id": "411110",
      "postDate": "10/27/2018 12:42:35",
      "content": "<p>Using a bigger image increases the model size and training time.</p>",
      "rawMarkdown": "Using a bigger image increases the model size and training time.",
      "votes": null
    },
    {
      "id": "411246",
      "postDate": "10/27/2018 17:27:42",
      "content": "<p>I'm planning to use even larger ones, maybe 1024*1024</p>",
      "rawMarkdown": "I'm planning to use even larger ones, maybe 1024*1024",
      "votes": null
    },
    {
      "id": "414117",
      "postDate": "11/02/2018 07:05:28",
      "content": "<p>What do you mean by 'masks'?</p>",
      "rawMarkdown": "What do you mean by 'masks'?",
      "votes": null
    },
    {
      "id": "414128",
      "postDate": "11/02/2018 07:28:59",
      "content": "<p>Oh I must have put the wrong name in there. I was working on the TGS at the same time I started this and that was predicting image masks. It should be images instead of masks. I take the training images and a 10% split for validation. Then I added 68093 additional input images. After that I duplicated some of the labels with very little data to end up with a total of 124632 images in my training data set, 3108 images in the validation set.</p>",
      "rawMarkdown": "Oh I must have put the wrong name in there. I was working on the TGS at the same time I started this and that was predicting image masks. It should be images instead of masks. I take the training images and a 10% split for validation. Then I added 68093 additional input images. After that I duplicated some of the labels with very little data to end up with a total of 124632 images in my training data set, 3108 images in the validation set.",
      "votes": null
    },
    {
      "id": "414374",
      "postDate": "11/02/2018 16:09:00",
      "content": "<p>Here's a good classification model library on: <a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a>.  You can either pip install or git clone.  In regards to image size for pretrained models.  You can modify the default_size and min_size to fit your needs. That can be done in the builder.py file. </p>",
      "rawMarkdown": "Here's a good classification model library on: https://github.com/qubvel/classification_models.  You can either pip install or git clone.  In regards to image size for pretrained models.  You can modify the default_size and min_size to fit your needs. That can be done in the builder.py file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 408727,
      "author_name": "comratvlad",
      "author_url": "",
      "post_date": "10/23/2018 11:15:10",
      "content": "<p>Another interesting question is: are 250Gb-dataset version any good?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 408745,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "10/23/2018 11:40:57",
      "content": "<p>I guess yes. \nFrom this kernel: <a href=\"https://www.kaggle.com/allunia/protein-atlas-exploration-and-baseline\">https://www.kaggle.com/allunia/protein-atlas-exploration-and-baseline</a></p>\n\n<blockquote>\n  <p>You can see that we have lost a lot of information by downscaling the image!  </p>\n</blockquote>\n\n<p>There are some targets (Ex: MICROTUBULE ORGANIZING CENTER) have a small pattern recognition. So, I think image size is important\n<img src=\"https://www.proteinatlas.org/images_dictionary/microtubule-organizing_center__2__52269__1_red_green.jpg\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 409209,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "10/24/2018 00:43:48",
          "content": "<p>Interesting insights, thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 408774,
      "author_name": "bhamin",
      "author_url": "",
      "post_date": "10/23/2018 12:43:18",
      "content": "<p>Definitely larger images worth training as they capture very low level details. During resizing one might loose few details.  If you fear that you might loose data, do multiple crops and use these to augment your input data, so that the original image will be split into different images of smaller sizes while you will preserve details of original image.</p>",
      "votes": null,
      "replies": [
        {
          "id": 408867,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/23/2018 14:54:55",
          "content": "<p>How to ensure that the crops are focusing on the area containing the target proteins?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 409297,
          "author_name": "bhamin",
          "author_url": "",
          "post_date": "10/24/2018 04:36:50",
          "content": "<p>If you are using crops which covers more than 50% of your original image(e.g. create 299x299 crops for 512x512 image) then i think it would be reasonably fair to apply same target of image to its crops. This is an experiment but i think using this way allow you to preserve original details also with augmentation of data. I tried this and found 0.018% improvement from 0.15 to 0.168 though its not significant but its all about experimenting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409152,
      "author_name": "rejpalcz",
      "author_url": "",
      "post_date": "10/23/2018 21:01:52",
      "content": "<p>They are definitely good. The problem is, it's much harder to use 512x512 images for at least 2 reasons:\n 1/ most pretrained networks have maximum size 299x299 and using them promises quick win without too much effort early on the competition \n 2/ 512x512 take 3 times larger space in memory than 299x299, take longer to load from hdd and longer to converge. And Kaggle resources are limited. So you need either dedicated desktop pc with GPU or your own gcloud instance to train network for this size and it's much harder to train something usuable on Kaggle public kernels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 409210,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "10/24/2018 00:46:20",
          "content": "<p>I'm trying to train one at full size right now and it is definitely taking much longer to converge than at the default size. We'll have to wait a few days and see.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410476,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "10/26/2018 05:37:00",
      "content": "<p>My latest networks use the 512x512x3 using RGB only. The most recent design I am trying turns out to work much better with the 512x512 images. It's not a huge model, I can run a batch size of 48 in 10GB ram. Training is from scratch, 20-30 minutes per epoch over 123K input images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 410497,
          "author_name": "rejpalcz",
          "author_url": "",
          "post_date": "10/26/2018 06:36:32",
          "content": "<p>Do you think that Y channel literally worsens the performance of your model or that the added precision doesn't weight out the longer training and memory consumption what would be neccessary?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410601,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "10/26/2018 10:27:43",
          "content": "<p>123K input images?</p>\n\n<p>Are they in the dataset? (Are Kernel datasets limited also in the amount of images? 30K) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410752,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/26/2018 15:48:54",
          "content": "<p>@Daniel, I am manipulating and duplicating images to help with training and class imbalance. At 512x512 it works better if augmentation is done before and then the results saved in a folder. Otherwise I become limited by the CPU augmenting the images. I then use Keras ImageDataGenerator to stream from the pre augmented files. If I am running a \"fast\" model I load subsets of the training into memory. </p>\n\n<p>First I split the validation set from the train set, then I augment the train set with a list of images. The final counts look like:</p>\n\n<pre>27964 original training\n68093 augmented data\n124632 training masks\n3108 validation masks\n</pre>\n\n<p>@Michal, Y channel didn't seem to help or hurt model performance. Longer training and memory usage were the main reasons for dropping it. Researching the problem domain I also found that typically the RGB filters are used and Y was rarely mentioned. The earlier papers only used R and G. I am also aiming for the special prize, so memory usage and model size is a large factor.  My current best model is only 3.5M parameters, and half of that is the dense output layers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414117,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "11/02/2018 07:05:28",
          "content": "<p>What do you mean by 'masks'?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414128,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/02/2018 07:28:59",
          "content": "<p>Oh I must have put the wrong name in there. I was working on the TGS at the same time I started this and that was predicting image masks. It should be images instead of masks. I take the training images and a 10% split for validation. Then I added 68093 additional input images. After that I duplicated some of the labels with very little data to end up with a total of 124632 images in my training data set, 3108 images in the validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 411110,
      "author_name": "romainbeaumont",
      "author_url": "",
      "post_date": "10/27/2018 12:42:35",
      "content": "<p>Using a bigger image increases the model size and training time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 411246,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "10/27/2018 17:27:42",
      "content": "<p>I'm planning to use even larger ones, maybe 1024*1024</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414374,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "11/02/2018 16:09:00",
      "content": "<p>Here's a good classification model library on: <a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a>.  You can either pip install or git clone.  In regards to image size for pretrained models.  You can modify the default_size and min_size to fit your needs. That can be done in the builder.py file. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "408667": "I've just joined the competition and I see that most (if not all) of the kernels are using a resized version of the data, usually 299x299. Aren't the 512x512 images worth training? Or was that just due to the time limit?",
    "408727": "Another interesting question is: are 250Gb-dataset version any good?",
    "408745": "I guess yes. \nFrom this kernel: https://www.kaggle.com/allunia/protein-atlas-exploration-and-baseline\n&gt;  You can see that we have lost a lot of information by downscaling the image!  \n\nThere are some targets (Ex: MICROTUBULE ORGANIZING CENTER) have a small pattern recognition. So, I think image size is important\n![enter image description here][1]\n\n[1]: https://www.proteinatlas.org/images_dictionary/microtubule-organizing_center__2__52269__1_red_green.jpg",
    "408774": "Definitely larger images worth training as they capture very low level details. During resizing one might loose few details.  If you fear that you might loose data, do multiple crops and use these to augment your input data, so that the original image will be split into different images of smaller sizes while you will preserve details of original image.",
    "408867": "How to ensure that the crops are focusing on the area containing the target proteins?",
    "409152": "They are definitely good. The problem is, it's much harder to use 512x512 images for at least 2 reasons:\n 1/ most pretrained networks have maximum size 299x299 and using them promises quick win without too much effort early on the competition \n 2/ 512x512 take 3 times larger space in memory than 299x299, take longer to load from hdd and longer to converge. And Kaggle resources are limited. So you need either dedicated desktop pc with GPU or your own gcloud instance to train network for this size and it's much harder to train something usuable on Kaggle public kernels.",
    "409209": "Interesting insights, thanks for sharing!",
    "409210": "I'm trying to train one at full size right now and it is definitely taking much longer to converge than at the default size. We'll have to wait a few days and see.",
    "409297": "If you are using crops which covers more than 50% of your original image(e.g. create 299x299 crops for 512x512 image) then i think it would be reasonably fair to apply same target of image to its crops. This is an experiment but i think using this way allow you to preserve original details also with augmentation of data. I tried this and found 0.018% improvement from 0.15 to 0.168 though its not significant but its all about experimenting.",
    "410476": "My latest networks use the 512x512x3 using RGB only. The most recent design I am trying turns out to work much better with the 512x512 images. It's not a huge model, I can run a batch size of 48 in 10GB ram. Training is from scratch, 20-30 minutes per epoch over 123K input images.",
    "410497": "Do you think that Y channel literally worsens the performance of your model or that the added precision doesn't weight out the longer training and memory consumption what would be neccessary?",
    "410601": "123K input images?\n\nAre they in the dataset? (Are Kernel datasets limited also in the amount of images? 30K)",
    "410752": "Daniel, I am manipulating and duplicating images to help with training and class imbalance. At 512x512 it works better if augmentation is done before and then the results saved in a folder. Otherwise I become limited by the CPU augmenting the images. I then use Keras ImageDataGenerator to stream from the pre augmented files. If I am running a \"fast\" model I load subsets of the training into memory. \n\nFirst I split the validation set from the train set, then I augment the train set with a list of images. The final counts look like:\n<pre>27964 original training\n68093 augmented data\n124632 training masks\n3108 validation masks\n</pre>\n\n@Michal, Y channel didn't seem to help or hurt model performance. Longer training and memory usage were the main reasons for dropping it. Researching the problem domain I also found that typically the RGB filters are used and Y was rarely mentioned. The earlier papers only used R and G. I am also aiming for the special prize, so memory usage and model size is a large factor.  My current best model is only 3.5M parameters, and half of that is the dense output layers.",
    "411110": "Using a bigger image increases the model size and training time.",
    "411246": "I'm planning to use even larger ones, maybe 1024*1024",
    "414117": "What do you mean by 'masks'?",
    "414128": "Oh I must have put the wrong name in there. I was working on the TGS at the same time I started this and that was predicting image masks. It should be images instead of masks. I take the training images and a 10% split for validation. Then I added 68093 additional input images. After that I duplicated some of the labels with very little data to end up with a total of 124632 images in my training data set, 3108 images in the validation set.",
    "414374": "Here's a good classification model library on: https://github.com/qubvel/classification_models.  You can either pip install or git clone.  In regards to image size for pretrained models.  You can modify the default_size and min_size to fit your needs. That can be done in the builder.py file."
  },
  "source": "meta"
}