{
  "id": 171745,
  "title": "Interesting solution for 224 x 224 images (0.9287, no external data)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171745",
  "author_name": "chris",
  "post_date": "2020-08-02T09:51:20.436000",
  "votes": 89,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I sort of gave up on trying to compete with the leaderboard after awhile and decided to just use this competition to experiment with various ideas I have. I ended up coming up with an approach that I figured I'd share. I have created a GitHub repo documenting my findings:</p>\n\n<p><a href=\"https://github.com/CCareaga/attention_guided_cropping\">Attention Guided Cropping</a></p>\n\n<p>Here is a the main idea ripped from the README of the repo:</p>\n\n<h3>Disclaimer</h3>\n\n<p>After working on this project I searched around a bit for related work, and found:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1801.09927.pdf\">Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification</a></p>\n\n<p>This paper uses a eerily similar approach, although it doesn't seem like they use the cropping technique to cross-reference original high-resolution images, and their attention technique varies from mine. Credit should definitely go to these authors for first experimenting with this idea. I also assume many others have tried methods like this (if so please let me know so I can add credit). I know that many competition participants likely also experimented with intelligent cropping, but I figured since I put time into this approach I would still document my findings.</p>\n\n<h3>Approach</h3>\n\n<p>The dermatoscopy images provided as part of the competition are very high resolution, the largest images in the dataset are 4000x6000 pixels. When using pre-trained networks, the images for the downstream task are typically resized to match the size of the pre-training task. For example when using an ImageNet pre-trained CNN, one will typically resize images to 224 x 224, as this is what the network is \"used\" to. As many found out during the competition, this was not necessary due to scale invariance, as well as the vast difference between pre-training tasks and the competition task. Participants began ensembling models trained on various image sizes, anywhere from 224 x 224 all the way up to 1024 x 1024. Due to my limited resources, and aversion to online notebooks, I decided I wouldn't really be able to compete with large models and large image sizes. I instead made a personal challenge to squeeze as much performance as I could using only 224 x 224 images. </p>\n\n<p>When examining the images, I observed that in many images, the lesion in question occupies a very small portion of the image. This means that a large part of the 224^2 pixels (presumably) do not provide a significant amount of information. I realized I could gain back information by cropping the images tightly around the lesion and before resizing, this would result in a much higher resolution version of the lesion while maintaining the 224 x 224 image size! </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F9648ec1b82dbe3f3f07f9dba06ad4e5a%2Fcropping.png?generation=1596361389623025&amp;alt=media\" alt=\"example of targeted image crop to increase lesion detail\"></p>\n\n<p>The amount of space occupied by the lesion varies greatly from image to image and other noise exists in the images (shadows, rulers, etc.) so this process cannot be trivially automated despite the simplicity of the images. I decided the best way to find salient regions is by having a model decide what information is important for classification. To accomplish this goal I decided to try to use attention. In this case I figured it would be pretty easy for the model to decide which pixels are lesion pixels and which pixels are regular skin without requiring explicit supervision. I devised the following model as a first attempt:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F4055fb6338fde6dfdc4a475d2f2394a0%2Fmodel_diagram.png?generation=1596361447826651&amp;alt=media\" alt=\"initial attention model\"></p>\n\n<p>The model is a simple extension of the ResNet architecture, allowing for the use of pre-trained weights. The changes consist of removing the global average pooling from the end of the ResNet and replacing it with an attention-weighted pooling operation. The attention weights are determined by a series of two convolutional layers over the unpooled feature maps to produce a single 7x7 map of scores. The scores are softmax for two reasons:</p>\n\n<ol>\n<li>They cause attention weights to sum to 1 allowing for a simple weighted average.</li>\n<li>They force the model to \"choose\" which pixels it wants to keep while attenuating the rest.</li>\n</ol>\n\n<p>The second point is what makes the model \"focus\" on certain regions. The model can still predict a uniform attention map, but when it does, the average pooling \"mixes\" the information causing a noisier representation. When the model predicts a concentrated attention map, the pooling selects the pixel features coming from discriminative regions of the image (in this case the skin lesion). All of this attention learning is driven by supervision from the task at hand, meaning that the regions that are selected are deemed important to the classification of the lesion. This means we can now extract salient regions from the image, as well as diagnose any spurious correlations picked up by the model.</p>\n\n<p>This initial attempt worked surprisingly well, the model did in fact determine important regions of the image, but unfortunately the attention maps where not exact. I conjectured that this was because spatial locations of the input don't perfectly correspond to the same spatial locations of the resulting feature maps. This is due to the large effective receptive field size of each feature map location. To fix this issue, I devised a way to constrict the information that contributes to each location of feature map. I decided to divide the image into square pieces, referred to as chips,  and feed each one into a ResNet separately. With the global average pooling, this results in a 512 dimensional embedding for each chip. I could then concatenate the chips back together to create a feature map where each position corresponding directly to a fixed region of the input image:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F94f8491eabe7625dd825ae5d950247f5%2Fchip_diagram.png?generation=1596361472945916&amp;alt=media\" alt=\"improved attention model\"></p>\n\n<p>This change resulted in much more consistent and accurate attention maps such as the following:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2Ffc37edb56413d616e74525528daa6a57%2Fattn_examples.png?generation=1596361507358931&amp;alt=media\" alt=\"examples of good attention maps generated\"></p>\n\n<p>These are some selected samples of the attention maps and their corresponding input images. It is obvious that the model is focusing on the pixels that contain the legion. While the maps look satisfactory a majority of the time, there are failure cases:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F23ae894ce6d277ffe51b4f52b956edf3%2Ffailures.png?generation=1596361541322606&amp;alt=media\" alt=\"examples of bad attention maps generated\"></p>\n\n<p>I was not able to track down why these images generate poor attention maps, but I did observe that larger models greatly improved the consistency of the attention maps. The examples shown above were generated using a ResNet-50 backbone and the chipping model discussed earlier. Once I was able to generate satisfactory attention maps, I developed a process to convert these attention maps into square regions for cropping:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F72a86105602f2889775cccda4fc772d4%2Fattn_to_crop.png?generation=1596361577931234&amp;alt=media\" alt=\"converting attention map to bounding box for cropping\"></p>\n\n<p>When the attention fails the model generally generates a pretty unifrom attention map. This means that the bounding box creted encapsulates a very large portion of the image and therefore the cropping process defaults to maintaining the original image. With cropped versions of both the training set and the test set, a new model can be trained. By ensembling the predictions of the a ResNet trained on the original images, and a ResNet trained on the cropped images, we can get a decent boost in performance without having to increase the image size!</p>\n\n<h3>Implementation</h3>\n\nPreprocessing\n\n<p>To start I center-cropped all images into a square using the size of their shortest side.  Next I resized each image to be 224 x 224 and store the training set and test set as .npy files (for easy loading into RAM). Before going into any of the models the images go through a series of augmentations:\n<code>\ntrain_transform = transforms.Compose([\n    DrawHair(),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.RandomRotation(360),\n    transforms.ColorJitter(\n        brightness=[0.8, 1.2],\n        contrast=[0.8, 1.2],\n        saturation=[0.8, 1.2],\n    ),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n</code>\nAll of these are built into PyTorch except the <code>DrawHair</code> transform, which is a simple class I wrote which draws random curved lines over the image simulating hair. </p>\n\n<p>The meta-data fields are converted to one-hot encodings. The one-hot encodings include an encoding for unknown/null. This process results in a 29-dimensional vector.</p>\n\nModels\n\n<p>Three models are trained as part of the previously described approach:\n1. resnet: A regular ResNet trained on the original images and provided metadata\n2. chipnet: The chip model used to generate attention maps\n3. cropnet: A regular ResNet trained on images cropped using attention</p>\n\n<p>Each of these models has a very similar training process. The training data consists of 33,126  images of varying size (32,542 benign, 584 malignant).  Metadata is included for each image, the meta-data fields are: age, anatomical site, gender, and specific diagnosis. These meta-data fields are converted to features and used as additional input to the models. Each model is trained and validated using 3-fold stratified cross-validation, this maintains the class distribution in each fold. The test set predictions of each fold are ensembled to produce the final predictions for each of the models. All models use focal loss to train, but would likely perform just as well with vanilla binary cross entropy. Additionally, test set predictions are generated by combining three rounds of test time augmentation using the augmentations described previously.</p>\n\nResNet50 model (resnet)\n\n<p>This model consists of a ResNet50 to process the image input, along with a two linear layers to process the meta-data features. The embeddings from these two branches are concatenated and sent through a final linear layer to produce a single score. This model is implemented in <code>resnet.py</code> by the class <code>MetaResNet</code>. The only variation between this model and cropnet model is the images used to train. The model is trained using the Adam optimizer, and a balanced class sampler (oversampling), meaning each batch contains an even number of each class. The ResNet portion of the model is trained with a learning rate of 1x10  and the rest of the parameters have a learning rate of 5x10. Training proceeds until the validation AUC does not improve for 3 epochs. </p>\n\nChipping model (chipnet)\n\n<p>This model also consists of a ResNet50 backbone, but has no additional meta-data branch as it's goal is simply to produce attention maps. As explained before, each 224 x 224 image is split into 32 x 32 non-overlapping chips. Each chip is sent through the ResNet50 and the resulting embeddings are re-stacked to create 7 x 7 feature maps with some number of channels depending on the backbone. To compute attention maps, two 3 x 3 same convolutions are applied (with ReLU and batch norm in between). The result of this is a 7 x 7 map of scores. These scores are softmaxed to sum to one and used to take a weighted average of the original feature maps. This process results in a single embedding that can then be used for classification after the application of a final linear layer. This model is implemented in <code>resnet.py</code> by the class <code>ChipNet</code>. The model is trained the same way as the resnet model except  the ResNet backbone is trained with a learning rate of 5x10 while the attention sub-net and final classifier are trained with a learning rate of 1x10.</p>\n\nCropped image model (cropnet)\n\n<p>As previously stated this model is trained identically to the resnet model except it uses the image cropped by the attention maps generated from the chipnet model.</p>\n\n<h3>Results</h3>\n\n<p>I tried multiple ensembling techniques against the public leaderboard of the competition. Here are the results for the predictions of each model trained as part of the previously described approach:\n|model|public LB score (AUC ROC)  |\n|--|--|\n|ResNet-50 on original images (resnet)| <strong>0.9203</strong> |\n|Chip model (chipnet) | 0.8931|\n|ResNet-50 on cropped images (cropnet) | 0.9177 |</p>\n\n<p>The resnet model achieved the highest performance on the public leaderboard, while chipnet model scored the lowest. This makes sense as the chipnet only looks at single tiles and therefore has a very constrained receptive field corresponding to each feature map location. This hinders performance while also improving attention map consistency and accuracy. The cropnet model achieves performance comparable to the resnet model, despite the attention maps sometimes producing poor crops of the original image. With a better chipnet model, I believe there would be a corresponding improvement in the performance of the cropnet model. I used these numbers to  generate ensemble weights for the each model's predictions (bad practice; leads to overfitting public LB), to see how much new information the chipnet and cropnet models add.</p>\n\n<p>|ensemble|public LB|\n|--|--|\n|(0.4 * resnet) + (0.2 * chipnet) + (0.4 * cropnet) | 0.9263|\n|(0.7 * resnet) + (0.3 * chipnet) | 0.9203 |\n|(0.5 * resnet) + (0.5 cropnet) | <strong>0.9287</strong> |</p>\n\n<p>Turns out using the predictions of the chipnet does not improve results and the best ensemble is a straight average of the resnet model and the cropnet model. Although public LB may not be the best performance indicator, I argue that it at least shows that this method is viable and extra information can be gleaned from the cropped images. </p>\n\n<h3>Conclusion</h3>\n\n<p>Although the approach to medical image classification isn't necessarily novel, I believe there is still a lot of exploration to be done in this area. I could see these methods being applied to other vision tasks (for example landmark recognition). Additionally I hope this repository is at least slightly informational and encourages others to experiment beyond the typical image classification approaches. I would also like to note that this approach does not make use of pixel level segmentation supervision (although this type of data exists for this task) so it could be considered a \"weakly\" supervised technique. This means the approach could be used on tasks where no segmentation annotations are available.</p>",
  "messages": [
    {
      "id": 955067,
      "postDate": "2020-08-02T09:51:20.437Z",
      "content": "<p>I sort of gave up on trying to compete with the leaderboard after awhile and decided to just use this competition to experiment with various ideas I have. I ended up coming up with an approach that I figured I'd share. I have created a GitHub repo documenting my findings:</p>\n\n<p><a href=\"https://github.com/CCareaga/attention_guided_cropping\">Attention Guided Cropping</a></p>\n\n<p>Here is a the main idea ripped from the README of the repo:</p>\n\n<h3>Disclaimer</h3>\n\n<p>After working on this project I searched around a bit for related work, and found:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1801.09927.pdf\">Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification</a></p>\n\n<p>This paper uses a eerily similar approach, although it doesn't seem like they use the cropping technique to cross-reference original high-resolution images, and their attention technique varies from mine. Credit should definitely go to these authors for first experimenting with this idea. I also assume many others have tried methods like this (if so please let me know so I can add credit). I know that many competition participants likely also experimented with intelligent cropping, but I figured since I put time into this approach I would still document my findings.</p>\n\n<h3>Approach</h3>\n\n<p>The dermatoscopy images provided as part of the competition are very high resolution, the largest images in the dataset are 4000x6000 pixels. When using pre-trained networks, the images for the downstream task are typically resized to match the size of the pre-training task. For example when using an ImageNet pre-trained CNN, one will typically resize images to 224 x 224, as this is what the network is \"used\" to. As many found out during the competition, this was not necessary due to scale invariance, as well as the vast difference between pre-training tasks and the competition task. Participants began ensembling models trained on various image sizes, anywhere from 224 x 224 all the way up to 1024 x 1024. Due to my limited resources, and aversion to online notebooks, I decided I wouldn't really be able to compete with large models and large image sizes. I instead made a personal challenge to squeeze as much performance as I could using only 224 x 224 images. </p>\n\n<p>When examining the images, I observed that in many images, the lesion in question occupies a very small portion of the image. This means that a large part of the 224^2 pixels (presumably) do not provide a significant amount of information. I realized I could gain back information by cropping the images tightly around the lesion and before resizing, this would result in a much higher resolution version of the lesion while maintaining the 224 x 224 image size! </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F9648ec1b82dbe3f3f07f9dba06ad4e5a%2Fcropping.png?generation=1596361389623025&amp;alt=media\" alt=\"example of targeted image crop to increase lesion detail\"></p>\n\n<p>The amount of space occupied by the lesion varies greatly from image to image and other noise exists in the images (shadows, rulers, etc.) so this process cannot be trivially automated despite the simplicity of the images. I decided the best way to find salient regions is by having a model decide what information is important for classification. To accomplish this goal I decided to try to use attention. In this case I figured it would be pretty easy for the model to decide which pixels are lesion pixels and which pixels are regular skin without requiring explicit supervision. I devised the following model as a first attempt:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F4055fb6338fde6dfdc4a475d2f2394a0%2Fmodel_diagram.png?generation=1596361447826651&amp;alt=media\" alt=\"initial attention model\"></p>\n\n<p>The model is a simple extension of the ResNet architecture, allowing for the use of pre-trained weights. The changes consist of removing the global average pooling from the end of the ResNet and replacing it with an attention-weighted pooling operation. The attention weights are determined by a series of two convolutional layers over the unpooled feature maps to produce a single 7x7 map of scores. The scores are softmax for two reasons:</p>\n\n<ol>\n<li>They cause attention weights to sum to 1 allowing for a simple weighted average.</li>\n<li>They force the model to \"choose\" which pixels it wants to keep while attenuating the rest.</li>\n</ol>\n\n<p>The second point is what makes the model \"focus\" on certain regions. The model can still predict a uniform attention map, but when it does, the average pooling \"mixes\" the information causing a noisier representation. When the model predicts a concentrated attention map, the pooling selects the pixel features coming from discriminative regions of the image (in this case the skin lesion). All of this attention learning is driven by supervision from the task at hand, meaning that the regions that are selected are deemed important to the classification of the lesion. This means we can now extract salient regions from the image, as well as diagnose any spurious correlations picked up by the model.</p>\n\n<p>This initial attempt worked surprisingly well, the model did in fact determine important regions of the image, but unfortunately the attention maps where not exact. I conjectured that this was because spatial locations of the input don't perfectly correspond to the same spatial locations of the resulting feature maps. This is due to the large effective receptive field size of each feature map location. To fix this issue, I devised a way to constrict the information that contributes to each location of feature map. I decided to divide the image into square pieces, referred to as chips,  and feed each one into a ResNet separately. With the global average pooling, this results in a 512 dimensional embedding for each chip. I could then concatenate the chips back together to create a feature map where each position corresponding directly to a fixed region of the input image:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F94f8491eabe7625dd825ae5d950247f5%2Fchip_diagram.png?generation=1596361472945916&amp;alt=media\" alt=\"improved attention model\"></p>\n\n<p>This change resulted in much more consistent and accurate attention maps such as the following:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2Ffc37edb56413d616e74525528daa6a57%2Fattn_examples.png?generation=1596361507358931&amp;alt=media\" alt=\"examples of good attention maps generated\"></p>\n\n<p>These are some selected samples of the attention maps and their corresponding input images. It is obvious that the model is focusing on the pixels that contain the legion. While the maps look satisfactory a majority of the time, there are failure cases:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F23ae894ce6d277ffe51b4f52b956edf3%2Ffailures.png?generation=1596361541322606&amp;alt=media\" alt=\"examples of bad attention maps generated\"></p>\n\n<p>I was not able to track down why these images generate poor attention maps, but I did observe that larger models greatly improved the consistency of the attention maps. The examples shown above were generated using a ResNet-50 backbone and the chipping model discussed earlier. Once I was able to generate satisfactory attention maps, I developed a process to convert these attention maps into square regions for cropping:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F72a86105602f2889775cccda4fc772d4%2Fattn_to_crop.png?generation=1596361577931234&amp;alt=media\" alt=\"converting attention map to bounding box for cropping\"></p>\n\n<p>When the attention fails the model generally generates a pretty unifrom attention map. This means that the bounding box creted encapsulates a very large portion of the image and therefore the cropping process defaults to maintaining the original image. With cropped versions of both the training set and the test set, a new model can be trained. By ensembling the predictions of the a ResNet trained on the original images, and a ResNet trained on the cropped images, we can get a decent boost in performance without having to increase the image size!</p>\n\n<h3>Implementation</h3>\n\nPreprocessing\n\n<p>To start I center-cropped all images into a square using the size of their shortest side.  Next I resized each image to be 224 x 224 and store the training set and test set as .npy files (for easy loading into RAM). Before going into any of the models the images go through a series of augmentations:\n<code>\ntrain_transform = transforms.Compose([\n    DrawHair(),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.RandomRotation(360),\n    transforms.ColorJitter(\n        brightness=[0.8, 1.2],\n        contrast=[0.8, 1.2],\n        saturation=[0.8, 1.2],\n    ),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n</code>\nAll of these are built into PyTorch except the <code>DrawHair</code> transform, which is a simple class I wrote which draws random curved lines over the image simulating hair. </p>\n\n<p>The meta-data fields are converted to one-hot encodings. The one-hot encodings include an encoding for unknown/null. This process results in a 29-dimensional vector.</p>\n\nModels\n\n<p>Three models are trained as part of the previously described approach:\n1. resnet: A regular ResNet trained on the original images and provided metadata\n2. chipnet: The chip model used to generate attention maps\n3. cropnet: A regular ResNet trained on images cropped using attention</p>\n\n<p>Each of these models has a very similar training process. The training data consists of 33,126  images of varying size (32,542 benign, 584 malignant).  Metadata is included for each image, the meta-data fields are: age, anatomical site, gender, and specific diagnosis. These meta-data fields are converted to features and used as additional input to the models. Each model is trained and validated using 3-fold stratified cross-validation, this maintains the class distribution in each fold. The test set predictions of each fold are ensembled to produce the final predictions for each of the models. All models use focal loss to train, but would likely perform just as well with vanilla binary cross entropy. Additionally, test set predictions are generated by combining three rounds of test time augmentation using the augmentations described previously.</p>\n\nResNet50 model (resnet)\n\n<p>This model consists of a ResNet50 to process the image input, along with a two linear layers to process the meta-data features. The embeddings from these two branches are concatenated and sent through a final linear layer to produce a single score. This model is implemented in <code>resnet.py</code> by the class <code>MetaResNet</code>. The only variation between this model and cropnet model is the images used to train. The model is trained using the Adam optimizer, and a balanced class sampler (oversampling), meaning each batch contains an even number of each class. The ResNet portion of the model is trained with a learning rate of 1x10  and the rest of the parameters have a learning rate of 5x10. Training proceeds until the validation AUC does not improve for 3 epochs. </p>\n\nChipping model (chipnet)\n\n<p>This model also consists of a ResNet50 backbone, but has no additional meta-data branch as it's goal is simply to produce attention maps. As explained before, each 224 x 224 image is split into 32 x 32 non-overlapping chips. Each chip is sent through the ResNet50 and the resulting embeddings are re-stacked to create 7 x 7 feature maps with some number of channels depending on the backbone. To compute attention maps, two 3 x 3 same convolutions are applied (with ReLU and batch norm in between). The result of this is a 7 x 7 map of scores. These scores are softmaxed to sum to one and used to take a weighted average of the original feature maps. This process results in a single embedding that can then be used for classification after the application of a final linear layer. This model is implemented in <code>resnet.py</code> by the class <code>ChipNet</code>. The model is trained the same way as the resnet model except  the ResNet backbone is trained with a learning rate of 5x10 while the attention sub-net and final classifier are trained with a learning rate of 1x10.</p>\n\nCropped image model (cropnet)\n\n<p>As previously stated this model is trained identically to the resnet model except it uses the image cropped by the attention maps generated from the chipnet model.</p>\n\n<h3>Results</h3>\n\n<p>I tried multiple ensembling techniques against the public leaderboard of the competition. Here are the results for the predictions of each model trained as part of the previously described approach:\n|model|public LB score (AUC ROC)  |\n|--|--|\n|ResNet-50 on original images (resnet)| <strong>0.9203</strong> |\n|Chip model (chipnet) | 0.8931|\n|ResNet-50 on cropped images (cropnet) | 0.9177 |</p>\n\n<p>The resnet model achieved the highest performance on the public leaderboard, while chipnet model scored the lowest. This makes sense as the chipnet only looks at single tiles and therefore has a very constrained receptive field corresponding to each feature map location. This hinders performance while also improving attention map consistency and accuracy. The cropnet model achieves performance comparable to the resnet model, despite the attention maps sometimes producing poor crops of the original image. With a better chipnet model, I believe there would be a corresponding improvement in the performance of the cropnet model. I used these numbers to  generate ensemble weights for the each model's predictions (bad practice; leads to overfitting public LB), to see how much new information the chipnet and cropnet models add.</p>\n\n<p>|ensemble|public LB|\n|--|--|\n|(0.4 * resnet) + (0.2 * chipnet) + (0.4 * cropnet) | 0.9263|\n|(0.7 * resnet) + (0.3 * chipnet) | 0.9203 |\n|(0.5 * resnet) + (0.5 cropnet) | <strong>0.9287</strong> |</p>\n\n<p>Turns out using the predictions of the chipnet does not improve results and the best ensemble is a straight average of the resnet model and the cropnet model. Although public LB may not be the best performance indicator, I argue that it at least shows that this method is viable and extra information can be gleaned from the cropped images. </p>\n\n<h3>Conclusion</h3>\n\n<p>Although the approach to medical image classification isn't necessarily novel, I believe there is still a lot of exploration to be done in this area. I could see these methods being applied to other vision tasks (for example landmark recognition). Additionally I hope this repository is at least slightly informational and encourages others to experiment beyond the typical image classification approaches. I would also like to note that this approach does not make use of pixel level segmentation supervision (although this type of data exists for this task) so it could be considered a \"weakly\" supervised technique. This means the approach could be used on tasks where no segmentation annotations are available.</p>",
      "rawMarkdown": "I sort of gave up on trying to compete with the leaderboard after awhile and decided to just use this competition to experiment with various ideas I have. I ended up coming up with an approach that I figured I'd share. I have created a GitHub repo documenting my findings:\n\n[Attention Guided Cropping](https://github.com/CCareaga/attention_guided_cropping)\n\nHere is a the main idea ripped from the README of the repo:\n\n### Disclaimer\nAfter working on this project I searched around a bit for related work, and found:\n\n[Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification](https://arxiv.org/pdf/1801.09927.pdf)\n\nThis paper uses a eerily similar approach, although it doesn't seem like they use the cropping technique to cross-reference original high-resolution images, and their attention technique varies from mine. Credit should definitely go to these authors for first experimenting with this idea. I also assume many others have tried methods like this (if so please let me know so I can add credit). I know that many competition participants likely also experimented with intelligent cropping, but I figured since I put time into this approach I would still document my findings.\n\n### Approach\n\nThe dermatoscopy images provided as part of the competition are very high resolution, the largest images in the dataset are 4000x6000 pixels. When using pre-trained networks, the images for the downstream task are typically resized to match the size of the pre-training task. For example when using an ImageNet pre-trained CNN, one will typically resize images to 224 x 224, as this is what the network is \"used\" to. As many found out during the competition, this was not necessary due to scale invariance, as well as the vast difference between pre-training tasks and the competition task. Participants began ensembling models trained on various image sizes, anywhere from 224 x 224 all the way up to 1024 x 1024. Due to my limited resources, and aversion to online notebooks, I decided I wouldn't really be able to compete with large models and large image sizes. I instead made a personal challenge to squeeze as much performance as I could using only 224 x 224 images. \n\nWhen examining the images, I observed that in many images, the lesion in question occupies a very small portion of the image. This means that a large part of the 224^2 pixels (presumably) do not provide a significant amount of information. I realized I could gain back information by cropping the images tightly around the lesion and before resizing, this would result in a much higher resolution version of the lesion while maintaining the 224 x 224 image size! \n\n![example of targeted image crop to increase lesion detail](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F9648ec1b82dbe3f3f07f9dba06ad4e5a%2Fcropping.png?generation=1596361389623025&amp;alt=media)\n\nThe amount of space occupied by the lesion varies greatly from image to image and other noise exists in the images (shadows, rulers, etc.) so this process cannot be trivially automated despite the simplicity of the images. I decided the best way to find salient regions is by having a model decide what information is important for classification. To accomplish this goal I decided to try to use attention. In this case I figured it would be pretty easy for the model to decide which pixels are lesion pixels and which pixels are regular skin without requiring explicit supervision. I devised the following model as a first attempt:\n\n![initial attention model](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F4055fb6338fde6dfdc4a475d2f2394a0%2Fmodel_diagram.png?generation=1596361447826651&amp;alt=media)\n\nThe model is a simple extension of the ResNet architecture, allowing for the use of pre-trained weights. The changes consist of removing the global average pooling from the end of the ResNet and replacing it with an attention-weighted pooling operation. The attention weights are determined by a series of two convolutional layers over the unpooled feature maps to produce a single 7x7 map of scores. The scores are softmax for two reasons:\n\n1. They cause attention weights to sum to 1 allowing for a simple weighted average.\n2. They force the model to \"choose\" which pixels it wants to keep while attenuating the rest.\n\nThe second point is what makes the model \"focus\" on certain regions. The model can still predict a uniform attention map, but when it does, the average pooling \"mixes\" the information causing a noisier representation. When the model predicts a concentrated attention map, the pooling selects the pixel features coming from discriminative regions of the image (in this case the skin lesion). All of this attention learning is driven by supervision from the task at hand, meaning that the regions that are selected are deemed important to the classification of the lesion. This means we can now extract salient regions from the image, as well as diagnose any spurious correlations picked up by the model.\n\nThis initial attempt worked surprisingly well, the model did in fact determine important regions of the image, but unfortunately the attention maps where not exact. I conjectured that this was because spatial locations of the input don't perfectly correspond to the same spatial locations of the resulting feature maps. This is due to the large effective receptive field size of each feature map location. To fix this issue, I devised a way to constrict the information that contributes to each location of feature map. I decided to divide the image into square pieces, referred to as chips,  and feed each one into a ResNet separately. With the global average pooling, this results in a 512 dimensional embedding for each chip. I could then concatenate the chips back together to create a feature map where each position corresponding directly to a fixed region of the input image:\n\n![improved attention model](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F94f8491eabe7625dd825ae5d950247f5%2Fchip_diagram.png?generation=1596361472945916&amp;alt=media)\n\nThis change resulted in much more consistent and accurate attention maps such as the following:\n\n![examples of good attention maps generated](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2Ffc37edb56413d616e74525528daa6a57%2Fattn_examples.png?generation=1596361507358931&amp;alt=media)\n\nThese are some selected samples of the attention maps and their corresponding input images. It is obvious that the model is focusing on the pixels that contain the legion. While the maps look satisfactory a majority of the time, there are failure cases:\n\n![examples of bad attention maps generated](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F23ae894ce6d277ffe51b4f52b956edf3%2Ffailures.png?generation=1596361541322606&amp;alt=media)\n\nI was not able to track down why these images generate poor attention maps, but I did observe that larger models greatly improved the consistency of the attention maps. The examples shown above were generated using a ResNet-50 backbone and the chipping model discussed earlier. Once I was able to generate satisfactory attention maps, I developed a process to convert these attention maps into square regions for cropping:\n\n![converting attention map to bounding box for cropping](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F72a86105602f2889775cccda4fc772d4%2Fattn_to_crop.png?generation=1596361577931234&amp;alt=media)\n\nWhen the attention fails the model generally generates a pretty unifrom attention map. This means that the bounding box creted encapsulates a very large portion of the image and therefore the cropping process defaults to maintaining the original image. With cropped versions of both the training set and the test set, a new model can be trained. By ensembling the predictions of the a ResNet trained on the original images, and a ResNet trained on the cropped images, we can get a decent boost in performance without having to increase the image size!\n\n### Implementation\n\n#### Preprocessing\n\nTo start I center-cropped all images into a square using the size of their shortest side.  Next I resized each image to be 224 x 224 and store the training set and test set as .npy files (for easy loading into RAM). Before going into any of the models the images go through a series of augmentations:\n```\ntrain_transform = transforms.Compose([\n    DrawHair(),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.RandomRotation(360),\n    transforms.ColorJitter(\n        brightness=[0.8, 1.2],\n        contrast=[0.8, 1.2],\n        saturation=[0.8, 1.2],\n    ),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n```\nAll of these are built into PyTorch except the `DrawHair` transform, which is a simple class I wrote which draws random curved lines over the image simulating hair. \n\nThe meta-data fields are converted to one-hot encodings. The one-hot encodings include an encoding for unknown/null. This process results in a 29-dimensional vector.\n\n#### Models\nThree models are trained as part of the previously described approach:\n1. resnet: A regular ResNet trained on the original images and provided metadata\n2. chipnet: The chip model used to generate attention maps\n3. cropnet: A regular ResNet trained on images cropped using attention\n\nEach of these models has a very similar training process. The training data consists of 33,126  images of varying size (32,542 benign, 584 malignant).  Metadata is included for each image, the meta-data fields are: age, anatomical site, gender, and specific diagnosis. These meta-data fields are converted to features and used as additional input to the models. Each model is trained and validated using 3-fold stratified cross-validation, this maintains the class distribution in each fold. The test set predictions of each fold are ensembled to produce the final predictions for each of the models. All models use focal loss to train, but would likely perform just as well with vanilla binary cross entropy. Additionally, test set predictions are generated by combining three rounds of test time augmentation using the augmentations described previously.\n\n##### ResNet50 model (resnet)\nThis model consists of a ResNet50 to process the image input, along with a two linear layers to process the meta-data features. The embeddings from these two branches are concatenated and sent through a final linear layer to produce a single score. This model is implemented in `resnet.py` by the class `MetaResNet`. The only variation between this model and cropnet model is the images used to train. The model is trained using the Adam optimizer, and a balanced class sampler (oversampling), meaning each batch contains an even number of each class. The ResNet portion of the model is trained with a learning rate of 1x10  and the rest of the parameters have a learning rate of 5x10. Training proceeds until the validation AUC does not improve for 3 epochs. \n\n##### Chipping model (chipnet)\nThis model also consists of a ResNet50 backbone, but has no additional meta-data branch as it's goal is simply to produce attention maps. As explained before, each 224 x 224 image is split into 32 x 32 non-overlapping chips. Each chip is sent through the ResNet50 and the resulting embeddings are re-stacked to create 7 x 7 feature maps with some number of channels depending on the backbone. To compute attention maps, two 3 x 3 same convolutions are applied (with ReLU and batch norm in between). The result of this is a 7 x 7 map of scores. These scores are softmaxed to sum to one and used to take a weighted average of the original feature maps. This process results in a single embedding that can then be used for classification after the application of a final linear layer. This model is implemented in `resnet.py` by the class `ChipNet`. The model is trained the same way as the resnet model except  the ResNet backbone is trained with a learning rate of 5x10 while the attention sub-net and final classifier are trained with a learning rate of 1x10.\n\n##### Cropped image model (cropnet)\nAs previously stated this model is trained identically to the resnet model except it uses the image cropped by the attention maps generated from the chipnet model.\n\n### Results\n\nI tried multiple ensembling techniques against the public leaderboard of the competition. Here are the results for the predictions of each model trained as part of the previously described approach:\n|model|public LB score (AUC ROC)  |\n|--|--|\n|ResNet-50 on original images (resnet)| **0.9203** |\n|Chip model (chipnet) | 0.8931|\n|ResNet-50 on cropped images (cropnet) | 0.9177 |\n\nThe resnet model achieved the highest performance on the public leaderboard, while chipnet model scored the lowest. This makes sense as the chipnet only looks at single tiles and therefore has a very constrained receptive field corresponding to each feature map location. This hinders performance while also improving attention map consistency and accuracy. The cropnet model achieves performance comparable to the resnet model, despite the attention maps sometimes producing poor crops of the original image. With a better chipnet model, I believe there would be a corresponding improvement in the performance of the cropnet model. I used these numbers to  generate ensemble weights for the each model's predictions (bad practice; leads to overfitting public LB), to see how much new information the chipnet and cropnet models add.\n\n|ensemble|public LB|\n|--|--|\n|(0.4 * resnet) + (0.2 * chipnet) + (0.4 * cropnet) | 0.9263|\n|(0.7 * resnet) + (0.3 * chipnet) | 0.9203 |\n|(0.5 * resnet) + (0.5 cropnet) | **0.9287** |\n\nTurns out using the predictions of the chipnet does not improve results and the best ensemble is a straight average of the resnet model and the cropnet model. Although public LB may not be the best performance indicator, I argue that it at least shows that this method is viable and extra information can be gleaned from the cropped images. \n\n### Conclusion\nAlthough the approach to medical image classification isn't necessarily novel, I believe there is still a lot of exploration to be done in this area. I could see these methods being applied to other vision tasks (for example landmark recognition). Additionally I hope this repository is at least slightly informational and encourages others to experiment beyond the typical image classification approaches. I would also like to note that this approach does not make use of pixel level segmentation supervision (although this type of data exists for this task) so it could be considered a \"weakly\" supervised technique. This means the approach could be used on tasks where no segmentation annotations are available.",
      "votes": 89
    },
    {
      "id": 955091,
      "postDate": "2020-08-02T10:29:31.820Z",
      "content": "<p>Really interesting approach. </p>\n\n<p>I  think you should pursue further the study and write a paper</p>",
      "rawMarkdown": "Really interesting approach. \n\nI  think you should pursue further the study and write a paper",
      "votes": 5
    },
    {
      "id": 965227,
      "postDate": "2020-08-10T13:34:48.410Z",
      "content": "<p>The awesome strategy of approach to medical image classification.</p>",
      "rawMarkdown": "The awesome strategy of approach to medical image classification.",
      "votes": 1
    },
    {
      "id": 964646,
      "postDate": "2020-08-10T04:27:53.793Z",
      "content": "<p>This is so cool! I was thinking about something like this but didn't think I would have time to explore it.</p>\n<p>I want to say that I really think this idea is impressive for this competition!</p>",
      "rawMarkdown": "This is so cool! I was thinking about something like this but didn't think I would have time to explore it.\n\nI want to say that I really think this idea is impressive for this competition!",
      "votes": 1
    },
    {
      "id": 960118,
      "postDate": "2020-08-06T06:41:54.117Z",
      "content": "<p><a href=\"/chriscareaga\">@chriscareaga</a> Awesome! Do you have any comment on why using 7x7 attention map? Can we increase it into a much larger size?</p>",
      "rawMarkdown": "@chriscareaga Awesome! Do you have any comment on why using 7x7 attention map? Can we increase it into a much larger size?",
      "votes": 1,
      "replies": [
        {
          "id": 960122,
          "postDate": "2020-08-06T06:49:32.033Z",
          "content": "<p>7x7 is the final layer output (and final layer is used to generate attention maps) of all state of the art CNN. You start with 224x224 then it gets halved 112x112 then half 56x56 then half 28x28 then half 14x14 then half 7x7</p>",
          "rawMarkdown": "7x7 is the final layer output (and final layer is used to generate attention maps) of all state of the art CNN. You start with 224x224 then it gets halved 112x112 then half 56x56 then half 28x28 then half 14x14 then half 7x7",
          "votes": 3
        },
        {
          "id": 960926,
          "postDate": "2020-08-06T19:46:50.407Z",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> </p>\n\n<p><a href=\"/cdeotte\">@cdeotte</a> is correct, with my original model I was using the penultimate feature maps of the ResNet (14x14) to generate the attention maps, but I found the the maps were not always precise even though they were more granular. When I switched to the chipnet model, 32x32 chips sounded like a good size and 224/32 = 7, so we end up with a 7x7 attention map. So in this case, they are not 7x7 because of the ResNet architecture, but rather because of the way I chose to split the input. You could very trivially alter the code to generate more granular attention maps, but at the cost of decreasing the chip size (and therefore decreasing the information in each chip). Alternatively, you could maintain the chip size, but adjust the stride (currently the chips are disjoint so stride = size). This would result in a more granular attention map but the mapping back to the image size is not as simple because the chips now overlap... All this to say, there are a lot of considerations to make, but I think it would be interesting to experiment with!</p>",
          "rawMarkdown": "@khahuras \n\n@cdeotte is correct, with my original model I was using the penultimate feature maps of the ResNet (14x14) to generate the attention maps, but I found the the maps were not always precise even though they were more granular. When I switched to the chipnet model, 32x32 chips sounded like a good size and 224/32 = 7, so we end up with a 7x7 attention map. So in this case, they are not 7x7 because of the ResNet architecture, but rather because of the way I chose to split the input. You could very trivially alter the code to generate more granular attention maps, but at the cost of decreasing the chip size (and therefore decreasing the information in each chip). Alternatively, you could maintain the chip size, but adjust the stride (currently the chips are disjoint so stride = size). This would result in a more granular attention map but the mapping back to the image size is not as simple because the chips now overlap... All this to say, there are a lot of considerations to make, but I think it would be interesting to experiment with!",
          "votes": 3
        }
      ]
    },
    {
      "id": 957638,
      "postDate": "2020-08-04T13:07:52.327Z",
      "content": "<p>Wow, This is some interesting work. Great stacking (ensembling) technique as well. I feel you can further improve this by trying out some Attention models and also maybe think about ways of improving the chipnet model</p>",
      "rawMarkdown": "Wow, This is some interesting work. Great stacking (ensembling) technique as well. I feel you can further improve this by trying out some Attention models and also maybe think about ways of improving the chipnet model",
      "votes": 1
    },
    {
      "id": 959406,
      "postDate": "2020-08-05T14:59:04.947Z",
      "content": "<p>amazing!</p>",
      "rawMarkdown": "amazing!"
    },
    {
      "id": 957552,
      "postDate": "2020-08-04T11:58:16.390Z",
      "content": "<p>Awesome strategy\nI have learned a lot from this discussion!</p>",
      "rawMarkdown": "Awesome strategy\nI have learned a lot from this discussion!\n"
    },
    {
      "id": 957487,
      "postDate": "2020-08-04T10:58:58.030Z",
      "content": "<p>This is pretty interesting. The chipping method is commonly used when processing large whole-slide microscopy scans. Maybe you can try sequence attention such as Transformer with your cropped patches to improve its performance.</p>",
      "rawMarkdown": "This is pretty interesting. The chipping method is commonly used when processing large whole-slide microscopy scans. Maybe you can try sequence attention such as Transformer with your cropped patches to improve its performance."
    },
    {
      "id": 971569,
      "postDate": "2020-08-15T17:16:12.983Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 955091,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-08-02T10:29:31.820000",
      "content": "<p>Really interesting approach. </p>\n\n<p>I  think you should pursue further the study and write a paper</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 965227,
      "author_name": "Henrique Silva",
      "author_url": "",
      "post_date": "2020-08-10T13:34:48.410000",
      "content": "<p>The awesome strategy of approach to medical image classification.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 964646,
      "author_name": "Campbell Hutcheson",
      "author_url": "",
      "post_date": "2020-08-10T04:27:53.793000",
      "content": "<p>This is so cool! I was thinking about something like this but didn't think I would have time to explore it.</p>\n<p>I want to say that I really think this idea is impressive for this competition!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 960118,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-08-06T06:41:54.117000",
      "content": "<p><a href=\"/chriscareaga\">@chriscareaga</a> Awesome! Do you have any comment on why using 7x7 attention map? Can we increase it into a much larger size?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 960122,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-06T06:49:32.033000",
          "content": "<p>7x7 is the final layer output (and final layer is used to generate attention maps) of all state of the art CNN. You start with 224x224 then it gets halved 112x112 then half 56x56 then half 28x28 then half 14x14 then half 7x7</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 960926,
          "author_name": "chris",
          "author_url": "",
          "post_date": "2020-08-06T19:46:50.407000",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> </p>\n\n<p><a href=\"/cdeotte\">@cdeotte</a> is correct, with my original model I was using the penultimate feature maps of the ResNet (14x14) to generate the attention maps, but I found the the maps were not always precise even though they were more granular. When I switched to the chipnet model, 32x32 chips sounded like a good size and 224/32 = 7, so we end up with a 7x7 attention map. So in this case, they are not 7x7 because of the ResNet architecture, but rather because of the way I chose to split the input. You could very trivially alter the code to generate more granular attention maps, but at the cost of decreasing the chip size (and therefore decreasing the information in each chip). Alternatively, you could maintain the chip size, but adjust the stride (currently the chips are disjoint so stride = size). This would result in a more granular attention map but the mapping back to the image size is not as simple because the chips now overlap... All this to say, there are a lot of considerations to make, but I think it would be interesting to experiment with!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 957638,
      "author_name": "Shyam R",
      "author_url": "",
      "post_date": "2020-08-04T13:07:52.327000",
      "content": "<p>Wow, This is some interesting work. Great stacking (ensembling) technique as well. I feel you can further improve this by trying out some Attention models and also maybe think about ways of improving the chipnet model</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 959406,
      "author_name": "Lakshmi47",
      "author_url": "",
      "post_date": "2020-08-05T14:59:04.947000",
      "content": "<p>amazing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 957552,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2020-08-04T11:58:16.390000",
      "content": "<p>Awesome strategy\nI have learned a lot from this discussion!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 957487,
      "author_name": "Kaiming Kuang",
      "author_url": "",
      "post_date": "2020-08-04T10:58:58.030000",
      "content": "<p>This is pretty interesting. The chipping method is commonly used when processing large whole-slide microscopy scans. Maybe you can try sequence attention such as Transformer with your cropped patches to improve its performance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 971569,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-15T17:16:12.983000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "955067": "I sort of gave up on trying to compete with the leaderboard after awhile and decided to just use this competition to experiment with various ideas I have. I ended up coming up with an approach that I figured I'd share. I have created a GitHub repo documenting my findings:\n\n[Attention Guided Cropping](https://github.com/CCareaga/attention_guided_cropping)\n\nHere is a the main idea ripped from the README of the repo:\n\n### Disclaimer\nAfter working on this project I searched around a bit for related work, and found:\n\n[Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification](https://arxiv.org/pdf/1801.09927.pdf)\n\nThis paper uses a eerily similar approach, although it doesn't seem like they use the cropping technique to cross-reference original high-resolution images, and their attention technique varies from mine. Credit should definitely go to these authors for first experimenting with this idea. I also assume many others have tried methods like this (if so please let me know so I can add credit). I know that many competition participants likely also experimented with intelligent cropping, but I figured since I put time into this approach I would still document my findings.\n\n### Approach\n\nThe dermatoscopy images provided as part of the competition are very high resolution, the largest images in the dataset are 4000x6000 pixels. When using pre-trained networks, the images for the downstream task are typically resized to match the size of the pre-training task. For example when using an ImageNet pre-trained CNN, one will typically resize images to 224 x 224, as this is what the network is \"used\" to. As many found out during the competition, this was not necessary due to scale invariance, as well as the vast difference between pre-training tasks and the competition task. Participants began ensembling models trained on various image sizes, anywhere from 224 x 224 all the way up to 1024 x 1024. Due to my limited resources, and aversion to online notebooks, I decided I wouldn't really be able to compete with large models and large image sizes. I instead made a personal challenge to squeeze as much performance as I could using only 224 x 224 images. \n\nWhen examining the images, I observed that in many images, the lesion in question occupies a very small portion of the image. This means that a large part of the 224^2 pixels (presumably) do not provide a significant amount of information. I realized I could gain back information by cropping the images tightly around the lesion and before resizing, this would result in a much higher resolution version of the lesion while maintaining the 224 x 224 image size! \n\n![example of targeted image crop to increase lesion detail](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F9648ec1b82dbe3f3f07f9dba06ad4e5a%2Fcropping.png?generation=1596361389623025&amp;alt=media)\n\nThe amount of space occupied by the lesion varies greatly from image to image and other noise exists in the images (shadows, rulers, etc.) so this process cannot be trivially automated despite the simplicity of the images. I decided the best way to find salient regions is by having a model decide what information is important for classification. To accomplish this goal I decided to try to use attention. In this case I figured it would be pretty easy for the model to decide which pixels are lesion pixels and which pixels are regular skin without requiring explicit supervision. I devised the following model as a first attempt:\n\n![initial attention model](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F4055fb6338fde6dfdc4a475d2f2394a0%2Fmodel_diagram.png?generation=1596361447826651&amp;alt=media)\n\nThe model is a simple extension of the ResNet architecture, allowing for the use of pre-trained weights. The changes consist of removing the global average pooling from the end of the ResNet and replacing it with an attention-weighted pooling operation. The attention weights are determined by a series of two convolutional layers over the unpooled feature maps to produce a single 7x7 map of scores. The scores are softmax for two reasons:\n\n1. They cause attention weights to sum to 1 allowing for a simple weighted average.\n2. They force the model to \"choose\" which pixels it wants to keep while attenuating the rest.\n\nThe second point is what makes the model \"focus\" on certain regions. The model can still predict a uniform attention map, but when it does, the average pooling \"mixes\" the information causing a noisier representation. When the model predicts a concentrated attention map, the pooling selects the pixel features coming from discriminative regions of the image (in this case the skin lesion). All of this attention learning is driven by supervision from the task at hand, meaning that the regions that are selected are deemed important to the classification of the lesion. This means we can now extract salient regions from the image, as well as diagnose any spurious correlations picked up by the model.\n\nThis initial attempt worked surprisingly well, the model did in fact determine important regions of the image, but unfortunately the attention maps where not exact. I conjectured that this was because spatial locations of the input don't perfectly correspond to the same spatial locations of the resulting feature maps. This is due to the large effective receptive field size of each feature map location. To fix this issue, I devised a way to constrict the information that contributes to each location of feature map. I decided to divide the image into square pieces, referred to as chips,  and feed each one into a ResNet separately. With the global average pooling, this results in a 512 dimensional embedding for each chip. I could then concatenate the chips back together to create a feature map where each position corresponding directly to a fixed region of the input image:\n\n![improved attention model](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F94f8491eabe7625dd825ae5d950247f5%2Fchip_diagram.png?generation=1596361472945916&amp;alt=media)\n\nThis change resulted in much more consistent and accurate attention maps such as the following:\n\n![examples of good attention maps generated](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2Ffc37edb56413d616e74525528daa6a57%2Fattn_examples.png?generation=1596361507358931&amp;alt=media)\n\nThese are some selected samples of the attention maps and their corresponding input images. It is obvious that the model is focusing on the pixels that contain the legion. While the maps look satisfactory a majority of the time, there are failure cases:\n\n![examples of bad attention maps generated](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F23ae894ce6d277ffe51b4f52b956edf3%2Ffailures.png?generation=1596361541322606&amp;alt=media)\n\nI was not able to track down why these images generate poor attention maps, but I did observe that larger models greatly improved the consistency of the attention maps. The examples shown above were generated using a ResNet-50 backbone and the chipping model discussed earlier. Once I was able to generate satisfactory attention maps, I developed a process to convert these attention maps into square regions for cropping:\n\n![converting attention map to bounding box for cropping](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3095100%2F72a86105602f2889775cccda4fc772d4%2Fattn_to_crop.png?generation=1596361577931234&amp;alt=media)\n\nWhen the attention fails the model generally generates a pretty unifrom attention map. This means that the bounding box creted encapsulates a very large portion of the image and therefore the cropping process defaults to maintaining the original image. With cropped versions of both the training set and the test set, a new model can be trained. By ensembling the predictions of the a ResNet trained on the original images, and a ResNet trained on the cropped images, we can get a decent boost in performance without having to increase the image size!\n\n### Implementation\n\n#### Preprocessing\n\nTo start I center-cropped all images into a square using the size of their shortest side.  Next I resized each image to be 224 x 224 and store the training set and test set as .npy files (for easy loading into RAM). Before going into any of the models the images go through a series of augmentations:\n```\ntrain_transform = transforms.Compose([\n    DrawHair(),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.RandomRotation(360),\n    transforms.ColorJitter(\n        brightness=[0.8, 1.2],\n        contrast=[0.8, 1.2],\n        saturation=[0.8, 1.2],\n    ),\n    transforms.ToTensor(),\n    transforms.Normalize(mean=[0.485, 0.456, 0.406],std=[0.229, 0.224, 0.225])\n])\n```\nAll of these are built into PyTorch except the `DrawHair` transform, which is a simple class I wrote which draws random curved lines over the image simulating hair. \n\nThe meta-data fields are converted to one-hot encodings. The one-hot encodings include an encoding for unknown/null. This process results in a 29-dimensional vector.\n\n#### Models\nThree models are trained as part of the previously described approach:\n1. resnet: A regular ResNet trained on the original images and provided metadata\n2. chipnet: The chip model used to generate attention maps\n3. cropnet: A regular ResNet trained on images cropped using attention\n\nEach of these models has a very similar training process. The training data consists of 33,126  images of varying size (32,542 benign, 584 malignant).  Metadata is included for each image, the meta-data fields are: age, anatomical site, gender, and specific diagnosis. These meta-data fields are converted to features and used as additional input to the models. Each model is trained and validated using 3-fold stratified cross-validation, this maintains the class distribution in each fold. The test set predictions of each fold are ensembled to produce the final predictions for each of the models. All models use focal loss to train, but would likely perform just as well with vanilla binary cross entropy. Additionally, test set predictions are generated by combining three rounds of test time augmentation using the augmentations described previously.\n\n##### ResNet50 model (resnet)\nThis model consists of a ResNet50 to process the image input, along with a two linear layers to process the meta-data features. The embeddings from these two branches are concatenated and sent through a final linear layer to produce a single score. This model is implemented in `resnet.py` by the class `MetaResNet`. The only variation between this model and cropnet model is the images used to train. The model is trained using the Adam optimizer, and a balanced class sampler (oversampling), meaning each batch contains an even number of each class. The ResNet portion of the model is trained with a learning rate of 1x10  and the rest of the parameters have a learning rate of 5x10. Training proceeds until the validation AUC does not improve for 3 epochs. \n\n##### Chipping model (chipnet)\nThis model also consists of a ResNet50 backbone, but has no additional meta-data branch as it's goal is simply to produce attention maps. As explained before, each 224 x 224 image is split into 32 x 32 non-overlapping chips. Each chip is sent through the ResNet50 and the resulting embeddings are re-stacked to create 7 x 7 feature maps with some number of channels depending on the backbone. To compute attention maps, two 3 x 3 same convolutions are applied (with ReLU and batch norm in between). The result of this is a 7 x 7 map of scores. These scores are softmaxed to sum to one and used to take a weighted average of the original feature maps. This process results in a single embedding that can then be used for classification after the application of a final linear layer. This model is implemented in `resnet.py` by the class `ChipNet`. The model is trained the same way as the resnet model except  the ResNet backbone is trained with a learning rate of 5x10 while the attention sub-net and final classifier are trained with a learning rate of 1x10.\n\n##### Cropped image model (cropnet)\nAs previously stated this model is trained identically to the resnet model except it uses the image cropped by the attention maps generated from the chipnet model.\n\n### Results\n\nI tried multiple ensembling techniques against the public leaderboard of the competition. Here are the results for the predictions of each model trained as part of the previously described approach:\n|model|public LB score (AUC ROC)  |\n|--|--|\n|ResNet-50 on original images (resnet)| **0.9203** |\n|Chip model (chipnet) | 0.8931|\n|ResNet-50 on cropped images (cropnet) | 0.9177 |\n\nThe resnet model achieved the highest performance on the public leaderboard, while chipnet model scored the lowest. This makes sense as the chipnet only looks at single tiles and therefore has a very constrained receptive field corresponding to each feature map location. This hinders performance while also improving attention map consistency and accuracy. The cropnet model achieves performance comparable to the resnet model, despite the attention maps sometimes producing poor crops of the original image. With a better chipnet model, I believe there would be a corresponding improvement in the performance of the cropnet model. I used these numbers to  generate ensemble weights for the each model's predictions (bad practice; leads to overfitting public LB), to see how much new information the chipnet and cropnet models add.\n\n|ensemble|public LB|\n|--|--|\n|(0.4 * resnet) + (0.2 * chipnet) + (0.4 * cropnet) | 0.9263|\n|(0.7 * resnet) + (0.3 * chipnet) | 0.9203 |\n|(0.5 * resnet) + (0.5 cropnet) | **0.9287** |\n\nTurns out using the predictions of the chipnet does not improve results and the best ensemble is a straight average of the resnet model and the cropnet model. Although public LB may not be the best performance indicator, I argue that it at least shows that this method is viable and extra information can be gleaned from the cropped images. \n\n### Conclusion\nAlthough the approach to medical image classification isn't necessarily novel, I believe there is still a lot of exploration to be done in this area. I could see these methods being applied to other vision tasks (for example landmark recognition). Additionally I hope this repository is at least slightly informational and encourages others to experiment beyond the typical image classification approaches. I would also like to note that this approach does not make use of pixel level segmentation supervision (although this type of data exists for this task) so it could be considered a \"weakly\" supervised technique. This means the approach could be used on tasks where no segmentation annotations are available.",
    "955091": "Really interesting approach. \n\nI  think you should pursue further the study and write a paper",
    "965227": "The awesome strategy of approach to medical image classification.",
    "964646": "This is so cool! I was thinking about something like this but didn't think I would have time to explore it.\n\nI want to say that I really think this idea is impressive for this competition!",
    "960118": "@chriscareaga Awesome! Do you have any comment on why using 7x7 attention map? Can we increase it into a much larger size?",
    "957638": "Wow, This is some interesting work. Great stacking (ensembling) technique as well. I feel you can further improve this by trying out some Attention models and also maybe think about ways of improving the chipnet model",
    "959406": "amazing!",
    "957552": "Awesome strategy\nI have learned a lot from this discussion!\n",
    "957487": "This is pretty interesting. The chipping method is commonly used when processing large whole-slide microscopy scans. Maybe you can try sequence attention such as Transformer with your cropped patches to improve its performance.",
    "971569": ""
  }
}