{
  "id": 181830,
  "title": "27th place solution - 2nd with context prize (edited)",
  "url": "/competitions/siim-isic-melanoma-classification/writeups/yuval-nosound-27th-place-solution-2nd-with-context",
  "author_name": "",
  "post_date": "2020-09-15T14:29:27.747Z",
  "votes": 34,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This is a summery of \"Yuval and nosound\" model (27th place)</p>\n<p>You can find more in:</p>\n<ul>\n<li><a href=\"https://github.com/yuval6957/SIIM-Transformer\" target=\"_blank\">our github repository</a> </li>\n<li><a href=\"https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20paper.pdf\" target=\"_blank\">Paper</a></li>\n<li><a href=\"https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20presentation.pdf\" target=\"_blank\">Presentation</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=qa6zimKQcno&amp;t=15s\" target=\"_blank\">video</a></li>\n</ul>\n<h2>1. Summary</h2>\n<p>Our solution is based on two step model + Ensemble:</p>\n<ol>\n<li>Base model for feature extraction per image</li>\n<li>Transformer model - combining all the output features from a patient and predict per image. </li>\n<li>The 2nd stage also included some post - processing and ensembling.</li>\n</ol>\n<h3>Base Model:</h3>\n<p>As base model we used a models from the <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet</a> family[6]:</p>\n<ul>\n<li>EfficientNet b3 </li>\n<li>EfficientNet b4 </li>\n<li>EfficientNet b5 </li>\n<li>EfficientNet b6 </li>\n<li>EfficientNet b7 </li>\n</ul>\n<p>All models were pre-trained on imagenet using noisy student algorithm. The models and weights are from  <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch\" target=\"_blank\">gen-efficientnet-pytorch</a>[2].</p>\n<p>The input to these model is the image and meta-data such as age, sex, and anatomic Site. The meta-data is  processed by a small fully connected network and it’s output is concatenated to the input of the classification layer of the original EfficientNet network. This vector is going through a linear layer with output size of 256 to create the “features”, and then after an activation layer to the final linear classification layer. </p>\n<p>This network has 8 outputs and tries to classify the diagnosis label (there are actually more than 8 possible diagnoses, but some don’t have enough examples). </p>\n<h3>Transformer Models:</h3>\n<p>The input to the Transformer models are a stack of features from all images belonging to the same patient + the metadata for these images.</p>\n<p>The transformer is a stack of 4 transformer encoder layers with self attention as described in <a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">Attention Is All You Need</a> [1]. Each transformer encoder layer uses 4 self attention heads. </p>\n<p>The output of the transformer is a N*C, where N is the number of input feature vectors (the number of images) and C is the number of classes (8 in this case). Hence, the transformer predicts the class of each feature vector simultaneously, using the information from all other feature vectors.</p>\n<p>The metadata is added using a “transformer style”, i.e. each parameter is transformed to a vector (size 256) using an embedding matrix and then added to the feature vector. for continuous values (like age) the embedding matrix was replaced by a 2 layer fully connected network.   </p>\n<h3>Ensembling the output of all networks:</h3>\n<p>The data was split to 3 folds, 3 times (using 3 different seeds for splitting), and the inference was done using 16 (or 12) TTAs. giving 144 predictions from each model. These were averaged and then the outputs of all the models were averaged. All averaging was done on the outputs before softmax and there form it is actually geometric averaging.</p>\n<h3>Training</h3>\n<p>The heavy lifting was the training and inference of the base models. This was done on a server with 2 GPUs – Tesla V100, Titan RTX that worked in parallel on different tasks. Training one fold of one model took 3H (B3)  to 11H (B7, B6 large images) on the Tesla and 20% more on the Titan, this sums up to about one day for B3 and 3.5 days for B7. Inferencing all the training data  for 12 TTA’s + test data for 16 TTA’s to get the features for the next level took  another 4h - 14h. The transformer training took less than 1H for the full model (3 folds*3seed). </p>\n<p>The total time it took to train all models and folds is about 2.5W for one Tesla (~1.5W using the 2 GPUs).</p>\n<h2>2. Models and features</h2>\n<h3>Base models</h3>\n<p>As base models we tried various types of models (pre-trained on Imagenet):</p>\n<ul>\n<li>Densenets – 121, 161, 169, 201</li>\n<li>EfficientNet B0, B3 , B4, B5, B6, B7 with and with noisy student pre-training and with normal pretraining </li>\n<li>ResNet 101</li>\n<li>Xception </li>\n</ul>\n<p>At the end we used EfficientNet as it was best when judging accuracy/time</p>\n<p>The noisy student version performed better than the normal one.</p>\n<p>We also tried different image sizes and ended up using a <code>400*600</code> images in most cases, except one were we used <code>600*900</code> with the B6 network.</p>\n<h4>Metadata</h4>\n<p>As was described above the metadata was processed by a small fully connected nn and its output was concatenated to the output of the EfficientNet network (after removing the original top layer).</p>\n<p>We also tried a network without metadata, and used the metadata as targets, i.e. this network predicted the diagnosis, but also the sex, age and anatomic site. The final predictions (including transformer) when using this approach weren't as good as the metadata as input approach. </p>\n<h4>Model’s output</h4>\n<p>Although the task at hand is to predict melanoma yes/no, it is better to let the network choose the diagnosis among a few possible options. This lets the network “understand” more about the image. The final prediction is the value of the Melanoma output after doing softmax on the output vector. </p>\n<h4>Features</h4>\n<p>The final layer in this model is a linear layer with 256 inputs and 8 outputs, we use the input to this layer as features. </p>\n<h4>Augmentation</h4>\n<p>The following augmentations where used while training and inference:</p>\n<p>Random resize + crop</p>\n<p>Random rotation</p>\n<p>Random flip</p>\n<p>Random color jitter (brightness, contrast, saturation, hue)</p>\n<p><a href=\"https://arxiv.org/abs/1708.04552\" target=\"_blank\">Cutout</a>[3] - erasing a small rectangle in the image</p>\n<p>Hair - Randomly adding “hair like” lines to the image</p>\n<p>Metadata augmentation - adding random noise to the metadata as was done in the <a href=\"https://www.sciencedirect.com/science/article/pii/S2215016120300832?via%3Dihub\" target=\"_blank\">1st place solution in ISIC 2019 challenge</a> [7].</p>\n<h4>TTA</h4>\n<p>For inference each image was augmented differently 16 times and the final prediction was the average. These augmentations were also used for extracting 16 different features vectors per test image.</p>\n<p>The same was done to extract 12 features vectors for the train images (12 and not 16 because of time limits).</p>\n<h3>Transformer Network</h3>\n<p>The input to the transformer network is the features from all the images from one patient.</p>\n<p>The inspiration for this kind of model came from a previous competition in which we participated in RSNA<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection\" target=\"_blank\"> Intracranial Hemorrhage Detection</a>. In that competition, all the top solutions (including ours) used a two stage network approach (although none of them were transformers).</p>\n<p>Using a transformer seems appropriate in this case because transformers are built to seek relationships between embedding (feature) vectors in it’s input.</p>\n<p>As this is not a full seq2seq task, we only used encoder layers. The transformer is a stack of 4 encoder layers with 4 attention heads in each layer. (we also tested higher numbers of layers and attention heads - no performance improvement).</p>\n<h4>metadata</h4>\n<p>The metadata was incorporated in the network by adding “metadata vectors” to the input vectors - each value was transformed to a vector size 256 and added. The discrete values’  transformation was done using a trainable embedding matrix and the continuous values using a small nn.</p>\n<h4>output</h4>\n<p>The output of this network is a matrix of size N*C where N - the number of images, C - the number of classes. Which means it decides on all the images of the patient at once.     </p>\n<h4>Limit the input size</h4>\n<p>A transformer can be trained on different number of feature vectors, by using padding. But when the range of numbers is very large, from a couple of hundred images for some patients to a handful for others, this may cause some implementation issues (like in calculating the loss). To simplify these issues, we limited N to 24 feature vectors, and for each patient we randomly divided the images to groups of size up to 24. </p>\n<p>This might degrade the prediction as the most “similar” images might accidentally fall into different groups, but as we use TTA, this issue is almost solved. </p>\n<h4>Augmentation</h4>\n<p>From the base model we extract a number of feature vectors (12 for train and 16 for test) using different augmentation for the images and metadata. In the training and inference steps of the transformer model we randomly choose one of these vectors.</p>\n<p>Another augmentation is the random grouping as stated above.</p>\n<h2>3. Training and Inferencing</h2>\n<p>The original 2020 competition data is highly unbalanced, there are only 2-3% of positive targets in the train and test data. Although we were able to train the base model using uneven sampling, the best way to get good training was to add the data from ISIC 2019 competition which has a much higher percentage of melanoma images. </p>\n<p>We split the training data to 3 folds keeping all the images from the same patient in the same fold, and making sure each fold has a similar number of patients with melanoma. The ISIC2019’s data was also split evenly between the folds. The same folds were kept for the base and the transformer models. </p>\n<p>To get more diversity we had 3 different splits using 3 seeds </p>\n<h3>Preprocessing</h3>\n<p>All images were resized to an aspect ratio of 1:1.5, which was the most popular aspect ratio of the images in the original dataset. We prepared 3 image datasets of sizes <code>300*450</code>, <code>400*600</code>, <code>600*900</code>. Most of the models were trained using the <code>400*600</code> dataset, as 300*450 gave inferior results and the <code>600*900</code> didn’t improve the results enough.</p>\n<p>For the metadata we had to set the same terminology for the 2020 and 2019 datasets.</p>\n<h3>Loss Function</h3>\n<p>The loss function we used was cross entropy. Although the task is to predict only melanoma we found it is better to predict the diagnosis which was split to 8 different classes one of which was melanoma. The final prediction was the value for the melanoma class, after a softmax function on all classes. We also tried a binary cross entropy on the melanoma class alone and a combination between the two, but using cross entropy gave the best results.</p>\n<p>The same loss was used for the base model and the transformer, but in the transformer we needed to regularize for the different number of predictions in each batch resulting from the different number of images for each patient. </p>\n<p>We also tried using focal loss which didn’t improve the results, but we left one of the transformer models which was trained with focal loss in the ensemble (A model with cross entropy loss gave the similar CV and LB).</p>\n<h4>Training the transformer model</h4>\n<p>The transformer model was trained in two steps. For the first step we used the data from both competitions (2019, 2020). For the 2019 competition we don’t have information about the patient, and each image got a different dummy patient_id, meaning the transformer didn’t learn much from these images. In the 2nd stage we fine-tuned the transformer using only the 2020 competition’s data.</p>\n<p>In both steps we used a sampler that over sampled the larger groups.</p>\n<h4>Inference</h4>\n<p>As stated above, the inference was done using TTA. For the base model we used 12-16 different augmentations and for the transformer model 32.</p>\n<h3>Ensembling</h3>\n<p>For our final submissions we used 2 ensembles:</p>\n<h5>Without Context Submission:</h5>\n<ol>\n<li>EfficientNet B3 noisy student image size 400*600</li>\n<li>EfficientNet B4 noisy student image size 400*600</li>\n<li>EfficientNet B5 noisy student image size 400*600</li>\n<li>EfficientNet B6 noisy student image size <strong>600*900</strong></li>\n<li>EfficientNet B7 noisy student image size 400*600</li>\n</ol>\n<h5>With Context</h5>\n<p>All the “without context” model +</p>\n<ul>\n<li>Transformer on features from A.</li>\n<li>Transformer on features from B.</li>\n<li>Transformer on features from C using focal loss</li>\n<li>Transformer on features from D.</li>\n<li>Transformer on features from E.</li>\n</ul>\n<h2>4. Interesting findings</h2>\n<h3>CV and LB</h3>\n<p>Although the number of images in the competition was very large, the number of different patients wasn’t large enough to give a stable and reliable CV and LB. And the correlation between the two was low. At the end we trusted neither and we submitted the models we felt were the most robust.</p>\n<h3>What didn’t work</h3>\n<ol>\n<li>Using different sizes of images as suggested in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">CNN Input Size Explained</a> didn’t show any improvement.</li>\n<li>Using Mixup[5] and MixCut[4] augmentation didn’t work</li>\n<li>Larger transformers didn’t improve the results. </li>\n</ol>\n<h2>5. Better “Real World” model</h2>\n<p>This model can be much simpler if we won’t ensemble and take only one base model, the EfficientNet B5 is probably the best compromise.</p>\n<p>In the real world scenario the images from previous years will probably be tagged already. In that case we can use a full transformer with encoder and decoder which will perform a seq2seq operation. </p>",
  "messages": [
    {
      "id": "1005162",
      "postDate": "09/10/2020 09:45:59",
      "content": "<p>This is a summery of \"Yuval and nosound\" model (27th place)</p>\n<p>You can find more in:</p>\n<ul>\n<li><a href=\"https://github.com/yuval6957/SIIM-Transformer\" target=\"_blank\">our github repository</a> </li>\n<li><a href=\"https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20paper.pdf\" target=\"_blank\">Paper</a></li>\n<li><a href=\"https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20presentation.pdf\" target=\"_blank\">Presentation</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=qa6zimKQcno&amp;t=15s\" target=\"_blank\">video</a></li>\n</ul>\n<h2>1. Summary</h2>\n<p>Our solution is based on two step model + Ensemble:</p>\n<ol>\n<li>Base model for feature extraction per image</li>\n<li>Transformer model - combining all the output features from a patient and predict per image. </li>\n<li>The 2nd stage also included some post - processing and ensembling.</li>\n</ol>\n<h3>Base Model:</h3>\n<p>As base model we used a models from the <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet</a> family[6]:</p>\n<ul>\n<li>EfficientNet b3 </li>\n<li>EfficientNet b4 </li>\n<li>EfficientNet b5 </li>\n<li>EfficientNet b6 </li>\n<li>EfficientNet b7 </li>\n</ul>\n<p>All models were pre-trained on imagenet using noisy student algorithm. The models and weights are from  <a href=\"https://github.com/rwightman/gen-efficientnet-pytorch\" target=\"_blank\">gen-efficientnet-pytorch</a>[2].</p>\n<p>The input to these model is the image and meta-data such as age, sex, and anatomic Site. The meta-data is  processed by a small fully connected network and it’s output is concatenated to the input of the classification layer of the original EfficientNet network. This vector is going through a linear layer with output size of 256 to create the “features”, and then after an activation layer to the final linear classification layer. </p>\n<p>This network has 8 outputs and tries to classify the diagnosis label (there are actually more than 8 possible diagnoses, but some don’t have enough examples). </p>\n<h3>Transformer Models:</h3>\n<p>The input to the Transformer models are a stack of features from all images belonging to the same patient + the metadata for these images.</p>\n<p>The transformer is a stack of 4 transformer encoder layers with self attention as described in <a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">Attention Is All You Need</a> [1]. Each transformer encoder layer uses 4 self attention heads. </p>\n<p>The output of the transformer is a N*C, where N is the number of input feature vectors (the number of images) and C is the number of classes (8 in this case). Hence, the transformer predicts the class of each feature vector simultaneously, using the information from all other feature vectors.</p>\n<p>The metadata is added using a “transformer style”, i.e. each parameter is transformed to a vector (size 256) using an embedding matrix and then added to the feature vector. for continuous values (like age) the embedding matrix was replaced by a 2 layer fully connected network.   </p>\n<h3>Ensembling the output of all networks:</h3>\n<p>The data was split to 3 folds, 3 times (using 3 different seeds for splitting), and the inference was done using 16 (or 12) TTAs. giving 144 predictions from each model. These were averaged and then the outputs of all the models were averaged. All averaging was done on the outputs before softmax and there form it is actually geometric averaging.</p>\n<h3>Training</h3>\n<p>The heavy lifting was the training and inference of the base models. This was done on a server with 2 GPUs – Tesla V100, Titan RTX that worked in parallel on different tasks. Training one fold of one model took 3H (B3)  to 11H (B7, B6 large images) on the Tesla and 20% more on the Titan, this sums up to about one day for B3 and 3.5 days for B7. Inferencing all the training data  for 12 TTA’s + test data for 16 TTA’s to get the features for the next level took  another 4h - 14h. The transformer training took less than 1H for the full model (3 folds*3seed). </p>\n<p>The total time it took to train all models and folds is about 2.5W for one Tesla (~1.5W using the 2 GPUs).</p>\n<h2>2. Models and features</h2>\n<h3>Base models</h3>\n<p>As base models we tried various types of models (pre-trained on Imagenet):</p>\n<ul>\n<li>Densenets – 121, 161, 169, 201</li>\n<li>EfficientNet B0, B3 , B4, B5, B6, B7 with and with noisy student pre-training and with normal pretraining </li>\n<li>ResNet 101</li>\n<li>Xception </li>\n</ul>\n<p>At the end we used EfficientNet as it was best when judging accuracy/time</p>\n<p>The noisy student version performed better than the normal one.</p>\n<p>We also tried different image sizes and ended up using a <code>400*600</code> images in most cases, except one were we used <code>600*900</code> with the B6 network.</p>\n<h4>Metadata</h4>\n<p>As was described above the metadata was processed by a small fully connected nn and its output was concatenated to the output of the EfficientNet network (after removing the original top layer).</p>\n<p>We also tried a network without metadata, and used the metadata as targets, i.e. this network predicted the diagnosis, but also the sex, age and anatomic site. The final predictions (including transformer) when using this approach weren't as good as the metadata as input approach. </p>\n<h4>Model’s output</h4>\n<p>Although the task at hand is to predict melanoma yes/no, it is better to let the network choose the diagnosis among a few possible options. This lets the network “understand” more about the image. The final prediction is the value of the Melanoma output after doing softmax on the output vector. </p>\n<h4>Features</h4>\n<p>The final layer in this model is a linear layer with 256 inputs and 8 outputs, we use the input to this layer as features. </p>\n<h4>Augmentation</h4>\n<p>The following augmentations where used while training and inference:</p>\n<p>Random resize + crop</p>\n<p>Random rotation</p>\n<p>Random flip</p>\n<p>Random color jitter (brightness, contrast, saturation, hue)</p>\n<p><a href=\"https://arxiv.org/abs/1708.04552\" target=\"_blank\">Cutout</a>[3] - erasing a small rectangle in the image</p>\n<p>Hair - Randomly adding “hair like” lines to the image</p>\n<p>Metadata augmentation - adding random noise to the metadata as was done in the <a href=\"https://www.sciencedirect.com/science/article/pii/S2215016120300832?via%3Dihub\" target=\"_blank\">1st place solution in ISIC 2019 challenge</a> [7].</p>\n<h4>TTA</h4>\n<p>For inference each image was augmented differently 16 times and the final prediction was the average. These augmentations were also used for extracting 16 different features vectors per test image.</p>\n<p>The same was done to extract 12 features vectors for the train images (12 and not 16 because of time limits).</p>\n<h3>Transformer Network</h3>\n<p>The input to the transformer network is the features from all the images from one patient.</p>\n<p>The inspiration for this kind of model came from a previous competition in which we participated in RSNA<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection\" target=\"_blank\"> Intracranial Hemorrhage Detection</a>. In that competition, all the top solutions (including ours) used a two stage network approach (although none of them were transformers).</p>\n<p>Using a transformer seems appropriate in this case because transformers are built to seek relationships between embedding (feature) vectors in it’s input.</p>\n<p>As this is not a full seq2seq task, we only used encoder layers. The transformer is a stack of 4 encoder layers with 4 attention heads in each layer. (we also tested higher numbers of layers and attention heads - no performance improvement).</p>\n<h4>metadata</h4>\n<p>The metadata was incorporated in the network by adding “metadata vectors” to the input vectors - each value was transformed to a vector size 256 and added. The discrete values’  transformation was done using a trainable embedding matrix and the continuous values using a small nn.</p>\n<h4>output</h4>\n<p>The output of this network is a matrix of size N*C where N - the number of images, C - the number of classes. Which means it decides on all the images of the patient at once.     </p>\n<h4>Limit the input size</h4>\n<p>A transformer can be trained on different number of feature vectors, by using padding. But when the range of numbers is very large, from a couple of hundred images for some patients to a handful for others, this may cause some implementation issues (like in calculating the loss). To simplify these issues, we limited N to 24 feature vectors, and for each patient we randomly divided the images to groups of size up to 24. </p>\n<p>This might degrade the prediction as the most “similar” images might accidentally fall into different groups, but as we use TTA, this issue is almost solved. </p>\n<h4>Augmentation</h4>\n<p>From the base model we extract a number of feature vectors (12 for train and 16 for test) using different augmentation for the images and metadata. In the training and inference steps of the transformer model we randomly choose one of these vectors.</p>\n<p>Another augmentation is the random grouping as stated above.</p>\n<h2>3. Training and Inferencing</h2>\n<p>The original 2020 competition data is highly unbalanced, there are only 2-3% of positive targets in the train and test data. Although we were able to train the base model using uneven sampling, the best way to get good training was to add the data from ISIC 2019 competition which has a much higher percentage of melanoma images. </p>\n<p>We split the training data to 3 folds keeping all the images from the same patient in the same fold, and making sure each fold has a similar number of patients with melanoma. The ISIC2019’s data was also split evenly between the folds. The same folds were kept for the base and the transformer models. </p>\n<p>To get more diversity we had 3 different splits using 3 seeds </p>\n<h3>Preprocessing</h3>\n<p>All images were resized to an aspect ratio of 1:1.5, which was the most popular aspect ratio of the images in the original dataset. We prepared 3 image datasets of sizes <code>300*450</code>, <code>400*600</code>, <code>600*900</code>. Most of the models were trained using the <code>400*600</code> dataset, as 300*450 gave inferior results and the <code>600*900</code> didn’t improve the results enough.</p>\n<p>For the metadata we had to set the same terminology for the 2020 and 2019 datasets.</p>\n<h3>Loss Function</h3>\n<p>The loss function we used was cross entropy. Although the task is to predict only melanoma we found it is better to predict the diagnosis which was split to 8 different classes one of which was melanoma. The final prediction was the value for the melanoma class, after a softmax function on all classes. We also tried a binary cross entropy on the melanoma class alone and a combination between the two, but using cross entropy gave the best results.</p>\n<p>The same loss was used for the base model and the transformer, but in the transformer we needed to regularize for the different number of predictions in each batch resulting from the different number of images for each patient. </p>\n<p>We also tried using focal loss which didn’t improve the results, but we left one of the transformer models which was trained with focal loss in the ensemble (A model with cross entropy loss gave the similar CV and LB).</p>\n<h4>Training the transformer model</h4>\n<p>The transformer model was trained in two steps. For the first step we used the data from both competitions (2019, 2020). For the 2019 competition we don’t have information about the patient, and each image got a different dummy patient_id, meaning the transformer didn’t learn much from these images. In the 2nd stage we fine-tuned the transformer using only the 2020 competition’s data.</p>\n<p>In both steps we used a sampler that over sampled the larger groups.</p>\n<h4>Inference</h4>\n<p>As stated above, the inference was done using TTA. For the base model we used 12-16 different augmentations and for the transformer model 32.</p>\n<h3>Ensembling</h3>\n<p>For our final submissions we used 2 ensembles:</p>\n<h5>Without Context Submission:</h5>\n<ol>\n<li>EfficientNet B3 noisy student image size 400*600</li>\n<li>EfficientNet B4 noisy student image size 400*600</li>\n<li>EfficientNet B5 noisy student image size 400*600</li>\n<li>EfficientNet B6 noisy student image size <strong>600*900</strong></li>\n<li>EfficientNet B7 noisy student image size 400*600</li>\n</ol>\n<h5>With Context</h5>\n<p>All the “without context” model +</p>\n<ul>\n<li>Transformer on features from A.</li>\n<li>Transformer on features from B.</li>\n<li>Transformer on features from C using focal loss</li>\n<li>Transformer on features from D.</li>\n<li>Transformer on features from E.</li>\n</ul>\n<h2>4. Interesting findings</h2>\n<h3>CV and LB</h3>\n<p>Although the number of images in the competition was very large, the number of different patients wasn’t large enough to give a stable and reliable CV and LB. And the correlation between the two was low. At the end we trusted neither and we submitted the models we felt were the most robust.</p>\n<h3>What didn’t work</h3>\n<ol>\n<li>Using different sizes of images as suggested in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">CNN Input Size Explained</a> didn’t show any improvement.</li>\n<li>Using Mixup[5] and MixCut[4] augmentation didn’t work</li>\n<li>Larger transformers didn’t improve the results. </li>\n</ol>\n<h2>5. Better “Real World” model</h2>\n<p>This model can be much simpler if we won’t ensemble and take only one base model, the EfficientNet B5 is probably the best compromise.</p>\n<p>In the real world scenario the images from previous years will probably be tagged already. In that case we can use a full transformer with encoder and decoder which will perform a seq2seq operation. </p>",
      "rawMarkdown": "This is a summery of \"Yuval and nosound\" model (27th place)\n\nYou can find more in:\n* [our github repository](https://github.com/yuval6957/SIIM-Transformer) \n* [Paper](https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20paper.pdf)\n* [Presentation](https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20presentation.pdf)\n* [video](https://www.youtube.com/watch?v=qa6zimKQcno&t=15s)\n\n## 1. Summary\n\nOur solution is based on two step model + Ensemble:\n\n\n\n1. Base model for feature extraction per image\n2. Transformer model - combining all the output features from a patient and predict per image. \n3. The 2nd stage also included some post - processing and ensembling.\n\n\n### Base Model:\n\nAs base model we used a models from the [EfficientNet](https://arxiv.org/abs/1905.11946) family[6]:\n\n\n\n*   EfficientNet b3 \n*   EfficientNet b4 \n*   EfficientNet b5 \n*   EfficientNet b6 \n*   EfficientNet b7 \n\nAll models were pre-trained on imagenet using noisy student algorithm. The models and weights are from  [gen-efficientnet-pytorch](https://github.com/rwightman/gen-efficientnet-pytorch)[2].\n\nThe input to these model is the image and meta-data such as age, sex, and anatomic Site. The meta-data is  processed by a small fully connected network and it’s output is concatenated to the input of the classification layer of the original EfficientNet network. This vector is going through a linear layer with output size of 256 to create the “features”, and then after an activation layer to the final linear classification layer. \n\nThis network has 8 outputs and tries to classify the diagnosis label (there are actually more than 8 possible diagnoses, but some don’t have enough examples). \n\n\n### Transformer Models:\n\nThe input to the Transformer models are a stack of features from all images belonging to the same patient + the metadata for these images.\n\nThe transformer is a stack of 4 transformer encoder layers with self attention as described in [Attention Is All You Need](https://arxiv.org/abs/1706.03762) [1]. Each transformer encoder layer uses 4 self attention heads. \n\nThe output of the transformer is a N*C, where N is the number of input feature vectors (the number of images) and C is the number of classes (8 in this case). Hence, the transformer predicts the class of each feature vector simultaneously, using the information from all other feature vectors.\n\nThe metadata is added using a “transformer style”, i.e. each parameter is transformed to a vector (size 256) using an embedding matrix and then added to the feature vector. for continuous values (like age) the embedding matrix was replaced by a 2 layer fully connected network.   \n\n\n### Ensembling the output of all networks:\n\nThe data was split to 3 folds, 3 times (using 3 different seeds for splitting), and the inference was done using 16 (or 12) TTAs. giving 144 predictions from each model. These were averaged and then the outputs of all the models were averaged. All averaging was done on the outputs before softmax and there form it is actually geometric averaging.\n\n\n### Training\n\nThe heavy lifting was the training and inference of the base models. This was done on a server with 2 GPUs – Tesla V100, Titan RTX that worked in parallel on different tasks. Training one fold of one model took 3H (B3)  to 11H (B7, B6 large images) on the Tesla and 20% more on the Titan, this sums up to about one day for B3 and 3.5 days for B7. Inferencing all the training data  for 12 TTA’s + test data for 16 TTA’s to get the features for the next level took  another 4h - 14h. The transformer training took less than 1H for the full model (3 folds*3seed). \n\nThe total time it took to train all models and folds is about 2.5W for one Tesla (~1.5W using the 2 GPUs).\n\n\n## 2. Models and features\n\n\n### Base models\n\nAs base models we tried various types of models (pre-trained on Imagenet):\n\n\n\n*   Densenets – 121, 161, 169, 201\n*   EfficientNet B0, B3 , B4, B5, B6, B7 with and with noisy student pre-training and with normal pretraining \n*   ResNet 101\n*   Xception \n\nAt the end we used EfficientNet as it was best when judging accuracy/time\n\nThe noisy student version performed better than the normal one.\n\nWe also tried different image sizes and ended up using a `400*600` images in most cases, except one were we used `600*900` with the B6 network.\n\n\n\n#### Metadata\n\nAs was described above the metadata was processed by a small fully connected nn and its output was concatenated to the output of the EfficientNet network (after removing the original top layer).\n\nWe also tried a network without metadata, and used the metadata as targets, i.e. this network predicted the diagnosis, but also the sex, age and anatomic site. The final predictions (including transformer) when using this approach weren't as good as the metadata as input approach. \n\n\n#### Model’s output\n\nAlthough the task at hand is to predict melanoma yes/no, it is better to let the network choose the diagnosis among a few possible options. This lets the network “understand” more about the image. The final prediction is the value of the Melanoma output after doing softmax on the output vector. \n\n\n#### Features\n\nThe final layer in this model is a linear layer with 256 inputs and 8 outputs, we use the input to this layer as features. \n\n\n#### Augmentation\n\nThe following augmentations where used while training and inference:\n\nRandom resize + crop\n\nRandom rotation\n\nRandom flip\n\nRandom color jitter (brightness, contrast, saturation, hue)\n\n[Cutout](https://arxiv.org/abs/1708.04552)[3] - erasing a small rectangle in the image\n\nHair - Randomly adding “hair like” lines to the image\n\nMetadata augmentation - adding random noise to the metadata as was done in the [1st place solution in ISIC 2019 challenge](https://www.sciencedirect.com/science/article/pii/S2215016120300832?via%3Dihub) [7].\n\n\n#### TTA\n\nFor inference each image was augmented differently 16 times and the final prediction was the average. These augmentations were also used for extracting 16 different features vectors per test image.\n\nThe same was done to extract 12 features vectors for the train images (12 and not 16 because of time limits).\n\n\n### Transformer Network\n\nThe input to the transformer network is the features from all the images from one patient.\n\nThe inspiration for this kind of model came from a previous competition in which we participated in RSNA[ Intracranial Hemorrhage Detection](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection). In that competition, all the top solutions (including ours) used a two stage network approach (although none of them were transformers).\n\nUsing a transformer seems appropriate in this case because transformers are built to seek relationships between embedding (feature) vectors in it’s input.\n\nAs this is not a full seq2seq task, we only used encoder layers. The transformer is a stack of 4 encoder layers with 4 attention heads in each layer. (we also tested higher numbers of layers and attention heads - no performance improvement).\n\n\n#### metadata\n\nThe metadata was incorporated in the network by adding “metadata vectors” to the input vectors - each value was transformed to a vector size 256 and added. The discrete values’  transformation was done using a trainable embedding matrix and the continuous values using a small nn.\n\n\n#### output\n\nThe output of this network is a matrix of size N*C where N - the number of images, C - the number of classes. Which means it decides on all the images of the patient at once.     \n\n\n#### Limit the input size\n\nA transformer can be trained on different number of feature vectors, by using padding. But when the range of numbers is very large, from a couple of hundred images for some patients to a handful for others, this may cause some implementation issues (like in calculating the loss). To simplify these issues, we limited N to 24 feature vectors, and for each patient we randomly divided the images to groups of size up to 24. \n\nThis might degrade the prediction as the most “similar” images might accidentally fall into different groups, but as we use TTA, this issue is almost solved. \n\n\n#### Augmentation\n\nFrom the base model we extract a number of feature vectors (12 for train and 16 for test) using different augmentation for the images and metadata. In the training and inference steps of the transformer model we randomly choose one of these vectors.\n\nAnother augmentation is the random grouping as stated above.\n\n\n## 3. Training and Inferencing\n\nThe original 2020 competition data is highly unbalanced, there are only 2-3% of positive targets in the train and test data. Although we were able to train the base model using uneven sampling, the best way to get good training was to add the data from ISIC 2019 competition which has a much higher percentage of melanoma images. \n\nWe split the training data to 3 folds keeping all the images from the same patient in the same fold, and making sure each fold has a similar number of patients with melanoma. The ISIC2019’s data was also split evenly between the folds. The same folds were kept for the base and the transformer models. \n\nTo get more diversity we had 3 different splits using 3 seeds \n\n\n### Preprocessing\n\nAll images were resized to an aspect ratio of 1:1.5, which was the most popular aspect ratio of the images in the original dataset. We prepared 3 image datasets of sizes `300*450`, `400*600`, `600*900`. Most of the models were trained using the `400*600` dataset, as 300*450 gave inferior results and the `600*900` didn’t improve the results enough.\n\nFor the metadata we had to set the same terminology for the 2020 and 2019 datasets.\n\n\n### Loss Function\n\nThe loss function we used was cross entropy. Although the task is to predict only melanoma we found it is better to predict the diagnosis which was split to 8 different classes one of which was melanoma. The final prediction was the value for the melanoma class, after a softmax function on all classes. We also tried a binary cross entropy on the melanoma class alone and a combination between the two, but using cross entropy gave the best results.\n\nThe same loss was used for the base model and the transformer, but in the transformer we needed to regularize for the different number of predictions in each batch resulting from the different number of images for each patient. \n\nWe also tried using focal loss which didn’t improve the results, but we left one of the transformer models which was trained with focal loss in the ensemble (A model with cross entropy loss gave the similar CV and LB).\n\n\n#### Training the transformer model \n\nThe transformer model was trained in two steps. For the first step we used the data from both competitions (2019, 2020). For the 2019 competition we don’t have information about the patient, and each image got a different dummy patient_id, meaning the transformer didn’t learn much from these images. In the 2nd stage we fine-tuned the transformer using only the 2020 competition’s data.\n\nIn both steps we used a sampler that over sampled the larger groups.\n\n\n#### Inference \n\nAs stated above, the inference was done using TTA. For the base model we used 12-16 different augmentations and for the transformer model 32.\n\n\n### Ensembling\n\nFor our final submissions we used 2 ensembles:\n\n\n##### Without Context Submission:\n\n\n\n1. EfficientNet B3 noisy student image size 400*600\n2. EfficientNet B4 noisy student image size 400*600\n3. EfficientNet B5 noisy student image size 400*600\n4. EfficientNet B6 noisy student image size **600*900**\n5. EfficientNet B7 noisy student image size 400*600\n\n\n##### With Context\n\nAll the “without context” model +\n\n\n\n*   Transformer on features from A.\n*   Transformer on features from B.\n*   Transformer on features from C using focal loss\n*   Transformer on features from D.\n*   Transformer on features from E.\n\n\n## 4. Interesting findings\n\n\n### CV and LB\n\nAlthough the number of images in the competition was very large, the number of different patients wasn’t large enough to give a stable and reliable CV and LB. And the correlation between the two was low. At the end we trusted neither and we submitted the models we felt were the most robust.\n\n\n### What didn’t work\n\n\n\n1. Using different sizes of images as suggested in [CNN Input Size Explained](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147) didn’t show any improvement.\n2. Using Mixup[5] and MixCut[4] augmentation didn’t work\n3. Larger transformers didn’t improve the results. \n\n\n## 5. Better “Real World” model\n\nThis model can be much simpler if we won’t ensemble and take only one base model, the EfficientNet B5 is probably the best compromise.\n\nIn the real world scenario the images from previous years will probably be tagged already. In that case we can use a full transformer with encoder and decoder which will perform a seq2seq operation.",
      "votes": null
    },
    {
      "id": "1005643",
      "postDate": "09/10/2020 16:12:45",
      "content": "<p>Grats and big thanks for publishing!!!<br>\nGives me, as a beginner, tons of stuff to learn from!</p>",
      "rawMarkdown": "Grats and big thanks for publishing!!!\nGives me, as a beginner, tons of stuff to learn from!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1005643,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "09/10/2020 16:12:45",
      "content": "<p>Grats and big thanks for publishing!!!<br>\nGives me, as a beginner, tons of stuff to learn from!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1005162": "This is a summery of \"Yuval and nosound\" model (27th place)\n\nYou can find more in:\n* [our github repository](https://github.com/yuval6957/SIIM-Transformer) \n* [Paper](https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20paper.pdf)\n* [Presentation](https://github.com/yuval6957/SIIM-Transformer/blob/master/SIIM%20presentation.pdf)\n* [video](https://www.youtube.com/watch?v=qa6zimKQcno&t=15s)\n\n## 1. Summary\n\nOur solution is based on two step model + Ensemble:\n\n\n\n1. Base model for feature extraction per image\n2. Transformer model - combining all the output features from a patient and predict per image. \n3. The 2nd stage also included some post - processing and ensembling.\n\n\n### Base Model:\n\nAs base model we used a models from the [EfficientNet](https://arxiv.org/abs/1905.11946) family[6]:\n\n\n\n*   EfficientNet b3 \n*   EfficientNet b4 \n*   EfficientNet b5 \n*   EfficientNet b6 \n*   EfficientNet b7 \n\nAll models were pre-trained on imagenet using noisy student algorithm. The models and weights are from  [gen-efficientnet-pytorch](https://github.com/rwightman/gen-efficientnet-pytorch)[2].\n\nThe input to these model is the image and meta-data such as age, sex, and anatomic Site. The meta-data is  processed by a small fully connected network and it’s output is concatenated to the input of the classification layer of the original EfficientNet network. This vector is going through a linear layer with output size of 256 to create the “features”, and then after an activation layer to the final linear classification layer. \n\nThis network has 8 outputs and tries to classify the diagnosis label (there are actually more than 8 possible diagnoses, but some don’t have enough examples). \n\n\n### Transformer Models:\n\nThe input to the Transformer models are a stack of features from all images belonging to the same patient + the metadata for these images.\n\nThe transformer is a stack of 4 transformer encoder layers with self attention as described in [Attention Is All You Need](https://arxiv.org/abs/1706.03762) [1]. Each transformer encoder layer uses 4 self attention heads. \n\nThe output of the transformer is a N*C, where N is the number of input feature vectors (the number of images) and C is the number of classes (8 in this case). Hence, the transformer predicts the class of each feature vector simultaneously, using the information from all other feature vectors.\n\nThe metadata is added using a “transformer style”, i.e. each parameter is transformed to a vector (size 256) using an embedding matrix and then added to the feature vector. for continuous values (like age) the embedding matrix was replaced by a 2 layer fully connected network.   \n\n\n### Ensembling the output of all networks:\n\nThe data was split to 3 folds, 3 times (using 3 different seeds for splitting), and the inference was done using 16 (or 12) TTAs. giving 144 predictions from each model. These were averaged and then the outputs of all the models were averaged. All averaging was done on the outputs before softmax and there form it is actually geometric averaging.\n\n\n### Training\n\nThe heavy lifting was the training and inference of the base models. This was done on a server with 2 GPUs – Tesla V100, Titan RTX that worked in parallel on different tasks. Training one fold of one model took 3H (B3)  to 11H (B7, B6 large images) on the Tesla and 20% more on the Titan, this sums up to about one day for B3 and 3.5 days for B7. Inferencing all the training data  for 12 TTA’s + test data for 16 TTA’s to get the features for the next level took  another 4h - 14h. The transformer training took less than 1H for the full model (3 folds*3seed). \n\nThe total time it took to train all models and folds is about 2.5W for one Tesla (~1.5W using the 2 GPUs).\n\n\n## 2. Models and features\n\n\n### Base models\n\nAs base models we tried various types of models (pre-trained on Imagenet):\n\n\n\n*   Densenets – 121, 161, 169, 201\n*   EfficientNet B0, B3 , B4, B5, B6, B7 with and with noisy student pre-training and with normal pretraining \n*   ResNet 101\n*   Xception \n\nAt the end we used EfficientNet as it was best when judging accuracy/time\n\nThe noisy student version performed better than the normal one.\n\nWe also tried different image sizes and ended up using a `400*600` images in most cases, except one were we used `600*900` with the B6 network.\n\n\n\n#### Metadata\n\nAs was described above the metadata was processed by a small fully connected nn and its output was concatenated to the output of the EfficientNet network (after removing the original top layer).\n\nWe also tried a network without metadata, and used the metadata as targets, i.e. this network predicted the diagnosis, but also the sex, age and anatomic site. The final predictions (including transformer) when using this approach weren't as good as the metadata as input approach. \n\n\n#### Model’s output\n\nAlthough the task at hand is to predict melanoma yes/no, it is better to let the network choose the diagnosis among a few possible options. This lets the network “understand” more about the image. The final prediction is the value of the Melanoma output after doing softmax on the output vector. \n\n\n#### Features\n\nThe final layer in this model is a linear layer with 256 inputs and 8 outputs, we use the input to this layer as features. \n\n\n#### Augmentation\n\nThe following augmentations where used while training and inference:\n\nRandom resize + crop\n\nRandom rotation\n\nRandom flip\n\nRandom color jitter (brightness, contrast, saturation, hue)\n\n[Cutout](https://arxiv.org/abs/1708.04552)[3] - erasing a small rectangle in the image\n\nHair - Randomly adding “hair like” lines to the image\n\nMetadata augmentation - adding random noise to the metadata as was done in the [1st place solution in ISIC 2019 challenge](https://www.sciencedirect.com/science/article/pii/S2215016120300832?via%3Dihub) [7].\n\n\n#### TTA\n\nFor inference each image was augmented differently 16 times and the final prediction was the average. These augmentations were also used for extracting 16 different features vectors per test image.\n\nThe same was done to extract 12 features vectors for the train images (12 and not 16 because of time limits).\n\n\n### Transformer Network\n\nThe input to the transformer network is the features from all the images from one patient.\n\nThe inspiration for this kind of model came from a previous competition in which we participated in RSNA[ Intracranial Hemorrhage Detection](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection). In that competition, all the top solutions (including ours) used a two stage network approach (although none of them were transformers).\n\nUsing a transformer seems appropriate in this case because transformers are built to seek relationships between embedding (feature) vectors in it’s input.\n\nAs this is not a full seq2seq task, we only used encoder layers. The transformer is a stack of 4 encoder layers with 4 attention heads in each layer. (we also tested higher numbers of layers and attention heads - no performance improvement).\n\n\n#### metadata\n\nThe metadata was incorporated in the network by adding “metadata vectors” to the input vectors - each value was transformed to a vector size 256 and added. The discrete values’  transformation was done using a trainable embedding matrix and the continuous values using a small nn.\n\n\n#### output\n\nThe output of this network is a matrix of size N*C where N - the number of images, C - the number of classes. Which means it decides on all the images of the patient at once.     \n\n\n#### Limit the input size\n\nA transformer can be trained on different number of feature vectors, by using padding. But when the range of numbers is very large, from a couple of hundred images for some patients to a handful for others, this may cause some implementation issues (like in calculating the loss). To simplify these issues, we limited N to 24 feature vectors, and for each patient we randomly divided the images to groups of size up to 24. \n\nThis might degrade the prediction as the most “similar” images might accidentally fall into different groups, but as we use TTA, this issue is almost solved. \n\n\n#### Augmentation\n\nFrom the base model we extract a number of feature vectors (12 for train and 16 for test) using different augmentation for the images and metadata. In the training and inference steps of the transformer model we randomly choose one of these vectors.\n\nAnother augmentation is the random grouping as stated above.\n\n\n## 3. Training and Inferencing\n\nThe original 2020 competition data is highly unbalanced, there are only 2-3% of positive targets in the train and test data. Although we were able to train the base model using uneven sampling, the best way to get good training was to add the data from ISIC 2019 competition which has a much higher percentage of melanoma images. \n\nWe split the training data to 3 folds keeping all the images from the same patient in the same fold, and making sure each fold has a similar number of patients with melanoma. The ISIC2019’s data was also split evenly between the folds. The same folds were kept for the base and the transformer models. \n\nTo get more diversity we had 3 different splits using 3 seeds \n\n\n### Preprocessing\n\nAll images were resized to an aspect ratio of 1:1.5, which was the most popular aspect ratio of the images in the original dataset. We prepared 3 image datasets of sizes `300*450`, `400*600`, `600*900`. Most of the models were trained using the `400*600` dataset, as 300*450 gave inferior results and the `600*900` didn’t improve the results enough.\n\nFor the metadata we had to set the same terminology for the 2020 and 2019 datasets.\n\n\n### Loss Function\n\nThe loss function we used was cross entropy. Although the task is to predict only melanoma we found it is better to predict the diagnosis which was split to 8 different classes one of which was melanoma. The final prediction was the value for the melanoma class, after a softmax function on all classes. We also tried a binary cross entropy on the melanoma class alone and a combination between the two, but using cross entropy gave the best results.\n\nThe same loss was used for the base model and the transformer, but in the transformer we needed to regularize for the different number of predictions in each batch resulting from the different number of images for each patient. \n\nWe also tried using focal loss which didn’t improve the results, but we left one of the transformer models which was trained with focal loss in the ensemble (A model with cross entropy loss gave the similar CV and LB).\n\n\n#### Training the transformer model \n\nThe transformer model was trained in two steps. For the first step we used the data from both competitions (2019, 2020). For the 2019 competition we don’t have information about the patient, and each image got a different dummy patient_id, meaning the transformer didn’t learn much from these images. In the 2nd stage we fine-tuned the transformer using only the 2020 competition’s data.\n\nIn both steps we used a sampler that over sampled the larger groups.\n\n\n#### Inference \n\nAs stated above, the inference was done using TTA. For the base model we used 12-16 different augmentations and for the transformer model 32.\n\n\n### Ensembling\n\nFor our final submissions we used 2 ensembles:\n\n\n##### Without Context Submission:\n\n\n\n1. EfficientNet B3 noisy student image size 400*600\n2. EfficientNet B4 noisy student image size 400*600\n3. EfficientNet B5 noisy student image size 400*600\n4. EfficientNet B6 noisy student image size **600*900**\n5. EfficientNet B7 noisy student image size 400*600\n\n\n##### With Context\n\nAll the “without context” model +\n\n\n\n*   Transformer on features from A.\n*   Transformer on features from B.\n*   Transformer on features from C using focal loss\n*   Transformer on features from D.\n*   Transformer on features from E.\n\n\n## 4. Interesting findings\n\n\n### CV and LB\n\nAlthough the number of images in the competition was very large, the number of different patients wasn’t large enough to give a stable and reliable CV and LB. And the correlation between the two was low. At the end we trusted neither and we submitted the models we felt were the most robust.\n\n\n### What didn’t work\n\n\n\n1. Using different sizes of images as suggested in [CNN Input Size Explained](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147) didn’t show any improvement.\n2. Using Mixup[5] and MixCut[4] augmentation didn’t work\n3. Larger transformers didn’t improve the results. \n\n\n## 5. Better “Real World” model\n\nThis model can be much simpler if we won’t ensemble and take only one base model, the EfficientNet B5 is probably the best compromise.\n\nIn the real world scenario the images from previous years will probably be tagged already. In that case we can use a full transformer with encoder and decoder which will perform a seq2seq operation.",
    "1005643": "Grats and big thanks for publishing!!!\nGives me, as a beginner, tons of stuff to learn from!"
  },
  "source": "meta"
}