{
  "id": 110346,
  "title": "Low Budget LB 0.9759 kaggle-kernels-only solution",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110346",
  "author_name": "Henrique Mendonça",
  "post_date": "2019-09-27T02:08:47.465000",
  "votes": 58,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I will try to summarise here my part of the 2 months battle with Kaggle kernels in this comp.\nOur team's solution was basically entirely run on kaggle GPU's (mostly prior to the new quotas).\nThe final kernel is here: <a href=\"https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759\">https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759</a></p>\n\n<p><strong>Quick outline in chronological order:</strong>\n 1. CV split by experiments about ⅓ of the data, 3 folds.\n 2. Trained a quick and dirty classifier (left it training while analysing the data) with all 1108 classes, 6 channels and no augmentation. OK: <strong>LB 0.10-0.25</strong> and established a reasonable learning rate schedule which was used in the rest of the comp.\n 3. Normalised the data by experiment mean and std, added random crops: <strong>LB 0.32</strong>\nAdded more aug: no improvement.\n 4. Added ArcNet head with 512 features and default params, using pre-trained model from above, eventually <strong>LB 0.35</strong>\n 5. Started training with each cell type as a different class 1108 * 4 = 4432 classes: <strong>LB 0.44</strong>\n 6. Cell type balancing in training data: <strong>LB 0.45</strong>\n 7. Added the 31 controls to the training data (but not to validation) 1139 * 4 = 4556 classes: LB 0.46\n 8. Added a custom cosine distance loss to the ArcNet features (same class should be close to each other, different classes far...) <strong>LB 0.47</strong>\n 9. PCA showed that 100 dimensions explained all the variance in features, so I reduced the ArcNet features to 200 (double, just in case ;) <strong>LB 0.52</strong>\n 10. Applied plates leak, loads of normalisation and assigned single treatment per plate LB 0.82 !!\nMore normalisation, and lots of multi-fold ensembles later: <strong>LB 0.88</strong> (Thanks to <a href=\"/aharless\">@aharless</a> and <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a>)\n 11. Used the former predictions as pseudo labels and trained further with train + test data, eventually <strong>LB 0.93</strong> (this process was repeated a few times, always training the latest models with pseudo labels from the best possible submission/ensemble in an amazing team effort! Thanks <a href=\"/aharless\">@aharless</a> <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a> <a href=\"/stillsut\">@stillsut</a> <a href=\"/lastlegion\">@lastlegion</a>)\n 12. Failed badly trying to blend the new best models (single models had better LB than the blends) probably due to models overfitting to the pseudo labels?\nReduce the amount of pseudo labels in the training data (from 50/50% to 33/66% now) and improved feature distance metric: <strong>LB 0.943</strong>\n 13. Out of gas, $$ and quota, but I have never learned so much in a comp before! Super happy with the result</p>\n\n<p>The code for all that is in the <a href=\"https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759\">kernel linked above</a> (including some visualisations).</p>\n\n<p><strong>In the beginning</strong>\nI basically started looking at the data and decided to just try a normal classifier without any fanciness. I ported an EfficientNetB4 kernel from another comp and watched the batch loss for a few different learning rates.\nThe loss during cosine annealing with warm restarts is very informative, almost as much as the LR finder scheduler used in fast.ai, so after a couple of runs you can work out what is too high and what is too low, and also how much you need to decrease it from epoch to epoch.\nI was training 3 folds and later 3 architectures (EfficientNet B4, B5 and seResNet-xt101). However, the journey described above is actually mainly from the initial B4 I used, and that also ended up being the best model and submission. Its weights were carved slowly by training several different heads one after another, with several branches used to experiment with different ideas.\nUnfortunately, we didn't have the budget to do clean experiments and try each idea with their optimal hyper-params and initial conditions. Therefore, some changes used here might not have been optimal but still ended up improving the score, so I had to keep them in the pipeline for the lack of a better reference.</p>\n\n<p><strong>Normalisation and Post-processing:</strong>\nThe experiment setup let out a lot of information, and thanks to the public kernels that was made clear for everyone. That played a huge role for us to improve the submission score and create the pseudo labels. Cheers <a href=\"/zaharch\">@zaharch</a> and <a href=\"/christopherberner\">@christopherberner</a> !\nLooking at the raw model predictions we could see that the initial models were obviously preferring some treatments to others. The predictions weren't balanced across the different classes, but we knew that they should, so a basic idea was to normalise across all treatments. Each treatment should appear exactly the same number of times in each experiment. Normalising the prediction probabilities across the siRNA already helped increasing the validation and LB score.\nAndy and Giulia went further and created normalisation loop iteratively normalising the probabilities vertically and horizontally until convergence.\nThis process showed great improvements in both validation and LB scores and was crucial to our solution. </p>\n\n<p><strong>Platform</strong>\nAs I mentioned before, the solution was run on GPU only.\nTPU's would have been amazing but I had a look initially and wasn't super excited about its pytorch support. I have also trained a ResNet101 model using Colab and the provided notebook in the comp's intro, but I found it a bit too annoying to work with all the deprecated TF code (hopefully it will get better soon as they're fixing the new TF2.0), so I gave up after a few days and went back to PyTorch. Maybe that was a mistake?\nWe had a few GCP credits but $400 goes really quick, and we only made a few experiments in the beginning. I also used some more (including a bit of my own pocket) to train another few models improving the pseudo labels a bit. However, our best model was trained exclusively in kaggle in about 13 runs in total (100hrs approximately, excluding all abandoned branches)</p>\n\n<p>I tried Colab again after the kaggle quotas were introduced, but I found their GPU notebook quite flaky, as it keeps disconnecting my runs during training :(\nMaybe I should buy a GPU lol</p>\n\n<p>Many thanks to all the organisers and the amazing community of kagglers always sharing some much!</p>",
  "messages": [
    {
      "id": 634962,
      "postDate": "2019-09-27T02:08:47.467Z",
      "content": "<p>I will try to summarise here my part of the 2 months battle with Kaggle kernels in this comp.\nOur team's solution was basically entirely run on kaggle GPU's (mostly prior to the new quotas).\nThe final kernel is here: <a href=\"https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759\">https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759</a></p>\n\n<p><strong>Quick outline in chronological order:</strong>\n 1. CV split by experiments about ⅓ of the data, 3 folds.\n 2. Trained a quick and dirty classifier (left it training while analysing the data) with all 1108 classes, 6 channels and no augmentation. OK: <strong>LB 0.10-0.25</strong> and established a reasonable learning rate schedule which was used in the rest of the comp.\n 3. Normalised the data by experiment mean and std, added random crops: <strong>LB 0.32</strong>\nAdded more aug: no improvement.\n 4. Added ArcNet head with 512 features and default params, using pre-trained model from above, eventually <strong>LB 0.35</strong>\n 5. Started training with each cell type as a different class 1108 * 4 = 4432 classes: <strong>LB 0.44</strong>\n 6. Cell type balancing in training data: <strong>LB 0.45</strong>\n 7. Added the 31 controls to the training data (but not to validation) 1139 * 4 = 4556 classes: LB 0.46\n 8. Added a custom cosine distance loss to the ArcNet features (same class should be close to each other, different classes far...) <strong>LB 0.47</strong>\n 9. PCA showed that 100 dimensions explained all the variance in features, so I reduced the ArcNet features to 200 (double, just in case ;) <strong>LB 0.52</strong>\n 10. Applied plates leak, loads of normalisation and assigned single treatment per plate LB 0.82 !!\nMore normalisation, and lots of multi-fold ensembles later: <strong>LB 0.88</strong> (Thanks to <a href=\"/aharless\">@aharless</a> and <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a>)\n 11. Used the former predictions as pseudo labels and trained further with train + test data, eventually <strong>LB 0.93</strong> (this process was repeated a few times, always training the latest models with pseudo labels from the best possible submission/ensemble in an amazing team effort! Thanks <a href=\"/aharless\">@aharless</a> <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a> <a href=\"/stillsut\">@stillsut</a> <a href=\"/lastlegion\">@lastlegion</a>)\n 12. Failed badly trying to blend the new best models (single models had better LB than the blends) probably due to models overfitting to the pseudo labels?\nReduce the amount of pseudo labels in the training data (from 50/50% to 33/66% now) and improved feature distance metric: <strong>LB 0.943</strong>\n 13. Out of gas, $$ and quota, but I have never learned so much in a comp before! Super happy with the result</p>\n\n<p>The code for all that is in the <a href=\"https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759\">kernel linked above</a> (including some visualisations).</p>\n\n<p><strong>In the beginning</strong>\nI basically started looking at the data and decided to just try a normal classifier without any fanciness. I ported an EfficientNetB4 kernel from another comp and watched the batch loss for a few different learning rates.\nThe loss during cosine annealing with warm restarts is very informative, almost as much as the LR finder scheduler used in fast.ai, so after a couple of runs you can work out what is too high and what is too low, and also how much you need to decrease it from epoch to epoch.\nI was training 3 folds and later 3 architectures (EfficientNet B4, B5 and seResNet-xt101). However, the journey described above is actually mainly from the initial B4 I used, and that also ended up being the best model and submission. Its weights were carved slowly by training several different heads one after another, with several branches used to experiment with different ideas.\nUnfortunately, we didn't have the budget to do clean experiments and try each idea with their optimal hyper-params and initial conditions. Therefore, some changes used here might not have been optimal but still ended up improving the score, so I had to keep them in the pipeline for the lack of a better reference.</p>\n\n<p><strong>Normalisation and Post-processing:</strong>\nThe experiment setup let out a lot of information, and thanks to the public kernels that was made clear for everyone. That played a huge role for us to improve the submission score and create the pseudo labels. Cheers <a href=\"/zaharch\">@zaharch</a> and <a href=\"/christopherberner\">@christopherberner</a> !\nLooking at the raw model predictions we could see that the initial models were obviously preferring some treatments to others. The predictions weren't balanced across the different classes, but we knew that they should, so a basic idea was to normalise across all treatments. Each treatment should appear exactly the same number of times in each experiment. Normalising the prediction probabilities across the siRNA already helped increasing the validation and LB score.\nAndy and Giulia went further and created normalisation loop iteratively normalising the probabilities vertically and horizontally until convergence.\nThis process showed great improvements in both validation and LB scores and was crucial to our solution. </p>\n\n<p><strong>Platform</strong>\nAs I mentioned before, the solution was run on GPU only.\nTPU's would have been amazing but I had a look initially and wasn't super excited about its pytorch support. I have also trained a ResNet101 model using Colab and the provided notebook in the comp's intro, but I found it a bit too annoying to work with all the deprecated TF code (hopefully it will get better soon as they're fixing the new TF2.0), so I gave up after a few days and went back to PyTorch. Maybe that was a mistake?\nWe had a few GCP credits but $400 goes really quick, and we only made a few experiments in the beginning. I also used some more (including a bit of my own pocket) to train another few models improving the pseudo labels a bit. However, our best model was trained exclusively in kaggle in about 13 runs in total (100hrs approximately, excluding all abandoned branches)</p>\n\n<p>I tried Colab again after the kaggle quotas were introduced, but I found their GPU notebook quite flaky, as it keeps disconnecting my runs during training :(\nMaybe I should buy a GPU lol</p>\n\n<p>Many thanks to all the organisers and the amazing community of kagglers always sharing some much!</p>",
      "rawMarkdown": "I will try to summarise here my part of the 2 months battle with Kaggle kernels in this comp.\nOur team's solution was basically entirely run on kaggle GPU's (mostly prior to the new quotas).\nThe final kernel is here: https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759\n\n**Quick outline in chronological order:**\n 1. CV split by experiments about ⅓ of the data, 3 folds.\n 2. Trained a quick and dirty classifier (left it training while analysing the data) with all 1108 classes, 6 channels and no augmentation. OK: **LB 0.10-0.25** and established a reasonable learning rate schedule which was used in the rest of the comp.\n 3. Normalised the data by experiment mean and std, added random crops: **LB 0.32**\nAdded more aug: no improvement.\n 4. Added ArcNet head with 512 features and default params, using pre-trained model from above, eventually **LB 0.35**\n 5. Started training with each cell type as a different class 1108 * 4 = 4432 classes: **LB 0.44**\n 6. Cell type balancing in training data: **LB 0.45**\n 7. Added the 31 controls to the training data (but not to validation) 1139 * 4 = 4556 classes: LB 0.46\n 8. Added a custom cosine distance loss to the ArcNet features (same class should be close to each other, different classes far...) **LB 0.47**\n 9. PCA showed that 100 dimensions explained all the variance in features, so I reduced the ArcNet features to 200 (double, just in case ;) **LB 0.52**\n 10. Applied plates leak, loads of normalisation and assigned single treatment per plate LB 0.82 !!\nMore normalisation, and lots of multi-fold ensembles later: **LB 0.88** (Thanks to @aharless and @giuliasavorgnan)\n 11. Used the former predictions as pseudo labels and trained further with train + test data, eventually **LB 0.93** (this process was repeated a few times, always training the latest models with pseudo labels from the best possible submission/ensemble in an amazing team effort! Thanks @aharless @giuliasavorgnan @stillsut @lastlegion)\n 12. Failed badly trying to blend the new best models (single models had better LB than the blends) probably due to models overfitting to the pseudo labels?\nReduce the amount of pseudo labels in the training data (from 50/50% to 33/66% now) and improved feature distance metric: **LB 0.943**\n 13. Out of gas, $$ and quota, but I have never learned so much in a comp before! Super happy with the result\n\nThe code for all that is in the [kernel linked above](https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759) (including some visualisations).\n\n\n**In the beginning**\nI basically started looking at the data and decided to just try a normal classifier without any fanciness. I ported an EfficientNetB4 kernel from another comp and watched the batch loss for a few different learning rates.\nThe loss during cosine annealing with warm restarts is very informative, almost as much as the LR finder scheduler used in fast.ai, so after a couple of runs you can work out what is too high and what is too low, and also how much you need to decrease it from epoch to epoch.\nI was training 3 folds and later 3 architectures (EfficientNet B4, B5 and seResNet-xt101). However, the journey described above is actually mainly from the initial B4 I used, and that also ended up being the best model and submission. Its weights were carved slowly by training several different heads one after another, with several branches used to experiment with different ideas.\nUnfortunately, we didn't have the budget to do clean experiments and try each idea with their optimal hyper-params and initial conditions. Therefore, some changes used here might not have been optimal but still ended up improving the score, so I had to keep them in the pipeline for the lack of a better reference.\n\n**Normalisation and Post-processing:**\nThe experiment setup let out a lot of information, and thanks to the public kernels that was made clear for everyone. That played a huge role for us to improve the submission score and create the pseudo labels. Cheers @zaharch and @christopherberner !\nLooking at the raw model predictions we could see that the initial models were obviously preferring some treatments to others. The predictions weren't balanced across the different classes, but we knew that they should, so a basic idea was to normalise across all treatments. Each treatment should appear exactly the same number of times in each experiment. Normalising the prediction probabilities across the siRNA already helped increasing the validation and LB score.\nAndy and Giulia went further and created normalisation loop iteratively normalising the probabilities vertically and horizontally until convergence.\nThis process showed great improvements in both validation and LB scores and was crucial to our solution. \n\n**Platform**\nAs I mentioned before, the solution was run on GPU only.\nTPU's would have been amazing but I had a look initially and wasn't super excited about its pytorch support. I have also trained a ResNet101 model using Colab and the provided notebook in the comp's intro, but I found it a bit too annoying to work with all the deprecated TF code (hopefully it will get better soon as they're fixing the new TF2.0), so I gave up after a few days and went back to PyTorch. Maybe that was a mistake?\nWe had a few GCP credits but $400 goes really quick, and we only made a few experiments in the beginning. I also used some more (including a bit of my own pocket) to train another few models improving the pseudo labels a bit. However, our best model was trained exclusively in kaggle in about 13 runs in total (100hrs approximately, excluding all abandoned branches)\n\nI tried Colab again after the kaggle quotas were introduced, but I found their GPU notebook quite flaky, as it keeps disconnecting my runs during training :(\nMaybe I should buy a GPU lol\n\nMany thanks to all the organisers and the amazing community of kagglers always sharing some much!",
      "votes": 58
    },
    {
      "id": 635125,
      "postDate": "2019-09-27T07:21:36.597Z",
      "content": "<p>I'm going to add a note on what did not work for me during model experimentation. <a href=\"/zaharch\">@zaharch</a>, this also partly addresses your question. I experimented with a siamese network + cosine embedding loss. I normalized the images to the mean and stddev of their experiment. I split the train set into 4 cell-types and trained 4 separate models. I used all train treatment+control siRNAs for training and the test control for validation (I was planning to use the CV folds of <a href=\"/hmendonca\">@hmendonca</a> later, but the model never achieved any decent results). A resnet50 backbone was starting to overfit almost immediately, so I discarded it. Efficientnet and densenet121 with dropout (~0.3-0.5) helped overcame the overfitting problem. In addition I added a very severe scheduling to reduce the learning rate, dropping by ~10% roughly every 10-20 iterations (not epochs!) without loss improvement; this helped the model avoid getting a nan loss. I modified the first backbone layer to accept the 6-channel inputs, and added a 128 encoder at the head. I initialized the network with ImageNet pretrained weights except in the first layer and in the encoder, where I left the random initialization. I chose a 128 encoder after experimenting with different bigger sizes that turned out to overfit (this seems in agreement with <a href=\"/hmendonca\">@hmendonca</a>'s PCA test). I trained on 6x128x128 inputs first, achieved convergence, then continued on 6x256x256. Augmentation was done with 2/3 random cropping and random rotation. The batch was composed of half positive and half negative pairs of images. The model was learning to distinguish positive and negative matches with an accuracy of ~70%. Then I started selecting the hardest negative pairs \"online\" (batch by batch), and the siamese accuracy improved to ~80-90%, however the model was not able at all to correctly classify the 1139 siRNAs (naively: calculate embeddings for all train samples xN augmentations, calculate embeddings for test samples, find the closest match). I trained a lightGBM classifier on the embeddings, achieved max 20% accuracy on the test control samples, but basically 0 accuracy on the LB score. Turned out the classifier was overconfident on the control classes only, probably because I did not balance the dataset. I added a crossentropy loss to the siamese network (so loss = cosine_embedding_loss + crossentropy_loss) and I observed an improvement in the distribution of the cosine_embedding_loss, but again very low accuracy on the test controls. At this stage I should have started using proper CV folds, a more clever LR scheduler, maybe triplet loss + crossentropy, but... <strong>I ran out of my portion of GCP credits. I had tried to use the kaggle kernels at the beginning but I gave up pretty quickly, frustrated with the too frequent disconnection errors. I remember I found completely impossible to use the kernels over the weekend, I was waiting 30 minutes to start a new kernel, another 30 to load the data, another hour in case I had to restart the kernel, and when the kernel was restarting by itself I often just closed the laptop lid and went for a hike...</strong> Unfortunately I did not have access to a local GPU not even for testing and debugging the code. It was my first time experimenting with such a model, I was probably slow and needed more computing resources to go a bit further. If anyone has any insight as to why my methodology did not work, I would love to hear you!</p>",
      "rawMarkdown": "I'm going to add a note on what did not work for me during model experimentation. @zaharch, this also partly addresses your question. I experimented with a siamese network + cosine embedding loss. I normalized the images to the mean and stddev of their experiment. I split the train set into 4 cell-types and trained 4 separate models. I used all train treatment+control siRNAs for training and the test control for validation (I was planning to use the CV folds of @hmendonca later, but the model never achieved any decent results). A resnet50 backbone was starting to overfit almost immediately, so I discarded it. Efficientnet and densenet121 with dropout (~0.3-0.5) helped overcame the overfitting problem. In addition I added a very severe scheduling to reduce the learning rate, dropping by ~10% roughly every 10-20 iterations (not epochs!) without loss improvement; this helped the model avoid getting a nan loss. I modified the first backbone layer to accept the 6-channel inputs, and added a 128 encoder at the head. I initialized the network with ImageNet pretrained weights except in the first layer and in the encoder, where I left the random initialization. I chose a 128 encoder after experimenting with different bigger sizes that turned out to overfit (this seems in agreement with @hmendonca's PCA test). I trained on 6x128x128 inputs first, achieved convergence, then continued on 6x256x256. Augmentation was done with 2/3 random cropping and random rotation. The batch was composed of half positive and half negative pairs of images. The model was learning to distinguish positive and negative matches with an accuracy of ~70%. Then I started selecting the hardest negative pairs \"online\" (batch by batch), and the siamese accuracy improved to ~80-90%, however the model was not able at all to correctly classify the 1139 siRNAs (naively: calculate embeddings for all train samples xN augmentations, calculate embeddings for test samples, find the closest match). I trained a lightGBM classifier on the embeddings, achieved max 20% accuracy on the test control samples, but basically 0 accuracy on the LB score. Turned out the classifier was overconfident on the control classes only, probably because I did not balance the dataset. I added a crossentropy loss to the siamese network (so loss = cosine\\_embedding\\_loss + crossentropy\\_loss) and I observed an improvement in the distribution of the cosine\\_embedding\\_loss, but again very low accuracy on the test controls. At this stage I should have started using proper CV folds, a more clever LR scheduler, maybe triplet loss + crossentropy, but... **I ran out of my portion of GCP credits. I had tried to use the kaggle kernels at the beginning but I gave up pretty quickly, frustrated with the too frequent disconnection errors. I remember I found completely impossible to use the kernels over the weekend, I was waiting 30 minutes to start a new kernel, another 30 to load the data, another hour in case I had to restart the kernel, and when the kernel was restarting by itself I often just closed the laptop lid and went for a hike...** Unfortunately I did not have access to a local GPU not even for testing and debugging the code. It was my first time experimenting with such a model, I was probably slow and needed more computing resources to go a bit further. If anyone has any insight as to why my methodology did not work, I would love to hear you!",
      "votes": 3,
      "replies": [
        {
          "id": 635173,
          "postDate": "2019-09-27T08:23:50.967Z",
          "content": "<p>Great work! I was looking into siamese networks in the beginning too, but then realized that it was probably an outdated approach, with ArcFace and co outperforming it on few-short learning. Definitely you needed normal CV folds from the beginning, would have saved you a lot of time. Have you used GPUs with GCP on a virtual machine? And thank you for the contribution in this competition.</p>",
          "rawMarkdown": "Great work! I was looking into siamese networks in the beginning too, but then realized that it was probably an outdated approach, with ArcFace and co outperforming it on few-short learning. Definitely you needed normal CV folds from the beginning, would have saved you a lot of time. Have you used GPUs with GCP on a virtual machine? And thank you for the contribution in this competition.",
          "votes": 1
        },
        {
          "id": 635202,
          "postDate": "2019-09-27T08:50:28.403Z",
          "content": "<p>Yes, I spent my credits on a GPU VM. Good learning for the future. The only regret is not having pushed the model to the end because of kaggle kernels frustration :/ Now, with the GPU quota, I'm basically out of game. Thanks to you for your awesome sharing!</p>",
          "rawMarkdown": "Yes, I spent my credits on a GPU VM. Good learning for the future. The only regret is not having pushed the model to the end because of kaggle kernels frustration :/ Now, with the GPU quota, I'm basically out of game. Thanks to you for your awesome sharing!",
          "votes": 2
        },
        {
          "id": 635559,
          "postDate": "2019-09-27T20:22:19.420Z",
          "content": "<p>I am \"glad\" to see that others had such difficulty with a metric network, as it failed quite miserably for me.  I started a bit late on this competition and was quite focused on getting a metric learning method to work.  I would love to see if anyone had success in implementing such an approach.</p>\n\n<p>In short, I trained a VGG-19 with ImageNet weights, \"doubling\" the initial conv layer (depth-wise) to accept 6-channel images.  Following the base VGG, I added a FC (2048), dropout, FC(256);  this final 256 vector was my embedding.  I used N-pairs loss (had also tried triplet, but was swayed by the apparent improvements using n-pairs).  I started by training the FC layers, and then began unfreezing the conv layers.  </p>\n\n<p>I trained a different model for each cell-line, and my training sets were created by random sampling (likely not a very efficient way to learn).  My positives were the two sites from same well AND the same siRNA from different experiments, while negatives were random different siRNA.  By taking positive examples from different plates, I hoped that the model would learn the \"biology\" rather than any batch effects.</p>\n\n<p>Once I got the learning rate correct, I was getting very nice separation of the embeddings that I trained on.  A 2-d PCA showed very distinct clusters, which led me to believe I was training well.</p>\n\n<p>However, once applied to test set, it was quite poor.  After generating embeddings for the test set, I naively tried to assign based on KNN (3-6), nearest centroid etc....nothing got above 1% LB!!!  I then tried a number of transformations on the embeddings to see if it was just a simple \"shift\", due to the fact that the network had not seen these samples before.  Even comparing the embeddings of the control siRNA, none were \"displaced\" consistently between training and test sets (e.g. it was not something simple like a simple translation of the coordinate system).  My test embeddings showed some weak segregation by siRNA (via 2-d PCA), but they were still quite mixed. I ultimately became stuck at this last step and did not have time to try a classification based approach.  </p>\n\n<p>If anyone got a metric network to perform well in test, I'd love to hear how!  Thanks!</p>",
          "rawMarkdown": "I am \"glad\" to see that others had such difficulty with a metric network, as it failed quite miserably for me.  I started a bit late on this competition and was quite focused on getting a metric learning method to work.  I would love to see if anyone had success in implementing such an approach.\n\nIn short, I trained a VGG-19 with ImageNet weights, \"doubling\" the initial conv layer (depth-wise) to accept 6-channel images.  Following the base VGG, I added a FC (2048), dropout, FC(256);  this final 256 vector was my embedding.  I used N-pairs loss (had also tried triplet, but was swayed by the apparent improvements using n-pairs).  I started by training the FC layers, and then began unfreezing the conv layers.  \n\nI trained a different model for each cell-line, and my training sets were created by random sampling (likely not a very efficient way to learn).  My positives were the two sites from same well AND the same siRNA from different experiments, while negatives were random different siRNA.  By taking positive examples from different plates, I hoped that the model would learn the \"biology\" rather than any batch effects.\n\nOnce I got the learning rate correct, I was getting very nice separation of the embeddings that I trained on.  A 2-d PCA showed very distinct clusters, which led me to believe I was training well.\n\nHowever, once applied to test set, it was quite poor.  After generating embeddings for the test set, I naively tried to assign based on KNN (3-6), nearest centroid etc....nothing got above 1% LB!!!  I then tried a number of transformations on the embeddings to see if it was just a simple \"shift\", due to the fact that the network had not seen these samples before.  Even comparing the embeddings of the control siRNA, none were \"displaced\" consistently between training and test sets (e.g. it was not something simple like a simple translation of the coordinate system).  My test embeddings showed some weak segregation by siRNA (via 2-d PCA), but they were still quite mixed. I ultimately became stuck at this last step and did not have time to try a classification based approach.  \n\nIf anyone got a metric network to perform well in test, I'd love to hear how!  Thanks!"
        },
        {
          "id": 635819,
          "postDate": "2019-09-28T09:24:33.397Z",
          "content": "<p><a href=\"/bioin4\">@bioin4</a> Many top solutions used ArcNet, which is considered metric learning (including us, please see the topic above). </p>\n\n<p>I believe it doesn't work very well from scratch if you have a large ArcNet margin ('m' param), as it probably need to have good features to start with, and imageNet wasn't good enough here.  But it worked fine for me with margin 0 and imageNet weights.</p>\n\n<p>I also found that hard mining didn't help the validation score much, as it tends to quickly overfit to the training data. Hard mining is used in most metric learning pipelines if I'm not wrong.</p>",
          "rawMarkdown": "@bioin4 Many top solutions used ArcNet, which is considered metric learning (including us, please see the topic above). \n\nI believe it doesn't work very well from scratch if you have a large ArcNet margin ('m' param), as it probably need to have good features to start with, and imageNet wasn't good enough here.  But it worked fine for me with margin 0 and imageNet weights.\n\nI also found that hard mining didn't help the validation score much, as it tends to quickly overfit to the training data. Hard mining is used in most metric learning pipelines if I'm not wrong."
        }
      ]
    },
    {
      "id": 635319,
      "postDate": "2019-09-27T11:30:28.140Z",
      "content": "<p>I would want to emphasize the role of <a href=\"/christopherberner\">@christopherberner</a> 's \"Hungarian algorithm\" kernel in helping us to exploit the plates leak (and to exploit the overall structure, which would have been helpful even without the leak).  I had an intuitive sense that we needed to find an optimal way to assign siRNAs within a plate given a set of probabilities, but I didn't know the appropriate method until I found that kernel.  There's a lesson here about learning from public kernels: we were already far ahead of that kernel's LB score when I found it, but the method we learned from it produced a critical improvement in our score.</p>",
      "rawMarkdown": "I would want to emphasize the role of @christopherberner 's \"Hungarian algorithm\" kernel in helping us to exploit the plates leak (and to exploit the overall structure, which would have been helpful even without the leak).  I had an intuitive sense that we needed to find an optimal way to assign siRNAs within a plate given a set of probabilities, but I didn't know the appropriate method until I found that kernel.  There's a lesson here about learning from public kernels: we were already far ahead of that kernel's LB score when I found it, but the method we learned from it produced a critical improvement in our score.",
      "votes": 4
    },
    {
      "id": 635453,
      "postDate": "2019-09-27T15:03:58.383Z",
      "content": "<p>This the PCA explained variance of the ArcNet features as I described above:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F3a5e5eb0aa2f983a663cc5d63e2fee3d%2Fvar_components.png?generation=1569596328000292&amp;alt=media\" alt=\"ArcNet features PCA explained variance\"></p>\n\n<p><code>\npca.fit(arcnet_features)\nplt.plot(pca.explained_variance_.cumsum())\n</code></p>",
      "rawMarkdown": "This the PCA explained variance of the ArcNet features as I described above:\n![ArcNet features PCA explained variance](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F3a5e5eb0aa2f983a663cc5d63e2fee3d%2Fvar_components.png?generation=1569596328000292&amp;alt=media)\n\n```\npca.fit(arcnet_features)\nplt.plot(pca.explained_variance_.cumsum())\n```\n",
      "votes": 1,
      "replies": [
        {
          "id": 635897,
          "postDate": "2019-09-28T12:34:03.203Z",
          "content": "<p>Thanks for the write-up <a href=\"/hmendonca\">@hmendonca</a> \nWhy did you use 200 features then? Shouldn't 100 be enough?</p>",
          "rawMarkdown": "Thanks for the write-up @hmendonca \nWhy did you use 200 features then? Shouldn't 100 be enough?"
        },
        {
          "id": 636880,
          "postDate": "2019-09-30T09:53:07.087Z",
          "content": "<p>Good question\nI wanted to leave some room for model learn further features, as the validation score was about 80% at that point.\nI did try to slowly reduce the number of features (to 180 and then 150), but didn't seem to help.\nAfter starting using pseudo labels, I have actually increased it to 320 to allow the models to [over]fit to the test data. We saw some improvements but didn't have time to test anything else...</p>",
          "rawMarkdown": "Good question\nI wanted to leave some room for model learn further features, as the validation score was about 80% at that point.\nI did try to slowly reduce the number of features (to 180 and then 150), but didn't seem to help.\nAfter starting using pseudo labels, I have actually increased it to 320 to allow the models to [over]fit to the test data. We saw some improvements but didn't have time to test anything else...",
          "votes": 1
        }
      ]
    },
    {
      "id": 635162,
      "postDate": "2019-09-27T08:03:10.973Z",
      "content": "<p>congratulations,it reminds me of the quote \"slow and steady wins the race\",,you richly deserved gold medal from this competition,,bad luck,better luck next time</p>",
      "rawMarkdown": "congratulations,it reminds me of the quote \"slow and steady wins the race\",,you richly deserved gold medal from this competition,,bad luck,better luck next time",
      "votes": 1
    },
    {
      "id": 634990,
      "postDate": "2019-09-27T03:08:38.790Z",
      "content": "<p>Many thanks for sharing your code solutions! Congratulations! <a href=\"/zaharch\">@zaharch</a>: Congratulations also to your team. Looking forward to your detail explanation/steps on how to recreate your model training via Pytorch/XLA TPU-based! :-)</p>",
      "rawMarkdown": "Many thanks for sharing your code solutions! Congratulations! @zaharch: Congratulations also to your team. Looking forward to your detail explanation/steps on how to recreate your model training via Pytorch/XLA TPU-based! :-)",
      "votes": 1
    },
    {
      "id": 634967,
      "postDate": "2019-09-27T02:20:24.720Z",
      "content": "<p>Awesome work, congratulations with the silver medal. Training with 4432 classes is a great idea. Have you experienced any issues working with pipelines of kernels? Like connection lost, any errors, memory\\disk limitations? Very interesting to here you experience on that.</p>",
      "rawMarkdown": "Awesome work, congratulations with the silver medal. Training with 4432 classes is a great idea. Have you experienced any issues working with pipelines of kernels? Like connection lost, any errors, memory\\disk limitations? Very interesting to here you experience on that.",
      "votes": 1,
      "replies": [
        {
          "id": 635269,
          "postDate": "2019-09-27T10:23:47.087Z",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> Cheers and congrats for your team's gold! ;) I also want to try using XLA next!</p>\n\n<p>I found normal jupyter notebooks quite reliable actually, if it disconnects it should reconnect fine and sync the outputs, but I don't really want to leave the browser open the whole time...\nNow Colab and kaggle on interactive/edit mode can be very frustrating!!\nI generally just run it for like 30min to check if things are going fine, then either commit (in kaggle) or generate the py script and run it with 'ipython -i' on top of 'screen' (great tool!)\nUsing 'ipython -i' allows you to still go in and debug it in case it crashes...</p>",
          "rawMarkdown": "@zaharch Cheers and congrats for your team's gold! ;) I also want to try using XLA next!\n\nI found normal jupyter notebooks quite reliable actually, if it disconnects it should reconnect fine and sync the outputs, but I don't really want to leave the browser open the whole time...\nNow Colab and kaggle on interactive/edit mode can be very frustrating!!\nI generally just run it for like 30min to check if things are going fine, then either commit (in kaggle) or generate the py script and run it with 'ipython -i' on top of 'screen' (great tool!)\nUsing 'ipython -i' allows you to still go in and debug it in case it crashes...",
          "votes": 2
        }
      ]
    },
    {
      "id": 636178,
      "postDate": "2019-09-29T00:00:33.183Z",
      "content": "<p>Really love this writeup. Particularly how you step through the different things you tried and how they mapped to improvements in your score. Very easy to follow. </p>",
      "rawMarkdown": "Really love this writeup. Particularly how you step through the different things you tried and how they mapped to improvements in your score. Very easy to follow. ",
      "votes": 2
    },
    {
      "id": 637331,
      "postDate": "2019-10-01T00:02:21.840Z",
      "content": "<p>How the \"new topic\" was started only 8 hours ago and the comments were made 4 days -14 hours ago? Is that the algorithm? \nI appreciated the \"Henrique's battle\" description.</p>",
      "rawMarkdown": "How the \"new topic\" was started only 8 hours ago and the comments were made 4 days -14 hours ago? Is that the algorithm? \nI appreciated the \"Henrique's battle\" description."
    },
    {
      "id": 635006,
      "postDate": "2019-09-27T03:49:39.150Z",
      "content": "<p>Great write-up!!\nThanks for your sharing!!\nCongrats <a href=\"/hmendonca\">@hmendonca</a> </p>",
      "rawMarkdown": "Great write-up!!\nThanks for your sharing!!\nCongrats @hmendonca "
    },
    {
      "id": 635028,
      "postDate": "2019-09-27T04:50:35.487Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 635074,
      "postDate": "2019-09-27T06:12:40.743Z",
      "content": "<p>thanks so much! Learn a lots from you</p>",
      "rawMarkdown": "thanks so much! Learn a lots from you"
    }
  ],
  "comments": [
    {
      "id": 635125,
      "author_name": "Giulia Savorgnan",
      "author_url": "",
      "post_date": "2019-09-27T07:21:36.597000",
      "content": "<p>I'm going to add a note on what did not work for me during model experimentation. <a href=\"/zaharch\">@zaharch</a>, this also partly addresses your question. I experimented with a siamese network + cosine embedding loss. I normalized the images to the mean and stddev of their experiment. I split the train set into 4 cell-types and trained 4 separate models. I used all train treatment+control siRNAs for training and the test control for validation (I was planning to use the CV folds of <a href=\"/hmendonca\">@hmendonca</a> later, but the model never achieved any decent results). A resnet50 backbone was starting to overfit almost immediately, so I discarded it. Efficientnet and densenet121 with dropout (~0.3-0.5) helped overcame the overfitting problem. In addition I added a very severe scheduling to reduce the learning rate, dropping by ~10% roughly every 10-20 iterations (not epochs!) without loss improvement; this helped the model avoid getting a nan loss. I modified the first backbone layer to accept the 6-channel inputs, and added a 128 encoder at the head. I initialized the network with ImageNet pretrained weights except in the first layer and in the encoder, where I left the random initialization. I chose a 128 encoder after experimenting with different bigger sizes that turned out to overfit (this seems in agreement with <a href=\"/hmendonca\">@hmendonca</a>'s PCA test). I trained on 6x128x128 inputs first, achieved convergence, then continued on 6x256x256. Augmentation was done with 2/3 random cropping and random rotation. The batch was composed of half positive and half negative pairs of images. The model was learning to distinguish positive and negative matches with an accuracy of ~70%. Then I started selecting the hardest negative pairs \"online\" (batch by batch), and the siamese accuracy improved to ~80-90%, however the model was not able at all to correctly classify the 1139 siRNAs (naively: calculate embeddings for all train samples xN augmentations, calculate embeddings for test samples, find the closest match). I trained a lightGBM classifier on the embeddings, achieved max 20% accuracy on the test control samples, but basically 0 accuracy on the LB score. Turned out the classifier was overconfident on the control classes only, probably because I did not balance the dataset. I added a crossentropy loss to the siamese network (so loss = cosine_embedding_loss + crossentropy_loss) and I observed an improvement in the distribution of the cosine_embedding_loss, but again very low accuracy on the test controls. At this stage I should have started using proper CV folds, a more clever LR scheduler, maybe triplet loss + crossentropy, but... <strong>I ran out of my portion of GCP credits. I had tried to use the kaggle kernels at the beginning but I gave up pretty quickly, frustrated with the too frequent disconnection errors. I remember I found completely impossible to use the kernels over the weekend, I was waiting 30 minutes to start a new kernel, another 30 to load the data, another hour in case I had to restart the kernel, and when the kernel was restarting by itself I often just closed the laptop lid and went for a hike...</strong> Unfortunately I did not have access to a local GPU not even for testing and debugging the code. It was my first time experimenting with such a model, I was probably slow and needed more computing resources to go a bit further. If anyone has any insight as to why my methodology did not work, I would love to hear you!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 635173,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-09-27T08:23:50.967000",
          "content": "<p>Great work! I was looking into siamese networks in the beginning too, but then realized that it was probably an outdated approach, with ArcFace and co outperforming it on few-short learning. Definitely you needed normal CV folds from the beginning, would have saved you a lot of time. Have you used GPUs with GCP on a virtual machine? And thank you for the contribution in this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 635202,
          "author_name": "Giulia Savorgnan",
          "author_url": "",
          "post_date": "2019-09-27T08:50:28.403000",
          "content": "<p>Yes, I spent my credits on a GPU VM. Good learning for the future. The only regret is not having pushed the model to the end because of kaggle kernels frustration :/ Now, with the GPU quota, I'm basically out of game. Thanks to you for your awesome sharing!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 635559,
          "author_name": "bioIn4",
          "author_url": "",
          "post_date": "2019-09-27T20:22:19.420000",
          "content": "<p>I am \"glad\" to see that others had such difficulty with a metric network, as it failed quite miserably for me.  I started a bit late on this competition and was quite focused on getting a metric learning method to work.  I would love to see if anyone had success in implementing such an approach.</p>\n\n<p>In short, I trained a VGG-19 with ImageNet weights, \"doubling\" the initial conv layer (depth-wise) to accept 6-channel images.  Following the base VGG, I added a FC (2048), dropout, FC(256);  this final 256 vector was my embedding.  I used N-pairs loss (had also tried triplet, but was swayed by the apparent improvements using n-pairs).  I started by training the FC layers, and then began unfreezing the conv layers.  </p>\n\n<p>I trained a different model for each cell-line, and my training sets were created by random sampling (likely not a very efficient way to learn).  My positives were the two sites from same well AND the same siRNA from different experiments, while negatives were random different siRNA.  By taking positive examples from different plates, I hoped that the model would learn the \"biology\" rather than any batch effects.</p>\n\n<p>Once I got the learning rate correct, I was getting very nice separation of the embeddings that I trained on.  A 2-d PCA showed very distinct clusters, which led me to believe I was training well.</p>\n\n<p>However, once applied to test set, it was quite poor.  After generating embeddings for the test set, I naively tried to assign based on KNN (3-6), nearest centroid etc....nothing got above 1% LB!!!  I then tried a number of transformations on the embeddings to see if it was just a simple \"shift\", due to the fact that the network had not seen these samples before.  Even comparing the embeddings of the control siRNA, none were \"displaced\" consistently between training and test sets (e.g. it was not something simple like a simple translation of the coordinate system).  My test embeddings showed some weak segregation by siRNA (via 2-d PCA), but they were still quite mixed. I ultimately became stuck at this last step and did not have time to try a classification based approach.  </p>\n\n<p>If anyone got a metric network to perform well in test, I'd love to hear how!  Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 635819,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2019-09-28T09:24:33.397000",
          "content": "<p><a href=\"/bioin4\">@bioin4</a> Many top solutions used ArcNet, which is considered metric learning (including us, please see the topic above). </p>\n\n<p>I believe it doesn't work very well from scratch if you have a large ArcNet margin ('m' param), as it probably need to have good features to start with, and imageNet wasn't good enough here.  But it worked fine for me with margin 0 and imageNet weights.</p>\n\n<p>I also found that hard mining didn't help the validation score much, as it tends to quickly overfit to the training data. Hard mining is used in most metric learning pipelines if I'm not wrong.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 635319,
      "author_name": "Andy Harless",
      "author_url": "",
      "post_date": "2019-09-27T11:30:28.140000",
      "content": "<p>I would want to emphasize the role of <a href=\"/christopherberner\">@christopherberner</a> 's \"Hungarian algorithm\" kernel in helping us to exploit the plates leak (and to exploit the overall structure, which would have been helpful even without the leak).  I had an intuitive sense that we needed to find an optimal way to assign siRNAs within a plate given a set of probabilities, but I didn't know the appropriate method until I found that kernel.  There's a lesson here about learning from public kernels: we were already far ahead of that kernel's LB score when I found it, but the method we learned from it produced a critical improvement in our score.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 635453,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2019-09-27T15:03:58.383000",
      "content": "<p>This the PCA explained variance of the ArcNet features as I described above:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F3a5e5eb0aa2f983a663cc5d63e2fee3d%2Fvar_components.png?generation=1569596328000292&amp;alt=media\" alt=\"ArcNet features PCA explained variance\"></p>\n\n<p><code>\npca.fit(arcnet_features)\nplt.plot(pca.explained_variance_.cumsum())\n</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 635897,
          "author_name": "Kim Wilson",
          "author_url": "",
          "post_date": "2019-09-28T12:34:03.203000",
          "content": "<p>Thanks for the write-up <a href=\"/hmendonca\">@hmendonca</a> \nWhy did you use 200 features then? Shouldn't 100 be enough?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 636880,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2019-09-30T09:53:07.087000",
          "content": "<p>Good question\nI wanted to leave some room for model learn further features, as the validation score was about 80% at that point.\nI did try to slowly reduce the number of features (to 180 and then 150), but didn't seem to help.\nAfter starting using pseudo labels, I have actually increased it to 320 to allow the models to [over]fit to the test data. We saw some improvements but didn't have time to test anything else...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 635162,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-09-27T08:03:10.973000",
      "content": "<p>congratulations,it reminds me of the quote \"slow and steady wins the race\",,you richly deserved gold medal from this competition,,bad luck,better luck next time</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 634990,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-09-27T03:08:38.790000",
      "content": "<p>Many thanks for sharing your code solutions! Congratulations! <a href=\"/zaharch\">@zaharch</a>: Congratulations also to your team. Looking forward to your detail explanation/steps on how to recreate your model training via Pytorch/XLA TPU-based! :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 634967,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2019-09-27T02:20:24.720000",
      "content": "<p>Awesome work, congratulations with the silver medal. Training with 4432 classes is a great idea. Have you experienced any issues working with pipelines of kernels? Like connection lost, any errors, memory\\disk limitations? Very interesting to here you experience on that.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 635269,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2019-09-27T10:23:47.087000",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> Cheers and congrats for your team's gold! ;) I also want to try using XLA next!</p>\n\n<p>I found normal jupyter notebooks quite reliable actually, if it disconnects it should reconnect fine and sync the outputs, but I don't really want to leave the browser open the whole time...\nNow Colab and kaggle on interactive/edit mode can be very frustrating!!\nI generally just run it for like 30min to check if things are going fine, then either commit (in kaggle) or generate the py script and run it with 'ipython -i' on top of 'screen' (great tool!)\nUsing 'ipython -i' allows you to still go in and debug it in case it crashes...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 636178,
      "author_name": "Anthony Goldbloom",
      "author_url": "",
      "post_date": "2019-09-29T00:00:33.183000",
      "content": "<p>Really love this writeup. Particularly how you step through the different things you tried and how they mapped to improvements in your score. Very easy to follow. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 637331,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2019-10-01T00:02:21.840000",
      "content": "<p>How the \"new topic\" was started only 8 hours ago and the comments were made 4 days -14 hours ago? Is that the algorithm? \nI appreciated the \"Henrique's battle\" description.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635006,
      "author_name": "0xFunky",
      "author_url": "",
      "post_date": "2019-09-27T03:49:39.150000",
      "content": "<p>Great write-up!!\nThanks for your sharing!!\nCongrats <a href=\"/hmendonca\">@hmendonca</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635028,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T04:50:35.487000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635074,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2019-09-27T06:12:40.743000",
      "content": "<p>thanks so much! Learn a lots from you</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "634962": "I will try to summarise here my part of the 2 months battle with Kaggle kernels in this comp.\nOur team's solution was basically entirely run on kaggle GPU's (mostly prior to the new quotas).\nThe final kernel is here: https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759\n\n**Quick outline in chronological order:**\n 1. CV split by experiments about ⅓ of the data, 3 folds.\n 2. Trained a quick and dirty classifier (left it training while analysing the data) with all 1108 classes, 6 channels and no augmentation. OK: **LB 0.10-0.25** and established a reasonable learning rate schedule which was used in the rest of the comp.\n 3. Normalised the data by experiment mean and std, added random crops: **LB 0.32**\nAdded more aug: no improvement.\n 4. Added ArcNet head with 512 features and default params, using pre-trained model from above, eventually **LB 0.35**\n 5. Started training with each cell type as a different class 1108 * 4 = 4432 classes: **LB 0.44**\n 6. Cell type balancing in training data: **LB 0.45**\n 7. Added the 31 controls to the training data (but not to validation) 1139 * 4 = 4556 classes: LB 0.46\n 8. Added a custom cosine distance loss to the ArcNet features (same class should be close to each other, different classes far...) **LB 0.47**\n 9. PCA showed that 100 dimensions explained all the variance in features, so I reduced the ArcNet features to 200 (double, just in case ;) **LB 0.52**\n 10. Applied plates leak, loads of normalisation and assigned single treatment per plate LB 0.82 !!\nMore normalisation, and lots of multi-fold ensembles later: **LB 0.88** (Thanks to @aharless and @giuliasavorgnan)\n 11. Used the former predictions as pseudo labels and trained further with train + test data, eventually **LB 0.93** (this process was repeated a few times, always training the latest models with pseudo labels from the best possible submission/ensemble in an amazing team effort! Thanks @aharless @giuliasavorgnan @stillsut @lastlegion)\n 12. Failed badly trying to blend the new best models (single models had better LB than the blends) probably due to models overfitting to the pseudo labels?\nReduce the amount of pseudo labels in the training data (from 50/50% to 33/66% now) and improved feature distance metric: **LB 0.943**\n 13. Out of gas, $$ and quota, but I have never learned so much in a comp before! Super happy with the result\n\nThe code for all that is in the [kernel linked above](https://www.kaggle.com/hmendonca/fold1h4r3-arcenetb4-2-256px-rcic-lb-0-9759) (including some visualisations).\n\n\n**In the beginning**\nI basically started looking at the data and decided to just try a normal classifier without any fanciness. I ported an EfficientNetB4 kernel from another comp and watched the batch loss for a few different learning rates.\nThe loss during cosine annealing with warm restarts is very informative, almost as much as the LR finder scheduler used in fast.ai, so after a couple of runs you can work out what is too high and what is too low, and also how much you need to decrease it from epoch to epoch.\nI was training 3 folds and later 3 architectures (EfficientNet B4, B5 and seResNet-xt101). However, the journey described above is actually mainly from the initial B4 I used, and that also ended up being the best model and submission. Its weights were carved slowly by training several different heads one after another, with several branches used to experiment with different ideas.\nUnfortunately, we didn't have the budget to do clean experiments and try each idea with their optimal hyper-params and initial conditions. Therefore, some changes used here might not have been optimal but still ended up improving the score, so I had to keep them in the pipeline for the lack of a better reference.\n\n**Normalisation and Post-processing:**\nThe experiment setup let out a lot of information, and thanks to the public kernels that was made clear for everyone. That played a huge role for us to improve the submission score and create the pseudo labels. Cheers @zaharch and @christopherberner !\nLooking at the raw model predictions we could see that the initial models were obviously preferring some treatments to others. The predictions weren't balanced across the different classes, but we knew that they should, so a basic idea was to normalise across all treatments. Each treatment should appear exactly the same number of times in each experiment. Normalising the prediction probabilities across the siRNA already helped increasing the validation and LB score.\nAndy and Giulia went further and created normalisation loop iteratively normalising the probabilities vertically and horizontally until convergence.\nThis process showed great improvements in both validation and LB scores and was crucial to our solution. \n\n**Platform**\nAs I mentioned before, the solution was run on GPU only.\nTPU's would have been amazing but I had a look initially and wasn't super excited about its pytorch support. I have also trained a ResNet101 model using Colab and the provided notebook in the comp's intro, but I found it a bit too annoying to work with all the deprecated TF code (hopefully it will get better soon as they're fixing the new TF2.0), so I gave up after a few days and went back to PyTorch. Maybe that was a mistake?\nWe had a few GCP credits but $400 goes really quick, and we only made a few experiments in the beginning. I also used some more (including a bit of my own pocket) to train another few models improving the pseudo labels a bit. However, our best model was trained exclusively in kaggle in about 13 runs in total (100hrs approximately, excluding all abandoned branches)\n\nI tried Colab again after the kaggle quotas were introduced, but I found their GPU notebook quite flaky, as it keeps disconnecting my runs during training :(\nMaybe I should buy a GPU lol\n\nMany thanks to all the organisers and the amazing community of kagglers always sharing some much!",
    "635125": "I'm going to add a note on what did not work for me during model experimentation. @zaharch, this also partly addresses your question. I experimented with a siamese network + cosine embedding loss. I normalized the images to the mean and stddev of their experiment. I split the train set into 4 cell-types and trained 4 separate models. I used all train treatment+control siRNAs for training and the test control for validation (I was planning to use the CV folds of @hmendonca later, but the model never achieved any decent results). A resnet50 backbone was starting to overfit almost immediately, so I discarded it. Efficientnet and densenet121 with dropout (~0.3-0.5) helped overcame the overfitting problem. In addition I added a very severe scheduling to reduce the learning rate, dropping by ~10% roughly every 10-20 iterations (not epochs!) without loss improvement; this helped the model avoid getting a nan loss. I modified the first backbone layer to accept the 6-channel inputs, and added a 128 encoder at the head. I initialized the network with ImageNet pretrained weights except in the first layer and in the encoder, where I left the random initialization. I chose a 128 encoder after experimenting with different bigger sizes that turned out to overfit (this seems in agreement with @hmendonca's PCA test). I trained on 6x128x128 inputs first, achieved convergence, then continued on 6x256x256. Augmentation was done with 2/3 random cropping and random rotation. The batch was composed of half positive and half negative pairs of images. The model was learning to distinguish positive and negative matches with an accuracy of ~70%. Then I started selecting the hardest negative pairs \"online\" (batch by batch), and the siamese accuracy improved to ~80-90%, however the model was not able at all to correctly classify the 1139 siRNAs (naively: calculate embeddings for all train samples xN augmentations, calculate embeddings for test samples, find the closest match). I trained a lightGBM classifier on the embeddings, achieved max 20% accuracy on the test control samples, but basically 0 accuracy on the LB score. Turned out the classifier was overconfident on the control classes only, probably because I did not balance the dataset. I added a crossentropy loss to the siamese network (so loss = cosine\\_embedding\\_loss + crossentropy\\_loss) and I observed an improvement in the distribution of the cosine\\_embedding\\_loss, but again very low accuracy on the test controls. At this stage I should have started using proper CV folds, a more clever LR scheduler, maybe triplet loss + crossentropy, but... **I ran out of my portion of GCP credits. I had tried to use the kaggle kernels at the beginning but I gave up pretty quickly, frustrated with the too frequent disconnection errors. I remember I found completely impossible to use the kernels over the weekend, I was waiting 30 minutes to start a new kernel, another 30 to load the data, another hour in case I had to restart the kernel, and when the kernel was restarting by itself I often just closed the laptop lid and went for a hike...** Unfortunately I did not have access to a local GPU not even for testing and debugging the code. It was my first time experimenting with such a model, I was probably slow and needed more computing resources to go a bit further. If anyone has any insight as to why my methodology did not work, I would love to hear you!",
    "635319": "I would want to emphasize the role of @christopherberner 's \"Hungarian algorithm\" kernel in helping us to exploit the plates leak (and to exploit the overall structure, which would have been helpful even without the leak).  I had an intuitive sense that we needed to find an optimal way to assign siRNAs within a plate given a set of probabilities, but I didn't know the appropriate method until I found that kernel.  There's a lesson here about learning from public kernels: we were already far ahead of that kernel's LB score when I found it, but the method we learned from it produced a critical improvement in our score.",
    "635453": "This the PCA explained variance of the ArcNet features as I described above:\n![ArcNet features PCA explained variance](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F3a5e5eb0aa2f983a663cc5d63e2fee3d%2Fvar_components.png?generation=1569596328000292&amp;alt=media)\n\n```\npca.fit(arcnet_features)\nplt.plot(pca.explained_variance_.cumsum())\n```\n",
    "635162": "congratulations,it reminds me of the quote \"slow and steady wins the race\",,you richly deserved gold medal from this competition,,bad luck,better luck next time",
    "634990": "Many thanks for sharing your code solutions! Congratulations! @zaharch: Congratulations also to your team. Looking forward to your detail explanation/steps on how to recreate your model training via Pytorch/XLA TPU-based! :-)",
    "634967": "Awesome work, congratulations with the silver medal. Training with 4432 classes is a great idea. Have you experienced any issues working with pipelines of kernels? Like connection lost, any errors, memory\\disk limitations? Very interesting to here you experience on that.",
    "636178": "Really love this writeup. Particularly how you step through the different things you tried and how they mapped to improvements in your score. Very easy to follow. ",
    "637331": "How the \"new topic\" was started only 8 hours ago and the comments were made 4 days -14 hours ago? Is that the algorithm? \nI appreciated the \"Henrique's battle\" description.",
    "635006": "Great write-up!!\nThanks for your sharing!!\nCongrats @hmendonca ",
    "635028": "",
    "635074": "thanks so much! Learn a lots from you"
  }
}