{
  "id": 110393,
  "title": "DenseNet-121 on a single GPU [LB 0.97520]",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110393",
  "author_name": "Mikhail Papkov",
  "post_date": "2019-09-27T09:54:45.656000",
  "votes": 17,
  "comment_count": 1,
  "views": 0,
  "content": "<h2>Disclaimer</h2>\n\n<p>From the title, it might look like a voluntary challenge but originally it was not. Before the competition, I knew very little about modern approaches in classification: all kinds of Face-family losses, modern architectures for classification and so on. My main motivation was to work with microscopy images (since it is kinda my area of expertise) and to learn things starting from PyTorch which I have never used before and decided to familiarize myself with. So I would like to thank competition organizers for such a delightful journey, it was very informative for me. Beware of noob approaches below.</p>\n\n<h2>Highlights</h2>\n\n<ul>\n<li>A single V100, no GCP, no Kaggle kernels</li>\n<li>DenseNet-121 + softmax (standard cross-entropy loss)</li>\n<li>6 channels</li>\n<li>Normalization by plate (worked better than by experiment)</li>\n<li>Simple augmentations: rotate, flip, crop 256/384, jitter color</li>\n<li>Use controls</li>\n<li>SGD+momentum, warm restarts from lr 0.1, batch size 32</li>\n<li>Hard pseudolabeling</li>\n</ul>\n\n<h2>Pipeline</h2>\n\n<h3>Stage 1: train a master model for all experiments</h3>\n\n<p>Here I have struggled a lot to get my first model converged. At first, I normalized all the data together disregarding experimental layout. This decision was poor but I later reused this model to continue training with other types of normalization (by cell line, by experiment and then by plate). Progressive resizing also helped the network to converge a bit faster. I included all the controls (train+val+test) in the train set, it improved the performance. </p>\n\n<h3>Stage 2: train cell line models</h3>\n\n<p>From the pretrained master model it is natural to proceed with line-specific models. Here I have just continued to train with the same hyperparameters and warm restarts.</p>\n\n<h3>Stage 3: fine-tune using <em>only</em> controls from each experiment</h3>\n\n<p>From the very beginning, it was clear that using control experiments is crucial for success. Although this stage still seems the most counter-intuitive to me, it worked. For each experiment in val/test set take its control samples and train the respective cell line model for 1 epoch without augmentation (no random crops, no color jitter) with lower learning rate (I used 0.01). It always gave me +3-5% accuracy on the validation set. </p>\n\n<p>A similar approach can be applied to metric learning but without retraining. First, extract features from train controls and test controls, calculate per-class mean (\"template\").  Next, subtract train template from test template, take the mean by feature. As a result, you should receive a correction vector of shape (N features, ). Subtract this vector from your test features and you will have the same +3-5%. Although I never used this approach in my final submission, it seems interesting despite being super-naive.</p>\n\n<h3>Stage 4: prediction aggregation</h3>\n\n<p>First, I normalized network outputs (before softmax) by mean and std. It gave +1-2% validation accuracy. Predictions were averaged by site. Experimental design (here it was called a leak) helped a lot. Every siRNA appeared only once in an experiment in groups 277 by plate. So I softmaxed outputs within the respective group (others set to 0) and assigned classes based on probabilities.  I wrote my own algorithm for the assignment which iteratively selects a max probability, picks a class and removes the row from further selection. Hungarian algorithm should be more efficient.</p>\n\n<h3>Stage 5: iterative hard pseudolabeling</h3>\n\n<p>Because our final predictions are much stronger than the raw network outputs, we can use them to label test instances and therefore obtain better predictions going back to Stage 4. I trained three models with different validation folds (leaving out experiments 01, 02 and 03 from each cell line), aggregated their predictions, used ensemble labels as pseudolabels for each of those and repeated the process. It gave my final result 0.97520.</p>\n\n<h2>Failures</h2>\n\n<ul>\n<li>ArcFace: I have managed to squeeze decent results from it 2 days before the deadline, it obviously did not help and distracted me from pedantic ensemble construction. I am pretty sure that my position could be higher either with more pseudolabeling iterations or with metric learning approach if I would use it from the very beginning. But I have had hard times understanding how it should work.</li>\n<li>Mixup: I did not spend too much time here, but results were disappointing (given original paper promises)</li>\n<li>Not using free resources: this one was really stupid. Being too lazy to migrate somewhere from your machine obviously will not make results better.</li>\n<li>Poor validation set selection: people did smart things, leaving out hard ceases like HUVEC-05, I did not.</li>\n</ul>\n\n<p>Now I would have done many things differently, and, I think, this feeling was the main competition purpose for me after all. My best single model would finish 31st with a score of 0.97367. </p>",
  "messages": [
    {
      "id": 635246,
      "postDate": "2019-09-27T09:54:45.657Z",
      "content": "<h2>Disclaimer</h2>\n\n<p>From the title, it might look like a voluntary challenge but originally it was not. Before the competition, I knew very little about modern approaches in classification: all kinds of Face-family losses, modern architectures for classification and so on. My main motivation was to work with microscopy images (since it is kinda my area of expertise) and to learn things starting from PyTorch which I have never used before and decided to familiarize myself with. So I would like to thank competition organizers for such a delightful journey, it was very informative for me. Beware of noob approaches below.</p>\n\n<h2>Highlights</h2>\n\n<ul>\n<li>A single V100, no GCP, no Kaggle kernels</li>\n<li>DenseNet-121 + softmax (standard cross-entropy loss)</li>\n<li>6 channels</li>\n<li>Normalization by plate (worked better than by experiment)</li>\n<li>Simple augmentations: rotate, flip, crop 256/384, jitter color</li>\n<li>Use controls</li>\n<li>SGD+momentum, warm restarts from lr 0.1, batch size 32</li>\n<li>Hard pseudolabeling</li>\n</ul>\n\n<h2>Pipeline</h2>\n\n<h3>Stage 1: train a master model for all experiments</h3>\n\n<p>Here I have struggled a lot to get my first model converged. At first, I normalized all the data together disregarding experimental layout. This decision was poor but I later reused this model to continue training with other types of normalization (by cell line, by experiment and then by plate). Progressive resizing also helped the network to converge a bit faster. I included all the controls (train+val+test) in the train set, it improved the performance. </p>\n\n<h3>Stage 2: train cell line models</h3>\n\n<p>From the pretrained master model it is natural to proceed with line-specific models. Here I have just continued to train with the same hyperparameters and warm restarts.</p>\n\n<h3>Stage 3: fine-tune using <em>only</em> controls from each experiment</h3>\n\n<p>From the very beginning, it was clear that using control experiments is crucial for success. Although this stage still seems the most counter-intuitive to me, it worked. For each experiment in val/test set take its control samples and train the respective cell line model for 1 epoch without augmentation (no random crops, no color jitter) with lower learning rate (I used 0.01). It always gave me +3-5% accuracy on the validation set. </p>\n\n<p>A similar approach can be applied to metric learning but without retraining. First, extract features from train controls and test controls, calculate per-class mean (\"template\").  Next, subtract train template from test template, take the mean by feature. As a result, you should receive a correction vector of shape (N features, ). Subtract this vector from your test features and you will have the same +3-5%. Although I never used this approach in my final submission, it seems interesting despite being super-naive.</p>\n\n<h3>Stage 4: prediction aggregation</h3>\n\n<p>First, I normalized network outputs (before softmax) by mean and std. It gave +1-2% validation accuracy. Predictions were averaged by site. Experimental design (here it was called a leak) helped a lot. Every siRNA appeared only once in an experiment in groups 277 by plate. So I softmaxed outputs within the respective group (others set to 0) and assigned classes based on probabilities.  I wrote my own algorithm for the assignment which iteratively selects a max probability, picks a class and removes the row from further selection. Hungarian algorithm should be more efficient.</p>\n\n<h3>Stage 5: iterative hard pseudolabeling</h3>\n\n<p>Because our final predictions are much stronger than the raw network outputs, we can use them to label test instances and therefore obtain better predictions going back to Stage 4. I trained three models with different validation folds (leaving out experiments 01, 02 and 03 from each cell line), aggregated their predictions, used ensemble labels as pseudolabels for each of those and repeated the process. It gave my final result 0.97520.</p>\n\n<h2>Failures</h2>\n\n<ul>\n<li>ArcFace: I have managed to squeeze decent results from it 2 days before the deadline, it obviously did not help and distracted me from pedantic ensemble construction. I am pretty sure that my position could be higher either with more pseudolabeling iterations or with metric learning approach if I would use it from the very beginning. But I have had hard times understanding how it should work.</li>\n<li>Mixup: I did not spend too much time here, but results were disappointing (given original paper promises)</li>\n<li>Not using free resources: this one was really stupid. Being too lazy to migrate somewhere from your machine obviously will not make results better.</li>\n<li>Poor validation set selection: people did smart things, leaving out hard ceases like HUVEC-05, I did not.</li>\n</ul>\n\n<p>Now I would have done many things differently, and, I think, this feeling was the main competition purpose for me after all. My best single model would finish 31st with a score of 0.97367. </p>",
      "rawMarkdown": "## Disclaimer\nFrom the title, it might look like a voluntary challenge but originally it was not. Before the competition, I knew very little about modern approaches in classification: all kinds of Face-family losses, modern architectures for classification and so on. My main motivation was to work with microscopy images (since it is kinda my area of expertise) and to learn things starting from PyTorch which I have never used before and decided to familiarize myself with. So I would like to thank competition organizers for such a delightful journey, it was very informative for me. Beware of noob approaches below.\n\n\n## Highlights\n- A single V100, no GCP, no Kaggle kernels\n- DenseNet-121 + softmax (standard cross-entropy loss)\n- 6 channels\n- Normalization by plate (worked better than by experiment)\n- Simple augmentations: rotate, flip, crop 256/384, jitter color\n- Use controls\n- SGD+momentum, warm restarts from lr 0.1, batch size 32\n- Hard pseudolabeling\n\n## Pipeline\n### Stage 1: train a master model for all experiments\nHere I have struggled a lot to get my first model converged. At first, I normalized all the data together disregarding experimental layout. This decision was poor but I later reused this model to continue training with other types of normalization (by cell line, by experiment and then by plate). Progressive resizing also helped the network to converge a bit faster. I included all the controls (train+val+test) in the train set, it improved the performance. \n\n### Stage 2: train cell line models\nFrom the pretrained master model it is natural to proceed with line-specific models. Here I have just continued to train with the same hyperparameters and warm restarts.\n\n### Stage 3: fine-tune using _only_ controls from each experiment\nFrom the very beginning, it was clear that using control experiments is crucial for success. Although this stage still seems the most counter-intuitive to me, it worked. For each experiment in val/test set take its control samples and train the respective cell line model for 1 epoch without augmentation (no random crops, no color jitter) with lower learning rate (I used 0.01). It always gave me +3-5% accuracy on the validation set. \n\nA similar approach can be applied to metric learning but without retraining. First, extract features from train controls and test controls, calculate per-class mean (\"template\").  Next, subtract train template from test template, take the mean by feature. As a result, you should receive a correction vector of shape (N features, ). Subtract this vector from your test features and you will have the same +3-5%. Although I never used this approach in my final submission, it seems interesting despite being super-naive.\n\n###Stage 4: prediction aggregation\nFirst, I normalized network outputs (before softmax) by mean and std. It gave +1-2% validation accuracy. Predictions were averaged by site. Experimental design (here it was called a leak) helped a lot. Every siRNA appeared only once in an experiment in groups 277 by plate. So I softmaxed outputs within the respective group (others set to 0) and assigned classes based on probabilities.  I wrote my own algorithm for the assignment which iteratively selects a max probability, picks a class and removes the row from further selection. Hungarian algorithm should be more efficient.\n\n###Stage 5: iterative hard pseudolabeling\nBecause our final predictions are much stronger than the raw network outputs, we can use them to label test instances and therefore obtain better predictions going back to Stage 4. I trained three models with different validation folds (leaving out experiments 01, 02 and 03 from each cell line), aggregated their predictions, used ensemble labels as pseudolabels for each of those and repeated the process. It gave my final result 0.97520.\n\n## Failures\n- ArcFace: I have managed to squeeze decent results from it 2 days before the deadline, it obviously did not help and distracted me from pedantic ensemble construction. I am pretty sure that my position could be higher either with more pseudolabeling iterations or with metric learning approach if I would use it from the very beginning. But I have had hard times understanding how it should work.\n- Mixup: I did not spend too much time here, but results were disappointing (given original paper promises)\n- Not using free resources: this one was really stupid. Being too lazy to migrate somewhere from your machine obviously will not make results better.\n- Poor validation set selection: people did smart things, leaving out hard ceases like HUVEC-05, I did not.\n\nNow I would have done many things differently, and, I think, this feeling was the main competition purpose for me after all. My best single model would finish 31st with a score of 0.97367. \n",
      "votes": 17
    },
    {
      "id": 635266,
      "postDate": "2019-09-27T10:22:24.637Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 635266,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T10:22:24.637000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635246": "## Disclaimer\nFrom the title, it might look like a voluntary challenge but originally it was not. Before the competition, I knew very little about modern approaches in classification: all kinds of Face-family losses, modern architectures for classification and so on. My main motivation was to work with microscopy images (since it is kinda my area of expertise) and to learn things starting from PyTorch which I have never used before and decided to familiarize myself with. So I would like to thank competition organizers for such a delightful journey, it was very informative for me. Beware of noob approaches below.\n\n\n## Highlights\n- A single V100, no GCP, no Kaggle kernels\n- DenseNet-121 + softmax (standard cross-entropy loss)\n- 6 channels\n- Normalization by plate (worked better than by experiment)\n- Simple augmentations: rotate, flip, crop 256/384, jitter color\n- Use controls\n- SGD+momentum, warm restarts from lr 0.1, batch size 32\n- Hard pseudolabeling\n\n## Pipeline\n### Stage 1: train a master model for all experiments\nHere I have struggled a lot to get my first model converged. At first, I normalized all the data together disregarding experimental layout. This decision was poor but I later reused this model to continue training with other types of normalization (by cell line, by experiment and then by plate). Progressive resizing also helped the network to converge a bit faster. I included all the controls (train+val+test) in the train set, it improved the performance. \n\n### Stage 2: train cell line models\nFrom the pretrained master model it is natural to proceed with line-specific models. Here I have just continued to train with the same hyperparameters and warm restarts.\n\n### Stage 3: fine-tune using _only_ controls from each experiment\nFrom the very beginning, it was clear that using control experiments is crucial for success. Although this stage still seems the most counter-intuitive to me, it worked. For each experiment in val/test set take its control samples and train the respective cell line model for 1 epoch without augmentation (no random crops, no color jitter) with lower learning rate (I used 0.01). It always gave me +3-5% accuracy on the validation set. \n\nA similar approach can be applied to metric learning but without retraining. First, extract features from train controls and test controls, calculate per-class mean (\"template\").  Next, subtract train template from test template, take the mean by feature. As a result, you should receive a correction vector of shape (N features, ). Subtract this vector from your test features and you will have the same +3-5%. Although I never used this approach in my final submission, it seems interesting despite being super-naive.\n\n###Stage 4: prediction aggregation\nFirst, I normalized network outputs (before softmax) by mean and std. It gave +1-2% validation accuracy. Predictions were averaged by site. Experimental design (here it was called a leak) helped a lot. Every siRNA appeared only once in an experiment in groups 277 by plate. So I softmaxed outputs within the respective group (others set to 0) and assigned classes based on probabilities.  I wrote my own algorithm for the assignment which iteratively selects a max probability, picks a class and removes the row from further selection. Hungarian algorithm should be more efficient.\n\n###Stage 5: iterative hard pseudolabeling\nBecause our final predictions are much stronger than the raw network outputs, we can use them to label test instances and therefore obtain better predictions going back to Stage 4. I trained three models with different validation folds (leaving out experiments 01, 02 and 03 from each cell line), aggregated their predictions, used ensemble labels as pseudolabels for each of those and repeated the process. It gave my final result 0.97520.\n\n## Failures\n- ArcFace: I have managed to squeeze decent results from it 2 days before the deadline, it obviously did not help and distracted me from pedantic ensemble construction. I am pretty sure that my position could be higher either with more pseudolabeling iterations or with metric learning approach if I would use it from the very beginning. But I have had hard times understanding how it should work.\n- Mixup: I did not spend too much time here, but results were disappointing (given original paper promises)\n- Not using free resources: this one was really stupid. Being too lazy to migrate somewhere from your machine obviously will not make results better.\n- Poor validation set selection: people did smart things, leaving out hard ceases like HUVEC-05, I did not.\n\nNow I would have done many things differently, and, I think, this feeling was the main competition purpose for me after all. My best single model would finish 31st with a score of 0.97367. \n",
    "635266": ""
  }
}