{
  "id": 110335,
  "title": "7th place solution summary",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110335",
  "author_name": "nosound",
  "post_date": "2019-09-27T00:53:33.685000",
  "votes": 48,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thank you Recursion and Kaggle for organizing this unique and challenging competition! When I decided to participate in it, I was impressed by the quality of data and organization, <a href=\"https://www.rxrx.ai/\">this beautiful site</a> alone gave me a lot of motivation. GCP and TPU credits were also indispensable for me, hope to see more in the future. The leaks that were found were handled smoothly by Kaggle, specifically thanks to <a href=\"https://www.kaggle.com/sohier\">Sohier Dane</a>. And credits to <a href=\"https://www.kaggle.com/giuliasavorgnan\">Giulia Savorgnan</a> for reporting the second leak. </p>\n\n<p>I was very fortunate to team up with Yuval on this one. His models scored and maintained top10 for very long time, and I was able to contribute meaningfully only much later in the competition. I still can't wrap my head around how it was possible for him to put so much work in this competition, work full-time at a day job, but also prepare for and run UTMB 171k ultra-marathon race at the end of August. He finished after 50k due to injury, but in my eyes it is already a super-human level of toughness. </p>\n\n<h1>Setup</h1>\n\n<p>All my models I trained with pytorch and TPUs, and used exactly all 600$ of the free GCP credit that I had. My setup I described <a href=\"https://towardsdatascience.com/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4\">in this post</a>. Despite being able to get overall good training speeds, the experience using TPUs with pytorch was rough. The main problem was that it hangs unexpectedly after a few hours of training. At times it was so annoying that I questioned myself if I lived this life correctly. Pytorch/XLA guys tried to help on the forum, but at the moment I think Pytorch/XLA is just not quite production ready (but I will still be happy to receive free TPUs for the next competition!).</p>\n\n<h1>The model</h1>\n\n<p>From the beginning I started with resnet50 and didn’t have an opportunity to successfully try anything else. In hindsight, it was a good choice for this competition. Yuval has used other backbones, I believe he will describe his work in a separate post. </p>\n\n<ul>\n<li>Resnet50 backbone</li>\n<li>ArcFace (m=0.2, s=30, 512 features), used “as is” from the beginning of training, no adjustments. A vanilla implementation <a href=\"https://arxiv.org/abs/1801.07698\">from the original article</a>.</li>\n<li>384x384 images cropped from original 512x512 with random shifts, flips and all angles rotations (albumentations package)</li>\n<li>Normalization of images per experiment and channel, with small randomization</li>\n<li>Not using well, plate and experiment meta info.</li>\n<li>3 folds each containing full experiments</li>\n<li>Training for about 10 epochs together and then a separate model per cell type.</li>\n<li>Each site is treated as a separate sample</li>\n<li>Adam optimizer, decreasing learning rate</li>\n<li>HUVEC-05 is removed, as it is too different from HUVEC test experiments</li>\n<li>HUVEC-18 is moved to the train set</li>\n</ul>\n\n<p>Besides this standard structure described above I had 3 special features in the model. I have not done any serious comparisons with and without them, but my feeling is that the first gives me a major boost, and second and third some additional smaller improvements (but a fair comparison vs regular pseudo-labeling is definitely missing).</p>\n\n<ol>\n<li><strong>Training with test</strong>. The model is trained on all train, test and control images together, predicting 1139 classes (1108 + 31 controls). For the test images I took their softmax output as a target for the log-loss. In this case it can be shown that the gradient of the features before the softmax is zero, that is, these samples have no effect. It makes sense intuitively, I take the predicted probability distribution as a target, - this is already the perfect prediction according to log-loss. But then I modify this target probabilities vector for each sample according to the rich structural information that we hold, - and this is quite unique for this competition. First, - <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905\">the plates leak</a>, all components besides 277 are zeroed out. Second, - each set of 277 samples contains every sirna only once. This second one was not easy to enforce, but as a soft constraint in the loss it worked reasonably well. Additional note here, - validation set data was in the train as well, with hidden sirnas, to simulate similar setting to test. One more thing, - ArcFace didn’t work well here, so it was turned off for test/validation samples. And the last point, - I think this approach in general can be thought of as a dynamic type of pseudo-labelling.</li>\n<li><strong>Mean normalization of ArcFace features</strong>, kudos to Yuval who pioneered this idea. I save EMA of 512 features before ArcFace for each sample, calculate their mean per experiment and overall mean (vectors of length 512 again), and to each sample on the forward pass to these features I add the overall mean and subtract its experiment mean. This way the features entering ArcFace don’t have per-experiment bias and we improve this side of domain adaptation problem. Theoretically.</li>\n<li><strong>Incremental hard pseudo-labelling (PL)</strong>. The PL in point 1 above is a probabilistic PL, in the sense that the target is a probabilities vector of size 277, but not forced to be a specific value, anything is good as long as it is in the set. In this addition to the model at the end of the training I try to push all reasonably confident test and validation samples to lock into one sirna, which becomes its target. I can then zero this sirna for all other samples in the 277-set, which helps them to converge. Once locked, this sample also gets into ArcFace (remember, it is turned off for the test). In practice, judging by validation experiments, first 200 samples out of 277 can be easily locked this way without a single error for all cell types. They were confident anyway, so probably not a huge benefit, but helpful. </li>\n</ol>\n\n<h1>Test-time-augmentation (TTA) and ensemble</h1>\n\n<p>For each sample and each site I run 16 predefined transformations, aggregating those 16 by quantile 75, and then aggregating by gmean between the sites. I then aggregated also by gmean between the folds, and again by gmean between different runs (both me and Yuval had 2 runs). Guess what aggregation function we used to merge the results between me and Yuval’s probabilities? That’s right, gmean again. It just worked very well everywhere, despite me trying 7-8 competitors. We selected aggregation weights for all of this by CV and public LB, different per cell type.</p>\n\n<h1>Mixed integer programming (MIP)</h1>\n\n<p>The output of the above procedure is a set of <code>18*4=72</code> matrices 277 by 277 where each row sums up to 1 (samples) but column sums range from 0.5 to 3 (sirnas). We want columns to sum up to one as well, as we know that each set contains every sirna only once. To achieve this we divided the matrices iteratively by sum of columns and then sum of rows several times (e.g. 10), to force it into desired <a href=\"https://en.wikipedia.org/wiki/Stochastic_matrix\">stochastic matrix structure</a>. This procedure does reliably improve score. It doesn’t sound ideal, but we couldn’t do any better. </p>\n\n<p>And then the Hungarian algorithm, to get the solution. The name Hungarian algorithm was introduced to us by <a href=\"https://www.kaggle.com/christopherberner\">Christopher Berner</a> in <a href=\"https://www.kaggle.com/christopherberner/hungarian-algorithm-to-optimize-sirna-prediction\">this great kernel</a>. But we solve this matching problem with pulp package with Cplex solver, with a mixed integer programming formulation so short that I can paste it here. Cplex solves each such matrix in a couple of seconds.</p>\n\n<p>```\nprob = LpProblem(\"Recursion\",LpMaximize) <br>\np_vars = LpVariable.dicts(\"match\",(range(L),range(L)),0,1,LpInteger)</p>\n\n<h1>objective</h1>\n\n<p>prob += lpSum(lpSum(p_vars[d][i] * mat[d,i] for i in range(L)) for d in range(L))</p>\n\n<h1>constraints</h1>\n\n<p>for d in range(L):\n    prob += lpSum(p_vars[d][i] for i in range(L)) == 1, \"OneSelected_%i\"%d\nfor i in range(L):\n    prob += lpSum(p_vars[d][i] for d in range(L)) &lt;= 1, \"NoDuplication_%i\"%i\n```</p>\n\n<h1>LB probing (1st place public explained)</h1>\n\n<p>The public set is comprised of experiments HUVEC-17, HEPG2-08, RPE-08, U2OS-04, as can be easily verified and initially <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075\">discussed here</a>.</p>\n\n<p>The LB probing idea stems from the observation that the above MIP formulation does not solve the problem with all the available information. Specifically, we know the scores of our own submissions, so we can decrease the feasible set with two constraints per submission:</p>\n\n<p><code>\nfor s in range(S):\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &amp;gt;= vals[s], \"SubLow_%i\"%s\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &amp;lt;= vals[s] + 0.001, \"SubHigh_%i\"%s\n</code></p>\n\n<p>Note that here we need to solve for 16 sets, 4431 public samples simultaneously, because the constraints use all of them together. The LB probing procedure is simple, we generate a new solution with this formulation, it is a feasible solution in the sense that this solution can be the actual one given all the constraints. After submitting it we get some non-perfect score, which makes the found solution to be infeasible now, and we start from the beginning again.</p>\n\n<p>This is the same approach that I used in LANL earthquake prediction competition, <a href=\"https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe\">this is a post about it</a>. It contains some interesting insights about the approach, and it is the same as in this case, so take a look if you are interested in more details.</p>\n\n<p>Besides the constraints and the feasibility question, there is of course still the objective to maximize the likelihood of the matching, same as in initial MIP formulation. It helps to look for a solution in the right region, as this combinatoric problem is huge. Without good objective function which already solves most of the problem by itself (gives high public LB scores), it is hopeless to advance anywhere with this approach. For good and for bad, the problem here is easy enough for the approach to work. When we got to un-constrained score 0.979 (about 93 errors) the probing got traction and it solved the puzzle in 18 submissions from that. Of course building also on all 200+ submissions that we did before, both probing and regular ones, but getting a high score by itself is when it really clicked. </p>\n\n<p>Two comments on the LB probing\n1. It would not be possible to get anything with the LB probing, if not for the plates leak. It just simplified the problem too much, so it worked. Otherwise it would have been searching in the infinite feasible set forever. Therefore I think that the original design of the competition was correct, - probing not possible.\n2. We corrected 7 HEPG2, 8 RPE, and 45 U2OS predictions thanks to probing and added it all to the training. Note that this is very few. U2OS-04 is one of the hardest, while U2OS-05 is easy, and it is the only U2OS test experiment, so the benefit for it is small. Overall, my estimation is that it gave us private LB boost of 0.001 at the absolute maximum. But it was fun!</p>\n\n<p>Finally, congratulations to the winners and kudos to all kagglers who participated in the kernels and forums discussions!</p>\n\n<p>And I will drop this random but nice image from my notebook here, to remember what it was all about</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2243477%2F44cc7ebcc0b45cbd684aae0ef7fdb019%2Fhuvec05.png?generation=1569543490620681&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 634931,
      "postDate": "2019-09-27T00:53:33.687Z",
      "content": "<p>Thank you Recursion and Kaggle for organizing this unique and challenging competition! When I decided to participate in it, I was impressed by the quality of data and organization, <a href=\"https://www.rxrx.ai/\">this beautiful site</a> alone gave me a lot of motivation. GCP and TPU credits were also indispensable for me, hope to see more in the future. The leaks that were found were handled smoothly by Kaggle, specifically thanks to <a href=\"https://www.kaggle.com/sohier\">Sohier Dane</a>. And credits to <a href=\"https://www.kaggle.com/giuliasavorgnan\">Giulia Savorgnan</a> for reporting the second leak. </p>\n\n<p>I was very fortunate to team up with Yuval on this one. His models scored and maintained top10 for very long time, and I was able to contribute meaningfully only much later in the competition. I still can't wrap my head around how it was possible for him to put so much work in this competition, work full-time at a day job, but also prepare for and run UTMB 171k ultra-marathon race at the end of August. He finished after 50k due to injury, but in my eyes it is already a super-human level of toughness. </p>\n\n<h1>Setup</h1>\n\n<p>All my models I trained with pytorch and TPUs, and used exactly all 600$ of the free GCP credit that I had. My setup I described <a href=\"https://towardsdatascience.com/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4\">in this post</a>. Despite being able to get overall good training speeds, the experience using TPUs with pytorch was rough. The main problem was that it hangs unexpectedly after a few hours of training. At times it was so annoying that I questioned myself if I lived this life correctly. Pytorch/XLA guys tried to help on the forum, but at the moment I think Pytorch/XLA is just not quite production ready (but I will still be happy to receive free TPUs for the next competition!).</p>\n\n<h1>The model</h1>\n\n<p>From the beginning I started with resnet50 and didn’t have an opportunity to successfully try anything else. In hindsight, it was a good choice for this competition. Yuval has used other backbones, I believe he will describe his work in a separate post. </p>\n\n<ul>\n<li>Resnet50 backbone</li>\n<li>ArcFace (m=0.2, s=30, 512 features), used “as is” from the beginning of training, no adjustments. A vanilla implementation <a href=\"https://arxiv.org/abs/1801.07698\">from the original article</a>.</li>\n<li>384x384 images cropped from original 512x512 with random shifts, flips and all angles rotations (albumentations package)</li>\n<li>Normalization of images per experiment and channel, with small randomization</li>\n<li>Not using well, plate and experiment meta info.</li>\n<li>3 folds each containing full experiments</li>\n<li>Training for about 10 epochs together and then a separate model per cell type.</li>\n<li>Each site is treated as a separate sample</li>\n<li>Adam optimizer, decreasing learning rate</li>\n<li>HUVEC-05 is removed, as it is too different from HUVEC test experiments</li>\n<li>HUVEC-18 is moved to the train set</li>\n</ul>\n\n<p>Besides this standard structure described above I had 3 special features in the model. I have not done any serious comparisons with and without them, but my feeling is that the first gives me a major boost, and second and third some additional smaller improvements (but a fair comparison vs regular pseudo-labeling is definitely missing).</p>\n\n<ol>\n<li><strong>Training with test</strong>. The model is trained on all train, test and control images together, predicting 1139 classes (1108 + 31 controls). For the test images I took their softmax output as a target for the log-loss. In this case it can be shown that the gradient of the features before the softmax is zero, that is, these samples have no effect. It makes sense intuitively, I take the predicted probability distribution as a target, - this is already the perfect prediction according to log-loss. But then I modify this target probabilities vector for each sample according to the rich structural information that we hold, - and this is quite unique for this competition. First, - <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905\">the plates leak</a>, all components besides 277 are zeroed out. Second, - each set of 277 samples contains every sirna only once. This second one was not easy to enforce, but as a soft constraint in the loss it worked reasonably well. Additional note here, - validation set data was in the train as well, with hidden sirnas, to simulate similar setting to test. One more thing, - ArcFace didn’t work well here, so it was turned off for test/validation samples. And the last point, - I think this approach in general can be thought of as a dynamic type of pseudo-labelling.</li>\n<li><strong>Mean normalization of ArcFace features</strong>, kudos to Yuval who pioneered this idea. I save EMA of 512 features before ArcFace for each sample, calculate their mean per experiment and overall mean (vectors of length 512 again), and to each sample on the forward pass to these features I add the overall mean and subtract its experiment mean. This way the features entering ArcFace don’t have per-experiment bias and we improve this side of domain adaptation problem. Theoretically.</li>\n<li><strong>Incremental hard pseudo-labelling (PL)</strong>. The PL in point 1 above is a probabilistic PL, in the sense that the target is a probabilities vector of size 277, but not forced to be a specific value, anything is good as long as it is in the set. In this addition to the model at the end of the training I try to push all reasonably confident test and validation samples to lock into one sirna, which becomes its target. I can then zero this sirna for all other samples in the 277-set, which helps them to converge. Once locked, this sample also gets into ArcFace (remember, it is turned off for the test). In practice, judging by validation experiments, first 200 samples out of 277 can be easily locked this way without a single error for all cell types. They were confident anyway, so probably not a huge benefit, but helpful. </li>\n</ol>\n\n<h1>Test-time-augmentation (TTA) and ensemble</h1>\n\n<p>For each sample and each site I run 16 predefined transformations, aggregating those 16 by quantile 75, and then aggregating by gmean between the sites. I then aggregated also by gmean between the folds, and again by gmean between different runs (both me and Yuval had 2 runs). Guess what aggregation function we used to merge the results between me and Yuval’s probabilities? That’s right, gmean again. It just worked very well everywhere, despite me trying 7-8 competitors. We selected aggregation weights for all of this by CV and public LB, different per cell type.</p>\n\n<h1>Mixed integer programming (MIP)</h1>\n\n<p>The output of the above procedure is a set of <code>18*4=72</code> matrices 277 by 277 where each row sums up to 1 (samples) but column sums range from 0.5 to 3 (sirnas). We want columns to sum up to one as well, as we know that each set contains every sirna only once. To achieve this we divided the matrices iteratively by sum of columns and then sum of rows several times (e.g. 10), to force it into desired <a href=\"https://en.wikipedia.org/wiki/Stochastic_matrix\">stochastic matrix structure</a>. This procedure does reliably improve score. It doesn’t sound ideal, but we couldn’t do any better. </p>\n\n<p>And then the Hungarian algorithm, to get the solution. The name Hungarian algorithm was introduced to us by <a href=\"https://www.kaggle.com/christopherberner\">Christopher Berner</a> in <a href=\"https://www.kaggle.com/christopherberner/hungarian-algorithm-to-optimize-sirna-prediction\">this great kernel</a>. But we solve this matching problem with pulp package with Cplex solver, with a mixed integer programming formulation so short that I can paste it here. Cplex solves each such matrix in a couple of seconds.</p>\n\n<p>```\nprob = LpProblem(\"Recursion\",LpMaximize) <br>\np_vars = LpVariable.dicts(\"match\",(range(L),range(L)),0,1,LpInteger)</p>\n\n<h1>objective</h1>\n\n<p>prob += lpSum(lpSum(p_vars[d][i] * mat[d,i] for i in range(L)) for d in range(L))</p>\n\n<h1>constraints</h1>\n\n<p>for d in range(L):\n    prob += lpSum(p_vars[d][i] for i in range(L)) == 1, \"OneSelected_%i\"%d\nfor i in range(L):\n    prob += lpSum(p_vars[d][i] for d in range(L)) &lt;= 1, \"NoDuplication_%i\"%i\n```</p>\n\n<h1>LB probing (1st place public explained)</h1>\n\n<p>The public set is comprised of experiments HUVEC-17, HEPG2-08, RPE-08, U2OS-04, as can be easily verified and initially <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075\">discussed here</a>.</p>\n\n<p>The LB probing idea stems from the observation that the above MIP formulation does not solve the problem with all the available information. Specifically, we know the scores of our own submissions, so we can decrease the feasible set with two constraints per submission:</p>\n\n<p><code>\nfor s in range(S):\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &amp;gt;= vals[s], \"SubLow_%i\"%s\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &amp;lt;= vals[s] + 0.001, \"SubHigh_%i\"%s\n</code></p>\n\n<p>Note that here we need to solve for 16 sets, 4431 public samples simultaneously, because the constraints use all of them together. The LB probing procedure is simple, we generate a new solution with this formulation, it is a feasible solution in the sense that this solution can be the actual one given all the constraints. After submitting it we get some non-perfect score, which makes the found solution to be infeasible now, and we start from the beginning again.</p>\n\n<p>This is the same approach that I used in LANL earthquake prediction competition, <a href=\"https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe\">this is a post about it</a>. It contains some interesting insights about the approach, and it is the same as in this case, so take a look if you are interested in more details.</p>\n\n<p>Besides the constraints and the feasibility question, there is of course still the objective to maximize the likelihood of the matching, same as in initial MIP formulation. It helps to look for a solution in the right region, as this combinatoric problem is huge. Without good objective function which already solves most of the problem by itself (gives high public LB scores), it is hopeless to advance anywhere with this approach. For good and for bad, the problem here is easy enough for the approach to work. When we got to un-constrained score 0.979 (about 93 errors) the probing got traction and it solved the puzzle in 18 submissions from that. Of course building also on all 200+ submissions that we did before, both probing and regular ones, but getting a high score by itself is when it really clicked. </p>\n\n<p>Two comments on the LB probing\n1. It would not be possible to get anything with the LB probing, if not for the plates leak. It just simplified the problem too much, so it worked. Otherwise it would have been searching in the infinite feasible set forever. Therefore I think that the original design of the competition was correct, - probing not possible.\n2. We corrected 7 HEPG2, 8 RPE, and 45 U2OS predictions thanks to probing and added it all to the training. Note that this is very few. U2OS-04 is one of the hardest, while U2OS-05 is easy, and it is the only U2OS test experiment, so the benefit for it is small. Overall, my estimation is that it gave us private LB boost of 0.001 at the absolute maximum. But it was fun!</p>\n\n<p>Finally, congratulations to the winners and kudos to all kagglers who participated in the kernels and forums discussions!</p>\n\n<p>And I will drop this random but nice image from my notebook here, to remember what it was all about</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2243477%2F44cc7ebcc0b45cbd684aae0ef7fdb019%2Fhuvec05.png?generation=1569543490620681&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thank you Recursion and Kaggle for organizing this unique and challenging competition! When I decided to participate in it, I was impressed by the quality of data and organization, [this beautiful site](https://www.rxrx.ai/) alone gave me a lot of motivation. GCP and TPU credits were also indispensable for me, hope to see more in the future. The leaks that were found were handled smoothly by Kaggle, specifically thanks to [Sohier Dane](https://www.kaggle.com/sohier). And credits to [Giulia Savorgnan](https://www.kaggle.com/giuliasavorgnan) for reporting the second leak. \n\nI was very fortunate to team up with Yuval on this one. His models scored and maintained top10 for very long time, and I was able to contribute meaningfully only much later in the competition. I still can't wrap my head around how it was possible for him to put so much work in this competition, work full-time at a day job, but also prepare for and run UTMB 171k ultra-marathon race at the end of August. He finished after 50k due to injury, but in my eyes it is already a super-human level of toughness. \n\n# Setup\n\nAll my models I trained with pytorch and TPUs, and used exactly all 600$ of the free GCP credit that I had. My setup I described [in this post](https://towardsdatascience.com/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4). Despite being able to get overall good training speeds, the experience using TPUs with pytorch was rough. The main problem was that it hangs unexpectedly after a few hours of training. At times it was so annoying that I questioned myself if I lived this life correctly. Pytorch/XLA guys tried to help on the forum, but at the moment I think Pytorch/XLA is just not quite production ready (but I will still be happy to receive free TPUs for the next competition!).\n\n# The model\n\nFrom the beginning I started with resnet50 and didn’t have an opportunity to successfully try anything else. In hindsight, it was a good choice for this competition. Yuval has used other backbones, I believe he will describe his work in a separate post. \n\n- Resnet50 backbone\n- ArcFace (m=0.2, s=30, 512 features), used “as is” from the beginning of training, no adjustments. A vanilla implementation [from the original article](https://arxiv.org/abs/1801.07698).\n- 384x384 images cropped from original 512x512 with random shifts, flips and all angles rotations (albumentations package)\n- Normalization of images per experiment and channel, with small randomization\n- Not using well, plate and experiment meta info.\n- 3 folds each containing full experiments\n- Training for about 10 epochs together and then a separate model per cell type.\n- Each site is treated as a separate sample\n- Adam optimizer, decreasing learning rate\n- HUVEC-05 is removed, as it is too different from HUVEC test experiments\n- HUVEC-18 is moved to the train set\n\nBesides this standard structure described above I had 3 special features in the model. I have not done any serious comparisons with and without them, but my feeling is that the first gives me a major boost, and second and third some additional smaller improvements (but a fair comparison vs regular pseudo-labeling is definitely missing).\n\n1. **Training with test**. The model is trained on all train, test and control images together, predicting 1139 classes (1108 + 31 controls). For the test images I took their softmax output as a target for the log-loss. In this case it can be shown that the gradient of the features before the softmax is zero, that is, these samples have no effect. It makes sense intuitively, I take the predicted probability distribution as a target, - this is already the perfect prediction according to log-loss. But then I modify this target probabilities vector for each sample according to the rich structural information that we hold, - and this is quite unique for this competition. First, - [the plates leak](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905), all components besides 277 are zeroed out. Second, - each set of 277 samples contains every sirna only once. This second one was not easy to enforce, but as a soft constraint in the loss it worked reasonably well. Additional note here, - validation set data was in the train as well, with hidden sirnas, to simulate similar setting to test. One more thing, - ArcFace didn’t work well here, so it was turned off for test/validation samples. And the last point, - I think this approach in general can be thought of as a dynamic type of pseudo-labelling.\n2. **Mean normalization of ArcFace features**, kudos to Yuval who pioneered this idea. I save EMA of 512 features before ArcFace for each sample, calculate their mean per experiment and overall mean (vectors of length 512 again), and to each sample on the forward pass to these features I add the overall mean and subtract its experiment mean. This way the features entering ArcFace don’t have per-experiment bias and we improve this side of domain adaptation problem. Theoretically.\n3. **Incremental hard pseudo-labelling (PL)**. The PL in point 1 above is a probabilistic PL, in the sense that the target is a probabilities vector of size 277, but not forced to be a specific value, anything is good as long as it is in the set. In this addition to the model at the end of the training I try to push all reasonably confident test and validation samples to lock into one sirna, which becomes its target. I can then zero this sirna for all other samples in the 277-set, which helps them to converge. Once locked, this sample also gets into ArcFace (remember, it is turned off for the test). In practice, judging by validation experiments, first 200 samples out of 277 can be easily locked this way without a single error for all cell types. They were confident anyway, so probably not a huge benefit, but helpful. \n\n# Test-time-augmentation (TTA) and ensemble\n\nFor each sample and each site I run 16 predefined transformations, aggregating those 16 by quantile 75, and then aggregating by gmean between the sites. I then aggregated also by gmean between the folds, and again by gmean between different runs (both me and Yuval had 2 runs). Guess what aggregation function we used to merge the results between me and Yuval’s probabilities? That’s right, gmean again. It just worked very well everywhere, despite me trying 7-8 competitors. We selected aggregation weights for all of this by CV and public LB, different per cell type.\n\n# Mixed integer programming (MIP)\n\nThe output of the above procedure is a set of `18*4=72` matrices 277 by 277 where each row sums up to 1 (samples) but column sums range from 0.5 to 3 (sirnas). We want columns to sum up to one as well, as we know that each set contains every sirna only once. To achieve this we divided the matrices iteratively by sum of columns and then sum of rows several times (e.g. 10), to force it into desired [stochastic matrix structure](https://en.wikipedia.org/wiki/Stochastic_matrix). This procedure does reliably improve score. It doesn’t sound ideal, but we couldn’t do any better. \n\nAnd then the Hungarian algorithm, to get the solution. The name Hungarian algorithm was introduced to us by [Christopher Berner](https://www.kaggle.com/christopherberner) in [this great kernel](https://www.kaggle.com/christopherberner/hungarian-algorithm-to-optimize-sirna-prediction). But we solve this matching problem with pulp package with Cplex solver, with a mixed integer programming formulation so short that I can paste it here. Cplex solves each such matrix in a couple of seconds.\n\n```\nprob = LpProblem(\"Recursion\",LpMaximize)    \np_vars = LpVariable.dicts(\"match\",(range(L),range(L)),0,1,LpInteger)\n    \n# objective\nprob += lpSum(lpSum(p_vars[d][i] * mat[d,i] for i in range(L)) for d in range(L))\n    \n# constraints\nfor d in range(L):\n    prob += lpSum(p_vars[d][i] for i in range(L)) == 1, \"OneSelected_%i\"%d\nfor i in range(L):\n    prob += lpSum(p_vars[d][i] for d in range(L)) &lt;= 1, \"NoDuplication_%i\"%i\n```\n\n# LB probing (1st place public explained)\n\nThe public set is comprised of experiments HUVEC-17, HEPG2-08, RPE-08, U2OS-04, as can be easily verified and initially [discussed here](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075).\n\nThe LB probing idea stems from the observation that the above MIP formulation does not solve the problem with all the available information. Specifically, we know the scores of our own submissions, so we can decrease the feasible set with two constraints per submission:\n\n```\nfor s in range(S):\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &gt;= vals[s], \"SubLow_%i\"%s\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &lt;= vals[s] + 0.001, \"SubHigh_%i\"%s\n```\n\nNote that here we need to solve for 16 sets, 4431 public samples simultaneously, because the constraints use all of them together. The LB probing procedure is simple, we generate a new solution with this formulation, it is a feasible solution in the sense that this solution can be the actual one given all the constraints. After submitting it we get some non-perfect score, which makes the found solution to be infeasible now, and we start from the beginning again.\n\nThis is the same approach that I used in LANL earthquake prediction competition, [this is a post about it](https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe). It contains some interesting insights about the approach, and it is the same as in this case, so take a look if you are interested in more details.\n\nBesides the constraints and the feasibility question, there is of course still the objective to maximize the likelihood of the matching, same as in initial MIP formulation. It helps to look for a solution in the right region, as this combinatoric problem is huge. Without good objective function which already solves most of the problem by itself (gives high public LB scores), it is hopeless to advance anywhere with this approach. For good and for bad, the problem here is easy enough for the approach to work. When we got to un-constrained score 0.979 (about 93 errors) the probing got traction and it solved the puzzle in 18 submissions from that. Of course building also on all 200+ submissions that we did before, both probing and regular ones, but getting a high score by itself is when it really clicked. \n\nTwo comments on the LB probing\n1. It would not be possible to get anything with the LB probing, if not for the plates leak. It just simplified the problem too much, so it worked. Otherwise it would have been searching in the infinite feasible set forever. Therefore I think that the original design of the competition was correct, - probing not possible.\n2. We corrected 7 HEPG2, 8 RPE, and 45 U2OS predictions thanks to probing and added it all to the training. Note that this is very few. U2OS-04 is one of the hardest, while U2OS-05 is easy, and it is the only U2OS test experiment, so the benefit for it is small. Overall, my estimation is that it gave us private LB boost of 0.001 at the absolute maximum. But it was fun!\n\nFinally, congratulations to the winners and kudos to all kagglers who participated in the kernels and forums discussions!\n\nAnd I will drop this random but nice image from my notebook here, to remember what it was all about\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2243477%2F44cc7ebcc0b45cbd684aae0ef7fdb019%2Fhuvec05.png?generation=1569543490620681&amp;alt=media)\n",
      "votes": 48
    },
    {
      "id": 635825,
      "postDate": "2019-09-28T09:49:07.867Z",
      "content": "<h3>The Other Part of The Solution</h3>\n\n<p>First I want to thank <a href=\"/zaharch\">@zaharch</a> who is a great teammate, and I learned a lot from him. His extensive data exploration changed this competition, after he discovered the plate leak (and ,of course, we shared it with everybody).</p>\n\n<p>In my part of the solution I used the Densenet architecture, I used 121, 169, 201, 161. <strong>Surprisingly, the 121 (smallest one) gave the best results</strong>. I believe this is because it's last layer was the smallest - 1024, which is still to long, as about half of it ended up as zeros.</p>\n\n<p>The training steps: (for every step, I used the model from the previous step as the base model)\n1. Train on train + test + control, targets - SiRNAs, Cell type, Experiment, Plate.   Loss - weighted cross entropy (every target got a different weight). The test + control test SiRNA was set to -1 and didn't participate in the loss calculation</p>\n\n<ol>\n<li><p>Train on train dataset, with different model for each cell type. Targets - SiRNA only. Loss - Cross Entropy</p></li>\n<li><p>Like 2 but with ArcFace loss - (s=30, m=0.1)</p></li>\n<li><p>Repeat 3 but use test with Weighted Pseudo Labeling - The weight of the image loss was relative to the prediction confidence. images from train set, got weight = 1</p></li>\n</ol>\n\n<p>Inference - <strong>For inference I didn't use the classifier output</strong>, but the output of the previous layer - the features layer.\nI calculated the features for all the train dataset,  and the probability for a specific SiRAN was the average of the correlation between the test image's features and all the images of this SiRNA from the train set (for the same cell type).</p>\n\n<p>The final prediction was done using the MIP procedure <a href=\"/zaharch\">@zaharch</a> described above.</p>\n\n<p>Other post processing: Form every feature vector I substracted the mean of all feature vectors from the same plate</p>\n\n<p>Augmentation and TTA - I used the original 512 size images,  used only 90 deg rotations as augmentation. For TTA I averaged the features from all possible 90 deg rotation and flip for the 2 sites.</p>\n\n<p>I used 4 splits, where I made sure the SiRNA are divided equally.</p>\n\n<p>The final result for the Densnet121 on public LB was 0.954.\nWe averaged the 121 and 169 models with <a href=\"/zaharch\">@zaharch</a> 's model. My models got lower weight (about 0.2).</p>\n\n<p>As final attempt, I retraind the models on <a href=\"/zaharch\">@zaharch</a> probing results - at the end this didn't improve our private score.</p>",
      "rawMarkdown": "### The Other Part of The Solution\n\nFirst I want to thank @zaharch who is a great teammate, and I learned a lot from him. His extensive data exploration changed this competition, after he discovered the plate leak (and ,of course, we shared it with everybody).\n\nIn my part of the solution I used the Densenet architecture, I used 121, 169, 201, 161. **Surprisingly, the 121 (smallest one) gave the best results**. I believe this is because it's last layer was the smallest - 1024, which is still to long, as about half of it ended up as zeros.\n\nThe training steps: (for every step, I used the model from the previous step as the base model)\n1. Train on train + test + control, targets - SiRNAs, Cell type, Experiment, Plate.   Loss - weighted cross entropy (every target got a different weight). The test + control test SiRNA was set to -1 and didn't participate in the loss calculation\n\n2. Train on train dataset, with different model for each cell type. Targets - SiRNA only. Loss - Cross Entropy\n\n3. Like 2 but with ArcFace loss - (s=30, m=0.1)\n\n4. Repeat 3 but use test with Weighted Pseudo Labeling - The weight of the image loss was relative to the prediction confidence. images from train set, got weight = 1\n\nInference - **For inference I didn't use the classifier output**, but the output of the previous layer - the features layer.\nI calculated the features for all the train dataset,  and the probability for a specific SiRAN was the average of the correlation between the test image's features and all the images of this SiRNA from the train set (for the same cell type).\n\nThe final prediction was done using the MIP procedure @zaharch described above.\n\nOther post processing: Form every feature vector I substracted the mean of all feature vectors from the same plate\n\nAugmentation and TTA - I used the original 512 size images,  used only 90 deg rotations as augmentation. For TTA I averaged the features from all possible 90 deg rotation and flip for the 2 sites.\n\nI used 4 splits, where I made sure the SiRNA are divided equally.\n\nThe final result for the Densnet121 on public LB was 0.954.\nWe averaged the 121 and 169 models with @zaharch 's model. My models got lower weight (about 0.2).\n\nAs final attempt, I retraind the models on @zaharch probing results - at the end this didn't improve our private score.",
      "votes": 5
    },
    {
      "id": 636228,
      "postDate": "2019-09-29T04:06:58.460Z",
      "content": "<p>Nice overview. Congrats on placing!</p>",
      "rawMarkdown": "Nice overview. Congrats on placing!",
      "votes": 1
    },
    {
      "id": 635140,
      "postDate": "2019-09-27T07:37:23.703Z",
      "content": "<p>Congrats <a href=\"/zaharch\">@zaharch</a> and thanks for sharing!</p>",
      "rawMarkdown": "Congrats @zaharch and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 634937,
      "postDate": "2019-09-27T01:09:11.767Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 1
    },
    {
      "id": 635013,
      "postDate": "2019-09-27T04:10:52.510Z",
      "content": "<p>I'm so awed with how close you look into the details of data!</p>\n\n<p>Thanks for sharing</p>",
      "rawMarkdown": "I'm so awed with how close you look into the details of data!\n\nThanks for sharing",
      "votes": 2
    },
    {
      "id": 634934,
      "postDate": "2019-09-27T01:07:06.890Z",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> - Thanks for sharing the approach. It would be nice to share the GitHub link/notebook for the solution. Integer programming for solving a classification problem is actually an excellent approach which I had not heard off till now. Just curious to know how much was difference between before and after the integer programming. Would also be interested to look at your tpu code for the same..</p>",
      "rawMarkdown": "@zaharch - Thanks for sharing the approach. It would be nice to share the GitHub link/notebook for the solution. Integer programming for solving a classification problem is actually an excellent approach which I had not heard off till now. Just curious to know how much was difference between before and after the integer programming. Would also be interested to look at your tpu code for the same..",
      "votes": 2,
      "replies": [
        {
          "id": 634953,
          "postDate": "2019-09-27T02:01:47.930Z",
          "content": "<p>Thanks, I think I will share my TPU model notebook later today. To the question on before/after integer programming, the original MIP (not LB probing) formulation is just a way to solve the assignment problem, one of many ways that give the same solution. And it does improve the score reliably above argmax() and another more involved approach that we tried. In the kernel of Christopher Berner that I mentioned he also gives the score how much it helped him.</p>",
          "rawMarkdown": "Thanks, I think I will share my TPU model notebook later today. To the question on before/after integer programming, the original MIP (not LB probing) formulation is just a way to solve the assignment problem, one of many ways that give the same solution. And it does improve the score reliably above argmax() and another more involved approach that we tried. In the kernel of Christopher Berner that I mentioned he also gives the score how much it helped him."
        },
        {
          "id": 635212,
          "postDate": "2019-09-27T08:58:55.953Z",
          "content": "<p>Wow. Simply amazing. Looking forward to your code. <a href=\"/zaharch\">@zaharch</a> . Congrats once again.  </p>",
          "rawMarkdown": "Wow. Simply amazing. Looking forward to your code. @zaharch . Congrats once again.  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 635030,
      "postDate": "2019-09-27T04:54:29.953Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 635825,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2019-09-28T09:49:07.867000",
      "content": "<h3>The Other Part of The Solution</h3>\n\n<p>First I want to thank <a href=\"/zaharch\">@zaharch</a> who is a great teammate, and I learned a lot from him. His extensive data exploration changed this competition, after he discovered the plate leak (and ,of course, we shared it with everybody).</p>\n\n<p>In my part of the solution I used the Densenet architecture, I used 121, 169, 201, 161. <strong>Surprisingly, the 121 (smallest one) gave the best results</strong>. I believe this is because it's last layer was the smallest - 1024, which is still to long, as about half of it ended up as zeros.</p>\n\n<p>The training steps: (for every step, I used the model from the previous step as the base model)\n1. Train on train + test + control, targets - SiRNAs, Cell type, Experiment, Plate.   Loss - weighted cross entropy (every target got a different weight). The test + control test SiRNA was set to -1 and didn't participate in the loss calculation</p>\n\n<ol>\n<li><p>Train on train dataset, with different model for each cell type. Targets - SiRNA only. Loss - Cross Entropy</p></li>\n<li><p>Like 2 but with ArcFace loss - (s=30, m=0.1)</p></li>\n<li><p>Repeat 3 but use test with Weighted Pseudo Labeling - The weight of the image loss was relative to the prediction confidence. images from train set, got weight = 1</p></li>\n</ol>\n\n<p>Inference - <strong>For inference I didn't use the classifier output</strong>, but the output of the previous layer - the features layer.\nI calculated the features for all the train dataset,  and the probability for a specific SiRAN was the average of the correlation between the test image's features and all the images of this SiRNA from the train set (for the same cell type).</p>\n\n<p>The final prediction was done using the MIP procedure <a href=\"/zaharch\">@zaharch</a> described above.</p>\n\n<p>Other post processing: Form every feature vector I substracted the mean of all feature vectors from the same plate</p>\n\n<p>Augmentation and TTA - I used the original 512 size images,  used only 90 deg rotations as augmentation. For TTA I averaged the features from all possible 90 deg rotation and flip for the 2 sites.</p>\n\n<p>I used 4 splits, where I made sure the SiRNA are divided equally.</p>\n\n<p>The final result for the Densnet121 on public LB was 0.954.\nWe averaged the 121 and 169 models with <a href=\"/zaharch\">@zaharch</a> 's model. My models got lower weight (about 0.2).</p>\n\n<p>As final attempt, I retraind the models on <a href=\"/zaharch\">@zaharch</a> probing results - at the end this didn't improve our private score.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 636228,
      "author_name": "Psiczyk",
      "author_url": "",
      "post_date": "2019-09-29T04:06:58.460000",
      "content": "<p>Nice overview. Congrats on placing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 635140,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2019-09-27T07:37:23.703000",
      "content": "<p>Congrats <a href=\"/zaharch\">@zaharch</a> and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 634937,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2019-09-27T01:09:11.767000",
      "content": "<p>Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 635013,
      "author_name": "XC Zhang (Raynard)",
      "author_url": "",
      "post_date": "2019-09-27T04:10:52.510000",
      "content": "<p>I'm so awed with how close you look into the details of data!</p>\n\n<p>Thanks for sharing</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 634934,
      "author_name": "Manoj Prabhakar",
      "author_url": "",
      "post_date": "2019-09-27T01:07:06.890000",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> - Thanks for sharing the approach. It would be nice to share the GitHub link/notebook for the solution. Integer programming for solving a classification problem is actually an excellent approach which I had not heard off till now. Just curious to know how much was difference between before and after the integer programming. Would also be interested to look at your tpu code for the same..</p>",
      "votes": 2,
      "replies": [
        {
          "id": 634953,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-09-27T02:01:47.930000",
          "content": "<p>Thanks, I think I will share my TPU model notebook later today. To the question on before/after integer programming, the original MIP (not LB probing) formulation is just a way to solve the assignment problem, one of many ways that give the same solution. And it does improve the score reliably above argmax() and another more involved approach that we tried. In the kernel of Christopher Berner that I mentioned he also gives the score how much it helped him.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 635212,
          "author_name": "Manoj Prabhakar",
          "author_url": "",
          "post_date": "2019-09-27T08:58:55.953000",
          "content": "<p>Wow. Simply amazing. Looking forward to your code. <a href=\"/zaharch\">@zaharch</a> . Congrats once again.  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 635030,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T04:54:29.953000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "634931": "Thank you Recursion and Kaggle for organizing this unique and challenging competition! When I decided to participate in it, I was impressed by the quality of data and organization, [this beautiful site](https://www.rxrx.ai/) alone gave me a lot of motivation. GCP and TPU credits were also indispensable for me, hope to see more in the future. The leaks that were found were handled smoothly by Kaggle, specifically thanks to [Sohier Dane](https://www.kaggle.com/sohier). And credits to [Giulia Savorgnan](https://www.kaggle.com/giuliasavorgnan) for reporting the second leak. \n\nI was very fortunate to team up with Yuval on this one. His models scored and maintained top10 for very long time, and I was able to contribute meaningfully only much later in the competition. I still can't wrap my head around how it was possible for him to put so much work in this competition, work full-time at a day job, but also prepare for and run UTMB 171k ultra-marathon race at the end of August. He finished after 50k due to injury, but in my eyes it is already a super-human level of toughness. \n\n# Setup\n\nAll my models I trained with pytorch and TPUs, and used exactly all 600$ of the free GCP credit that I had. My setup I described [in this post](https://towardsdatascience.com/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4). Despite being able to get overall good training speeds, the experience using TPUs with pytorch was rough. The main problem was that it hangs unexpectedly after a few hours of training. At times it was so annoying that I questioned myself if I lived this life correctly. Pytorch/XLA guys tried to help on the forum, but at the moment I think Pytorch/XLA is just not quite production ready (but I will still be happy to receive free TPUs for the next competition!).\n\n# The model\n\nFrom the beginning I started with resnet50 and didn’t have an opportunity to successfully try anything else. In hindsight, it was a good choice for this competition. Yuval has used other backbones, I believe he will describe his work in a separate post. \n\n- Resnet50 backbone\n- ArcFace (m=0.2, s=30, 512 features), used “as is” from the beginning of training, no adjustments. A vanilla implementation [from the original article](https://arxiv.org/abs/1801.07698).\n- 384x384 images cropped from original 512x512 with random shifts, flips and all angles rotations (albumentations package)\n- Normalization of images per experiment and channel, with small randomization\n- Not using well, plate and experiment meta info.\n- 3 folds each containing full experiments\n- Training for about 10 epochs together and then a separate model per cell type.\n- Each site is treated as a separate sample\n- Adam optimizer, decreasing learning rate\n- HUVEC-05 is removed, as it is too different from HUVEC test experiments\n- HUVEC-18 is moved to the train set\n\nBesides this standard structure described above I had 3 special features in the model. I have not done any serious comparisons with and without them, but my feeling is that the first gives me a major boost, and second and third some additional smaller improvements (but a fair comparison vs regular pseudo-labeling is definitely missing).\n\n1. **Training with test**. The model is trained on all train, test and control images together, predicting 1139 classes (1108 + 31 controls). For the test images I took their softmax output as a target for the log-loss. In this case it can be shown that the gradient of the features before the softmax is zero, that is, these samples have no effect. It makes sense intuitively, I take the predicted probability distribution as a target, - this is already the perfect prediction according to log-loss. But then I modify this target probabilities vector for each sample according to the rich structural information that we hold, - and this is quite unique for this competition. First, - [the plates leak](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905), all components besides 277 are zeroed out. Second, - each set of 277 samples contains every sirna only once. This second one was not easy to enforce, but as a soft constraint in the loss it worked reasonably well. Additional note here, - validation set data was in the train as well, with hidden sirnas, to simulate similar setting to test. One more thing, - ArcFace didn’t work well here, so it was turned off for test/validation samples. And the last point, - I think this approach in general can be thought of as a dynamic type of pseudo-labelling.\n2. **Mean normalization of ArcFace features**, kudos to Yuval who pioneered this idea. I save EMA of 512 features before ArcFace for each sample, calculate their mean per experiment and overall mean (vectors of length 512 again), and to each sample on the forward pass to these features I add the overall mean and subtract its experiment mean. This way the features entering ArcFace don’t have per-experiment bias and we improve this side of domain adaptation problem. Theoretically.\n3. **Incremental hard pseudo-labelling (PL)**. The PL in point 1 above is a probabilistic PL, in the sense that the target is a probabilities vector of size 277, but not forced to be a specific value, anything is good as long as it is in the set. In this addition to the model at the end of the training I try to push all reasonably confident test and validation samples to lock into one sirna, which becomes its target. I can then zero this sirna for all other samples in the 277-set, which helps them to converge. Once locked, this sample also gets into ArcFace (remember, it is turned off for the test). In practice, judging by validation experiments, first 200 samples out of 277 can be easily locked this way without a single error for all cell types. They were confident anyway, so probably not a huge benefit, but helpful. \n\n# Test-time-augmentation (TTA) and ensemble\n\nFor each sample and each site I run 16 predefined transformations, aggregating those 16 by quantile 75, and then aggregating by gmean between the sites. I then aggregated also by gmean between the folds, and again by gmean between different runs (both me and Yuval had 2 runs). Guess what aggregation function we used to merge the results between me and Yuval’s probabilities? That’s right, gmean again. It just worked very well everywhere, despite me trying 7-8 competitors. We selected aggregation weights for all of this by CV and public LB, different per cell type.\n\n# Mixed integer programming (MIP)\n\nThe output of the above procedure is a set of `18*4=72` matrices 277 by 277 where each row sums up to 1 (samples) but column sums range from 0.5 to 3 (sirnas). We want columns to sum up to one as well, as we know that each set contains every sirna only once. To achieve this we divided the matrices iteratively by sum of columns and then sum of rows several times (e.g. 10), to force it into desired [stochastic matrix structure](https://en.wikipedia.org/wiki/Stochastic_matrix). This procedure does reliably improve score. It doesn’t sound ideal, but we couldn’t do any better. \n\nAnd then the Hungarian algorithm, to get the solution. The name Hungarian algorithm was introduced to us by [Christopher Berner](https://www.kaggle.com/christopherberner) in [this great kernel](https://www.kaggle.com/christopherberner/hungarian-algorithm-to-optimize-sirna-prediction). But we solve this matching problem with pulp package with Cplex solver, with a mixed integer programming formulation so short that I can paste it here. Cplex solves each such matrix in a couple of seconds.\n\n```\nprob = LpProblem(\"Recursion\",LpMaximize)    \np_vars = LpVariable.dicts(\"match\",(range(L),range(L)),0,1,LpInteger)\n    \n# objective\nprob += lpSum(lpSum(p_vars[d][i] * mat[d,i] for i in range(L)) for d in range(L))\n    \n# constraints\nfor d in range(L):\n    prob += lpSum(p_vars[d][i] for i in range(L)) == 1, \"OneSelected_%i\"%d\nfor i in range(L):\n    prob += lpSum(p_vars[d][i] for d in range(L)) &lt;= 1, \"NoDuplication_%i\"%i\n```\n\n# LB probing (1st place public explained)\n\nThe public set is comprised of experiments HUVEC-17, HEPG2-08, RPE-08, U2OS-04, as can be easily verified and initially [discussed here](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075).\n\nThe LB probing idea stems from the observation that the above MIP formulation does not solve the problem with all the available information. Specifically, we know the scores of our own submissions, so we can decrease the feasible set with two constraints per submission:\n\n```\nfor s in range(S):\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &gt;= vals[s], \"SubLow_%i\"%s\n    prob += lpSum(lpSum(p_vars[d][i] for i in range(L) if subs_sel[s,d,i])\n        for d in range(D)) / D &lt;= vals[s] + 0.001, \"SubHigh_%i\"%s\n```\n\nNote that here we need to solve for 16 sets, 4431 public samples simultaneously, because the constraints use all of them together. The LB probing procedure is simple, we generate a new solution with this formulation, it is a feasible solution in the sense that this solution can be the actual one given all the constraints. After submitting it we get some non-perfect score, which makes the found solution to be infeasible now, and we start from the beginning again.\n\nThis is the same approach that I used in LANL earthquake prediction competition, [this is a post about it](https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe). It contains some interesting insights about the approach, and it is the same as in this case, so take a look if you are interested in more details.\n\nBesides the constraints and the feasibility question, there is of course still the objective to maximize the likelihood of the matching, same as in initial MIP formulation. It helps to look for a solution in the right region, as this combinatoric problem is huge. Without good objective function which already solves most of the problem by itself (gives high public LB scores), it is hopeless to advance anywhere with this approach. For good and for bad, the problem here is easy enough for the approach to work. When we got to un-constrained score 0.979 (about 93 errors) the probing got traction and it solved the puzzle in 18 submissions from that. Of course building also on all 200+ submissions that we did before, both probing and regular ones, but getting a high score by itself is when it really clicked. \n\nTwo comments on the LB probing\n1. It would not be possible to get anything with the LB probing, if not for the plates leak. It just simplified the problem too much, so it worked. Otherwise it would have been searching in the infinite feasible set forever. Therefore I think that the original design of the competition was correct, - probing not possible.\n2. We corrected 7 HEPG2, 8 RPE, and 45 U2OS predictions thanks to probing and added it all to the training. Note that this is very few. U2OS-04 is one of the hardest, while U2OS-05 is easy, and it is the only U2OS test experiment, so the benefit for it is small. Overall, my estimation is that it gave us private LB boost of 0.001 at the absolute maximum. But it was fun!\n\nFinally, congratulations to the winners and kudos to all kagglers who participated in the kernels and forums discussions!\n\nAnd I will drop this random but nice image from my notebook here, to remember what it was all about\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2243477%2F44cc7ebcc0b45cbd684aae0ef7fdb019%2Fhuvec05.png?generation=1569543490620681&amp;alt=media)\n",
    "635825": "### The Other Part of The Solution\n\nFirst I want to thank @zaharch who is a great teammate, and I learned a lot from him. His extensive data exploration changed this competition, after he discovered the plate leak (and ,of course, we shared it with everybody).\n\nIn my part of the solution I used the Densenet architecture, I used 121, 169, 201, 161. **Surprisingly, the 121 (smallest one) gave the best results**. I believe this is because it's last layer was the smallest - 1024, which is still to long, as about half of it ended up as zeros.\n\nThe training steps: (for every step, I used the model from the previous step as the base model)\n1. Train on train + test + control, targets - SiRNAs, Cell type, Experiment, Plate.   Loss - weighted cross entropy (every target got a different weight). The test + control test SiRNA was set to -1 and didn't participate in the loss calculation\n\n2. Train on train dataset, with different model for each cell type. Targets - SiRNA only. Loss - Cross Entropy\n\n3. Like 2 but with ArcFace loss - (s=30, m=0.1)\n\n4. Repeat 3 but use test with Weighted Pseudo Labeling - The weight of the image loss was relative to the prediction confidence. images from train set, got weight = 1\n\nInference - **For inference I didn't use the classifier output**, but the output of the previous layer - the features layer.\nI calculated the features for all the train dataset,  and the probability for a specific SiRAN was the average of the correlation between the test image's features and all the images of this SiRNA from the train set (for the same cell type).\n\nThe final prediction was done using the MIP procedure @zaharch described above.\n\nOther post processing: Form every feature vector I substracted the mean of all feature vectors from the same plate\n\nAugmentation and TTA - I used the original 512 size images,  used only 90 deg rotations as augmentation. For TTA I averaged the features from all possible 90 deg rotation and flip for the 2 sites.\n\nI used 4 splits, where I made sure the SiRNA are divided equally.\n\nThe final result for the Densnet121 on public LB was 0.954.\nWe averaged the 121 and 169 models with @zaharch 's model. My models got lower weight (about 0.2).\n\nAs final attempt, I retraind the models on @zaharch probing results - at the end this didn't improve our private score.",
    "636228": "Nice overview. Congrats on placing!",
    "635140": "Congrats @zaharch and thanks for sharing!",
    "634937": "Congrats!",
    "635013": "I'm so awed with how close you look into the details of data!\n\nThanks for sharing",
    "634934": "@zaharch - Thanks for sharing the approach. It would be nice to share the GitHub link/notebook for the solution. Integer programming for solving a classification problem is actually an excellent approach which I had not heard off till now. Just curious to know how much was difference between before and after the integer programming. Would also be interested to look at your tpu code for the same..",
    "635030": ""
  }
}