{
  "id": 110457,
  "title": "Second Place Solution",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110457",
  "author_name": "Junonia",
  "post_date": "2019-09-28T02:38:19.368000",
  "votes": 44,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to Recursion and Kaggle for sponsoring this very interesting competition and congratulations to all that went through the journey.  We learned a lot and enjoyed the entire process.  Looking forward to learning from other teams’ solutions.</p>\n\n<h3>Input and preprocessing:</h3>\n\n<ul>\n<li>6 channel input, </li>\n<li>per image standardization (minus the mean and divide by the standard deviation), </li>\n<li>random crop 384x384,</li>\n<li>random flip, </li>\n<li>random rotation multiple of 90 degrees.  </li>\n</ul>\n\n<h3>Modeling:</h3>\n\n<p>We modified ResNet to limit the receptive field size of the output, as we suspect it is the individual cells and their immediate neighbors that contain the most discriminating information.  Here are the list of things we modified from the vanilla ResNet:\n- Fewer blocks. ResNet typically has 4 chunks of blocks, some of our models only has 2 chunks.\n- More 1x1 conv blocks \n- Average pooling from lower blocks, concatenated with average pooling from higher blocks\n- Remove the immediate max-pooling after the first convolution\n- Replace the 1x1 convolution with 3x3 convolution in the shortcut layer. This increased the smoothness of the test accuracy during training, but only increased the final testing accuracy slightly.  </p>\n\n<p>We used all the negative and positive controls (including those in the test plates) as part of the training set.  The output is 1139x4 logits.</p>\n\n<h3>Loss:</h3>\n\n<p>We used the [ArcFace loss posted by bestfitting] (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a>).  We made the <code>gamma</code> parameter adjustable.  We were struggling to get ArcFace to converge initially until we tuned the gamma parameter.  Our settings are <code>gamma = 0.2, m = 0.4, s = 32</code></p>\n\n<h3>Optimizer:</h3>\n\n<p>Adam optimizer with the following schedules:\n<code>python\n0-20: linear warm up to 3e-3\n20-60: 3e-3\n60-80: 9e-4\n80-90: 3e-4\n90-100: 3e-5\n100-120: 3e-6\n</code></p>\n\n<h3>Post-processing:</h3>\n\n<ul>\n<li>Center the embeddings by plate.</li>\n<li>Average the embeddings from both sites to obtain per-well embedding.</li>\n<li>For each cell line, obtain train center embeddings by averaging together the siRNA embeddings.</li>\n<li>Compute cosine-similarity of each well’s embedding to train center embeddings. </li>\n<li>Use LSA to compute label assignment based on the 277 leak.</li>\n</ul>\n\n<h3>Pseudo-labeling:</h3>\n\n<p>An ensemble of 5 models achieved a public LB score of 0.993 and a private LB score of 0.9957 without pseudo-labeling (single model 0.990 and 0.9947).  We then collected all our public LB 0.990+ predictions and identified 327 examples that were not consistent.  All the test predictions not in this set of 327 were then used as pseudo-labels.  An ensemble of models trained on this pseudo-label set achieved a public LB of 0.997, and a private LB of 0.9967.  We also tried iteratively adding more pseudo-labels to the training set (500, 700, 900 per experiment), but it did not improve our public LB score.</p>",
  "messages": [
    {
      "id": 635673,
      "postDate": "2019-09-28T02:38:19.370Z",
      "content": "<p>Thanks to Recursion and Kaggle for sponsoring this very interesting competition and congratulations to all that went through the journey.  We learned a lot and enjoyed the entire process.  Looking forward to learning from other teams’ solutions.</p>\n\n<h3>Input and preprocessing:</h3>\n\n<ul>\n<li>6 channel input, </li>\n<li>per image standardization (minus the mean and divide by the standard deviation), </li>\n<li>random crop 384x384,</li>\n<li>random flip, </li>\n<li>random rotation multiple of 90 degrees.  </li>\n</ul>\n\n<h3>Modeling:</h3>\n\n<p>We modified ResNet to limit the receptive field size of the output, as we suspect it is the individual cells and their immediate neighbors that contain the most discriminating information.  Here are the list of things we modified from the vanilla ResNet:\n- Fewer blocks. ResNet typically has 4 chunks of blocks, some of our models only has 2 chunks.\n- More 1x1 conv blocks \n- Average pooling from lower blocks, concatenated with average pooling from higher blocks\n- Remove the immediate max-pooling after the first convolution\n- Replace the 1x1 convolution with 3x3 convolution in the shortcut layer. This increased the smoothness of the test accuracy during training, but only increased the final testing accuracy slightly.  </p>\n\n<p>We used all the negative and positive controls (including those in the test plates) as part of the training set.  The output is 1139x4 logits.</p>\n\n<h3>Loss:</h3>\n\n<p>We used the [ArcFace loss posted by bestfitting] (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a>).  We made the <code>gamma</code> parameter adjustable.  We were struggling to get ArcFace to converge initially until we tuned the gamma parameter.  Our settings are <code>gamma = 0.2, m = 0.4, s = 32</code></p>\n\n<h3>Optimizer:</h3>\n\n<p>Adam optimizer with the following schedules:\n<code>python\n0-20: linear warm up to 3e-3\n20-60: 3e-3\n60-80: 9e-4\n80-90: 3e-4\n90-100: 3e-5\n100-120: 3e-6\n</code></p>\n\n<h3>Post-processing:</h3>\n\n<ul>\n<li>Center the embeddings by plate.</li>\n<li>Average the embeddings from both sites to obtain per-well embedding.</li>\n<li>For each cell line, obtain train center embeddings by averaging together the siRNA embeddings.</li>\n<li>Compute cosine-similarity of each well’s embedding to train center embeddings. </li>\n<li>Use LSA to compute label assignment based on the 277 leak.</li>\n</ul>\n\n<h3>Pseudo-labeling:</h3>\n\n<p>An ensemble of 5 models achieved a public LB score of 0.993 and a private LB score of 0.9957 without pseudo-labeling (single model 0.990 and 0.9947).  We then collected all our public LB 0.990+ predictions and identified 327 examples that were not consistent.  All the test predictions not in this set of 327 were then used as pseudo-labels.  An ensemble of models trained on this pseudo-label set achieved a public LB of 0.997, and a private LB of 0.9967.  We also tried iteratively adding more pseudo-labels to the training set (500, 700, 900 per experiment), but it did not improve our public LB score.</p>",
      "rawMarkdown": "Thanks to Recursion and Kaggle for sponsoring this very interesting competition and congratulations to all that went through the journey.  We learned a lot and enjoyed the entire process.  Looking forward to learning from other teams’ solutions.\n\n### Input and preprocessing:\n- 6 channel input, \n- per image standardization (minus the mean and divide by the standard deviation), \n- random crop 384x384,\n- random flip, \n- random rotation multiple of 90 degrees.  \n\n### Modeling:\nWe modified ResNet to limit the receptive field size of the output, as we suspect it is the individual cells and their immediate neighbors that contain the most discriminating information.  Here are the list of things we modified from the vanilla ResNet:\n- Fewer blocks. ResNet typically has 4 chunks of blocks, some of our models only has 2 chunks.\n- More 1x1 conv blocks \n- Average pooling from lower blocks, concatenated with average pooling from higher blocks\n- Remove the immediate max-pooling after the first convolution\n- Replace the 1x1 convolution with 3x3 convolution in the shortcut layer. This increased the smoothness of the test accuracy during training, but only increased the final testing accuracy slightly.  \n\nWe used all the negative and positive controls (including those in the test plates) as part of the training set.  The output is 1139x4 logits.\n\n### Loss:\nWe used the [ArcFace loss posted by bestfitting] (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109).  We made the `gamma` parameter adjustable.  We were struggling to get ArcFace to converge initially until we tuned the gamma parameter.  Our settings are `gamma = 0.2, m = 0.4, s = 32`\n\n### Optimizer:\nAdam optimizer with the following schedules:\n``` python\n0-20: linear warm up to 3e-3\n20-60: 3e-3\n60-80: 9e-4\n80-90: 3e-4\n90-100: 3e-5\n100-120: 3e-6\n```\n\n### Post-processing:\n- Center the embeddings by plate.\n- Average the embeddings from both sites to obtain per-well embedding.\n- For each cell line, obtain train center embeddings by averaging together the siRNA embeddings.\n- Compute cosine-similarity of each well’s embedding to train center embeddings. \n- Use LSA to compute label assignment based on the 277 leak.\n\n### Pseudo-labeling:\nAn ensemble of 5 models achieved a public LB score of 0.993 and a private LB score of 0.9957 without pseudo-labeling (single model 0.990 and 0.9947).  We then collected all our public LB 0.990+ predictions and identified 327 examples that were not consistent.  All the test predictions not in this set of 327 were then used as pseudo-labels.  An ensemble of models trained on this pseudo-label set achieved a public LB of 0.997, and a private LB of 0.9967.  We also tried iteratively adding more pseudo-labels to the training set (500, 700, 900 per experiment), but it did not improve our public LB score.\n",
      "votes": 44
    },
    {
      "id": 635739,
      "postDate": "2019-09-28T05:58:31.867Z",
      "content": "<p>Thanks for sharing! Congrats to your achievements and I think your public LB performance motivated many teams during the whole competition period.\nQuestion: </p>\n\n<blockquote>\n  <p>An ensemble of 5 models achieved a public LB score</p>\n</blockquote>\n\n<p>Here 5 models means 5 folds or 5 different backbones?</p>",
      "rawMarkdown": "Thanks for sharing! Congrats to your achievements and I think your public LB performance motivated many teams during the whole competition period.\nQuestion: \n&gt; An ensemble of 5 models achieved a public LB score\n\nHere 5 models means 5 folds or 5 different backbones?",
      "votes": 2,
      "replies": [
        {
          "id": 635828,
          "postDate": "2019-09-28T10:03:25.610Z",
          "content": "<p>It was 5 different backbones trained on all the data.  </p>\n\n<p>This is Kaggle, we all inspire each other to get better.  To be honest, if we were to work on the problem by ourselves, we probably will stop at 0.6 and be satisfied.  Team OrdenaLenina and others motivated us to move higher early on. </p>",
          "rawMarkdown": "It was 5 different backbones trained on all the data.  \n\nThis is Kaggle, we all inspire each other to get better.  To be honest, if we were to work on the problem by ourselves, we probably will stop at 0.6 and be satisfied.  Team OrdenaLenina and others motivated us to move higher early on. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 656159,
      "postDate": "2019-10-24T00:55:04.160Z",
      "content": "<p>Thank you for sharing. Your Model is very interesting.\nWhat is the channel, height and width of feature map before last pooling (e.g. GAP)?</p>",
      "rawMarkdown": "Thank you for sharing. Your Model is very interesting.\nWhat is the channel, height and width of feature map before last pooling (e.g. GAP)?\n"
    },
    {
      "id": 635820,
      "postDate": "2019-09-28T09:25:23.943Z",
      "content": "<p>Congratulations! do you have any plan to open source your code solutions?</p>",
      "rawMarkdown": "Congratulations! do you have any plan to open source your code solutions?",
      "replies": [
        {
          "id": 636160,
          "postDate": "2019-09-28T22:21:20.813Z",
          "content": "<p>Yes, we do plan to open source the code. </p>",
          "rawMarkdown": "Yes, we do plan to open source the code. ",
          "votes": 1
        },
        {
          "id": 664016,
          "postDate": "2019-11-03T03:24:11.917Z",
          "content": "<p>Any update on open-sourcing of code solutions? :-)</p>",
          "rawMarkdown": "Any update on open-sourcing of code solutions? :-)"
        }
      ]
    },
    {
      "id": 635786,
      "postDate": "2019-09-28T07:59:42.723Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 643003,
      "postDate": "2019-10-06T23:15:46.920Z",
      "content": "<p>Congrats, thank you for sharing.</p>",
      "rawMarkdown": "Congrats, thank you for sharing."
    },
    {
      "id": 636767,
      "postDate": "2019-09-30T07:08:46.233Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing\n"
    },
    {
      "id": 636675,
      "postDate": "2019-09-30T02:55:40.713Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 635768,
      "postDate": "2019-09-28T06:57:17.923Z",
      "content": "<p>Thanks for sharing! \nGreat solution</p>",
      "rawMarkdown": "Thanks for sharing! \nGreat solution"
    }
  ],
  "comments": [
    {
      "id": 635739,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2019-09-28T05:58:31.867000",
      "content": "<p>Thanks for sharing! Congrats to your achievements and I think your public LB performance motivated many teams during the whole competition period.\nQuestion: </p>\n\n<blockquote>\n  <p>An ensemble of 5 models achieved a public LB score</p>\n</blockquote>\n\n<p>Here 5 models means 5 folds or 5 different backbones?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 635828,
          "author_name": "Junonia",
          "author_url": "",
          "post_date": "2019-09-28T10:03:25.610000",
          "content": "<p>It was 5 different backbones trained on all the data.  </p>\n\n<p>This is Kaggle, we all inspire each other to get better.  To be honest, if we were to work on the problem by ourselves, we probably will stop at 0.6 and be satisfied.  Team OrdenaLenina and others motivated us to move higher early on. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 656159,
      "author_name": "Shudai",
      "author_url": "",
      "post_date": "2019-10-24T00:55:04.160000",
      "content": "<p>Thank you for sharing. Your Model is very interesting.\nWhat is the channel, height and width of feature map before last pooling (e.g. GAP)?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635820,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-09-28T09:25:23.943000",
      "content": "<p>Congratulations! do you have any plan to open source your code solutions?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 636160,
          "author_name": "Junonia",
          "author_url": "",
          "post_date": "2019-09-28T22:21:20.813000",
          "content": "<p>Yes, we do plan to open source the code. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 664016,
          "author_name": "FGPC",
          "author_url": "",
          "post_date": "2019-11-03T03:24:11.917000",
          "content": "<p>Any update on open-sourcing of code solutions? :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 635786,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-28T07:59:42.723000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 643003,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-10-06T23:15:46.920000",
      "content": "<p>Congrats, thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636767,
      "author_name": "sai kumar",
      "author_url": "",
      "post_date": "2019-09-30T07:08:46.233000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636675,
      "author_name": "SAS",
      "author_url": "",
      "post_date": "2019-09-30T02:55:40.713000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635768,
      "author_name": "Igor Krashenyi",
      "author_url": "",
      "post_date": "2019-09-28T06:57:17.923000",
      "content": "<p>Thanks for sharing! \nGreat solution</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635673": "Thanks to Recursion and Kaggle for sponsoring this very interesting competition and congratulations to all that went through the journey.  We learned a lot and enjoyed the entire process.  Looking forward to learning from other teams’ solutions.\n\n### Input and preprocessing:\n- 6 channel input, \n- per image standardization (minus the mean and divide by the standard deviation), \n- random crop 384x384,\n- random flip, \n- random rotation multiple of 90 degrees.  \n\n### Modeling:\nWe modified ResNet to limit the receptive field size of the output, as we suspect it is the individual cells and their immediate neighbors that contain the most discriminating information.  Here are the list of things we modified from the vanilla ResNet:\n- Fewer blocks. ResNet typically has 4 chunks of blocks, some of our models only has 2 chunks.\n- More 1x1 conv blocks \n- Average pooling from lower blocks, concatenated with average pooling from higher blocks\n- Remove the immediate max-pooling after the first convolution\n- Replace the 1x1 convolution with 3x3 convolution in the shortcut layer. This increased the smoothness of the test accuracy during training, but only increased the final testing accuracy slightly.  \n\nWe used all the negative and positive controls (including those in the test plates) as part of the training set.  The output is 1139x4 logits.\n\n### Loss:\nWe used the [ArcFace loss posted by bestfitting] (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109).  We made the `gamma` parameter adjustable.  We were struggling to get ArcFace to converge initially until we tuned the gamma parameter.  Our settings are `gamma = 0.2, m = 0.4, s = 32`\n\n### Optimizer:\nAdam optimizer with the following schedules:\n``` python\n0-20: linear warm up to 3e-3\n20-60: 3e-3\n60-80: 9e-4\n80-90: 3e-4\n90-100: 3e-5\n100-120: 3e-6\n```\n\n### Post-processing:\n- Center the embeddings by plate.\n- Average the embeddings from both sites to obtain per-well embedding.\n- For each cell line, obtain train center embeddings by averaging together the siRNA embeddings.\n- Compute cosine-similarity of each well’s embedding to train center embeddings. \n- Use LSA to compute label assignment based on the 277 leak.\n\n### Pseudo-labeling:\nAn ensemble of 5 models achieved a public LB score of 0.993 and a private LB score of 0.9957 without pseudo-labeling (single model 0.990 and 0.9947).  We then collected all our public LB 0.990+ predictions and identified 327 examples that were not consistent.  All the test predictions not in this set of 327 were then used as pseudo-labels.  An ensemble of models trained on this pseudo-label set achieved a public LB of 0.997, and a private LB of 0.9967.  We also tried iteratively adding more pseudo-labels to the training set (500, 700, 900 per experiment), but it did not improve our public LB score.\n",
    "635739": "Thanks for sharing! Congrats to your achievements and I think your public LB performance motivated many teams during the whole competition period.\nQuestion: \n&gt; An ensemble of 5 models achieved a public LB score\n\nHere 5 models means 5 folds or 5 different backbones?",
    "656159": "Thank you for sharing. Your Model is very interesting.\nWhat is the channel, height and width of feature map before last pooling (e.g. GAP)?\n",
    "635820": "Congratulations! do you have any plan to open source your code solutions?",
    "635786": "",
    "643003": "Congrats, thank you for sharing.",
    "636767": "Thanks for sharing\n",
    "636675": "Thanks for sharing.",
    "635768": "Thanks for sharing! \nGreat solution"
  }
}