{
  "id": 110362,
  "title": "5th place solution. AdaBN[domain==plate]",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110362",
  "author_name": "tascj",
  "post_date": "2019-09-27T05:17:52.738000",
  "votes": 35,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First, I would like to thank Recursion and Kaggle for organizing this interesting challenge, and thank team Double strand for their contribution.</p>\n\n<p>Here's a summary of my solution.</p>\n\n<h3>In this competition we have a multi-source and multi-target domain dataset. And we have domain labels.</h3>\n\n<h3>What is \"domain\" in this dataset?</h3>\n\n<p>cell, experiment, or plate</p>\n\n<p>My EDA suggests [domain==experiment]. But [domain==plate] worked better in practice. This is reasonable since batch effects and plate effects exists.</p>\n\n<h3>How to use domain labels? Simply <a href=\"https://arxiv.org/abs/1603.04779\">AdaBN</a>.</h3>\n\n<p>In short, do not challenge your model(with BN layers) with cross-domain batches.</p>\n\n<p>In training, use domain(plate) aware batch sampling.</p>\n\n<p>In testing, use domain batch statistics in BN layers.</p>\n\n<p>With batch norm done right, this competition <strong>IMMEDIATELY</strong> becomes a regular classification challenge: training converges smoothly and I had consistent validation/LB scores (except for HUVEC-05).</p>\n\n<p>Some results of my early experiments, ResNet50, 224x224 input, same hyperparameters\n- random batch sampling, val acc 40+%\n- sample batches in the same cell type, val acc 50+%\n- sample batches in the same experiment, val acc 60+%\n- sample batches in the same plate, val acc 70+%</p>\n\n<h3>Model</h3>\n\n<p>Sequential(BatchNorm2d(6), backbone, neck, head)</p>\n\n<p>backbone: DenseNet201, ResNeXt101_32x8d, HRNet-W18, HRNet-W30</p>\n\n<p>neck: gap or gap+bn</p>\n\n<p>head: 5 fc layers (1 shared and 4 for different cells)</p>\n\n<h3>Loss:</h3>\n\n<p>ArcFaceLoss(s=64, m=0.5) for gap+bn neck</p>\n\n<p>ArcFaceLoss(s=64, m=0.3) for gap neck</p>\n\n<h3>Exemplar Memory</h3>\n\n<p>In this dataset, we could get more supervision than siRNA labels.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F381412%2F27dd6f8424d2b8e2e3d48a55be5bf12a%2Fdataset.png?generation=1569561415969234&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://arxiv.org/abs/1904.01990\">Exemplar Memory</a> fits this structure perfectly.</p>\n\n<p>Fine tuning with 19 exemplar memory modules (HUVEC-05 and 18 test experiments) gave me ~1% LB boost, and HUVEC-05 validation accuracy increased from ~35% to ~46%(~65% with 277 linear assignment)</p>\n\n<h3>Training</h3>\n\n<p>1108-way classifier with treatment only.\ninput 512 -&gt; random crop 384 -&gt; random rot90 -&gt; random hflip\nloss = 0.5 * loss_fc_cell + 0.5 * loss_fc_shared</p>\n\n<p>No pseudo labeling was used.</p>\n\n<h3>Prediction</h3>\n\n<p>input 512\nUse fc_cell.\nNo TTA.\nAveraging two sites.\nlapjv for linear assignment.</p>\n\n<h3>DenseNet201 results</h3>\n\n<p>|  | Public | Private | Public (leak) | Private (leak) |\n| --- | --- | --- | --- | --- |\n| train data only | 0.92620 |0.97325 | 0.98307 | 0.99303 \n| + exemplar memory fine tune | 0.95531 | 0.98321 | 0.98691 | 0.99394</p>\n\n<h3>Some interesting finding</h3>\n\n<p>HUVEC-05 prediction of fc_shared is always better than fc_cell. What's wrong with this experiment?</p>",
  "messages": [
    {
      "id": 635044,
      "postDate": "2019-09-27T05:17:52.740Z",
      "content": "<p>First, I would like to thank Recursion and Kaggle for organizing this interesting challenge, and thank team Double strand for their contribution.</p>\n\n<p>Here's a summary of my solution.</p>\n\n<h3>In this competition we have a multi-source and multi-target domain dataset. And we have domain labels.</h3>\n\n<h3>What is \"domain\" in this dataset?</h3>\n\n<p>cell, experiment, or plate</p>\n\n<p>My EDA suggests [domain==experiment]. But [domain==plate] worked better in practice. This is reasonable since batch effects and plate effects exists.</p>\n\n<h3>How to use domain labels? Simply <a href=\"https://arxiv.org/abs/1603.04779\">AdaBN</a>.</h3>\n\n<p>In short, do not challenge your model(with BN layers) with cross-domain batches.</p>\n\n<p>In training, use domain(plate) aware batch sampling.</p>\n\n<p>In testing, use domain batch statistics in BN layers.</p>\n\n<p>With batch norm done right, this competition <strong>IMMEDIATELY</strong> becomes a regular classification challenge: training converges smoothly and I had consistent validation/LB scores (except for HUVEC-05).</p>\n\n<p>Some results of my early experiments, ResNet50, 224x224 input, same hyperparameters\n- random batch sampling, val acc 40+%\n- sample batches in the same cell type, val acc 50+%\n- sample batches in the same experiment, val acc 60+%\n- sample batches in the same plate, val acc 70+%</p>\n\n<h3>Model</h3>\n\n<p>Sequential(BatchNorm2d(6), backbone, neck, head)</p>\n\n<p>backbone: DenseNet201, ResNeXt101_32x8d, HRNet-W18, HRNet-W30</p>\n\n<p>neck: gap or gap+bn</p>\n\n<p>head: 5 fc layers (1 shared and 4 for different cells)</p>\n\n<h3>Loss:</h3>\n\n<p>ArcFaceLoss(s=64, m=0.5) for gap+bn neck</p>\n\n<p>ArcFaceLoss(s=64, m=0.3) for gap neck</p>\n\n<h3>Exemplar Memory</h3>\n\n<p>In this dataset, we could get more supervision than siRNA labels.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F381412%2F27dd6f8424d2b8e2e3d48a55be5bf12a%2Fdataset.png?generation=1569561415969234&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://arxiv.org/abs/1904.01990\">Exemplar Memory</a> fits this structure perfectly.</p>\n\n<p>Fine tuning with 19 exemplar memory modules (HUVEC-05 and 18 test experiments) gave me ~1% LB boost, and HUVEC-05 validation accuracy increased from ~35% to ~46%(~65% with 277 linear assignment)</p>\n\n<h3>Training</h3>\n\n<p>1108-way classifier with treatment only.\ninput 512 -&gt; random crop 384 -&gt; random rot90 -&gt; random hflip\nloss = 0.5 * loss_fc_cell + 0.5 * loss_fc_shared</p>\n\n<p>No pseudo labeling was used.</p>\n\n<h3>Prediction</h3>\n\n<p>input 512\nUse fc_cell.\nNo TTA.\nAveraging two sites.\nlapjv for linear assignment.</p>\n\n<h3>DenseNet201 results</h3>\n\n<p>|  | Public | Private | Public (leak) | Private (leak) |\n| --- | --- | --- | --- | --- |\n| train data only | 0.92620 |0.97325 | 0.98307 | 0.99303 \n| + exemplar memory fine tune | 0.95531 | 0.98321 | 0.98691 | 0.99394</p>\n\n<h3>Some interesting finding</h3>\n\n<p>HUVEC-05 prediction of fc_shared is always better than fc_cell. What's wrong with this experiment?</p>",
      "rawMarkdown": "First, I would like to thank Recursion and Kaggle for organizing this interesting challenge, and thank team Double strand for their contribution.\n\n\nHere's a summary of my solution.\n\n\n### In this competition we have a multi-source and multi-target domain dataset. And we have domain labels.\n\n### What is \"domain\" in this dataset?\n\ncell, experiment, or plate\n\nMy EDA suggests \\[domain==experiment\\]. But \\[domain==plate\\] worked better in practice. This is reasonable since batch effects and plate effects exists.\n\n### How to use domain labels? Simply [AdaBN](https://arxiv.org/abs/1603.04779).\nIn short, do not challenge your model(with BN layers) with cross-domain batches.\n\nIn training, use domain(plate) aware batch sampling.\n\nIn testing, use domain batch statistics in BN layers.\n\nWith batch norm done right, this competition **IMMEDIATELY** becomes a regular classification challenge: training converges smoothly and I had consistent validation/LB scores (except for HUVEC-05).\n\nSome results of my early experiments, ResNet50, 224x224 input, same hyperparameters\n- random batch sampling, val acc 40+%\n- sample batches in the same cell type, val acc 50+%\n- sample batches in the same experiment, val acc 60+%\n- sample batches in the same plate, val acc 70+%\n\n### Model\nSequential(BatchNorm2d(6), backbone, neck, head)\n\nbackbone: DenseNet201, ResNeXt101_32x8d, HRNet-W18, HRNet-W30\n\nneck: gap or gap+bn\n\nhead: 5 fc layers (1 shared and 4 for different cells)\n\n### Loss:\n\nArcFaceLoss(s=64, m=0.5) for gap+bn neck\n\nArcFaceLoss(s=64, m=0.3) for gap neck\n\n### Exemplar Memory\nIn this dataset, we could get more supervision than siRNA labels.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F381412%2F27dd6f8424d2b8e2e3d48a55be5bf12a%2Fdataset.png?generation=1569561415969234&amp;alt=media)\n\n\n[Exemplar Memory](https://arxiv.org/abs/1904.01990) fits this structure perfectly.\n\nFine tuning with 19 exemplar memory modules (HUVEC-05 and 18 test experiments) gave me ~1% LB boost, and HUVEC-05 validation accuracy increased from ~35% to ~46%(~65% with 277 linear assignment)\n\n### Training\n1108-way classifier with treatment only.\ninput 512 -&gt; random crop 384 -&gt; random rot90 -&gt; random hflip\nloss = 0.5 * loss\\_fc\\_cell + 0.5 * loss\\_fc\\_shared\n\nNo pseudo labeling was used.\n\n### Prediction\ninput 512\nUse fc_cell.\nNo TTA.\nAveraging two sites.\nlapjv for linear assignment.\n\n### DenseNet201 results\n\n|  | Public | Private | Public (leak) | Private (leak) |\n| --- | --- | --- | --- | --- |\n| train data only | 0.92620 |0.97325 | 0.98307 | 0.99303 \n| + exemplar memory fine tune | 0.95531 | 0.98321 | 0.98691 | 0.99394\n\n### Some interesting finding\n\nHUVEC-05 prediction of fc\\_shared is always better than fc\\_cell. What's wrong with this experiment?",
      "votes": 35
    },
    {
      "id": 635151,
      "postDate": "2019-09-27T07:51:07.173Z",
      "content": "<p>Awesome solution, congratulations. Thanks for sharing! It's great to see that there are alternatives to pseudo labeling &amp; huge ensembles.</p>\n\n<p>&gt; In testing, use domain batch statistics in BN layers.</p>\n\n<p>How does this look like in practice? A snippet would be great. Do you run forward 2 times? I.e first to compute the statistics and then use these statistics.</p>",
      "rawMarkdown": "Awesome solution, congratulations. Thanks for sharing! It's great to see that there are alternatives to pseudo labeling &amp; huge ensembles.\n\n&gt; In testing, use domain batch statistics in BN layers.\n\nHow does this look like in practice? A snippet would be great. Do you run forward 2 times? I.e first to compute the statistics and then use these statistics.",
      "votes": 1,
      "replies": [
        {
          "id": 635190,
          "postDate": "2019-09-27T08:43:16.770Z",
          "content": "<p>In practice, I simply set <code>track_running_stats</code> of all BN layers to <code>False</code> and use large batch size (111) in val/test. Prediction is stable if batch size is large enough. I tried to fit a whole plate into a batch with FP16 and SyncBN, but the accuracy did not change compared to batch size 111.</p>\n\n<p>You could estimate domain statistics with an extra forward, but that may not be accurate (think about several sequential BN layers). There is a good example in the repo of <a href=\"https://github.com/timgaripov/swa/blob/master/utils.py\">SWA</a>. Actually, I noticed the domain shift when trying SWA.</p>",
          "rawMarkdown": "In practice, I simply set `track_running_stats` of all BN layers to `False` and use large batch size (111) in val/test. Prediction is stable if batch size is large enough. I tried to fit a whole plate into a batch with FP16 and SyncBN, but the accuracy did not change compared to batch size 111.\n\nYou could estimate domain statistics with an extra forward, but that may not be accurate (think about several sequential BN layers). There is a good example in the repo of [SWA](https://github.com/timgaripov/swa/blob/master/utils.py). Actually, I noticed the domain shift when trying SWA.",
          "votes": 6
        },
        {
          "id": 635467,
          "postDate": "2019-09-27T15:46:51.693Z",
          "content": "<p>If you want to try AdaBN in pytorch,</p>\n\n<p>```\nfrom torch.nn.modules.batchnorm import _BatchNorm</p>\n\n<p>for m in net.modules():\n    if isinstance(m, _BatchNorm):\n        m.track_running_stats = False\n```</p>\n\n<p>```\nimport random\nimport itertools\nfrom torch.utils.data.sampler import BatchSampler</p>\n\n<p>class TrainBatchSampler(BatchSampler):\n    def <strong>init</strong>(self, dataframe, batch_size):\n        self.batch_size = batch_size</p>\n\n<pre><code>    dataframe = dataframe.copy().reset_index(drop=True)\n    index_groups = []\n    for _, df in dataframe.groupby(['experiment', 'plate']):\n        index_groups.append(df.index.values)\n    self.group_sizes = [len(g) // self.batch_size for g in index_groups]\n\n    self.index_groups = [\n        self._take_every(self._cycle_with_shuffle(g), self.batch_size)\n        for g in index_groups\n    ]\n    self.length = sum(self.group_sizes)\n\ndef __len__(self):\n    return self.length\n\ndef __iter__(self):\n    batches = []\n    for size, group in zip(self.group_sizes, self.index_groups):\n        for _ in range(size):\n            batches.append(next(group))\n\n    random.shuffle(batches)\n    return iter(batches)\n\ndef _cycle_with_shuffle(self, xs):\n    while True:\n        random.shuffle(xs)\n        yield from xs\n\ndef _take_every(self, it, n):\n    while True:\n        chunk = []\n        for _ in range(n):\n            chunk.append(next(it))\n        yield chunk\n</code></pre>\n\n<p>class TestBatchSampler(BatchSampler):\n    def <strong>init</strong>(self, dataframe, batch_size):\n        self.batch_size = batch_size</p>\n\n<pre><code>    dataframe = dataframe.copy().reset_index(drop=True)\n    index_groups = []\n    for _, df in dataframe.groupby(['experiment', 'plate']):\n        index_groups.append(df.index.values)\n\n    self.batches = []\n    for g in index_groups:\n        self.batches.extend(self._split_every(g, self.batch_size))\n\ndef __len__(self):\n    return len(self.batches)\n\ndef __iter__(self):\n    return iter(self.batches)\n\ndef _split_every(self, xs, n):\n    ret = []\n    chunk = []\n    for x in xs:\n        chunk.append(x)\n        if len(chunk) == n:\n            ret.append(chunk)\n            chunk = []\n    if chunk:\n        ret.append(chunk)\n    return ret\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "If you want to try AdaBN in pytorch,\n\n```\nfrom torch.nn.modules.batchnorm import _BatchNorm\n\nfor m in net.modules():\n    if isinstance(m, _BatchNorm):\n        m.track_running_stats = False\n```\n\n```\nimport random\nimport itertools\nfrom torch.utils.data.sampler import BatchSampler\n\n\nclass TrainBatchSampler(BatchSampler):\n    def __init__(self, dataframe, batch_size):\n        self.batch_size = batch_size\n\n        dataframe = dataframe.copy().reset_index(drop=True)\n        index_groups = []\n        for _, df in dataframe.groupby(['experiment', 'plate']):\n            index_groups.append(df.index.values)\n        self.group_sizes = [len(g) // self.batch_size for g in index_groups]\n\n        self.index_groups = [\n            self._take_every(self._cycle_with_shuffle(g), self.batch_size)\n            for g in index_groups\n        ]\n        self.length = sum(self.group_sizes)\n\n    def __len__(self):\n        return self.length\n\n    def __iter__(self):\n        batches = []\n        for size, group in zip(self.group_sizes, self.index_groups):\n            for _ in range(size):\n                batches.append(next(group))\n\n        random.shuffle(batches)\n        return iter(batches)\n\n    def _cycle_with_shuffle(self, xs):\n        while True:\n            random.shuffle(xs)\n            yield from xs\n\n    def _take_every(self, it, n):\n        while True:\n            chunk = []\n            for _ in range(n):\n                chunk.append(next(it))\n            yield chunk\n\n\nclass TestBatchSampler(BatchSampler):\n    def __init__(self, dataframe, batch_size):\n        self.batch_size = batch_size\n\n        dataframe = dataframe.copy().reset_index(drop=True)\n        index_groups = []\n        for _, df in dataframe.groupby(['experiment', 'plate']):\n            index_groups.append(df.index.values)\n\n        self.batches = []\n        for g in index_groups:\n            self.batches.extend(self._split_every(g, self.batch_size))\n\n    def __len__(self):\n        return len(self.batches)\n\n    def __iter__(self):\n        return iter(self.batches)\n\n    def _split_every(self, xs, n):\n        ret = []\n        chunk = []\n        for x in xs:\n            chunk.append(x)\n            if len(chunk) == n:\n                ret.append(chunk)\n                chunk = []\n        if chunk:\n            ret.append(chunk)\n        return ret\n```",
          "votes": 4
        }
      ]
    },
    {
      "id": 635113,
      "postDate": "2019-09-27T07:10:51.537Z",
      "content": "<p>Wow, AdaBN and exemplar memory beautifully deals with batch effect and controls. learned a lot. Thanks for sharing!</p>",
      "rawMarkdown": "Wow, AdaBN and exemplar memory beautifully deals with batch effect and controls. learned a lot. Thanks for sharing!",
      "votes": 1,
      "replies": [
        {
          "id": 635466,
          "postDate": "2019-09-27T15:41:23.667Z",
          "content": "<p>Sorry I forgot to add some important details. </p>\n\n<p>Actually, I ended up with 1108-way classification models. Positive controls does not contribute to accuracy. Once negative controls added to batches, training became unstable.</p>",
          "rawMarkdown": "Sorry I forgot to add some important details. \n\nActually, I ended up with 1108-way classification models. Positive controls does not contribute to accuracy. Once negative controls added to batches, training became unstable.",
          "votes": 1
        }
      ]
    },
    {
      "id": 657214,
      "postDate": "2019-10-25T00:57:14.637Z",
      "content": "<p><a href=\"/tascj0\">@tascj0</a> Excellent solution! I really like this elegant approach.</p>\n\n<blockquote>\n  <p>In training, use domain(plate) aware batch sampling.</p>\n</blockquote>\n\n<p>Did you try to use different BN layer for each domains (plates) in training?\nI think it is natural to use different BN layer for different domains even in training in addition to domain-aware batch sampling.</p>",
      "rawMarkdown": "@tascj0 Excellent solution! I really like this elegant approach.\n\n&gt; In training, use domain(plate) aware batch sampling.\n\nDid you try to use different BN layer for each domains (plates) in training?\nI think it is natural to use different BN layer for different domains even in training in addition to domain-aware batch sampling.",
      "replies": [
        {
          "id": 657346,
          "postDate": "2019-10-25T03:52:10.487Z",
          "content": "<p>I had the idea of conditional BN gamma/beta for different domains, but this would not work for unseen domains.\nI've also tried some structures of \"BN gamma/beta from negative controls\". I believe that could be a real solution to this challenge. But no luck with that.</p>",
          "rawMarkdown": "I had the idea of conditional BN gamma/beta for different domains, but this would not work for unseen domains.\nI've also tried some structures of \"BN gamma/beta from negative controls\". I believe that could be a real solution to this challenge. But no luck with that.",
          "votes": 1
        },
        {
          "id": 658126,
          "postDate": "2019-10-25T17:29:36.413Z",
          "content": "<p>I see, thx!</p>",
          "rawMarkdown": "I see, thx!"
        }
      ]
    },
    {
      "id": 635055,
      "postDate": "2019-09-27T05:29:47.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 791520,
      "postDate": "2020-03-30T13:36:28.863Z",
      "content": "<p>Congrats Thanks for sharing !</p>",
      "rawMarkdown": "Congrats Thanks for sharing !",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 635151,
      "author_name": "See--",
      "author_url": "",
      "post_date": "2019-09-27T07:51:07.173000",
      "content": "<p>Awesome solution, congratulations. Thanks for sharing! It's great to see that there are alternatives to pseudo labeling &amp; huge ensembles.</p>\n\n<p>&gt; In testing, use domain batch statistics in BN layers.</p>\n\n<p>How does this look like in practice? A snippet would be great. Do you run forward 2 times? I.e first to compute the statistics and then use these statistics.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 635190,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2019-09-27T08:43:16.770000",
          "content": "<p>In practice, I simply set <code>track_running_stats</code> of all BN layers to <code>False</code> and use large batch size (111) in val/test. Prediction is stable if batch size is large enough. I tried to fit a whole plate into a batch with FP16 and SyncBN, but the accuracy did not change compared to batch size 111.</p>\n\n<p>You could estimate domain statistics with an extra forward, but that may not be accurate (think about several sequential BN layers). There is a good example in the repo of <a href=\"https://github.com/timgaripov/swa/blob/master/utils.py\">SWA</a>. Actually, I noticed the domain shift when trying SWA.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 635467,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2019-09-27T15:46:51.693000",
          "content": "<p>If you want to try AdaBN in pytorch,</p>\n\n<p>```\nfrom torch.nn.modules.batchnorm import _BatchNorm</p>\n\n<p>for m in net.modules():\n    if isinstance(m, _BatchNorm):\n        m.track_running_stats = False\n```</p>\n\n<p>```\nimport random\nimport itertools\nfrom torch.utils.data.sampler import BatchSampler</p>\n\n<p>class TrainBatchSampler(BatchSampler):\n    def <strong>init</strong>(self, dataframe, batch_size):\n        self.batch_size = batch_size</p>\n\n<pre><code>    dataframe = dataframe.copy().reset_index(drop=True)\n    index_groups = []\n    for _, df in dataframe.groupby(['experiment', 'plate']):\n        index_groups.append(df.index.values)\n    self.group_sizes = [len(g) // self.batch_size for g in index_groups]\n\n    self.index_groups = [\n        self._take_every(self._cycle_with_shuffle(g), self.batch_size)\n        for g in index_groups\n    ]\n    self.length = sum(self.group_sizes)\n\ndef __len__(self):\n    return self.length\n\ndef __iter__(self):\n    batches = []\n    for size, group in zip(self.group_sizes, self.index_groups):\n        for _ in range(size):\n            batches.append(next(group))\n\n    random.shuffle(batches)\n    return iter(batches)\n\ndef _cycle_with_shuffle(self, xs):\n    while True:\n        random.shuffle(xs)\n        yield from xs\n\ndef _take_every(self, it, n):\n    while True:\n        chunk = []\n        for _ in range(n):\n            chunk.append(next(it))\n        yield chunk\n</code></pre>\n\n<p>class TestBatchSampler(BatchSampler):\n    def <strong>init</strong>(self, dataframe, batch_size):\n        self.batch_size = batch_size</p>\n\n<pre><code>    dataframe = dataframe.copy().reset_index(drop=True)\n    index_groups = []\n    for _, df in dataframe.groupby(['experiment', 'plate']):\n        index_groups.append(df.index.values)\n\n    self.batches = []\n    for g in index_groups:\n        self.batches.extend(self._split_every(g, self.batch_size))\n\ndef __len__(self):\n    return len(self.batches)\n\ndef __iter__(self):\n    return iter(self.batches)\n\ndef _split_every(self, xs, n):\n    ret = []\n    chunk = []\n    for x in xs:\n        chunk.append(x)\n        if len(chunk) == n:\n            ret.append(chunk)\n            chunk = []\n    if chunk:\n        ret.append(chunk)\n    return ret\n</code></pre>\n\n<p>```</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 635113,
      "author_name": "dohlee",
      "author_url": "",
      "post_date": "2019-09-27T07:10:51.537000",
      "content": "<p>Wow, AdaBN and exemplar memory beautifully deals with batch effect and controls. learned a lot. Thanks for sharing!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 635466,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2019-09-27T15:41:23.667000",
          "content": "<p>Sorry I forgot to add some important details. </p>\n\n<p>Actually, I ended up with 1108-way classification models. Positive controls does not contribute to accuracy. Once negative controls added to batches, training became unstable.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 657214,
      "author_name": "yu4u",
      "author_url": "",
      "post_date": "2019-10-25T00:57:14.637000",
      "content": "<p><a href=\"/tascj0\">@tascj0</a> Excellent solution! I really like this elegant approach.</p>\n\n<blockquote>\n  <p>In training, use domain(plate) aware batch sampling.</p>\n</blockquote>\n\n<p>Did you try to use different BN layer for each domains (plates) in training?\nI think it is natural to use different BN layer for different domains even in training in addition to domain-aware batch sampling.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 657346,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2019-10-25T03:52:10.487000",
          "content": "<p>I had the idea of conditional BN gamma/beta for different domains, but this would not work for unseen domains.\nI've also tried some structures of \"BN gamma/beta from negative controls\". I believe that could be a real solution to this challenge. But no luck with that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 658126,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2019-10-25T17:29:36.413000",
          "content": "<p>I see, thx!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 635055,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T05:29:47.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 791520,
      "author_name": "Muhammet Ikbal Elek",
      "author_url": "",
      "post_date": "2020-03-30T13:36:28.863000",
      "content": "<p>Congrats Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635044": "First, I would like to thank Recursion and Kaggle for organizing this interesting challenge, and thank team Double strand for their contribution.\n\n\nHere's a summary of my solution.\n\n\n### In this competition we have a multi-source and multi-target domain dataset. And we have domain labels.\n\n### What is \"domain\" in this dataset?\n\ncell, experiment, or plate\n\nMy EDA suggests \\[domain==experiment\\]. But \\[domain==plate\\] worked better in practice. This is reasonable since batch effects and plate effects exists.\n\n### How to use domain labels? Simply [AdaBN](https://arxiv.org/abs/1603.04779).\nIn short, do not challenge your model(with BN layers) with cross-domain batches.\n\nIn training, use domain(plate) aware batch sampling.\n\nIn testing, use domain batch statistics in BN layers.\n\nWith batch norm done right, this competition **IMMEDIATELY** becomes a regular classification challenge: training converges smoothly and I had consistent validation/LB scores (except for HUVEC-05).\n\nSome results of my early experiments, ResNet50, 224x224 input, same hyperparameters\n- random batch sampling, val acc 40+%\n- sample batches in the same cell type, val acc 50+%\n- sample batches in the same experiment, val acc 60+%\n- sample batches in the same plate, val acc 70+%\n\n### Model\nSequential(BatchNorm2d(6), backbone, neck, head)\n\nbackbone: DenseNet201, ResNeXt101_32x8d, HRNet-W18, HRNet-W30\n\nneck: gap or gap+bn\n\nhead: 5 fc layers (1 shared and 4 for different cells)\n\n### Loss:\n\nArcFaceLoss(s=64, m=0.5) for gap+bn neck\n\nArcFaceLoss(s=64, m=0.3) for gap neck\n\n### Exemplar Memory\nIn this dataset, we could get more supervision than siRNA labels.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F381412%2F27dd6f8424d2b8e2e3d48a55be5bf12a%2Fdataset.png?generation=1569561415969234&amp;alt=media)\n\n\n[Exemplar Memory](https://arxiv.org/abs/1904.01990) fits this structure perfectly.\n\nFine tuning with 19 exemplar memory modules (HUVEC-05 and 18 test experiments) gave me ~1% LB boost, and HUVEC-05 validation accuracy increased from ~35% to ~46%(~65% with 277 linear assignment)\n\n### Training\n1108-way classifier with treatment only.\ninput 512 -&gt; random crop 384 -&gt; random rot90 -&gt; random hflip\nloss = 0.5 * loss\\_fc\\_cell + 0.5 * loss\\_fc\\_shared\n\nNo pseudo labeling was used.\n\n### Prediction\ninput 512\nUse fc_cell.\nNo TTA.\nAveraging two sites.\nlapjv for linear assignment.\n\n### DenseNet201 results\n\n|  | Public | Private | Public (leak) | Private (leak) |\n| --- | --- | --- | --- | --- |\n| train data only | 0.92620 |0.97325 | 0.98307 | 0.99303 \n| + exemplar memory fine tune | 0.95531 | 0.98321 | 0.98691 | 0.99394\n\n### Some interesting finding\n\nHUVEC-05 prediction of fc\\_shared is always better than fc\\_cell. What's wrong with this experiment?",
    "635151": "Awesome solution, congratulations. Thanks for sharing! It's great to see that there are alternatives to pseudo labeling &amp; huge ensembles.\n\n&gt; In testing, use domain batch statistics in BN layers.\n\nHow does this look like in practice? A snippet would be great. Do you run forward 2 times? I.e first to compute the statistics and then use these statistics.",
    "635113": "Wow, AdaBN and exemplar memory beautifully deals with batch effect and controls. learned a lot. Thanks for sharing!",
    "657214": "@tascj0 Excellent solution! I really like this elegant approach.\n\n&gt; In training, use domain(plate) aware batch sampling.\n\nDid you try to use different BN layer for each domains (plates) in training?\nI think it is natural to use different BN layer for different domains even in training in addition to domain-aware batch sampling.",
    "635055": "",
    "791520": "Congrats Thanks for sharing !"
  }
}