{
  "id": 102905,
  "title": "Treatment siRNAs occur in groups of 277 per plate",
  "url": "/competitions/recursion-cellular-image-classification/discussion/102905",
  "author_name": "",
  "post_date": "2019-08-05T22:30:52.255728600Z",
  "votes": 63,
  "comment_count": 22,
  "views": 0,
  "content": "<p>We want to make everyone aware of structure in the data that we have not previously disclosed, and thank team Double Strand for encouraging us to do so.  As you may know by now, each experiment contains at most one well for each of the 1,108 treatment siRNAs (any treatment siRNA not appearing in an experiment is due to an operational error that occurred while attempting to transfer siRNA from “source” plates to “destination\" plates used in an experiment).  What you may not have noticed yet is that these treatment siRNAs are divided into four groups of 277, and each group always occurs together in a single plate per experiment.  While the plate the group is assigned within an experiment is randomly chosen, and the individual well each siRNA occupies is randomly assigned per plate, the plate will nevertheless contain the same 277 treatment siRNAs as a corresponding plate in another experiment.  This structure occurs as the result of simple operational efficiency — Recursion stores the 1,108 treatment siRNAs in source plates containing 277 siRNA each, and transfers siRNA from these source plates to the destination plates used in an experiment <em>by plate</em>.  Exploiting this non-essential structure in your models can provide some lift to your results, and thus we wanted all to be aware of it.</p>",
  "messages": [
    {
      "id": "592853",
      "postDate": "08/05/2019 22:30:52",
      "content": "<p>We want to make everyone aware of structure in the data that we have not previously disclosed, and thank team Double Strand for encouraging us to do so.  As you may know by now, each experiment contains at most one well for each of the 1,108 treatment siRNAs (any treatment siRNA not appearing in an experiment is due to an operational error that occurred while attempting to transfer siRNA from “source” plates to “destination\" plates used in an experiment).  What you may not have noticed yet is that these treatment siRNAs are divided into four groups of 277, and each group always occurs together in a single plate per experiment.  While the plate the group is assigned within an experiment is randomly chosen, and the individual well each siRNA occupies is randomly assigned per plate, the plate will nevertheless contain the same 277 treatment siRNAs as a corresponding plate in another experiment.  This structure occurs as the result of simple operational efficiency — Recursion stores the 1,108 treatment siRNAs in source plates containing 277 siRNA each, and transfers siRNA from these source plates to the destination plates used in an experiment <em>by plate</em>.  Exploiting this non-essential structure in your models can provide some lift to your results, and thus we wanted all to be aware of it.</p>",
      "rawMarkdown": "We want to make everyone aware of structure in the data that we have not previously disclosed, and thank team Double Strand for encouraging us to do so.  As you may know by now, each experiment contains at most one well for each of the 1,108 treatment siRNAs (any treatment siRNA not appearing in an experiment is due to an operational error that occurred while attempting to transfer siRNA from “source” plates to “destination\" plates used in an experiment).  What you may not have noticed yet is that these treatment siRNAs are divided into four groups of 277, and each group always occurs together in a single plate per experiment.  While the plate the group is assigned within an experiment is randomly chosen, and the individual well each siRNA occupies is randomly assigned per plate, the plate will nevertheless contain the same 277 treatment siRNAs as a corresponding plate in another experiment.  This structure occurs as the result of simple operational efficiency — Recursion stores the 1,108 treatment siRNAs in source plates containing 277 siRNA each, and transfers siRNA from these source plates to the destination plates used in an experiment *by plate*.  Exploiting this non-essential structure in your models can provide some lift to your results, and thus we wanted all to be aware of it.",
      "votes": null
    },
    {
      "id": "592878",
      "postDate": "08/06/2019 00:05:18",
      "content": "<p>Another point here is that out of <code>4*3*2=24</code> possible group assignments to plates there are only 3 in the train data (can be easily observed), and 4 in the test data (our speculation based on some evidence, see the kernel). It seems like Recursion uses some kind of rotation operation on plates, therefore only 4 options.</p>\n\n<p><a href=\"https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak\">There is a kernel</a> that I prepared to demonstrate the leak and show the way to exploit it, it improves the selected public kernel result from 0.113 to 0.207. I can also report that when we found it, it improved our LB 0.661 to 0.769, a similar boost. </p>\n\n<p>Thanks to Recursion for making the public announcement, - this pattern is somewhat tricky to spot, but the competition is not about that. So we now have a problem of assigning 277 sirnas to 277 wells <code>4*18=72</code> times, still a very difficult problem. Good luck to all the participants with the updated challenge!</p>",
      "rawMarkdown": "Another point here is that out of `4*3*2=24` possible group assignments to plates there are only 3 in the train data (can be easily observed), and 4 in the test data (our speculation based on some evidence, see the kernel). It seems like Recursion uses some kind of rotation operation on plates, therefore only 4 options.\n\n[There is a kernel](https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak) that I prepared to demonstrate the leak and show the way to exploit it, it improves the selected public kernel result from 0.113 to 0.207. I can also report that when we found it, it improved our LB 0.661 to 0.769, a similar boost. \n\nThanks to Recursion for making the public announcement, - this pattern is somewhat tricky to spot, but the competition is not about that. So we now have a problem of assigning 277 sirnas to 277 wells `4*18=72` times, still a very difficult problem. Good luck to all the participants with the updated challenge!",
      "votes": null
    },
    {
      "id": "593239",
      "postDate": "08/06/2019 10:38:39",
      "content": "<p>A visualization here: <a href=\"https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization\">https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization</a></p>",
      "rawMarkdown": "A visualization here: https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization",
      "votes": null
    },
    {
      "id": "593315",
      "postDate": "08/06/2019 12:50:48",
      "content": "<p>Thanks for the announcement, this is important to know and should save some people some frustration.. </p>\n\n<p>But people are calling this a leak - is this really considered a data leak?\nI found this like half an hour into my initial EDA - it is more just the structure of the problem than a data leak.</p>",
      "rawMarkdown": "Thanks for the announcement, this is important to know and should save some people some frustration.. \n\nBut people are calling this a leak - is this really considered a data leak?\nI found this like half an hour into my initial EDA - it is more just the structure of the problem than a data leak.",
      "votes": null
    },
    {
      "id": "593319",
      "postDate": "08/06/2019 12:56:17",
      "content": "<p>How is this NOT a leak? Now that you know that the treatments always occur in groups of 277, you can filter your predictions knowing this important hidden point. Is this pattern occurring in the test set? I would almost surely expect so and <a href=\"/zaharch\">@zaharch</a>'s <a href=\"https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak\">kernel</a> seems to confirm this. </p>",
      "rawMarkdown": "How is this NOT a leak? Now that you know that the treatments always occur in groups of 277, you can filter your predictions knowing this important hidden point. Is this pattern occurring in the test set? I would almost surely expect so and @zaharch's [kernel](https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak) seems to confirm this.",
      "votes": null
    },
    {
      "id": "593356",
      "postDate": "08/06/2019 13:41:37",
      "content": "<p>Right, sorry I take that back haha I was a bit quick to write that without doing some reading. </p>\n\n<p>Yes it is definitely a leak and my half an hour of EDA definitely did not find this. Just found the pattern in the train set but not the test set :) So this definitely saved me a bit of time - in fact I just cancelled three training jobs :p</p>",
      "rawMarkdown": "Right, sorry I take that back haha I was a bit quick to write that without doing some reading. \n\nYes it is definitely a leak and my half an hour of EDA definitely did not find this. Just found the pattern in the train set but not the test set :) So this definitely saved me a bit of time - in fact I just cancelled three training jobs :p",
      "votes": null
    },
    {
      "id": "593706",
      "postDate": "08/07/2019 02:18:18",
      "content": "<p>Thanks for taking the time to find, explain, and demonstrate this aspect of the data.</p>",
      "rawMarkdown": "Thanks for taking the time to find, explain, and demonstrate this aspect of the data.",
      "votes": null
    },
    {
      "id": "593721",
      "postDate": "08/07/2019 02:49:25",
      "content": "<p>How many scores may be increased? For your reference, the figure shows the changes of LB scores before and after the \"277 leaks\" is disclosed:\n---update 0807---\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F682501%2Faa9cd93cb1be379349c7465c1f01a528%2Flb_changes.png?generation=1565166644397459&amp;alt=media\" alt=\"\"></p>\n\n<p>Related Code is in the attach file.</p>",
      "rawMarkdown": "How many scores may be increased? For your reference, the figure shows the changes of LB scores before and after the \"277 leaks\" is disclosed:\n---update 0807---\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F682501%2Faa9cd93cb1be379349c7465c1f01a528%2Flb_changes.png?generation=1565166644397459&amp;alt=media)\n\nRelated Code is in the attach file.",
      "votes": null
    },
    {
      "id": "593925",
      "postDate": "08/07/2019 09:47:32",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F2369c4f4b0e07c1750d1ec76ea8569fe%2FScreenshot%202019-08-07%2010.04.36.png?generation=1565171003470494&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F0a73b8c363433db395d7f7c3a2c237cf%2FScreenshot%202019-08-07%2010.04.42.png?generation=1565171028206946&amp;alt=media\" alt=\"\"></p>\n\n<p>These two experiments have the same pattern of controls, but not only... also the treatments are in the same pattern, but the plates are rotated. </p>\n\n<p>Now, there is an experiment in the test set that has the same pattern of controls... HUVEC-18. You can check what I said in <a href=\"https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization\">this kernel</a></p>\n\n<p>Does HUVEC-18 also have the same pattern of treatments in some plate rotation?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F2369c4f4b0e07c1750d1ec76ea8569fe%2FScreenshot%202019-08-07%2010.04.36.png?generation=1565171003470494&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F0a73b8c363433db395d7f7c3a2c237cf%2FScreenshot%202019-08-07%2010.04.42.png?generation=1565171028206946&amp;alt=media)\n\nThese two experiments have the same pattern of controls, but not only... also the treatments are in the same pattern, but the plates are rotated. \n\nNow, there is an experiment in the test set that has the same pattern of controls... HUVEC-18. You can check what I said in [this kernel](https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization)\n\nDoes HUVEC-18 also have the same pattern of treatments in some plate rotation?",
      "votes": null
    },
    {
      "id": "593933",
      "postDate": "08/07/2019 10:00:09",
      "content": "<p>Wow, this is a crazy finding, I hope the competition is not in danger. Recursion has to take a serious look at their random number generator.</p>",
      "rawMarkdown": "Wow, this is a crazy finding, I hope the competition is not in danger. Recursion has to take a serious look at their random number generator.",
      "votes": null
    },
    {
      "id": "593942",
      "postDate": "08/07/2019 10:12:52",
      "content": "<p>I figured it would be better to disclose any leak if aware of it, before this comp becomes a combinatory problem instead of a data science one...</p>",
      "rawMarkdown": "I figured it would be better to disclose any leak if aware of it, before this comp becomes a combinatory problem instead of a data science one...",
      "votes": null
    },
    {
      "id": "593965",
      "postDate": "08/07/2019 11:03:19",
      "content": "<p>The good news is: this seems to be the only case in the train set.</p>",
      "rawMarkdown": "The good news is: this seems to be the only case in the train set.",
      "votes": null
    },
    {
      "id": "594004",
      "postDate": "08/07/2019 12:23:07",
      "content": "<p>Since HUVEC-18 is in private test set (according to this post: <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075#latest-567690\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075#latest-567690</a>), we can't check if the same treatment placement holds. If it is indeed the case, then everyone should basically modify 1 final sub just to account for this possibility. Maybe organisers should remove this experiment from scoring? </p>",
      "rawMarkdown": "Since HUVEC-18 is in private test set (according to this post: https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075#latest-567690), we can't check if the same treatment placement holds. If it is indeed the case, then everyone should basically modify 1 final sub just to account for this possibility. Maybe organisers should remove this experiment from scoring?",
      "votes": null
    },
    {
      "id": "594008",
      "postDate": "08/07/2019 12:32:25",
      "content": "<p>Good point, I had missed that post. Agreed on the removal.</p>",
      "rawMarkdown": "Good point, I had missed that post. Agreed on the removal.",
      "votes": null
    },
    {
      "id": "594018",
      "postDate": "08/07/2019 12:54:26",
      "content": "<p>Could we get a stage 2 testset without well&amp;plate information?</p>",
      "rawMarkdown": "Could we get a stage 2 testset without well&amp;plate information?",
      "votes": null
    },
    {
      "id": "595252",
      "postDate": "08/09/2019 02:11:38",
      "content": "<p>Can someone explain in layman terms how having 277 siRNA on a plate help increase model accuracy? Is it that you will have to train only for 277 siRNA per plate rather 1108?</p>",
      "rawMarkdown": "Can someone explain in layman terms how having 277 siRNA on a plate help increase model accuracy? Is it that you will have to train only for 277 siRNA per plate rather 1108?",
      "votes": null
    },
    {
      "id": "595305",
      "postDate": "08/09/2019 04:54:23",
      "content": "<p>Yes, you can train models that have only 277 classes. However it can benefit models that have already been trained as you can artificially set 3*277 of the predicted results to 0 and then just look at the strongest prediction for the remaining 277 predictions.</p>",
      "rawMarkdown": "Yes, you can train models that have only 277 classes. However it can benefit models that have already been trained as you can artificially set 3*277 of the predicted results to 0 and then just look at the strongest prediction for the remaining 277 predictions.",
      "votes": null
    },
    {
      "id": "596369",
      "postDate": "08/10/2019 14:41:39",
      "content": "<p>Thanks for the explanation!</p>",
      "rawMarkdown": "Thanks for the explanation!",
      "votes": null
    },
    {
      "id": "602180",
      "postDate": "08/18/2019 17:56:10",
      "content": "<p>FYI, It looks like this leakage has since been removed from the leaderboard scoring <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/104011#latest-598578\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/104011#latest-598578</a> Great find <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a>!</p>",
      "rawMarkdown": "FYI, It looks like this leakage has since been removed from the leaderboard scoring https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/104011#latest-598578 Great find @giuliasavorgnan!",
      "votes": null
    },
    {
      "id": "602240",
      "postDate": "08/18/2019 20:03:31",
      "content": "<p>Hi <a href=\"/bearnshaw\">@bearnshaw</a> </p>\n\n<p>What subcellular structure does each <strong>channel</strong> represent? and is the channel meaning constant? Because I'm pretty sure you're using DAPI. For example:</p>\n\n<p><code>\n1.  Nucleus \n2.  Nucleoplasm\n3.  Nucleoli\n4.  Nucleoli fibrillar center\n5.  Nuclear speckles\n6.  Cytoskeleton\n</code></p>\n\n<p>In this case I can see clearly:</p>\n\n<p><code>\n1. Nucleus\n3. Cytoskeleton\n6. Cytoskeleton\n</code></p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fbe1dd2074cf8cfd4d5caa2b67a5cbbab%2FScreenshot%20from%202019-08-18%2022-01-21.png?generation=1566158538027825&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fbe1dd2074cf8cfd4d5caa2b67a5cbbab%2FScreenshot%20from%202019-08-18%2022-01-21.png?generation=1566158538027825&amp;alt=media</a> =700x*)</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @bearnshaw \n\nWhat subcellular structure does each **channel** represent? and is the channel meaning constant? Because I'm pretty sure you're using DAPI. For example:\n\n```\n1.  Nucleus \n2.  Nucleoplasm\n3.  Nucleoli\n4.  Nucleoli fibrillar center\n5.  Nuclear speckles\n6.  Cytoskeleton\n```\n\nIn this case I can see clearly:\n\n```\n1. Nucleus\n3. Cytoskeleton\n6. Cytoskeleton\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fbe1dd2074cf8cfd4d5caa2b67a5cbbab%2FScreenshot%20from%202019-08-18%2022-01-21.png?generation=1566158538027825&amp;alt=media =700x*)\n\n\nThanks!",
      "votes": null
    },
    {
      "id": "603270",
      "postDate": "08/20/2019 04:54:13",
      "content": "<p>Just guessing, but according to the provided info,</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/overview/resources\">https://www.kaggle.com/c/recursion-cellular-image-classification/overview/resources</a>\n(the color map at <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py\">https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py</a>)</li>\n<li>Fig 6 of <a href=\"https://www.rxrx.ai/#the-data\">https://www.rxrx.ai/#the-data</a></li>\n</ul>\n\n<p>it seems</p>\n\n<ol>\n<li>nuclei (blue, 'rgb': [19, 0, 249])</li>\n<li>endoplasmic reticuli (green, 'rgb': [42, 255, 31])</li>\n<li>actin (red, 'rgb': [255, 0, 25])</li>\n<li>nucleoli (cyan, 'rgb': [45, 255, 252])</li>\n<li>mitochondria (magenta, 'rgb': [250, 0, 253])</li>\n<li>golgi apparatus (yellow, 'rgb': [254, 255, 40])</li>\n</ol>",
      "rawMarkdown": "Just guessing, but according to the provided info,\n\n- https://www.kaggle.com/c/recursion-cellular-image-classification/overview/resources\n  (the color map at https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py)\n- Fig 6 of https://www.rxrx.ai/#the-data\n\nit seems\n\n1. nuclei (blue, 'rgb': [19, 0, 249])\n2. endoplasmic reticuli (green, 'rgb': [42, 255, 31])\n3. actin (red, 'rgb': [255, 0, 25])\n4. nucleoli (cyan, 'rgb': [45, 255, 252])\n5. mitochondria (magenta, 'rgb': [250, 0, 253])\n6. golgi apparatus (yellow, 'rgb': [254, 255, 40])",
      "votes": null
    },
    {
      "id": "624515",
      "postDate": "09/12/2019 06:45:17",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> Thx for the detailed explanation! excuse me for not understanding, why are there 4*3*2 group assignments? Also 4*18</p>",
      "rawMarkdown": "zaharch Thx for the detailed explanation! excuse me for not understanding, why are there 4*3*2 group assignments? Also 4*18",
      "votes": null
    },
    {
      "id": "624588",
      "postDate": "09/12/2019 07:40:41",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a> , there are 4 numbered groups of sirnas 1,2,3,4 and 4 numbered plates. It is just combinatorics, there are 24 different ways to match them.  4*18 is number of plates (4) and number of test experiments (18), for each we need to match 277 sirnas into 277 wells.</p>",
      "rawMarkdown": "roguekk007 , there are 4 numbered groups of sirnas 1,2,3,4 and 4 numbered plates. It is just combinatorics, there are 24 different ways to match them.  4*18 is number of plates (4) and number of test experiments (18), for each we need to match 277 sirnas into 277 wells.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 592878,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/06/2019 00:05:18",
      "content": "<p>Another point here is that out of <code>4*3*2=24</code> possible group assignments to plates there are only 3 in the train data (can be easily observed), and 4 in the test data (our speculation based on some evidence, see the kernel). It seems like Recursion uses some kind of rotation operation on plates, therefore only 4 options.</p>\n\n<p><a href=\"https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak\">There is a kernel</a> that I prepared to demonstrate the leak and show the way to exploit it, it improves the selected public kernel result from 0.113 to 0.207. I can also report that when we found it, it improved our LB 0.661 to 0.769, a similar boost. </p>\n\n<p>Thanks to Recursion for making the public announcement, - this pattern is somewhat tricky to spot, but the competition is not about that. So we now have a problem of assigning 277 sirnas to 277 wells <code>4*18=72</code> times, still a very difficult problem. Good luck to all the participants with the updated challenge!</p>",
      "votes": null,
      "replies": [
        {
          "id": 593706,
          "author_name": "npschafer",
          "author_url": "",
          "post_date": "08/07/2019 02:18:18",
          "content": "<p>Thanks for taking the time to find, explain, and demonstrate this aspect of the data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 624515,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "09/12/2019 06:45:17",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> Thx for the detailed explanation! excuse me for not understanding, why are there 4*3*2 group assignments? Also 4*18</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 624588,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/12/2019 07:40:41",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> , there are 4 numbered groups of sirnas 1,2,3,4 and 4 numbered plates. It is just combinatorics, there are 24 different ways to match them.  4*18 is number of plates (4) and number of test experiments (18), for each we need to match 277 sirnas into 277 wells.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 593239,
      "author_name": "giuliasavorgnan",
      "author_url": "",
      "post_date": "08/06/2019 10:38:39",
      "content": "<p>A visualization here: <a href=\"https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization\">https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 593315,
      "author_name": "cherring",
      "author_url": "",
      "post_date": "08/06/2019 12:50:48",
      "content": "<p>Thanks for the announcement, this is important to know and should save some people some frustration.. </p>\n\n<p>But people are calling this a leak - is this really considered a data leak?\nI found this like half an hour into my initial EDA - it is more just the structure of the problem than a data leak.</p>",
      "votes": null,
      "replies": [
        {
          "id": 593319,
          "author_name": "giuliasavorgnan",
          "author_url": "",
          "post_date": "08/06/2019 12:56:17",
          "content": "<p>How is this NOT a leak? Now that you know that the treatments always occur in groups of 277, you can filter your predictions knowing this important hidden point. Is this pattern occurring in the test set? I would almost surely expect so and <a href=\"/zaharch\">@zaharch</a>'s <a href=\"https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak\">kernel</a> seems to confirm this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 593356,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "08/06/2019 13:41:37",
          "content": "<p>Right, sorry I take that back haha I was a bit quick to write that without doing some reading. </p>\n\n<p>Yes it is definitely a leak and my half an hour of EDA definitely did not find this. Just found the pattern in the train set but not the test set :) So this definitely saved me a bit of time - in fact I just cancelled three training jobs :p</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 593721,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "08/07/2019 02:49:25",
      "content": "<p>How many scores may be increased? For your reference, the figure shows the changes of LB scores before and after the \"277 leaks\" is disclosed:\n---update 0807---\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F682501%2Faa9cd93cb1be379349c7465c1f01a528%2Flb_changes.png?generation=1565166644397459&amp;alt=media\" alt=\"\"></p>\n\n<p>Related Code is in the attach file.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 593925,
      "author_name": "giuliasavorgnan",
      "author_url": "",
      "post_date": "08/07/2019 09:47:32",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F2369c4f4b0e07c1750d1ec76ea8569fe%2FScreenshot%202019-08-07%2010.04.36.png?generation=1565171003470494&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F0a73b8c363433db395d7f7c3a2c237cf%2FScreenshot%202019-08-07%2010.04.42.png?generation=1565171028206946&amp;alt=media\" alt=\"\"></p>\n\n<p>These two experiments have the same pattern of controls, but not only... also the treatments are in the same pattern, but the plates are rotated. </p>\n\n<p>Now, there is an experiment in the test set that has the same pattern of controls... HUVEC-18. You can check what I said in <a href=\"https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization\">this kernel</a></p>\n\n<p>Does HUVEC-18 also have the same pattern of treatments in some plate rotation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 593933,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "08/07/2019 10:00:09",
          "content": "<p>Wow, this is a crazy finding, I hope the competition is not in danger. Recursion has to take a serious look at their random number generator.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 593942,
          "author_name": "giuliasavorgnan",
          "author_url": "",
          "post_date": "08/07/2019 10:12:52",
          "content": "<p>I figured it would be better to disclose any leak if aware of it, before this comp becomes a combinatory problem instead of a data science one...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 593965,
          "author_name": "giuliasavorgnan",
          "author_url": "",
          "post_date": "08/07/2019 11:03:19",
          "content": "<p>The good news is: this seems to be the only case in the train set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 594004,
          "author_name": "lukeeee",
          "author_url": "",
          "post_date": "08/07/2019 12:23:07",
          "content": "<p>Since HUVEC-18 is in private test set (according to this post: <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075#latest-567690\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075#latest-567690</a>), we can't check if the same treatment placement holds. If it is indeed the case, then everyone should basically modify 1 final sub just to account for this possibility. Maybe organisers should remove this experiment from scoring? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 594008,
          "author_name": "giuliasavorgnan",
          "author_url": "",
          "post_date": "08/07/2019 12:32:25",
          "content": "<p>Good point, I had missed that post. Agreed on the removal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602180,
          "author_name": "pflood314",
          "author_url": "",
          "post_date": "08/18/2019 17:56:10",
          "content": "<p>FYI, It looks like this leakage has since been removed from the leaderboard scoring <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/104011#latest-598578\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/104011#latest-598578</a> Great find <a href=\"/giuliasavorgnan\">@giuliasavorgnan</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 594018,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "08/07/2019 12:54:26",
      "content": "<p>Could we get a stage 2 testset without well&amp;plate information?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 595252,
      "author_name": "harshpatel1692",
      "author_url": "",
      "post_date": "08/09/2019 02:11:38",
      "content": "<p>Can someone explain in layman terms how having 277 siRNA on a plate help increase model accuracy? Is it that you will have to train only for 277 siRNA per plate rather 1108?</p>",
      "votes": null,
      "replies": [
        {
          "id": 595305,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "08/09/2019 04:54:23",
          "content": "<p>Yes, you can train models that have only 277 classes. However it can benefit models that have already been trained as you can artificially set 3*277 of the predicted results to 0 and then just look at the strongest prediction for the remaining 277 predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596369,
          "author_name": "harshpatel1692",
          "author_url": "",
          "post_date": "08/10/2019 14:41:39",
          "content": "<p>Thanks for the explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 602240,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "08/18/2019 20:03:31",
      "content": "<p>Hi <a href=\"/bearnshaw\">@bearnshaw</a> </p>\n\n<p>What subcellular structure does each <strong>channel</strong> represent? and is the channel meaning constant? Because I'm pretty sure you're using DAPI. For example:</p>\n\n<p><code>\n1.  Nucleus \n2.  Nucleoplasm\n3.  Nucleoli\n4.  Nucleoli fibrillar center\n5.  Nuclear speckles\n6.  Cytoskeleton\n</code></p>\n\n<p>In this case I can see clearly:</p>\n\n<p><code>\n1. Nucleus\n3. Cytoskeleton\n6. Cytoskeleton\n</code></p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fbe1dd2074cf8cfd4d5caa2b67a5cbbab%2FScreenshot%20from%202019-08-18%2022-01-21.png?generation=1566158538027825&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fbe1dd2074cf8cfd4d5caa2b67a5cbbab%2FScreenshot%20from%202019-08-18%2022-01-21.png?generation=1566158538027825&amp;alt=media</a> =700x*)</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 603270,
          "author_name": "harahelix",
          "author_url": "",
          "post_date": "08/20/2019 04:54:13",
          "content": "<p>Just guessing, but according to the provided info,</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/overview/resources\">https://www.kaggle.com/c/recursion-cellular-image-classification/overview/resources</a>\n(the color map at <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py\">https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py</a>)</li>\n<li>Fig 6 of <a href=\"https://www.rxrx.ai/#the-data\">https://www.rxrx.ai/#the-data</a></li>\n</ul>\n\n<p>it seems</p>\n\n<ol>\n<li>nuclei (blue, 'rgb': [19, 0, 249])</li>\n<li>endoplasmic reticuli (green, 'rgb': [42, 255, 31])</li>\n<li>actin (red, 'rgb': [255, 0, 25])</li>\n<li>nucleoli (cyan, 'rgb': [45, 255, 252])</li>\n<li>mitochondria (magenta, 'rgb': [250, 0, 253])</li>\n<li>golgi apparatus (yellow, 'rgb': [254, 255, 40])</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "592853": "We want to make everyone aware of structure in the data that we have not previously disclosed, and thank team Double Strand for encouraging us to do so.  As you may know by now, each experiment contains at most one well for each of the 1,108 treatment siRNAs (any treatment siRNA not appearing in an experiment is due to an operational error that occurred while attempting to transfer siRNA from “source” plates to “destination\" plates used in an experiment).  What you may not have noticed yet is that these treatment siRNAs are divided into four groups of 277, and each group always occurs together in a single plate per experiment.  While the plate the group is assigned within an experiment is randomly chosen, and the individual well each siRNA occupies is randomly assigned per plate, the plate will nevertheless contain the same 277 treatment siRNAs as a corresponding plate in another experiment.  This structure occurs as the result of simple operational efficiency — Recursion stores the 1,108 treatment siRNAs in source plates containing 277 siRNA each, and transfers siRNA from these source plates to the destination plates used in an experiment *by plate*.  Exploiting this non-essential structure in your models can provide some lift to your results, and thus we wanted all to be aware of it.",
    "592878": "Another point here is that out of `4*3*2=24` possible group assignments to plates there are only 3 in the train data (can be easily observed), and 4 in the test data (our speculation based on some evidence, see the kernel). It seems like Recursion uses some kind of rotation operation on plates, therefore only 4 options.\n\n[There is a kernel](https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak) that I prepared to demonstrate the leak and show the way to exploit it, it improves the selected public kernel result from 0.113 to 0.207. I can also report that when we found it, it improved our LB 0.661 to 0.769, a similar boost. \n\nThanks to Recursion for making the public announcement, - this pattern is somewhat tricky to spot, but the competition is not about that. So we now have a problem of assigning 277 sirnas to 277 wells `4*18=72` times, still a very difficult problem. Good luck to all the participants with the updated challenge!",
    "593239": "A visualization here: https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization",
    "593315": "Thanks for the announcement, this is important to know and should save some people some frustration.. \n\nBut people are calling this a leak - is this really considered a data leak?\nI found this like half an hour into my initial EDA - it is more just the structure of the problem than a data leak.",
    "593319": "How is this NOT a leak? Now that you know that the treatments always occur in groups of 277, you can filter your predictions knowing this important hidden point. Is this pattern occurring in the test set? I would almost surely expect so and @zaharch's [kernel](https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak) seems to confirm this.",
    "593356": "Right, sorry I take that back haha I was a bit quick to write that without doing some reading. \n\nYes it is definitely a leak and my half an hour of EDA definitely did not find this. Just found the pattern in the train set but not the test set :) So this definitely saved me a bit of time - in fact I just cancelled three training jobs :p",
    "593706": "Thanks for taking the time to find, explain, and demonstrate this aspect of the data.",
    "593721": "How many scores may be increased? For your reference, the figure shows the changes of LB scores before and after the \"277 leaks\" is disclosed:\n---update 0807---\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F682501%2Faa9cd93cb1be379349c7465c1f01a528%2Flb_changes.png?generation=1565166644397459&amp;alt=media)\n\nRelated Code is in the attach file.",
    "593925": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F2369c4f4b0e07c1750d1ec76ea8569fe%2FScreenshot%202019-08-07%2010.04.36.png?generation=1565171003470494&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F455216%2F0a73b8c363433db395d7f7c3a2c237cf%2FScreenshot%202019-08-07%2010.04.42.png?generation=1565171028206946&amp;alt=media)\n\nThese two experiments have the same pattern of controls, but not only... also the treatments are in the same pattern, but the plates are rotated. \n\nNow, there is an experiment in the test set that has the same pattern of controls... HUVEC-18. You can check what I said in [this kernel](https://www.kaggle.com/giuliasavorgnan/plates-leak-clear-visualization)\n\nDoes HUVEC-18 also have the same pattern of treatments in some plate rotation?",
    "593933": "Wow, this is a crazy finding, I hope the competition is not in danger. Recursion has to take a serious look at their random number generator.",
    "593942": "I figured it would be better to disclose any leak if aware of it, before this comp becomes a combinatory problem instead of a data science one...",
    "593965": "The good news is: this seems to be the only case in the train set.",
    "594004": "Since HUVEC-18 is in private test set (according to this post: https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98075#latest-567690), we can't check if the same treatment placement holds. If it is indeed the case, then everyone should basically modify 1 final sub just to account for this possibility. Maybe organisers should remove this experiment from scoring?",
    "594008": "Good point, I had missed that post. Agreed on the removal.",
    "594018": "Could we get a stage 2 testset without well&amp;plate information?",
    "595252": "Can someone explain in layman terms how having 277 siRNA on a plate help increase model accuracy? Is it that you will have to train only for 277 siRNA per plate rather 1108?",
    "595305": "Yes, you can train models that have only 277 classes. However it can benefit models that have already been trained as you can artificially set 3*277 of the predicted results to 0 and then just look at the strongest prediction for the remaining 277 predictions.",
    "596369": "Thanks for the explanation!",
    "602180": "FYI, It looks like this leakage has since been removed from the leaderboard scoring https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/104011#latest-598578 Great find @giuliasavorgnan!",
    "602240": "Hi @bearnshaw \n\nWhat subcellular structure does each **channel** represent? and is the channel meaning constant? Because I'm pretty sure you're using DAPI. For example:\n\n```\n1.  Nucleus \n2.  Nucleoplasm\n3.  Nucleoli\n4.  Nucleoli fibrillar center\n5.  Nuclear speckles\n6.  Cytoskeleton\n```\n\nIn this case I can see clearly:\n\n```\n1. Nucleus\n3. Cytoskeleton\n6. Cytoskeleton\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fbe1dd2074cf8cfd4d5caa2b67a5cbbab%2FScreenshot%20from%202019-08-18%2022-01-21.png?generation=1566158538027825&amp;alt=media =700x*)\n\n\nThanks!",
    "603270": "Just guessing, but according to the provided info,\n\n- https://www.kaggle.com/c/recursion-cellular-image-classification/overview/resources\n  (the color map at https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py)\n- Fig 6 of https://www.rxrx.ai/#the-data\n\nit seems\n\n1. nuclei (blue, 'rgb': [19, 0, 249])\n2. endoplasmic reticuli (green, 'rgb': [42, 255, 31])\n3. actin (red, 'rgb': [255, 0, 25])\n4. nucleoli (cyan, 'rgb': [45, 255, 252])\n5. mitochondria (magenta, 'rgb': [250, 0, 253])\n6. golgi apparatus (yellow, 'rgb': [254, 255, 40])",
    "624515": "zaharch Thx for the detailed explanation! excuse me for not understanding, why are there 4*3*2 group assignments? Also 4*18",
    "624588": "roguekk007 , there are 4 numbered groups of sirnas 1,2,3,4 and 4 numbered plates. It is just combinatorics, there are 24 different ways to match them.  4*18 is number of plates (4) and number of test experiments (18), for each we need to match 277 sirnas into 277 wells."
  },
  "source": "meta"
}