{
  "id": 98075,
  "title": "Did anyone check the distribution of cell types in public/private?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/98075",
  "author_name": "",
  "post_date": "2019-07-01T06:28:31.280163100Z",
  "votes": 10,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Given that public is just 22% of all test data, it should have a very different distribution of cell types from private, if private/public are split by experiment. This is something that could be checked in around 5 submissions, did anyone try this?</p>",
  "messages": [
    {
      "id": "565586",
      "postDate": "07/01/2019 06:28:31",
      "content": "<p>Given that public is just 22% of all test data, it should have a very different distribution of cell types from private, if private/public are split by experiment. This is something that could be checked in around 5 submissions, did anyone try this?</p>",
      "rawMarkdown": "Given that public is just 22% of all test data, it should have a very different distribution of cell types from private, if private/public are split by experiment. This is something that could be checked in around 5 submissions, did anyone try this?",
      "votes": null
    },
    {
      "id": "565851",
      "postDate": "07/01/2019 13:05:29",
      "content": "<p>22% seems to be one of each cell types, 4 in total. If it was splitted this way, I think our private score should be higher than their public counterpart. But if it was all from the type we have less data to train for, then the public board is going to be useless.</p>",
      "rawMarkdown": "22% seems to be one of each cell types, 4 in total. If it was splitted this way, I think our private score should be higher than their public counterpart. But if it was all from the type we have less data to train for, then the public board is going to be useless.",
      "votes": null
    },
    {
      "id": "565903",
      "postDate": "07/01/2019 14:28:59",
      "content": "<p>So far I checked that removing U2OS-04 or U2OS-05 didn't affect the score (looks like they're all in private, or my predictions for them are too bad but that looks unlikely), and removing all RPE gives a drop from 0.9 to 0.79:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F20095a0c262ff4dd6b8160b048f7f803%2FScreenshot%202019-07-01%20at%2017.02.03.png?generation=1561991320430533&amp;alt=media\" alt=\"\"></p>\n\n<p>That's all my submissions for today, if someone can continue probing I'd appreciate that :)</p>",
      "rawMarkdown": "So far I checked that removing U2OS-04 or U2OS-05 didn't affect the score (looks like they're all in private, or my predictions for them are too bad but that looks unlikely), and removing all RPE gives a drop from 0.9 to 0.79:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F20095a0c262ff4dd6b8160b048f7f803%2FScreenshot%202019-07-01%20at%2017.02.03.png?generation=1561991320430533&amp;alt=media)\n\n\nThat's all my submissions for today, if someone can continue probing I'd appreciate that :)",
      "votes": null
    },
    {
      "id": "565918",
      "postDate": "07/01/2019 14:58:46",
      "content": "<p>remove U2OS does not drop score of my model either, remove either HEPG or RPE would lower the score, and I ran out of subs to test HUVEC.  </p>\n\n<p>I think public board is 1 RPE 1 HUVEC 2 HEPG</p>",
      "rawMarkdown": "remove U2OS does not drop score of my model either, remove either HEPG or RPE would lower the score, and I ran out of subs to test HUVEC.  \n\nI think public board is 1 RPE 1 HUVEC 2 HEPG",
      "votes": null
    },
    {
      "id": "565933",
      "postDate": "07/01/2019 15:14:44",
      "content": "<p>I wonder what would happen to other cell types</p>",
      "rawMarkdown": "I wonder what would happen to other cell types",
      "votes": null
    },
    {
      "id": "567016",
      "postDate": "07/03/2019 00:01:18",
      "content": "<p>On LB ( in total 4 experiments ):\nPRE-08,\n2 HUVEC - HUVEC-17, and ... (not 18)\nHEPG-08</p>",
      "rawMarkdown": "On LB ( in total 4 experiments ):\nPRE-08,\n2 HUVEC - HUVEC-17, and ... (not 18)\nHEPG-08",
      "votes": null
    },
    {
      "id": "567175",
      "postDate": "07/03/2019 06:25:24",
      "content": "<p>Unless I messed something up, HUVEC-17 is in Public - I wonder if you checked it? But I can't find the 4th experiment that should be in public so far. And I can confirm the other two you found (was lucky that I checked them first).</p>",
      "rawMarkdown": "Unless I messed something up, HUVEC-17 is in Public - I wonder if you checked it? But I can't find the 4th experiment that should be in public so far. And I can confirm the other two you found (was lucky that I checked them first).",
      "votes": null
    },
    {
      "id": "567304",
      "postDate": "07/03/2019 10:18:59",
      "content": "<p>My score drops from .362 to .205 after I mask HUVEC-17 predictions all to 1108. It is looking real to me that the public/private is split by certain cell types.</p>\n\n<p>If Kaggle is really splitting public/private by cell type, I couldn't understand the reason. I don't know how to yet, but this split info can protentially be leveraged in solution.</p>",
      "rawMarkdown": "My score drops from .362 to .205 after I mask HUVEC-17 predictions all to 1108. It is looking real to me that the public/private is split by certain cell types.\n\nIf Kaggle is really splitting public/private by cell type, I couldn't understand the reason. I don't know how to yet, but this split info can protentially be leveraged in solution.",
      "votes": null
    },
    {
      "id": "567312",
      "postDate": "07/03/2019 10:28:27",
      "content": "<p>Because total random split give leak, each exp have mostly 1108 classes, so know what on LB u can find part on PB</p>",
      "rawMarkdown": "Because total random split give leak, each exp have mostly 1108 classes, so know what on LB u can find part on PB",
      "votes": null
    },
    {
      "id": "567314",
      "postDate": "07/03/2019 10:37:52",
      "content": "<p>Yeah, that makes sense too, but at least it would be much harder to prob that way I think.</p>",
      "rawMarkdown": "Yeah, that makes sense too, but at least it would be much harder to prob that way I think.",
      "votes": null
    },
    {
      "id": "567448",
      "postDate": "07/03/2019 14:32:26",
      "content": "<p>I think splitting by experiment or cell type is the best kaggle can do here - if they split randomly or by sirna, then public and private LB would be almost the same, so they are forcing us to build a robust solution.\nI'm probing LB mostly to check if my local validation is good enough.</p>",
      "rawMarkdown": "I think splitting by experiment or cell type is the best kaggle can do here - if they split randomly or by sirna, then public and private LB would be almost the same, so they are forcing us to build a robust solution.\nI'm probing LB mostly to check if my local validation is good enough.",
      "votes": null
    },
    {
      "id": "567512",
      "postDate": "07/03/2019 16:08:54",
      "content": "<p>Hard LB probing + data leak or easy probing + no leak. I think what avoid leak more important.</p>",
      "rawMarkdown": "Hard LB probing + data leak or easy probing + no leak. I think what avoid leak more important.",
      "votes": null
    },
    {
      "id": "567690",
      "postDate": "07/03/2019 20:59:02",
      "content": "<p>Thanks! I see now.</p>",
      "rawMarkdown": "Thanks! I see now.",
      "votes": null
    },
    {
      "id": "629270",
      "postDate": "09/18/2019 15:31:54",
      "content": "<p>I did another round of probing with my .687 pb score model. I believe the four experiments in pb set are:</p>\n\n<p>HUVEC-17, RPE-08, HEPG2-08 and U2OS-04  </p>\n\n<p>The reason that U2OS-04 is not stated as in pb set previously is that for a low score model, it got 0 correct for this experiment. </p>",
      "rawMarkdown": "I did another round of probing with my .687 pb score model. I believe the four experiments in pb set are:\n\nHUVEC-17, RPE-08, HEPG2-08 and U2OS-04  \n\nThe reason that U2OS-04 is not stated as in pb set previously is that for a low score model, it got 0 correct for this experiment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 565851,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "07/01/2019 13:05:29",
      "content": "<p>22% seems to be one of each cell types, 4 in total. If it was splitted this way, I think our private score should be higher than their public counterpart. But if it was all from the type we have less data to train for, then the public board is going to be useless.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 565903,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "07/01/2019 14:28:59",
      "content": "<p>So far I checked that removing U2OS-04 or U2OS-05 didn't affect the score (looks like they're all in private, or my predictions for them are too bad but that looks unlikely), and removing all RPE gives a drop from 0.9 to 0.79:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F20095a0c262ff4dd6b8160b048f7f803%2FScreenshot%202019-07-01%20at%2017.02.03.png?generation=1561991320430533&amp;alt=media\" alt=\"\"></p>\n\n<p>That's all my submissions for today, if someone can continue probing I'd appreciate that :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 565918,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "07/01/2019 14:58:46",
          "content": "<p>remove U2OS does not drop score of my model either, remove either HEPG or RPE would lower the score, and I ran out of subs to test HUVEC.  </p>\n\n<p>I think public board is 1 RPE 1 HUVEC 2 HEPG</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 565933,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/01/2019 15:14:44",
          "content": "<p>I wonder what would happen to other cell types</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 567016,
      "author_name": "leighplt",
      "author_url": "",
      "post_date": "07/03/2019 00:01:18",
      "content": "<p>On LB ( in total 4 experiments ):\nPRE-08,\n2 HUVEC - HUVEC-17, and ... (not 18)\nHEPG-08</p>",
      "votes": null,
      "replies": [
        {
          "id": 567175,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/03/2019 06:25:24",
          "content": "<p>Unless I messed something up, HUVEC-17 is in Public - I wonder if you checked it? But I can't find the 4th experiment that should be in public so far. And I can confirm the other two you found (was lucky that I checked them first).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567304,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "07/03/2019 10:18:59",
          "content": "<p>My score drops from .362 to .205 after I mask HUVEC-17 predictions all to 1108. It is looking real to me that the public/private is split by certain cell types.</p>\n\n<p>If Kaggle is really splitting public/private by cell type, I couldn't understand the reason. I don't know how to yet, but this split info can protentially be leveraged in solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567312,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "07/03/2019 10:28:27",
          "content": "<p>Because total random split give leak, each exp have mostly 1108 classes, so know what on LB u can find part on PB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567314,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "07/03/2019 10:37:52",
          "content": "<p>Yeah, that makes sense too, but at least it would be much harder to prob that way I think.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567448,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/03/2019 14:32:26",
          "content": "<p>I think splitting by experiment or cell type is the best kaggle can do here - if they split randomly or by sirna, then public and private LB would be almost the same, so they are forcing us to build a robust solution.\nI'm probing LB mostly to check if my local validation is good enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567512,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "07/03/2019 16:08:54",
          "content": "<p>Hard LB probing + data leak or easy probing + no leak. I think what avoid leak more important.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 567690,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "07/03/2019 20:59:02",
      "content": "<p>Thanks! I see now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 629270,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "09/18/2019 15:31:54",
      "content": "<p>I did another round of probing with my .687 pb score model. I believe the four experiments in pb set are:</p>\n\n<p>HUVEC-17, RPE-08, HEPG2-08 and U2OS-04  </p>\n\n<p>The reason that U2OS-04 is not stated as in pb set previously is that for a low score model, it got 0 correct for this experiment. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "565586": "Given that public is just 22% of all test data, it should have a very different distribution of cell types from private, if private/public are split by experiment. This is something that could be checked in around 5 submissions, did anyone try this?",
    "565851": "22% seems to be one of each cell types, 4 in total. If it was splitted this way, I think our private score should be higher than their public counterpart. But if it was all from the type we have less data to train for, then the public board is going to be useless.",
    "565903": "So far I checked that removing U2OS-04 or U2OS-05 didn't affect the score (looks like they're all in private, or my predictions for them are too bad but that looks unlikely), and removing all RPE gives a drop from 0.9 to 0.79:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F20095a0c262ff4dd6b8160b048f7f803%2FScreenshot%202019-07-01%20at%2017.02.03.png?generation=1561991320430533&amp;alt=media)\n\n\nThat's all my submissions for today, if someone can continue probing I'd appreciate that :)",
    "565918": "remove U2OS does not drop score of my model either, remove either HEPG or RPE would lower the score, and I ran out of subs to test HUVEC.  \n\nI think public board is 1 RPE 1 HUVEC 2 HEPG",
    "565933": "I wonder what would happen to other cell types",
    "567016": "On LB ( in total 4 experiments ):\nPRE-08,\n2 HUVEC - HUVEC-17, and ... (not 18)\nHEPG-08",
    "567175": "Unless I messed something up, HUVEC-17 is in Public - I wonder if you checked it? But I can't find the 4th experiment that should be in public so far. And I can confirm the other two you found (was lucky that I checked them first).",
    "567304": "My score drops from .362 to .205 after I mask HUVEC-17 predictions all to 1108. It is looking real to me that the public/private is split by certain cell types.\n\nIf Kaggle is really splitting public/private by cell type, I couldn't understand the reason. I don't know how to yet, but this split info can protentially be leveraged in solution.",
    "567312": "Because total random split give leak, each exp have mostly 1108 classes, so know what on LB u can find part on PB",
    "567314": "Yeah, that makes sense too, but at least it would be much harder to prob that way I think.",
    "567448": "I think splitting by experiment or cell type is the best kaggle can do here - if they split randomly or by sirna, then public and private LB would be almost the same, so they are forcing us to build a robust solution.\nI'm probing LB mostly to check if my local validation is good enough.",
    "567512": "Hard LB probing + data leak or easy probing + no leak. I think what avoid leak more important.",
    "567690": "Thanks! I see now.",
    "629270": "I did another round of probing with my .687 pb score model. I believe the four experiments in pb set are:\n\nHUVEC-17, RPE-08, HEPG2-08 and U2OS-04  \n\nThe reason that U2OS-04 is not stated as in pb set previously is that for a low score model, it got 0 correct for this experiment."
  },
  "source": "meta"
}