{
  "id": 98116,
  "title": "CV/Public LB correlation",
  "url": "/competitions/recursion-cellular-image-classification/discussion/98116",
  "author_name": "",
  "post_date": "2019-07-01T11:20:17.212857Z",
  "votes": 11,
  "comment_count": 36,
  "views": 0,
  "content": "<p>How are your scores on CV vs. the leaderboard? Posting mine below.</p>",
  "messages": [
    {
      "id": "565774",
      "postDate": "07/01/2019 11:20:17",
      "content": "<p>How are your scores on CV vs. the leaderboard? Posting mine below.</p>",
      "rawMarkdown": "How are your scores on CV vs. the leaderboard? Posting mine below.",
      "votes": null
    },
    {
      "id": "565775",
      "postDate": "07/01/2019 11:21:51",
      "content": "<p>I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)</p>\n\n<p><code>\nHEPG2-01   0.100\nHEPG2-04   0.107\nHUVEC-01   0.434\nHUVEC-04   0.647\nHUVEC-07   0.353\nHUVEC-10   0.397\nHUVEC-13   0.502\nHUVEC-16   0.266\nRPE-01     0.338\nRPE-04     0.322\nU2OS-01    0.098\nAll        0.324\n</code></p>\n\n<p>But only 0.090 on the leaderboard. I wonder if this is just me, or it's the same for everyone?</p>",
      "rawMarkdown": "I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)\n\n```\nHEPG2-01   0.100\nHEPG2-04   0.107\nHUVEC-01   0.434\nHUVEC-04   0.647\nHUVEC-07   0.353\nHUVEC-10   0.397\nHUVEC-13   0.502\nHUVEC-16   0.266\nRPE-01     0.338\nRPE-04     0.322\nU2OS-01    0.098\nAll        0.324\n```\n\nBut only 0.090 on the leaderboard. I wonder if this is just me, or it's the same for everyone?",
      "votes": null
    },
    {
      "id": "565821",
      "postDate": "07/01/2019 12:18:20",
      "content": "<p>same for me. I got 0.4+ on CV (one fold) but only 0.094 on LB. Further train the model with more epoch only make the test accuracy worse.</p>",
      "rawMarkdown": "same for me. I got 0.4+ on CV (one fold) but only 0.094 on LB. Further train the model with more epoch only make the test accuracy worse.",
      "votes": null
    },
    {
      "id": "565870",
      "postDate": "07/01/2019 13:34:33",
      "content": "<p>All CV: 22.4% LB 11.7%</p>",
      "rawMarkdown": "All CV: 22.4% LB 11.7%",
      "votes": null
    },
    {
      "id": "565985",
      "postDate": "07/01/2019 16:53:52",
      "content": "<p>We have really strong correlation between CV and LB: \nALL CV: 49% -&gt; LB: 33% \nALL CV: 53% -&gt; LB: 37% \nALL CV: 57% -&gt; LB: 40% </p>",
      "rawMarkdown": "We have really strong correlation between CV and LB: \nALL CV: 49% -&gt; LB: 33% \nALL CV: 53% -&gt; LB: 37% \nALL CV: 57% -&gt; LB: 40%",
      "votes": null
    },
    {
      "id": "566000",
      "postDate": "07/01/2019 17:22:17",
      "content": "<p>Nice progress 🚀 </p>",
      "rawMarkdown": "Nice progress 🚀",
      "votes": null
    },
    {
      "id": "566390",
      "postDate": "07/02/2019 06:00:30",
      "content": "<p>submission file order?</p>",
      "rawMarkdown": "submission file order?",
      "votes": null
    },
    {
      "id": "566718",
      "postDate": "07/02/2019 14:04:23",
      "content": "<p>I confirm that after training the model to a certain epoch the score started decreasing. </p>",
      "rawMarkdown": "I confirm that after training the model to a certain epoch the score started decreasing.",
      "votes": null
    },
    {
      "id": "566836",
      "postDate": "07/02/2019 16:41:14",
      "content": "<p>One fold: \nCV: 33% -&gt; LB: 14.2% \nUpdated:\nCV: 34%, LB: 17.6%</p>",
      "rawMarkdown": "One fold: \nCV: 33% -&gt; LB: 14.2% \nUpdated:\nCV: 34%, LB: 17.6%",
      "votes": null
    },
    {
      "id": "566918",
      "postDate": "07/02/2019 19:20:34",
      "content": "<p>0.39 CV -&gt; 0.069 LB</p>",
      "rawMarkdown": "0.39 CV -&gt; 0.069 LB",
      "votes": null
    },
    {
      "id": "566951",
      "postDate": "07/02/2019 20:28:23",
      "content": "<blockquote>\n  <p>submission file order?</p>\n</blockquote>\n\n<p>Sorry <a href=\"/petewills\">@petewills</a> could you please clarify? Do you mean the score might depend on the order of rows in submission file? Or something else?</p>",
      "rawMarkdown": "&gt; submission file order?\n\nSorry @petewills could you please clarify? Do you mean the score might depend on the order of rows in submission file? Or something else?",
      "votes": null
    },
    {
      "id": "566984",
      "postDate": "07/02/2019 21:59:32",
      "content": "<p>CV: 26.2% -&gt; LB: 11.1%</p>",
      "rawMarkdown": "CV: 26.2% -&gt; LB: 11.1%",
      "votes": null
    },
    {
      "id": "566988",
      "postDate": "07/02/2019 22:07:28",
      "content": "<p>Agree. But it seems training for over 50+ epoch could make a difference. </p>",
      "rawMarkdown": "Agree. But it seems training for over 50+ epoch could make a difference.",
      "votes": null
    },
    {
      "id": "567210",
      "postDate": "07/03/2019 07:16:03",
      "content": "<p>Sorry to be so cryptic. I meant that in other competitions, it mattered as to the order of the entries in the submission - it had to exactly match the example submission.</p>",
      "rawMarkdown": "Sorry to be so cryptic. I meant that in other competitions, it mattered as to the order of the entries in the submission - it had to exactly match the example submission.",
      "votes": null
    },
    {
      "id": "567689",
      "postDate": "07/03/2019 20:49:39",
      "content": "<p>I've found being smart about creating a validation set is important for getting CV/LB correlation.</p>\n\n<p>Same model, similar training:\nRandom sample validation set: 40% CV, 13% LB\nSeparate validation by experiment: 20% CV, 15% LB</p>\n\n<p>I think for this competition it's important to validate your model on plates it hasn't seen during training.</p>",
      "rawMarkdown": "I've found being smart about creating a validation set is important for getting CV/LB correlation.\n\nSame model, similar training:\nRandom sample validation set: 40% CV, 13% LB\nSeparate validation by experiment: 20% CV, 15% LB\n\nI think for this competition it's important to validate your model on plates it hasn't seen during training.",
      "votes": null
    },
    {
      "id": "567828",
      "postDate": "07/04/2019 03:55:57",
      "content": "<p>I think that was already clear</p>\n\n<blockquote>\n  <p><strong>Konstantin Lopuhin wrote</strong></p>\n  \n  <blockquote>\n    <p>I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)</p>\n  </blockquote>\n</blockquote>",
      "rawMarkdown": "I think that was already clear\n\n&gt; **Konstantin Lopuhin wrote**\n&gt; \n&gt; &gt;  I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)",
      "votes": null
    },
    {
      "id": "567829",
      "postDate": "07/04/2019 03:57:12",
      "content": "<blockquote>\n  <p>Separate validation by experiment: 20% CV, 15% LB</p>\n</blockquote>\n\n<p>is that  Kfold or a single fold?</p>",
      "rawMarkdown": "&gt; Separate validation by experiment: 20% CV, 15% LB\n\nis that  Kfold or a single fold?",
      "votes": null
    },
    {
      "id": "567885",
      "postDate": "07/04/2019 06:04:31",
      "content": "<p><a href=\"/sawseen\">@sawseen</a>  I guess your results are KFold. Am I correct?</p>",
      "rawMarkdown": "sawseen  I guess your results are KFold. Am I correct?",
      "votes": null
    },
    {
      "id": "568029",
      "postDate": "07/04/2019 10:09:01",
      "content": "<p>Actually, it's one fold with bag of tricks:)</p>",
      "rawMarkdown": "Actually, it's one fold with bag of tricks:)",
      "votes": null
    },
    {
      "id": "568065",
      "postDate": "07/04/2019 10:55:29",
      "content": "<blockquote>\n  <p>Actually, it's one fold with bag of tricks:)</p>\n</blockquote>\n\n<p>Is metric learning in that bag? :)</p>",
      "rawMarkdown": "&gt; Actually, it's one fold with bag of tricks:)\n\nIs metric learning in that bag? :)",
      "votes": null
    },
    {
      "id": "568094",
      "postDate": "07/04/2019 11:46:40",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a>  I think he found another magic trick :) </p>",
      "rawMarkdown": "lopuhin  I think he found another magic trick :)",
      "votes": null
    },
    {
      "id": "568101",
      "postDate": "07/04/2019 11:59:03",
      "content": "<p>CV: 29% -&gt; LB: 16.8%. A simple stratified random split.</p>",
      "rawMarkdown": "CV: 29% -&gt; LB: 16.8%. A simple stratified random split.",
      "votes": null
    },
    {
      "id": "568469",
      "postDate": "07/05/2019 01:21:19",
      "content": "<p>Random validation: CV 0.33 -&gt; LB 0.20 </p>",
      "rawMarkdown": "Random validation: CV 0.33 -&gt; LB 0.20",
      "votes": null
    },
    {
      "id": "568721",
      "postDate": "07/05/2019 10:29:35",
      "content": "<p>&gt; Is metric learning in that bag? :)</p>\n\n<p><a href=\"/lopuhin\">@lopuhin</a> I can only say that you should give it a try :)</p>",
      "rawMarkdown": "&gt; Is metric learning in that bag? :)\n\n@lopuhin I can only say that you should give it a try :)",
      "votes": null
    },
    {
      "id": "568734",
      "postDate": "07/05/2019 10:52:25",
      "content": "<blockquote>\n  <p>I can only say that you should give it a try :)</p>\n</blockquote>\n\n<p>Haha thanks for encouragement! I tried ArcFace and it didn't converge well, but most likely it's my fault.   So far aiming at more conservative losses ;)</p>",
      "rawMarkdown": "&gt; I can only say that you should give it a try :)\n\nHaha thanks for encouragement! I tried ArcFace and it didn't converge well, but most likely it's my fault.   So far aiming at more conservative losses ;)",
      "votes": null
    },
    {
      "id": "569484",
      "postDate": "07/06/2019 18:45:06",
      "content": "<p>3 fold, CV: 0.619, LB: 0.505\n3fold, CV: 0.697, LB: 0.595</p>",
      "rawMarkdown": "3 fold, CV: 0.619, LB: 0.505\n3fold, CV: 0.697, LB: 0.595",
      "votes": null
    },
    {
      "id": "571286",
      "postDate": "07/09/2019 13:00:29",
      "content": "<p>Thanks! I will check that.</p>",
      "rawMarkdown": "Thanks! I will check that.",
      "votes": null
    },
    {
      "id": "577126",
      "postDate": "07/16/2019 11:45:05",
      "content": "<p><a href=\"/amitkumarjaiswal\">@amitkumarjaiswal</a> what architecture are you using? isnt 50 epochs a bit too much?</p>",
      "rawMarkdown": "amitkumarjaiswal what architecture are you using? isnt 50 epochs a bit too much?",
      "votes": null
    },
    {
      "id": "577174",
      "postDate": "07/16/2019 12:42:24",
      "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> I am using ResNet50 as of now, and my current LB score is based on the model trained over 20 epochs. In my case, training over more epochs getting better.</p>",
      "rawMarkdown": "wjshenggggg I am using ResNet50 as of now, and my current LB score is based on the model trained over 20 epochs. In my case, training over more epochs getting better.",
      "votes": null
    },
    {
      "id": "577261",
      "postDate": "07/16/2019 14:46:31",
      "content": "<p>thank you <a href=\"/amitkumarjaiswal\">@amitkumarjaiswal</a> may I also know what is your validation strategy? Thank you.  :)</p>",
      "rawMarkdown": "thank you @amitkumarjaiswal may I also know what is your validation strategy? Thank you.  :)",
      "votes": null
    },
    {
      "id": "577324",
      "postDate": "07/16/2019 15:49:13",
      "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> Currently using face loss, but that's still I'm not sure about. Anyway, good 👍  that you started a dedicated thread on this.</p>",
      "rawMarkdown": "wjshenggggg Currently using face loss, but that's still I'm not sure about. Anyway, good 👍  that you started a dedicated thread on this.",
      "votes": null
    },
    {
      "id": "577461",
      "postDate": "07/16/2019 17:52:51",
      "content": "<p>thanks! hope to collect some ideas :) </p>",
      "rawMarkdown": "thanks! hope to collect some ideas :)",
      "votes": null
    },
    {
      "id": "579590",
      "postDate": "07/19/2019 02:17:46",
      "content": "<p>hi <a href=\"/amitkumarjaiswal\">@amitkumarjaiswal</a> i tried implementing <code>ArcFaceLoss</code> taken from bestfitting's post here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973</a> and it does not seem to converge at all. The loss increases and the accuracy stays the same after 3 epochs... I wonder am I doing it wrongly (having wrong hyperparameters) or it takes a longer time to converge? Thanks! </p>",
      "rawMarkdown": "hi @amitkumarjaiswal i tried implementing `ArcFaceLoss` taken from bestfitting's post here: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973 and it does not seem to converge at all. The loss increases and the accuracy stays the same after 3 epochs... I wonder am I doing it wrongly (having wrong hyperparameters) or it takes a longer time to converge? Thanks!",
      "votes": null
    },
    {
      "id": "579643",
      "postDate": "07/19/2019 03:53:14",
      "content": "<p>StratifiledKFold(k=3)  CV:0.289 -&gt; LB: 0.110\nI split equally about only sirna.</p>",
      "rawMarkdown": "StratifiledKFold(k=3)  CV:0.289 -&gt; LB: 0.110\nI split equally about only sirna.",
      "votes": null
    },
    {
      "id": "579780",
      "postDate": "07/19/2019 07:34:00",
      "content": "<p>One fold: \nCV: 0.642 -&gt; LB: 0.551 </p>",
      "rawMarkdown": "One fold: \nCV: 0.642 -&gt; LB: 0.551",
      "votes": null
    },
    {
      "id": "579835",
      "postDate": "07/19/2019 09:35:47",
      "content": "<p>I assume it is metric learning and not classification </p>",
      "rawMarkdown": "I assume it is metric learning and not classification",
      "votes": null
    },
    {
      "id": "607796",
      "postDate": "08/25/2019 22:12:35",
      "content": "<p>One-fold:\n0.512 CV --&gt; 0.313 LB</p>",
      "rawMarkdown": "One-fold:\n0.512 CV --&gt; 0.313 LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 565775,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "07/01/2019 11:21:51",
      "content": "<p>I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)</p>\n\n<p><code>\nHEPG2-01   0.100\nHEPG2-04   0.107\nHUVEC-01   0.434\nHUVEC-04   0.647\nHUVEC-07   0.353\nHUVEC-10   0.397\nHUVEC-13   0.502\nHUVEC-16   0.266\nRPE-01     0.338\nRPE-04     0.322\nU2OS-01    0.098\nAll        0.324\n</code></p>\n\n<p>But only 0.090 on the leaderboard. I wonder if this is just me, or it's the same for everyone?</p>",
      "votes": null,
      "replies": [
        {
          "id": 566390,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "07/02/2019 06:00:30",
          "content": "<p>submission file order?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 566951,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/02/2019 20:28:23",
          "content": "<blockquote>\n  <p>submission file order?</p>\n</blockquote>\n\n<p>Sorry <a href=\"/petewills\">@petewills</a> could you please clarify? Do you mean the score might depend on the order of rows in submission file? Or something else?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567210,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "07/03/2019 07:16:03",
          "content": "<p>Sorry to be so cryptic. I meant that in other competitions, it mattered as to the order of the entries in the submission - it had to exactly match the example submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 565821,
      "author_name": "dexter6855",
      "author_url": "",
      "post_date": "07/01/2019 12:18:20",
      "content": "<p>same for me. I got 0.4+ on CV (one fold) but only 0.094 on LB. Further train the model with more epoch only make the test accuracy worse.</p>",
      "votes": null,
      "replies": [
        {
          "id": 566718,
          "author_name": "tayorm",
          "author_url": "",
          "post_date": "07/02/2019 14:04:23",
          "content": "<p>I confirm that after training the model to a certain epoch the score started decreasing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 566988,
          "author_name": "amitkumarjaiswal",
          "author_url": "",
          "post_date": "07/02/2019 22:07:28",
          "content": "<p>Agree. But it seems training for over 50+ epoch could make a difference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 571286,
          "author_name": "tayorm",
          "author_url": "",
          "post_date": "07/09/2019 13:00:29",
          "content": "<p>Thanks! I will check that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577126,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/16/2019 11:45:05",
          "content": "<p><a href=\"/amitkumarjaiswal\">@amitkumarjaiswal</a> what architecture are you using? isnt 50 epochs a bit too much?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577174,
          "author_name": "amitkumarjaiswal",
          "author_url": "",
          "post_date": "07/16/2019 12:42:24",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> I am using ResNet50 as of now, and my current LB score is based on the model trained over 20 epochs. In my case, training over more epochs getting better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577261,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/16/2019 14:46:31",
          "content": "<p>thank you <a href=\"/amitkumarjaiswal\">@amitkumarjaiswal</a> may I also know what is your validation strategy? Thank you.  :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577324,
          "author_name": "amitkumarjaiswal",
          "author_url": "",
          "post_date": "07/16/2019 15:49:13",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> Currently using face loss, but that's still I'm not sure about. Anyway, good 👍  that you started a dedicated thread on this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577461,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/16/2019 17:52:51",
          "content": "<p>thanks! hope to collect some ideas :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 579590,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/19/2019 02:17:46",
          "content": "<p>hi <a href=\"/amitkumarjaiswal\">@amitkumarjaiswal</a> i tried implementing <code>ArcFaceLoss</code> taken from bestfitting's post here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973</a> and it does not seem to converge at all. The loss increases and the accuracy stays the same after 3 epochs... I wonder am I doing it wrongly (having wrong hyperparameters) or it takes a longer time to converge? Thanks! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 565870,
      "author_name": "leighplt",
      "author_url": "",
      "post_date": "07/01/2019 13:34:33",
      "content": "<p>All CV: 22.4% LB 11.7%</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 565985,
      "author_name": "sawseen",
      "author_url": "",
      "post_date": "07/01/2019 16:53:52",
      "content": "<p>We have really strong correlation between CV and LB: \nALL CV: 49% -&gt; LB: 33% \nALL CV: 53% -&gt; LB: 37% \nALL CV: 57% -&gt; LB: 40% </p>",
      "votes": null,
      "replies": [
        {
          "id": 566000,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/01/2019 17:22:17",
          "content": "<p>Nice progress 🚀 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567885,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "07/04/2019 06:04:31",
          "content": "<p><a href=\"/sawseen\">@sawseen</a>  I guess your results are KFold. Am I correct?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568029,
          "author_name": "sawseen",
          "author_url": "",
          "post_date": "07/04/2019 10:09:01",
          "content": "<p>Actually, it's one fold with bag of tricks:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568065,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/04/2019 10:55:29",
          "content": "<blockquote>\n  <p>Actually, it's one fold with bag of tricks:)</p>\n</blockquote>\n\n<p>Is metric learning in that bag? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568094,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "07/04/2019 11:46:40",
          "content": "<p><a href=\"/lopuhin\">@lopuhin</a>  I think he found another magic trick :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568721,
          "author_name": "sawseen",
          "author_url": "",
          "post_date": "07/05/2019 10:29:35",
          "content": "<p>&gt; Is metric learning in that bag? :)</p>\n\n<p><a href=\"/lopuhin\">@lopuhin</a> I can only say that you should give it a try :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568734,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "07/05/2019 10:52:25",
          "content": "<blockquote>\n  <p>I can only say that you should give it a try :)</p>\n</blockquote>\n\n<p>Haha thanks for encouragement! I tried ArcFace and it didn't converge well, but most likely it's my fault.   So far aiming at more conservative losses ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 566836,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "07/02/2019 16:41:14",
      "content": "<p>One fold: \nCV: 33% -&gt; LB: 14.2% \nUpdated:\nCV: 34%, LB: 17.6%</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 566918,
      "author_name": "vshmyhlo",
      "author_url": "",
      "post_date": "07/02/2019 19:20:34",
      "content": "<p>0.39 CV -&gt; 0.069 LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 566984,
      "author_name": "amitkumarjaiswal",
      "author_url": "",
      "post_date": "07/02/2019 21:59:32",
      "content": "<p>CV: 26.2% -&gt; LB: 11.1%</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 567689,
      "author_name": "towardsentropy",
      "author_url": "",
      "post_date": "07/03/2019 20:49:39",
      "content": "<p>I've found being smart about creating a validation set is important for getting CV/LB correlation.</p>\n\n<p>Same model, similar training:\nRandom sample validation set: 40% CV, 13% LB\nSeparate validation by experiment: 20% CV, 15% LB</p>\n\n<p>I think for this competition it's important to validate your model on plates it hasn't seen during training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 567828,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "07/04/2019 03:55:57",
          "content": "<p>I think that was already clear</p>\n\n<blockquote>\n  <p><strong>Konstantin Lopuhin wrote</strong></p>\n  \n  <blockquote>\n    <p>I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)</p>\n  </blockquote>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 567829,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "07/04/2019 03:57:12",
          "content": "<blockquote>\n  <p>Separate validation by experiment: 20% CV, 15% LB</p>\n</blockquote>\n\n<p>is that  Kfold or a single fold?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 568101,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "07/04/2019 11:59:03",
      "content": "<p>CV: 29% -&gt; LB: 16.8%. A simple stratified random split.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 568469,
      "author_name": "yoshitaka1105",
      "author_url": "",
      "post_date": "07/05/2019 01:21:19",
      "content": "<p>Random validation: CV 0.33 -&gt; LB 0.20 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 569484,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "07/06/2019 18:45:06",
      "content": "<p>3 fold, CV: 0.619, LB: 0.505\n3fold, CV: 0.697, LB: 0.595</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 579643,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "07/19/2019 03:53:14",
      "content": "<p>StratifiledKFold(k=3)  CV:0.289 -&gt; LB: 0.110\nI split equally about only sirna.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 579780,
      "author_name": "mathormad",
      "author_url": "",
      "post_date": "07/19/2019 07:34:00",
      "content": "<p>One fold: \nCV: 0.642 -&gt; LB: 0.551 </p>",
      "votes": null,
      "replies": [
        {
          "id": 579835,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "07/19/2019 09:35:47",
          "content": "<p>I assume it is metric learning and not classification </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 607796,
      "author_name": "michelml",
      "author_url": "",
      "post_date": "08/25/2019 22:12:35",
      "content": "<p>One-fold:\n0.512 CV --&gt; 0.313 LB</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "565774": "How are your scores on CV vs. the leaderboard? Posting mine below.",
    "565775": "I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)\n\n```\nHEPG2-01   0.100\nHEPG2-04   0.107\nHUVEC-01   0.434\nHUVEC-04   0.647\nHUVEC-07   0.353\nHUVEC-10   0.397\nHUVEC-13   0.502\nHUVEC-16   0.266\nRPE-01     0.338\nRPE-04     0.322\nU2OS-01    0.098\nAll        0.324\n```\n\nBut only 0.090 on the leaderboard. I wonder if this is just me, or it's the same for everyone?",
    "565821": "same for me. I got 0.4+ on CV (one fold) but only 0.094 on LB. Further train the model with more epoch only make the test accuracy worse.",
    "565870": "All CV: 22.4% LB 11.7%",
    "565985": "We have really strong correlation between CV and LB: \nALL CV: 49% -&gt; LB: 33% \nALL CV: 53% -&gt; LB: 37% \nALL CV: 57% -&gt; LB: 40%",
    "566000": "Nice progress 🚀",
    "566390": "submission file order?",
    "566718": "I confirm that after training the model to a certain epoch the score started decreasing.",
    "566836": "One fold: \nCV: 33% -&gt; LB: 14.2% \nUpdated:\nCV: 34%, LB: 17.6%",
    "566918": "0.39 CV -&gt; 0.069 LB",
    "566951": "&gt; submission file order?\n\nSorry @petewills could you please clarify? Do you mean the score might depend on the order of rows in submission file? Or something else?",
    "566984": "CV: 26.2% -&gt; LB: 11.1%",
    "566988": "Agree. But it seems training for over 50+ epoch could make a difference.",
    "567210": "Sorry to be so cryptic. I meant that in other competitions, it mattered as to the order of the entries in the submission - it had to exactly match the example submission.",
    "567689": "I've found being smart about creating a validation set is important for getting CV/LB correlation.\n\nSame model, similar training:\nRandom sample validation set: 40% CV, 13% LB\nSeparate validation by experiment: 20% CV, 15% LB\n\nI think for this competition it's important to validate your model on plates it hasn't seen during training.",
    "567828": "I think that was already clear\n\n&gt; **Konstantin Lopuhin wrote**\n&gt; \n&gt; &gt;  I split dataset by experiment, evenly distributing cell types, and got this on CV (one fold)",
    "567829": "&gt; Separate validation by experiment: 20% CV, 15% LB\n\nis that  Kfold or a single fold?",
    "567885": "sawseen  I guess your results are KFold. Am I correct?",
    "568029": "Actually, it's one fold with bag of tricks:)",
    "568065": "&gt; Actually, it's one fold with bag of tricks:)\n\nIs metric learning in that bag? :)",
    "568094": "lopuhin  I think he found another magic trick :)",
    "568101": "CV: 29% -&gt; LB: 16.8%. A simple stratified random split.",
    "568469": "Random validation: CV 0.33 -&gt; LB 0.20",
    "568721": "&gt; Is metric learning in that bag? :)\n\n@lopuhin I can only say that you should give it a try :)",
    "568734": "&gt; I can only say that you should give it a try :)\n\nHaha thanks for encouragement! I tried ArcFace and it didn't converge well, but most likely it's my fault.   So far aiming at more conservative losses ;)",
    "569484": "3 fold, CV: 0.619, LB: 0.505\n3fold, CV: 0.697, LB: 0.595",
    "571286": "Thanks! I will check that.",
    "577126": "amitkumarjaiswal what architecture are you using? isnt 50 epochs a bit too much?",
    "577174": "wjshenggggg I am using ResNet50 as of now, and my current LB score is based on the model trained over 20 epochs. In my case, training over more epochs getting better.",
    "577261": "thank you @amitkumarjaiswal may I also know what is your validation strategy? Thank you.  :)",
    "577324": "wjshenggggg Currently using face loss, but that's still I'm not sure about. Anyway, good 👍  that you started a dedicated thread on this.",
    "577461": "thanks! hope to collect some ideas :)",
    "579590": "hi @amitkumarjaiswal i tried implementing `ArcFaceLoss` taken from bestfitting's post here: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973 and it does not seem to converge at all. The loss increases and the accuracy stays the same after 3 epochs... I wonder am I doing it wrongly (having wrong hyperparameters) or it takes a longer time to converge? Thanks!",
    "579643": "StratifiledKFold(k=3)  CV:0.289 -&gt; LB: 0.110\nI split equally about only sirna.",
    "579780": "One fold: \nCV: 0.642 -&gt; LB: 0.551",
    "579835": "I assume it is metric learning and not classification",
    "607796": "One-fold:\n0.512 CV --&gt; 0.313 LB"
  },
  "source": "meta"
}