{
  "id": 266652,
  "title": "Adversarial validation improves cv-lb gap",
  "url": "/competitions/seti-breakthrough-listen/discussion/266652",
  "author_name": "imori",
  "post_date": "2021-08-19T21:41:33.846000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This competiton has CV-LB gap around 0.1 in spite of the model size.<br>\nex)</p>\n<ul>\n<li>5fold tf_efficientnetv2_m_in21ft1k 512x512   CV:0.8749  LB:0.7740  gap:0.1009</li>\n<li>5fold seresnext50_32x4d  512x512  CV:0.8619  LB:0.7620   gap:0.0999</li>\n</ul>\n<p>There are above gap, but I believed improving CV because CV is correlated to LB</p>\n<p>The day before deatline, I tried adversrial validation as soon as seeing this interesting dicussion <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265921\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265921</a>, and I succeeded in filling some cv-lb gap.</p>\n<p>Process (below process may contain mistake)</p>\n<ol>\n<li>label training data to 0 and test data to 1</li>\n<li>train efficientnet-b0 with above data (input size 256x256)</li>\n<li>cv around 0.995</li>\n<li>predict oof training data</li>\n<li>split training data to train and val using above probabilities<br>\nprob &gt;= 0.4  is valid<br>\nprob &lt; 0.4 is train<br>\nI implemented 3 fold cv.</li>\n</ol>\n<p>Result</p>\n<ul>\n<li>efficientnet-b1 (512x512) 15 epoch  CV:0.7845   LB:0.74556  PB :0.74587  gap:0.0389</li>\n<li>efficientnet-b0 (512x512) 15 epoch  CV:0.7842   LB 0.73481  PB:0.73305  gap:0.0494</li>\n<li>ensemble 2 models   LB:0.74650 PB:0.74674</li>\n</ul>\n<p>It looks like descreasing CV-LB gap.<br>\nThese are still some gap because test dataset contains some data appeared only test.</p>\n<p>I tried adversarial validation for the first time, but it is interesting technique.</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": 1482099,
      "postDate": "2021-08-19T21:41:33.847Z",
      "content": "<p>This competiton has CV-LB gap around 0.1 in spite of the model size.<br>\nex)</p>\n<ul>\n<li>5fold tf_efficientnetv2_m_in21ft1k 512x512   CV:0.8749  LB:0.7740  gap:0.1009</li>\n<li>5fold seresnext50_32x4d  512x512  CV:0.8619  LB:0.7620   gap:0.0999</li>\n</ul>\n<p>There are above gap, but I believed improving CV because CV is correlated to LB</p>\n<p>The day before deatline, I tried adversrial validation as soon as seeing this interesting dicussion <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265921\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265921</a>, and I succeeded in filling some cv-lb gap.</p>\n<p>Process (below process may contain mistake)</p>\n<ol>\n<li>label training data to 0 and test data to 1</li>\n<li>train efficientnet-b0 with above data (input size 256x256)</li>\n<li>cv around 0.995</li>\n<li>predict oof training data</li>\n<li>split training data to train and val using above probabilities<br>\nprob &gt;= 0.4  is valid<br>\nprob &lt; 0.4 is train<br>\nI implemented 3 fold cv.</li>\n</ol>\n<p>Result</p>\n<ul>\n<li>efficientnet-b1 (512x512) 15 epoch  CV:0.7845   LB:0.74556  PB :0.74587  gap:0.0389</li>\n<li>efficientnet-b0 (512x512) 15 epoch  CV:0.7842   LB 0.73481  PB:0.73305  gap:0.0494</li>\n<li>ensemble 2 models   LB:0.74650 PB:0.74674</li>\n</ul>\n<p>It looks like descreasing CV-LB gap.<br>\nThese are still some gap because test dataset contains some data appeared only test.</p>\n<p>I tried adversarial validation for the first time, but it is interesting technique.</p>\n<p>Thanks</p>",
      "rawMarkdown": "This competiton has CV-LB gap around 0.1 in spite of the model size.\nex)\n- 5fold tf_efficientnetv2_m_in21ft1k 512x512   CV:0.8749  LB:0.7740  gap:0.1009\n- 5fold seresnext50_32x4d  512x512  CV:0.8619  LB:0.7620   gap:0.0999\n\nThere are above gap, but I believed improving CV because CV is correlated to LB\n\nThe day before deatline, I tried adversrial validation as soon as seeing this interesting dicussion https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265921, and I succeeded in filling some cv-lb gap.\n\n\nProcess (below process may contain mistake)\n1. label training data to 0 and test data to 1\n2. train efficientnet-b0 with above data (input size 256x256)\n3. cv around 0.995\n4. predict oof training data\n5. split training data to train and val using above probabilities\n    prob >= 0.4  is valid\n    prob < 0.4 is train\n   I implemented 3 fold cv.\n\nResult\n- efficientnet-b1 (512x512) 15 epoch  CV:0.7845   LB:0.74556  PB :0.74587  gap:0.0389\n- efficientnet-b0 (512x512) 15 epoch  CV:0.7842   LB 0.73481  PB:0.73305  gap:0.0494\n- ensemble 2 models   LB:0.74650 PB:0.74674\n\nIt looks like descreasing CV-LB gap.\nThese are still some gap because test dataset contains some data appeared only test.\n\nI tried adversarial validation for the first time, but it is interesting technique.\n\nThanks",
      "votes": 3
    },
    {
      "id": 1482134,
      "postDate": "2021-08-19T22:20:25.157Z",
      "content": "<p>Very nice post, upvoted it!</p>",
      "rawMarkdown": "Very nice post, upvoted it!",
      "replies": [
        {
          "id": 1482151,
          "postDate": "2021-08-19T22:39:10.090Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1482134,
      "author_name": "Rishiraj Acharya",
      "author_url": "",
      "post_date": "2021-08-19T22:20:25.157000",
      "content": "<p>Very nice post, upvoted it!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1482151,
          "author_name": "imori",
          "author_url": "",
          "post_date": "2021-08-19T22:39:10.090000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1482099": "This competiton has CV-LB gap around 0.1 in spite of the model size.\nex)\n- 5fold tf_efficientnetv2_m_in21ft1k 512x512   CV:0.8749  LB:0.7740  gap:0.1009\n- 5fold seresnext50_32x4d  512x512  CV:0.8619  LB:0.7620   gap:0.0999\n\nThere are above gap, but I believed improving CV because CV is correlated to LB\n\nThe day before deatline, I tried adversrial validation as soon as seeing this interesting dicussion https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265921, and I succeeded in filling some cv-lb gap.\n\n\nProcess (below process may contain mistake)\n1. label training data to 0 and test data to 1\n2. train efficientnet-b0 with above data (input size 256x256)\n3. cv around 0.995\n4. predict oof training data\n5. split training data to train and val using above probabilities\n    prob >= 0.4  is valid\n    prob < 0.4 is train\n   I implemented 3 fold cv.\n\nResult\n- efficientnet-b1 (512x512) 15 epoch  CV:0.7845   LB:0.74556  PB :0.74587  gap:0.0389\n- efficientnet-b0 (512x512) 15 epoch  CV:0.7842   LB 0.73481  PB:0.73305  gap:0.0494\n- ensemble 2 models   LB:0.74650 PB:0.74674\n\nIt looks like descreasing CV-LB gap.\nThese are still some gap because test dataset contains some data appeared only test.\n\nI tried adversarial validation for the first time, but it is interesting technique.\n\nThanks",
    "1482134": "Very nice post, upvoted it!"
  }
}