{
  "id": 257417,
  "title": "What is a good CV Strategy here?",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/257417",
  "author_name": "Vincent Wang",
  "post_date": "2021-08-03T02:08:17.478000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I just came from the CommonLit contest and found the data size here is much larger than that and the model also takes quite long to train. So instead of 5 fold CV strategy there, I thought maybe we should do something different here: Like training 1 model using 80% train, 10% valid, 10% test?  I saw the popular notebook <a href=\"https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">G2Net / efficientnet_b7 / baseline [training]</a> use 80% to train and 20% to valid but I am not sure if that's the best practice. Thank you.</p>",
  "messages": [
    {
      "id": 1439478,
      "postDate": "2021-08-03T17:02:53.307Z",
      "content": "<p>I've conducted almost all of my experiments using 4 Fold CV at this point and gotten pretty close to the LB. My local CV Score was <code>0.865</code> and the LB score was <code>0.866</code>.</p>",
      "rawMarkdown": "I've conducted almost all of my experiments using 4 Fold CV at this point and gotten pretty close to the LB. My local CV Score was `0.865` and the LB score was `0.866`.",
      "votes": 1,
      "replies": [
        {
          "id": 1443148,
          "postDate": "2021-08-04T01:54:48.473Z",
          "content": "<p>Thank you Saurav for sharing</p>",
          "rawMarkdown": "Thank you Saurav for sharing"
        }
      ]
    },
    {
      "id": 1419168,
      "postDate": "2021-08-03T02:08:17.480Z",
      "content": "<p>I just came from the CommonLit contest and found the data size here is much larger than that and the model also takes quite long to train. So instead of 5 fold CV strategy there, I thought maybe we should do something different here: Like training 1 model using 80% train, 10% valid, 10% test?  I saw the popular notebook <a href=\"https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training\" target=\"_blank\">G2Net / efficientnet_b7 / baseline [training]</a> use 80% to train and 20% to valid but I am not sure if that's the best practice. Thank you.</p>",
      "rawMarkdown": "I just came from the CommonLit contest and found the data size here is much larger than that and the model also takes quite long to train. So instead of 5 fold CV strategy there, I thought maybe we should do something different here: Like training 1 model using 80% train, 10% valid, 10% test?  I saw the popular notebook [G2Net / efficientnet_b7 / baseline [training]](https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training) use 80% to train and 20% to valid but I am not sure if that's the best practice. Thank you.",
      "votes": 1
    },
    {
      "id": 1561298,
      "postDate": "2021-10-27T13:13:05.380Z",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
    },
    {
      "id": 1427758,
      "postDate": "2021-08-03T08:46:10.277Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1430404,
          "postDate": "2021-08-03T11:37:03.830Z",
          "content": "<p>Thank you for your reply! What do you mean by 50% discount and 7% online here? I may have missed something..</p>",
          "rawMarkdown": "Thank you for your reply! What do you mean by 50% discount and 7% online here? I may have missed something.."
        },
        {
          "id": 1431000,
          "postDate": "2021-08-03T12:03:49.007Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1448220,
          "postDate": "2021-08-04T16:57:14.197Z",
          "content": "<p>I tried using a 5 fold CV with <strong><code>EfficientNetB7</code></strong> but ran out of the notebook time limit after training for a single fold. But even with this model, I was able to get my current best score of <code>0.866</code></p>",
          "rawMarkdown": "I tried using a 5 fold CV with **`EfficientNetB7`** but ran out of the notebook time limit after training for a single fold. But even with this model, I was able to get my current best score of `0.866`"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1439478,
      "author_name": "Saurav Maheshkar ☕️",
      "author_url": "",
      "post_date": "2021-08-03T17:02:53.307000",
      "content": "<p>I've conducted almost all of my experiments using 4 Fold CV at this point and gotten pretty close to the LB. My local CV Score was <code>0.865</code> and the LB score was <code>0.866</code>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1443148,
          "author_name": "Vincent Wang",
          "author_url": "",
          "post_date": "2021-08-04T01:54:48.473000",
          "content": "<p>Thank you Saurav for sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1561298,
      "author_name": "ChristopherZerafa",
      "author_url": "",
      "post_date": "2021-10-27T13:13:05.380000",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1427758,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-03T08:46:10.277000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1430404,
          "author_name": "Vincent Wang",
          "author_url": "",
          "post_date": "2021-08-03T11:37:03.830000",
          "content": "<p>Thank you for your reply! What do you mean by 50% discount and 7% online here? I may have missed something..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1431000,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-03T12:03:49.007000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1448220,
          "author_name": "Saurav Maheshkar ☕️",
          "author_url": "",
          "post_date": "2021-08-04T16:57:14.197000",
          "content": "<p>I tried using a 5 fold CV with <strong><code>EfficientNetB7</code></strong> but ran out of the notebook time limit after training for a single fold. But even with this model, I was able to get my current best score of <code>0.866</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1439478": "I've conducted almost all of my experiments using 4 Fold CV at this point and gotten pretty close to the LB. My local CV Score was `0.865` and the LB score was `0.866`.",
    "1419168": "I just came from the CommonLit contest and found the data size here is much larger than that and the model also takes quite long to train. So instead of 5 fold CV strategy there, I thought maybe we should do something different here: Like training 1 model using 80% train, 10% valid, 10% test?  I saw the popular notebook [G2Net / efficientnet_b7 / baseline [training]](https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-training) use 80% to train and 20% to valid but I am not sure if that's the best practice. Thank you.",
    "1561298": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1427758": ""
  }
}