{
  "id": 508966,
  "title": "Some general questions about stratified k-fold",
  "url": "/competitions/birdclef-2024/discussion/508966",
  "author_name": "SakuraYuyuko",
  "post_date": "2024-05-31T15:04:30.202000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi guys, I'm a kaggle rookie and thanks for help. Here are some context:<br>\nI trained 5 models using 5 folds without any data argumentation. From the CV results, fold 1 works best for me, so I submit that model to kaggle and get LB 0.64 (amost 0.65 because I'm now the top1 in 0.64 range). </p>\n<h4>Question 1</h4>\n<p>Should I submit the rest of folds to kaggle because the gap between CV and LB is huge? The CV between folds are very similar (I guess? about 0.95-0.97 auc).</p>\n<h4>Question 2</h4>\n<p>I'm planning to add argumentations (like xy_mask). <br>\nShould I retrain all 5 folds using that argumentation or I should use the fold with highest CV/LB and build on that?</p>\n<p>Any help will be appreciated! XD</p>",
  "messages": [
    {
      "id": 2847503,
      "postDate": "2024-05-31T15:04:30.203Z",
      "content": "<p>Hi guys, I'm a kaggle rookie and thanks for help. Here are some context:<br>\nI trained 5 models using 5 folds without any data argumentation. From the CV results, fold 1 works best for me, so I submit that model to kaggle and get LB 0.64 (amost 0.65 because I'm now the top1 in 0.64 range). </p>\n<h4>Question 1</h4>\n<p>Should I submit the rest of folds to kaggle because the gap between CV and LB is huge? The CV between folds are very similar (I guess? about 0.95-0.97 auc).</p>\n<h4>Question 2</h4>\n<p>I'm planning to add argumentations (like xy_mask). <br>\nShould I retrain all 5 folds using that argumentation or I should use the fold with highest CV/LB and build on that?</p>\n<p>Any help will be appreciated! XD</p>",
      "rawMarkdown": "Hi guys, I'm a kaggle rookie and thanks for help. Here are some context:\nI trained 5 models using 5 folds without any data argumentation. From the CV results, fold 1 works best for me, so I submit that model to kaggle and get LB 0.64 (amost 0.65 because I'm now the top1 in 0.64 range). \n\n\n#### Question 1\nShould I submit the rest of folds to kaggle because the gap between CV and LB is huge? The CV between folds are very similar (I guess? about 0.95-0.97 auc).\n\n#### Question 2\nI'm planning to add argumentations (like xy_mask). \nShould I retrain all 5 folds using that argumentation or I should use the fold with highest CV/LB and build on that?\n\nAny help will be appreciated! XD",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2847503": "Hi guys, I'm a kaggle rookie and thanks for help. Here are some context:\nI trained 5 models using 5 folds without any data argumentation. From the CV results, fold 1 works best for me, so I submit that model to kaggle and get LB 0.64 (amost 0.65 because I'm now the top1 in 0.64 range). \n\n\n#### Question 1\nShould I submit the rest of folds to kaggle because the gap between CV and LB is huge? The CV between folds are very similar (I guess? about 0.95-0.97 auc).\n\n#### Question 2\nI'm planning to add argumentations (like xy_mask). \nShould I retrain all 5 folds using that argumentation or I should use the fold with highest CV/LB and build on that?\n\nAny help will be appreciated! XD"
  }
}