{
  "id": 109102,
  "title": "Validation strategy",
  "url": "/competitions/understanding_cloud_organization/discussion/109102",
  "author_name": "",
  "post_date": "2019-09-16T13:50:37.386431400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Let me start by saying that it's the first time in computer vision competitions for me.</p>\n\n<p>I'd like to ask about the validation strategy, which seems to be a harder matter that in tabular competitions. Running a full training with 0.8/0.2 train/test split, even using two cards takes hours to complete. So I can imagine that one could run the training phase 5 times to understand the model performance on the entire dataset, but is it efficient?</p>\n\n<p>Or maybe do you simply run one training and take the score on, for example, 20 % as the CV score?</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": "627865",
      "postDate": "09/16/2019 13:50:37",
      "content": "<p>Let me start by saying that it's the first time in computer vision competitions for me.</p>\n\n<p>I'd like to ask about the validation strategy, which seems to be a harder matter that in tabular competitions. Running a full training with 0.8/0.2 train/test split, even using two cards takes hours to complete. So I can imagine that one could run the training phase 5 times to understand the model performance on the entire dataset, but is it efficient?</p>\n\n<p>Or maybe do you simply run one training and take the score on, for example, 20 % as the CV score?</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "Let me start by saying that it's the first time in computer vision competitions for me.\n\nI'd like to ask about the validation strategy, which seems to be a harder matter that in tabular competitions. Running a full training with 0.8/0.2 train/test split, even using two cards takes hours to complete. So I can imagine that one could run the training phase 5 times to understand the model performance on the entire dataset, but is it efficient?\n\nOr maybe do you simply run one training and take the score on, for example, 20 % as the CV score?\n\nCheers",
      "votes": null
    },
    {
      "id": "628461",
      "postDate": "09/17/2019 11:27:33",
      "content": "<p>Obviously, there is a cost/accuracy tradeoff between simple holdout split and n-folds.\nSo, holdout can be enough at earlier stages, and you could switch to n-folds later.\nHoldout seems to be giving me quite good CV/LB correlation so far for this competition.</p>",
      "rawMarkdown": "Obviously, there is a cost/accuracy tradeoff between simple holdout split and n-folds.\nSo, holdout can be enough at earlier stages, and you could switch to n-folds later.\nHoldout seems to be giving me quite good CV/LB correlation so far for this competition.",
      "votes": null
    },
    {
      "id": "628470",
      "postDate": "09/17/2019 11:35:03",
      "content": "<p>I agree. For now I also use holdout.</p>",
      "rawMarkdown": "I agree. For now I also use holdout.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 628461,
      "author_name": "vzaguskin",
      "author_url": "",
      "post_date": "09/17/2019 11:27:33",
      "content": "<p>Obviously, there is a cost/accuracy tradeoff between simple holdout split and n-folds.\nSo, holdout can be enough at earlier stages, and you could switch to n-folds later.\nHoldout seems to be giving me quite good CV/LB correlation so far for this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 628470,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "09/17/2019 11:35:03",
          "content": "<p>I agree. For now I also use holdout.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "627865": "Let me start by saying that it's the first time in computer vision competitions for me.\n\nI'd like to ask about the validation strategy, which seems to be a harder matter that in tabular competitions. Running a full training with 0.8/0.2 train/test split, even using two cards takes hours to complete. So I can imagine that one could run the training phase 5 times to understand the model performance on the entire dataset, but is it efficient?\n\nOr maybe do you simply run one training and take the score on, for example, 20 % as the CV score?\n\nCheers",
    "628461": "Obviously, there is a cost/accuracy tradeoff between simple holdout split and n-folds.\nSo, holdout can be enough at earlier stages, and you could switch to n-folds later.\nHoldout seems to be giving me quite good CV/LB correlation so far for this competition.",
    "628470": "I agree. For now I also use holdout."
  },
  "source": "meta"
}