{
  "id": 279841,
  "title": "Stratified Sampling into Training and Validation set",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/279841",
  "author_name": "Andreas Horlbeck",
  "post_date": "2021-10-19T09:56:21.498000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi first of all, congrats to all of you that had fun in the competition and learned something, that are the most important things..;) and of course congrats to all winners.</p>\n<p>While approaching this classification task, i thought it would be very important to divide the training set into train and validation set along the patients, so that a patient is only found in the train or test set - not both. Because if not there were images of one patient in both sets and so clearly the patients information could be extracted in the training and validation process -  on other images found in the validation set, that should be very similar to those in the train set - causing a lot of overfitting (which is not noticable on the validation set, but later in the test set)  but very good results, so no overfitting seen in the validation set.</p>\n<p>Then I had AUCs about 0.7 and thats why I thoguht I should discard this approch and spread one patients images over the train and validation set - motivated by the good LB results.<br>\nI think this was the problem.</p>\n<p>How have you dealt with this issue?<br>\nAre you using another houldout set, cross validation?<br>\nAnd have you divided the patients in train and validation set and didn't mix it up?</p>",
  "messages": [
    {
      "id": 1549933,
      "postDate": "2021-10-19T09:56:21.497Z",
      "content": "<p>Hi first of all, congrats to all of you that had fun in the competition and learned something, that are the most important things..;) and of course congrats to all winners.</p>\n<p>While approaching this classification task, i thought it would be very important to divide the training set into train and validation set along the patients, so that a patient is only found in the train or test set - not both. Because if not there were images of one patient in both sets and so clearly the patients information could be extracted in the training and validation process -  on other images found in the validation set, that should be very similar to those in the train set - causing a lot of overfitting (which is not noticable on the validation set, but later in the test set)  but very good results, so no overfitting seen in the validation set.</p>\n<p>Then I had AUCs about 0.7 and thats why I thoguht I should discard this approch and spread one patients images over the train and validation set - motivated by the good LB results.<br>\nI think this was the problem.</p>\n<p>How have you dealt with this issue?<br>\nAre you using another houldout set, cross validation?<br>\nAnd have you divided the patients in train and validation set and didn't mix it up?</p>",
      "rawMarkdown": "Hi first of all, congrats to all of you that had fun in the competition and learned something, that are the most important things..;) and of course congrats to all winners.\n\nWhile approaching this classification task, i thought it would be very important to divide the training set into train and validation set along the patients, so that a patient is only found in the train or test set - not both. Because if not there were images of one patient in both sets and so clearly the patients information could be extracted in the training and validation process -  on other images found in the validation set, that should be very similar to those in the train set - causing a lot of overfitting (which is not noticable on the validation set, but later in the test set)  but very good results, so no overfitting seen in the validation set.\n\nThen I had AUCs about 0.7 and thats why I thoguht I should discard this approch and spread one patients images over the train and validation set - motivated by the good LB results.\nI think this was the problem.\n\nHow have you dealt with this issue?\nAre you using another houldout set, cross validation?\nAnd have you divided the patients in train and validation set and didn't mix it up?\n\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1549933": "Hi first of all, congrats to all of you that had fun in the competition and learned something, that are the most important things..;) and of course congrats to all winners.\n\nWhile approaching this classification task, i thought it would be very important to divide the training set into train and validation set along the patients, so that a patient is only found in the train or test set - not both. Because if not there were images of one patient in both sets and so clearly the patients information could be extracted in the training and validation process -  on other images found in the validation set, that should be very similar to those in the train set - causing a lot of overfitting (which is not noticable on the validation set, but later in the test set)  but very good results, so no overfitting seen in the validation set.\n\nThen I had AUCs about 0.7 and thats why I thoguht I should discard this approch and spread one patients images over the train and validation set - motivated by the good LB results.\nI think this was the problem.\n\nHow have you dealt with this issue?\nAre you using another houldout set, cross validation?\nAnd have you divided the patients in train and validation set and didn't mix it up?\n\n"
  }
}