{
  "id": 245023,
  "title": "How to tackle overfitting? Am I doing it right?",
  "url": "/competitions/seti-breakthrough-listen/discussion/245023",
  "author_name": "gao-hongnan",
  "post_date": "2021-06-09T10:12:55.886000",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have had success with small models too like <code>efficientnetb0</code>. I stratified the folds into 5 and train on them. I observe there is variance in folds, especially in fold 5 (for my seed).</p>\n<p>The roc score is monotonously increasing. great? no, the validation loss seems to be rather unstable and can hover up and down. I call this fold my bad fold and I want to tune it's learning rate, or apply some more regularization to this particular fold.</p>\n<p>Now the question is, is it blatantly wrong to use different folds when inferencing? In my case, I trained model A for F1_A-F4_A (subscript to show its model A), they have good cv, and perform well on LB, but for F5, it is pulling the score down. I therefore retrain F5 using some tricks, almost everything is the same (image size etc) but I added some custom layers, some LR changes, I call this fold F5_M (M=modified) and then in my inference it becomes F1_A to F4_A and F5_M. My score improved by quite a bit. Is there any data leakage? For the record, I also tested F1_M - F4_M, not all of them perform as well as F1_A-F4_A</p>",
  "messages": [
    {
      "id": 1342272,
      "postDate": "2021-06-09T10:12:55.887Z",
      "content": "<p>I have had success with small models too like <code>efficientnetb0</code>. I stratified the folds into 5 and train on them. I observe there is variance in folds, especially in fold 5 (for my seed).</p>\n<p>The roc score is monotonously increasing. great? no, the validation loss seems to be rather unstable and can hover up and down. I call this fold my bad fold and I want to tune it's learning rate, or apply some more regularization to this particular fold.</p>\n<p>Now the question is, is it blatantly wrong to use different folds when inferencing? In my case, I trained model A for F1_A-F4_A (subscript to show its model A), they have good cv, and perform well on LB, but for F5, it is pulling the score down. I therefore retrain F5 using some tricks, almost everything is the same (image size etc) but I added some custom layers, some LR changes, I call this fold F5_M (M=modified) and then in my inference it becomes F1_A to F4_A and F5_M. My score improved by quite a bit. Is there any data leakage? For the record, I also tested F1_M - F4_M, not all of them perform as well as F1_A-F4_A</p>",
      "rawMarkdown": "I have had success with small models too like `efficientnetb0`. I stratified the folds into 5 and train on them. I observe there is variance in folds, especially in fold 5 (for my seed).\n\nThe roc score is monotonously increasing. great? no, the validation loss seems to be rather unstable and can hover up and down. I call this fold my bad fold and I want to tune it's learning rate, or apply some more regularization to this particular fold.\n\nNow the question is, is it blatantly wrong to use different folds when inferencing? In my case, I trained model A for F1_A-F4_A (subscript to show its model A), they have good cv, and perform well on LB, but for F5, it is pulling the score down. I therefore retrain F5 using some tricks, almost everything is the same (image size etc) but I added some custom layers, some LR changes, I call this fold F5_M (M=modified) and then in my inference it becomes F1_A to F4_A and F5_M. My score improved by quite a bit. Is there any data leakage? For the record, I also tested F1_M - F4_M, not all of them perform as well as F1_A-F4_A",
      "votes": 13
    },
    {
      "id": 1342386,
      "postDate": "2021-06-09T11:57:57.943Z",
      "content": "<p>My view is that if one is being pious, then if you consider that your CV is the test of the model there is a leak here. You've chosen model M over model A because you've already looked at the results of model A for fold 5 and seen that they aren't good. However, this is something that happens a lot, a kind of indirect model selection bias which means that we end up choosing the model that does best on our validation, whereas ideally we would do the model selection independently of the validation, or perhaps do it based on a previous internal validation set that is specificaly designed for parameter adjustment or model selection.</p>\n<p>However, it's equally possible to say that the real test set is the private LB, and that we can do whatever we like (provided we don't hack Kaggle or Berkeley's servers to steal the ground truth - or outer space truth - labels), since everything including your CV and the public LB is available for feedback, and we may be prepared take the risk of the overfitting that could happen if we rely on leaky feedback too much.</p>",
      "rawMarkdown": "My view is that if one is being pious, then if you consider that your CV is the test of the model there is a leak here. You've chosen model M over model A because you've already looked at the results of model A for fold 5 and seen that they aren't good. However, this is something that happens a lot, a kind of indirect model selection bias which means that we end up choosing the model that does best on our validation, whereas ideally we would do the model selection independently of the validation, or perhaps do it based on a previous internal validation set that is specificaly designed for parameter adjustment or model selection.\n\nHowever, it's equally possible to say that the real test set is the private LB, and that we can do whatever we like (provided we don't hack Kaggle or Berkeley's servers to steal the ground truth - or outer space truth - labels), since everything including your CV and the public LB is available for feedback, and we may be prepared take the risk of the overfitting that could happen if we rely on leaky feedback too much.",
      "votes": 5
    },
    {
      "id": 1342934,
      "postDate": "2021-06-09T20:48:32.487Z",
      "content": "<p>I am trying regularization by adding a hyper-parameter.<br>\nIt is also essential to balance the images after clustering them.</p>\n<p>The neural network model often makes mistakes with datasets which they haven't experienced. <br>\nWe have to increase a few classes with augmentation. <br>\nFor example, we can create a mesh with spec-augmentation, apply cut mix after resize. </p>\n<p>Also, in the fully connected layer, a lot of information duplicated. <br>\nIt is necessary to group the layer so that each layer has independent features.</p>",
      "rawMarkdown": "I am trying regularization by adding a hyper-parameter.\nIt is also essential to balance the images after clustering them.\n\nThe neural network model often makes mistakes with datasets which they haven't experienced. \nWe have to increase a few classes with augmentation. \nFor example, we can create a mesh with spec-augmentation, apply cut mix after resize. \n\nAlso, in the fully connected layer, a lot of information duplicated. \nIt is necessary to group the layer so that each layer has independent features.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1342386,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2021-06-09T11:57:57.943000",
      "content": "<p>My view is that if one is being pious, then if you consider that your CV is the test of the model there is a leak here. You've chosen model M over model A because you've already looked at the results of model A for fold 5 and seen that they aren't good. However, this is something that happens a lot, a kind of indirect model selection bias which means that we end up choosing the model that does best on our validation, whereas ideally we would do the model selection independently of the validation, or perhaps do it based on a previous internal validation set that is specificaly designed for parameter adjustment or model selection.</p>\n<p>However, it's equally possible to say that the real test set is the private LB, and that we can do whatever we like (provided we don't hack Kaggle or Berkeley's servers to steal the ground truth - or outer space truth - labels), since everything including your CV and the public LB is available for feedback, and we may be prepared take the risk of the overfitting that could happen if we rely on leaky feedback too much.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1342934,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-06-09T20:48:32.487000",
      "content": "<p>I am trying regularization by adding a hyper-parameter.<br>\nIt is also essential to balance the images after clustering them.</p>\n<p>The neural network model often makes mistakes with datasets which they haven't experienced. <br>\nWe have to increase a few classes with augmentation. <br>\nFor example, we can create a mesh with spec-augmentation, apply cut mix after resize. </p>\n<p>Also, in the fully connected layer, a lot of information duplicated. <br>\nIt is necessary to group the layer so that each layer has independent features.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1342272": "I have had success with small models too like `efficientnetb0`. I stratified the folds into 5 and train on them. I observe there is variance in folds, especially in fold 5 (for my seed).\n\nThe roc score is monotonously increasing. great? no, the validation loss seems to be rather unstable and can hover up and down. I call this fold my bad fold and I want to tune it's learning rate, or apply some more regularization to this particular fold.\n\nNow the question is, is it blatantly wrong to use different folds when inferencing? In my case, I trained model A for F1_A-F4_A (subscript to show its model A), they have good cv, and perform well on LB, but for F5, it is pulling the score down. I therefore retrain F5 using some tricks, almost everything is the same (image size etc) but I added some custom layers, some LR changes, I call this fold F5_M (M=modified) and then in my inference it becomes F1_A to F4_A and F5_M. My score improved by quite a bit. Is there any data leakage? For the record, I also tested F1_M - F4_M, not all of them perform as well as F1_A-F4_A",
    "1342386": "My view is that if one is being pious, then if you consider that your CV is the test of the model there is a leak here. You've chosen model M over model A because you've already looked at the results of model A for fold 5 and seen that they aren't good. However, this is something that happens a lot, a kind of indirect model selection bias which means that we end up choosing the model that does best on our validation, whereas ideally we would do the model selection independently of the validation, or perhaps do it based on a previous internal validation set that is specificaly designed for parameter adjustment or model selection.\n\nHowever, it's equally possible to say that the real test set is the private LB, and that we can do whatever we like (provided we don't hack Kaggle or Berkeley's servers to steal the ground truth - or outer space truth - labels), since everything including your CV and the public LB is available for feedback, and we may be prepared take the risk of the overfitting that could happen if we rely on leaky feedback too much.",
    "1342934": "I am trying regularization by adding a hyper-parameter.\nIt is also essential to balance the images after clustering them.\n\nThe neural network model often makes mistakes with datasets which they haven't experienced. \nWe have to increase a few classes with augmentation. \nFor example, we can create a mesh with spec-augmentation, apply cut mix after resize. \n\nAlso, in the fully connected layer, a lot of information duplicated. \nIt is necessary to group the layer so that each layer has independent features."
  }
}