{
  "id": 240500,
  "title": "Aren't People overfitting?",
  "url": "/competitions/seti-breakthrough-listen/discussion/240500",
  "author_name": "shiroe",
  "post_date": "2021-05-20T05:25:44.057000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Sorry if this is a silly question but how people high in LB making sure they are not just overfitting the test dataset</p>",
  "messages": [
    {
      "id": 1316142,
      "postDate": "2021-05-20T10:36:58.007Z",
      "content": "<p>I think it is too soon to say about overfitting as the competition just started. For now I see a correlation between my CV and LB scores.</p>",
      "rawMarkdown": "I think it is too soon to say about overfitting as the competition just started. For now I see a correlation between my CV and LB scores.",
      "votes": 3
    },
    {
      "id": 1315751,
      "postDate": "2021-05-20T05:28:39.313Z",
      "content": "<p>If your Cv is solid Then their is less chance of overfitting. </p>",
      "rawMarkdown": "If your Cv is solid Then their is less chance of overfitting. ",
      "votes": 2,
      "replies": [
        {
          "id": 1315778,
          "postDate": "2021-05-20T05:44:19.620Z",
          "content": "<p>I see, so they are believing in their cross-validation score </p>",
          "rawMarkdown": "I see, so they are believing in their cross-validation score "
        },
        {
          "id": 1315864,
          "postDate": "2021-05-20T06:43:04.830Z",
          "content": "<p>Yeah, although I just trained a mixup model with my best CV so far and it's dropped in LB from 0.98 to 0.97. I think it's reassuring when your CV and LB move in the right direction, although if you submit enough, you'll end up seeing even these movements by chance.</p>\n<p>Probably best not to submit every model and rely on CV/LB, as much as make adjustments you think make sense and then only submit the few with the best CVs.</p>",
          "rawMarkdown": "Yeah, although I just trained a mixup model with my best CV so far and it's dropped in LB from 0.98 to 0.97. I think it's reassuring when your CV and LB move in the right direction, although if you submit enough, you'll end up seeing even these movements by chance.\n\nProbably best not to submit every model and rely on CV/LB, as much as make adjustments you think make sense and then only submit the few with the best CVs.",
          "votes": 3
        },
        {
          "id": 1315879,
          "postDate": "2021-05-20T06:55:13.410Z",
          "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> Unrelated to This But how did you implement Mixup ?</p>",
          "rawMarkdown": "@jamesphoward Unrelated to This But how did you implement Mixup ?"
        },
        {
          "id": 1315885,
          "postDate": "2021-05-20T07:02:43.640Z",
          "content": "<p>I created my own one (I think) which worked on non-one-hot labels and uses the same default alpha as I found in the FastAI docs:</p>\n<pre><code>def mixup(x, y, alpha=0.4):\n    # A lot of people seem to use an alpha of 1 and then clip to e.g. 0.3 and 0.7.\n    # This is closer to what FastAI does and has more gentle mixing up\n    indices = torch.randperm(len(x))\n\n    x_shuffled = x[indices]\n    y_shuffled = y[indices]\n\n    lam = np.random.beta(alpha, alpha)\n    x = lam * x + (1 - lam) * x_shuffled\n    y = lam * y + (1 - lam) * y_shuffled\n\n    return x, y\n</code></pre>\n<p>Remember if you're using sklearn's AUC calculation for metrics you'll need to round() the y_true to the nearest integer or you'll get an error.</p>",
          "rawMarkdown": "I created my own one (I think) which worked on non-one-hot labels and uses the same default alpha as I found in the FastAI docs:\n\n```\ndef mixup(x, y, alpha=0.4):\n    # A lot of people seem to use an alpha of 1 and then clip to e.g. 0.3 and 0.7.\n    # This is closer to what FastAI does and has more gentle mixing up\n    indices = torch.randperm(len(x))\n\n    x_shuffled = x[indices]\n    y_shuffled = y[indices]\n\n    lam = np.random.beta(alpha, alpha)\n    x = lam * x + (1 - lam) * x_shuffled\n    y = lam * y + (1 - lam) * y_shuffled\n\n    return x, y\n```\n\nRemember if you're using sklearn's AUC calculation for metrics you'll need to round() the y_true to the nearest integer or you'll get an error.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1315747,
      "postDate": "2021-05-20T05:25:44.057Z",
      "content": "<p>Sorry if this is a silly question but how people high in LB making sure they are not just overfitting the test dataset</p>",
      "rawMarkdown": "Sorry if this is a silly question but how people high in LB making sure they are not just overfitting the test dataset"
    },
    {
      "id": 1316160,
      "postDate": "2021-05-20T10:50:57.987Z",
      "content": "<p>We should follow that 'IN CV WE TRUST' phrase, in my opinion :) All best!</p>",
      "rawMarkdown": "We should follow that 'IN CV WE TRUST' phrase, in my opinion :) All best!"
    }
  ],
  "comments": [
    {
      "id": 1316142,
      "author_name": "Oleg Panichev",
      "author_url": "",
      "post_date": "2021-05-20T10:36:58.007000",
      "content": "<p>I think it is too soon to say about overfitting as the competition just started. For now I see a correlation between my CV and LB scores.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1315751,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-05-20T05:28:39.313000",
      "content": "<p>If your Cv is solid Then their is less chance of overfitting. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1315778,
          "author_name": "shiroe",
          "author_url": "",
          "post_date": "2021-05-20T05:44:19.620000",
          "content": "<p>I see, so they are believing in their cross-validation score </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1315864,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2021-05-20T06:43:04.830000",
          "content": "<p>Yeah, although I just trained a mixup model with my best CV so far and it's dropped in LB from 0.98 to 0.97. I think it's reassuring when your CV and LB move in the right direction, although if you submit enough, you'll end up seeing even these movements by chance.</p>\n<p>Probably best not to submit every model and rely on CV/LB, as much as make adjustments you think make sense and then only submit the few with the best CVs.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1315879,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-05-20T06:55:13.410000",
          "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> Unrelated to This But how did you implement Mixup ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1315885,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2021-05-20T07:02:43.640000",
          "content": "<p>I created my own one (I think) which worked on non-one-hot labels and uses the same default alpha as I found in the FastAI docs:</p>\n<pre><code>def mixup(x, y, alpha=0.4):\n    # A lot of people seem to use an alpha of 1 and then clip to e.g. 0.3 and 0.7.\n    # This is closer to what FastAI does and has more gentle mixing up\n    indices = torch.randperm(len(x))\n\n    x_shuffled = x[indices]\n    y_shuffled = y[indices]\n\n    lam = np.random.beta(alpha, alpha)\n    x = lam * x + (1 - lam) * x_shuffled\n    y = lam * y + (1 - lam) * y_shuffled\n\n    return x, y\n</code></pre>\n<p>Remember if you're using sklearn's AUC calculation for metrics you'll need to round() the y_true to the nearest integer or you'll get an error.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1316160,
      "author_name": "s3nh",
      "author_url": "",
      "post_date": "2021-05-20T10:50:57.987000",
      "content": "<p>We should follow that 'IN CV WE TRUST' phrase, in my opinion :) All best!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1316142": "I think it is too soon to say about overfitting as the competition just started. For now I see a correlation between my CV and LB scores.",
    "1315751": "If your Cv is solid Then their is less chance of overfitting. ",
    "1315747": "Sorry if this is a silly question but how people high in LB making sure they are not just overfitting the test dataset",
    "1316160": "We should follow that 'IN CV WE TRUST' phrase, in my opinion :) All best!"
  }
}