{
  "id": 243013,
  "title": "Overfitting analysis",
  "url": "/competitions/seti-breakthrough-listen/discussion/243013",
  "author_name": "FelipeKitamura, MD, PhD",
  "post_date": "2021-05-31T23:31:02.875000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>To <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> and/or any Kaggle staff,</p>\n<p>We know shake-up happens in every Kaggle competition (to a greater or lesser extent) and it is something expected to happen.</p>\n<p>I wonder if anyone has already analyzed if there is any correlation between the number of submissions from each team and the increase/decrease in rank from public LB to private LB.</p>\n<p>My point is that maybe submitting a lot increases the risk of overfitting the public LB. I wanted to know if the real-world data supports that hypothesis.</p>",
  "messages": [
    {
      "id": 1335958,
      "postDate": "2021-06-04T14:45:35.317Z",
      "content": "<p>We do not think that only looking at total number of submissions is correlated with overfitting. Relying only on public leaderboard score alone may lead to this but each competition has many variables so it is hard to generalize.  </p>",
      "rawMarkdown": "We do not think that only looking at total number of submissions is correlated with overfitting. Relying only on public leaderboard score alone may lead to this but each competition has many variables so it is hard to generalize.  ",
      "votes": 5,
      "replies": [
        {
          "id": 1335977,
          "postDate": "2021-06-04T15:01:18.413Z",
          "content": "<p>Great point!</p>",
          "rawMarkdown": "Great point!"
        }
      ]
    },
    {
      "id": 1331290,
      "postDate": "2021-06-01T11:34:57.723Z",
      "content": "<p>depends on the competition and how much sense you make out of LB feedback</p>",
      "rawMarkdown": "depends on the competition and how much sense you make out of LB feedback",
      "votes": 3
    },
    {
      "id": 1330858,
      "postDate": "2021-06-01T06:15:14.897Z",
      "content": "<p>You can test your hypothesis using kaggle <a href=\"https://www.kaggle.com/kaggle/meta-kaggle\" target=\"_blank\">meta-data</a> and raw data from each competition leaderboard.</p>",
      "rawMarkdown": "You can test your hypothesis using kaggle [meta-data](https://www.kaggle.com/kaggle/meta-kaggle) and raw data from each competition leaderboard.",
      "votes": 1,
      "replies": [
        {
          "id": 1331166,
          "postDate": "2021-06-01T10:03:43.853Z",
          "content": "<p>Thanks for letting me know. </p>",
          "rawMarkdown": "Thanks for letting me know. "
        }
      ]
    },
    {
      "id": 1330721,
      "postDate": "2021-06-01T04:19:26.893Z",
      "content": "<p>Interesting project for someone to take on - but think there is a lot of noise in the data.  </p>\n<p>A more likely correlation for over-fitting would evaluate experience vs LB shakeup.  </p>\n<p>A team of Kaggle GM's might use all submissions every day and be exploring a very wide range of experiments.  A rookie might use all submissions every day and be doing a one factor at a time experiments using the LB as a guide.</p>\n<p>Speaking from my own experience - I jumped up 1000 places in a competition where I only had 3 submissions.   In my very first competition I had domain knowledge up the rear and had a gold winner on day 2 with my 5th submission.  A couple of hundred submissions later I was down several thousand at 3198 - over fit like crazy believing the LB.  </p>",
      "rawMarkdown": "Interesting project for someone to take on - but think there is a lot of noise in the data.  \n\nA more likely correlation for over-fitting would evaluate experience vs LB shakeup.  \n\nA team of Kaggle GM's might use all submissions every day and be exploring a very wide range of experiments.  A rookie might use all submissions every day and be doing a one factor at a time experiments using the LB as a guide.\n\nSpeaking from my own experience - I jumped up 1000 places in a competition where I only had 3 submissions.   In my very first competition I had domain knowledge up the rear and had a gold winner on day 2 with my 5th submission.  A couple of hundred submissions later I was down several thousand at 3198 - over fit like crazy believing the LB.  ",
      "votes": 1,
      "replies": [
        {
          "id": 1331170,
          "postDate": "2021-06-01T10:07:22.057Z",
          "content": "<p>Thanks for sharing your experience. Quite interesting. I also had the two kinds of surprise. Jumping up and winning gold in one competition. But plummeting in another. </p>",
          "rawMarkdown": "Thanks for sharing your experience. Quite interesting. I also had the two kinds of surprise. Jumping up and winning gold in one competition. But plummeting in another. "
        }
      ]
    },
    {
      "id": 1330526,
      "postDate": "2021-05-31T23:31:02.877Z",
      "content": "<p>To <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> and/or any Kaggle staff,</p>\n<p>We know shake-up happens in every Kaggle competition (to a greater or lesser extent) and it is something expected to happen.</p>\n<p>I wonder if anyone has already analyzed if there is any correlation between the number of submissions from each team and the increase/decrease in rank from public LB to private LB.</p>\n<p>My point is that maybe submitting a lot increases the risk of overfitting the public LB. I wanted to know if the real-world data supports that hypothesis.</p>",
      "rawMarkdown": "To @maggiemd and/or any Kaggle staff,\n\nWe know shake-up happens in every Kaggle competition (to a greater or lesser extent) and it is something expected to happen.\n\nI wonder if anyone has already analyzed if there is any correlation between the number of submissions from each team and the increase/decrease in rank from public LB to private LB.\n\nMy point is that maybe submitting a lot increases the risk of overfitting the public LB. I wanted to know if the real-world data supports that hypothesis."
    }
  ],
  "comments": [
    {
      "id": 1335958,
      "author_name": "Maggie",
      "author_url": "",
      "post_date": "2021-06-04T14:45:35.317000",
      "content": "<p>We do not think that only looking at total number of submissions is correlated with overfitting. Relying only on public leaderboard score alone may lead to this but each competition has many variables so it is hard to generalize.  </p>",
      "votes": 5,
      "replies": [
        {
          "id": 1335977,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2021-06-04T15:01:18.413000",
          "content": "<p>Great point!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1331290,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-06-01T11:34:57.723000",
      "content": "<p>depends on the competition and how much sense you make out of LB feedback</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1330858,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-06-01T06:15:14.897000",
      "content": "<p>You can test your hypothesis using kaggle <a href=\"https://www.kaggle.com/kaggle/meta-kaggle\" target=\"_blank\">meta-data</a> and raw data from each competition leaderboard.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1331166,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2021-06-01T10:03:43.853000",
          "content": "<p>Thanks for letting me know. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1330721,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-06-01T04:19:26.893000",
      "content": "<p>Interesting project for someone to take on - but think there is a lot of noise in the data.  </p>\n<p>A more likely correlation for over-fitting would evaluate experience vs LB shakeup.  </p>\n<p>A team of Kaggle GM's might use all submissions every day and be exploring a very wide range of experiments.  A rookie might use all submissions every day and be doing a one factor at a time experiments using the LB as a guide.</p>\n<p>Speaking from my own experience - I jumped up 1000 places in a competition where I only had 3 submissions.   In my very first competition I had domain knowledge up the rear and had a gold winner on day 2 with my 5th submission.  A couple of hundred submissions later I was down several thousand at 3198 - over fit like crazy believing the LB.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1331170,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2021-06-01T10:07:22.057000",
          "content": "<p>Thanks for sharing your experience. Quite interesting. I also had the two kinds of surprise. Jumping up and winning gold in one competition. But plummeting in another. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1335958": "We do not think that only looking at total number of submissions is correlated with overfitting. Relying only on public leaderboard score alone may lead to this but each competition has many variables so it is hard to generalize.  ",
    "1331290": "depends on the competition and how much sense you make out of LB feedback",
    "1330858": "You can test your hypothesis using kaggle [meta-data](https://www.kaggle.com/kaggle/meta-kaggle) and raw data from each competition leaderboard.",
    "1330721": "Interesting project for someone to take on - but think there is a lot of noise in the data.  \n\nA more likely correlation for over-fitting would evaluate experience vs LB shakeup.  \n\nA team of Kaggle GM's might use all submissions every day and be exploring a very wide range of experiments.  A rookie might use all submissions every day and be doing a one factor at a time experiments using the LB as a guide.\n\nSpeaking from my own experience - I jumped up 1000 places in a competition where I only had 3 submissions.   In my very first competition I had domain knowledge up the rear and had a gold winner on day 2 with my 5th submission.  A couple of hundred submissions later I was down several thousand at 3198 - over fit like crazy believing the LB.  ",
    "1330526": "To @maggiemd and/or any Kaggle staff,\n\nWe know shake-up happens in every Kaggle competition (to a greater or lesser extent) and it is something expected to happen.\n\nI wonder if anyone has already analyzed if there is any correlation between the number of submissions from each team and the increase/decrease in rank from public LB to private LB.\n\nMy point is that maybe submitting a lot increases the risk of overfitting the public LB. I wanted to know if the real-world data supports that hypothesis."
  }
}