{
  "id": 243034,
  "title": "why AUC similar, but LB different?",
  "url": "/competitions/seti-breakthrough-listen/discussion/243034",
  "author_name": "dragon zhang",
  "post_date": "2021-06-01T02:04:09.810000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>different efficientnet models, their AUC look similar, but LB differ above 0.01 more?</p>",
  "messages": [
    {
      "id": 1330604,
      "postDate": "2021-06-01T02:04:09.810Z",
      "content": "<p>different efficientnet models, their AUC look similar, but LB differ above 0.01 more?</p>",
      "rawMarkdown": "different efficientnet models, their AUC look similar, but LB differ above 0.01 more?",
      "votes": 1
    },
    {
      "id": 1330706,
      "postDate": "2021-06-01T04:06:14.010Z",
      "content": "<p>LB score based on 20% of 35, 847 samples or around 6000 samples.  Depending on how you do your training auc it has around 50,000 samples.  </p>\n<p>It's easy to fall in the trap of over-fitting when you believe the LB more than you believe your local auc.  All the things taught in basic statistics class do not go away just because we are using ML methods.  Uncertainty between populations is ruled by sample size.</p>",
      "rawMarkdown": "LB score based on 20% of 35, 847 samples or around 6000 samples.  Depending on how you do your training auc it has around 50,000 samples.  \n\nIt's easy to fall in the trap of over-fitting when you believe the LB more than you believe your local auc.  All the things taught in basic statistics class do not go away just because we are using ML methods.  Uncertainty between populations is ruled by sample size.",
      "votes": 2,
      "replies": [
        {
          "id": 1333112,
          "postDate": "2021-06-02T13:48:21.990Z",
          "content": "<p>I have built a model with CV = 0.978 when it is trained in 12 epochs. Then, I reset and trained the same model 12 times. All the results were in the interval [0.972, 0.981] (dispersion less than 0.01)… The LB I get for this model was 0.96</p>\n<p>My validation dataset is 15% of the train samples (about 7500 samples), a similar size to the 20% of Test samples used in LB. </p>\n<p>The difference between CV and LB is usually greater than 0.01, wich can be in the range of the statistic dispersion yet. </p>\n<p>But, if the test samples are representative of the train samples (beacuse train and test datasets were  created randomnly from the same original dataset), then,  I think the probability to pick these 20% of test samples between the least respresentative in order to achieve this gap is really low. </p>\n<p>It would be interesting to know if someone has get sometime a LB greater than the local score</p>",
          "rawMarkdown": "I have built a model with CV = 0.978 when it is trained in 12 epochs. Then, I reset and trained the same model 12 times. All the results were in the interval [0.972, 0.981] (dispersion less than 0.01)... The LB I get for this model was 0.96\n\nMy validation dataset is 15% of the train samples (about 7500 samples), a similar size to the 20% of Test samples used in LB. \n\nThe difference between CV and LB is usually greater than 0.01, wich can be in the range of the statistic dispersion yet. \n\nBut, if the test samples are representative of the train samples (beacuse train and test datasets were  created randomnly from the same original dataset), then,  I think the probability to pick these 20% of test samples between the least respresentative in order to achieve this gap is really low. \n\nIt would be interesting to know if someone has get sometime a LB greater than the local score\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1332323,
      "postDate": "2021-06-02T04:20:30.267Z",
      "content": "<p>The main point is to look at the data distribution of the training set and the test set. If they are similar, then we can roughly estimate the correlation between LB and CV. Then we can look at the distribution of 20% of the test set and 80% of the other test sets. If 80% has no special characteristics, then we can see that the distribution of 20% and 80% of the test sets is similar. Therefore, based on the above, we can judge that there is not much difference between LB and Pb scores, On the contrary, there will be huge fluctuations</p>",
      "rawMarkdown": "The main point is to look at the data distribution of the training set and the test set. If they are similar, then we can roughly estimate the correlation between LB and CV. Then we can look at the distribution of 20% of the test set and 80% of the other test sets. If 80% has no special characteristics, then we can see that the distribution of 20% and 80% of the test sets is similar. Therefore, based on the above, we can judge that there is not much difference between LB and Pb scores, On the contrary, there will be huge fluctuations"
    }
  ],
  "comments": [
    {
      "id": 1330706,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-06-01T04:06:14.010000",
      "content": "<p>LB score based on 20% of 35, 847 samples or around 6000 samples.  Depending on how you do your training auc it has around 50,000 samples.  </p>\n<p>It's easy to fall in the trap of over-fitting when you believe the LB more than you believe your local auc.  All the things taught in basic statistics class do not go away just because we are using ML methods.  Uncertainty between populations is ruled by sample size.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1333112,
          "author_name": "ePolaris",
          "author_url": "",
          "post_date": "2021-06-02T13:48:21.990000",
          "content": "<p>I have built a model with CV = 0.978 when it is trained in 12 epochs. Then, I reset and trained the same model 12 times. All the results were in the interval [0.972, 0.981] (dispersion less than 0.01)… The LB I get for this model was 0.96</p>\n<p>My validation dataset is 15% of the train samples (about 7500 samples), a similar size to the 20% of Test samples used in LB. </p>\n<p>The difference between CV and LB is usually greater than 0.01, wich can be in the range of the statistic dispersion yet. </p>\n<p>But, if the test samples are representative of the train samples (beacuse train and test datasets were  created randomnly from the same original dataset), then,  I think the probability to pick these 20% of test samples between the least respresentative in order to achieve this gap is really low. </p>\n<p>It would be interesting to know if someone has get sometime a LB greater than the local score</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1332323,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-06-02T04:20:30.267000",
      "content": "<p>The main point is to look at the data distribution of the training set and the test set. If they are similar, then we can roughly estimate the correlation between LB and CV. Then we can look at the distribution of 20% of the test set and 80% of the other test sets. If 80% has no special characteristics, then we can see that the distribution of 20% and 80% of the test sets is similar. Therefore, based on the above, we can judge that there is not much difference between LB and Pb scores, On the contrary, there will be huge fluctuations</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1330604": "different efficientnet models, their AUC look similar, but LB differ above 0.01 more?",
    "1330706": "LB score based on 20% of 35, 847 samples or around 6000 samples.  Depending on how you do your training auc it has around 50,000 samples.  \n\nIt's easy to fall in the trap of over-fitting when you believe the LB more than you believe your local auc.  All the things taught in basic statistics class do not go away just because we are using ML methods.  Uncertainty between populations is ruled by sample size.",
    "1332323": "The main point is to look at the data distribution of the training set and the test set. If they are similar, then we can roughly estimate the correlation between LB and CV. Then we can look at the distribution of 20% of the test set and 80% of the other test sets. If 80% has no special characteristics, then we can see that the distribution of 20% and 80% of the test sets is similar. Therefore, based on the above, we can judge that there is not much difference between LB and Pb scores, On the contrary, there will be huge fluctuations"
  }
}