{
  "id": 253227,
  "title": "Why is Everbody stuck at 75 ?",
  "url": "/competitions/seti-breakthrough-listen/discussion/253227",
  "author_name": "Mithil Salunkhe",
  "post_date": "2021-07-15T13:41:20.964000",
  "votes": 0,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Why is everybody stuck at AUC 75 right now ?</p>",
  "messages": [
    {
      "id": 1389398,
      "postDate": "2021-07-15T17:09:17.073Z",
      "content": "<p>There is now someone &gt; 75 :)</p>",
      "rawMarkdown": "There is now someone > 75 :)",
      "votes": 5,
      "replies": [
        {
          "id": 1389489,
          "postDate": "2021-07-15T18:43:20.823Z",
          "content": "<p>I was getting nervous ;)</p>",
          "rawMarkdown": "I was getting nervous ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1389174,
      "postDate": "2021-07-15T13:51:18.637Z",
      "content": "<p>metric.clip(0, 0.75) maybe? <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> 😂</p>\n<p>or maybe by chance, not many have submitted yet</p>",
      "rawMarkdown": "metric.clip(0, 0.75) maybe? @inversion 😂\n\nor maybe by chance, not many have submitted yet",
      "votes": 4,
      "replies": [
        {
          "id": 1389183,
          "postDate": "2021-07-15T14:01:13.850Z",
          "content": "<p>no way  Psi comments on  A topic started  by me 😂. I think it might be 2 option but who know </p>",
          "rawMarkdown": "no way  Psi comments on  A topic started  by me 😂. I think it might be 2 option but who know "
        },
        {
          "id": 1389238,
          "postDate": "2021-07-15T14:43:18.270Z",
          "content": "<p>unlikely a coincidence with now 6 identical scores and #1 and #2 staying in the same place (first one with identical score keeps the rank).</p>\n<p>though, i hope it's not another scoring issue</p>",
          "rawMarkdown": "unlikely a coincidence with now 6 identical scores and #1 and #2 staying in the same place (first one with identical score keeps the rank).\n\nthough, i hope it's not another scoring issue",
          "votes": 1
        },
        {
          "id": 1389243,
          "postDate": "2021-07-15T14:54:40.890Z",
          "content": "<p>This is the result of the best public kernel<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253233\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253233</a></p>",
          "rawMarkdown": "This is the result of the best public kernel\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/253233",
          "votes": 6
        },
        {
          "id": 1389246,
          "postDate": "2021-07-15T14:56:15.753Z",
          "content": "<p>Hope it just coincidence I don't want another competition  reset. </p>",
          "rawMarkdown": "Hope it just coincidence I don't want another competition  reset. "
        },
        {
          "id": 1389262,
          "postDate": "2021-07-15T15:03:49.353Z",
          "content": "<p>That makes sense <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> !</p>",
          "rawMarkdown": "That makes sense @vlomme !"
        },
        {
          "id": 1389295,
          "postDate": "2021-07-15T15:32:33.803Z",
          "content": "<p>good catch <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a>, hopefully, it's just that</p>",
          "rawMarkdown": "good catch @vlomme, hopefully, it's just that",
          "votes": 1
        },
        {
          "id": 1389373,
          "postDate": "2021-07-15T16:43:52.690Z",
          "content": "<p>fyi: I submitted one of my previously trained models (with old leaky data) that scored <code>LB 0.98x</code> before  - now <code>LB 0.750</code>  as well</p>",
          "rawMarkdown": "fyi: I submitted one of my previously trained models (with old leaky data) that scored `LB 0.98x` before  - now `LB 0.750`  as well"
        }
      ]
    },
    {
      "id": 1389568,
      "postDate": "2021-07-15T20:33:23.023Z",
      "content": "<p>Efficientnet b3 single fold,<br>\ncv: 0.885 <br>\nlb: 0.758</p>",
      "rawMarkdown": "Efficientnet b3 single fold,\ncv: 0.885 \nlb: 0.758",
      "votes": 1,
      "replies": [
        {
          "id": 1389629,
          "postDate": "2021-07-15T21:56:11.837Z",
          "content": "<p>Thanks for sharing! Interesting to see many people reporting a large CV-LB gap on the new data. </p>",
          "rawMarkdown": "Thanks for sharing! Interesting to see many people reporting a large CV-LB gap on the new data. ",
          "votes": 3
        },
        {
          "id": 1396386,
          "postDate": "2021-07-22T04:34:23.907Z",
          "content": "<p>Sorry if my understanding of CV and kfold cross validation is off, but if you have a single fold (i.e. K = 1), then that means you have no leftover data in your training set to make validation data. Therefore, CV doesn't apply when K = 1 ?? </p>",
          "rawMarkdown": "Sorry if my understanding of CV and kfold cross validation is off, but if you have a single fold (i.e. K = 1), then that means you have no leftover data in your training set to make validation data. Therefore, CV doesn't apply when K = 1 ?? "
        },
        {
          "id": 1396422,
          "postDate": "2021-07-22T05:30:13.240Z",
          "content": "<p>With single fold this is just like saying the training and predictions for just one of the 5 folds of K=5 since the training can take awhile so e.g. the results for Fold 0, rather than the average of all 5. </p>",
          "rawMarkdown": "With single fold this is just like saying the training and predictions for just one of the 5 folds of K=5 since the training can take awhile so e.g. the results for Fold 0, rather than the average of all 5. "
        }
      ]
    },
    {
      "id": 1389297,
      "postDate": "2021-07-15T15:35:39.020Z",
      "content": "<p>My singe fold score - 0.8684, LB - 0.748<br>\nI guess if average 4 folds, will get 0.75(most likely)<br>\nSo it can be some scoring issue actually</p>",
      "rawMarkdown": "My singe fold score - 0.8684, LB - 0.748\nI guess if average 4 folds, will get 0.75(most likely)\nSo it can be some scoring issue actually",
      "votes": 1
    },
    {
      "id": 1389185,
      "postDate": "2021-07-15T14:04:42.753Z",
      "content": "<p>Just guessing here, but maybe some people simply test one of the already available public notebooks for the first submission and it achieves 0.750 on the new data. </p>\n<p>In any case, it looks like the new data might be somewhat more difficult to classify. From a quick look, I am getting around 0.80 CV with a pipeline that was above 0.95 on the old data.</p>",
      "rawMarkdown": "Just guessing here, but maybe some people simply test one of the already available public notebooks for the first submission and it achieves 0.750 on the new data. \n\nIn any case, it looks like the new data might be somewhat more difficult to classify. From a quick look, I am getting around 0.80 CV with a pipeline that was above 0.95 on the old data.",
      "votes": 1,
      "replies": [
        {
          "id": 1389189,
          "postDate": "2021-07-15T14:08:30.947Z",
          "content": "<p>Same thing happening with me. earlier pipeline had .98 CV the same one with new data has .81 Cv. <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>  Any explanation </p>",
          "rawMarkdown": "Same thing happening with me. earlier pipeline had .98 CV the same one with new data has .81 Cv. @inversion  Any explanation "
        },
        {
          "id": 1389439,
          "postDate": "2021-07-15T17:46:57.600Z",
          "content": "<p>I believe that the previous models were leveraging the quantization leak under the hood, even though Kagglers may not have intended to utilize the leak. Lower scores should be expected now that we've (hopefully) fixed that leak.</p>",
          "rawMarkdown": "I believe that the previous models were leveraging the quantization leak under the hood, even though Kagglers may not have intended to utilize the leak. Lower scores should be expected now that we've (hopefully) fixed that leak.",
          "votes": 5
        },
        {
          "id": 1389550,
          "postDate": "2021-07-15T20:09:32.323Z",
          "content": "<p>To be fair, my EDA showed that the new data is much harder to classify (based on visual inspection).<br>\nThe old set had many more easy to spot artificial signals.<br>\njust three hard examples from the new set below:<br>\n<img src=\"https://i.imgur.com/ztoEWqZ.png\" alt=\"image1\"><br>\n<img src=\"https://i.imgur.com/O7mOvnX.png\" alt=\"image2\"><br>\n<img src=\"https://i.imgur.com/AjhbEbe.png\" alt=\"image3\"></p>",
          "rawMarkdown": "To be fair, my EDA showed that the new data is much harder to classify (based on visual inspection).\nThe old set had many more easy to spot artificial signals.\njust three hard examples from the new set below:\n![image1](https://i.imgur.com/ztoEWqZ.png)\n![image2](https://i.imgur.com/O7mOvnX.png)\n![image3](https://i.imgur.com/AjhbEbe.png)",
          "votes": 6
        },
        {
          "id": 1400178,
          "postDate": "2021-07-26T06:00:55.340Z",
          "content": "<p>You mean to say our neural networks were picking up on the timestamps of the .npy files even if we didn't intend them to?</p>",
          "rawMarkdown": "You mean to say our neural networks were picking up on the timestamps of the .npy files even if we didn't intend them to?"
        }
      ]
    },
    {
      "id": 1389692,
      "postDate": "2021-07-16T00:47:43.323Z",
      "content": "<p>Efficientnetv2_b1(5fold),<br>\ncv: 0.86362<br>\nlb: 0.756<br>\nwhat?Is there such a big gap between CV and LB?<br>\nIs it necessary to re analyze the new data? Just three weeks longer? Time is a little tight</p>",
      "rawMarkdown": "Efficientnetv2_b1(5fold),\ncv: 0.86362\nlb: 0.756\nwhat?Is there such a big gap between CV and LB?\nIs it necessary to re analyze the new data? Just three weeks longer? Time is a little tight",
      "votes": 2,
      "replies": [
        {
          "id": 1389732,
          "postDate": "2021-07-16T02:56:06.327Z",
          "content": "<p>Just wanted to ask if you are using Tensorflow how did you train on Efficientnetv2_b1 ?</p>",
          "rawMarkdown": "Just wanted to ask if you are using Tensorflow how did you train on Efficientnetv2_b1 ?"
        },
        {
          "id": 1389747,
          "postDate": "2021-07-16T03:33:17.663Z",
          "content": "<p>There are some code based on tensorflow/keras in the open source code. It seems that except for the imagenet weight of Efficientnetv2_b1, it is not clear where to obtain it. Changing the model is an easy job…</p>",
          "rawMarkdown": "There are some code based on tensorflow/keras in the open source code. It seems that except for the imagenet weight of Efficientnetv2_b1, it is not clear where to obtain it. Changing the model is an easy job..."
        }
      ]
    },
    {
      "id": 1389261,
      "postDate": "2021-07-15T15:02:52.813Z",
      "content": "<p>I think this can be a coincidence but the Cv score is not going beyond 81 so it might just be that the data is harder </p>",
      "rawMarkdown": "I think this can be a coincidence but the Cv score is not going beyond 81 so it might just be that the data is harder ",
      "votes": 2
    },
    {
      "id": 1389875,
      "postDate": "2021-07-16T07:01:53.640Z",
      "content": "<p>For comparison, a single fold 20 epochs trained using CPMP's How to train without leak, inference on new test data - <br>\nmodel trained  on leak data lb 0.741<br>\nretrained model on new data lb 0.749</p>\n<p>Not sure if it helps. Maybe need some new ideas for new train.</p>",
      "rawMarkdown": "For comparison, a single fold 20 epochs trained using CPMP's How to train without leak, inference on new test data - \nmodel trained  on leak data lb 0.741\nretrained model on new data lb 0.749\n\nNot sure if it helps. Maybe need some new ideas for new train."
    },
    {
      "id": 1389159,
      "postDate": "2021-07-15T13:41:20.963Z",
      "content": "<p>Why is everybody stuck at AUC 75 right now ?</p>",
      "rawMarkdown": "Why is everybody stuck at AUC 75 right now ?"
    },
    {
      "id": 1390521,
      "postDate": "2021-07-16T18:18:20.690Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true,
      "replies": [
        {
          "id": 1390535,
          "postDate": "2021-07-16T18:27:52.600Z",
          "content": "<p>metric is AUC, doesnt matter</p>",
          "rawMarkdown": "metric is AUC, doesnt matter",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1389398,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-07-15T17:09:17.073000",
      "content": "<p>There is now someone &gt; 75 :)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1389489,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-07-15T18:43:20.823000",
          "content": "<p>I was getting nervous ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1389174,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-07-15T13:51:18.637000",
      "content": "<p>metric.clip(0, 0.75) maybe? <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> 😂</p>\n<p>or maybe by chance, not many have submitted yet</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1389183,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-15T14:01:13.850000",
          "content": "<p>no way  Psi comments on  A topic started  by me 😂. I think it might be 2 option but who know </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1389238,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-07-15T14:43:18.270000",
          "content": "<p>unlikely a coincidence with now 6 identical scores and #1 and #2 staying in the same place (first one with identical score keeps the rank).</p>\n<p>though, i hope it's not another scoring issue</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1389243,
          "author_name": "Kramarenko Vladislav",
          "author_url": "",
          "post_date": "2021-07-15T14:54:40.890000",
          "content": "<p>This is the result of the best public kernel<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253233\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253233</a></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1389246,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-15T14:56:15.753000",
          "content": "<p>Hope it just coincidence I don't want another competition  reset. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1389262,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-07-15T15:03:49.353000",
          "content": "<p>That makes sense <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1389295,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-07-15T15:32:33.803000",
          "content": "<p>good catch <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a>, hopefully, it's just that</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1389373,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2021-07-15T16:43:52.690000",
          "content": "<p>fyi: I submitted one of my previously trained models (with old leaky data) that scored <code>LB 0.98x</code> before  - now <code>LB 0.750</code>  as well</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389568,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-07-15T20:33:23.023000",
      "content": "<p>Efficientnet b3 single fold,<br>\ncv: 0.885 <br>\nlb: 0.758</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1389629,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-07-15T21:56:11.837000",
          "content": "<p>Thanks for sharing! Interesting to see many people reporting a large CV-LB gap on the new data. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1396386,
          "author_name": "Erik Kaufman",
          "author_url": "",
          "post_date": "2021-07-22T04:34:23.907000",
          "content": "<p>Sorry if my understanding of CV and kfold cross validation is off, but if you have a single fold (i.e. K = 1), then that means you have no leftover data in your training set to make validation data. Therefore, CV doesn't apply when K = 1 ?? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1396422,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2021-07-22T05:30:13.240000",
          "content": "<p>With single fold this is just like saying the training and predictions for just one of the 5 folds of K=5 since the training can take awhile so e.g. the results for Fold 0, rather than the average of all 5. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389297,
      "author_name": "Ed Yanakov",
      "author_url": "",
      "post_date": "2021-07-15T15:35:39.020000",
      "content": "<p>My singe fold score - 0.8684, LB - 0.748<br>\nI guess if average 4 folds, will get 0.75(most likely)<br>\nSo it can be some scoring issue actually</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1389185,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2021-07-15T14:04:42.753000",
      "content": "<p>Just guessing here, but maybe some people simply test one of the already available public notebooks for the first submission and it achieves 0.750 on the new data. </p>\n<p>In any case, it looks like the new data might be somewhat more difficult to classify. From a quick look, I am getting around 0.80 CV with a pipeline that was above 0.95 on the old data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1389189,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-15T14:08:30.947000",
          "content": "<p>Same thing happening with me. earlier pipeline had .98 CV the same one with new data has .81 Cv. <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>  Any explanation </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1389439,
          "author_name": "Yuhong Chen",
          "author_url": "",
          "post_date": "2021-07-15T17:46:57.600000",
          "content": "<p>I believe that the previous models were leveraging the quantization leak under the hood, even though Kagglers may not have intended to utilize the leak. Lower scores should be expected now that we've (hopefully) fixed that leak.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1389550,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-07-15T20:09:32.323000",
          "content": "<p>To be fair, my EDA showed that the new data is much harder to classify (based on visual inspection).<br>\nThe old set had many more easy to spot artificial signals.<br>\njust three hard examples from the new set below:<br>\n<img src=\"https://i.imgur.com/ztoEWqZ.png\" alt=\"image1\"><br>\n<img src=\"https://i.imgur.com/O7mOvnX.png\" alt=\"image2\"><br>\n<img src=\"https://i.imgur.com/AjhbEbe.png\" alt=\"image3\"></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1400178,
          "author_name": "Erik Kaufman",
          "author_url": "",
          "post_date": "2021-07-26T06:00:55.340000",
          "content": "<p>You mean to say our neural networks were picking up on the timestamps of the .npy files even if we didn't intend them to?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389692,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-07-16T00:47:43.323000",
      "content": "<p>Efficientnetv2_b1(5fold),<br>\ncv: 0.86362<br>\nlb: 0.756<br>\nwhat?Is there such a big gap between CV and LB?<br>\nIs it necessary to re analyze the new data? Just three weeks longer? Time is a little tight</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1389732,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-16T02:56:06.327000",
          "content": "<p>Just wanted to ask if you are using Tensorflow how did you train on Efficientnetv2_b1 ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1389747,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-07-16T03:33:17.663000",
          "content": "<p>There are some code based on tensorflow/keras in the open source code. It seems that except for the imagenet weight of Efficientnetv2_b1, it is not clear where to obtain it. Changing the model is an easy job…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389261,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-07-15T15:02:52.813000",
      "content": "<p>I think this can be a coincidence but the Cv score is not going beyond 81 so it might just be that the data is harder </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1389875,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-07-16T07:01:53.640000",
      "content": "<p>For comparison, a single fold 20 epochs trained using CPMP's How to train without leak, inference on new test data - <br>\nmodel trained  on leak data lb 0.741<br>\nretrained model on new data lb 0.749</p>\n<p>Not sure if it helps. Maybe need some new ideas for new train.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1390521,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-16T18:18:20.690000",
      "content": "",
      "votes": -4,
      "replies": [
        {
          "id": 1390535,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-07-16T18:27:52.600000",
          "content": "<p>metric is AUC, doesnt matter</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1389398": "There is now someone > 75 :)",
    "1389174": "metric.clip(0, 0.75) maybe? @inversion 😂\n\nor maybe by chance, not many have submitted yet",
    "1389568": "Efficientnet b3 single fold,\ncv: 0.885 \nlb: 0.758",
    "1389297": "My singe fold score - 0.8684, LB - 0.748\nI guess if average 4 folds, will get 0.75(most likely)\nSo it can be some scoring issue actually",
    "1389185": "Just guessing here, but maybe some people simply test one of the already available public notebooks for the first submission and it achieves 0.750 on the new data. \n\nIn any case, it looks like the new data might be somewhat more difficult to classify. From a quick look, I am getting around 0.80 CV with a pipeline that was above 0.95 on the old data.",
    "1389692": "Efficientnetv2_b1(5fold),\ncv: 0.86362\nlb: 0.756\nwhat?Is there such a big gap between CV and LB?\nIs it necessary to re analyze the new data? Just three weeks longer? Time is a little tight",
    "1389261": "I think this can be a coincidence but the Cv score is not going beyond 81 so it might just be that the data is harder ",
    "1389875": "For comparison, a single fold 20 epochs trained using CPMP's How to train without leak, inference on new test data - \nmodel trained  on leak data lb 0.741\nretrained model on new data lb 0.749\n\nNot sure if it helps. Maybe need some new ideas for new train.",
    "1389159": "Why is everybody stuck at AUC 75 right now ?",
    "1390521": ""
  }
}