{
  "id": 271201,
  "title": "Mind the gap",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/271201",
  "author_name": "",
  "post_date": "2021-09-09T05:20:36.143290900Z",
  "votes": 10,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I observed that LB and CV has a relatively large gap(30~50bps) for most people's models(as appeared in discussion threads). Can anyone give some reasonable explanations? </p>\n<p>Furthermore, our team is working on a 1D model. A very weird thing is that our 1D model somehow only has a gap 10bps or less most of the time, while our 2D model frequently has a ~20-30 bps gap. These two kinds of models share the same preprocessing(high low pass and whiten). This is very confusing to me and I hope people can share some insights about this(and I can increase the gap without overfitting to LB :P   ) Thanks!</p>",
  "messages": [
    {
      "id": "1507328",
      "postDate": "09/09/2021 05:20:36",
      "content": "<p>I observed that LB and CV has a relatively large gap(30~50bps) for most people's models(as appeared in discussion threads). Can anyone give some reasonable explanations? </p>\n<p>Furthermore, our team is working on a 1D model. A very weird thing is that our 1D model somehow only has a gap 10bps or less most of the time, while our 2D model frequently has a ~20-30 bps gap. These two kinds of models share the same preprocessing(high low pass and whiten). This is very confusing to me and I hope people can share some insights about this(and I can increase the gap without overfitting to LB :P   ) Thanks!</p>",
      "rawMarkdown": "I observed that LB and CV has a relatively large gap(30~50bps) for most people's models(as appeared in discussion threads). Can anyone give some reasonable explanations? \n\nFurthermore, our team is working on a 1D model. A very weird thing is that our 1D model somehow only has a gap 10bps or less most of the time, while our 2D model frequently has a ~20-30 bps gap. These two kinds of models share the same preprocessing(high low pass and whiten). This is very confusing to me and I hope people can share some insights about this(and I can increase the gap without overfitting to LB :P   ) Thanks!",
      "votes": null
    },
    {
      "id": "1507334",
      "postDate": "09/09/2021 05:25:44",
      "content": "<p>What do you mean by BPS ?</p>",
      "rawMarkdown": "What do you mean by BPS ?",
      "votes": null
    },
    {
      "id": "1507335",
      "postDate": "09/09/2021 05:27:03",
      "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>  0.0001 is one bp</p>",
      "rawMarkdown": "mithilsalunkhe  0.0001 is one bp",
      "votes": null
    },
    {
      "id": "1507434",
      "postDate": "09/09/2021 07:43:10",
      "content": "<p>It's not always the case, difference between CV and LB for me was less than 10 bps, even though I'm using 2D model.<br>\nCV : 0.8757, LB : 0.8766<br>\nWith release of new 0.8758 kernel it's hard to properly compute CV for now, so I can't judge on that. </p>",
      "rawMarkdown": "It's not always the case, difference between CV and LB for me was less than 10 bps, even though I'm using 2D model.\nCV : 0.8757, LB : 0.8766\nWith release of new 0.8758 kernel it's hard to properly compute CV for now, so I can't judge on that.",
      "votes": null
    },
    {
      "id": "1507435",
      "postDate": "09/09/2021 07:43:42",
      "content": "<p>As calculated the test data has relative smaller mean value, I guess that's why most of us got higer LB than CV, maybe test data has less noise level?</p>\n<blockquote>\n  <p>Traind data: global_mean, global_std = 4.3110404142640676e-26, 6.150793058115705e-21<br>\n  Test data: global_mean, global_std = 1.647648937086022e-26, 6.150793058115705e-21</p>\n</blockquote>",
      "rawMarkdown": "As calculated the test data has relative smaller mean value, I guess that's why most of us got higer LB than CV, maybe test data has less noise level?\n\n> Traind data: global_mean, global_std = 4.3110404142640676e-26, 6.150793058115705e-21\nTest data: global_mean, global_std = 1.647648937086022e-26, 6.150793058115705e-21",
      "votes": null
    },
    {
      "id": "1507464",
      "postDate": "09/09/2021 08:16:35",
      "content": "<p>CV-LB difference looks reasonable considering that the same difference can be observed between different folds.</p>",
      "rawMarkdown": "CV-LB difference looks reasonable considering that the same difference can be observed between different folds.",
      "votes": null
    },
    {
      "id": "1507465",
      "postDate": "09/09/2021 08:18:52",
      "content": "<p>you have to tell us the performance of model.</p>\n<p>for less discriminative model, the generalization is of course better.</p>\n<p>as an extreme example, imagine we use a weak model like resnet18. then we have cv low cv and lb, say 0.840. the value are low but generalization gap is almost zero.</p>\n<p>if a model is highly discriminative (e.g. very deep CNN or transformer), you will learn the \"noise in the data\" and generalization drop and overfitting kicks in (unless you use regularisation methods)</p>",
      "rawMarkdown": "you have to tell us the performance of model.\n\nfor less discriminative model, the generalization is of course better.\n\nas an extreme example, imagine we use a weak model like resnet18. then we have cv low cv and lb, say 0.840. the value are low but generalization gap is almost zero.\n\nif a model is highly discriminative (e.g. very deep CNN or transformer), you will learn the \"noise in the data\" and generalization drop and overfitting kicks in (unless you use regularisation methods)",
      "votes": null
    },
    {
      "id": "1507476",
      "postDate": "09/09/2021 08:29:29",
      "content": "<p>in fact, resnet18 is not that weak. My current best single model is a resnet18d with local CV = 0.874047 and LB = 0.8759</p>",
      "rawMarkdown": "in fact, resnet18 is not that weak. My current best single model is a resnet18d with local CV = 0.874047 and LB = 0.8759",
      "votes": null
    },
    {
      "id": "1507774",
      "postDate": "09/09/2021 14:25:51",
      "content": "<p>Wow, nice CV-LB <a href=\"https://www.kaggle.com/fabiendaniel\" target=\"_blank\">@fabiendaniel</a> !! If you like to share, you use CWT or CQT ? </p>",
      "rawMarkdown": "Wow, nice CV-LB @fabiendaniel !! If you like to share, you use CWT or CQT ?",
      "votes": null
    },
    {
      "id": "1507810",
      "postDate": "09/09/2021 14:59:27",
      "content": "<p>We didn't see such big difference between folds either(std less than 5bps with only a few runs for 5Fold; we mostly did 1 fold for faster iteration of ideas). If you don't mind, what's your cv std for different folds?</p>",
      "rawMarkdown": "We didn't see such big difference between folds either(std less than 5bps with only a few runs for 5Fold; we mostly did 1 fold for faster iteration of ideas). If you don't mind, what's your cv std for different folds?",
      "votes": null
    },
    {
      "id": "1507816",
      "postDate": "09/09/2021 15:07:23",
      "content": "<p>I used CQT but according to past comments in the forum (I did not try CWT), there's not significant differences between the 2.</p>",
      "rawMarkdown": "I used CQT but according to past comments in the forum (I did not try CWT), there's not significant differences between the 2.",
      "votes": null
    },
    {
      "id": "1507817",
      "postDate": "09/09/2021 15:07:41",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks! For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730</p>",
      "rawMarkdown": "hengck23 Thanks! For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730",
      "votes": null
    },
    {
      "id": "1507822",
      "postDate": "09/09/2021 15:14:38",
      "content": "<p>Most people use N-fold training with out-of-fold CV, and therefore use one model per one example on train test. When they submit to LB, they use N models per one example.</p>\n<p>Thats all.</p>",
      "rawMarkdown": "Most people use N-fold training with out-of-fold CV, and therefore use one model per one example on train test. When they submit to LB, they use N models per one example.\n\nThats all.",
      "votes": null
    },
    {
      "id": "1507825",
      "postDate": "09/09/2021 15:16:45",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<blockquote>\n  <p>then we have cv low cv and lb, say 0.840. the value are low but generalisation gap is almost zero.</p>\n</blockquote>\n<p>so if I got it correct, what you suggest is that models with low CV-LB diff (gap) have better generalisation  - and hence we should trust more towards final selection (?)</p>",
      "rawMarkdown": "hengck23 \n>then we have cv low cv and lb, say 0.840. the value are low but generalisation gap is almost zero.\n\nso if I got it correct, what you suggest is that models with low CV-LB diff (gap) have better generalisation  - and hence we should trust more towards final selection (?)",
      "votes": null
    },
    {
      "id": "1507833",
      "postDate": "09/09/2021 15:25:01",
      "content": "<p><a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> congrats you managed to make 1D work! <br>\nI'm still struggling with wavenet/cnn1d models to make them converge - may I ask what is your input, i.e. if you use whole signal length (4096,3) or window of that?  </p>",
      "rawMarkdown": "richx86 congrats you managed to make 1D work! \nI'm still struggling with wavenet/cnn1d models to make them converge - may I ask what is your input, i.e. if you use whole signal length (4096,3) or window of that?",
      "votes": null
    },
    {
      "id": "1508157",
      "postDate": "09/10/2021 00:14:09",
      "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a>  Thanks! It's tricky to make 1D work. I used to have 0.5000 AUC. Since it's in the late stage of the competition, I will not share too much information here. Hope you would understand.</p>",
      "rawMarkdown": "imeintanis  Thanks! It's tricky to make 1D work. I used to have 0.5000 AUC. Since it's in the late stage of the competition, I will not share too much information here. Hope you would understand.",
      "votes": null
    },
    {
      "id": "1508158",
      "postDate": "09/10/2021 00:15:43",
      "content": "<p>The weird thing is that 5Fold doesn't seem to help with increasing the gap for our 1D model.</p>",
      "rawMarkdown": "The weird thing is that 5Fold doesn't seem to help with increasing the gap for our 1D model.",
      "votes": null
    },
    {
      "id": "1508182",
      "postDate": "09/10/2021 01:25:53",
      "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> </p>\n<p>\"hence we should trust more towards final selection\"</p>\n<p>if the model lb score is low it is not useful even it has a better generalization gap.</p>\n<p>the objective of model learning is to get BOTH the best accuracy AND best generalization.</p>",
      "rawMarkdown": "imeintanis \n\n\"hence we should trust more towards final selection\"\n\nif the model lb score is low it is not useful even it has a better generalization gap.\n\nthe objective of model learning is to get BOTH the best accuracy AND best generalization.",
      "votes": null
    },
    {
      "id": "1508187",
      "postDate": "09/10/2021 01:37:46",
      "content": "<p><a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> <br>\n\"For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730\"</p>\n<p>Then i would say 1D model generalizes better. I suppose you should also observe:</p>\n<ul>\n<li>std of AUC of k-fold model is lower </li>\n<li>if you compare AUC for 100%,50%,25%, etc … of the validation set, you should also see 1D model is less sensitive to the number of validation samples.</li>\n</ul>\n<p>the good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think</p>",
      "rawMarkdown": "richx86 \n\"For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730\"\n\nThen i would say 1D model generalizes better. I suppose you should also observe:\n- std of AUC of k-fold model is lower \n- if you compare AUC for 100%,50%,25%, etc ... of the validation set, you should also see 1D model is less sensitive to the number of validation samples.\n\nthe good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think",
      "votes": null
    },
    {
      "id": "1508209",
      "postDate": "09/10/2021 02:26:36",
      "content": "<blockquote>\n  <p>Then i would say 1D model generalizes better. I suppose you should also observe:</p>\n</blockquote>\n<pre><code>std of AUC of k-fold model is lower\nif you compare AUC for 100%,50%,25%, etc … of the validation set, you should also see 1D model is less sensitive to the number of validation samples.\n</code></pre>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks a lot for your insight! I will run a few more 5-folds of our models to check std and also AUC sensitivity for difference sizes of validation set. However, if 1D model generalizes better, shouldn't we see larger LB-CV gap for 1D? That's opposite of what we see right now. </p>\n<blockquote>\n  <p>the good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think</p>\n</blockquote>\n<p>We did see a nice boost for ensembling :).  The correlation between 1D and 2D model is around 0.95. I guess that evidence support your statement.</p>",
      "rawMarkdown": "> Then i would say 1D model generalizes better. I suppose you should also observe:\n\n    std of AUC of k-fold model is lower\n    if you compare AUC for 100%,50%,25%, etc … of the validation set, you should also see 1D model is less sensitive to the number of validation samples.\n\n@hengck23 Thanks a lot for your insight! I will run a few more 5-folds of our models to check std and also AUC sensitivity for difference sizes of validation set. However, if 1D model generalizes better, shouldn't we see larger LB-CV gap for 1D? That's opposite of what we see right now. \n\n> the good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think\n\nWe did see a nice boost for ensembling :).  The correlation between 1D and 2D model is around 0.95. I guess that evidence support your statement.",
      "votes": null
    },
    {
      "id": "1508217",
      "postDate": "09/10/2021 02:35:49",
      "content": "<p>better generalisation means lower CV-LB gap (assume same distribution of test and train data)</p>\n<p>\"also AUC sensitivity for difference sizes of validation set.\"<br>\ndo include this size:</p>\n<p>\"This leaderboard is calculated with approximately 16% of the test data.<br>\nThe final results will be based on the other 84%, so the final standings may be different\"</p>\n<hr>\n<p>on a side note the time localisation of both 1d CNN and 2D CNN should be the same. this setup consistency loss for training. maybe there can be some form of unsupervised, semi/weak unsupervised learning as well.</p>",
      "rawMarkdown": "better generalisation means lower CV-LB gap (assume same distribution of test and train data)\n\n\"also AUC sensitivity for difference sizes of validation set.\"\ndo include this size:\n\n\"This leaderboard is calculated with approximately 16% of the test data.\nThe final results will be based on the other 84%, so the final standings may be different\"\n\n---\n\non a side note the time localisation of both 1d CNN and 2D CNN should be the same. this setup consistency loss for training. maybe there can be some form of unsupervised, semi/weak unsupervised learning as well.",
      "votes": null
    },
    {
      "id": "1508294",
      "postDate": "09/10/2021 05:30:12",
      "content": "<p><a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> thanks! of course I understand, looking fwd to read your solution afterwards</p>",
      "rawMarkdown": "richx86 thanks! of course I understand, looking fwd to read your solution afterwards",
      "votes": null
    },
    {
      "id": "1508679",
      "postDate": "09/10/2021 13:00:49",
      "content": "<p>it is mentioned here:<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/271353\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/271353</a><br>\n(see section on matched filter CNN)</p>\n<p>noise is non stationary.<br>\n2d CNN can only learn stationary noise.<br>\n1d CNN however can work with nonstationary noise.</p>",
      "rawMarkdown": "it is mentioned here:\nhttps://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/271353\n(see section on matched filter CNN)\n\nnoise is non stationary.\n2d CNN can only learn stationary noise.\n1d CNN however can work with nonstationary noise.",
      "votes": null
    },
    {
      "id": "1530320",
      "postDate": "10/01/2021 05:18:44",
      "content": "<p>One hypothesis I have right now for my own question is that for 1D model and 2D model with similar performance(CV score), 1D outperform 2D on low SNR signals, while 2D outperform 1D on high SNR signals(which is the case for public LB). </p>",
      "rawMarkdown": "One hypothesis I have right now for my own question is that for 1D model and 2D model with similar performance(CV score), 1D outperform 2D on low SNR signals, while 2D outperform 1D on high SNR signals(which is the case for public LB).",
      "votes": null
    },
    {
      "id": "1559894",
      "postDate": "10/27/2021 08:06:44",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1507334,
      "author_name": "mithilsalunkhe",
      "author_url": "",
      "post_date": "09/09/2021 05:25:44",
      "content": "<p>What do you mean by BPS ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1507335,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "09/09/2021 05:27:03",
          "content": "<p><a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>  0.0001 is one bp</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1507434,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "09/09/2021 07:43:10",
      "content": "<p>It's not always the case, difference between CV and LB for me was less than 10 bps, even though I'm using 2D model.<br>\nCV : 0.8757, LB : 0.8766<br>\nWith release of new 0.8758 kernel it's hard to properly compute CV for now, so I can't judge on that. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1507435,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "09/09/2021 07:43:42",
      "content": "<p>As calculated the test data has relative smaller mean value, I guess that's why most of us got higer LB than CV, maybe test data has less noise level?</p>\n<blockquote>\n  <p>Traind data: global_mean, global_std = 4.3110404142640676e-26, 6.150793058115705e-21<br>\n  Test data: global_mean, global_std = 1.647648937086022e-26, 6.150793058115705e-21</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1507464,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "09/09/2021 08:16:35",
      "content": "<p>CV-LB difference looks reasonable considering that the same difference can be observed between different folds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1507810,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "09/09/2021 14:59:27",
          "content": "<p>We didn't see such big difference between folds either(std less than 5bps with only a few runs for 5Fold; we mostly did 1 fold for faster iteration of ideas). If you don't mind, what's your cv std for different folds?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1507465,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/09/2021 08:18:52",
      "content": "<p>you have to tell us the performance of model.</p>\n<p>for less discriminative model, the generalization is of course better.</p>\n<p>as an extreme example, imagine we use a weak model like resnet18. then we have cv low cv and lb, say 0.840. the value are low but generalization gap is almost zero.</p>\n<p>if a model is highly discriminative (e.g. very deep CNN or transformer), you will learn the \"noise in the data\" and generalization drop and overfitting kicks in (unless you use regularisation methods)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1507476,
          "author_name": "fabiendaniel",
          "author_url": "",
          "post_date": "09/09/2021 08:29:29",
          "content": "<p>in fact, resnet18 is not that weak. My current best single model is a resnet18d with local CV = 0.874047 and LB = 0.8759</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507774,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "09/09/2021 14:25:51",
          "content": "<p>Wow, nice CV-LB <a href=\"https://www.kaggle.com/fabiendaniel\" target=\"_blank\">@fabiendaniel</a> !! If you like to share, you use CWT or CQT ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507816,
          "author_name": "fabiendaniel",
          "author_url": "",
          "post_date": "09/09/2021 15:07:23",
          "content": "<p>I used CQT but according to past comments in the forum (I did not try CWT), there's not significant differences between the 2.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507817,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "09/09/2021 15:07:41",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks! For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507825,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "09/09/2021 15:16:45",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<blockquote>\n  <p>then we have cv low cv and lb, say 0.840. the value are low but generalisation gap is almost zero.</p>\n</blockquote>\n<p>so if I got it correct, what you suggest is that models with low CV-LB diff (gap) have better generalisation  - and hence we should trust more towards final selection (?)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1507833,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "09/09/2021 15:25:01",
          "content": "<p><a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> congrats you managed to make 1D work! <br>\nI'm still struggling with wavenet/cnn1d models to make them converge - may I ask what is your input, i.e. if you use whole signal length (4096,3) or window of that?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508157,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "09/10/2021 00:14:09",
          "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a>  Thanks! It's tricky to make 1D work. I used to have 0.5000 AUC. Since it's in the late stage of the competition, I will not share too much information here. Hope you would understand.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508182,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/10/2021 01:25:53",
          "content": "<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> </p>\n<p>\"hence we should trust more towards final selection\"</p>\n<p>if the model lb score is low it is not useful even it has a better generalization gap.</p>\n<p>the objective of model learning is to get BOTH the best accuracy AND best generalization.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508187,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/10/2021 01:37:46",
          "content": "<p><a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> <br>\n\"For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730\"</p>\n<p>Then i would say 1D model generalizes better. I suppose you should also observe:</p>\n<ul>\n<li>std of AUC of k-fold model is lower </li>\n<li>if you compare AUC for 100%,50%,25%, etc … of the validation set, you should also see 1D model is less sensitive to the number of validation samples.</li>\n</ul>\n<p>the good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508209,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "09/10/2021 02:26:36",
          "content": "<blockquote>\n  <p>Then i would say 1D model generalizes better. I suppose you should also observe:</p>\n</blockquote>\n<pre><code>std of AUC of k-fold model is lower\nif you compare AUC for 100%,50%,25%, etc … of the validation set, you should also see 1D model is less sensitive to the number of validation samples.\n</code></pre>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks a lot for your insight! I will run a few more 5-folds of our models to check std and also AUC sensitivity for difference sizes of validation set. However, if 1D model generalizes better, shouldn't we see larger LB-CV gap for 1D? That's opposite of what we see right now. </p>\n<blockquote>\n  <p>the good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think</p>\n</blockquote>\n<p>We did see a nice boost for ensembling :).  The correlation between 1D and 2D model is around 0.95. I guess that evidence support your statement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508217,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/10/2021 02:35:49",
          "content": "<p>better generalisation means lower CV-LB gap (assume same distribution of test and train data)</p>\n<p>\"also AUC sensitivity for difference sizes of validation set.\"<br>\ndo include this size:</p>\n<p>\"This leaderboard is calculated with approximately 16% of the test data.<br>\nThe final results will be based on the other 84%, so the final standings may be different\"</p>\n<hr>\n<p>on a side note the time localisation of both 1d CNN and 2D CNN should be the same. this setup consistency loss for training. maybe there can be some form of unsupervised, semi/weak unsupervised learning as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1508294,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "09/10/2021 05:30:12",
          "content": "<p><a href=\"https://www.kaggle.com/richx86\" target=\"_blank\">@richx86</a> thanks! of course I understand, looking fwd to read your solution afterwards</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1507822,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "09/09/2021 15:14:38",
      "content": "<p>Most people use N-fold training with out-of-fold CV, and therefore use one model per one example on train test. When they submit to LB, they use N models per one example.</p>\n<p>Thats all.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1508158,
          "author_name": "richx86",
          "author_url": "",
          "post_date": "09/10/2021 00:15:43",
          "content": "<p>The weird thing is that 5Fold doesn't seem to help with increasing the gap for our 1D model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1508679,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/10/2021 13:00:49",
      "content": "<p>it is mentioned here:<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/271353\" target=\"_blank\">https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/271353</a><br>\n(see section on matched filter CNN)</p>\n<p>noise is non stationary.<br>\n2d CNN can only learn stationary noise.<br>\n1d CNN however can work with nonstationary noise.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1530320,
      "author_name": "richx86",
      "author_url": "",
      "post_date": "10/01/2021 05:18:44",
      "content": "<p>One hypothesis I have right now for my own question is that for 1D model and 2D model with similar performance(CV score), 1D outperform 2D on low SNR signals, while 2D outperform 1D on high SNR signals(which is the case for public LB). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559894,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:06:44",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1507328": "I observed that LB and CV has a relatively large gap(30~50bps) for most people's models(as appeared in discussion threads). Can anyone give some reasonable explanations? \n\nFurthermore, our team is working on a 1D model. A very weird thing is that our 1D model somehow only has a gap 10bps or less most of the time, while our 2D model frequently has a ~20-30 bps gap. These two kinds of models share the same preprocessing(high low pass and whiten). This is very confusing to me and I hope people can share some insights about this(and I can increase the gap without overfitting to LB :P   ) Thanks!",
    "1507334": "What do you mean by BPS ?",
    "1507335": "mithilsalunkhe  0.0001 is one bp",
    "1507434": "It's not always the case, difference between CV and LB for me was less than 10 bps, even though I'm using 2D model.\nCV : 0.8757, LB : 0.8766\nWith release of new 0.8758 kernel it's hard to properly compute CV for now, so I can't judge on that.",
    "1507435": "As calculated the test data has relative smaller mean value, I guess that's why most of us got higer LB than CV, maybe test data has less noise level?\n\n> Traind data: global_mean, global_std = 4.3110404142640676e-26, 6.150793058115705e-21\nTest data: global_mean, global_std = 1.647648937086022e-26, 6.150793058115705e-21",
    "1507464": "CV-LB difference looks reasonable considering that the same difference can be observed between different folds.",
    "1507465": "you have to tell us the performance of model.\n\nfor less discriminative model, the generalization is of course better.\n\nas an extreme example, imagine we use a weak model like resnet18. then we have cv low cv and lb, say 0.840. the value are low but generalization gap is almost zero.\n\nif a model is highly discriminative (e.g. very deep CNN or transformer), you will learn the \"noise in the data\" and generalization drop and overfitting kicks in (unless you use regularisation methods)",
    "1507476": "in fact, resnet18 is not that weak. My current best single model is a resnet18d with local CV = 0.874047 and LB = 0.8759",
    "1507774": "Wow, nice CV-LB @fabiendaniel !! If you like to share, you use CWT or CQT ?",
    "1507810": "We didn't see such big difference between folds either(std less than 5bps with only a few runs for 5Fold; we mostly did 1 fold for faster iteration of ideas). If you don't mind, what's your cv std for different folds?",
    "1507816": "I used CQT but according to past comments in the forum (I did not try CWT), there's not significant differences between the 2.",
    "1507817": "hengck23 Thanks! For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730",
    "1507822": "Most people use N-fold training with out-of-fold CV, and therefore use one model per one example on train test. When they submit to LB, they use N models per one example.\n\nThats all.",
    "1507825": "hengck23 \n>then we have cv low cv and lb, say 0.840. the value are low but generalisation gap is almost zero.\n\nso if I got it correct, what you suggest is that models with low CV-LB diff (gap) have better generalisation  - and hence we should trust more towards final selection (?)",
    "1507833": "richx86 congrats you managed to make 1D work! \nI'm still struggling with wavenet/cnn1d models to make them converge - may I ask what is your input, i.e. if you use whole signal length (4096,3) or window of that?",
    "1508157": "imeintanis  Thanks! It's tricky to make 1D work. I used to have 0.5000 AUC. Since it's in the late stage of the competition, I will not share too much information here. Hope you would understand.",
    "1508158": "The weird thing is that 5Fold doesn't seem to help with increasing the gap for our 1D model.",
    "1508182": "imeintanis \n\n\"hence we should trust more towards final selection\"\n\nif the model lb score is low it is not useful even it has a better generalization gap.\n\nthe objective of model learning is to get BOTH the best accuracy AND best generalization.",
    "1508187": "richx86 \n\"For the ones we submitted with our 1D model, CV ranges from 0.8697 to 0.8730\"\n\nThen i would say 1D model generalizes better. I suppose you should also observe:\n- std of AUC of k-fold model is lower \n- if you compare AUC for 100%,50%,25%, etc ... of the validation set, you should also see 1D model is less sensitive to the number of validation samples.\n\nthe good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think",
    "1508209": "> Then i would say 1D model generalizes better. I suppose you should also observe:\n\n    std of AUC of k-fold model is lower\n    if you compare AUC for 100%,50%,25%, etc … of the validation set, you should also see 1D model is less sensitive to the number of validation samples.\n\n@hengck23 Thanks a lot for your insight! I will run a few more 5-folds of our models to check std and also AUC sensitivity for difference sizes of validation set. However, if 1D model generalizes better, shouldn't we see larger LB-CV gap for 1D? That's opposite of what we see right now. \n\n> the good news is that 1D and 2D focuses on different \"view\" of the data and you are likely to get a good boost at ensembling i think\n\nWe did see a nice boost for ensembling :).  The correlation between 1D and 2D model is around 0.95. I guess that evidence support your statement.",
    "1508217": "better generalisation means lower CV-LB gap (assume same distribution of test and train data)\n\n\"also AUC sensitivity for difference sizes of validation set.\"\ndo include this size:\n\n\"This leaderboard is calculated with approximately 16% of the test data.\nThe final results will be based on the other 84%, so the final standings may be different\"\n\n---\n\non a side note the time localisation of both 1d CNN and 2D CNN should be the same. this setup consistency loss for training. maybe there can be some form of unsupervised, semi/weak unsupervised learning as well.",
    "1508294": "richx86 thanks! of course I understand, looking fwd to read your solution afterwards",
    "1508679": "it is mentioned here:\nhttps://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/271353\n(see section on matched filter CNN)\n\nnoise is non stationary.\n2d CNN can only learn stationary noise.\n1d CNN however can work with nonstationary noise.",
    "1530320": "One hypothesis I have right now for my own question is that for 1D model and 2D model with similar performance(CV score), 1D outperform 2D on low SNR signals, while 2D outperform 1D on high SNR signals(which is the case for public LB).",
    "1559894": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}