{
  "id": 274995,
  "title": "Shake Estimation",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/274995",
  "author_name": "",
  "post_date": "2021-09-28T11:16:21.431734900Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In this competition, CV vs LB is very nice.<br>\nIt seems that no big shake will be happen as mentioned in previous discussion.<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/251549\" target=\"_blank\">[CV vs LB]</a><br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/272968\" target=\"_blank\">Shake up expectations ?</a></p>\n<p>I was just curious to try shake estimation, so I published the notebook.<br>\n<a href=\"https://www.kaggle.com/cocoinit23/shake-estimation\" target=\"_blank\">https://www.kaggle.com/cocoinit23/shake-estimation</a></p>\n<p>I would be happy to get your feedbacks.<br>\nThanks!</p>",
  "messages": [
    {
      "id": "1526832",
      "postDate": "09/28/2021 11:16:21",
      "content": "<p>In this competition, CV vs LB is very nice.<br>\nIt seems that no big shake will be happen as mentioned in previous discussion.<br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/251549\" target=\"_blank\">[CV vs LB]</a><br>\n<a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/272968\" target=\"_blank\">Shake up expectations ?</a></p>\n<p>I was just curious to try shake estimation, so I published the notebook.<br>\n<a href=\"https://www.kaggle.com/cocoinit23/shake-estimation\" target=\"_blank\">https://www.kaggle.com/cocoinit23/shake-estimation</a></p>\n<p>I would be happy to get your feedbacks.<br>\nThanks!</p>",
      "rawMarkdown": "In this competition, CV vs LB is very nice.\nIt seems that no big shake will be happen as mentioned in previous discussion.\n[[CV vs LB]](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/251549)\n[Shake up expectations ?](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/272968)\n\nI was just curious to try shake estimation, so I published the notebook.\n[https://www.kaggle.com/cocoinit23/shake-estimation](https://www.kaggle.com/cocoinit23/shake-estimation)\n\nI would be happy to get your feedbacks.\nThanks!",
      "votes": null
    },
    {
      "id": "1526846",
      "postDate": "09/28/2021 11:30:29",
      "content": "<p><img src=\"https://i.ibb.co/PjTj84K/Selection-916.png\" alt=\"https://i.ibb.co/PjTj84K/Selection-916.png\"></p>\n<p>our public board uses 16% of 226000 = 36160 samples.<br>\nthis is 33% (36160/112000) in my graph above.<br>\nthis is roughly sufficient.<br>\n(you should analyze for 5 folds, over different random subset)</p>\n<p>private/public shakeup should be within 0.001~0.002.</p>\n<p>If the current top team does not overfit, he should be a clear winner.</p>\n<hr>\n<p>it is interesting to make similar graphs if pretraining/normalisation/self-supervised learning is applied to both validation and training sets.  or one may think of a way to make the curves (3.75% 10% 20% 50% …) closer</p>",
      "rawMarkdown": "![https://i.ibb.co/PjTj84K/Selection-916.png](https://i.ibb.co/PjTj84K/Selection-916.png)\n\nour public board uses 16% of 226000 = 36160 samples.\nthis is 33% (36160/112000) in my graph above.\nthis is roughly sufficient.\n(you should analyze for 5 folds, over different random subset)\n\nprivate/public shakeup should be within 0.001~0.002.\n\nIf the current top team does not overfit, he should be a clear winner.\n\n---\n\nit is interesting to make similar graphs if pretraining/normalisation/self-supervised learning is applied to both validation and training sets.  or one may think of a way to make the curves (3.75% 10% 20% 50% ...) closer",
      "votes": null
    },
    {
      "id": "1526941",
      "postDate": "09/28/2021 12:38:20",
      "content": "<p>I don't think you can call a constant shift from public to private a shakeup, as everyone will or will not have this shift. A shakeup is a severe shuffle of ranks.</p>",
      "rawMarkdown": "I don't think you can call a constant shift from public to private a shakeup, as everyone will or will not have this shift. A shakeup is a severe shuffle of ranks.",
      "votes": null
    },
    {
      "id": "1527577",
      "postDate": "09/28/2021 22:11:24",
      "content": "<p>I think that your analysis is too pessimistic because your sampling doesn't account for SNR in samples.<br>\n(basically, some of your sample sets may contain only low-SNR samples which are hard or impossible to predict, and some sample sets may contain only high-SNR samples which are relatively easy to predict)</p>\n<p>I have an assumption that organizers will not only keep class balance the same, but also set of intrinsic parameters for generating GWs across public and private test (this will result in roughly the same PDFs of SNR of public-test/private-test and hence we won't experience the shakeup)</p>\n<p>I think if there is a way to measure SNR of samples, it would be better to compare different sample sets with close PDFs of SNR, with this approach we'll get more reasonable estimation of shake-up.</p>\n<p>PDF is a probability density function, if we can call it like that in this case.</p>\n<p>Is will be great to discuss what would happen if distribution shift of SNR occurs. </p>",
      "rawMarkdown": "I think that your analysis is too pessimistic because your sampling doesn't account for SNR in samples.\n(basically, some of your sample sets may contain only low-SNR samples which are hard or impossible to predict, and some sample sets may contain only high-SNR samples which are relatively easy to predict)\n \nI have an assumption that organizers will not only keep class balance the same, but also set of intrinsic parameters for generating GWs across public and private test (this will result in roughly the same PDFs of SNR of public-test/private-test and hence we won't experience the shakeup)\n\nI think if there is a way to measure SNR of samples, it would be better to compare different sample sets with close PDFs of SNR, with this approach we'll get more reasonable estimation of shake-up.\n\nPDF is a probability density function, if we can call it like that in this case.\n\nIs will be great to discuss what would happen if distribution shift of SNR occurs.",
      "votes": null
    },
    {
      "id": "1527650",
      "postDate": "09/29/2021 01:30:36",
      "content": "<p>You are right, SNR is an important factor.<br>\nIf we could calculate SNR of samples, the d prime based on Signal Detection Theory might be better metric to analyze difference of probability density distributions.</p>",
      "rawMarkdown": "You are right, SNR is an important factor.\nIf we could calculate SNR of samples, the d prime based on Signal Detection Theory might be better metric to analyze difference of probability density distributions.",
      "votes": null
    },
    {
      "id": "1527660",
      "postDate": "09/29/2021 01:43:32",
      "content": "<p>That’s a sharp opinion, AUC is affected by the order of predicted scores.<br>\nBut I re-run my estimation notebook with rank transformation below, the results were almost the same.<br>\n<code>df[“y_pred”] = df[\"y_pred\"].rank(pct=True)\n</code><br>\nMaybe I have overlooked something important…</p>",
      "rawMarkdown": "That’s a sharp opinion, AUC is affected by the order of predicted scores.\nBut I re-run my estimation notebook with rank transformation below, the results were almost the same.\n`df[“y_pred”] = df[\"y_pred\"].rank(pct=True)\n`\nMaybe I have overlooked something important…",
      "votes": null
    },
    {
      "id": "1527703",
      "postDate": "09/29/2021 03:03:28",
      "content": "<p>I think the point here is that the variance of score when you bootstrap dataset contains both fixed effect (constant shift as Psi mentioned) and random effect (this is what we call \"shakeup\").</p>",
      "rawMarkdown": "I think the point here is that the variance of score when you bootstrap dataset contains both fixed effect (constant shift as Psi mentioned) and random effect (this is what we call \"shakeup\").",
      "votes": null
    },
    {
      "id": "1527787",
      "postDate": "09/29/2021 05:08:31",
      "content": "<p>That make sense, I misunderstood.<br>\nThank you for elaboration!</p>",
      "rawMarkdown": "That make sense, I misunderstood.\nThank you for elaboration!",
      "votes": null
    },
    {
      "id": "1559739",
      "postDate": "10/27/2021 07:10:44",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1526846,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/28/2021 11:30:29",
      "content": "<p><img src=\"https://i.ibb.co/PjTj84K/Selection-916.png\" alt=\"https://i.ibb.co/PjTj84K/Selection-916.png\"></p>\n<p>our public board uses 16% of 226000 = 36160 samples.<br>\nthis is 33% (36160/112000) in my graph above.<br>\nthis is roughly sufficient.<br>\n(you should analyze for 5 folds, over different random subset)</p>\n<p>private/public shakeup should be within 0.001~0.002.</p>\n<p>If the current top team does not overfit, he should be a clear winner.</p>\n<hr>\n<p>it is interesting to make similar graphs if pretraining/normalisation/self-supervised learning is applied to both validation and training sets.  or one may think of a way to make the curves (3.75% 10% 20% 50% …) closer</p>",
      "votes": null,
      "replies": [
        {
          "id": 1526941,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "09/28/2021 12:38:20",
          "content": "<p>I don't think you can call a constant shift from public to private a shakeup, as everyone will or will not have this shift. A shakeup is a severe shuffle of ranks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1527660,
          "author_name": "cocoinit23",
          "author_url": "",
          "post_date": "09/29/2021 01:43:32",
          "content": "<p>That’s a sharp opinion, AUC is affected by the order of predicted scores.<br>\nBut I re-run my estimation notebook with rank transformation below, the results were almost the same.<br>\n<code>df[“y_pred”] = df[\"y_pred\"].rank(pct=True)\n</code><br>\nMaybe I have overlooked something important…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1527703,
          "author_name": "analokamus",
          "author_url": "",
          "post_date": "09/29/2021 03:03:28",
          "content": "<p>I think the point here is that the variance of score when you bootstrap dataset contains both fixed effect (constant shift as Psi mentioned) and random effect (this is what we call \"shakeup\").</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1527787,
          "author_name": "cocoinit23",
          "author_url": "",
          "post_date": "09/29/2021 05:08:31",
          "content": "<p>That make sense, I misunderstood.<br>\nThank you for elaboration!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1527577,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "09/28/2021 22:11:24",
      "content": "<p>I think that your analysis is too pessimistic because your sampling doesn't account for SNR in samples.<br>\n(basically, some of your sample sets may contain only low-SNR samples which are hard or impossible to predict, and some sample sets may contain only high-SNR samples which are relatively easy to predict)</p>\n<p>I have an assumption that organizers will not only keep class balance the same, but also set of intrinsic parameters for generating GWs across public and private test (this will result in roughly the same PDFs of SNR of public-test/private-test and hence we won't experience the shakeup)</p>\n<p>I think if there is a way to measure SNR of samples, it would be better to compare different sample sets with close PDFs of SNR, with this approach we'll get more reasonable estimation of shake-up.</p>\n<p>PDF is a probability density function, if we can call it like that in this case.</p>\n<p>Is will be great to discuss what would happen if distribution shift of SNR occurs. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1527650,
          "author_name": "cocoinit23",
          "author_url": "",
          "post_date": "09/29/2021 01:30:36",
          "content": "<p>You are right, SNR is an important factor.<br>\nIf we could calculate SNR of samples, the d prime based on Signal Detection Theory might be better metric to analyze difference of probability density distributions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559739,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:10:44",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1526832": "In this competition, CV vs LB is very nice.\nIt seems that no big shake will be happen as mentioned in previous discussion.\n[[CV vs LB]](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/251549)\n[Shake up expectations ?](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/272968)\n\nI was just curious to try shake estimation, so I published the notebook.\n[https://www.kaggle.com/cocoinit23/shake-estimation](https://www.kaggle.com/cocoinit23/shake-estimation)\n\nI would be happy to get your feedbacks.\nThanks!",
    "1526846": "![https://i.ibb.co/PjTj84K/Selection-916.png](https://i.ibb.co/PjTj84K/Selection-916.png)\n\nour public board uses 16% of 226000 = 36160 samples.\nthis is 33% (36160/112000) in my graph above.\nthis is roughly sufficient.\n(you should analyze for 5 folds, over different random subset)\n\nprivate/public shakeup should be within 0.001~0.002.\n\nIf the current top team does not overfit, he should be a clear winner.\n\n---\n\nit is interesting to make similar graphs if pretraining/normalisation/self-supervised learning is applied to both validation and training sets.  or one may think of a way to make the curves (3.75% 10% 20% 50% ...) closer",
    "1526941": "I don't think you can call a constant shift from public to private a shakeup, as everyone will or will not have this shift. A shakeup is a severe shuffle of ranks.",
    "1527577": "I think that your analysis is too pessimistic because your sampling doesn't account for SNR in samples.\n(basically, some of your sample sets may contain only low-SNR samples which are hard or impossible to predict, and some sample sets may contain only high-SNR samples which are relatively easy to predict)\n \nI have an assumption that organizers will not only keep class balance the same, but also set of intrinsic parameters for generating GWs across public and private test (this will result in roughly the same PDFs of SNR of public-test/private-test and hence we won't experience the shakeup)\n\nI think if there is a way to measure SNR of samples, it would be better to compare different sample sets with close PDFs of SNR, with this approach we'll get more reasonable estimation of shake-up.\n\nPDF is a probability density function, if we can call it like that in this case.\n\nIs will be great to discuss what would happen if distribution shift of SNR occurs.",
    "1527650": "You are right, SNR is an important factor.\nIf we could calculate SNR of samples, the d prime based on Signal Detection Theory might be better metric to analyze difference of probability density distributions.",
    "1527660": "That’s a sharp opinion, AUC is affected by the order of predicted scores.\nBut I re-run my estimation notebook with rank transformation below, the results were almost the same.\n`df[“y_pred”] = df[\"y_pred\"].rank(pct=True)\n`\nMaybe I have overlooked something important…",
    "1527703": "I think the point here is that the variance of score when you bootstrap dataset contains both fixed effect (constant shift as Psi mentioned) and random effect (this is what we call \"shakeup\").",
    "1527787": "That make sense, I misunderstood.\nThank you for elaboration!",
    "1559739": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}