{
  "id": 267588,
  "title": "Important PreProcessing Kernel",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/267588",
  "author_name": "",
  "post_date": "2021-08-23T23:05:06.535208800Z",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>So it's clear that preprocessing is the most important step here. I looked over all of the kernels and all of the discussion posts, and by far, <a href=\"https://www.kaggle.com/cnhung/kaggle-g2net-noise-analysis-feature-extraction\" target=\"_blank\">this particular one</a> by <a href=\"https://www.kaggle.com/cnhung\" target=\"_blank\">@cnhung</a> is orders of magnitude more consequential and informative than any of the others. It's unfortunate it has so few views and so few upvotes. I can only assume that since it was posted so early in the competition, most of us weren't even at the necessary minimal amount of domain understanding to digest it. Much effort went into building it.</p>",
  "messages": [
    {
      "id": "1487879",
      "postDate": "08/23/2021 23:05:06",
      "content": "<p>So it's clear that preprocessing is the most important step here. I looked over all of the kernels and all of the discussion posts, and by far, <a href=\"https://www.kaggle.com/cnhung/kaggle-g2net-noise-analysis-feature-extraction\" target=\"_blank\">this particular one</a> by <a href=\"https://www.kaggle.com/cnhung\" target=\"_blank\">@cnhung</a> is orders of magnitude more consequential and informative than any of the others. It's unfortunate it has so few views and so few upvotes. I can only assume that since it was posted so early in the competition, most of us weren't even at the necessary minimal amount of domain understanding to digest it. Much effort went into building it.</p>",
      "rawMarkdown": "So it's clear that preprocessing is the most important step here. I looked over all of the kernels and all of the discussion posts, and by far, [this particular one](https://www.kaggle.com/cnhung/kaggle-g2net-noise-analysis-feature-extraction) by @cnhung is orders of magnitude more consequential and informative than any of the others. It's unfortunate it has so few views and so few upvotes. I can only assume that since it was posted so early in the competition, most of us weren't even at the necessary minimal amount of domain understanding to digest it. Much effort went into building it.",
      "votes": null
    },
    {
      "id": "1490438",
      "postDate": "08/25/2021 16:17:25",
      "content": "<p>How do you assess importance here?</p>\n<p>The fact that the kernel author has not submitted probably explains why his notebook got little attention.  People look for what yields good public LB score in general.  This can be misleading in general, but here CV LB are well correlated, hence good public KB score for a notebook is a good indication of its importance IMHO.</p>\n<p>Back to the kernel you link to, it does contain a lot of EDA and preprocessing steps. Question is: are these helping solve te problem or not?  Answer to this is missing.</p>",
      "rawMarkdown": "How do you assess importance here?\n\nThe fact that the kernel author has not submitted probably explains why his notebook got little attention.  People look for what yields good public LB score in general.  This can be misleading in general, but here CV LB are well correlated, hence good public KB score for a notebook is a good indication of its importance IMHO.\n\nBack to the kernel you link to, it does contain a lot of EDA and preprocessing steps. Question is: are these helping solve te problem or not?  Answer to this is missing.",
      "votes": null
    },
    {
      "id": "1490476",
      "postDate": "08/25/2021 16:38:04",
      "content": "<p>A fair note. <a href=\"https://www.kaggle.com/cnhung\" target=\"_blank\">@cnhung</a>'s FE towards the end of that kernel was.. interesting, and TBH I didn't use those statistical features in my modeling. Here, I assess importance by the amount of domain knowledge brought to the table and probably should have used the term <em>interesting</em> instead.</p>\n<p>If we break down kernels, I feel they can be clustered into four groups:</p>\n<ul>\n<li>The highest ranked are the ensembles for the reasons you mentioned. Something can be learned from them.</li>\n<li>Then the majority are CQT/CWT/MEL type kernels. Somewhat more can be learned from them, especially those that include their own data pre-processing.</li>\n<li>A plethora of surface-level EDA kernels.</li>\n<li>Finally, a modest pinch of notebooks that get into the nitty gritty meat of domain-expert level GW analysis. Of them and I felt this was the best in terms of code.</li>\n</ul>",
      "rawMarkdown": "A fair note. @cnhung's FE towards the end of that kernel was.. interesting, and TBH I didn't use those statistical features in my modeling. Here, I assess importance by the amount of domain knowledge brought to the table and probably should have used the term *interesting* instead.\n\nIf we break down kernels, I feel they can be clustered into four groups:\n- The highest ranked are the ensembles for the reasons you mentioned. Something can be learned from them.\n- Then the majority are CQT/CWT/MEL type kernels. Somewhat more can be learned from them, especially those that include their own data pre-processing.\n- A plethora of surface-level EDA kernels.\n- Finally, a modest pinch of notebooks that get into the nitty gritty meat of domain-expert level GW analysis. Of them and I felt this was the best in terms of code.",
      "votes": null
    },
    {
      "id": "1491251",
      "postDate": "08/26/2021 09:05:31",
      "content": "<p>I am just trying to explain why this notebook didn't get much attention.  TL;DR its impact on solving the problem is not clear.</p>",
      "rawMarkdown": "I am just trying to explain why this notebook didn't get much attention.  TL;DR its impact on solving the problem is not clear.",
      "votes": null
    },
    {
      "id": "1559989",
      "postDate": "10/27/2021 08:53:20",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    },
    {
      "id": "1560999",
      "postDate": "10/27/2021 09:03:45",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1490438,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/25/2021 16:17:25",
      "content": "<p>How do you assess importance here?</p>\n<p>The fact that the kernel author has not submitted probably explains why his notebook got little attention.  People look for what yields good public LB score in general.  This can be misleading in general, but here CV LB are well correlated, hence good public KB score for a notebook is a good indication of its importance IMHO.</p>\n<p>Back to the kernel you link to, it does contain a lot of EDA and preprocessing steps. Question is: are these helping solve te problem or not?  Answer to this is missing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1490476,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/25/2021 16:38:04",
          "content": "<p>A fair note. <a href=\"https://www.kaggle.com/cnhung\" target=\"_blank\">@cnhung</a>'s FE towards the end of that kernel was.. interesting, and TBH I didn't use those statistical features in my modeling. Here, I assess importance by the amount of domain knowledge brought to the table and probably should have used the term <em>interesting</em> instead.</p>\n<p>If we break down kernels, I feel they can be clustered into four groups:</p>\n<ul>\n<li>The highest ranked are the ensembles for the reasons you mentioned. Something can be learned from them.</li>\n<li>Then the majority are CQT/CWT/MEL type kernels. Somewhat more can be learned from them, especially those that include their own data pre-processing.</li>\n<li>A plethora of surface-level EDA kernels.</li>\n<li>Finally, a modest pinch of notebooks that get into the nitty gritty meat of domain-expert level GW analysis. Of them and I felt this was the best in terms of code.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1491251,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/26/2021 09:05:31",
          "content": "<p>I am just trying to explain why this notebook didn't get much attention.  TL;DR its impact on solving the problem is not clear.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559989,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:53:20",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1560999,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 09:03:45",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1487879": "So it's clear that preprocessing is the most important step here. I looked over all of the kernels and all of the discussion posts, and by far, [this particular one](https://www.kaggle.com/cnhung/kaggle-g2net-noise-analysis-feature-extraction) by @cnhung is orders of magnitude more consequential and informative than any of the others. It's unfortunate it has so few views and so few upvotes. I can only assume that since it was posted so early in the competition, most of us weren't even at the necessary minimal amount of domain understanding to digest it. Much effort went into building it.",
    "1490438": "How do you assess importance here?\n\nThe fact that the kernel author has not submitted probably explains why his notebook got little attention.  People look for what yields good public LB score in general.  This can be misleading in general, but here CV LB are well correlated, hence good public KB score for a notebook is a good indication of its importance IMHO.\n\nBack to the kernel you link to, it does contain a lot of EDA and preprocessing steps. Question is: are these helping solve te problem or not?  Answer to this is missing.",
    "1490476": "A fair note. @cnhung's FE towards the end of that kernel was.. interesting, and TBH I didn't use those statistical features in my modeling. Here, I assess importance by the amount of domain knowledge brought to the table and probably should have used the term *interesting* instead.\n\nIf we break down kernels, I feel they can be clustered into four groups:\n- The highest ranked are the ensembles for the reasons you mentioned. Something can be learned from them.\n- Then the majority are CQT/CWT/MEL type kernels. Somewhat more can be learned from them, especially those that include their own data pre-processing.\n- A plethora of surface-level EDA kernels.\n- Finally, a modest pinch of notebooks that get into the nitty gritty meat of domain-expert level GW analysis. Of them and I felt this was the best in terms of code.",
    "1491251": "I am just trying to explain why this notebook didn't get much attention.  TL;DR its impact on solving the problem is not clear.",
    "1559989": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1560999": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}