{
  "id": 492243,
  "title": "high quality definition and test data distribution may be supported by SPaRCNet authors",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492243",
  "author_name": "",
  "post_date": "2024-04-09T02:33:08.251516500Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439\" target=\"_blank\">discussion in 2 months ago</a>, I've found the information about definition of low/high quality label and how to split test data in the host's paper, SPaRCNet. Note that the test data in this competition may be not same as the paper, but there may be a high likelihood of such a tendency.</p>\n<p>I’ve realized that <strong>it would be beneficial to read the host's paper in advance next competition</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F7d9c714f802c3c84470e15026ae3e7de%2F2024-04-09%2011.20.40.png?generation=1712629355247147&amp;alt=media\"></p>\n<p>reference</p>\n<ul>\n<li><a href=\"https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\" target=\"_blank\">https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf</a></li>\n<li><a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_SPaRCNet.pdf\" target=\"_blank\">https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_SPaRCNet.pdf</a></li>\n</ul>",
  "messages": [
    {
      "id": "2742647",
      "postDate": "04/09/2024 02:33:08",
      "content": "<p>In <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439\" target=\"_blank\">discussion in 2 months ago</a>, I've found the information about definition of low/high quality label and how to split test data in the host's paper, SPaRCNet. Note that the test data in this competition may be not same as the paper, but there may be a high likelihood of such a tendency.</p>\n<p>I’ve realized that <strong>it would be beneficial to read the host's paper in advance next competition</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F7d9c714f802c3c84470e15026ae3e7de%2F2024-04-09%2011.20.40.png?generation=1712629355247147&amp;alt=media\"></p>\n<p>reference</p>\n<ul>\n<li><a href=\"https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\" target=\"_blank\">https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf</a></li>\n<li><a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_SPaRCNet.pdf\" target=\"_blank\">https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_SPaRCNet.pdf</a></li>\n</ul>",
      "rawMarkdown": "In [discussion in 2 months ago](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439), I've found the information about definition of low/high quality label and how to split test data in the host's paper, SPaRCNet. Note that the test data in this competition may be not same as the paper, but there may be a high likelihood of such a tendency.\n\nI’ve realized that **it would be beneficial to read the host's paper in advance next competition**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F7d9c714f802c3c84470e15026ae3e7de%2F2024-04-09%2011.20.40.png?generation=1712629355247147&alt=media)\n\n\nreference\n- https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\n- https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_SPaRCNet.pdf",
      "votes": null
    },
    {
      "id": "2742652",
      "postDate": "04/09/2024 02:41:03",
      "content": "<p>I experimented with the “channel flipping” data augmentation method outlined in the SPaRCNet paper. However, I found that it only marginally enhanced my CV (&lt;0.01) while significantly prolonging my training duration. As a result, I decided to forgo its implementation. I’m curious to know if there might be a more effective approach to utilizing this technique.</p>",
      "rawMarkdown": "I experimented with the “channel flipping” data augmentation method outlined in the SPaRCNet paper. However, I found that it only marginally enhanced my CV (<0.01) while significantly prolonging my training duration. As a result, I decided to forgo its implementation. I’m curious to know if there might be a more effective approach to utilizing this technique.",
      "votes": null
    },
    {
      "id": "2742657",
      "postDate": "04/09/2024 02:44:50",
      "content": "<p>same result as me.</p>",
      "rawMarkdown": "same result as me.",
      "votes": null
    },
    {
      "id": "2742669",
      "postDate": "04/09/2024 02:50:43",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697#2710549\" target=\"_blank\">my visualization of oof again</a>.</p>\n<blockquote>\n  <p>Here is the prediction distribution by total_evaluators in oof_df.<br>\n  Y: kl divergence from label<br>\n  X: total_evaluators<br>\n  This prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.<br>\n  The baseline model is almost same as moth's published notebook.<br>\n  As you can see, kl divergence is lager in data with total_evaluators &lt;= 3.<br>\n  It seems that data with total_evaluators &lt;= 3 have difficult sample or noizy sample.<br>\n  In my experiments, this distribution often happens when training baseline efficient net model with all data.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fa60a02eaacbcddfdbfb7e72348af5d2f%2Finbox_4250230_9f2d180d88de96d03a41fa119a0703c2_viz_.png?generation=1712630970497468&amp;alt=media\"></p>",
      "rawMarkdown": "[my visualization of oof again](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697#2710549).\n\n>Here is the prediction distribution by total_evaluators in oof_df.\nY: kl divergence from label\nX: total_evaluators\nThis prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.\nThe baseline model is almost same as moth's published notebook.\nAs you can see, kl divergence is lager in data with total_evaluators <= 3.\nIt seems that data with total_evaluators <= 3 have difficult sample or noizy sample.\nIn my experiments, this distribution often happens when training baseline efficient net model with all data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fa60a02eaacbcddfdbfb7e72348af5d2f%2Finbox_4250230_9f2d180d88de96d03a41fa119a0703c2_viz_.png?generation=1712630970497468&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2742652,
      "author_name": "phoenix1203",
      "author_url": "",
      "post_date": "04/09/2024 02:41:03",
      "content": "<p>I experimented with the “channel flipping” data augmentation method outlined in the SPaRCNet paper. However, I found that it only marginally enhanced my CV (&lt;0.01) while significantly prolonging my training duration. As a result, I decided to forgo its implementation. I’m curious to know if there might be a more effective approach to utilizing this technique.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2742657,
          "author_name": "clearwaterkzk",
          "author_url": "",
          "post_date": "04/09/2024 02:44:50",
          "content": "<p>same result as me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2742669,
      "author_name": "clearwaterkzk",
      "author_url": "",
      "post_date": "04/09/2024 02:50:43",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697#2710549\" target=\"_blank\">my visualization of oof again</a>.</p>\n<blockquote>\n  <p>Here is the prediction distribution by total_evaluators in oof_df.<br>\n  Y: kl divergence from label<br>\n  X: total_evaluators<br>\n  This prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.<br>\n  The baseline model is almost same as moth's published notebook.<br>\n  As you can see, kl divergence is lager in data with total_evaluators &lt;= 3.<br>\n  It seems that data with total_evaluators &lt;= 3 have difficult sample or noizy sample.<br>\n  In my experiments, this distribution often happens when training baseline efficient net model with all data.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fa60a02eaacbcddfdbfb7e72348af5d2f%2Finbox_4250230_9f2d180d88de96d03a41fa119a0703c2_viz_.png?generation=1712630970497468&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742647": "In [discussion in 2 months ago](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439), I've found the information about definition of low/high quality label and how to split test data in the host's paper, SPaRCNet. Note that the test data in this competition may be not same as the paper, but there may be a high likelihood of such a tendency.\n\nI’ve realized that **it would be beneficial to read the host's paper in advance next competition**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F7d9c714f802c3c84470e15026ae3e7de%2F2024-04-09%2011.20.40.png?generation=1712629355247147&alt=media)\n\n\nreference\n- https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\n- https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_SPaRCNet.pdf",
    "2742652": "I experimented with the “channel flipping” data augmentation method outlined in the SPaRCNet paper. However, I found that it only marginally enhanced my CV (<0.01) while significantly prolonging my training duration. As a result, I decided to forgo its implementation. I’m curious to know if there might be a more effective approach to utilizing this technique.",
    "2742657": "same result as me.",
    "2742669": "[my visualization of oof again](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/485697#2710549).\n\n>Here is the prediction distribution by total_evaluators in oof_df.\nY: kl divergence from label\nX: total_evaluators\nThis prediction is made by baseline efficient net model trained with all data, cv score with all data is 0.5X.\nThe baseline model is almost same as moth's published notebook.\nAs you can see, kl divergence is lager in data with total_evaluators <= 3.\nIt seems that data with total_evaluators <= 3 have difficult sample or noizy sample.\nIn my experiments, this distribution often happens when training baseline efficient net model with all data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fa60a02eaacbcddfdbfb7e72348af5d2f%2Finbox_4250230_9f2d180d88de96d03a41fa119a0703c2_viz_.png?generation=1712630970497468&alt=media)"
  },
  "source": "meta"
}