{
  "id": 471439,
  "title": "Background of this competition",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/471439",
  "author_name": "",
  "post_date": "2024-01-28T10:56:34.977368700Z",
  "votes": 50,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi, kagglers.</p>\n<p>I want to bring in some discoveries through online surfing. In 2023, researchers from medical schools have jointly collected data and published it (<a href=\"https://bdsp.io/content/bdsp-sparcnet/1.1/\" target=\"_blank\">paper</a>). I am pretty sure the data described in there is the same as ours, since the targets and authors are exactly the same. Two models were built as an attempt to classify eegs into 6 patterns - i.e. seizure, lpd, gpd, lrda, grda, other. One is SPaRCNet <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file\" target=\"_blank\">(GitHub repo)</a> and the other is IRR <a href=\"https://github.com/bdsp-core/IIIC-IRR\" target=\"_blank\">(GitHub repo)</a> along with two papers (<a href=\"https://pubmed.ncbi.nlm.nih.gov/36878708/\" target=\"_blank\">p1</a> and <a href=\"https://arxiv.org/pdf/2211.05207.pdf\" target=\"_blank\">p2</a>). </p>\n<p>I am not sure of how effective the models are, but it doesn't seem to be very effective. So I assume that's why they host a competition on Kaggle (hopefully we can build a better solution). Nevertheless, there are more information provided on the papers, so it may help fellow kagglers better understanding the data and baseline.</p>\n<p>Last Note: An 40MB unlabelled data is provided in this <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file\" target=\"_blank\">repo</a>. Maybe use it for pseudolabeling, provided that it's not part of training data. </p>\n<p>A very interesting competition. Hope that everyone can enjoy it. 😀</p>",
  "messages": [
    {
      "id": "2623734",
      "postDate": "01/28/2024 10:56:34",
      "content": "<p>Hi, kagglers.</p>\n<p>I want to bring in some discoveries through online surfing. In 2023, researchers from medical schools have jointly collected data and published it (<a href=\"https://bdsp.io/content/bdsp-sparcnet/1.1/\" target=\"_blank\">paper</a>). I am pretty sure the data described in there is the same as ours, since the targets and authors are exactly the same. Two models were built as an attempt to classify eegs into 6 patterns - i.e. seizure, lpd, gpd, lrda, grda, other. One is SPaRCNet <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file\" target=\"_blank\">(GitHub repo)</a> and the other is IRR <a href=\"https://github.com/bdsp-core/IIIC-IRR\" target=\"_blank\">(GitHub repo)</a> along with two papers (<a href=\"https://pubmed.ncbi.nlm.nih.gov/36878708/\" target=\"_blank\">p1</a> and <a href=\"https://arxiv.org/pdf/2211.05207.pdf\" target=\"_blank\">p2</a>). </p>\n<p>I am not sure of how effective the models are, but it doesn't seem to be very effective. So I assume that's why they host a competition on Kaggle (hopefully we can build a better solution). Nevertheless, there are more information provided on the papers, so it may help fellow kagglers better understanding the data and baseline.</p>\n<p>Last Note: An 40MB unlabelled data is provided in this <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file\" target=\"_blank\">repo</a>. Maybe use it for pseudolabeling, provided that it's not part of training data. </p>\n<p>A very interesting competition. Hope that everyone can enjoy it. 😀</p>",
      "rawMarkdown": "Hi, kagglers.\n\nI want to bring in some discoveries through online surfing. In 2023, researchers from medical schools have jointly collected data and published it ([paper](https://bdsp.io/content/bdsp-sparcnet/1.1/)). I am pretty sure the data described in there is the same as ours, since the targets and authors are exactly the same. Two models were built as an attempt to classify eegs into 6 patterns - i.e. seizure, lpd, gpd, lrda, grda, other. One is SPaRCNet [(GitHub repo)](https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file) and the other is IRR [(GitHub repo)](https://github.com/bdsp-core/IIIC-IRR) along with two papers ([p1](https://pubmed.ncbi.nlm.nih.gov/36878708/) and [p2](https://arxiv.org/pdf/2211.05207.pdf)). \n\nI am not sure of how effective the models are, but it doesn't seem to be very effective. So I assume that's why they host a competition on Kaggle (hopefully we can build a better solution). Nevertheless, there are more information provided on the papers, so it may help fellow kagglers better understanding the data and baseline.\n\nLast Note: An 40MB unlabelled data is provided in this [repo](https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file). Maybe use it for pseudolabeling, provided that it's not part of training data. \n\nA very interesting competition. Hope that everyone can enjoy it. 😀",
      "votes": null
    },
    {
      "id": "2623743",
      "postDate": "01/28/2024 11:01:02",
      "content": "<p><a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> nice findings and thanks for sharing </p>",
      "rawMarkdown": "renyiwei nice findings and thanks for sharing",
      "votes": null
    },
    {
      "id": "2623864",
      "postDate": "01/28/2024 12:06:20",
      "content": "<p><a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> it having 25 labels samples too. will explore more and update you, if any extra data can be used for training.</p>",
      "rawMarkdown": "renyiwei it having 25 labels samples too. will explore more and update you, if any extra data can be used for training.",
      "votes": null
    },
    {
      "id": "2624892",
      "postDate": "01/29/2024 04:21:58",
      "content": "<p>thanks for sharing </p>",
      "rawMarkdown": "thanks for sharing",
      "votes": null
    },
    {
      "id": "2626346",
      "postDate": "01/29/2024 23:13:04",
      "content": "<p>Nice find <a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a>. Has anyone been able to download the <strong>complete</strong> dataset?</p>\n<p>I wonder if AWS open-data <a href=\"https://registry.opendata.aws/bdsp-sparcnet/\" target=\"_blank\">here</a> is the only source..</p>",
      "rawMarkdown": "Nice find @renyiwei. Has anyone been able to download the **complete** dataset?\n\nI wonder if AWS open-data [here](https://registry.opendata.aws/bdsp-sparcnet/) is the only source..",
      "votes": null
    },
    {
      "id": "2627521",
      "postDate": "01/30/2024 17:59:17",
      "content": "<p>SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.</p>",
      "rawMarkdown": "SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.",
      "votes": null
    },
    {
      "id": "2627543",
      "postDate": "01/30/2024 18:10:32",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, thanks for the clarification. Would the hosts be willing to share the dataset on Kaggle? </p>\n<p>It looks like an AWS account is required to access the data (which requires providing payment information when signing up for an account). I think a Kaggle dataset would make this more accessible for all competition participants.</p>",
      "rawMarkdown": "Hi @sohier, thanks for the clarification. Would the hosts be willing to share the dataset on Kaggle? \n\nIt looks like an AWS account is required to access the data (which requires providing payment information when signing up for an account). I think a Kaggle dataset would make this more accessible for all competition participants.",
      "votes": null
    },
    {
      "id": "2627705",
      "postDate": "01/30/2024 20:21:09",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Does the SPaRCNet dataset provide more labeled data (in addition to the train data that Kaggle provided)? If so, then those who sign up with Amazon will have an advantage compared with those who do not.</p>\n<p>If SPaRCNet dataset has more data, then I agree with Bartley that the best solution is to make the additional data available to everyone on Kaggle instead of requiring all competition participants to sign up for an Amazon account to access the additional data.</p>",
      "rawMarkdown": "sohier Does the SPaRCNet dataset provide more labeled data (in addition to the train data that Kaggle provided)? If so, then those who sign up with Amazon will have an advantage compared with those who do not.\n\nIf SPaRCNet dataset has more data, then I agree with Bartley that the best solution is to make the additional data available to everyone on Kaggle instead of requiring all competition participants to sign up for an Amazon account to access the additional data.",
      "votes": null
    },
    {
      "id": "2631550",
      "postDate": "02/01/2024 19:44:33",
      "content": "<p>Totally agree with Chris</p>",
      "rawMarkdown": "Totally agree with Chris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2623743,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/28/2024 11:01:02",
      "content": "<p><a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> nice findings and thanks for sharing </p>",
      "votes": null,
      "replies": [
        {
          "id": 2623864,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/28/2024 12:06:20",
          "content": "<p><a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> it having 25 labels samples too. will explore more and update you, if any extra data can be used for training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2624892,
      "author_name": "alanyjw999",
      "author_url": "",
      "post_date": "01/29/2024 04:21:58",
      "content": "<p>thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2626346,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "01/29/2024 23:13:04",
      "content": "<p>Nice find <a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a>. Has anyone been able to download the <strong>complete</strong> dataset?</p>\n<p>I wonder if AWS open-data <a href=\"https://registry.opendata.aws/bdsp-sparcnet/\" target=\"_blank\">here</a> is the only source..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2627521,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "01/30/2024 17:59:17",
      "content": "<p>SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2627543,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "01/30/2024 18:10:32",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, thanks for the clarification. Would the hosts be willing to share the dataset on Kaggle? </p>\n<p>It looks like an AWS account is required to access the data (which requires providing payment information when signing up for an account). I think a Kaggle dataset would make this more accessible for all competition participants.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2627705,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/30/2024 20:21:09",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Does the SPaRCNet dataset provide more labeled data (in addition to the train data that Kaggle provided)? If so, then those who sign up with Amazon will have an advantage compared with those who do not.</p>\n<p>If SPaRCNet dataset has more data, then I agree with Bartley that the best solution is to make the additional data available to everyone on Kaggle instead of requiring all competition participants to sign up for an Amazon account to access the additional data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2631550,
              "author_name": "rafaelzimmermann1",
              "author_url": "",
              "post_date": "02/01/2024 19:44:33",
              "content": "<p>Totally agree with Chris</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2623734": "Hi, kagglers.\n\nI want to bring in some discoveries through online surfing. In 2023, researchers from medical schools have jointly collected data and published it ([paper](https://bdsp.io/content/bdsp-sparcnet/1.1/)). I am pretty sure the data described in there is the same as ours, since the targets and authors are exactly the same. Two models were built as an attempt to classify eegs into 6 patterns - i.e. seizure, lpd, gpd, lrda, grda, other. One is SPaRCNet [(GitHub repo)](https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file) and the other is IRR [(GitHub repo)](https://github.com/bdsp-core/IIIC-IRR) along with two papers ([p1](https://pubmed.ncbi.nlm.nih.gov/36878708/) and [p2](https://arxiv.org/pdf/2211.05207.pdf)). \n\nI am not sure of how effective the models are, but it doesn't seem to be very effective. So I assume that's why they host a competition on Kaggle (hopefully we can build a better solution). Nevertheless, there are more information provided on the papers, so it may help fellow kagglers better understanding the data and baseline.\n\nLast Note: An 40MB unlabelled data is provided in this [repo](https://github.com/bdsp-core/IIIC-SPaRCNet#License-1-ov-file). Maybe use it for pseudolabeling, provided that it's not part of training data. \n\nA very interesting competition. Hope that everyone can enjoy it. 😀",
    "2623743": "renyiwei nice findings and thanks for sharing",
    "2623864": "renyiwei it having 25 labels samples too. will explore more and update you, if any extra data can be used for training.",
    "2624892": "thanks for sharing",
    "2626346": "Nice find @renyiwei. Has anyone been able to download the **complete** dataset?\n\nI wonder if AWS open-data [here](https://registry.opendata.aws/bdsp-sparcnet/) is the only source..",
    "2627521": "SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.",
    "2627543": "Hi @sohier, thanks for the clarification. Would the hosts be willing to share the dataset on Kaggle? \n\nIt looks like an AWS account is required to access the data (which requires providing payment information when signing up for an account). I think a Kaggle dataset would make this more accessible for all competition participants.",
    "2627705": "sohier Does the SPaRCNet dataset provide more labeled data (in addition to the train data that Kaggle provided)? If so, then those who sign up with Amazon will have an advantage compared with those who do not.\n\nIf SPaRCNet dataset has more data, then I agree with Bartley that the best solution is to make the additional data available to everyone on Kaggle instead of requiring all competition participants to sign up for an Amazon account to access the additional data.",
    "2631550": "Totally agree with Chris"
  },
  "source": "meta"
}