{
  "id": 475292,
  "title": "SPaRCNet Dataset Update",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/475292",
  "author_name": "Bartley",
  "post_date": "2024-02-07T19:08:50.030000",
  "votes": 23,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I want to start a discussion thread specific to the <a href=\"https://registry.opendata.aws/bdsp-sparcnet/\" target=\"_blank\">SPaRCNet Dataset</a>. This dataset was first mentioned by <a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439#2627521\" target=\"_blank\">here</a>, but there has been little activity in this thread over the past week.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> hints that it is a \"rough equivalent\" of the competition data, and therefore I suspect the data could be useful to competitors.</p>\n<p>Are there plans to release this data on Kaggle? If not, is it acceptable under the competition rules?</p>\n<pre><code>C.  . You may   other than the Competition  (“ ”) to develop and test your Submissions. However, you will ensure the   is publicly available and equally accessible to  by  participants of the Competition for purposes of the competition at no cost to the other participants. The ability to    under this Section C ( ) does not limit your other obligations under these Competition Rules, including but not limited to Section  (Winners Obligations).\n</code></pre>",
  "messages": [
    {
      "id": 2641941,
      "postDate": "2024-02-07T19:08:50.030Z",
      "content": "<p>I want to start a discussion thread specific to the <a href=\"https://registry.opendata.aws/bdsp-sparcnet/\" target=\"_blank\">SPaRCNet Dataset</a>. This dataset was first mentioned by <a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439#2627521\" target=\"_blank\">here</a>, but there has been little activity in this thread over the past week.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> hints that it is a \"rough equivalent\" of the competition data, and therefore I suspect the data could be useful to competitors.</p>\n<p>Are there plans to release this data on Kaggle? If not, is it acceptable under the competition rules?</p>\n<pre><code>C.  . You may   other than the Competition  (“ ”) to develop and test your Submissions. However, you will ensure the   is publicly available and equally accessible to  by  participants of the Competition for purposes of the competition at no cost to the other participants. The ability to    under this Section C ( ) does not limit your other obligations under these Competition Rules, including but not limited to Section  (Winners Obligations).\n</code></pre>",
      "rawMarkdown": "I want to start a discussion thread specific to the [SPaRCNet Dataset](https://registry.opendata.aws/bdsp-sparcnet/). This dataset was first mentioned by @renyiwei [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439#2627521), but there has been little activity in this thread over the past week.\n\n@sohier hints that it is a \"rough equivalent\" of the competition data, and therefore I suspect the data could be useful to competitors.\n\nAre there plans to release this data on Kaggle? If not, is it acceptable under the competition rules?\n\n```\nC. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n```",
      "votes": 23
    },
    {
      "id": 2642052,
      "postDate": "2024-02-07T21:35:13.643Z",
      "content": "<p>It's the rough equivalent in the sense that it's less processed. For example, the null values were all encoded as zeroes or <code>9999</code>. I don't expect that access would provide an advantage. Please respect the researcher's time by not spamming them with access requests unless you are planning on using the data for other research.</p>",
      "rawMarkdown": "It's the rough equivalent in the sense that it's less processed. For example, the null values were all encoded as zeroes or `9999`. I don't expect that access would provide an advantage. Please respect the researcher's time by not spamming them with access requests unless you are planning on using the data for other research.",
      "votes": 5,
      "replies": [
        {
          "id": 2642057,
          "postDate": "2024-02-07T21:41:14.353Z",
          "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>. Does this mean that the data is not acceptable under the Competition Rules?</p>",
          "rawMarkdown": "Thanks for the reply @sohier. Does this mean that the data is not acceptable under the Competition Rules?"
        },
        {
          "id": 2642546,
          "postDate": "2024-02-08T08:27:18.623Z",
          "content": "<p>Indeed. I think host must have fully considered the accessibility of data before deciding to cooperate with Kaggle. In the GitHub repo, they suggest reseachers may contact them for private dataset they've endeavored to collect for research purposes. That being said, spamming them at the hope of gaining more data is very unethical and prohibited. </p>",
          "rawMarkdown": "Indeed. I think host must have fully considered the accessibility of data before deciding to cooperate with Kaggle. In the GitHub repo, they suggest reseachers may contact them for private dataset they've endeavored to collect for research purposes. That being said, spamming them at the hope of gaining more data is very unethical and prohibited. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2645995,
      "postDate": "2024-02-10T15:44:04.113Z",
      "content": "<p>I think instead of the data it would be more helpful if someone could decipher their model and the way they processed the data. The structure of the SPaRCeNet has been mentioned in the appendix section of the paper but their data processing has not been discussed. </p>",
      "rawMarkdown": "I think instead of the data it would be more helpful if someone could decipher their model and the way they processed the data. The structure of the SPaRCeNet has been mentioned in the appendix section of the paper but their data processing has not been discussed. ",
      "votes": 1,
      "replies": [
        {
          "id": 2646009,
          "postDate": "2024-02-10T15:53:09.710Z",
          "content": "<p>Good point <a href=\"https://www.kaggle.com/swapnilbmebuet\" target=\"_blank\">@swapnilbmebuet</a>, there is some interesting info there.</p>\n<p><a href=\"https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\" target=\"_blank\">Here</a> is the appendix section if anyone if interested.</p>",
          "rawMarkdown": "Good point @swapnilbmebuet, there is some interesting info there.\n\n[Here](https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf) is the appendix section if anyone if interested.",
          "votes": 2
        },
        {
          "id": 2650689,
          "postDate": "2024-02-13T16:28:01.447Z",
          "content": "<p>I tried to inference with the model but I couldn't figure out their preprocessing and got the worst possible score from the predictions.</p>",
          "rawMarkdown": "I tried to inference with the model but I couldn't figure out their preprocessing and got the worst possible score from the predictions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2642040,
      "postDate": "2024-02-07T21:20:03.720Z",
      "content": "<p>I didn't study the dataset, nor did I see the work needed to make it usable in the competition, but one thing I can say, my models improve the more data they use, so yes, it makes sense</p>",
      "rawMarkdown": "I didn't study the dataset, nor did I see the work needed to make it usable in the competition, but one thing I can say, my models improve the more data they use, so yes, it makes sense",
      "votes": 2,
      "replies": [
        {
          "id": 2642044,
          "postDate": "2024-02-07T21:23:51.880Z",
          "content": "<p>Agreed, more data helps for me too. Here is quote from the SPaRCNet description.</p>\n<pre><code> IIIC dataset includes , labeled EEG samples from , patients' and , EEGs that were annotated by physician experts from  institutions\n</code></pre>\n<hr>\n<p>The EEG labelling section indicates that the data format is the same. Here are some relevant snippets from that section. </p>\n<pre><code>EEG Labeling: Labeling  - EEG  was done .... Experts could pan       target , change montages,  adjust  signal gain.... A -minute spectrogram was provided  additional context...\n</code></pre>",
          "rawMarkdown": "Agreed, more data helps for me too. Here is quote from the SPaRCNet description.\n\n```\nThe IIIC dataset includes 50,697 labeled EEG samples from 2,711 patients' and 6,095 EEGs that were annotated by physician experts from 18 institutions\n```\n\n---\n\nThe EEG labelling section indicates that the data format is the same. Here are some relevant snippets from that section. \n\n```\nEEG Labeling: Labeling of 10-second EEG segments was done using.... Experts could pan 20 seconds before or after the target segment, change montages, and adjust the signal gain.... A 10-minute spectrogram was provided for additional context...\n```",
          "votes": 1,
          "replies": [
            {
              "id": 2642137,
              "postDate": "2024-02-08T00:08:48.173Z",
              "content": "<p>Very interesting, it actually looks very similar<br>\nDoes it say anything about whether any stimulus is used during the experiments?</p>",
              "rawMarkdown": "Very interesting, it actually looks very similar\nDoes it say anything about whether any stimulus is used during the experiments?"
            }
          ]
        }
      ]
    },
    {
      "id": 2642691,
      "postDate": "2024-02-08T10:52:49.580Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2645518,
      "postDate": "2024-02-10T09:49:47.480Z",
      "content": "<p>thanks. really helpful</p>",
      "rawMarkdown": "thanks. really helpful"
    }
  ],
  "comments": [
    {
      "id": 2642052,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2024-02-07T21:35:13.643000",
      "content": "<p>It's the rough equivalent in the sense that it's less processed. For example, the null values were all encoded as zeroes or <code>9999</code>. I don't expect that access would provide an advantage. Please respect the researcher's time by not spamming them with access requests unless you are planning on using the data for other research.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2642057,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-02-07T21:41:14.353000",
          "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>. Does this mean that the data is not acceptable under the Competition Rules?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2642546,
          "author_name": "Roy Wei",
          "author_url": "",
          "post_date": "2024-02-08T08:27:18.623000",
          "content": "<p>Indeed. I think host must have fully considered the accessibility of data before deciding to cooperate with Kaggle. In the GitHub repo, they suggest reseachers may contact them for private dataset they've endeavored to collect for research purposes. That being said, spamming them at the hope of gaining more data is very unethical and prohibited. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2645995,
      "author_name": "Swapnil BME BUET",
      "author_url": "",
      "post_date": "2024-02-10T15:44:04.113000",
      "content": "<p>I think instead of the data it would be more helpful if someone could decipher their model and the way they processed the data. The structure of the SPaRCeNet has been mentioned in the appendix section of the paper but their data processing has not been discussed. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2646009,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-02-10T15:53:09.710000",
          "content": "<p>Good point <a href=\"https://www.kaggle.com/swapnilbmebuet\" target=\"_blank\">@swapnilbmebuet</a>, there is some interesting info there.</p>\n<p><a href=\"https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\" target=\"_blank\">Here</a> is the appendix section if anyone if interested.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2650689,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-02-13T16:28:01.447000",
          "content": "<p>I tried to inference with the model but I couldn't figure out their preprocessing and got the worst possible score from the predictions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2642040,
      "author_name": "Rafael Zimmermann",
      "author_url": "",
      "post_date": "2024-02-07T21:20:03.720000",
      "content": "<p>I didn't study the dataset, nor did I see the work needed to make it usable in the competition, but one thing I can say, my models improve the more data they use, so yes, it makes sense</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2642044,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-02-07T21:23:51.880000",
          "content": "<p>Agreed, more data helps for me too. Here is quote from the SPaRCNet description.</p>\n<pre><code> IIIC dataset includes , labeled EEG samples from , patients' and , EEGs that were annotated by physician experts from  institutions\n</code></pre>\n<hr>\n<p>The EEG labelling section indicates that the data format is the same. Here are some relevant snippets from that section. </p>\n<pre><code>EEG Labeling: Labeling  - EEG  was done .... Experts could pan       target , change montages,  adjust  signal gain.... A -minute spectrogram was provided  additional context...\n</code></pre>",
          "votes": 1,
          "replies": [
            {
              "id": 2642137,
              "author_name": "Rafael Zimmermann",
              "author_url": "",
              "post_date": "2024-02-08T00:08:48.173000",
              "content": "<p>Very interesting, it actually looks very similar<br>\nDoes it say anything about whether any stimulus is used during the experiments?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2642691,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-08T10:52:49.580000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2645518,
      "author_name": "Ganesh Talwar",
      "author_url": "",
      "post_date": "2024-02-10T09:49:47.480000",
      "content": "<p>thanks. really helpful</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2641941": "I want to start a discussion thread specific to the [SPaRCNet Dataset](https://registry.opendata.aws/bdsp-sparcnet/). This dataset was first mentioned by @renyiwei [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439#2627521), but there has been little activity in this thread over the past week.\n\n@sohier hints that it is a \"rough equivalent\" of the competition data, and therefore I suspect the data could be useful to competitors.\n\nAre there plans to release this data on Kaggle? If not, is it acceptable under the competition rules?\n\n```\nC. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n```",
    "2642052": "It's the rough equivalent in the sense that it's less processed. For example, the null values were all encoded as zeroes or `9999`. I don't expect that access would provide an advantage. Please respect the researcher's time by not spamming them with access requests unless you are planning on using the data for other research.",
    "2645995": "I think instead of the data it would be more helpful if someone could decipher their model and the way they processed the data. The structure of the SPaRCeNet has been mentioned in the appendix section of the paper but their data processing has not been discussed. ",
    "2642040": "I didn't study the dataset, nor did I see the work needed to make it usable in the competition, but one thing I can say, my models improve the more data they use, so yes, it makes sense",
    "2642691": "",
    "2645518": "thanks. really helpful"
  }
}