{
  "id": 492388,
  "title": "Details in 1D model in 4th place solution (bilzard's part)",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492388",
  "author_name": "",
  "post_date": "2024-04-09T13:52:50.106291Z",
  "votes": 40,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Big thanks to the hosts and the Kaggle team for setting up this competition. Diving into raw EEG signals was really fun. Also thanks to my teammates <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a>, <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a>, and <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> for being awesome partners.</p>\n<p>Here’s a detailed description of 1D model (bilzard's part) of our solution. The summary of our team solution is also available <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492240\" target=\"_blank\">here</a>.</p>\n<p>Edit:<br>\nWe open-sourced source code repository <a href=\"https://github.com/bilzard/kaggle-hms-bilzard\" target=\"_blank\">here</a> (Apache 2.0).</p>\n<p>Edit: we shared revision of this article with more detailed explanation of CQF here:<br>\n<a href=\"https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\" target=\"_blank\">https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf</a></p>\n<h3>1. Outline</h3>\n<p>Initially, my solution was an ensemble of 1D and 2D models. However, after merging teams and recognizing that my teammates had already developed strong 2D models, I shifted my focus primarily towards 1D models (i.e. the models process raw EEG signals directly).</p>\n<h3>2. Basic Concept</h3>\n<p>The approach to our 1D modeling is twofold:</p>\n<ol>\n<li><strong>L/R Symmetric Modeling</strong>: We aimed to maintain symmetry in the model to facilitate the detection of Laterality.</li>\n<li><strong>Channel Quality Factor (CQF)</strong>: This is used to evaluate the quality of EEG channels and identify any that are suboptimal.</li>\n</ol>\n<p>We will discuss these on the following sections.</p>\n<h3>3. L/R Symmetric modeling</h3>\n<p>In this section, we'll cover the basics of <em>L/R Symmetric Modeling</em>.</p>\n<p>Our inspiration came from textbooks[2], where we learned that the differences in signals from the left and right channels are key to telling General seizures apart from Lateral ones. That led us to design a model that treats left and right brain signals symmetrically, as illustrated in Fig. 1.</p>\n<p>Here’s how it works: The left and right brain signals are fed into a 1D CNN model separately, but they share the same parameters. The features we get from these are then processed in two ways: 1) through L/R invariant mapping, and 2) with a similarity encoder. We’ll dive into these parts more in the sections that follow.</p>\n<p>An important detail is our use of a late fusion approach. This means we input each channel separately into the 1D encoder to get a more abstract representation of the EEG channels.</p>\n<p>We also discovered that gradually blending the EEG channels partway through the 1D encoder improved our results. We call this the Channel Mixer. While we didn’t explore every possible architecture, adding a simple 1D convolution layer in each block was enough to see significant benefits.</p>\n<h4>3.1. L/R invariant mapping</h4>\n<p>This part of our model pulls out features that don’t change whether we swap left and right EEG channels. We do this using two mathematical functions, $f$ and $g$, defined as follows:</p>\n<p>$$<br>\nu = f(x, y) \\<br>\nv = g(x, y) \\<br>\n\\text{ where } f(x, y) = f(y, x) \\<br>\n\\text{ and } g(x, y) = g(y, x)<br>\n$$</p>\n<p>We considered several options for these functions, like:</p>\n<ol>\n<li>diff, mean</li>\n<li>prod, sum</li>\n<li>min, max</li>\n</ol>\n<p>In our initial tests, we found that the choice of function didn’t drastically change the outcome. So, we settled on min/max as our default functions.</p>\n<h4>3.2. Similarity Encoder</h4>\n<p>To pinpoint the differences between the left and right sides, we measure the cosine similarity between paired L/R channels (for example, (Fp1, Fp2), (O1, O2), etc.). After calculating the similarity, we project it onto a $d$-dimensional vector to form a feature vector. This approach proved to be more effective than using the one-dimensional cosine similarity directly.</p>\n<p><img src=\"https://i.postimg.cc/NGT1JcxK/hms-1d.png\" alt=\"hms-1d\"><br></p>\n<p><strong>Fig. 1 Architecture of 1D Model Proposed</strong></p>\n<h3>4. Channel Quality Factor (CQF)</h3>\n<p>During our initial EDA, we noticed a significant presence of bad channels in the EEG signals (see Fig. 2). This observation led us to devote effort to assessing channel quality, aiming to reduce potential confusion for our model.</p>\n<p>The challenge of identifying corrupted channels is a well-known issue in EEG research. We discovered a study [1] that addresses this problem by introducing the <em>Locality Factor (LOF)</em>. LOF assesses the quality of a channel based on its discrepancy from other (usually normal) channels, with discrepancy measured by distance metrics like the Euclidean distance.</p>\n<p>While [1] considers channels in a static manner, we adapted LOF for use with time-series data, allowing us to capture nuanced aspects of channel quality over time. We refer to this adapted feature as the Channel Quality Factor (CQF), illustrated in Fig. 3.</p>\n<p><img src=\"https://i.postimg.cc/tCjF8Mr2/bad-channel-sample01.png\" alt=\"bad-channel-sample01\"><br><br></p>\n<p><strong>Fig. 2: Example of Bad Channels</strong></p>\n<p><img src=\"https://i.postimg.cc/YqzQ7zZk/cqm002.jpg\" alt=\"cqm002\"><br><br></p>\n<p><strong>Fig. 3: CQF Illustration</strong>: Channels of lower quality are indicated in cooler colors.</p>\n<h3>6. Result</h3>\n<p>The cross-validation scores (CVs) of our 1D models in the final submissions are listed below. These CVs were calculated using samples with <code>num_votes &gt; 8</code>.</p>\n<pre><code> .\n</code></pre>\n<p>We should emphasize our 1D model is really light weight. It only has 863K parameters and took 1-2 minutes for infer with test set with 15 models (3 seed x 5 fold ensemble).</p>\n<h3>7. Reference</h3>\n<ul>\n<li>[1] Velu et. al., Adaptable and Robust EEG Bad Channel Detection Using Local Outlier Factor (LOF)</li>\n<li>[2] American Clinical Neurophysiology Society’s Standardized Critical Care EEG Terminology: 2021 Version</li>\n<li>[3] <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">https://github.com/huggingface/pytorch-image-models</a></li>\n</ul>\n<h3>Appendix</h3>\n<h4>A1. Training Configuration</h4>\n<p>Most participants 2-stage training scheme, however, we schedule min_vote during 1-stage training schedule.</p>\n<ol>\n<li>0-15 epoch: train with all data</li>\n<li>16-23 epoch: train with n_votes&gt;=8.4</li>\n</ol>\n<p>Other training configurations are:</p>\n<ul>\n<li>Optimizer: Adam</li>\n<li>lr: 1e-3</li>\n<li>batch_size: 32</li>\n<li>weight_decay: 1e-5</li>\n<li>Scheduler: cosine</li>\n</ul>\n<h4>A2. Augmentations</h4>\n<ul>\n<li>channel shuffling: shuffling channels while keeping L/R symmetry</li>\n<li>L/R swap</li>\n<li>dropout (&lt;=128 frames x 4)</li>\n<li>cutmix</li>\n</ul>\n<h4>A3. Backbone 1D CNN Architecture</h4>\n<p>We designed 1D-CNN based on timm[3]'s EfficientNet2d's building blocks. We called this architecture as <em>EfficientNet1d</em>.</p>\n<pre><code>========================================================================================================================\nLayer (type:depth-idx)                                                 Output Shape              Param #\n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n========================================================================================================================\n</code></pre>",
  "messages": [
    {
      "id": "2743516",
      "postDate": "04/09/2024 13:52:50",
      "content": "<p>Big thanks to the hosts and the Kaggle team for setting up this competition. Diving into raw EEG signals was really fun. Also thanks to my teammates <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a>, <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a>, and <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> for being awesome partners.</p>\n<p>Here’s a detailed description of 1D model (bilzard's part) of our solution. The summary of our team solution is also available <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492240\" target=\"_blank\">here</a>.</p>\n<p>Edit:<br>\nWe open-sourced source code repository <a href=\"https://github.com/bilzard/kaggle-hms-bilzard\" target=\"_blank\">here</a> (Apache 2.0).</p>\n<p>Edit: we shared revision of this article with more detailed explanation of CQF here:<br>\n<a href=\"https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\" target=\"_blank\">https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf</a></p>\n<h3>1. Outline</h3>\n<p>Initially, my solution was an ensemble of 1D and 2D models. However, after merging teams and recognizing that my teammates had already developed strong 2D models, I shifted my focus primarily towards 1D models (i.e. the models process raw EEG signals directly).</p>\n<h3>2. Basic Concept</h3>\n<p>The approach to our 1D modeling is twofold:</p>\n<ol>\n<li><strong>L/R Symmetric Modeling</strong>: We aimed to maintain symmetry in the model to facilitate the detection of Laterality.</li>\n<li><strong>Channel Quality Factor (CQF)</strong>: This is used to evaluate the quality of EEG channels and identify any that are suboptimal.</li>\n</ol>\n<p>We will discuss these on the following sections.</p>\n<h3>3. L/R Symmetric modeling</h3>\n<p>In this section, we'll cover the basics of <em>L/R Symmetric Modeling</em>.</p>\n<p>Our inspiration came from textbooks[2], where we learned that the differences in signals from the left and right channels are key to telling General seizures apart from Lateral ones. That led us to design a model that treats left and right brain signals symmetrically, as illustrated in Fig. 1.</p>\n<p>Here’s how it works: The left and right brain signals are fed into a 1D CNN model separately, but they share the same parameters. The features we get from these are then processed in two ways: 1) through L/R invariant mapping, and 2) with a similarity encoder. We’ll dive into these parts more in the sections that follow.</p>\n<p>An important detail is our use of a late fusion approach. This means we input each channel separately into the 1D encoder to get a more abstract representation of the EEG channels.</p>\n<p>We also discovered that gradually blending the EEG channels partway through the 1D encoder improved our results. We call this the Channel Mixer. While we didn’t explore every possible architecture, adding a simple 1D convolution layer in each block was enough to see significant benefits.</p>\n<h4>3.1. L/R invariant mapping</h4>\n<p>This part of our model pulls out features that don’t change whether we swap left and right EEG channels. We do this using two mathematical functions, $f$ and $g$, defined as follows:</p>\n<p>$$<br>\nu = f(x, y) \\<br>\nv = g(x, y) \\<br>\n\\text{ where } f(x, y) = f(y, x) \\<br>\n\\text{ and } g(x, y) = g(y, x)<br>\n$$</p>\n<p>We considered several options for these functions, like:</p>\n<ol>\n<li>diff, mean</li>\n<li>prod, sum</li>\n<li>min, max</li>\n</ol>\n<p>In our initial tests, we found that the choice of function didn’t drastically change the outcome. So, we settled on min/max as our default functions.</p>\n<h4>3.2. Similarity Encoder</h4>\n<p>To pinpoint the differences between the left and right sides, we measure the cosine similarity between paired L/R channels (for example, (Fp1, Fp2), (O1, O2), etc.). After calculating the similarity, we project it onto a $d$-dimensional vector to form a feature vector. This approach proved to be more effective than using the one-dimensional cosine similarity directly.</p>\n<p><img src=\"https://i.postimg.cc/NGT1JcxK/hms-1d.png\" alt=\"hms-1d\"><br></p>\n<p><strong>Fig. 1 Architecture of 1D Model Proposed</strong></p>\n<h3>4. Channel Quality Factor (CQF)</h3>\n<p>During our initial EDA, we noticed a significant presence of bad channels in the EEG signals (see Fig. 2). This observation led us to devote effort to assessing channel quality, aiming to reduce potential confusion for our model.</p>\n<p>The challenge of identifying corrupted channels is a well-known issue in EEG research. We discovered a study [1] that addresses this problem by introducing the <em>Locality Factor (LOF)</em>. LOF assesses the quality of a channel based on its discrepancy from other (usually normal) channels, with discrepancy measured by distance metrics like the Euclidean distance.</p>\n<p>While [1] considers channels in a static manner, we adapted LOF for use with time-series data, allowing us to capture nuanced aspects of channel quality over time. We refer to this adapted feature as the Channel Quality Factor (CQF), illustrated in Fig. 3.</p>\n<p><img src=\"https://i.postimg.cc/tCjF8Mr2/bad-channel-sample01.png\" alt=\"bad-channel-sample01\"><br><br></p>\n<p><strong>Fig. 2: Example of Bad Channels</strong></p>\n<p><img src=\"https://i.postimg.cc/YqzQ7zZk/cqm002.jpg\" alt=\"cqm002\"><br><br></p>\n<p><strong>Fig. 3: CQF Illustration</strong>: Channels of lower quality are indicated in cooler colors.</p>\n<h3>6. Result</h3>\n<p>The cross-validation scores (CVs) of our 1D models in the final submissions are listed below. These CVs were calculated using samples with <code>num_votes &gt; 8</code>.</p>\n<pre><code> .\n</code></pre>\n<p>We should emphasize our 1D model is really light weight. It only has 863K parameters and took 1-2 minutes for infer with test set with 15 models (3 seed x 5 fold ensemble).</p>\n<h3>7. Reference</h3>\n<ul>\n<li>[1] Velu et. al., Adaptable and Robust EEG Bad Channel Detection Using Local Outlier Factor (LOF)</li>\n<li>[2] American Clinical Neurophysiology Society’s Standardized Critical Care EEG Terminology: 2021 Version</li>\n<li>[3] <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">https://github.com/huggingface/pytorch-image-models</a></li>\n</ul>\n<h3>Appendix</h3>\n<h4>A1. Training Configuration</h4>\n<p>Most participants 2-stage training scheme, however, we schedule min_vote during 1-stage training schedule.</p>\n<ol>\n<li>0-15 epoch: train with all data</li>\n<li>16-23 epoch: train with n_votes&gt;=8.4</li>\n</ol>\n<p>Other training configurations are:</p>\n<ul>\n<li>Optimizer: Adam</li>\n<li>lr: 1e-3</li>\n<li>batch_size: 32</li>\n<li>weight_decay: 1e-5</li>\n<li>Scheduler: cosine</li>\n</ul>\n<h4>A2. Augmentations</h4>\n<ul>\n<li>channel shuffling: shuffling channels while keeping L/R symmetry</li>\n<li>L/R swap</li>\n<li>dropout (&lt;=128 frames x 4)</li>\n<li>cutmix</li>\n</ul>\n<h4>A3. Backbone 1D CNN Architecture</h4>\n<p>We designed 1D-CNN based on timm[3]'s EfficientNet2d's building blocks. We called this architecture as <em>EfficientNet1d</em>.</p>\n<pre><code>========================================================================================================================\nLayer (type:depth-idx)                                                 Output Shape              Param #\n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n========================================================================================================================\n</code></pre>",
      "rawMarkdown": "Big thanks to the hosts and the Kaggle team for setting up this competition. Diving into raw EEG signals was really fun. Also thanks to my teammates @yujiariyasu, @tattaka, and @ren4yu for being awesome partners.\n\nHere’s a detailed description of 1D model (bilzard's part) of our solution. The summary of our team solution is also available [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492240).\n\nEdit:\nWe open-sourced source code repository [here](https://github.com/bilzard/kaggle-hms-bilzard) (Apache 2.0).\n\nEdit: we shared revision of this article with more detailed explanation of CQF here:\nhttps://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\n\n### 1. Outline\n\nInitially, my solution was an ensemble of 1D and 2D models. However, after merging teams and recognizing that my teammates had already developed strong 2D models, I shifted my focus primarily towards 1D models (i.e. the models process raw EEG signals directly).\n\n### 2. Basic Concept\n\nThe approach to our 1D modeling is twofold:\n\n1. **L/R Symmetric Modeling**: We aimed to maintain symmetry in the model to facilitate the detection of Laterality.\n2. **Channel Quality Factor (CQF)**: This is used to evaluate the quality of EEG channels and identify any that are suboptimal.\n\nWe will discuss these on the following sections.\n\n### 3. L/R Symmetric modeling\n\nIn this section, we'll cover the basics of *L/R Symmetric Modeling*.\n\nOur inspiration came from textbooks[2], where we learned that the differences in signals from the left and right channels are key to telling General seizures apart from Lateral ones. That led us to design a model that treats left and right brain signals symmetrically, as illustrated in Fig. 1.\n\nHere’s how it works: The left and right brain signals are fed into a 1D CNN model separately, but they share the same parameters. The features we get from these are then processed in two ways: 1) through L/R invariant mapping, and 2) with a similarity encoder. We’ll dive into these parts more in the sections that follow.\n\nAn important detail is our use of a late fusion approach. This means we input each channel separately into the 1D encoder to get a more abstract representation of the EEG channels.\n\nWe also discovered that gradually blending the EEG channels partway through the 1D encoder improved our results. We call this the Channel Mixer. While we didn’t explore every possible architecture, adding a simple 1D convolution layer in each block was enough to see significant benefits.\n\n#### 3.1. L/R invariant mapping\n\nThis part of our model pulls out features that don’t change whether we swap left and right EEG channels. We do this using two mathematical functions, $f$ and $g$, defined as follows:\n\n$$\nu = f(x, y) \\\\\nv = g(x, y) \\\\\n\\text{ where } f(x, y) = f(y, x) \\\\\n\\text{ and } g(x, y) = g(y, x)\n$$\n\nWe considered several options for these functions, like:\n\n1. diff, mean\n1. prod, sum\n1. min, max\n\nIn our initial tests, we found that the choice of function didn’t drastically change the outcome. So, we settled on min/max as our default functions.\n\n#### 3.2. Similarity Encoder\n\nTo pinpoint the differences between the left and right sides, we measure the cosine similarity between paired L/R channels (for example, (Fp1, Fp2), (O1, O2), etc.). After calculating the similarity, we project it onto a $d$-dimensional vector to form a feature vector. This approach proved to be more effective than using the one-dimensional cosine similarity directly.\n\n<img src=\"https://i.postimg.cc/NGT1JcxK/hms-1d.png\" alt=\"hms-1d\" width=\"480\"/><br/>\n\n**Fig. 1 Architecture of 1D Model Proposed**\n\n### 4. Channel Quality Factor (CQF)\n\nDuring our initial EDA, we noticed a significant presence of bad channels in the EEG signals (see Fig. 2). This observation led us to devote effort to assessing channel quality, aiming to reduce potential confusion for our model.\n\nThe challenge of identifying corrupted channels is a well-known issue in EEG research. We discovered a study [1] that addresses this problem by introducing the *Locality Factor (LOF)*. LOF assesses the quality of a channel based on its discrepancy from other (usually normal) channels, with discrepancy measured by distance metrics like the Euclidean distance.\n\nWhile [1] considers channels in a static manner, we adapted LOF for use with time-series data, allowing us to capture nuanced aspects of channel quality over time. We refer to this adapted feature as the Channel Quality Factor (CQF), illustrated in Fig. 3.\n\n<img src=\"https://i.postimg.cc/tCjF8Mr2/bad-channel-sample01.png\" alt=\"bad-channel-sample01\" width=\"640\"/><br/><br/>\n\n**Fig. 2: Example of Bad Channels**\n\n<img src=\"https://i.postimg.cc/YqzQ7zZk/cqm002.jpg\" alt=\"cqm002\" width=\"640\"/><br/><br/>\n\n**Fig. 3: CQF Illustration**: Channels of lower quality are indicated in cooler colors.\n\n### 6. Result\n\nThe cross-validation scores (CVs) of our 1D models in the final submissions are listed below. These CVs were calculated using samples with `num_votes > 8`.\n\n```\nv5_eeg_24ep_cutmix 0.2477\n```\n\nWe should emphasize our 1D model is really light weight. It only has 863K parameters and took 1-2 minutes for infer with test set with 15 models (3 seed x 5 fold ensemble).\n\n### 7. Reference\n\n- [1] Velu et. al., Adaptable and Robust EEG Bad Channel Detection Using Local Outlier Factor (LOF)\n- [2] American Clinical Neurophysiology Society’s Standardized Critical Care EEG Terminology: 2021 Version\n- [3] https://github.com/huggingface/pytorch-image-models\n\n### Appendix\n\n#### A1. Training Configuration\n\nMost participants 2-stage training scheme, however, we schedule min_vote during 1-stage training schedule.\n\n1. 0-15 epoch: train with all data\n1. 16-23 epoch: train with n_votes>=8.4\n\nOther training configurations are:\n\n- Optimizer: Adam\n- lr: 1e-3\n- batch_size: 32\n- weight_decay: 1e-5\n- Scheduler: cosine\n\n#### A2. Augmentations\n\n- channel shuffling: shuffling channels while keeping L/R symmetry\n- L/R swap\n- dropout (<=128 frames x 4)\n- cutmix\n\n#### A3. Backbone 1D CNN Architecture\n\nWe designed 1D-CNN based on timm[3]'s EfficientNet2d's building blocks. We called this architecture as *EfficientNet1d*.\n\n```\n========================================================================================================================\nLayer (type:depth-idx)                                                 Output Shape              Param #\n========================================================================================================================\nEfficientNet1d                                                         [40, 64, 16]              --\n├─ConvBnAct2d: 1-1                                                     [4, 64, 10, 512]          --\n│    └─Sequential: 2-1                                                 [4, 64, 10, 512]          --\n│    │    └─Conv2d: 3-1                                                [4, 64, 10, 512]          384\n│    │    └─BatchNorm2d: 3-2                                           [4, 64, 10, 512]          128\n│    │    └─ELU: 3-3                                                   [4, 64, 10, 512]          --\n├─Sequential: 1-2                                                      --                        --\n│    └─ResBlock2d: 2-2                                                 [4, 64, 10, 256]          --\n│    │    └─MaxPool2d: 3-4                                             [4, 64, 10, 256]          --\n│    │    └─Sequential: 3-5                                            [4, 64, 10, 256]          --\n│    │    │    └─Sequential: 4-1                                       [4, 64, 10, 512]          --\n│    │    │    │    └─InvertedResidual: 5-1                            [4, 64, 10, 512]          55,104\n│    │    │    │    └─InvertedResidual: 5-2                            [4, 64, 10, 512]          55,104\n│    │    │    │    └─InvertedResidual: 5-3                            [4, 64, 10, 512]          55,104\n│    │    │    │    └─DepthWiseSeparableConv: 5-4                      [4, 64, 10, 512]          6,672\n│    │    │    └─MaxPool2d: 4-2                                        [4, 64, 10, 256]          --\n│    └─ResBlock2d: 2-3                                                 [4, 64, 10, 128]          --\n│    │    └─MaxPool2d: 3-6                                             [4, 64, 10, 128]          --\n│    │    └─Sequential: 3-7                                            [4, 64, 10, 128]          --\n│    │    │    └─Sequential: 4-3                                       [4, 64, 10, 256]          --\n│    │    │    │    └─InvertedResidual: 5-5                            [4, 64, 10, 256]          55,104\n│    │    │    │    └─InvertedResidual: 5-6                            [4, 64, 10, 256]          55,104\n│    │    │    │    └─InvertedResidual: 5-7                            [4, 64, 10, 256]          55,104\n│    │    │    │    └─DepthWiseSeparableConv: 5-8                      [4, 64, 10, 256]          6,672\n│    │    │    └─MaxPool2d: 4-4                                        [4, 64, 10, 128]          --\n│    └─ResBlock2d: 2-4                                                 [4, 64, 10, 64]           --\n│    │    └─MaxPool2d: 3-8                                             [4, 64, 10, 64]           --\n│    │    └─Sequential: 3-9                                            [4, 64, 10, 64]           --\n│    │    │    └─Sequential: 4-5                                       [4, 64, 10, 128]          --\n│    │    │    │    └─InvertedResidual: 5-9                            [4, 64, 10, 128]          55,616\n│    │    │    │    └─InvertedResidual: 5-10                           [4, 64, 10, 128]          55,616\n│    │    │    │    └─InvertedResidual: 5-11                           [4, 64, 10, 128]          55,616\n│    │    │    │    └─DepthWiseSeparableConv: 5-12                     [4, 64, 10, 128]          6,672\n│    │    │    └─MaxPool2d: 4-6                                        [4, 64, 10, 64]           --\n│    └─ResBlock2d: 2-5                                                 [4, 64, 10, 32]           --\n│    │    └─MaxPool2d: 3-10                                            [4, 64, 10, 32]           --\n│    │    └─Sequential: 3-11                                           [4, 64, 10, 32]           --\n│    │    │    └─Sequential: 4-7                                       [4, 64, 10, 64]           --\n│    │    │    │    └─InvertedResidual: 5-13                           [4, 64, 10, 64]           55,616\n│    │    │    │    └─InvertedResidual: 5-14                           [4, 64, 10, 64]           55,616\n│    │    │    │    └─InvertedResidual: 5-15                           [4, 64, 10, 64]           55,616\n│    │    │    │    └─DepthWiseSeparableConv: 5-16                     [4, 64, 10, 64]           6,672\n│    │    │    └─MaxPool2d: 4-8                                        [4, 64, 10, 32]           --\n│    └─ResBlock2d: 2-6                                                 [4, 64, 10, 16]           --\n│    │    └─MaxPool2d: 3-12                                            [4, 64, 10, 16]           --\n│    │    └─Sequential: 3-13                                           [4, 64, 10, 16]           --\n│    │    │    └─Sequential: 4-9                                       [4, 64, 10, 32]           --\n│    │    │    │    └─InvertedResidual: 5-17                           [4, 64, 10, 32]           55,104\n│    │    │    │    └─InvertedResidual: 5-18                           [4, 64, 10, 32]           55,104\n│    │    │    │    └─InvertedResidual: 5-19                           [4, 64, 10, 32]           55,104\n│    │    │    │    └─DepthWiseSeparableConv: 5-20                     [4, 64, 10, 32]           6,672\n│    │    │    └─MaxPool2d: 4-10                                       [4, 64, 10, 16]           --\n========================================================================================================================\nTotal params: 863,504\nTrainable params: 863,504\nNon-trainable params: 0\nTotal mult-adds (G): 2.72\n========================================================================================================================\nInput size (MB): 0.16\nForward/backward pass size (MB): 833.78\nParams size (MB): 3.45\nEstimated Total Size (MB): 837.40\n========================================================================================================================\n```",
      "votes": null
    },
    {
      "id": "2743554",
      "postDate": "04/09/2024 14:11:22",
      "content": "<p>This is incredible, I spent the whole comp on 1D data and this is so clever! Thanks for sharing your ideas in detail, this is more than inspirational!</p>\n<p>Best,<br>\nJan</p>",
      "rawMarkdown": "This is incredible, I spent the whole comp on 1D data and this is so clever! Thanks for sharing your ideas in detail, this is more than inspirational!\n\nBest,\nJan",
      "votes": null
    },
    {
      "id": "2743587",
      "postDate": "04/09/2024 14:28:15",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> great work and congratz on the prize zone finish, will you mind share the code, at least the preprocessing part?</p>",
      "rawMarkdown": "tatamikenn great work and congratz on the prize zone finish, will you mind share the code, at least the preprocessing part?",
      "votes": null
    },
    {
      "id": "2743601",
      "postDate": "04/09/2024 14:39:42",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> Thanks.</p>\n<blockquote>\n  <p>will you mind share the code, at least the preprocessing part?</p>\n</blockquote>\n<p>You mean CQM part? sure. Below is our snippet for generating CQM.<br>\nInput is supposed to be down-sampled to 40Hz.</p>\n<pre><code> () -&gt; pl.DataFrame:\n    eeg_df = eeg_df.with_columns(\n        \n        pl.col(p1)\n        .sub(pl.col(p2))\n        .()\n        .ewm_mean(half_life=kernel_size, min_periods=)\n        .add(eps)\n        .alias()\n         p1, p2  product(EEG_PROBES, EEG_PROBES)\n    )\n\n    idxs = []\n     p1  EEG_PROBES:\n        rx = eeg_df.select(\n              p2  EEG_PROBES\n        ).to_numpy()\n        top_k_indices = np.argsort(rx, axis=)[:,  : top_k + ]\n        top_k_values = np.take_along_axis(rx, top_k_indices, axis=)\n        top_k_dist = top_k_values.mean(axis=)\n        eeg_df = eeg_df.with_columns(\n            pl.Series(top_k_dist).alias(),\n        )\n        idxs.append(top_k_indices)\n\n    xs = eeg_df.select(  p1  EEG_PROBES).to_numpy()  \n    global_top_k_indices = np.argsort(xs, axis=)[:, :top_k]\n    global_top_k_values = np.take_along_axis(xs, global_top_k_indices, axis=)\n    global_top_k_dist = global_top_k_values.mean(axis=)\n    global_median_dist = np.median(xs, axis=)\n    idxs = np.stack(idxs, axis=)  \n\n     p1  EEG_PROBES:\n        local_top_k_idxs = idxs[:, PROBE2IDX[p1], :]\n        local_top_k_values = np.take_along_axis(xs, local_top_k_idxs, axis=)\n        local_top_k_dist = local_top_k_values.mean(axis=)\n        eeg_df = (\n            eeg_df.with_columns(\n                pl.Series(local_top_k_dist).alias(),\n                pl.Series(global_top_k_dist).alias(),\n                pl.Series(global_median_dist).alias(),\n            )\n            .with_columns(\n                pl.col()\n                .truediv(pl.col().add(eps))\n                .alias(),\n            )\n            .with_columns(\n                pl.col()\n                .truediv(pl.col().add(eps))\n                .alias(),\n            )\n        )\n\n     p  EEG_PROBES:\n        eeg_df = eeg_df.with_columns(\n            pl.col().lt(distance_threshold).alias(),\n            pl.lit()\n            .truediv(pl.col().clip().truediv(distance_threshold).add())\n            .alias(),\n        )\n\n     eeg_df\n</code></pre>",
      "rawMarkdown": "sergiosaharovskiy Thanks.\n\n> will you mind share the code, at least the preprocessing part?\n\nYou mean CQM part? sure. Below is our snippet for generating CQM.\nInput is supposed to be down-sampled to 40Hz.\n\n```python\ndef do_process_cqf(\n    eeg_df: pl.DataFrame,\n    kernel_size: int = 13,\n    top_k: int = 3,\n    eps: float = 1e-4,\n    distance_threshold: float = 10.0,\n    distance_metric: str = \"l2\",\n    normalize_type: str = \"top-k\",\n) -> pl.DataFrame:\n    eeg_df = eeg_df.with_columns(\n        # L2 distance of (p1, p2)\n        pl.col(p1)\n        .sub(pl.col(p2))\n        .pow(2)\n        .ewm_mean(half_life=kernel_size, min_periods=1)\n        .add(eps)\n        .alias(f\"l2-dist-{p1}-{p2}\")\n        for p1, p2 in product(EEG_PROBES, EEG_PROBES)\n    )\n\n    idxs = []\n    for p1 in EEG_PROBES:\n        rx = eeg_df.select(\n            f\"{distance_metric}-dist-{p1}-{p2}\" for p2 in EEG_PROBES\n        ).to_numpy()\n        top_k_indices = np.argsort(rx, axis=1)[:, 1 : top_k + 1]\n        top_k_values = np.take_along_axis(rx, top_k_indices, axis=1)\n        top_k_dist = top_k_values.mean(axis=1)\n        eeg_df = eeg_df.with_columns(\n            pl.Series(top_k_dist).alias(f\"top-k-dist-{p1}\"),\n        )\n        idxs.append(top_k_indices)\n\n    xs = eeg_df.select(f\"top-k-dist-{p1}\" for p1 in EEG_PROBES).to_numpy()  # (N, 19)\n    global_top_k_indices = np.argsort(xs, axis=1)[:, :top_k]\n    global_top_k_values = np.take_along_axis(xs, global_top_k_indices, axis=1)\n    global_top_k_dist = global_top_k_values.mean(axis=1)\n    global_median_dist = np.median(xs, axis=1)\n    idxs = np.stack(idxs, axis=1)  # (N, 19, top_k)\n\n    for p1 in EEG_PROBES:\n        local_top_k_idxs = idxs[:, PROBE2IDX[p1], :]\n        local_top_k_values = np.take_along_axis(xs, local_top_k_idxs, axis=1)\n        local_top_k_dist = local_top_k_values.mean(axis=1)\n        eeg_df = (\n            eeg_df.with_columns(\n                pl.Series(local_top_k_dist).alias(f\"local-top-k-dist-{p1}\"),\n                pl.Series(global_top_k_dist).alias(f\"global-top-k-dist-{p1}\"),\n                pl.Series(global_median_dist).alias(f\"global-median-dist-{p1}\"),\n            )\n            .with_columns(\n                pl.col(f\"top-k-dist-{p1}\")\n                .truediv(pl.col(f\"local-top-k-dist-{p1}\").add(eps))\n                .alias(f\"LOF-{p1}\"),\n            )\n            .with_columns(\n                pl.col(f\"top-k-dist-{p1}\")\n                .truediv(pl.col(f\"global-{normalize_type}-dist-{p1}\").add(eps))\n                .alias(f\"GOF-{p1}\"),\n            )\n        )\n\n    for p in EEG_PROBES:\n        eeg_df = eeg_df.with_columns(\n            pl.col(f\"GOF-{p}\").lt(distance_threshold).alias(f\"mask-{p}\"),\n            pl.lit(1.0)\n            .truediv(pl.col(f\"GOF-{p}\").clip(0).truediv(distance_threshold).add(1.0))\n            .alias(f\"CQF-{p}\"),\n        )\n\n    return eeg_df\n```",
      "votes": null
    },
    {
      "id": "2745017",
      "postDate": "04/10/2024 09:53:18",
      "content": "<p>Thanks for the great solution!<br>\nEspecially the approach to take into account the difference between left and right eeg is interesting!<br>\nLet me ask a question because I don't understand the L/R invariant mapping part.</p>\n<ol>\n<li>is x,y here embedding(B,64,10,8) of the output of cnn1d?</li>\n<li>I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.</li>\n<li>What will be the final output of this part?Is it binary whether or not each element is still equivalent after swap?</li>\n</ol>",
      "rawMarkdown": "Thanks for the great solution!\nEspecially the approach to take into account the difference between left and right eeg is interesting!\n\nLet me ask a question because I don't understand the L/R invariant mapping part.\n\n1. is x,y here embedding(B,64,10,8) of the output of cnn1d?\n2. I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.\n3. What will be the final output of this part?Is it binary whether or not each element is still equivalent after swap?",
      "votes": null
    },
    {
      "id": "2745067",
      "postDate": "04/10/2024 10:50:37",
      "content": "<p>Hi guys, we open-sourced our source code repository (bilzard part) with Apache 2.0 license.<br>\nSource code is available <a href=\"https://github.com/bilzard/kaggle-hms-bilzard\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Hi guys, we open-sourced our source code repository (bilzard part) with Apache 2.0 license.\nSource code is available [here](https://github.com/bilzard/kaggle-hms-bilzard).",
      "votes": null
    },
    {
      "id": "2745076",
      "postDate": "04/10/2024 11:01:15",
      "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> Thanks.</p>\n<blockquote>\n  <p>is x,y here embedding(B,64,10,8) of the output of cnn1d?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.</p>\n</blockquote>\n<p>Now you can refer to the source code here:<br>\n<a href=\"https://github.com/bilzard/kaggle-hms-bilzard/blob/main/src/model/basic_block/util.py#L17-L39\" target=\"_blank\">https://github.com/bilzard/kaggle-hms-bilzard/blob/main/src/model/basic_block/util.py#L17-L39</a></p>\n<p>For short answer, they takes min/max values when comparing x and y.</p>",
      "rawMarkdown": "kuto0633 Thanks.\n\n> is x,y here embedding(B,64,10,8) of the output of cnn1d?\n\nYes.\n\n> I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.\n\nNow you can refer to the source code here:\nhttps://github.com/bilzard/kaggle-hms-bilzard/blob/main/src/model/basic_block/util.py#L17-L39\n\nFor short answer, they takes min/max values when comparing x and y.",
      "votes": null
    },
    {
      "id": "2745142",
      "postDate": "04/10/2024 12:07:25",
      "content": "<p>I got it. Thanks for sharing!</p>",
      "rawMarkdown": "I got it. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2759872",
      "postDate": "04/19/2024 00:34:20",
      "content": "<p>Congratulations on your gold medal, and thank you for sharing your incredible knowledge!<br>\nI have a question. How’s the score changed by adding CQF information? I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.</p>",
      "rawMarkdown": "Congratulations on your gold medal, and thank you for sharing your incredible knowledge!\nI have a question. How’s the score changed by adding CQF information? I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.",
      "votes": null
    },
    {
      "id": "2759883",
      "postDate": "04/19/2024 01:09:54",
      "content": "<p><a href=\"https://www.kaggle.com/ludditep\" target=\"_blank\">@ludditep</a> Thanks.</p>\n<blockquote>\n  <p>How’s the score changed by adding CQF information?</p>\n</blockquote>\n<p>In the earlier experiments using 2D model shows adding CQF constantly improve CV/LB (-0.007~-0.008 gains). </p>\n<blockquote>\n  <p>I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.</p>\n</blockquote>\n<p>We consider the risk of confusing models by CQF is low. This is because CQF is calculated by absolute voltage (i.e. electrode of Fp1, F7, … etc.) instead of difference of voltages (Fp1-F7, … etc.). We believe the model still can successfully distinguish anomaly brain activity like seizure from corrupted channel.</p>\n<p>However, note that adding CQF to my teammate model does not improve performance, even some cases of deterioration. We conjecture this is because the model with enough parameter size and data, they might detect corrupted channel without CQF.</p>",
      "rawMarkdown": "ludditep Thanks.\n\n> How’s the score changed by adding CQF information?\n\nIn the earlier experiments using 2D model shows adding CQF constantly improve CV/LB (-0.007~-0.008 gains). \n\n> I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.\n\nWe consider the risk of confusing models by CQF is low. This is because CQF is calculated by absolute voltage (i.e. electrode of Fp1, F7, ... etc.) instead of difference of voltages (Fp1-F7, ... etc.). We believe the model still can successfully distinguish anomaly brain activity like seizure from corrupted channel.\n\nHowever, note that adding CQF to my teammate model does not improve performance, even some cases of deterioration. We conjecture this is because the model with enough parameter size and data, they might detect corrupted channel without CQF.",
      "votes": null
    },
    {
      "id": "2759891",
      "postDate": "04/19/2024 01:28:26",
      "content": "<p>The below picture illustrates how CQF successfully detect corrupted channels. From this picture, we can see the voltage of Fp1 gets apart from other normal channels in some time periods, and CQF gets low in these periods.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F4ba1f653256637507f55c3977930489c%2Fcqf.jpeg?generation=1713490264707152&amp;alt=media\"></p>",
      "rawMarkdown": "The below picture illustrates how CQF successfully detect corrupted channels. From this picture, we can see the voltage of Fp1 gets apart from other normal channels in some time periods, and CQF gets low in these periods.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F4ba1f653256637507f55c3977930489c%2Fcqf.jpeg?generation=1713490264707152&alt=media)",
      "votes": null
    },
    {
      "id": "2759892",
      "postDate": "04/19/2024 01:38:00",
      "content": "<p>Hi everyone.</p>\n<p>We shared revision of this article with more detailed explanation of CQF here:<br>\n<a href=\"https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\" target=\"_blank\">https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf</a></p>\n<p>Sorry for not sharing here, because some math equations are not correctly displayed on Kaggle discussion.</p>",
      "rawMarkdown": "Hi everyone.\n\nWe shared revision of this article with more detailed explanation of CQF here:\nhttps://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\n \nSorry for not sharing here, because some math equations are not correctly displayed on Kaggle discussion.",
      "votes": null
    },
    {
      "id": "2760348",
      "postDate": "04/19/2024 08:55:01",
      "content": "<p>I see, thank you again!</p>",
      "rawMarkdown": "I see, thank you again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2743554,
      "author_name": "janbrederecke",
      "author_url": "",
      "post_date": "04/09/2024 14:11:22",
      "content": "<p>This is incredible, I spent the whole comp on 1D data and this is so clever! Thanks for sharing your ideas in detail, this is more than inspirational!</p>\n<p>Best,<br>\nJan</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2743587,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "04/09/2024 14:28:15",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> great work and congratz on the prize zone finish, will you mind share the code, at least the preprocessing part?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2743601,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "04/09/2024 14:39:42",
          "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> Thanks.</p>\n<blockquote>\n  <p>will you mind share the code, at least the preprocessing part?</p>\n</blockquote>\n<p>You mean CQM part? sure. Below is our snippet for generating CQM.<br>\nInput is supposed to be down-sampled to 40Hz.</p>\n<pre><code> () -&gt; pl.DataFrame:\n    eeg_df = eeg_df.with_columns(\n        \n        pl.col(p1)\n        .sub(pl.col(p2))\n        .()\n        .ewm_mean(half_life=kernel_size, min_periods=)\n        .add(eps)\n        .alias()\n         p1, p2  product(EEG_PROBES, EEG_PROBES)\n    )\n\n    idxs = []\n     p1  EEG_PROBES:\n        rx = eeg_df.select(\n              p2  EEG_PROBES\n        ).to_numpy()\n        top_k_indices = np.argsort(rx, axis=)[:,  : top_k + ]\n        top_k_values = np.take_along_axis(rx, top_k_indices, axis=)\n        top_k_dist = top_k_values.mean(axis=)\n        eeg_df = eeg_df.with_columns(\n            pl.Series(top_k_dist).alias(),\n        )\n        idxs.append(top_k_indices)\n\n    xs = eeg_df.select(  p1  EEG_PROBES).to_numpy()  \n    global_top_k_indices = np.argsort(xs, axis=)[:, :top_k]\n    global_top_k_values = np.take_along_axis(xs, global_top_k_indices, axis=)\n    global_top_k_dist = global_top_k_values.mean(axis=)\n    global_median_dist = np.median(xs, axis=)\n    idxs = np.stack(idxs, axis=)  \n\n     p1  EEG_PROBES:\n        local_top_k_idxs = idxs[:, PROBE2IDX[p1], :]\n        local_top_k_values = np.take_along_axis(xs, local_top_k_idxs, axis=)\n        local_top_k_dist = local_top_k_values.mean(axis=)\n        eeg_df = (\n            eeg_df.with_columns(\n                pl.Series(local_top_k_dist).alias(),\n                pl.Series(global_top_k_dist).alias(),\n                pl.Series(global_median_dist).alias(),\n            )\n            .with_columns(\n                pl.col()\n                .truediv(pl.col().add(eps))\n                .alias(),\n            )\n            .with_columns(\n                pl.col()\n                .truediv(pl.col().add(eps))\n                .alias(),\n            )\n        )\n\n     p  EEG_PROBES:\n        eeg_df = eeg_df.with_columns(\n            pl.col().lt(distance_threshold).alias(),\n            pl.lit()\n            .truediv(pl.col().clip().truediv(distance_threshold).add())\n            .alias(),\n        )\n\n     eeg_df\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2745017,
      "author_name": "kuto0633",
      "author_url": "",
      "post_date": "04/10/2024 09:53:18",
      "content": "<p>Thanks for the great solution!<br>\nEspecially the approach to take into account the difference between left and right eeg is interesting!<br>\nLet me ask a question because I don't understand the L/R invariant mapping part.</p>\n<ol>\n<li>is x,y here embedding(B,64,10,8) of the output of cnn1d?</li>\n<li>I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.</li>\n<li>What will be the final output of this part?Is it binary whether or not each element is still equivalent after swap?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2745076,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "04/10/2024 11:01:15",
          "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> Thanks.</p>\n<blockquote>\n  <p>is x,y here embedding(B,64,10,8) of the output of cnn1d?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.</p>\n</blockquote>\n<p>Now you can refer to the source code here:<br>\n<a href=\"https://github.com/bilzard/kaggle-hms-bilzard/blob/main/src/model/basic_block/util.py#L17-L39\" target=\"_blank\">https://github.com/bilzard/kaggle-hms-bilzard/blob/main/src/model/basic_block/util.py#L17-L39</a></p>\n<p>For short answer, they takes min/max values when comparing x and y.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2745142,
              "author_name": "kuto0633",
              "author_url": "",
              "post_date": "04/10/2024 12:07:25",
              "content": "<p>I got it. Thanks for sharing!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2745067,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "04/10/2024 10:50:37",
      "content": "<p>Hi guys, we open-sourced our source code repository (bilzard part) with Apache 2.0 license.<br>\nSource code is available <a href=\"https://github.com/bilzard/kaggle-hms-bilzard\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2759872,
      "author_name": "ludditep",
      "author_url": "",
      "post_date": "04/19/2024 00:34:20",
      "content": "<p>Congratulations on your gold medal, and thank you for sharing your incredible knowledge!<br>\nI have a question. How’s the score changed by adding CQF information? I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2759883,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "04/19/2024 01:09:54",
          "content": "<p><a href=\"https://www.kaggle.com/ludditep\" target=\"_blank\">@ludditep</a> Thanks.</p>\n<blockquote>\n  <p>How’s the score changed by adding CQF information?</p>\n</blockquote>\n<p>In the earlier experiments using 2D model shows adding CQF constantly improve CV/LB (-0.007~-0.008 gains). </p>\n<blockquote>\n  <p>I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.</p>\n</blockquote>\n<p>We consider the risk of confusing models by CQF is low. This is because CQF is calculated by absolute voltage (i.e. electrode of Fp1, F7, … etc.) instead of difference of voltages (Fp1-F7, … etc.). We believe the model still can successfully distinguish anomaly brain activity like seizure from corrupted channel.</p>\n<p>However, note that adding CQF to my teammate model does not improve performance, even some cases of deterioration. We conjecture this is because the model with enough parameter size and data, they might detect corrupted channel without CQF.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2759891,
              "author_name": "tatamikenn",
              "author_url": "",
              "post_date": "04/19/2024 01:28:26",
              "content": "<p>The below picture illustrates how CQF successfully detect corrupted channels. From this picture, we can see the voltage of Fp1 gets apart from other normal channels in some time periods, and CQF gets low in these periods.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F4ba1f653256637507f55c3977930489c%2Fcqf.jpeg?generation=1713490264707152&amp;alt=media\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2760348,
                  "author_name": "ludditep",
                  "author_url": "",
                  "post_date": "04/19/2024 08:55:01",
                  "content": "<p>I see, thank you again!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2759892,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "04/19/2024 01:38:00",
      "content": "<p>Hi everyone.</p>\n<p>We shared revision of this article with more detailed explanation of CQF here:<br>\n<a href=\"https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\" target=\"_blank\">https://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf</a></p>\n<p>Sorry for not sharing here, because some math equations are not correctly displayed on Kaggle discussion.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2743516": "Big thanks to the hosts and the Kaggle team for setting up this competition. Diving into raw EEG signals was really fun. Also thanks to my teammates @yujiariyasu, @tattaka, and @ren4yu for being awesome partners.\n\nHere’s a detailed description of 1D model (bilzard's part) of our solution. The summary of our team solution is also available [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492240).\n\nEdit:\nWe open-sourced source code repository [here](https://github.com/bilzard/kaggle-hms-bilzard) (Apache 2.0).\n\nEdit: we shared revision of this article with more detailed explanation of CQF here:\nhttps://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\n\n### 1. Outline\n\nInitially, my solution was an ensemble of 1D and 2D models. However, after merging teams and recognizing that my teammates had already developed strong 2D models, I shifted my focus primarily towards 1D models (i.e. the models process raw EEG signals directly).\n\n### 2. Basic Concept\n\nThe approach to our 1D modeling is twofold:\n\n1. **L/R Symmetric Modeling**: We aimed to maintain symmetry in the model to facilitate the detection of Laterality.\n2. **Channel Quality Factor (CQF)**: This is used to evaluate the quality of EEG channels and identify any that are suboptimal.\n\nWe will discuss these on the following sections.\n\n### 3. L/R Symmetric modeling\n\nIn this section, we'll cover the basics of *L/R Symmetric Modeling*.\n\nOur inspiration came from textbooks[2], where we learned that the differences in signals from the left and right channels are key to telling General seizures apart from Lateral ones. That led us to design a model that treats left and right brain signals symmetrically, as illustrated in Fig. 1.\n\nHere’s how it works: The left and right brain signals are fed into a 1D CNN model separately, but they share the same parameters. The features we get from these are then processed in two ways: 1) through L/R invariant mapping, and 2) with a similarity encoder. We’ll dive into these parts more in the sections that follow.\n\nAn important detail is our use of a late fusion approach. This means we input each channel separately into the 1D encoder to get a more abstract representation of the EEG channels.\n\nWe also discovered that gradually blending the EEG channels partway through the 1D encoder improved our results. We call this the Channel Mixer. While we didn’t explore every possible architecture, adding a simple 1D convolution layer in each block was enough to see significant benefits.\n\n#### 3.1. L/R invariant mapping\n\nThis part of our model pulls out features that don’t change whether we swap left and right EEG channels. We do this using two mathematical functions, $f$ and $g$, defined as follows:\n\n$$\nu = f(x, y) \\\\\nv = g(x, y) \\\\\n\\text{ where } f(x, y) = f(y, x) \\\\\n\\text{ and } g(x, y) = g(y, x)\n$$\n\nWe considered several options for these functions, like:\n\n1. diff, mean\n1. prod, sum\n1. min, max\n\nIn our initial tests, we found that the choice of function didn’t drastically change the outcome. So, we settled on min/max as our default functions.\n\n#### 3.2. Similarity Encoder\n\nTo pinpoint the differences between the left and right sides, we measure the cosine similarity between paired L/R channels (for example, (Fp1, Fp2), (O1, O2), etc.). After calculating the similarity, we project it onto a $d$-dimensional vector to form a feature vector. This approach proved to be more effective than using the one-dimensional cosine similarity directly.\n\n<img src=\"https://i.postimg.cc/NGT1JcxK/hms-1d.png\" alt=\"hms-1d\" width=\"480\"/><br/>\n\n**Fig. 1 Architecture of 1D Model Proposed**\n\n### 4. Channel Quality Factor (CQF)\n\nDuring our initial EDA, we noticed a significant presence of bad channels in the EEG signals (see Fig. 2). This observation led us to devote effort to assessing channel quality, aiming to reduce potential confusion for our model.\n\nThe challenge of identifying corrupted channels is a well-known issue in EEG research. We discovered a study [1] that addresses this problem by introducing the *Locality Factor (LOF)*. LOF assesses the quality of a channel based on its discrepancy from other (usually normal) channels, with discrepancy measured by distance metrics like the Euclidean distance.\n\nWhile [1] considers channels in a static manner, we adapted LOF for use with time-series data, allowing us to capture nuanced aspects of channel quality over time. We refer to this adapted feature as the Channel Quality Factor (CQF), illustrated in Fig. 3.\n\n<img src=\"https://i.postimg.cc/tCjF8Mr2/bad-channel-sample01.png\" alt=\"bad-channel-sample01\" width=\"640\"/><br/><br/>\n\n**Fig. 2: Example of Bad Channels**\n\n<img src=\"https://i.postimg.cc/YqzQ7zZk/cqm002.jpg\" alt=\"cqm002\" width=\"640\"/><br/><br/>\n\n**Fig. 3: CQF Illustration**: Channels of lower quality are indicated in cooler colors.\n\n### 6. Result\n\nThe cross-validation scores (CVs) of our 1D models in the final submissions are listed below. These CVs were calculated using samples with `num_votes > 8`.\n\n```\nv5_eeg_24ep_cutmix 0.2477\n```\n\nWe should emphasize our 1D model is really light weight. It only has 863K parameters and took 1-2 minutes for infer with test set with 15 models (3 seed x 5 fold ensemble).\n\n### 7. Reference\n\n- [1] Velu et. al., Adaptable and Robust EEG Bad Channel Detection Using Local Outlier Factor (LOF)\n- [2] American Clinical Neurophysiology Society’s Standardized Critical Care EEG Terminology: 2021 Version\n- [3] https://github.com/huggingface/pytorch-image-models\n\n### Appendix\n\n#### A1. Training Configuration\n\nMost participants 2-stage training scheme, however, we schedule min_vote during 1-stage training schedule.\n\n1. 0-15 epoch: train with all data\n1. 16-23 epoch: train with n_votes>=8.4\n\nOther training configurations are:\n\n- Optimizer: Adam\n- lr: 1e-3\n- batch_size: 32\n- weight_decay: 1e-5\n- Scheduler: cosine\n\n#### A2. Augmentations\n\n- channel shuffling: shuffling channels while keeping L/R symmetry\n- L/R swap\n- dropout (<=128 frames x 4)\n- cutmix\n\n#### A3. Backbone 1D CNN Architecture\n\nWe designed 1D-CNN based on timm[3]'s EfficientNet2d's building blocks. We called this architecture as *EfficientNet1d*.\n\n```\n========================================================================================================================\nLayer (type:depth-idx)                                                 Output Shape              Param #\n========================================================================================================================\nEfficientNet1d                                                         [40, 64, 16]              --\n├─ConvBnAct2d: 1-1                                                     [4, 64, 10, 512]          --\n│    └─Sequential: 2-1                                                 [4, 64, 10, 512]          --\n│    │    └─Conv2d: 3-1                                                [4, 64, 10, 512]          384\n│    │    └─BatchNorm2d: 3-2                                           [4, 64, 10, 512]          128\n│    │    └─ELU: 3-3                                                   [4, 64, 10, 512]          --\n├─Sequential: 1-2                                                      --                        --\n│    └─ResBlock2d: 2-2                                                 [4, 64, 10, 256]          --\n│    │    └─MaxPool2d: 3-4                                             [4, 64, 10, 256]          --\n│    │    └─Sequential: 3-5                                            [4, 64, 10, 256]          --\n│    │    │    └─Sequential: 4-1                                       [4, 64, 10, 512]          --\n│    │    │    │    └─InvertedResidual: 5-1                            [4, 64, 10, 512]          55,104\n│    │    │    │    └─InvertedResidual: 5-2                            [4, 64, 10, 512]          55,104\n│    │    │    │    └─InvertedResidual: 5-3                            [4, 64, 10, 512]          55,104\n│    │    │    │    └─DepthWiseSeparableConv: 5-4                      [4, 64, 10, 512]          6,672\n│    │    │    └─MaxPool2d: 4-2                                        [4, 64, 10, 256]          --\n│    └─ResBlock2d: 2-3                                                 [4, 64, 10, 128]          --\n│    │    └─MaxPool2d: 3-6                                             [4, 64, 10, 128]          --\n│    │    └─Sequential: 3-7                                            [4, 64, 10, 128]          --\n│    │    │    └─Sequential: 4-3                                       [4, 64, 10, 256]          --\n│    │    │    │    └─InvertedResidual: 5-5                            [4, 64, 10, 256]          55,104\n│    │    │    │    └─InvertedResidual: 5-6                            [4, 64, 10, 256]          55,104\n│    │    │    │    └─InvertedResidual: 5-7                            [4, 64, 10, 256]          55,104\n│    │    │    │    └─DepthWiseSeparableConv: 5-8                      [4, 64, 10, 256]          6,672\n│    │    │    └─MaxPool2d: 4-4                                        [4, 64, 10, 128]          --\n│    └─ResBlock2d: 2-4                                                 [4, 64, 10, 64]           --\n│    │    └─MaxPool2d: 3-8                                             [4, 64, 10, 64]           --\n│    │    └─Sequential: 3-9                                            [4, 64, 10, 64]           --\n│    │    │    └─Sequential: 4-5                                       [4, 64, 10, 128]          --\n│    │    │    │    └─InvertedResidual: 5-9                            [4, 64, 10, 128]          55,616\n│    │    │    │    └─InvertedResidual: 5-10                           [4, 64, 10, 128]          55,616\n│    │    │    │    └─InvertedResidual: 5-11                           [4, 64, 10, 128]          55,616\n│    │    │    │    └─DepthWiseSeparableConv: 5-12                     [4, 64, 10, 128]          6,672\n│    │    │    └─MaxPool2d: 4-6                                        [4, 64, 10, 64]           --\n│    └─ResBlock2d: 2-5                                                 [4, 64, 10, 32]           --\n│    │    └─MaxPool2d: 3-10                                            [4, 64, 10, 32]           --\n│    │    └─Sequential: 3-11                                           [4, 64, 10, 32]           --\n│    │    │    └─Sequential: 4-7                                       [4, 64, 10, 64]           --\n│    │    │    │    └─InvertedResidual: 5-13                           [4, 64, 10, 64]           55,616\n│    │    │    │    └─InvertedResidual: 5-14                           [4, 64, 10, 64]           55,616\n│    │    │    │    └─InvertedResidual: 5-15                           [4, 64, 10, 64]           55,616\n│    │    │    │    └─DepthWiseSeparableConv: 5-16                     [4, 64, 10, 64]           6,672\n│    │    │    └─MaxPool2d: 4-8                                        [4, 64, 10, 32]           --\n│    └─ResBlock2d: 2-6                                                 [4, 64, 10, 16]           --\n│    │    └─MaxPool2d: 3-12                                            [4, 64, 10, 16]           --\n│    │    └─Sequential: 3-13                                           [4, 64, 10, 16]           --\n│    │    │    └─Sequential: 4-9                                       [4, 64, 10, 32]           --\n│    │    │    │    └─InvertedResidual: 5-17                           [4, 64, 10, 32]           55,104\n│    │    │    │    └─InvertedResidual: 5-18                           [4, 64, 10, 32]           55,104\n│    │    │    │    └─InvertedResidual: 5-19                           [4, 64, 10, 32]           55,104\n│    │    │    │    └─DepthWiseSeparableConv: 5-20                     [4, 64, 10, 32]           6,672\n│    │    │    └─MaxPool2d: 4-10                                       [4, 64, 10, 16]           --\n========================================================================================================================\nTotal params: 863,504\nTrainable params: 863,504\nNon-trainable params: 0\nTotal mult-adds (G): 2.72\n========================================================================================================================\nInput size (MB): 0.16\nForward/backward pass size (MB): 833.78\nParams size (MB): 3.45\nEstimated Total Size (MB): 837.40\n========================================================================================================================\n```",
    "2743554": "This is incredible, I spent the whole comp on 1D data and this is so clever! Thanks for sharing your ideas in detail, this is more than inspirational!\n\nBest,\nJan",
    "2743587": "tatamikenn great work and congratz on the prize zone finish, will you mind share the code, at least the preprocessing part?",
    "2743601": "sergiosaharovskiy Thanks.\n\n> will you mind share the code, at least the preprocessing part?\n\nYou mean CQM part? sure. Below is our snippet for generating CQM.\nInput is supposed to be down-sampled to 40Hz.\n\n```python\ndef do_process_cqf(\n    eeg_df: pl.DataFrame,\n    kernel_size: int = 13,\n    top_k: int = 3,\n    eps: float = 1e-4,\n    distance_threshold: float = 10.0,\n    distance_metric: str = \"l2\",\n    normalize_type: str = \"top-k\",\n) -> pl.DataFrame:\n    eeg_df = eeg_df.with_columns(\n        # L2 distance of (p1, p2)\n        pl.col(p1)\n        .sub(pl.col(p2))\n        .pow(2)\n        .ewm_mean(half_life=kernel_size, min_periods=1)\n        .add(eps)\n        .alias(f\"l2-dist-{p1}-{p2}\")\n        for p1, p2 in product(EEG_PROBES, EEG_PROBES)\n    )\n\n    idxs = []\n    for p1 in EEG_PROBES:\n        rx = eeg_df.select(\n            f\"{distance_metric}-dist-{p1}-{p2}\" for p2 in EEG_PROBES\n        ).to_numpy()\n        top_k_indices = np.argsort(rx, axis=1)[:, 1 : top_k + 1]\n        top_k_values = np.take_along_axis(rx, top_k_indices, axis=1)\n        top_k_dist = top_k_values.mean(axis=1)\n        eeg_df = eeg_df.with_columns(\n            pl.Series(top_k_dist).alias(f\"top-k-dist-{p1}\"),\n        )\n        idxs.append(top_k_indices)\n\n    xs = eeg_df.select(f\"top-k-dist-{p1}\" for p1 in EEG_PROBES).to_numpy()  # (N, 19)\n    global_top_k_indices = np.argsort(xs, axis=1)[:, :top_k]\n    global_top_k_values = np.take_along_axis(xs, global_top_k_indices, axis=1)\n    global_top_k_dist = global_top_k_values.mean(axis=1)\n    global_median_dist = np.median(xs, axis=1)\n    idxs = np.stack(idxs, axis=1)  # (N, 19, top_k)\n\n    for p1 in EEG_PROBES:\n        local_top_k_idxs = idxs[:, PROBE2IDX[p1], :]\n        local_top_k_values = np.take_along_axis(xs, local_top_k_idxs, axis=1)\n        local_top_k_dist = local_top_k_values.mean(axis=1)\n        eeg_df = (\n            eeg_df.with_columns(\n                pl.Series(local_top_k_dist).alias(f\"local-top-k-dist-{p1}\"),\n                pl.Series(global_top_k_dist).alias(f\"global-top-k-dist-{p1}\"),\n                pl.Series(global_median_dist).alias(f\"global-median-dist-{p1}\"),\n            )\n            .with_columns(\n                pl.col(f\"top-k-dist-{p1}\")\n                .truediv(pl.col(f\"local-top-k-dist-{p1}\").add(eps))\n                .alias(f\"LOF-{p1}\"),\n            )\n            .with_columns(\n                pl.col(f\"top-k-dist-{p1}\")\n                .truediv(pl.col(f\"global-{normalize_type}-dist-{p1}\").add(eps))\n                .alias(f\"GOF-{p1}\"),\n            )\n        )\n\n    for p in EEG_PROBES:\n        eeg_df = eeg_df.with_columns(\n            pl.col(f\"GOF-{p}\").lt(distance_threshold).alias(f\"mask-{p}\"),\n            pl.lit(1.0)\n            .truediv(pl.col(f\"GOF-{p}\").clip(0).truediv(distance_threshold).add(1.0))\n            .alias(f\"CQF-{p}\"),\n        )\n\n    return eeg_df\n```",
    "2745017": "Thanks for the great solution!\nEspecially the approach to take into account the difference between left and right eeg is interesting!\n\nLet me ask a question because I don't understand the L/R invariant mapping part.\n\n1. is x,y here embedding(B,64,10,8) of the output of cnn1d?\n2. I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.\n3. What will be the final output of this part?Is it binary whether or not each element is still equivalent after swap?",
    "2745067": "Hi guys, we open-sourced our source code repository (bilzard part) with Apache 2.0 license.\nSource code is available [here](https://github.com/bilzard/kaggle-hms-bilzard).",
    "2745076": "kuto0633 Thanks.\n\n> is x,y here embedding(B,64,10,8) of the output of cnn1d?\n\nYes.\n\n> I am not sure about the f(x,y) process. you said you chose f=min,g=max. does this mean that you take the minimum value of each element of the x,y tensor?I wasn't sure about the equivalence condition since in that case, even if x,y are swapped, they will have the same value.\n\nNow you can refer to the source code here:\nhttps://github.com/bilzard/kaggle-hms-bilzard/blob/main/src/model/basic_block/util.py#L17-L39\n\nFor short answer, they takes min/max values when comparing x and y.",
    "2745142": "I got it. Thanks for sharing!",
    "2759872": "Congratulations on your gold medal, and thank you for sharing your incredible knowledge!\nI have a question. How’s the score changed by adding CQF information? I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.",
    "2759883": "ludditep Thanks.\n\n> How’s the score changed by adding CQF information?\n\nIn the earlier experiments using 2D model shows adding CQF constantly improve CV/LB (-0.007~-0.008 gains). \n\n> I thought that bad brain activities like seizures can also create discrepancies from other (usually normal) channels, and discriminating them as bad quality may confuse the model.\n\nWe consider the risk of confusing models by CQF is low. This is because CQF is calculated by absolute voltage (i.e. electrode of Fp1, F7, ... etc.) instead of difference of voltages (Fp1-F7, ... etc.). We believe the model still can successfully distinguish anomaly brain activity like seizure from corrupted channel.\n\nHowever, note that adding CQF to my teammate model does not improve performance, even some cases of deterioration. We conjecture this is because the model with enough parameter size and data, they might detect corrupted channel without CQF.",
    "2759891": "The below picture illustrates how CQF successfully detect corrupted channels. From this picture, we can see the voltage of Fp1 gets apart from other normal channels in some time periods, and CQF gets low in these periods.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F4ba1f653256637507f55c3977930489c%2Fcqf.jpeg?generation=1713490264707152&alt=media)",
    "2759892": "Hi everyone.\n\nWe shared revision of this article with more detailed explanation of CQF here:\nhttps://github.com/bilzard/kaggle-hms-public/blob/main/doc/02_solution_bilzard.pdf\n \nSorry for not sharing here, because some math equations are not correctly displayed on Kaggle discussion.",
    "2760348": "I see, thank you again!"
  },
  "source": "meta"
}