{
  "id": 492189,
  "title": "60th Place Solution: Single Multimodal Attention Model (0.26 CV | 0.28 Public LB | 0.33 Private LB)",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492189",
  "author_name": "ryan",
  "post_date": "2024-04-09T00:14:54.939000",
  "votes": 24,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Huge thanks to the Hosts and Kaggle, this competition was by far one of the most fun competitions I've competed in awhile. </p>\n<p><br></p>\n<p><strong>TL;DR:</strong></p>\n<ul>\n<li>My final model was a single multimodal attention-style model with CNN and CNN-RNN encoders. I treat each <em>node</em> (ex: Fp1-F7, or LP for spectrograms) as an independent item that needs to be embedded (27 items total), then use additive attention with learned positional encoding to attend to each of the samples, then pool and decode. I also used sample-dependent adaptive temperature scaling (for interpretability, neither increased nor decreased CV/LB score). </li>\n<li>Designed a custom kNN knowledge distillation scheme to pseudo-label samples with &lt;10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas. </li>\n<li>No pretrained models, all CNN's were trained from scratch.</li>\n<li>Robust normalization the 1D EEGs by the mean absolute deviation (MAD) instead of standard deviation.</li>\n</ul>\n<p><strong>Scores:</strong></p>\n<ul>\n<li>Single 1D RAW EEG: <ul>\n<li>0.28 CV | 0.28 Public LB | 0.34 Private LB</li>\n<li>~2M params</li></ul></li>\n<li>Single 2D EEG spectrograms: <ul>\n<li>0.32 CV | 0.33 Public LB | 0.4 Private LB</li>\n<li>~2M params</li></ul></li>\n<li>Single 2D Kaggle Spectrograms: <ul>\n<li>0.48 CV | 0.51 Public LB | 0.59 Private LB</li>\n<li>~2M params</li></ul></li>\n<li>Final Multimodal Model: <ul>\n<li>0.26 CV | 0.28 Public LB | 0.33 Private LB</li>\n<li>10M params</li></ul></li>\n</ul>\n<p>I selected two models, one for best CV on the whole dataset and one for best CV on samples &gt;10 in case there was a shakeup. No shakeup, trusted my personal CV and it payed off :). Given that my original goal was ~0.29 when I join this competition a month ago (a bit late, really had to grind), I am very happy with my performance.</p>\n<p>I feel if I had a bit more time, ~0.25 would be achievable (local CV / public LB) with my current solution, but that something was missing between what I have now and a sub 0.25 CV/LB score (probably some form of data psuedo-labeling, preprocessing, and/or selection). Can't wait to read everyone elses solutions and see how sub 0.25 scores were achieved. </p>\n<h3>Data Preprocessing</h3>\n<p>For the EEG, I used the banana montage (including the middle chain and the EKG). Nothing here is too different from what I saw other people doing, other than the fact that I normalize all 1D EEG signals by the MAD for robustness.</p>\n<p>Proprocessing (EEG): </p>\n<ul>\n<li>Butter bandpass filter to only keep frequencies between 0.5 and 50 Hz. </li>\n<li>4x mean downsampling to 50 Hz.</li>\n<li>Mean subtraction, and division by the <em>median</em> mean absolute deviation (MAD) across <em>all</em> channels for robust estimation of the standard deviation (see next section). That is, global normalization and not per signal normalization. I choose to do this because I think it's important to keep all of the relative scales the same. </li>\n<li>Clip between -10 and 10.</li>\n</ul>\n<p>Proprocessing (EEG Spectrogram): </p>\n<ul>\n<li>Butter bandpass filter of the raw signal to keep frequencies between 0.5 and 40 Hz.</li>\n<li>Normalized the signal by mean subtraction and MAD normalization first.</li>\n<li>Used the torchaudio Spectrogram function with hop_length 44, win_length 256, and n_fft 800. </li>\n<li>Clippd between $e^{-4}$ and $e^{7}$.</li>\n<li>Took the log.</li>\n</ul>\n<p>Proprocessing (Spectrogram): </p>\n<ul>\n<li>Clippd between $e^{-4}$ and $e^{7}$.</li>\n<li>Took the log.</li>\n<li>Mean/std normalized each spectrogram (NOT global).</li>\n</ul>\n<h3>Robust Normalization</h3>\n<p>In this competition especially, I found it very important to make sure that your normalize your data in a way that was robust to artifacts. And after playing around with quite a lot of different normalizations and visualizing the raw data, it was clear to me that the traditional standard deviation was NOT a robust measure and that it was extremely susceptible to artifacts in the case of the 1D EEGs. To this end, I instead chose to divide by the mean absolute deviation (MAD) and clip values between -10 and 10. This spead up convergence quite a bit, reduce overfitting, and also seemed to generalize better to my validation set. </p>\n<h3>Augmentations and TTA</h3>\n<p>1D EEGs: </p>\n<ul>\n<li>Horizontally and vertically flip (each with p = 0.5) the ordering of the nodes in the banana montage. Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.</li>\n<li>Randomly zero out 1-4 channels at a time with probability 0.5. </li>\n</ul>\n<p>2D Spectroggrams:</p>\n<ul>\n<li>During training, I did minimal augmentations to the spectrograms, other than swapping the direction of the chains (both ll &lt;-&gt; rl and lp &lt;-&gt; rp). This is the analog of the horizontal flip mentioned in the 1D EEGs.</li>\n</ul>\n<p>I tried working with test-time-augmentations (TTA) like flipping the orientation of chains, but it neither increased or decreased my CV/LB score in my later models. I think this is because I treat each node as an independent sample and used the same TTAs as augmentations during training. Which means by model tends to be robust to these shifts.</p>\n<h3>Model Architecture</h3>\n<p>When I started this competition, I had a couple of goals for my model architecture that I wanted to achieve.</p>\n<ol>\n<li>Scalable </li>\n<li>The general architecture made no hard assumptions about the spatial layout OR modality of the data, but could instead learn both the spatial layout and to incorperate new modalities on it's own. </li>\n<li>Semi-interpretable (not sure my solution achieved this, but I have ideas on how to achieve this going forward, see the <em>Extra</em> section at the bottom).</li>\n</ol>\n<p>To achieve these goals, my model architecture for each modality followed the same set of encoding/decoding principles:</p>\n<ol>\n<li>Split the data up into nodes (ex: F7-Fp1 or LP).</li>\n<li>Embed each node.</li>\n<li>Use an attention style decoder with learned positional encodings to model the relationship between each node. </li>\n<li>Pool and then linear decoding.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9041364%2F17977936ff43f954b6ee62fca3e706d4%2FUntitled%20drawing%20(8).png?generation=1712622281054757&amp;alt=media\"></p>\n<h3>Training Scheme</h3>\n<p>My training stages were an absolute disaster. I had 5-6 stages where I would start with easy samples to get the model to converge quick, then moved on to hard samples, and finalized with training on just the high confidence (&gt;10 votes) samples. </p>\n<h3>Extra: Future Ideas, Interpretability, and Kaggle</h3>\n<p>Since the beginning of this competition, I've put a bit of thought into the future directions of this project after the competition and have some ideas for the competition hosts regarding how I could imagine technology like this being used in practice. </p>\n<p>One of a couple ideas that I've had would be to design an interpretable model with a kNN decoder. For example, imagine you deploy a model like this in the real world as an aid to doctors who are labeling or interpreting this data. Then, upon making a prediction, the model not only gives you a probabilistic output of what it thinks is true, but also looks up the nearest N neighbors neighbors in the embedding space (of the training set) and shows you those N samples as to say \"here is my prediction, but here are also 10 (for example) other EEGs and spectrograms that have similar patterns to this one and hence why I labeled it this way.\"</p>\n<p>I have a lot of ideas like this that I'm going to be exploring over this summer as a side project (including reading/implementing other <em>unique</em> top submissions). I'll also be open sourcing it, so feel free to monitor my GitHub for updates (as well as the code for this project, which should be up in 2-3 days).</p>\n<p>Congrats to the winners!</p>\n<p>Thanks Kaggle, and Hosts! This competition was by far one of the most fun I have joined. I learned a ton.</p>\n<p>GitHub Training Code: <a href=\"https://github.com/ryanirl/hbac\" target=\"_blank\">https://github.com/ryanirl/hbac</a><br>\nKaggle Final Inference: <a href=\"https://www.kaggle.com/code/ryanirl/hms-baseline-sub\" target=\"_blank\">https://www.kaggle.com/code/ryanirl/hms-baseline-sub</a></p>\n<p>Ryan Peters<br>\n<a href=\"mailto:RyanIRL@icloud.com\">RyanIRL@icloud.com</a></p>",
  "messages": [
    {
      "id": 2742494,
      "postDate": "2024-04-09T00:14:54.940Z",
      "content": "<p>Huge thanks to the Hosts and Kaggle, this competition was by far one of the most fun competitions I've competed in awhile. </p>\n<p><br></p>\n<p><strong>TL;DR:</strong></p>\n<ul>\n<li>My final model was a single multimodal attention-style model with CNN and CNN-RNN encoders. I treat each <em>node</em> (ex: Fp1-F7, or LP for spectrograms) as an independent item that needs to be embedded (27 items total), then use additive attention with learned positional encoding to attend to each of the samples, then pool and decode. I also used sample-dependent adaptive temperature scaling (for interpretability, neither increased nor decreased CV/LB score). </li>\n<li>Designed a custom kNN knowledge distillation scheme to pseudo-label samples with &lt;10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas. </li>\n<li>No pretrained models, all CNN's were trained from scratch.</li>\n<li>Robust normalization the 1D EEGs by the mean absolute deviation (MAD) instead of standard deviation.</li>\n</ul>\n<p><strong>Scores:</strong></p>\n<ul>\n<li>Single 1D RAW EEG: <ul>\n<li>0.28 CV | 0.28 Public LB | 0.34 Private LB</li>\n<li>~2M params</li></ul></li>\n<li>Single 2D EEG spectrograms: <ul>\n<li>0.32 CV | 0.33 Public LB | 0.4 Private LB</li>\n<li>~2M params</li></ul></li>\n<li>Single 2D Kaggle Spectrograms: <ul>\n<li>0.48 CV | 0.51 Public LB | 0.59 Private LB</li>\n<li>~2M params</li></ul></li>\n<li>Final Multimodal Model: <ul>\n<li>0.26 CV | 0.28 Public LB | 0.33 Private LB</li>\n<li>10M params</li></ul></li>\n</ul>\n<p>I selected two models, one for best CV on the whole dataset and one for best CV on samples &gt;10 in case there was a shakeup. No shakeup, trusted my personal CV and it payed off :). Given that my original goal was ~0.29 when I join this competition a month ago (a bit late, really had to grind), I am very happy with my performance.</p>\n<p>I feel if I had a bit more time, ~0.25 would be achievable (local CV / public LB) with my current solution, but that something was missing between what I have now and a sub 0.25 CV/LB score (probably some form of data psuedo-labeling, preprocessing, and/or selection). Can't wait to read everyone elses solutions and see how sub 0.25 scores were achieved. </p>\n<h3>Data Preprocessing</h3>\n<p>For the EEG, I used the banana montage (including the middle chain and the EKG). Nothing here is too different from what I saw other people doing, other than the fact that I normalize all 1D EEG signals by the MAD for robustness.</p>\n<p>Proprocessing (EEG): </p>\n<ul>\n<li>Butter bandpass filter to only keep frequencies between 0.5 and 50 Hz. </li>\n<li>4x mean downsampling to 50 Hz.</li>\n<li>Mean subtraction, and division by the <em>median</em> mean absolute deviation (MAD) across <em>all</em> channels for robust estimation of the standard deviation (see next section). That is, global normalization and not per signal normalization. I choose to do this because I think it's important to keep all of the relative scales the same. </li>\n<li>Clip between -10 and 10.</li>\n</ul>\n<p>Proprocessing (EEG Spectrogram): </p>\n<ul>\n<li>Butter bandpass filter of the raw signal to keep frequencies between 0.5 and 40 Hz.</li>\n<li>Normalized the signal by mean subtraction and MAD normalization first.</li>\n<li>Used the torchaudio Spectrogram function with hop_length 44, win_length 256, and n_fft 800. </li>\n<li>Clippd between $e^{-4}$ and $e^{7}$.</li>\n<li>Took the log.</li>\n</ul>\n<p>Proprocessing (Spectrogram): </p>\n<ul>\n<li>Clippd between $e^{-4}$ and $e^{7}$.</li>\n<li>Took the log.</li>\n<li>Mean/std normalized each spectrogram (NOT global).</li>\n</ul>\n<h3>Robust Normalization</h3>\n<p>In this competition especially, I found it very important to make sure that your normalize your data in a way that was robust to artifacts. And after playing around with quite a lot of different normalizations and visualizing the raw data, it was clear to me that the traditional standard deviation was NOT a robust measure and that it was extremely susceptible to artifacts in the case of the 1D EEGs. To this end, I instead chose to divide by the mean absolute deviation (MAD) and clip values between -10 and 10. This spead up convergence quite a bit, reduce overfitting, and also seemed to generalize better to my validation set. </p>\n<h3>Augmentations and TTA</h3>\n<p>1D EEGs: </p>\n<ul>\n<li>Horizontally and vertically flip (each with p = 0.5) the ordering of the nodes in the banana montage. Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.</li>\n<li>Randomly zero out 1-4 channels at a time with probability 0.5. </li>\n</ul>\n<p>2D Spectroggrams:</p>\n<ul>\n<li>During training, I did minimal augmentations to the spectrograms, other than swapping the direction of the chains (both ll &lt;-&gt; rl and lp &lt;-&gt; rp). This is the analog of the horizontal flip mentioned in the 1D EEGs.</li>\n</ul>\n<p>I tried working with test-time-augmentations (TTA) like flipping the orientation of chains, but it neither increased or decreased my CV/LB score in my later models. I think this is because I treat each node as an independent sample and used the same TTAs as augmentations during training. Which means by model tends to be robust to these shifts.</p>\n<h3>Model Architecture</h3>\n<p>When I started this competition, I had a couple of goals for my model architecture that I wanted to achieve.</p>\n<ol>\n<li>Scalable </li>\n<li>The general architecture made no hard assumptions about the spatial layout OR modality of the data, but could instead learn both the spatial layout and to incorperate new modalities on it's own. </li>\n<li>Semi-interpretable (not sure my solution achieved this, but I have ideas on how to achieve this going forward, see the <em>Extra</em> section at the bottom).</li>\n</ol>\n<p>To achieve these goals, my model architecture for each modality followed the same set of encoding/decoding principles:</p>\n<ol>\n<li>Split the data up into nodes (ex: F7-Fp1 or LP).</li>\n<li>Embed each node.</li>\n<li>Use an attention style decoder with learned positional encodings to model the relationship between each node. </li>\n<li>Pool and then linear decoding.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9041364%2F17977936ff43f954b6ee62fca3e706d4%2FUntitled%20drawing%20(8).png?generation=1712622281054757&amp;alt=media\"></p>\n<h3>Training Scheme</h3>\n<p>My training stages were an absolute disaster. I had 5-6 stages where I would start with easy samples to get the model to converge quick, then moved on to hard samples, and finalized with training on just the high confidence (&gt;10 votes) samples. </p>\n<h3>Extra: Future Ideas, Interpretability, and Kaggle</h3>\n<p>Since the beginning of this competition, I've put a bit of thought into the future directions of this project after the competition and have some ideas for the competition hosts regarding how I could imagine technology like this being used in practice. </p>\n<p>One of a couple ideas that I've had would be to design an interpretable model with a kNN decoder. For example, imagine you deploy a model like this in the real world as an aid to doctors who are labeling or interpreting this data. Then, upon making a prediction, the model not only gives you a probabilistic output of what it thinks is true, but also looks up the nearest N neighbors neighbors in the embedding space (of the training set) and shows you those N samples as to say \"here is my prediction, but here are also 10 (for example) other EEGs and spectrograms that have similar patterns to this one and hence why I labeled it this way.\"</p>\n<p>I have a lot of ideas like this that I'm going to be exploring over this summer as a side project (including reading/implementing other <em>unique</em> top submissions). I'll also be open sourcing it, so feel free to monitor my GitHub for updates (as well as the code for this project, which should be up in 2-3 days).</p>\n<p>Congrats to the winners!</p>\n<p>Thanks Kaggle, and Hosts! This competition was by far one of the most fun I have joined. I learned a ton.</p>\n<p>GitHub Training Code: <a href=\"https://github.com/ryanirl/hbac\" target=\"_blank\">https://github.com/ryanirl/hbac</a><br>\nKaggle Final Inference: <a href=\"https://www.kaggle.com/code/ryanirl/hms-baseline-sub\" target=\"_blank\">https://www.kaggle.com/code/ryanirl/hms-baseline-sub</a></p>\n<p>Ryan Peters<br>\n<a href=\"mailto:RyanIRL@icloud.com\">RyanIRL@icloud.com</a></p>",
      "rawMarkdown": "Huge thanks to the Hosts and Kaggle, this competition was by far one of the most fun competitions I've competed in awhile. \n\n<br/>\n\n**TL;DR:**\n - My final model was a single multimodal attention-style model with CNN and CNN-RNN encoders. I treat each *node* (ex: Fp1-F7, or LP for spectrograms) as an independent item that needs to be embedded (27 items total), then use additive attention with learned positional encoding to attend to each of the samples, then pool and decode. I also used sample-dependent adaptive temperature scaling (for interpretability, neither increased nor decreased CV/LB score). \n - Designed a custom kNN knowledge distillation scheme to pseudo-label samples with <10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas. \n - No pretrained models, all CNN's were trained from scratch.\n - Robust normalization the 1D EEGs by the mean absolute deviation (MAD) instead of standard deviation.\n\n**Scores:**\n\n - Single 1D RAW EEG: \n   - 0.28 CV | 0.28 Public LB | 0.34 Private LB\n   - ~2M params\n - Single 2D EEG spectrograms: \n   - 0.32 CV | 0.33 Public LB | 0.4 Private LB\n   - ~2M params\n - Single 2D Kaggle Spectrograms: \n   - 0.48 CV | 0.51 Public LB | 0.59 Private LB\n   - ~2M params\n - Final Multimodal Model: \n   - 0.26 CV | 0.28 Public LB | 0.33 Private LB\n   - 10M params\n\n\nI selected two models, one for best CV on the whole dataset and one for best CV on samples >10 in case there was a shakeup. No shakeup, trusted my personal CV and it payed off :). Given that my original goal was ~0.29 when I join this competition a month ago (a bit late, really had to grind), I am very happy with my performance.\n\nI feel if I had a bit more time, ~0.25 would be achievable (local CV / public LB) with my current solution, but that something was missing between what I have now and a sub 0.25 CV/LB score (probably some form of data psuedo-labeling, preprocessing, and/or selection). Can't wait to read everyone elses solutions and see how sub 0.25 scores were achieved. \n\n\n### Data Preprocessing \n\nFor the EEG, I used the banana montage (including the middle chain and the EKG). Nothing here is too different from what I saw other people doing, other than the fact that I normalize all 1D EEG signals by the MAD for robustness.\n\nProprocessing (EEG): \n  - Butter bandpass filter to only keep frequencies between 0.5 and 50 Hz. \n  - 4x mean downsampling to 50 Hz.\n  - Mean subtraction, and division by the *median* mean absolute deviation (MAD) across *all* channels for robust estimation of the standard deviation (see next section). That is, global normalization and not per signal normalization. I choose to do this because I think it's important to keep all of the relative scales the same. \n  - Clip between -10 and 10.\n\nProprocessing (EEG Spectrogram): \n  - Butter bandpass filter of the raw signal to keep frequencies between 0.5 and 40 Hz.\n  - Normalized the signal by mean subtraction and MAD normalization first.\n  - Used the torchaudio Spectrogram function with hop_length 44, win_length 256, and n_fft 800. \n  - Clippd between $e^{-4}$ and $e^{7}$.\n  - Took the log.\n\nProprocessing (Spectrogram): \n  - Clippd between $e^{-4}$ and $e^{7}$.\n  - Took the log.\n  - Mean/std normalized each spectrogram (NOT global).\n\n\n### Robust Normalization\n\nIn this competition especially, I found it very important to make sure that your normalize your data in a way that was robust to artifacts. And after playing around with quite a lot of different normalizations and visualizing the raw data, it was clear to me that the traditional standard deviation was NOT a robust measure and that it was extremely susceptible to artifacts in the case of the 1D EEGs. To this end, I instead chose to divide by the mean absolute deviation (MAD) and clip values between -10 and 10. This spead up convergence quite a bit, reduce overfitting, and also seemed to generalize better to my validation set. \n\n\n### Augmentations and TTA\n\n1D EEGs: \n  - Horizontally and vertically flip (each with p = 0.5) the ordering of the nodes in the banana montage. Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.\n  - Randomly zero out 1-4 channels at a time with probability 0.5. \n\n2D Spectroggrams:\n  - During training, I did minimal augmentations to the spectrograms, other than swapping the direction of the chains (both ll <-> rl and lp <-> rp). This is the analog of the horizontal flip mentioned in the 1D EEGs.\n\nI tried working with test-time-augmentations (TTA) like flipping the orientation of chains, but it neither increased or decreased my CV/LB score in my later models. I think this is because I treat each node as an independent sample and used the same TTAs as augmentations during training. Which means by model tends to be robust to these shifts.\n\n### Model Architecture\n\nWhen I started this competition, I had a couple of goals for my model architecture that I wanted to achieve.\n 1. Scalable \n 2. The general architecture made no hard assumptions about the spatial layout OR modality of the data, but could instead learn both the spatial layout and to incorperate new modalities on it's own. \n 3. Semi-interpretable (not sure my solution achieved this, but I have ideas on how to achieve this going forward, see the *Extra* section at the bottom).\n\nTo achieve these goals, my model architecture for each modality followed the same set of encoding/decoding principles:\n 1. Split the data up into nodes (ex: F7-Fp1 or LP).\n 2. Embed each node.\n 3. Use an attention style decoder with learned positional encodings to model the relationship between each node. \n 4. Pool and then linear decoding.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9041364%2F17977936ff43f954b6ee62fca3e706d4%2FUntitled%20drawing%20(8).png?generation=1712622281054757&alt=media)\n\n### Training Scheme\n\nMy training stages were an absolute disaster. I had 5-6 stages where I would start with easy samples to get the model to converge quick, then moved on to hard samples, and finalized with training on just the high confidence (>10 votes) samples. \n\n### Extra: Future Ideas, Interpretability, and Kaggle\n\nSince the beginning of this competition, I've put a bit of thought into the future directions of this project after the competition and have some ideas for the competition hosts regarding how I could imagine technology like this being used in practice. \n\nOne of a couple ideas that I've had would be to design an interpretable model with a kNN decoder. For example, imagine you deploy a model like this in the real world as an aid to doctors who are labeling or interpreting this data. Then, upon making a prediction, the model not only gives you a probabilistic output of what it thinks is true, but also looks up the nearest N neighbors neighbors in the embedding space (of the training set) and shows you those N samples as to say \"here is my prediction, but here are also 10 (for example) other EEGs and spectrograms that have similar patterns to this one and hence why I labeled it this way.\"\n\nI have a lot of ideas like this that I'm going to be exploring over this summer as a side project (including reading/implementing other *unique* top submissions). I'll also be open sourcing it, so feel free to monitor my GitHub for updates (as well as the code for this project, which should be up in 2-3 days).\n\nCongrats to the winners!\n\nThanks Kaggle, and Hosts! This competition was by far one of the most fun I have joined. I learned a ton.\n\nGitHub Training Code: https://github.com/ryanirl/hbac\nKaggle Final Inference: https://www.kaggle.com/code/ryanirl/hms-baseline-sub\n\nRyan Peters\nRyanIRL@icloud.com",
      "votes": 24
    },
    {
      "id": 2744531,
      "postDate": "2024-04-10T01:18:27.393Z",
      "content": "<p>Designed a custom kNN knowledge distillation scheme to pseudo-label samples with &lt;10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas.</p>\n<p>What does this mean? </p>",
      "rawMarkdown": "Designed a custom kNN knowledge distillation scheme to pseudo-label samples with <10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas.\n\nWhat does this mean? ",
      "votes": 1,
      "replies": [
        {
          "id": 2766481,
          "postDate": "2024-04-21T18:00:28.650Z",
          "content": "<p>Sorry for the late response. Basically what I did for pseudo-labeling was the following:</p>\n<ul>\n<li>Take each of my models for every fold that I trained on, and generate a database of embedding-label pairs for each sample in the validation set for that model (so data the model has NOT seen yet). Some pseudo-code for this might look like the following:</li>\n</ul>\n<pre><code>\n\n x, y  validation_dataloader: \n    embed = model.generate_embedding(x)\n    db.((embed, y))\n</code></pre>\n<ul>\n<li>Then during training, embed every item in the batch and look up it's nearest k-Neighbors in the saved database of embeddings, and use the mean of the ground truth labels stored in the database. Pseudo-code for this would be: </li>\n</ul>\n<pre><code> x, y  train_dataloader:\n    embed = model.generate\n    db_ys, db_embeds = db.k\n    new_y = db_ys.mean(axis = ) # `new_y` becomes our pseudo-label.\n</code></pre>\n<p>The <em>key</em> here is that when generating the embeddings in the first step (validation step)… we ONLY use the samples with count &gt;= 10. This ensures that we are only sampling from high quality labels.</p>\n<p>The reason I chose kNN pseudo-labeling over traditional knowledge-distillation was two-fold:</p>\n<ol>\n<li>The entropy of labels derived from knowledge-distillation is different, causing a distribution shift in the training losses. For example, using kNN you could get a new y that looks like [1, 0, 0, 0, 0, 0]… But using knowledge-distillation the new y would probably look more like [0.95, 0.1, 0.1, 0.1, 0.1. 0.1]. </li>\n<li>You can guarentee through kNN that all of the y's you select from are high-quality. Whereas in the knowledge distillation scheme your model still might be skewed towards the low quality data.</li>\n</ol>",
          "rawMarkdown": "Sorry for the late response. Basically what I did for pseudo-labeling was the following:\n * Take each of my models for every fold that I trained on, and generate a database of embedding-label pairs for each sample in the validation set for that model (so data the model has NOT seen yet). Some pseudo-code for this might look like the following:\n\n```\n# One key-point here is to only use high quality labels in \n# this step. See below for more details.\nfor x, y in validation_dataloader: \n    embed = model.generate_embedding(x)\n    db.add((embed, y))\n```\n\n * Then during training, embed every item in the batch and look up it's nearest k-Neighbors in the saved database of embeddings, and use the mean of the ground truth labels stored in the database. Pseudo-code for this would be: \n\n```\nfor x, y in train_dataloader:\n    embed = model.generate_embedding(x)\n    db_ys, db_embeds = db.k_nearest(embed, n = 11)\n    new_y = db_ys.mean(axis = 0) # `new_y` becomes our pseudo-label.\n```\n\nThe *key* here is that when generating the embeddings in the first step (validation step)... we ONLY use the samples with count >= 10. This ensures that we are only sampling from high quality labels.\n\nThe reason I chose kNN pseudo-labeling over traditional knowledge-distillation was two-fold:\n1. The entropy of labels derived from knowledge-distillation is different, causing a distribution shift in the training losses. For example, using kNN you could get a new y that looks like [1, 0, 0, 0, 0, 0]... But using knowledge-distillation the new y would probably look more like [0.95, 0.1, 0.1, 0.1, 0.1. 0.1]. \n2. You can guarentee through kNN that all of the y's you select from are high-quality. Whereas in the knowledge distillation scheme your model still might be skewed towards the low quality data."
        }
      ]
    },
    {
      "id": 2742767,
      "postDate": "2024-04-09T04:10:59.667Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": 1
    },
    {
      "id": 2742704,
      "postDate": "2024-04-09T03:14:29.767Z",
      "content": "<p>great work!</p>",
      "rawMarkdown": "great work!",
      "votes": 1
    },
    {
      "id": 2742622,
      "postDate": "2024-04-09T01:58:40.993Z",
      "content": "<p>this is an amazing solution, congrats</p>",
      "rawMarkdown": "this is an amazing solution, congrats",
      "votes": 1
    },
    {
      "id": 2742582,
      "postDate": "2024-04-09T01:15:53.693Z",
      "content": "<p>Great work!</p>\n<blockquote>\n  <p>Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.</p>\n</blockquote>\n<p>What does this mean?</p>\n<blockquote>\n  <p>Unfreeze with very small LR and retrain again</p>\n</blockquote>\n<p>How do you achieve this?</p>",
      "rawMarkdown": "Great work!\n \n> Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.\n\nWhat does this mean?\n\n> Unfreeze with very small LR and retrain again\n\nHow do you achieve this?",
      "votes": 1,
      "replies": [
        {
          "id": 2742612,
          "postDate": "2024-04-09T01:46:47.457Z",
          "content": "<blockquote>\n  <p>Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.</p>\n</blockquote>\n<p>I am using the <a href=\"https://www.learningeeg.com/montages-and-technical-components\" target=\"_blank\">bipolar double banana montage</a> which means that for each channel in the EEG data (let's say Fp1 and F7), we normalize each channel by another channel to get a relative value. This is done by subtracting them to get F7-Fp1 (or the other way around). But when you vertically flip the montage, the ordering of subtraction is also reversed (because I normalize vertically). So after vertically flipping, Fp7-Fp1 becomes Fp1-F7. So to fix this, you just flip the sign. At the end of the day, I don't think the flipping of signs is really all that crutial. You could probably even have an augmentation that reverses the sign of the whole montage and it would most likely still work out to be about the same. I just did it for consistency. Here is also an image to demonstrate the idea.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9041364%2F0852e7fba6dc02ae97821bc3c168285f%2FUntitled%20drawing%20(10).png?generation=1712626749233600&amp;alt=media\"></p>\n<blockquote>\n  <p>Unfreeze with very small LR and retrain again.</p>\n</blockquote>\n<p>Not sure exactly what your asking, but I probably didn't word this too well 😅. First, I trained 3 models for each of the modalities that I am considering (EEG, EEG spectrograms, and the Kaggle provided spectrograms). Then I take each of the 3 models that are pretrained and I combine them into my final multimodal model explained in my post. But I only train the decoder head (attention layer on the end), as I freeze the weights of the 3 models I previously trained. The reason being that I want the model to focus on learning how to combine the embeddings in the final attention layer and not learn to encode the data (which is the frozen and pretrained backbones job). Once the final attention head is performing well, I unfreeze the whole model (meaning training the whole model now and no longer freezing the backbone) and retrain on the data again but with a really small learning rate (LR) so the model doesn't make any dramatic changes as this is just a fine-tuning step.</p>",
          "rawMarkdown": "> Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.\n\nI am using the [bipolar double banana montage](https://www.learningeeg.com/montages-and-technical-components) which means that for each channel in the EEG data (let's say Fp1 and F7), we normalize each channel by another channel to get a relative value. This is done by subtracting them to get F7-Fp1 (or the other way around). But when you vertically flip the montage, the ordering of subtraction is also reversed (because I normalize vertically). So after vertically flipping, Fp7-Fp1 becomes Fp1-F7. So to fix this, you just flip the sign. At the end of the day, I don't think the flipping of signs is really all that crutial. You could probably even have an augmentation that reverses the sign of the whole montage and it would most likely still work out to be about the same. I just did it for consistency. Here is also an image to demonstrate the idea.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9041364%2F0852e7fba6dc02ae97821bc3c168285f%2FUntitled%20drawing%20(10).png?generation=1712626749233600&alt=media)\n\n> Unfreeze with very small LR and retrain again.\n\nNot sure exactly what your asking, but I probably didn't word this too well 😅. First, I trained 3 models for each of the modalities that I am considering (EEG, EEG spectrograms, and the Kaggle provided spectrograms). Then I take each of the 3 models that are pretrained and I combine them into my final multimodal model explained in my post. But I only train the decoder head (attention layer on the end), as I freeze the weights of the 3 models I previously trained. The reason being that I want the model to focus on learning how to combine the embeddings in the final attention layer and not learn to encode the data (which is the frozen and pretrained backbones job). Once the final attention head is performing well, I unfreeze the whole model (meaning training the whole model now and no longer freezing the backbone) and retrain on the data again but with a really small learning rate (LR) so the model doesn't make any dramatic changes as this is just a fine-tuning step.",
          "replies": [
            {
              "id": 2742618,
              "postDate": "2024-04-09T01:54:52.853Z",
              "content": "<p>Oh I see. Thanks for clarification!</p>",
              "rawMarkdown": "Oh I see. Thanks for clarification!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2742504,
      "postDate": "2024-04-09T00:22:14.533Z",
      "content": "<p>That was fast! Thank you so much for this quick sharing! This is really cool multimodal model. Congrats, well deserved!</p>",
      "rawMarkdown": "That was fast! Thank you so much for this quick sharing! This is really cool multimodal model. Congrats, well deserved!",
      "votes": 1,
      "replies": [
        {
          "id": 2742511,
          "postDate": "2024-04-09T00:27:06.450Z",
          "content": "<p>Thanks! I've been writing it for the past couple of hours now 😅.</p>",
          "rawMarkdown": "Thanks! I've been writing it for the past couple of hours now 😅.",
          "votes": 1,
          "replies": [
            {
              "id": 2742533,
              "postDate": "2024-04-09T00:40:03.793Z",
              "content": "<p>Thanks for the great writeup. I am a little curious of how do you make such a beautiful graph of the model architecture?🤩</p>",
              "rawMarkdown": "Thanks for the great writeup. I am a little curious of how do you make such a beautiful graph of the model architecture?🤩"
            },
            {
              "id": 2742537,
              "postDate": "2024-04-09T00:43:08.277Z",
              "content": "<p>Just spent a couple of hours in google drawings 😂.</p>",
              "rawMarkdown": "Just spent a couple of hours in google drawings 😂."
            }
          ]
        }
      ]
    },
    {
      "id": 2742503,
      "postDate": "2024-04-09T00:22:06.693Z",
      "content": "<p>Congratulations, impressive solution! I'd love to see how you accelerated it with KNN. How can I follow you on GitHub? Thank you.</p>",
      "rawMarkdown": "Congratulations, impressive solution! I'd love to see how you accelerated it with KNN. How can I follow you on GitHub? Thank you.",
      "votes": 1,
      "replies": [
        {
          "id": 2742514,
          "postDate": "2024-04-09T00:28:46.777Z",
          "content": "<p>Here's the link: <a href=\"https://github.com/ryanirl\" target=\"_blank\">https://github.com/ryanirl</a></p>\n<p>I'll have all of the code posted within 2-3 days, as soon as I finish migrating all of the code from my private repo to a public one.</p>",
          "rawMarkdown": "Here's the link: https://github.com/ryanirl\n\nI'll have all of the code posted within 2-3 days, as soon as I finish migrating all of the code from my private repo to a public one.",
          "votes": 2,
          "replies": [
            {
              "id": 2742676,
              "postDate": "2024-04-09T02:53:22.180Z",
              "content": "<p>Big thanks</p>",
              "rawMarkdown": "Big thanks"
            },
            {
              "id": 2756661,
              "postDate": "2024-04-17T06:03:56.337Z",
              "content": "<p>Looking forward to！Thank you.</p>",
              "rawMarkdown": "Looking forward to！Thank you."
            }
          ]
        }
      ]
    },
    {
      "id": 3159058,
      "postDate": "2025-03-25T08:04:27.267Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2744531,
      "author_name": "MeiYuxin",
      "author_url": "",
      "post_date": "2024-04-10T01:18:27.393000",
      "content": "<p>Designed a custom kNN knowledge distillation scheme to pseudo-label samples with &lt;10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas.</p>\n<p>What does this mean? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2766481,
          "author_name": "ryan",
          "author_url": "",
          "post_date": "2024-04-21T18:00:28.650000",
          "content": "<p>Sorry for the late response. Basically what I did for pseudo-labeling was the following:</p>\n<ul>\n<li>Take each of my models for every fold that I trained on, and generate a database of embedding-label pairs for each sample in the validation set for that model (so data the model has NOT seen yet). Some pseudo-code for this might look like the following:</li>\n</ul>\n<pre><code>\n\n x, y  validation_dataloader: \n    embed = model.generate_embedding(x)\n    db.((embed, y))\n</code></pre>\n<ul>\n<li>Then during training, embed every item in the batch and look up it's nearest k-Neighbors in the saved database of embeddings, and use the mean of the ground truth labels stored in the database. Pseudo-code for this would be: </li>\n</ul>\n<pre><code> x, y  train_dataloader:\n    embed = model.generate\n    db_ys, db_embeds = db.k\n    new_y = db_ys.mean(axis = ) # `new_y` becomes our pseudo-label.\n</code></pre>\n<p>The <em>key</em> here is that when generating the embeddings in the first step (validation step)… we ONLY use the samples with count &gt;= 10. This ensures that we are only sampling from high quality labels.</p>\n<p>The reason I chose kNN pseudo-labeling over traditional knowledge-distillation was two-fold:</p>\n<ol>\n<li>The entropy of labels derived from knowledge-distillation is different, causing a distribution shift in the training losses. For example, using kNN you could get a new y that looks like [1, 0, 0, 0, 0, 0]… But using knowledge-distillation the new y would probably look more like [0.95, 0.1, 0.1, 0.1, 0.1. 0.1]. </li>\n<li>You can guarentee through kNN that all of the y's you select from are high-quality. Whereas in the knowledge distillation scheme your model still might be skewed towards the low quality data.</li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2742767,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-09T04:10:59.667000",
      "content": "<p>Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2742704,
      "author_name": "MiHu",
      "author_url": "",
      "post_date": "2024-04-09T03:14:29.767000",
      "content": "<p>great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2742622,
      "author_name": "MengYe",
      "author_url": "",
      "post_date": "2024-04-09T01:58:40.993000",
      "content": "<p>this is an amazing solution, congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2742582,
      "author_name": "Roy Wei",
      "author_url": "",
      "post_date": "2024-04-09T01:15:53.693000",
      "content": "<p>Great work!</p>\n<blockquote>\n  <p>Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.</p>\n</blockquote>\n<p>What does this mean?</p>\n<blockquote>\n  <p>Unfreeze with very small LR and retrain again</p>\n</blockquote>\n<p>How do you achieve this?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2742612,
          "author_name": "ryan",
          "author_url": "",
          "post_date": "2024-04-09T01:46:47.457000",
          "content": "<blockquote>\n  <p>Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.</p>\n</blockquote>\n<p>I am using the <a href=\"https://www.learningeeg.com/montages-and-technical-components\" target=\"_blank\">bipolar double banana montage</a> which means that for each channel in the EEG data (let's say Fp1 and F7), we normalize each channel by another channel to get a relative value. This is done by subtracting them to get F7-Fp1 (or the other way around). But when you vertically flip the montage, the ordering of subtraction is also reversed (because I normalize vertically). So after vertically flipping, Fp7-Fp1 becomes Fp1-F7. So to fix this, you just flip the sign. At the end of the day, I don't think the flipping of signs is really all that crutial. You could probably even have an augmentation that reverses the sign of the whole montage and it would most likely still work out to be about the same. I just did it for consistency. Here is also an image to demonstrate the idea.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9041364%2F0852e7fba6dc02ae97821bc3c168285f%2FUntitled%20drawing%20(10).png?generation=1712626749233600&amp;alt=media\"></p>\n<blockquote>\n  <p>Unfreeze with very small LR and retrain again.</p>\n</blockquote>\n<p>Not sure exactly what your asking, but I probably didn't word this too well 😅. First, I trained 3 models for each of the modalities that I am considering (EEG, EEG spectrograms, and the Kaggle provided spectrograms). Then I take each of the 3 models that are pretrained and I combine them into my final multimodal model explained in my post. But I only train the decoder head (attention layer on the end), as I freeze the weights of the 3 models I previously trained. The reason being that I want the model to focus on learning how to combine the embeddings in the final attention layer and not learn to encode the data (which is the frozen and pretrained backbones job). Once the final attention head is performing well, I unfreeze the whole model (meaning training the whole model now and no longer freezing the backbone) and retrain on the data again but with a really small learning rate (LR) so the model doesn't make any dramatic changes as this is just a fine-tuning step.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2742618,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2024-04-09T01:54:52.853000",
              "content": "<p>Oh I see. Thanks for clarification!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2742504,
      "author_name": "Steven_Y",
      "author_url": "",
      "post_date": "2024-04-09T00:22:14.533000",
      "content": "<p>That was fast! Thank you so much for this quick sharing! This is really cool multimodal model. Congrats, well deserved!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2742511,
          "author_name": "ryan",
          "author_url": "",
          "post_date": "2024-04-09T00:27:06.450000",
          "content": "<p>Thanks! I've been writing it for the past couple of hours now 😅.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2742533,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2024-04-09T00:40:03.793000",
              "content": "<p>Thanks for the great writeup. I am a little curious of how do you make such a beautiful graph of the model architecture?🤩</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2742537,
              "author_name": "ryan",
              "author_url": "",
              "post_date": "2024-04-09T00:43:08.277000",
              "content": "<p>Just spent a couple of hours in google drawings 😂.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2742503,
      "author_name": "kerry sun",
      "author_url": "",
      "post_date": "2024-04-09T00:22:06.693000",
      "content": "<p>Congratulations, impressive solution! I'd love to see how you accelerated it with KNN. How can I follow you on GitHub? Thank you.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2742514,
          "author_name": "ryan",
          "author_url": "",
          "post_date": "2024-04-09T00:28:46.777000",
          "content": "<p>Here's the link: <a href=\"https://github.com/ryanirl\" target=\"_blank\">https://github.com/ryanirl</a></p>\n<p>I'll have all of the code posted within 2-3 days, as soon as I finish migrating all of the code from my private repo to a public one.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2742676,
              "author_name": "kerry sun",
              "author_url": "",
              "post_date": "2024-04-09T02:53:22.180000",
              "content": "<p>Big thanks</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2756661,
              "author_name": "Kris_Lan",
              "author_url": "",
              "post_date": "2024-04-17T06:03:56.337000",
              "content": "<p>Looking forward to！Thank you.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3159058,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-25T08:04:27.267000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742494": "Huge thanks to the Hosts and Kaggle, this competition was by far one of the most fun competitions I've competed in awhile. \n\n<br/>\n\n**TL;DR:**\n - My final model was a single multimodal attention-style model with CNN and CNN-RNN encoders. I treat each *node* (ex: Fp1-F7, or LP for spectrograms) as an independent item that needs to be embedded (27 items total), then use additive attention with learned positional encoding to attend to each of the samples, then pool and decode. I also used sample-dependent adaptive temperature scaling (for interpretability, neither increased nor decreased CV/LB score). \n - Designed a custom kNN knowledge distillation scheme to pseudo-label samples with <10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas. \n - No pretrained models, all CNN's were trained from scratch.\n - Robust normalization the 1D EEGs by the mean absolute deviation (MAD) instead of standard deviation.\n\n**Scores:**\n\n - Single 1D RAW EEG: \n   - 0.28 CV | 0.28 Public LB | 0.34 Private LB\n   - ~2M params\n - Single 2D EEG spectrograms: \n   - 0.32 CV | 0.33 Public LB | 0.4 Private LB\n   - ~2M params\n - Single 2D Kaggle Spectrograms: \n   - 0.48 CV | 0.51 Public LB | 0.59 Private LB\n   - ~2M params\n - Final Multimodal Model: \n   - 0.26 CV | 0.28 Public LB | 0.33 Private LB\n   - 10M params\n\n\nI selected two models, one for best CV on the whole dataset and one for best CV on samples >10 in case there was a shakeup. No shakeup, trusted my personal CV and it payed off :). Given that my original goal was ~0.29 when I join this competition a month ago (a bit late, really had to grind), I am very happy with my performance.\n\nI feel if I had a bit more time, ~0.25 would be achievable (local CV / public LB) with my current solution, but that something was missing between what I have now and a sub 0.25 CV/LB score (probably some form of data psuedo-labeling, preprocessing, and/or selection). Can't wait to read everyone elses solutions and see how sub 0.25 scores were achieved. \n\n\n### Data Preprocessing \n\nFor the EEG, I used the banana montage (including the middle chain and the EKG). Nothing here is too different from what I saw other people doing, other than the fact that I normalize all 1D EEG signals by the MAD for robustness.\n\nProprocessing (EEG): \n  - Butter bandpass filter to only keep frequencies between 0.5 and 50 Hz. \n  - 4x mean downsampling to 50 Hz.\n  - Mean subtraction, and division by the *median* mean absolute deviation (MAD) across *all* channels for robust estimation of the standard deviation (see next section). That is, global normalization and not per signal normalization. I choose to do this because I think it's important to keep all of the relative scales the same. \n  - Clip between -10 and 10.\n\nProprocessing (EEG Spectrogram): \n  - Butter bandpass filter of the raw signal to keep frequencies between 0.5 and 40 Hz.\n  - Normalized the signal by mean subtraction and MAD normalization first.\n  - Used the torchaudio Spectrogram function with hop_length 44, win_length 256, and n_fft 800. \n  - Clippd between $e^{-4}$ and $e^{7}$.\n  - Took the log.\n\nProprocessing (Spectrogram): \n  - Clippd between $e^{-4}$ and $e^{7}$.\n  - Took the log.\n  - Mean/std normalized each spectrogram (NOT global).\n\n\n### Robust Normalization\n\nIn this competition especially, I found it very important to make sure that your normalize your data in a way that was robust to artifacts. And after playing around with quite a lot of different normalizations and visualizing the raw data, it was clear to me that the traditional standard deviation was NOT a robust measure and that it was extremely susceptible to artifacts in the case of the 1D EEGs. To this end, I instead chose to divide by the mean absolute deviation (MAD) and clip values between -10 and 10. This spead up convergence quite a bit, reduce overfitting, and also seemed to generalize better to my validation set. \n\n\n### Augmentations and TTA\n\n1D EEGs: \n  - Horizontally and vertically flip (each with p = 0.5) the ordering of the nodes in the banana montage. Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.\n  - Randomly zero out 1-4 channels at a time with probability 0.5. \n\n2D Spectroggrams:\n  - During training, I did minimal augmentations to the spectrograms, other than swapping the direction of the chains (both ll <-> rl and lp <-> rp). This is the analog of the horizontal flip mentioned in the 1D EEGs.\n\nI tried working with test-time-augmentations (TTA) like flipping the orientation of chains, but it neither increased or decreased my CV/LB score in my later models. I think this is because I treat each node as an independent sample and used the same TTAs as augmentations during training. Which means by model tends to be robust to these shifts.\n\n### Model Architecture\n\nWhen I started this competition, I had a couple of goals for my model architecture that I wanted to achieve.\n 1. Scalable \n 2. The general architecture made no hard assumptions about the spatial layout OR modality of the data, but could instead learn both the spatial layout and to incorperate new modalities on it's own. \n 3. Semi-interpretable (not sure my solution achieved this, but I have ideas on how to achieve this going forward, see the *Extra* section at the bottom).\n\nTo achieve these goals, my model architecture for each modality followed the same set of encoding/decoding principles:\n 1. Split the data up into nodes (ex: F7-Fp1 or LP).\n 2. Embed each node.\n 3. Use an attention style decoder with learned positional encodings to model the relationship between each node. \n 4. Pool and then linear decoding.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9041364%2F17977936ff43f954b6ee62fca3e706d4%2FUntitled%20drawing%20(8).png?generation=1712622281054757&alt=media)\n\n### Training Scheme\n\nMy training stages were an absolute disaster. I had 5-6 stages where I would start with easy samples to get the model to converge quick, then moved on to hard samples, and finalized with training on just the high confidence (>10 votes) samples. \n\n### Extra: Future Ideas, Interpretability, and Kaggle\n\nSince the beginning of this competition, I've put a bit of thought into the future directions of this project after the competition and have some ideas for the competition hosts regarding how I could imagine technology like this being used in practice. \n\nOne of a couple ideas that I've had would be to design an interpretable model with a kNN decoder. For example, imagine you deploy a model like this in the real world as an aid to doctors who are labeling or interpreting this data. Then, upon making a prediction, the model not only gives you a probabilistic output of what it thinks is true, but also looks up the nearest N neighbors neighbors in the embedding space (of the training set) and shows you those N samples as to say \"here is my prediction, but here are also 10 (for example) other EEGs and spectrograms that have similar patterns to this one and hence why I labeled it this way.\"\n\nI have a lot of ideas like this that I'm going to be exploring over this summer as a side project (including reading/implementing other *unique* top submissions). I'll also be open sourcing it, so feel free to monitor my GitHub for updates (as well as the code for this project, which should be up in 2-3 days).\n\nCongrats to the winners!\n\nThanks Kaggle, and Hosts! This competition was by far one of the most fun I have joined. I learned a ton.\n\nGitHub Training Code: https://github.com/ryanirl/hbac\nKaggle Final Inference: https://www.kaggle.com/code/ryanirl/hms-baseline-sub\n\nRyan Peters\nRyanIRL@icloud.com",
    "2744531": "Designed a custom kNN knowledge distillation scheme to pseudo-label samples with <10 votes. I won't go into it here, but I choose kNN so the entropy of labels would match almost exactly. I couldn't get this to increase my personal CV, but it did speed up training/convergence a lot which allowed me to rapidly test out ideas.\n\nWhat does this mean? ",
    "2742767": "Great work!",
    "2742704": "great work!",
    "2742622": "this is an amazing solution, congrats",
    "2742582": "Great work!\n \n> Notice that when vertically flipping, you must also reverse the sign of the data because the order of subtraction is changed.\n\nWhat does this mean?\n\n> Unfreeze with very small LR and retrain again\n\nHow do you achieve this?",
    "2742504": "That was fast! Thank you so much for this quick sharing! This is really cool multimodal model. Congrats, well deserved!",
    "2742503": "Congratulations, impressive solution! I'd love to see how you accelerated it with KNN. How can I follow you on GitHub? Thank you.",
    "3159058": ""
  }
}