{
  "id": 469666,
  "title": "Features+Head TF Starter - LB 0.34 (ENSEMBLE)",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/469666",
  "author_name": "",
  "post_date": "2024-01-21T15:27:33.174633600Z",
  "votes": 29,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Greetings everyone,</p>\n<p>I'm excited to share my <a href=\"https://www.kaggle.com/code/nartaa/features-head-starter\" target=\"_blank\">Features+Head Starter</a>, which achieves a competitive [LB 0.34]. This notebook is built upon the foundation laid by Chris in his remarkable work. Along side a great insight from <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a> in their post <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461\" target=\"_blank\">here</a> about two stage training!</p>\n<p>With this starter notebook, we can train single models, save them, then submit them for thier individual LBs, then submit them with weighted ensemble. (All models included)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F84e82be8053a16525be76ec7ae1c39c2%2FScreenshot%20from%202024-02-28%2023-51-16.png?generation=1709153551808393&amp;alt=media\"></p>\n<p>Potential improvements:</p>\n<ol>\n<li><strong>Highlight the 10 seconds in the middle of the spectograms</strong> since experts pay attention to the 10 seconds in the middle, I could bring that section of the image to the model's attention.</li>\n<li><strong>Use custom loss</strong> I could use a weighted loss that calculates both Binary Cross Entropy and Kullback–Leibler Divergence.</li>\n<li><strong>Data Augmentation</strong> use <a href=\"https://keras.io/examples/vision/mixup/\" target=\"_blank\">MixUp</a>, and Implement various data augmentation techniques, such as flipping, rotating the spectrogram images.</li>\n<li><strong>Weighted voting</strong> Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.<br>\nExamples: (votes -&gt; weighted probability)<br>\n1- [2, 2, 0, 0, 0, 0] -&gt; [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]<br>\n2- [10, 10, 0, 0, 0, 0] -&gt; [0.5, 0.5, 0, 0, 0, 0]<br>\nSome math formula is needed.</li>\n<li><strong><a href=\"https://keras.io/guides/distribution/\" target=\"_blank\">Distributed Learning</a></strong> I should implement custom training lopps to support multi-GPU and TPU, because freezing the weights for the backbone model doesn't seem to work with the fit method.</li>\n<li><strong>Stem for backbone</strong> I could build a custom stem and replace part of the original stem in the backbone(reconnect the normalization layers from the original to the new added stem), this way a (IMG_SIZE,IMG_SIZE,8) shape can be passed and reduced to 3 channels with trainable weights.</li>\n<li><strong>Dimensionality Reduction with Autoencoders</strong> I could train an autoencoder to learn the hidden representaion of the 8 channel images to a 3 channel images.</li>\n<li><strong>Sequential Model</strong> Since EGG is a time series signal, I could transfer learn and fine tune a Transformer Encoder for the classification task.</li>\n<li><strong>Machine Learning for Classification:</strong> With the head and feature models separated, we can explore employing machine learning techniques for classification, including random forest, gradient boost models, stacking, and blending.</li>\n<li><strong><a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">Freezing Backbone Weights</a>:</strong> Consider freezing most of the trainable weights of the EfficientNetB2 backbone, while selectively keeping specific layers trainable on the upper layers. This approach allows for prolonged training of the backend, focusing on layers with higher-level features.</li>\n<li><strong>Experimenting with Different Backbones:</strong> Explore training and ensembling other backbone models to further enhance performance.</li>\n<li><strong>Remove Outliers and Noise</strong> further investigation to preprocess the data, looking at the UMAP, it seems outliers exist that can be removed, also missing values (Nans) exist in spectrograms, and possibly EEG.</li>\n<li><strong>Reproduce EEG Spectrograms</strong> I could recalculate EEG spectograms, fill NANs before any operation, remove log transformation and standardization and do that when generating data.</li>\n<li><strong>Input Dropout</strong> Since the input image is a monotone that is repeated on the three channels, I could implement a random dropout in the generator, I shaould try dropout of complete channels, keep one since it has all the information, another approach is partial dropout from each channel.</li>\n</ol>\n<p>Let me know if you come across any issues, or have any suggestions.</p>",
  "messages": [
    {
      "id": "2612730",
      "postDate": "01/21/2024 15:27:33",
      "content": "<p>Greetings everyone,</p>\n<p>I'm excited to share my <a href=\"https://www.kaggle.com/code/nartaa/features-head-starter\" target=\"_blank\">Features+Head Starter</a>, which achieves a competitive [LB 0.34]. This notebook is built upon the foundation laid by Chris in his remarkable work. Along side a great insight from <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a> in their post <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461\" target=\"_blank\">here</a> about two stage training!</p>\n<p>With this starter notebook, we can train single models, save them, then submit them for thier individual LBs, then submit them with weighted ensemble. (All models included)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F84e82be8053a16525be76ec7ae1c39c2%2FScreenshot%20from%202024-02-28%2023-51-16.png?generation=1709153551808393&amp;alt=media\"></p>\n<p>Potential improvements:</p>\n<ol>\n<li><strong>Highlight the 10 seconds in the middle of the spectograms</strong> since experts pay attention to the 10 seconds in the middle, I could bring that section of the image to the model's attention.</li>\n<li><strong>Use custom loss</strong> I could use a weighted loss that calculates both Binary Cross Entropy and Kullback–Leibler Divergence.</li>\n<li><strong>Data Augmentation</strong> use <a href=\"https://keras.io/examples/vision/mixup/\" target=\"_blank\">MixUp</a>, and Implement various data augmentation techniques, such as flipping, rotating the spectrogram images.</li>\n<li><strong>Weighted voting</strong> Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.<br>\nExamples: (votes -&gt; weighted probability)<br>\n1- [2, 2, 0, 0, 0, 0] -&gt; [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]<br>\n2- [10, 10, 0, 0, 0, 0] -&gt; [0.5, 0.5, 0, 0, 0, 0]<br>\nSome math formula is needed.</li>\n<li><strong><a href=\"https://keras.io/guides/distribution/\" target=\"_blank\">Distributed Learning</a></strong> I should implement custom training lopps to support multi-GPU and TPU, because freezing the weights for the backbone model doesn't seem to work with the fit method.</li>\n<li><strong>Stem for backbone</strong> I could build a custom stem and replace part of the original stem in the backbone(reconnect the normalization layers from the original to the new added stem), this way a (IMG_SIZE,IMG_SIZE,8) shape can be passed and reduced to 3 channels with trainable weights.</li>\n<li><strong>Dimensionality Reduction with Autoencoders</strong> I could train an autoencoder to learn the hidden representaion of the 8 channel images to a 3 channel images.</li>\n<li><strong>Sequential Model</strong> Since EGG is a time series signal, I could transfer learn and fine tune a Transformer Encoder for the classification task.</li>\n<li><strong>Machine Learning for Classification:</strong> With the head and feature models separated, we can explore employing machine learning techniques for classification, including random forest, gradient boost models, stacking, and blending.</li>\n<li><strong><a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">Freezing Backbone Weights</a>:</strong> Consider freezing most of the trainable weights of the EfficientNetB2 backbone, while selectively keeping specific layers trainable on the upper layers. This approach allows for prolonged training of the backend, focusing on layers with higher-level features.</li>\n<li><strong>Experimenting with Different Backbones:</strong> Explore training and ensembling other backbone models to further enhance performance.</li>\n<li><strong>Remove Outliers and Noise</strong> further investigation to preprocess the data, looking at the UMAP, it seems outliers exist that can be removed, also missing values (Nans) exist in spectrograms, and possibly EEG.</li>\n<li><strong>Reproduce EEG Spectrograms</strong> I could recalculate EEG spectograms, fill NANs before any operation, remove log transformation and standardization and do that when generating data.</li>\n<li><strong>Input Dropout</strong> Since the input image is a monotone that is repeated on the three channels, I could implement a random dropout in the generator, I shaould try dropout of complete channels, keep one since it has all the information, another approach is partial dropout from each channel.</li>\n</ol>\n<p>Let me know if you come across any issues, or have any suggestions.</p>",
      "rawMarkdown": "Greetings everyone,\n\nI'm excited to share my [Features+Head Starter](https://www.kaggle.com/code/nartaa/features-head-starter), which achieves a competitive [LB 0.34]. This notebook is built upon the foundation laid by Chris in his remarkable work. Along side a great insight from @zijiangyang1116 in their post [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461) about two stage training!\n\nWith this starter notebook, we can train single models, save them, then submit them for thier individual LBs, then submit them with weighted ensemble. (All models included)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F84e82be8053a16525be76ec7ae1c39c2%2FScreenshot%20from%202024-02-28%2023-51-16.png?generation=1709153551808393&alt=media)\n\nPotential improvements:\n1. **Highlight the 10 seconds in the middle of the spectograms** since experts pay attention to the 10 seconds in the middle, I could bring that section of the image to the model's attention.\n2. **Use custom loss** I could use a weighted loss that calculates both Binary Cross Entropy and Kullback–Leibler Divergence.\n3. **Data Augmentation** use [MixUp](https://keras.io/examples/vision/mixup/), and Implement various data augmentation techniques, such as flipping, rotating the spectrogram images.\n4. **Weighted voting** Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.\nExamples: (votes -> weighted probability)\n1- [2, 2, 0, 0, 0, 0] -> [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]\n2- [10, 10, 0, 0, 0, 0] -> [0.5, 0.5, 0, 0, 0, 0]\nSome math formula is needed.\n5. **[Distributed Learning](https://keras.io/guides/distribution/)** I should implement custom training lopps to support multi-GPU and TPU, because freezing the weights for the backbone model doesn't seem to work with the fit method.\n6. **Stem for backbone** I could build a custom stem and replace part of the original stem in the backbone(reconnect the normalization layers from the original to the new added stem), this way a (IMG_SIZE,IMG_SIZE,8) shape can be passed and reduced to 3 channels with trainable weights.\n7. **Dimensionality Reduction with Autoencoders** I could train an autoencoder to learn the hidden representaion of the 8 channel images to a 3 channel images.\n8. **Sequential Model** Since EGG is a time series signal, I could transfer learn and fine tune a Transformer Encoder for the classification task.\n9. **Machine Learning for Classification:** With the head and feature models separated, we can explore employing machine learning techniques for classification, including random forest, gradient boost models, stacking, and blending.\n10. **[Freezing Backbone Weights](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/):** Consider freezing most of the trainable weights of the EfficientNetB2 backbone, while selectively keeping specific layers trainable on the upper layers. This approach allows for prolonged training of the backend, focusing on layers with higher-level features.\n11. **Experimenting with Different Backbones:** Explore training and ensembling other backbone models to further enhance performance.\n12. **Remove Outliers and Noise** further investigation to preprocess the data, looking at the UMAP, it seems outliers exist that can be removed, also missing values (Nans) exist in spectrograms, and possibly EEG.\n13. **Reproduce EEG Spectrograms** I could recalculate EEG spectograms, fill NANs before any operation, remove log transformation and standardization and do that when generating data.\n14. **Input Dropout** Since the input image is a monotone that is repeated on the three channels, I could implement a random dropout in the generator, I shaould try dropout of complete channels, keep one since it has all the information, another approach is partial dropout from each channel.\n\nLet me know if you come across any issues, or have any suggestions.",
      "votes": null
    },
    {
      "id": "2612765",
      "postDate": "01/21/2024 16:01:59",
      "content": "<p>Would I be correct in my assumption that the only reason for the higher score is using votes instead of consensus? Or is there an additional reason for the improved accuracy?<br>\nAnyway, it is nice to see a higher score TF notebook. TY.</p>",
      "rawMarkdown": "Would I be correct in my assumption that the only reason for the higher score is using votes instead of consensus? Or is there an additional reason for the improved accuracy?\nAnyway, it is nice to see a higher score TF notebook. TY.",
      "votes": null
    },
    {
      "id": "2612854",
      "postDate": "01/21/2024 17:00:21",
      "content": "<p>Hello! Why did you choose the P100 instead of the T4 x2? With these modifications to the code, how many hours were required for training? Chris's EfficientNet training was completed in 1 hour.</p>",
      "rawMarkdown": "Hello! Why did you choose the P100 instead of the T4 x2? With these modifications to the code, how many hours were required for training? Chris's EfficientNet training was completed in 1 hour.",
      "votes": null
    },
    {
      "id": "2612863",
      "postDate": "01/21/2024 17:07:58",
      "content": "<p>I had some issues utilizing the T4x2 which I couldn't resolve quickly, but I will probably try to use it again. <br>\nWith P100 it took 1.5 hours.</p>",
      "rawMarkdown": "I had some issues utilizing the T4x2 which I couldn't resolve quickly, but I will probably try to use it again. \nWith P100 it took 1.5 hours.",
      "votes": null
    },
    {
      "id": "2612872",
      "postDate": "01/21/2024 17:15:12",
      "content": "<p>Using the votes reduced LB from 0.52 to 0.49<br>\nTraining the head separately reduced LB from 0.57 to 0.52<br>\nYW.</p>",
      "rawMarkdown": "Using the votes reduced LB from 0.52 to 0.49\nTraining the head separately reduced LB from 0.57 to 0.52\nYW.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2612765,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "01/21/2024 16:01:59",
      "content": "<p>Would I be correct in my assumption that the only reason for the higher score is using votes instead of consensus? Or is there an additional reason for the improved accuracy?<br>\nAnyway, it is nice to see a higher score TF notebook. TY.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2612872,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "01/21/2024 17:15:12",
          "content": "<p>Using the votes reduced LB from 0.52 to 0.49<br>\nTraining the head separately reduced LB from 0.57 to 0.52<br>\nYW.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2612854,
      "author_name": "yantxx",
      "author_url": "",
      "post_date": "01/21/2024 17:00:21",
      "content": "<p>Hello! Why did you choose the P100 instead of the T4 x2? With these modifications to the code, how many hours were required for training? Chris's EfficientNet training was completed in 1 hour.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2612863,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "01/21/2024 17:07:58",
          "content": "<p>I had some issues utilizing the T4x2 which I couldn't resolve quickly, but I will probably try to use it again. <br>\nWith P100 it took 1.5 hours.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2612730": "Greetings everyone,\n\nI'm excited to share my [Features+Head Starter](https://www.kaggle.com/code/nartaa/features-head-starter), which achieves a competitive [LB 0.34]. This notebook is built upon the foundation laid by Chris in his remarkable work. Along side a great insight from @zijiangyang1116 in their post [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461) about two stage training!\n\nWith this starter notebook, we can train single models, save them, then submit them for thier individual LBs, then submit them with weighted ensemble. (All models included)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F84e82be8053a16525be76ec7ae1c39c2%2FScreenshot%20from%202024-02-28%2023-51-16.png?generation=1709153551808393&alt=media)\n\nPotential improvements:\n1. **Highlight the 10 seconds in the middle of the spectograms** since experts pay attention to the 10 seconds in the middle, I could bring that section of the image to the model's attention.\n2. **Use custom loss** I could use a weighted loss that calculates both Binary Cross Entropy and Kullback–Leibler Divergence.\n3. **Data Augmentation** use [MixUp](https://keras.io/examples/vision/mixup/), and Implement various data augmentation techniques, such as flipping, rotating the spectrogram images.\n4. **Weighted voting** Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.\nExamples: (votes -> weighted probability)\n1- [2, 2, 0, 0, 0, 0] -> [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]\n2- [10, 10, 0, 0, 0, 0] -> [0.5, 0.5, 0, 0, 0, 0]\nSome math formula is needed.\n5. **[Distributed Learning](https://keras.io/guides/distribution/)** I should implement custom training lopps to support multi-GPU and TPU, because freezing the weights for the backbone model doesn't seem to work with the fit method.\n6. **Stem for backbone** I could build a custom stem and replace part of the original stem in the backbone(reconnect the normalization layers from the original to the new added stem), this way a (IMG_SIZE,IMG_SIZE,8) shape can be passed and reduced to 3 channels with trainable weights.\n7. **Dimensionality Reduction with Autoencoders** I could train an autoencoder to learn the hidden representaion of the 8 channel images to a 3 channel images.\n8. **Sequential Model** Since EGG is a time series signal, I could transfer learn and fine tune a Transformer Encoder for the classification task.\n9. **Machine Learning for Classification:** With the head and feature models separated, we can explore employing machine learning techniques for classification, including random forest, gradient boost models, stacking, and blending.\n10. **[Freezing Backbone Weights](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/):** Consider freezing most of the trainable weights of the EfficientNetB2 backbone, while selectively keeping specific layers trainable on the upper layers. This approach allows for prolonged training of the backend, focusing on layers with higher-level features.\n11. **Experimenting with Different Backbones:** Explore training and ensembling other backbone models to further enhance performance.\n12. **Remove Outliers and Noise** further investigation to preprocess the data, looking at the UMAP, it seems outliers exist that can be removed, also missing values (Nans) exist in spectrograms, and possibly EEG.\n13. **Reproduce EEG Spectrograms** I could recalculate EEG spectograms, fill NANs before any operation, remove log transformation and standardization and do that when generating data.\n14. **Input Dropout** Since the input image is a monotone that is repeated on the three channels, I could implement a random dropout in the generator, I shaould try dropout of complete channels, keep one since it has all the information, another approach is partial dropout from each channel.\n\nLet me know if you come across any issues, or have any suggestions.",
    "2612765": "Would I be correct in my assumption that the only reason for the higher score is using votes instead of consensus? Or is there an additional reason for the improved accuracy?\nAnyway, it is nice to see a higher score TF notebook. TY.",
    "2612854": "Hello! Why did you choose the P100 instead of the T4 x2? With these modifications to the code, how many hours were required for training? Chris's EfficientNet training was completed in 1 hour.",
    "2612863": "I had some issues utilizing the T4x2 which I couldn't resolve quickly, but I will probably try to use it again. \nWith P100 it took 1.5 hours.",
    "2612872": "Using the votes reduced LB from 0.52 to 0.49\nTraining the head separately reduced LB from 0.57 to 0.52\nYW."
  },
  "source": "meta"
}