{
  "id": 473306,
  "title": "Experimenting with different models for EEG classification",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/473306",
  "author_name": "",
  "post_date": "2024-02-04T09:00:09.841233200Z",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello kagglers,</p>\n<p>Alongside EfficientNetB0 and ResNet34d, I have (successfully) trained <a href=\"https://www.kaggle.com/code/andreasbis/hms-train-efficientnetb1\" target=\"_blank\">EfficientNetB1</a>, using the same batch size of 16. Since most participants use the first two models, I have been experimenting for the previous days with different models and optimizers: EfficientNetB1 (AdamW optimizer), EfficientNetB3 (RMSProp optimizer) and ResNet50 (SGD optimizer, with momentum). From the first epoch, I concluded that larger models, namely EfficientNetB3 and ResNet50 perform worse compared to lighter models, scoring around 1.25 on the Kullback Liebler divergence loss. </p>\n<p>As for EfficientNetB1, despite doing better on the CV than EfficientNetB0 and ResNet34d, having a best score of 0.53 compared to 0.55 for the other two, it did slightly worse on the public LB: 0.46 compared to 0.42 (for EfficientNetB0) and 0.45 (for ResNet34d).</p>\n<p>At this juncture, my approach is to create a weighted ensembler from the three models (ResNet34d, EffNetB0 and EffNetB1), hoping to get a better score than 0.42 on the public LB. </p>\n<p>If you want to experiment with the weights of EfficientNetB1, you can visit <a href=\"https://www.kaggle.com/datasets/andreasbis/efficientnetb1-for-eeg-classification-weights\" target=\"_blank\">this dataset</a>.</p>",
  "messages": [
    {
      "id": "2635260",
      "postDate": "02/04/2024 09:00:09",
      "content": "<p>Hello kagglers,</p>\n<p>Alongside EfficientNetB0 and ResNet34d, I have (successfully) trained <a href=\"https://www.kaggle.com/code/andreasbis/hms-train-efficientnetb1\" target=\"_blank\">EfficientNetB1</a>, using the same batch size of 16. Since most participants use the first two models, I have been experimenting for the previous days with different models and optimizers: EfficientNetB1 (AdamW optimizer), EfficientNetB3 (RMSProp optimizer) and ResNet50 (SGD optimizer, with momentum). From the first epoch, I concluded that larger models, namely EfficientNetB3 and ResNet50 perform worse compared to lighter models, scoring around 1.25 on the Kullback Liebler divergence loss. </p>\n<p>As for EfficientNetB1, despite doing better on the CV than EfficientNetB0 and ResNet34d, having a best score of 0.53 compared to 0.55 for the other two, it did slightly worse on the public LB: 0.46 compared to 0.42 (for EfficientNetB0) and 0.45 (for ResNet34d).</p>\n<p>At this juncture, my approach is to create a weighted ensembler from the three models (ResNet34d, EffNetB0 and EffNetB1), hoping to get a better score than 0.42 on the public LB. </p>\n<p>If you want to experiment with the weights of EfficientNetB1, you can visit <a href=\"https://www.kaggle.com/datasets/andreasbis/efficientnetb1-for-eeg-classification-weights\" target=\"_blank\">this dataset</a>.</p>",
      "rawMarkdown": "Hello kagglers,\n\nAlongside EfficientNetB0 and ResNet34d, I have (successfully) trained [EfficientNetB1](https://www.kaggle.com/code/andreasbis/hms-train-efficientnetb1), using the same batch size of 16. Since most participants use the first two models, I have been experimenting for the previous days with different models and optimizers: EfficientNetB1 (AdamW optimizer), EfficientNetB3 (RMSProp optimizer) and ResNet50 (SGD optimizer, with momentum). From the first epoch, I concluded that larger models, namely EfficientNetB3 and ResNet50 perform worse compared to lighter models, scoring around 1.25 on the Kullback Liebler divergence loss. \n\nAs for EfficientNetB1, despite doing better on the CV than EfficientNetB0 and ResNet34d, having a best score of 0.53 compared to 0.55 for the other two, it did slightly worse on the public LB: 0.46 compared to 0.42 (for EfficientNetB0) and 0.45 (for ResNet34d).\n\nAt this juncture, my approach is to create a weighted ensembler from the three models (ResNet34d, EffNetB0 and EffNetB1), hoping to get a better score than 0.42 on the public LB. \n\nIf you want to experiment with the weights of EfficientNetB1, you can visit [this dataset](https://www.kaggle.com/datasets/andreasbis/efficientnetb1-for-eeg-classification-weights).",
      "votes": null
    },
    {
      "id": "2635421",
      "postDate": "02/04/2024 11:45:56",
      "content": "<p>I think this is the effect of ensembling, you are using only one model for your CV and the average of 5 for your LB.<br>\nThere is a variation ('underfitting') for a smaller architecture which can give a nice boost when ensembled, a bigger architecture tends to overfit very quickly<br>\nIt might be better to work with effnetb0 for now and optimize your hyperparameters or the way you process the data</p>",
      "rawMarkdown": "I think this is the effect of ensembling, you are using only one model for your CV and the average of 5 for your LB.\nThere is a variation ('underfitting') for a smaller architecture which can give a nice boost when ensembled, a bigger architecture tends to overfit very quickly\nIt might be better to work with effnetb0 for now and optimize your hyperparameters or the way you process the data",
      "votes": null
    },
    {
      "id": "2635644",
      "postDate": "02/04/2024 14:33:57",
      "content": "<p>Indeed, a weighted ensembler of the three models (ResNet34d, EffNetB0, EffNetB1) did better than the inidividual models, scoring 0.41 on the public LB. </p>\n<p>Nonetheless, considering your place in the competition, how do you approach data preprocessing? </p>",
      "rawMarkdown": "Indeed, a weighted ensembler of the three models (ResNet34d, EffNetB0, EffNetB1) did better than the inidividual models, scoring 0.41 on the public LB. \n\nNonetheless, considering your place in the competition, how do you approach data preprocessing?",
      "votes": null
    },
    {
      "id": "2635725",
      "postDate": "02/04/2024 15:07:42",
      "content": "<p>There are a lot of things that can be done, for example try to create new spectrograms and see how that might affect your cv.<br>\nIf you are stuck, you can read the top solutions from the <br>\nprevious image classification competitions, that will give you some inspiration and new ideas to try.</p>",
      "rawMarkdown": "There are a lot of things that can be done, for example try to create new spectrograms and see how that might affect your cv.\nIf you are stuck, you can read the top solutions from the \nprevious image classification competitions, that will give you some inspiration and new ideas to try.",
      "votes": null
    },
    {
      "id": "2635909",
      "postDate": "02/04/2024 17:24:07",
      "content": "<p>My experience is that the larger the model, the faster it overfit, I tried a few things to try to reduce the overfit, the only one that had a good effect was reducing the batch, but when doing this the training became slower.</p>",
      "rawMarkdown": "My experience is that the larger the model, the faster it overfit, I tried a few things to try to reduce the overfit, the only one that had a good effect was reducing the batch, but when doing this the training became slower.",
      "votes": null
    },
    {
      "id": "2635987",
      "postDate": "02/04/2024 18:18:20",
      "content": "<p>What batch sizes are you using? </p>",
      "rawMarkdown": "What batch sizes are you using?",
      "votes": null
    },
    {
      "id": "2638671",
      "postDate": "02/06/2024 12:30:50",
      "content": "<p>The best result in pythorch model was 16 </p>",
      "rawMarkdown": "The best result in pythorch model was 16",
      "votes": null
    },
    {
      "id": "2638905",
      "postDate": "02/06/2024 14:51:31",
      "content": "<p>Depends also on the optimizer, e.g. Adam is known to be more batch-size over-fit sensitive </p>",
      "rawMarkdown": "Depends also on the optimizer, e.g. Adam is known to be more batch-size over-fit sensitive",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2635421,
      "author_name": "ahmedelfazouan",
      "author_url": "",
      "post_date": "02/04/2024 11:45:56",
      "content": "<p>I think this is the effect of ensembling, you are using only one model for your CV and the average of 5 for your LB.<br>\nThere is a variation ('underfitting') for a smaller architecture which can give a nice boost when ensembled, a bigger architecture tends to overfit very quickly<br>\nIt might be better to work with effnetb0 for now and optimize your hyperparameters or the way you process the data</p>",
      "votes": null,
      "replies": [
        {
          "id": 2635644,
          "author_name": "andreasbis",
          "author_url": "",
          "post_date": "02/04/2024 14:33:57",
          "content": "<p>Indeed, a weighted ensembler of the three models (ResNet34d, EffNetB0, EffNetB1) did better than the inidividual models, scoring 0.41 on the public LB. </p>\n<p>Nonetheless, considering your place in the competition, how do you approach data preprocessing? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2635725,
              "author_name": "ahmedelfazouan",
              "author_url": "",
              "post_date": "02/04/2024 15:07:42",
              "content": "<p>There are a lot of things that can be done, for example try to create new spectrograms and see how that might affect your cv.<br>\nIf you are stuck, you can read the top solutions from the <br>\nprevious image classification competitions, that will give you some inspiration and new ideas to try.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2635909,
      "author_name": "rafaelzimmermann1",
      "author_url": "",
      "post_date": "02/04/2024 17:24:07",
      "content": "<p>My experience is that the larger the model, the faster it overfit, I tried a few things to try to reduce the overfit, the only one that had a good effect was reducing the batch, but when doing this the training became slower.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2635987,
          "author_name": "andreasbis",
          "author_url": "",
          "post_date": "02/04/2024 18:18:20",
          "content": "<p>What batch sizes are you using? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2638671,
              "author_name": "rafaelzimmermann1",
              "author_url": "",
              "post_date": "02/06/2024 12:30:50",
              "content": "<p>The best result in pythorch model was 16 </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2638905,
          "author_name": "danielwicz",
          "author_url": "",
          "post_date": "02/06/2024 14:51:31",
          "content": "<p>Depends also on the optimizer, e.g. Adam is known to be more batch-size over-fit sensitive </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2635260": "Hello kagglers,\n\nAlongside EfficientNetB0 and ResNet34d, I have (successfully) trained [EfficientNetB1](https://www.kaggle.com/code/andreasbis/hms-train-efficientnetb1), using the same batch size of 16. Since most participants use the first two models, I have been experimenting for the previous days with different models and optimizers: EfficientNetB1 (AdamW optimizer), EfficientNetB3 (RMSProp optimizer) and ResNet50 (SGD optimizer, with momentum). From the first epoch, I concluded that larger models, namely EfficientNetB3 and ResNet50 perform worse compared to lighter models, scoring around 1.25 on the Kullback Liebler divergence loss. \n\nAs for EfficientNetB1, despite doing better on the CV than EfficientNetB0 and ResNet34d, having a best score of 0.53 compared to 0.55 for the other two, it did slightly worse on the public LB: 0.46 compared to 0.42 (for EfficientNetB0) and 0.45 (for ResNet34d).\n\nAt this juncture, my approach is to create a weighted ensembler from the three models (ResNet34d, EffNetB0 and EffNetB1), hoping to get a better score than 0.42 on the public LB. \n\nIf you want to experiment with the weights of EfficientNetB1, you can visit [this dataset](https://www.kaggle.com/datasets/andreasbis/efficientnetb1-for-eeg-classification-weights).",
    "2635421": "I think this is the effect of ensembling, you are using only one model for your CV and the average of 5 for your LB.\nThere is a variation ('underfitting') for a smaller architecture which can give a nice boost when ensembled, a bigger architecture tends to overfit very quickly\nIt might be better to work with effnetb0 for now and optimize your hyperparameters or the way you process the data",
    "2635644": "Indeed, a weighted ensembler of the three models (ResNet34d, EffNetB0, EffNetB1) did better than the inidividual models, scoring 0.41 on the public LB. \n\nNonetheless, considering your place in the competition, how do you approach data preprocessing?",
    "2635725": "There are a lot of things that can be done, for example try to create new spectrograms and see how that might affect your cv.\nIf you are stuck, you can read the top solutions from the \nprevious image classification competitions, that will give you some inspiration and new ideas to try.",
    "2635909": "My experience is that the larger the model, the faster it overfit, I tried a few things to try to reduce the overfit, the only one that had a good effect was reducing the batch, but when doing this the training became slower.",
    "2635987": "What batch sizes are you using?",
    "2638671": "The best result in pythorch model was 16",
    "2638905": "Depends also on the optimizer, e.g. Adam is known to be more batch-size over-fit sensitive"
  },
  "source": "meta"
}