{
  "id": 473817,
  "title": "my approach to this competition",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/473817",
  "author_name": "sol lxr",
  "post_date": "2024-02-06T07:03:37.371000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>hey all! this is sol, a ML engineeer recently starting to work on kaggle competition. I had some experience in ML and DL from my time at Edinburgh and ETH, but didn't really get to Kaggle despite hearing all about it. now that I have some free time I felt it's a good time to start working on one of these and see how far I'd go. This post is intended as a summary page of my progress and findings and a todo list. Obviously, I'm still a newbie at this, so any advice or team forming invitations are well welcomed!  now let's start with the todo list. </p>\n<p>stage 1 preparation <br>\n~~1. go through the competition summary page  </p>\n<p>~~2. go through the data, have a general understanding of the structure and sample size. </p>\n<p>~~3. go through the discussion page, note down potential useful resources, problems to be aware of and explore further,  code to be replicated. </p>\n<ol>\n<li>go through the linked paper and youtube video on terminology, and pay special attention to how they classify each pattern. identify potential features and verify them later during eda. </li>\n<li>what are the key papers in EEG modeling? we can start by using research gpt and connected paper on the linked paper. </li>\n<li>see what kind of open source feature selection process is already available. </li>\n<li>can we augment the data or generate synthetic data for this  dataset? </li>\n<li>summarise my findings on miro, and share the link here. </li>\n</ol>\n<p>stage1.5 eda<br>\ndata cleaning , are any of these necessary<br>\nmissing value?<br>\nregularization?<br>\nfeature selection?<br>\nnoise reduction? non max impression, thresholding?<br>\ncorrelation and representation？ what are the long term and statistical features of these eeg?<br>\ntransformation? different window size?</p>\n<p>stage 2 . base model building. <br>\nright off the bet I think we can try transformer based ,  CNN based, maybe ViT too on the DL side. Efficient net seems like a nice model to use. graph network can be considered too, given the graphical nature of eeg sensors. recently there are some development in using large model in EEG, we can try to see if any of those are open source and can be fine-tuned for this purpose…. <br>\nbut remember always start simple. least you spend too much time on uselessly complicated model. so </p>\n<p>baseline:</p>\n<ol>\n<li>conv1d.</li>\n<li>transformer / rnn / lstm </li>\n<li>feature engineering and traditional ML, </li>\n<li>start to try out more recent publications, </li>\n<li>add more weak classifiers . </li>\n</ol>\n<p>stage 3 ensambling<br>\nxgboost, extra tree, lgb etc, </p>\n<p>stage 4 . merge with other teams for better ensambling</p>",
  "messages": [
    {
      "id": 2638255,
      "postDate": "2024-02-06T07:03:37.370Z",
      "content": "<p>hey all! this is sol, a ML engineeer recently starting to work on kaggle competition. I had some experience in ML and DL from my time at Edinburgh and ETH, but didn't really get to Kaggle despite hearing all about it. now that I have some free time I felt it's a good time to start working on one of these and see how far I'd go. This post is intended as a summary page of my progress and findings and a todo list. Obviously, I'm still a newbie at this, so any advice or team forming invitations are well welcomed!  now let's start with the todo list. </p>\n<p>stage 1 preparation <br>\n~~1. go through the competition summary page  </p>\n<p>~~2. go through the data, have a general understanding of the structure and sample size. </p>\n<p>~~3. go through the discussion page, note down potential useful resources, problems to be aware of and explore further,  code to be replicated. </p>\n<ol>\n<li>go through the linked paper and youtube video on terminology, and pay special attention to how they classify each pattern. identify potential features and verify them later during eda. </li>\n<li>what are the key papers in EEG modeling? we can start by using research gpt and connected paper on the linked paper. </li>\n<li>see what kind of open source feature selection process is already available. </li>\n<li>can we augment the data or generate synthetic data for this  dataset? </li>\n<li>summarise my findings on miro, and share the link here. </li>\n</ol>\n<p>stage1.5 eda<br>\ndata cleaning , are any of these necessary<br>\nmissing value?<br>\nregularization?<br>\nfeature selection?<br>\nnoise reduction? non max impression, thresholding?<br>\ncorrelation and representation？ what are the long term and statistical features of these eeg?<br>\ntransformation? different window size?</p>\n<p>stage 2 . base model building. <br>\nright off the bet I think we can try transformer based ,  CNN based, maybe ViT too on the DL side. Efficient net seems like a nice model to use. graph network can be considered too, given the graphical nature of eeg sensors. recently there are some development in using large model in EEG, we can try to see if any of those are open source and can be fine-tuned for this purpose…. <br>\nbut remember always start simple. least you spend too much time on uselessly complicated model. so </p>\n<p>baseline:</p>\n<ol>\n<li>conv1d.</li>\n<li>transformer / rnn / lstm </li>\n<li>feature engineering and traditional ML, </li>\n<li>start to try out more recent publications, </li>\n<li>add more weak classifiers . </li>\n</ol>\n<p>stage 3 ensambling<br>\nxgboost, extra tree, lgb etc, </p>\n<p>stage 4 . merge with other teams for better ensambling</p>",
      "rawMarkdown": "hey all! this is sol, a ML engineeer recently starting to work on kaggle competition. I had some experience in ML and DL from my time at Edinburgh and ETH, but didn't really get to Kaggle despite hearing all about it. now that I have some free time I felt it's a good time to start working on one of these and see how far I'd go. This post is intended as a summary page of my progress and findings and a todo list. Obviously, I'm still a newbie at this, so any advice or team forming invitations are well welcomed!  now let's start with the todo list. \n\nstage 1 preparation \n~~1. go through the competition summary page  \n\n~~2. go through the data, have a general understanding of the structure and sample size. \n\n~~3. go through the discussion page, note down potential useful resources, problems to be aware of and explore further,  code to be replicated. \n4. go through the linked paper and youtube video on terminology, and pay special attention to how they classify each pattern. identify potential features and verify them later during eda. \n5. what are the key papers in EEG modeling? we can start by using research gpt and connected paper on the linked paper. \n6. see what kind of open source feature selection process is already available. \n7. can we augment the data or generate synthetic data for this  dataset? \n6. summarise my findings on miro, and share the link here. \n\nstage1.5 eda\ndata cleaning , are any of these necessary\nmissing value?\nregularization?\nfeature selection?\nnoise reduction? non max impression, thresholding?\ncorrelation and representation？ what are the long term and statistical features of these eeg?\ntransformation? different window size?\n\nstage 2 . base model building. \nright off the bet I think we can try transformer based ,  CNN based, maybe ViT too on the DL side. Efficient net seems like a nice model to use. graph network can be considered too, given the graphical nature of eeg sensors. recently there are some development in using large model in EEG, we can try to see if any of those are open source and can be fine-tuned for this purpose.... \nbut remember always start simple. least you spend too much time on uselessly complicated model. so \n\nbaseline:\n1. conv1d.\n2. transformer / rnn / lstm \n3. feature engineering and traditional ML, \n4. start to try out more recent publications, \n5. add more weak classifiers . \n\nstage 3 ensambling\nxgboost, extra tree, lgb etc, \n\nstage 4 . merge with other teams for better ensambling\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2638255": "hey all! this is sol, a ML engineeer recently starting to work on kaggle competition. I had some experience in ML and DL from my time at Edinburgh and ETH, but didn't really get to Kaggle despite hearing all about it. now that I have some free time I felt it's a good time to start working on one of these and see how far I'd go. This post is intended as a summary page of my progress and findings and a todo list. Obviously, I'm still a newbie at this, so any advice or team forming invitations are well welcomed!  now let's start with the todo list. \n\nstage 1 preparation \n~~1. go through the competition summary page  \n\n~~2. go through the data, have a general understanding of the structure and sample size. \n\n~~3. go through the discussion page, note down potential useful resources, problems to be aware of and explore further,  code to be replicated. \n4. go through the linked paper and youtube video on terminology, and pay special attention to how they classify each pattern. identify potential features and verify them later during eda. \n5. what are the key papers in EEG modeling? we can start by using research gpt and connected paper on the linked paper. \n6. see what kind of open source feature selection process is already available. \n7. can we augment the data or generate synthetic data for this  dataset? \n6. summarise my findings on miro, and share the link here. \n\nstage1.5 eda\ndata cleaning , are any of these necessary\nmissing value?\nregularization?\nfeature selection?\nnoise reduction? non max impression, thresholding?\ncorrelation and representation？ what are the long term and statistical features of these eeg?\ntransformation? different window size?\n\nstage 2 . base model building. \nright off the bet I think we can try transformer based ,  CNN based, maybe ViT too on the DL side. Efficient net seems like a nice model to use. graph network can be considered too, given the graphical nature of eeg sensors. recently there are some development in using large model in EEG, we can try to see if any of those are open source and can be fine-tuned for this purpose.... \nbut remember always start simple. least you spend too much time on uselessly complicated model. so \n\nbaseline:\n1. conv1d.\n2. transformer / rnn / lstm \n3. feature engineering and traditional ML, \n4. start to try out more recent publications, \n5. add more weak classifiers . \n\nstage 3 ensambling\nxgboost, extra tree, lgb etc, \n\nstage 4 . merge with other teams for better ensambling\n"
  }
}