{
  "id": 476184,
  "title": "Looking For Advice For Experimentation",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/476184",
  "author_name": "",
  "post_date": "2024-02-11T12:57:45.158308900Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>This looks like a really fun problem to work on! It's my first real exposure to working with spectrograms, so I'm excited to learn more about how we can get the most information out of them.</p>\n<p>I'd like some advice on how to approach experimentation in competitions like this. I'm grateful that I have a RTX 3060 (12 gb) I can train with locally, but it is a little slow compared to the Kaggle GPUs so I need to be efficient with my training time. <em>My Request:</em> can anyone share how they approach these sorts of problems and perhaps provide guidance on where my strategy (see below) may go awry? I'm sure there are things I'm not considering and I'd love to hear about your advice and general strategy when approaching competitions.</p>\n<p>My current approach is the following: find the smallest model (convnets for now) that seems to fit the data at least reasonably well with no augmentation. Then experiment with preprocessing and augmentation to see what has meaningful impacts on the CV. After identifying which preprocessing an augmentation techniques work well for this model architecture, deepen the model to see if it learns more and anticipate that as the model gets deeper, heavier augmentations may be needed to prevent overfitting. However, the learnings about <em>which</em> augmentations are useful is assumed to translate so I don't have to do as many experiments with large, slow to train models. Thoughts?</p>",
  "messages": [
    {
      "id": "2647263",
      "postDate": "02/11/2024 12:57:45",
      "content": "<p>Hi Everyone,</p>\n<p>This looks like a really fun problem to work on! It's my first real exposure to working with spectrograms, so I'm excited to learn more about how we can get the most information out of them.</p>\n<p>I'd like some advice on how to approach experimentation in competitions like this. I'm grateful that I have a RTX 3060 (12 gb) I can train with locally, but it is a little slow compared to the Kaggle GPUs so I need to be efficient with my training time. <em>My Request:</em> can anyone share how they approach these sorts of problems and perhaps provide guidance on where my strategy (see below) may go awry? I'm sure there are things I'm not considering and I'd love to hear about your advice and general strategy when approaching competitions.</p>\n<p>My current approach is the following: find the smallest model (convnets for now) that seems to fit the data at least reasonably well with no augmentation. Then experiment with preprocessing and augmentation to see what has meaningful impacts on the CV. After identifying which preprocessing an augmentation techniques work well for this model architecture, deepen the model to see if it learns more and anticipate that as the model gets deeper, heavier augmentations may be needed to prevent overfitting. However, the learnings about <em>which</em> augmentations are useful is assumed to translate so I don't have to do as many experiments with large, slow to train models. Thoughts?</p>",
      "rawMarkdown": "Hi Everyone,\n\nThis looks like a really fun problem to work on! It's my first real exposure to working with spectrograms, so I'm excited to learn more about how we can get the most information out of them.\n\nI'd like some advice on how to approach experimentation in competitions like this. I'm grateful that I have a RTX 3060 (12 gb) I can train with locally, but it is a little slow compared to the Kaggle GPUs so I need to be efficient with my training time. *My Request:* can anyone share how they approach these sorts of problems and perhaps provide guidance on where my strategy (see below) may go awry? I'm sure there are things I'm not considering and I'd love to hear about your advice and general strategy when approaching competitions.\n\nMy current approach is the following: find the smallest model (convnets for now) that seems to fit the data at least reasonably well with no augmentation. Then experiment with preprocessing and augmentation to see what has meaningful impacts on the CV. After identifying which preprocessing an augmentation techniques work well for this model architecture, deepen the model to see if it learns more and anticipate that as the model gets deeper, heavier augmentations may be needed to prevent overfitting. However, the learnings about *which* augmentations are useful is assumed to translate so I don't have to do as many experiments with large, slow to train models. Thoughts?",
      "votes": null
    },
    {
      "id": "2651250",
      "postDate": "02/14/2024 03:39:49",
      "content": "<p>I do this for all competitions: sample the data so I have a small training file. Then do train/test split on that file. This way I can look for improvements without waiting a long time for training to complete. Only when I am sure I have a useful idea do I train on the whole dataset. In my case I mostly only use the kaggle allocation. In the unlikely event that I have a potentially winning notebook, I train it on a cloud provider. </p>",
      "rawMarkdown": "I do this for all competitions: sample the data so I have a small training file. Then do train/test split on that file. This way I can look for improvements without waiting a long time for training to complete. Only when I am sure I have a useful idea do I train on the whole dataset. In my case I mostly only use the kaggle allocation. In the unlikely event that I have a potentially winning notebook, I train it on a cloud provider.",
      "votes": null
    },
    {
      "id": "2651253",
      "postDate": "02/14/2024 03:41:20",
      "content": "<p>Generally people in competitions all follow each other, looking for 0.0001% improvement. So I ask myself what approaches might work that have not been tried. </p>",
      "rawMarkdown": "Generally people in competitions all follow each other, looking for 0.0001% improvement. So I ask myself what approaches might work that have not been tried.",
      "votes": null
    },
    {
      "id": "2651296",
      "postDate": "02/14/2024 04:45:16",
      "content": "<p>Understand the data.  Gain some domain knowledge.</p>\n<p>If your chasing cv without understanding the data or having a feel for the domain than I would predict a short and not so glorious career as a data scientist.  Cranking thru models and doing augmentation is a very small part of doing the job in real life.</p>\n<p>Look at every EDA that's been shared.  Create your own and ASK questions of the data as you go.  Read all the documents supplied by the host and spend some time on their web site.  Read the full ACNS guidelines.  </p>\n<p>Then - Follow <a href=\"https://www.kaggle.com/ajenningsfrankston\" target=\"_blank\">andy jennings</a> advice using the knowledge gained from the EDA and host materials.</p>",
      "rawMarkdown": "Understand the data.  Gain some domain knowledge.\n\nIf your chasing cv without understanding the data or having a feel for the domain than I would predict a short and not so glorious career as a data scientist.  Cranking thru models and doing augmentation is a very small part of doing the job in real life.\n\nLook at every EDA that's been shared.  Create your own and ASK questions of the data as you go.  Read all the documents supplied by the host and spend some time on their web site.  Read the full ACNS guidelines.  \n\nThen - Follow [andy jennings](https://www.kaggle.com/ajenningsfrankston) advice using the knowledge gained from the EDA and host materials.",
      "votes": null
    },
    {
      "id": "2655503",
      "postDate": "02/17/2024 00:20:57",
      "content": "<p>Thanks for the advice <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> :). Point well taken on understanding the data and getting some domain knowledge. I do quite a bit of that already, but it can't hurt to do more.</p>",
      "rawMarkdown": "Thanks for the advice @pcjimmmy :). Point well taken on understanding the data and getting some domain knowledge. I do quite a bit of that already, but it can't hurt to do more.",
      "votes": null
    },
    {
      "id": "2655505",
      "postDate": "02/17/2024 00:24:53",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ajenningsfrankston\" target=\"_blank\">@ajenningsfrankston</a>! I'll consider trying this as well </p>",
      "rawMarkdown": "Thanks @ajenningsfrankston! I'll consider trying this as well",
      "votes": null
    },
    {
      "id": "2655575",
      "postDate": "02/17/2024 03:28:52",
      "content": "<p>Also - time is not your friend.  The competition will almost always be over before you know it.   Try to get two or more submissions every day.   Take little steps every day.  </p>",
      "rawMarkdown": "Also - time is not your friend.  The competition will almost always be over before you know it.   Try to get two or more submissions every day.   Take little steps every day.",
      "votes": null
    },
    {
      "id": "2656188",
      "postDate": "02/17/2024 14:25:06",
      "content": "<p>It does go fast. I've been doing eda and getting my model training pipeline set up w/experiment capture. Now I'm getting my inference pipeline put together so I can do just that. Didn't want to make submissions before I had experiment capture though as I'd like to rely more on internal validation than lb score (while still trying for a good lb score of course).</p>",
      "rawMarkdown": "It does go fast. I've been doing eda and getting my model training pipeline set up w/experiment capture. Now I'm getting my inference pipeline put together so I can do just that. Didn't want to make submissions before I had experiment capture though as I'd like to rely more on internal validation than lb score (while still trying for a good lb score of course).",
      "votes": null
    },
    {
      "id": "2656312",
      "postDate": "02/17/2024 15:34:32",
      "content": "<p>Even when focused on doing stuff on local machine and following your path - I still make submissions as early in the 3 months as I can - I often have lots of learning and issues related to the hidden test file(s).  Error trapping might not be a huge piece of this competition but it has been on past ones. </p>\n<p>You don't want to end up with a process you love local that hits you with 'submission errors'.   A submission a day keeps grief away.</p>",
      "rawMarkdown": "Even when focused on doing stuff on local machine and following your path - I still make submissions as early in the 3 months as I can - I often have lots of learning and issues related to the hidden test file(s).  Error trapping might not be a huge piece of this competition but it has been on past ones. \n\nYou don't want to end up with a process you love local that hits you with 'submission errors'.   A submission a day keeps grief away.",
      "votes": null
    },
    {
      "id": "2656374",
      "postDate": "02/17/2024 16:41:14",
      "content": "<p>True, fighting with submission errors can be quite frustrating. Perhaps I'll start submitting sooner next time.</p>",
      "rawMarkdown": "True, fighting with submission errors can be quite frustrating. Perhaps I'll start submitting sooner next time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2651250,
      "author_name": "ajenningsfrankston",
      "author_url": "",
      "post_date": "02/14/2024 03:39:49",
      "content": "<p>I do this for all competitions: sample the data so I have a small training file. Then do train/test split on that file. This way I can look for improvements without waiting a long time for training to complete. Only when I am sure I have a useful idea do I train on the whole dataset. In my case I mostly only use the kaggle allocation. In the unlikely event that I have a potentially winning notebook, I train it on a cloud provider. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2651253,
          "author_name": "ajenningsfrankston",
          "author_url": "",
          "post_date": "02/14/2024 03:41:20",
          "content": "<p>Generally people in competitions all follow each other, looking for 0.0001% improvement. So I ask myself what approaches might work that have not been tried. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2655505,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "02/17/2024 00:24:53",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/ajenningsfrankston\" target=\"_blank\">@ajenningsfrankston</a>! I'll consider trying this as well </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2651296,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/14/2024 04:45:16",
      "content": "<p>Understand the data.  Gain some domain knowledge.</p>\n<p>If your chasing cv without understanding the data or having a feel for the domain than I would predict a short and not so glorious career as a data scientist.  Cranking thru models and doing augmentation is a very small part of doing the job in real life.</p>\n<p>Look at every EDA that's been shared.  Create your own and ASK questions of the data as you go.  Read all the documents supplied by the host and spend some time on their web site.  Read the full ACNS guidelines.  </p>\n<p>Then - Follow <a href=\"https://www.kaggle.com/ajenningsfrankston\" target=\"_blank\">andy jennings</a> advice using the knowledge gained from the EDA and host materials.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2655503,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "02/17/2024 00:20:57",
          "content": "<p>Thanks for the advice <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> :). Point well taken on understanding the data and getting some domain knowledge. I do quite a bit of that already, but it can't hurt to do more.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2655575,
              "author_name": "pcjimmmy",
              "author_url": "",
              "post_date": "02/17/2024 03:28:52",
              "content": "<p>Also - time is not your friend.  The competition will almost always be over before you know it.   Try to get two or more submissions every day.   Take little steps every day.  </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2656188,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "02/17/2024 14:25:06",
                  "content": "<p>It does go fast. I've been doing eda and getting my model training pipeline set up w/experiment capture. Now I'm getting my inference pipeline put together so I can do just that. Didn't want to make submissions before I had experiment capture though as I'd like to rely more on internal validation than lb score (while still trying for a good lb score of course).</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2656312,
                      "author_name": "pcjimmmy",
                      "author_url": "",
                      "post_date": "02/17/2024 15:34:32",
                      "content": "<p>Even when focused on doing stuff on local machine and following your path - I still make submissions as early in the 3 months as I can - I often have lots of learning and issues related to the hidden test file(s).  Error trapping might not be a huge piece of this competition but it has been on past ones. </p>\n<p>You don't want to end up with a process you love local that hits you with 'submission errors'.   A submission a day keeps grief away.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2656374,
                          "author_name": "chemdatafarmer",
                          "author_url": "",
                          "post_date": "02/17/2024 16:41:14",
                          "content": "<p>True, fighting with submission errors can be quite frustrating. Perhaps I'll start submitting sooner next time.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2647263": "Hi Everyone,\n\nThis looks like a really fun problem to work on! It's my first real exposure to working with spectrograms, so I'm excited to learn more about how we can get the most information out of them.\n\nI'd like some advice on how to approach experimentation in competitions like this. I'm grateful that I have a RTX 3060 (12 gb) I can train with locally, but it is a little slow compared to the Kaggle GPUs so I need to be efficient with my training time. *My Request:* can anyone share how they approach these sorts of problems and perhaps provide guidance on where my strategy (see below) may go awry? I'm sure there are things I'm not considering and I'd love to hear about your advice and general strategy when approaching competitions.\n\nMy current approach is the following: find the smallest model (convnets for now) that seems to fit the data at least reasonably well with no augmentation. Then experiment with preprocessing and augmentation to see what has meaningful impacts on the CV. After identifying which preprocessing an augmentation techniques work well for this model architecture, deepen the model to see if it learns more and anticipate that as the model gets deeper, heavier augmentations may be needed to prevent overfitting. However, the learnings about *which* augmentations are useful is assumed to translate so I don't have to do as many experiments with large, slow to train models. Thoughts?",
    "2651250": "I do this for all competitions: sample the data so I have a small training file. Then do train/test split on that file. This way I can look for improvements without waiting a long time for training to complete. Only when I am sure I have a useful idea do I train on the whole dataset. In my case I mostly only use the kaggle allocation. In the unlikely event that I have a potentially winning notebook, I train it on a cloud provider.",
    "2651253": "Generally people in competitions all follow each other, looking for 0.0001% improvement. So I ask myself what approaches might work that have not been tried.",
    "2651296": "Understand the data.  Gain some domain knowledge.\n\nIf your chasing cv without understanding the data or having a feel for the domain than I would predict a short and not so glorious career as a data scientist.  Cranking thru models and doing augmentation is a very small part of doing the job in real life.\n\nLook at every EDA that's been shared.  Create your own and ASK questions of the data as you go.  Read all the documents supplied by the host and spend some time on their web site.  Read the full ACNS guidelines.  \n\nThen - Follow [andy jennings](https://www.kaggle.com/ajenningsfrankston) advice using the knowledge gained from the EDA and host materials.",
    "2655503": "Thanks for the advice @pcjimmmy :). Point well taken on understanding the data and getting some domain knowledge. I do quite a bit of that already, but it can't hurt to do more.",
    "2655505": "Thanks @ajenningsfrankston! I'll consider trying this as well",
    "2655575": "Also - time is not your friend.  The competition will almost always be over before you know it.   Try to get two or more submissions every day.   Take little steps every day.",
    "2656188": "It does go fast. I've been doing eda and getting my model training pipeline set up w/experiment capture. Now I'm getting my inference pipeline put together so I can do just that. Didn't want to make submissions before I had experiment capture though as I'd like to rely more on internal validation than lb score (while still trying for a good lb score of course).",
    "2656312": "Even when focused on doing stuff on local machine and following your path - I still make submissions as early in the 3 months as I can - I often have lots of learning and issues related to the hidden test file(s).  Error trapping might not be a huge piece of this competition but it has been on past ones. \n\nYou don't want to end up with a process you love local that hits you with 'submission errors'.   A submission a day keeps grief away.",
    "2656374": "True, fighting with submission errors can be quite frustrating. Perhaps I'll start submitting sooner next time."
  },
  "source": "meta"
}