{
  "id": 481436,
  "title": "Augmentation leading to lower LB",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/481436",
  "author_name": "",
  "post_date": "2024-03-03T16:33:19.079075100Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As of my experimentation any popular augmentations didn't give me any major boost as of now, could it indicate that the public test data is similar to the training data? if that's the case then it would mean that we are having am overfitting problem, just my speculation.</p>",
  "messages": [
    {
      "id": "2679618",
      "postDate": "03/03/2024 16:33:19",
      "content": "<p>As of my experimentation any popular augmentations didn't give me any major boost as of now, could it indicate that the public test data is similar to the training data? if that's the case then it would mean that we are having am overfitting problem, just my speculation.</p>",
      "rawMarkdown": "As of my experimentation any popular augmentations didn't give me any major boost as of now, could it indicate that the public test data is similar to the training data? if that's the case then it would mean that we are having am overfitting problem, just my speculation.",
      "votes": null
    },
    {
      "id": "2679766",
      "postDate": "03/03/2024 18:15:28",
      "content": "<p>This is purely my own speculation based on the kinds of augmentations that have and have not worked for me. Since we're not predicting a single class, but rather the discrete probability distribution of expert opinions, it's possible that if we modify the image in the wrong way we might change the apparent distribution and therefore the KLDivergence. I believe that's why mixup has done so well for some folks, since it modifies the labels along with the image. So, we likely need to focus on augmentations that wouldn't change the distribution of expert opinions, or find ways to express that difference of opinion in the labels.</p>",
      "rawMarkdown": "This is purely my own speculation based on the kinds of augmentations that have and have not worked for me. Since we're not predicting a single class, but rather the discrete probability distribution of expert opinions, it's possible that if we modify the image in the wrong way we might change the apparent distribution and therefore the KLDivergence. I believe that's why mixup has done so well for some folks, since it modifies the labels along with the image. So, we likely need to focus on augmentations that wouldn't change the distribution of expert opinions, or find ways to express that difference of opinion in the labels.",
      "votes": null
    },
    {
      "id": "2679779",
      "postDate": "03/03/2024 18:20:11",
      "content": "<p>I'll comment on the EEG signals, even though you may be talking about spectrograms.</p>\n<p>A big problem is that augmentations like time warping, scaling, shifting, etc. would change the actual classification. If EEG signals from a GRDA example are scaled up more in one hemisphere, the class would change to LRDA. Same with shifting since 'lateralized' can also come from asynchrony. Time warping is bad too since the local frequency matters. For instance, if an RDA example is time warped, especially if dilated in a trough or contracted in a peak, the RDA would start to look more like PD.</p>\n<p>My point is that, unlike typical image classification tasks, augmentations of the data are less likely to leave the true class labels alone, which is essential for augmentations to be of any value.</p>\n<p>I think adding noise can also be bad, but for a different reason. Based on the nature of both spectrograms and EEG signals, it seems like the global perspective of a deep model is more focused on quantifying local features rather than stitching them together, at least relative to other classification tasks. For example, knowing how many times delta activity was found in dilated filters is more important than having very strong delta activity only a couple times. Another point: the time series data, though discrete, represents continuous data. So, it's possible that adding even a little bit of noise may erase too much of the essential, local information.</p>\n<p>I hope this adds some perspective. I'm not completely ruling out augmentations, but I think special care should be taken with this particular classification problem.</p>\n<p>Almost forgot! To your general concerns about overfitting, I've found dropout works well.</p>",
      "rawMarkdown": "I'll comment on the EEG signals, even though you may be talking about spectrograms.\n\nA big problem is that augmentations like time warping, scaling, shifting, etc. would change the actual classification. If EEG signals from a GRDA example are scaled up more in one hemisphere, the class would change to LRDA. Same with shifting since 'lateralized' can also come from asynchrony. Time warping is bad too since the local frequency matters. For instance, if an RDA example is time warped, especially if dilated in a trough or contracted in a peak, the RDA would start to look more like PD.\n\nMy point is that, unlike typical image classification tasks, augmentations of the data are less likely to leave the true class labels alone, which is essential for augmentations to be of any value.\n\nI think adding noise can also be bad, but for a different reason. Based on the nature of both spectrograms and EEG signals, it seems like the global perspective of a deep model is more focused on quantifying local features rather than stitching them together, at least relative to other classification tasks. For example, knowing how many times delta activity was found in dilated filters is more important than having very strong delta activity only a couple times. Another point: the time series data, though discrete, represents continuous data. So, it's possible that adding even a little bit of noise may erase too much of the essential, local information.\n\nI hope this adds some perspective. I'm not completely ruling out augmentations, but I think special care should be taken with this particular classification problem.\n\nAlmost forgot! To your general concerns about overfitting, I've found dropout works well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2679766,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "03/03/2024 18:15:28",
      "content": "<p>This is purely my own speculation based on the kinds of augmentations that have and have not worked for me. Since we're not predicting a single class, but rather the discrete probability distribution of expert opinions, it's possible that if we modify the image in the wrong way we might change the apparent distribution and therefore the KLDivergence. I believe that's why mixup has done so well for some folks, since it modifies the labels along with the image. So, we likely need to focus on augmentations that wouldn't change the distribution of expert opinions, or find ways to express that difference of opinion in the labels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2679779,
      "author_name": "tylerfeemster",
      "author_url": "",
      "post_date": "03/03/2024 18:20:11",
      "content": "<p>I'll comment on the EEG signals, even though you may be talking about spectrograms.</p>\n<p>A big problem is that augmentations like time warping, scaling, shifting, etc. would change the actual classification. If EEG signals from a GRDA example are scaled up more in one hemisphere, the class would change to LRDA. Same with shifting since 'lateralized' can also come from asynchrony. Time warping is bad too since the local frequency matters. For instance, if an RDA example is time warped, especially if dilated in a trough or contracted in a peak, the RDA would start to look more like PD.</p>\n<p>My point is that, unlike typical image classification tasks, augmentations of the data are less likely to leave the true class labels alone, which is essential for augmentations to be of any value.</p>\n<p>I think adding noise can also be bad, but for a different reason. Based on the nature of both spectrograms and EEG signals, it seems like the global perspective of a deep model is more focused on quantifying local features rather than stitching them together, at least relative to other classification tasks. For example, knowing how many times delta activity was found in dilated filters is more important than having very strong delta activity only a couple times. Another point: the time series data, though discrete, represents continuous data. So, it's possible that adding even a little bit of noise may erase too much of the essential, local information.</p>\n<p>I hope this adds some perspective. I'm not completely ruling out augmentations, but I think special care should be taken with this particular classification problem.</p>\n<p>Almost forgot! To your general concerns about overfitting, I've found dropout works well.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2679618": "As of my experimentation any popular augmentations didn't give me any major boost as of now, could it indicate that the public test data is similar to the training data? if that's the case then it would mean that we are having am overfitting problem, just my speculation.",
    "2679766": "This is purely my own speculation based on the kinds of augmentations that have and have not worked for me. Since we're not predicting a single class, but rather the discrete probability distribution of expert opinions, it's possible that if we modify the image in the wrong way we might change the apparent distribution and therefore the KLDivergence. I believe that's why mixup has done so well for some folks, since it modifies the labels along with the image. So, we likely need to focus on augmentations that wouldn't change the distribution of expert opinions, or find ways to express that difference of opinion in the labels.",
    "2679779": "I'll comment on the EEG signals, even though you may be talking about spectrograms.\n\nA big problem is that augmentations like time warping, scaling, shifting, etc. would change the actual classification. If EEG signals from a GRDA example are scaled up more in one hemisphere, the class would change to LRDA. Same with shifting since 'lateralized' can also come from asynchrony. Time warping is bad too since the local frequency matters. For instance, if an RDA example is time warped, especially if dilated in a trough or contracted in a peak, the RDA would start to look more like PD.\n\nMy point is that, unlike typical image classification tasks, augmentations of the data are less likely to leave the true class labels alone, which is essential for augmentations to be of any value.\n\nI think adding noise can also be bad, but for a different reason. Based on the nature of both spectrograms and EEG signals, it seems like the global perspective of a deep model is more focused on quantifying local features rather than stitching them together, at least relative to other classification tasks. For example, knowing how many times delta activity was found in dilated filters is more important than having very strong delta activity only a couple times. Another point: the time series data, though discrete, represents continuous data. So, it's possible that adding even a little bit of noise may erase too much of the essential, local information.\n\nI hope this adds some perspective. I'm not completely ruling out augmentations, but I think special care should be taken with this particular classification problem.\n\nAlmost forgot! To your general concerns about overfitting, I've found dropout works well."
  },
  "source": "meta"
}