{
  "id": 191042,
  "title": "How important is augmentation when you’ve got a massive dataset?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/191042",
  "author_name": "",
  "post_date": "2020-10-14T11:09:16.480151700Z",
  "votes": 17,
  "comment_count": 9,
  "views": 0,
  "content": "<p>So far, I haven’t seen much of an impact from augmentation. I’ve seen very slight improvements (0.2 - 0.5), but overall nothing that has made the effort that worthwhile. </p>\n<p>Perhaps, given the huge dataset, there is limited need for augmentation if we can instead train on different images?</p>\n<p>Or maybe I’ve just not happened on the right approach.</p>\n<p>So far I’ve played around with:</p>\n<ul>\n<li>Cutout</li>\n<li>Blur</li>\n<li>Random ego_center placement</li>\n</ul>",
  "messages": [
    {
      "id": "1049356",
      "postDate": "10/14/2020 11:09:16",
      "content": "<p>So far, I haven’t seen much of an impact from augmentation. I’ve seen very slight improvements (0.2 - 0.5), but overall nothing that has made the effort that worthwhile. </p>\n<p>Perhaps, given the huge dataset, there is limited need for augmentation if we can instead train on different images?</p>\n<p>Or maybe I’ve just not happened on the right approach.</p>\n<p>So far I’ve played around with:</p>\n<ul>\n<li>Cutout</li>\n<li>Blur</li>\n<li>Random ego_center placement</li>\n</ul>",
      "rawMarkdown": "So far, I haven’t seen much of an impact from augmentation. I’ve seen very slight improvements (0.2 - 0.5), but overall nothing that has made the effort that worthwhile. \n\nPerhaps, given the huge dataset, there is limited need for augmentation if we can instead train on different images?\n \nOr maybe I’ve just not happened on the right approach.\n\nSo far I’ve played around with:\n\n- Cutout\n- Blur\n- Random ego_center placement",
      "votes": null
    },
    {
      "id": "1049371",
      "postDate": "10/14/2020 11:24:53",
      "content": "<p>Image enhancement is a way of data expansion, in other words it is not so important when there is enough data.</p>",
      "rawMarkdown": "Image enhancement is a way of data expansion, in other words it is not so important when there is enough data.",
      "votes": null
    },
    {
      "id": "1049433",
      "postDate": "10/14/2020 12:15:00",
      "content": "<p>I dont see a need for data augmentation, due to the fact that we have a huge amount of data and there is no such thing as imbalanced dataset. </p>\n<p>The only need I can see using it for, is for training on kaggle with multiple runs and keeping shuffle on. If you dont slice you dataset there is the possibility where you show the model multiple times the same data, augmentation could change the data and thus helps the model learn more robustly.</p>",
      "rawMarkdown": "I dont see a need for data augmentation, due to the fact that we have a huge amount of data and there is no such thing as imbalanced dataset. \n\nThe only need I can see using it for, is for training on kaggle with multiple runs and keeping shuffle on. If you dont slice you dataset there is the possibility where you show the model multiple times the same data, augmentation could change the data and thus helps the model learn more robustly.",
      "votes": null
    },
    {
      "id": "1049458",
      "postDate": "10/14/2020 12:37:34",
      "content": "<p>I agree. </p>\n<p>What constitutes 'enough' data in this case (or in any case!), though. I think that data scientists will always find a way of using all the data they can get. We're greedy in that regard :)</p>\n<p>I think that within the confines of this competition, the cases where augmentation could theoretically add value are:</p>\n<ul>\n<li><p>Speed. If it's possible to rasterize and save a subset of train_full/train.zarr to disk then effective augmentation would allow you to iterate through ideas more efficiently.</p></li>\n<li><p>Resources. If you have access to exceptionally powerful machines with large numbers of cores /GPUs you can probably make it through train_full.zarr in a reasonable amount of time. It would be useful to squeeze more out of the data if you happen to be in that position.</p></li>\n</ul>",
      "rawMarkdown": "I agree. \n\nWhat constitutes 'enough' data in this case (or in any case!), though. I think that data scientists will always find a way of using all the data they can get. We're greedy in that regard :)\n\nI think that within the confines of this competition, the cases where augmentation could theoretically add value are:\n\n- Speed. If it's possible to rasterize and save a subset of train_full/train.zarr to disk then effective augmentation would allow you to iterate through ideas more efficiently.\n\n- Resources. If you have access to exceptionally powerful machines with large numbers of cores /GPUs you can probably make it through train_full.zarr in a reasonable amount of time. It would be useful to squeeze more out of the data if you happen to be in that position.",
      "votes": null
    },
    {
      "id": "1050140",
      "postDate": "10/15/2020 05:40:20",
      "content": "<p>I think the difficult part is creating augmentations that are actually meaningful and carry over to the test set and don't destroy information. </p>\n<p>For example, random ego_center placement might help make the model generalize to scenarios where the ego_center is not at the default location but we know that it will always be there so it doesn't really simulate a useful scenario. </p>\n<p>Rotations and flips are strong in other image tasks because our image dataset might be deficient in samples where an object is at a certain angle, but with how much data we have here and how standardized it is I think it is difficult  to prove any benefit beyond just training on more of the data we have</p>",
      "rawMarkdown": "I think the difficult part is creating augmentations that are actually meaningful and carry over to the test set and don't destroy information. \n\nFor example, random ego_center placement might help make the model generalize to scenarios where the ego_center is not at the default location but we know that it will always be there so it doesn't really simulate a useful scenario. \n\nRotations and flips are strong in other image tasks because our image dataset might be deficient in samples where an object is at a certain angle, but with how much data we have here and how standardized it is I think it is difficult  to prove any benefit beyond just training on more of the data we have",
      "votes": null
    },
    {
      "id": "1050160",
      "postDate": "10/15/2020 06:01:39",
      "content": "<p>That's also what I thought…since we are predicting motion instead of recognizing the agents, I did not even consider this.</p>\n<p>Also we are predicting a pretty stable scenario - no random pedestrians/cyclists crossing and traffic is quite civilized. Maybe if our task included more extreme situations (eg, a deer suddenly crossing the road, a car quickly crossing on red), augmentation could be useful (especially moving ego center), but even then I don't think its the most relevant technique.</p>",
      "rawMarkdown": "That's also what I thought...since we are predicting motion instead of recognizing the agents, I did not even consider this.\n\nAlso we are predicting a pretty stable scenario - no random pedestrians/cyclists crossing and traffic is quite civilized. Maybe if our task included more extreme situations (eg, a deer suddenly crossing the road, a car quickly crossing on red), augmentation could be useful (especially moving ego center), but even then I don't think its the most relevant technique.",
      "votes": null
    },
    {
      "id": "1050165",
      "postDate": "10/15/2020 06:06:39",
      "content": "<p>I agree that we have a huge amount of data and there is no need for general augmentation approaches. As <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> suggest, we should experiment and find relevant augmentation kernels that can add meaningful information to existing data.</p>",
      "rawMarkdown": "I agree that we have a huge amount of data and there is no need for general augmentation approaches. As @ryches suggest, we should experiment and find relevant augmentation kernels that can add meaningful information to existing data.",
      "votes": null
    },
    {
      "id": "1050541",
      "postDate": "10/15/2020 14:17:33",
      "content": "<p>Assume that we've used all the data - or whatever subset of data we can manage - and are looping through it for a second time (I think we've got to assume this at a minimum, otherwise no augmentation is necessary). </p>\n<p>I've been thinking about the distinction between </p>\n<ol>\n<li>augmentation =&gt; creates additional information =&gt; generalises better on test set</li>\n<li>augmentation =&gt; makes it more difficult to predict on the training set =&gt; prevents overfitting =&gt; generalises better on the test set</li>\n</ol>\n<p>1: seems difficult in this scenario (your <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188368\" target=\"_blank\">discussion topic</a> outlines this really well, <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>). It would be interesting to play around with adding agents/traffic light sequences and changing the ground truth accordingly, but difficult and time consuming to pull this off, I'd guess!</p>\n<p>2: I think is possible with items like cutout/blur/noise. But also possible by selecting a simple model architecture, being careful about how many epochs you train for, etc.</p>",
      "rawMarkdown": "Assume that we've used all the data - or whatever subset of data we can manage - and are looping through it for a second time (I think we've got to assume this at a minimum, otherwise no augmentation is necessary). \n\nI've been thinking about the distinction between \n\n1. augmentation => creates additional information => generalises better on test set\n2. augmentation => makes it more difficult to predict on the training set => prevents overfitting => generalises better on the test set\n\n\n1: seems difficult in this scenario (your [discussion topic](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188368) outlines this really well, @ryches). It would be interesting to play around with adding agents/traffic light sequences and changing the ground truth accordingly, but difficult and time consuming to pull this off, I'd guess!\n\n2: I think is possible with items like cutout/blur/noise. But also possible by selecting a simple model architecture, being careful about how many epochs you train for, etc.",
      "votes": null
    },
    {
      "id": "1050738",
      "postDate": "10/15/2020 17:17:15",
      "content": "<p>What about Test Time Augmentation?</p>\n<p>Does it add value in this problem?</p>",
      "rawMarkdown": "What about Test Time Augmentation?\n\nDoes it add value in this problem?",
      "votes": null
    },
    {
      "id": "1050796",
      "postDate": "10/15/2020 18:27:59",
      "content": "<p>I haven't got to that point yet.</p>",
      "rawMarkdown": "I haven't got to that point yet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1049371,
      "author_name": "zzy990106",
      "author_url": "",
      "post_date": "10/14/2020 11:24:53",
      "content": "<p>Image enhancement is a way of data expansion, in other words it is not so important when there is enough data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1049458,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "10/14/2020 12:37:34",
          "content": "<p>I agree. </p>\n<p>What constitutes 'enough' data in this case (or in any case!), though. I think that data scientists will always find a way of using all the data they can get. We're greedy in that regard :)</p>\n<p>I think that within the confines of this competition, the cases where augmentation could theoretically add value are:</p>\n<ul>\n<li><p>Speed. If it's possible to rasterize and save a subset of train_full/train.zarr to disk then effective augmentation would allow you to iterate through ideas more efficiently.</p></li>\n<li><p>Resources. If you have access to exceptionally powerful machines with large numbers of cores /GPUs you can probably make it through train_full.zarr in a reasonable amount of time. It would be useful to squeeze more out of the data if you happen to be in that position.</p></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1049433,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "10/14/2020 12:15:00",
      "content": "<p>I dont see a need for data augmentation, due to the fact that we have a huge amount of data and there is no such thing as imbalanced dataset. </p>\n<p>The only need I can see using it for, is for training on kaggle with multiple runs and keeping shuffle on. If you dont slice you dataset there is the possibility where you show the model multiple times the same data, augmentation could change the data and thus helps the model learn more robustly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1050140,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "10/15/2020 05:40:20",
      "content": "<p>I think the difficult part is creating augmentations that are actually meaningful and carry over to the test set and don't destroy information. </p>\n<p>For example, random ego_center placement might help make the model generalize to scenarios where the ego_center is not at the default location but we know that it will always be there so it doesn't really simulate a useful scenario. </p>\n<p>Rotations and flips are strong in other image tasks because our image dataset might be deficient in samples where an object is at a certain angle, but with how much data we have here and how standardized it is I think it is difficult  to prove any benefit beyond just training on more of the data we have</p>",
      "votes": null,
      "replies": [
        {
          "id": 1050160,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "10/15/2020 06:01:39",
          "content": "<p>That's also what I thought…since we are predicting motion instead of recognizing the agents, I did not even consider this.</p>\n<p>Also we are predicting a pretty stable scenario - no random pedestrians/cyclists crossing and traffic is quite civilized. Maybe if our task included more extreme situations (eg, a deer suddenly crossing the road, a car quickly crossing on red), augmentation could be useful (especially moving ego center), but even then I don't think its the most relevant technique.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1050541,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "10/15/2020 14:17:33",
          "content": "<p>Assume that we've used all the data - or whatever subset of data we can manage - and are looping through it for a second time (I think we've got to assume this at a minimum, otherwise no augmentation is necessary). </p>\n<p>I've been thinking about the distinction between </p>\n<ol>\n<li>augmentation =&gt; creates additional information =&gt; generalises better on test set</li>\n<li>augmentation =&gt; makes it more difficult to predict on the training set =&gt; prevents overfitting =&gt; generalises better on the test set</li>\n</ol>\n<p>1: seems difficult in this scenario (your <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188368\" target=\"_blank\">discussion topic</a> outlines this really well, <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>). It would be interesting to play around with adding agents/traffic light sequences and changing the ground truth accordingly, but difficult and time consuming to pull this off, I'd guess!</p>\n<p>2: I think is possible with items like cutout/blur/noise. But also possible by selecting a simple model architecture, being careful about how many epochs you train for, etc.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1050165,
      "author_name": "mineshjethva",
      "author_url": "",
      "post_date": "10/15/2020 06:06:39",
      "content": "<p>I agree that we have a huge amount of data and there is no need for general augmentation approaches. As <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> suggest, we should experiment and find relevant augmentation kernels that can add meaningful information to existing data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1050738,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "10/15/2020 17:17:15",
      "content": "<p>What about Test Time Augmentation?</p>\n<p>Does it add value in this problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1050796,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "10/15/2020 18:27:59",
          "content": "<p>I haven't got to that point yet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1049356": "So far, I haven’t seen much of an impact from augmentation. I’ve seen very slight improvements (0.2 - 0.5), but overall nothing that has made the effort that worthwhile. \n\nPerhaps, given the huge dataset, there is limited need for augmentation if we can instead train on different images?\n \nOr maybe I’ve just not happened on the right approach.\n\nSo far I’ve played around with:\n\n- Cutout\n- Blur\n- Random ego_center placement",
    "1049371": "Image enhancement is a way of data expansion, in other words it is not so important when there is enough data.",
    "1049433": "I dont see a need for data augmentation, due to the fact that we have a huge amount of data and there is no such thing as imbalanced dataset. \n\nThe only need I can see using it for, is for training on kaggle with multiple runs and keeping shuffle on. If you dont slice you dataset there is the possibility where you show the model multiple times the same data, augmentation could change the data and thus helps the model learn more robustly.",
    "1049458": "I agree. \n\nWhat constitutes 'enough' data in this case (or in any case!), though. I think that data scientists will always find a way of using all the data they can get. We're greedy in that regard :)\n\nI think that within the confines of this competition, the cases where augmentation could theoretically add value are:\n\n- Speed. If it's possible to rasterize and save a subset of train_full/train.zarr to disk then effective augmentation would allow you to iterate through ideas more efficiently.\n\n- Resources. If you have access to exceptionally powerful machines with large numbers of cores /GPUs you can probably make it through train_full.zarr in a reasonable amount of time. It would be useful to squeeze more out of the data if you happen to be in that position.",
    "1050140": "I think the difficult part is creating augmentations that are actually meaningful and carry over to the test set and don't destroy information. \n\nFor example, random ego_center placement might help make the model generalize to scenarios where the ego_center is not at the default location but we know that it will always be there so it doesn't really simulate a useful scenario. \n\nRotations and flips are strong in other image tasks because our image dataset might be deficient in samples where an object is at a certain angle, but with how much data we have here and how standardized it is I think it is difficult  to prove any benefit beyond just training on more of the data we have",
    "1050160": "That's also what I thought...since we are predicting motion instead of recognizing the agents, I did not even consider this.\n\nAlso we are predicting a pretty stable scenario - no random pedestrians/cyclists crossing and traffic is quite civilized. Maybe if our task included more extreme situations (eg, a deer suddenly crossing the road, a car quickly crossing on red), augmentation could be useful (especially moving ego center), but even then I don't think its the most relevant technique.",
    "1050165": "I agree that we have a huge amount of data and there is no need for general augmentation approaches. As @ryches suggest, we should experiment and find relevant augmentation kernels that can add meaningful information to existing data.",
    "1050541": "Assume that we've used all the data - or whatever subset of data we can manage - and are looping through it for a second time (I think we've got to assume this at a minimum, otherwise no augmentation is necessary). \n\nI've been thinking about the distinction between \n\n1. augmentation => creates additional information => generalises better on test set\n2. augmentation => makes it more difficult to predict on the training set => prevents overfitting => generalises better on the test set\n\n\n1: seems difficult in this scenario (your [discussion topic](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/188368) outlines this really well, @ryches). It would be interesting to play around with adding agents/traffic light sequences and changing the ground truth accordingly, but difficult and time consuming to pull this off, I'd guess!\n\n2: I think is possible with items like cutout/blur/noise. But also possible by selecting a simple model architecture, being careful about how many epochs you train for, etc.",
    "1050738": "What about Test Time Augmentation?\n\nDoes it add value in this problem?",
    "1050796": "I haven't got to that point yet."
  },
  "source": "meta"
}