{
  "id": 18443,
  "title": "Feature selection / removal of information",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18443",
  "author_name": "",
  "post_date": "2016-01-18T13:03:18.210Z",
  "votes": null,
  "comment_count": 1,
  "views": 655,
  "content": "<p>I am applying a convolutional neural net approach to this challenge. I am currently training, similar to the MxNet tutoral, with all the image sax folders and all 30 images per sax / stack. I would be curious to hear, from those with deep learning expertise and experience, whether it makes theoretical or practical sense to train on a subset of this information (e.g. only a subset of the most relevant sax / cross-cuts, or only part of the 30 seconds). I wonder whether this might improve performance, or whether the architecture of the CNN, if properly tuned, already hones in on the most relevant information and discards superfluous additional information? </p>",
  "messages": [
    {
      "id": "104968",
      "postDate": "01/18/2016 13:03:18",
      "content": "<p>I am applying a convolutional neural net approach to this challenge. I am currently training, similar to the MxNet tutoral, with all the image sax folders and all 30 images per sax / stack. I would be curious to hear, from those with deep learning expertise and experience, whether it makes theoretical or practical sense to train on a subset of this information (e.g. only a subset of the most relevant sax / cross-cuts, or only part of the 30 seconds). I wonder whether this might improve performance, or whether the architecture of the CNN, if properly tuned, already hones in on the most relevant information and discards superfluous additional information? </p>",
      "rawMarkdown": "I am applying a convolutional neural net approach to this challenge. I am currently training, similar to the MxNet tutoral, with all the image sax folders and all 30 images per sax / stack. I would be curious to hear, from those with deep learning expertise and experience, whether it makes theoretical or practical sense to train on a subset of this information (e.g. only a subset of the most relevant sax / cross-cuts, or only part of the 30 seconds). I wonder whether this might improve performance, or whether the architecture of the CNN, if properly tuned, already hones in on the most relevant information and discards superfluous additional information?",
      "votes": null
    },
    {
      "id": "105014",
      "postDate": "01/18/2016 20:40:24",
      "content": "<p>@WD, it very much depends. If you can find the right subset of data to train on, it can be very helpful since it can help reduce overfitting and speed things up.  Like pretty much all machine learning methods, the more features you throw at a net the easier it is to end up overfitting.  Of course, the trick is to find features that have little to no predictive value, and get rid of those while keeping the rest.  Often that's a hard problem in which case one can just throw everything at the net, take the speed hit, and try to control overfitting in other ways (dropout, data augmentation, l2/l1 normalization [aka weight decay], etc).</p>",
      "rawMarkdown": "WD, it very much depends. If you can find the right subset of data to train on, it can be very helpful since it can help reduce overfitting and speed things up.  Like pretty much all machine learning methods, the more features you throw at a net the easier it is to end up overfitting.  Of course, the trick is to find features that have little to no predictive value, and get rid of those while keeping the rest.  Often that's a hard problem in which case one can just throw everything at the net, take the speed hit, and try to control overfitting in other ways (dropout, data augmentation, l2/l1 normalization [aka weight decay], etc).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 105014,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "01/18/2016 20:40:24",
      "content": "<p>@WD, it very much depends. If you can find the right subset of data to train on, it can be very helpful since it can help reduce overfitting and speed things up.  Like pretty much all machine learning methods, the more features you throw at a net the easier it is to end up overfitting.  Of course, the trick is to find features that have little to no predictive value, and get rid of those while keeping the rest.  Often that's a hard problem in which case one can just throw everything at the net, take the speed hit, and try to control overfitting in other ways (dropout, data augmentation, l2/l1 normalization [aka weight decay], etc).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "104968": "I am applying a convolutional neural net approach to this challenge. I am currently training, similar to the MxNet tutoral, with all the image sax folders and all 30 images per sax / stack. I would be curious to hear, from those with deep learning expertise and experience, whether it makes theoretical or practical sense to train on a subset of this information (e.g. only a subset of the most relevant sax / cross-cuts, or only part of the 30 seconds). I wonder whether this might improve performance, or whether the architecture of the CNN, if properly tuned, already hones in on the most relevant information and discards superfluous additional information?",
    "105014": "WD, it very much depends. If you can find the right subset of data to train on, it can be very helpful since it can help reduce overfitting and speed things up.  Like pretty much all machine learning methods, the more features you throw at a net the easier it is to end up overfitting.  Of course, the trick is to find features that have little to no predictive value, and get rid of those while keeping the rest.  Often that's a hard problem in which case one can just throw everything at the net, take the speed hit, and try to control overfitting in other ways (dropout, data augmentation, l2/l1 normalization [aka weight decay], etc)."
  },
  "source": "meta"
}