{
  "id": 13177,
  "title": "recommended reading for neural network architecture selection?",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/13177",
  "author_name": "",
  "post_date": "2015-04-01T05:24:49.780Z",
  "votes": 8,
  "comment_count": 8,
  "views": 4123,
  "content": "<p>I would really like to get a better understanding of how to go about selecting and optimizing neural network architecture - or even whether there is much value to be gained from being smart about choosing the network architecture. So far I've found http://cs231n.github.io/ helpful as a basic introduction. However, I would like a better understanding of how to go about selecting network architectures that is more sophisticated that &quot;we picked some parameters and they seem to work OK&quot;.</p>\n<p>How many layers? How many filters for each convolutional layer? What should the dimensions of the conv filters be?</p>\n\n<p>Any reading recommendations?</p>",
  "messages": [
    {
      "id": "69263",
      "postDate": "04/01/2015 05:24:49",
      "content": "<p>I would really like to get a better understanding of how to go about selecting and optimizing neural network architecture - or even whether there is much value to be gained from being smart about choosing the network architecture. So far I've found http://cs231n.github.io/ helpful as a basic introduction. However, I would like a better understanding of how to go about selecting network architectures that is more sophisticated that &quot;we picked some parameters and they seem to work OK&quot;.</p>\n<p>How many layers? How many filters for each convolutional layer? What should the dimensions of the conv filters be?</p>\n\n<p>Any reading recommendations?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69405",
      "postDate": "04/02/2015 05:02:17",
      "content": "<p>Hi! I think you should read Sander Dieleman's post about his winning solution for Galaxy Zoo [1] and, obviously, Alex Krizhevsky's paper on Imagenet 2012 [2]. Try to understand why they did what they did. It should give you a better understanding of what you can and cannot do in order to conquer overfitting.</p>\n<p>[1]&nbsp;http://benanne.github.io/2014/04/05/galaxy-zoo.html</p>\n<p>[2]&nbsp;http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional</p>\n<p>Cheers!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69451",
      "postDate": "04/02/2015 17:44:37",
      "content": "<p>[quote=Ionel Hosu;69405]</p>\n<p>Hi! I think you should read Sander Dieleman's post about his winning solution for Galaxy Zoo [1] and, obviously, Alex Krizhevsky's paper on Imagenet 2012 [2]. Try to understand why they did what they did. It should give you a better understanding of what you can and cannot do in order to conquer overfitting.</p>\n<p>[1]&nbsp;http://benanne.github.io/2014/04/05/galaxy-zoo.html</p>\n<p>[2]&nbsp;http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional</p>\n<p>Cheers!</p>\n<p>[/quote]</p>\n<p>Thanks for the plug :) I also have a post about the national data science bowl now, with some info about the architectures we used for that one:&nbsp;http://benanne.github.io/2015/03/17/plankton.html</p>\n\n<p>[quote=small yellow duck;69263]</p>\n<p>However, I would like a better understanding of how to go about selecting network architectures that is more sophisticated that &quot;we picked some parameters and they seem to work OK&quot;.</p>\n<p>[/quote]</p>\n<p>There are a bunch of heuristics, but many times it really just boils down to trying every possibility and using what works. You also learn a lot from that experience and often you can derive new heuristics from it, to speed up architecture selection in the future.&nbsp;But really, the key ingredient is to try everything and see what sticks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69458",
      "postDate": "04/02/2015 19:25:42",
      "content": "<p>@sedielem It is interesting that deep learning is touted as the end-to-end learning and there is huge hype surrounding it. I cannot deny that you almost cannot beat convnets on image and video recognition&nbsp;tasks these days, but, I feel like at the end of the day, you are doing the huge heavy weight lifting of the learning by exploring all the possible heuristics, data augmentation, network architecture, etc... and it begs the question if this is what end-to-end learning is supposed to be in the future.</p>\n<p>Btw, great work on the bowl. I was so surprised to see that you guys even attempted to learn the best data augmentation parameters! Awesome job!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69460",
      "postDate": "04/02/2015 19:38:49",
      "content": "<p>What can I say, there is no such thing as a free lunch ;)</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69558",
      "postDate": "04/03/2015 20:16:04",
      "content": "<p>@sedielem, I saw your back and forth with Fred on how to make theano work with torque back in 2013. I am assuming that you guys have a torque cluster. Can I ask you a question?</p>\n<p>If you have two computer with GPUs where&nbsp;one runs during the day time and not at night and the other during the night, is there anyway to put&nbsp;these two computers in a cluster and transparently run a CNN code that may run more than a day or so without interruption?&nbsp;</p>\n<p>In other words, once you&nbsp;start to train using the GPU of one of the machines in the cluster, is there any way to make that same training job to switch context to other available GPUs on other machines in the cluster without interruption of the training job?</p>\n<p>Note: This is of course with out abandoning Theano and co and resorting to write the GPU kernels and the internal details yourself.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69613",
      "postDate": "04/04/2015 09:24:20",
      "content": "<p>I don't think it's possible to switch GPUs in the middle of a run unless you actually abort it, start it on the other GPU and load up the parameters from a backup.</p>\n<p>We no longer use torque for our GPU machines, but back then all of them had two identical GPUs each, so we just set them to run in exclusive mode (only one compute process per GPU) and that was sufficient. I don't have any experience with the GPU-specific features of torque, we didn't use those.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69628",
      "postDate": "04/04/2015 15:14:16",
      "content": "<p>Thanks for the reply. That's what I thought, just wanted to confirm .</p>\n<p>Did you switch to another cluster package or you just stopped using clusters and just run jobs manually specifying machines and GPUs? We are trying to asses if it is worth putting the effort to setup a cluster vs, just manually scheduling jobs in each machine specifying GPUs manually. From your experience, what's the benefit of setting up a cluster if availability is not transparent?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69644",
      "postDate": "04/04/2015 18:11:20",
      "content": "<p>Our little GPU farm has become a little too heterogeneous for our old approach (4 different types of GPUs now), and with long running experiments it's just more convenient to have them running in screen/tmux so you can check up on them easily. So right now we just have a google drive spreadsheet where everyone indicates which GPUs they're using. Pretty low-tech but it works OK for now. If we end up expanding it further this may become unworkable though.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 69405,
      "author_name": "ionelhosu",
      "author_url": "",
      "post_date": "04/02/2015 05:02:17",
      "content": "<p>Hi! I think you should read Sander Dieleman's post about his winning solution for Galaxy Zoo [1] and, obviously, Alex Krizhevsky's paper on Imagenet 2012 [2]. Try to understand why they did what they did. It should give you a better understanding of what you can and cannot do in order to conquer overfitting.</p>\n<p>[1]&nbsp;http://benanne.github.io/2014/04/05/galaxy-zoo.html</p>\n<p>[2]&nbsp;http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional</p>\n<p>Cheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69451,
      "author_name": "sedielem",
      "author_url": "",
      "post_date": "04/02/2015 17:44:37",
      "content": "<p>[quote=Ionel Hosu;69405]</p>\n<p>Hi! I think you should read Sander Dieleman's post about his winning solution for Galaxy Zoo [1] and, obviously, Alex Krizhevsky's paper on Imagenet 2012 [2]. Try to understand why they did what they did. It should give you a better understanding of what you can and cannot do in order to conquer overfitting.</p>\n<p>[1]&nbsp;http://benanne.github.io/2014/04/05/galaxy-zoo.html</p>\n<p>[2]&nbsp;http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional</p>\n<p>Cheers!</p>\n<p>[/quote]</p>\n<p>Thanks for the plug :) I also have a post about the national data science bowl now, with some info about the architectures we used for that one:&nbsp;http://benanne.github.io/2015/03/17/plankton.html</p>\n\n<p>[quote=small yellow duck;69263]</p>\n<p>However, I would like a better understanding of how to go about selecting network architectures that is more sophisticated that &quot;we picked some parameters and they seem to work OK&quot;.</p>\n<p>[/quote]</p>\n<p>There are a bunch of heuristics, but many times it really just boils down to trying every possibility and using what works. You also learn a lot from that experience and often you can derive new heuristics from it, to speed up architecture selection in the future.&nbsp;But really, the key ingredient is to try everything and see what sticks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69458,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "04/02/2015 19:25:42",
      "content": "<p>@sedielem It is interesting that deep learning is touted as the end-to-end learning and there is huge hype surrounding it. I cannot deny that you almost cannot beat convnets on image and video recognition&nbsp;tasks these days, but, I feel like at the end of the day, you are doing the huge heavy weight lifting of the learning by exploring all the possible heuristics, data augmentation, network architecture, etc... and it begs the question if this is what end-to-end learning is supposed to be in the future.</p>\n<p>Btw, great work on the bowl. I was so surprised to see that you guys even attempted to learn the best data augmentation parameters! Awesome job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69460,
      "author_name": "sedielem",
      "author_url": "",
      "post_date": "04/02/2015 19:38:49",
      "content": "<p>What can I say, there is no such thing as a free lunch ;)</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69558,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "04/03/2015 20:16:04",
      "content": "<p>@sedielem, I saw your back and forth with Fred on how to make theano work with torque back in 2013. I am assuming that you guys have a torque cluster. Can I ask you a question?</p>\n<p>If you have two computer with GPUs where&nbsp;one runs during the day time and not at night and the other during the night, is there anyway to put&nbsp;these two computers in a cluster and transparently run a CNN code that may run more than a day or so without interruption?&nbsp;</p>\n<p>In other words, once you&nbsp;start to train using the GPU of one of the machines in the cluster, is there any way to make that same training job to switch context to other available GPUs on other machines in the cluster without interruption of the training job?</p>\n<p>Note: This is of course with out abandoning Theano and co and resorting to write the GPU kernels and the internal details yourself.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69613,
      "author_name": "sedielem",
      "author_url": "",
      "post_date": "04/04/2015 09:24:20",
      "content": "<p>I don't think it's possible to switch GPUs in the middle of a run unless you actually abort it, start it on the other GPU and load up the parameters from a backup.</p>\n<p>We no longer use torque for our GPU machines, but back then all of them had two identical GPUs each, so we just set them to run in exclusive mode (only one compute process per GPU) and that was sufficient. I don't have any experience with the GPU-specific features of torque, we didn't use those.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69628,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "04/04/2015 15:14:16",
      "content": "<p>Thanks for the reply. That's what I thought, just wanted to confirm .</p>\n<p>Did you switch to another cluster package or you just stopped using clusters and just run jobs manually specifying machines and GPUs? We are trying to asses if it is worth putting the effort to setup a cluster vs, just manually scheduling jobs in each machine specifying GPUs manually. From your experience, what's the benefit of setting up a cluster if availability is not transparent?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69644,
      "author_name": "sedielem",
      "author_url": "",
      "post_date": "04/04/2015 18:11:20",
      "content": "<p>Our little GPU farm has become a little too heterogeneous for our old approach (4 different types of GPUs now), and with long running experiments it's just more convenient to have them running in screen/tmux so you can check up on them easily. So right now we just have a google drive spreadsheet where everyone indicates which GPUs they're using. Pretty low-tech but it works OK for now. If we end up expanding it further this may become unworkable though.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "69263": "",
    "69405": "",
    "69451": "",
    "69458": "",
    "69460": "",
    "69558": "",
    "69613": "",
    "69628": "",
    "69644": ""
  },
  "source": "meta"
}