{
  "id": 15845,
  "title": "3rd Place Solution Report",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/15845",
  "author_name": "John Dunavent",
  "post_date": "2015-08-07T23:18:26.063000",
  "votes": 26,
  "comment_count": 12,
  "views": 6262,
  "content": "<p>We&#8217;d like to thank the California Healthcare Foundation and EyePACs for sponsoring and providing data for the competition. Many thanks to the Kaggle community. We leaned heavily on work shared by past competition winners.</p>\n\n<p>Our solution was an ensemble of 9 convolutional neural networks. We used a variety of model architectures. Our best performing were variations of <a href=\"http://arxiv.org/abs/1409.1556\">Simonyan and Zisserman</a>, closely followed by models featuring the <a href=\"http://arxiv.org/abs/1412.6071\">fractional max pooling layers</a> developed by the 1st place finisher (<a href=\"https://www.kaggle.com/btgraham\">Graham</a>). Also included was a model utilizing <a href=\"http://benanne.github.io/2015/03/17/plankton.html\">cyclic pooling</a> as described in the winning solution to the National Data Science Bowl (<a href=\"https://www.kaggle.com/sedielem\">Dieleman</a>). We found that using large image sizes and combining information from both eyes were key to getting good performance.</p>\n\n<p>A more in-depth summary and submission code are attached.</p>\n\n<p>Edit: Fixed documentation around cyclic pooling model. See v2 of the PDF.</p>",
  "messages": [
    {
      "id": 88866,
      "postDate": "2015-08-07T23:18:26.063Z",
      "content": "<p>We&#8217;d like to thank the California Healthcare Foundation and EyePACs for sponsoring and providing data for the competition. Many thanks to the Kaggle community. We leaned heavily on work shared by past competition winners.</p>\n\n<p>Our solution was an ensemble of 9 convolutional neural networks. We used a variety of model architectures. Our best performing were variations of <a href=\"http://arxiv.org/abs/1409.1556\">Simonyan and Zisserman</a>, closely followed by models featuring the <a href=\"http://arxiv.org/abs/1412.6071\">fractional max pooling layers</a> developed by the 1st place finisher (<a href=\"https://www.kaggle.com/btgraham\">Graham</a>). Also included was a model utilizing <a href=\"http://benanne.github.io/2015/03/17/plankton.html\">cyclic pooling</a> as described in the winning solution to the National Data Science Bowl (<a href=\"https://www.kaggle.com/sedielem\">Dieleman</a>). We found that using large image sizes and combining information from both eyes were key to getting good performance.</p>\n\n<p>A more in-depth summary and submission code are attached.</p>\n\n<p>Edit: Fixed documentation around cyclic pooling model. See v2 of the PDF.</p>",
      "rawMarkdown": "We’d like to thank the California Healthcare Foundation and EyePACs for sponsoring and providing data for the competition. Many thanks to the Kaggle community. We leaned heavily on work shared by past competition winners.\r\n\r\nOur solution was an ensemble of 9 convolutional neural networks. We used a variety of model architectures. Our best performing were variations of [Simonyan and Zisserman][1], closely followed by models featuring the [fractional max pooling layers][2] developed by the 1st place finisher ([Graham][3]). Also included was a model utilizing [cyclic pooling][4] as described in the winning solution to the National Data Science Bowl ([Dieleman][5]). We found that using large image sizes and combining information from both eyes were key to getting good performance.\r\n\r\nA more in-depth summary and submission code are attached.\r\n\r\nEdit: Fixed documentation around cyclic pooling model. See v2 of the PDF.\r\n\r\n\r\n  [1]: http://arxiv.org/abs/1409.1556\r\n  [2]: http://arxiv.org/abs/1412.6071\r\n  [3]: https://www.kaggle.com/btgraham\r\n  [4]: http://benanne.github.io/2015/03/17/plankton.html\r\n  [5]: https://www.kaggle.com/sedielem",
      "votes": 26
    },
    {
      "id": 93510,
      "postDate": "2015-09-26T22:10:36.007Z",
      "content": "<p>[quote=LKT;93499]</p>\n\n<p>Jun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. </p>\n\n<p>/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)</p>\n\n<p>I've tried a couple different formats for the meta.tsv file but haven't solved it.</p>\n\n<p>Congratulations again!</p>\n\n<p>[/quote]</p>\n\n<p>I attached the meta.tsv and val01.list for you to test. Let me know if you still have problems.</p>\n\n<p>Jun</p>",
      "rawMarkdown": "[quote=LKT;93499]\r\n\r\nJun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. \r\n\r\n/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)\r\n\r\nI've tried a couple different formats for the meta.tsv file but haven't solved it.\r\n\r\nCongratulations again!\r\n\r\n[/quote]\r\n\r\nI attached the meta.tsv and val01.list for you to test. Let me know if you still have problems.\r\n\r\nJun\r\n",
      "votes": 1
    },
    {
      "id": 88899,
      "postDate": "2015-08-08T14:00:05.147Z",
      "content": "<p>[quote=sedielem;88895]</p>\n\n<p>Nice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?</p>\n\n<p>Another thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much.</p>\n\n<p>[/quote]</p>\n\n<p>That is a mistake in our documentation, only the first CS layer should be there and later ones are copy &amp; paste mistakes.  Thanks for point this out and we will correct it.  We wish we had more time to play with the cyclic architecture since it is a really innovative idea.</p>\n\n<p>For your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.</p>\n\n<p>Hopefully, you will begin to write your new packages in Lua/Torch framework, so we can directly use them without implementing ourselves :-).</p>\n\n<p>Jun</p>",
      "rawMarkdown": "[quote=sedielem;88895]\r\n\r\nNice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?\r\n\r\nAnother thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much.\r\n\r\n[/quote]\r\n\r\nThat is a mistake in our documentation, only the first CS layer should be there and later ones are copy & paste mistakes.  Thanks for point this out and we will correct it.  We wish we had more time to play with the cyclic architecture since it is a really innovative idea.\r\n\r\nFor your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.\r\n\r\nHopefully, you will begin to write your new packages in Lua/Torch framework, so we can directly use them without implementing ourselves :-).\r\n\r\nJun",
      "votes": 2
    },
    {
      "id": 147181,
      "postDate": "2016-11-29T08:59:43.730Z",
      "rawMarkdown": ""
    },
    {
      "id": 103767,
      "postDate": "2016-01-06T11:41:48.413Z",
      "content": "<p>Kiran,</p>\n\n<p>I suspect you are using too large a batch size and ran out of memory.  Try a batch size 4 or 8 first (even 2) to see whether you can reproduce the problem.  On the other hand, this model is likely too large for K20 to train in a meaningful time.  A TitanX coupled with cudnn4 will be much much better.</p>\n\n<p>Jun</p>",
      "rawMarkdown": "Kiran,\r\n\r\nI suspect you are using too large a batch size and ran out of memory.  Try a batch size 4 or 8 first (even 2) to see whether you can reproduce the problem.  On the other hand, this model is likely too large for K20 to train in a meaningful time.  A TitanX coupled with cudnn4 will be much much better.\r\n\r\nJun"
    },
    {
      "id": 93499,
      "postDate": "2015-09-26T17:52:49.947Z",
      "content": "<p>Jun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. </p>\n\n<p>/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)</p>\n\n<p>I've tried a couple different formats for the meta.tsv file but haven't solved it.</p>\n\n<p>Congratulations again!</p>",
      "rawMarkdown": "Jun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. \r\n\r\n/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)\r\n\r\nI've tried a couple different formats for the meta.tsv file but haven't solved it.\r\n\r\nCongratulations again!"
    },
    {
      "id": 88992,
      "postDate": "2015-08-10T01:21:20.783Z",
      "content": "<p>[quote=Keaton;88986]</p>\n\n<p>Congratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.</p>\n\n<p>So far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.</p>\n\n<p>Thanks in advanced for any feedback.</p>\n\n<p>[/quote]</p>\n\n<p>We used Torch framework from the beginning to the end.  As you pointed out, each framework has its pros and cons.  Some framework may have more development/support for a particular problem than others. I would suggest to use the one that you feel most comfortable with.  Visiting the forum for each of the framework may also help you make the decision.</p>",
      "rawMarkdown": "[quote=Keaton;88986]\r\n\r\nCongratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.\r\n\r\nSo far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.\r\n\r\nThanks in advanced for any feedback.\r\n\r\n[/quote]\r\n\r\nWe used Torch framework from the beginning to the end.  As you pointed out, each framework has its pros and cons.  Some framework may have more development/support for a particular problem than others. I would suggest to use the one that you feel most comfortable with.  Visiting the forum for each of the framework may also help you make the decision."
    },
    {
      "id": 88986,
      "postDate": "2015-08-09T23:54:11.630Z",
      "content": "<p>Congratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.</p>\n\n<p>So far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.</p>\n\n<p>Thanks in advanced for any feedback.</p>",
      "rawMarkdown": "Congratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.\r\n\r\nSo far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.\r\n\r\nThanks in advanced for any feedback."
    },
    {
      "id": 88918,
      "postDate": "2015-08-08T18:35:03.407Z",
      "content": "<p>[quote=Junonia;88899]</p>\n\n<p>For your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.\n[/quote]</p>\n\n<p>This is a very interesting finding! Usually most of the parameters in a net like this are precisely between those two fully connected layers, so that is where regularization is needed the most. Then again, adding dropout anywhere else in the network will of course have a slight regularization effect on all of the other layers too, maybe that sufficed in this case.</p>",
      "rawMarkdown": "[quote=Junonia;88899]\r\n\r\nFor your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.\r\n[/quote]\r\n\r\nThis is a very interesting finding! Usually most of the parameters in a net like this are precisely between those two fully connected layers, so that is where regularization is needed the most. Then again, adding dropout anywhere else in the network will of course have a slight regularization effect on all of the other layers too, maybe that sufficed in this case.\r\n"
    },
    {
      "id": 88898,
      "postDate": "2015-08-08T13:31:04.813Z",
      "content": "<p>[quote=Julian de Wit;88892]</p>\n\n<p>Hello and congrats.\nThank you for your clear report.</p>\n\n<p>I tried clipped MSE at one moment but did not see much change in score. \nDo you have an estimate how much it helped ? Did it also help in faster training for you ?</p>\n\n<p>Thanks for pointing out to Nvidia NPP. That was something I really liked to have. </p>\n\n<p>Regards,\n  Julian.</p>\n\n<p>[/quote]</p>\n\n<p>Julian,  </p>\n\n<p>We did not have a comprehensive evaluation of the effect of clipped MSE.  Since it did not appear to hurt score and had minimum effect on running time, we have kept using it.</p>\n\n<p>Thanks</p>\n\n<p>Jun </p>",
      "rawMarkdown": "[quote=Julian de Wit;88892]\r\n\r\nHello and congrats.\r\nThank you for your clear report.\r\n\r\nI tried clipped MSE at one moment but did not see much change in score. \r\nDo you have an estimate how much it helped ? Did it also help in faster training for you ?\r\n\r\nThanks for pointing out to Nvidia NPP. That was something I really liked to have. \r\n\r\nRegards,\r\n  Julian.\r\n\r\n\r\n\r\n\r\n[/quote]\r\n\r\nJulian,  \r\n\r\nWe did not have a comprehensive evaluation of the effect of clipped MSE.  Since it did not appear to hurt score and had minimum effect on running time, we have kept using it.\r\n\r\nThanks\r\n\r\nJun "
    },
    {
      "id": 88895,
      "postDate": "2015-08-08T11:53:18.140Z",
      "content": "<p>Nice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?</p>\n\n<p>Another thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much.</p>",
      "rawMarkdown": "Nice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?\r\n\r\nAnother thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much."
    },
    {
      "id": 88892,
      "postDate": "2015-08-08T09:15:01.563Z",
      "content": "<p>Hello and congrats.\nThank you for your clear report.</p>\n\n<p>I tried clipped MSE at one moment but did not see much change in score. \nDo you have an estimate how much it helped ? Did it also help in faster training for you ?</p>\n\n<p>Thanks for pointing out to Nvidia NPP. That was something I really liked to have. </p>\n\n<p>Regards,\n  Julian.</p>",
      "rawMarkdown": "Hello and congrats.\r\nThank you for your clear report.\r\n\r\nI tried clipped MSE at one moment but did not see much change in score. \r\nDo you have an estimate how much it helped ? Did it also help in faster training for you ?\r\n\r\nThanks for pointing out to Nvidia NPP. That was something I really liked to have. \r\n\r\nRegards,\r\n  Julian.\r\n\r\n\r\n"
    },
    {
      "id": 103752,
      "postDate": "2016-01-06T06:15:56.490Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 93510,
      "author_name": "Junonia",
      "author_url": "",
      "post_date": "2015-09-26T22:10:36.007000",
      "content": "<p>[quote=LKT;93499]</p>\n\n<p>Jun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. </p>\n\n<p>/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)</p>\n\n<p>I've tried a couple different formats for the meta.tsv file but haven't solved it.</p>\n\n<p>Congratulations again!</p>\n\n<p>[/quote]</p>\n\n<p>I attached the meta.tsv and val01.list for you to test. Let me know if you still have problems.</p>\n\n<p>Jun</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 88899,
      "author_name": "Junonia",
      "author_url": "",
      "post_date": "2015-08-08T14:00:05.147000",
      "content": "<p>[quote=sedielem;88895]</p>\n\n<p>Nice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?</p>\n\n<p>Another thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much.</p>\n\n<p>[/quote]</p>\n\n<p>That is a mistake in our documentation, only the first CS layer should be there and later ones are copy &amp; paste mistakes.  Thanks for point this out and we will correct it.  We wish we had more time to play with the cyclic architecture since it is a really innovative idea.</p>\n\n<p>For your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.</p>\n\n<p>Hopefully, you will begin to write your new packages in Lua/Torch framework, so we can directly use them without implementing ourselves :-).</p>\n\n<p>Jun</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 147181,
      "author_name": "samaneh",
      "author_url": "",
      "post_date": "2016-11-29T08:59:43.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 103767,
      "author_name": "Junonia",
      "author_url": "",
      "post_date": "2016-01-06T11:41:48.413000",
      "content": "<p>Kiran,</p>\n\n<p>I suspect you are using too large a batch size and ran out of memory.  Try a batch size 4 or 8 first (even 2) to see whether you can reproduce the problem.  On the other hand, this model is likely too large for K20 to train in a meaningful time.  A TitanX coupled with cudnn4 will be much much better.</p>\n\n<p>Jun</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 93499,
      "author_name": "Leo Tam",
      "author_url": "",
      "post_date": "2015-09-26T17:52:49.947000",
      "content": "<p>Jun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. </p>\n\n<p>/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)</p>\n\n<p>I've tried a couple different formats for the meta.tsv file but haven't solved it.</p>\n\n<p>Congratulations again!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 88992,
      "author_name": "Junonia",
      "author_url": "",
      "post_date": "2015-08-10T01:21:20.783000",
      "content": "<p>[quote=Keaton;88986]</p>\n\n<p>Congratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.</p>\n\n<p>So far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.</p>\n\n<p>Thanks in advanced for any feedback.</p>\n\n<p>[/quote]</p>\n\n<p>We used Torch framework from the beginning to the end.  As you pointed out, each framework has its pros and cons.  Some framework may have more development/support for a particular problem than others. I would suggest to use the one that you feel most comfortable with.  Visiting the forum for each of the framework may also help you make the decision.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 88986,
      "author_name": "Keaton",
      "author_url": "",
      "post_date": "2015-08-09T23:54:11.630000",
      "content": "<p>Congratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.</p>\n\n<p>So far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.</p>\n\n<p>Thanks in advanced for any feedback.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 88918,
      "author_name": "sedielem",
      "author_url": "",
      "post_date": "2015-08-08T18:35:03.407000",
      "content": "<p>[quote=Junonia;88899]</p>\n\n<p>For your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.\n[/quote]</p>\n\n<p>This is a very interesting finding! Usually most of the parameters in a net like this are precisely between those two fully connected layers, so that is where regularization is needed the most. Then again, adding dropout anywhere else in the network will of course have a slight regularization effect on all of the other layers too, maybe that sufficed in this case.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 88898,
      "author_name": "Junonia",
      "author_url": "",
      "post_date": "2015-08-08T13:31:04.813000",
      "content": "<p>[quote=Julian de Wit;88892]</p>\n\n<p>Hello and congrats.\nThank you for your clear report.</p>\n\n<p>I tried clipped MSE at one moment but did not see much change in score. \nDo you have an estimate how much it helped ? Did it also help in faster training for you ?</p>\n\n<p>Thanks for pointing out to Nvidia NPP. That was something I really liked to have. </p>\n\n<p>Regards,\n  Julian.</p>\n\n<p>[/quote]</p>\n\n<p>Julian,  </p>\n\n<p>We did not have a comprehensive evaluation of the effect of clipped MSE.  Since it did not appear to hurt score and had minimum effect on running time, we have kept using it.</p>\n\n<p>Thanks</p>\n\n<p>Jun </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 88895,
      "author_name": "sedielem",
      "author_url": "",
      "post_date": "2015-08-08T11:53:18.140000",
      "content": "<p>Nice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?</p>\n\n<p>Another thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 88892,
      "author_name": "Julian de Wit",
      "author_url": "",
      "post_date": "2015-08-08T09:15:01.563000",
      "content": "<p>Hello and congrats.\nThank you for your clear report.</p>\n\n<p>I tried clipped MSE at one moment but did not see much change in score. \nDo you have an estimate how much it helped ? Did it also help in faster training for you ?</p>\n\n<p>Thanks for pointing out to Nvidia NPP. That was something I really liked to have. </p>\n\n<p>Regards,\n  Julian.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 103752,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-06T06:15:56.490000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "88866": "We’d like to thank the California Healthcare Foundation and EyePACs for sponsoring and providing data for the competition. Many thanks to the Kaggle community. We leaned heavily on work shared by past competition winners.\r\n\r\nOur solution was an ensemble of 9 convolutional neural networks. We used a variety of model architectures. Our best performing were variations of [Simonyan and Zisserman][1], closely followed by models featuring the [fractional max pooling layers][2] developed by the 1st place finisher ([Graham][3]). Also included was a model utilizing [cyclic pooling][4] as described in the winning solution to the National Data Science Bowl ([Dieleman][5]). We found that using large image sizes and combining information from both eyes were key to getting good performance.\r\n\r\nA more in-depth summary and submission code are attached.\r\n\r\nEdit: Fixed documentation around cyclic pooling model. See v2 of the PDF.\r\n\r\n\r\n  [1]: http://arxiv.org/abs/1409.1556\r\n  [2]: http://arxiv.org/abs/1412.6071\r\n  [3]: https://www.kaggle.com/btgraham\r\n  [4]: http://benanne.github.io/2015/03/17/plankton.html\r\n  [5]: https://www.kaggle.com/sedielem",
    "93510": "[quote=LKT;93499]\r\n\r\nJun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. \r\n\r\n/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)\r\n\r\nI've tried a couple different formats for the meta.tsv file but haven't solved it.\r\n\r\nCongratulations again!\r\n\r\n[/quote]\r\n\r\nI attached the meta.tsv and val01.list for you to test. Let me know if you still have problems.\r\n\r\nJun\r\n",
    "88899": "[quote=sedielem;88895]\r\n\r\nNice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?\r\n\r\nAnother thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much.\r\n\r\n[/quote]\r\n\r\nThat is a mistake in our documentation, only the first CS layer should be there and later ones are copy & paste mistakes.  Thanks for point this out and we will correct it.  We wish we had more time to play with the cyclic architecture since it is a really innovative idea.\r\n\r\nFor your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.\r\n\r\nHopefully, you will begin to write your new packages in Lua/Torch framework, so we can directly use them without implementing ourselves :-).\r\n\r\nJun",
    "147181": "",
    "103767": "Kiran,\r\n\r\nI suspect you are using too large a batch size and ran out of memory.  Try a batch size 4 or 8 first (even 2) to see whether you can reproduce the problem.  On the other hand, this model is likely too large for K20 to train in a meaningful time.  A TitanX coupled with cudnn4 will be much much better.\r\n\r\nJun",
    "93499": "Jun, thank you very much for the documentation and code.  Might you be able to upload a sample meta.tsv file?  I've tried building a CSV with the list of images in meta.tsv but I'm still receiving the error below. \r\n\r\n/usr/local/bin/luajit: /usr/local/share/lua/5.1/csvigo/init.lua:102: bad argument #1 to 'ipairs' (table expected, got nil)\r\n\r\nI've tried a couple different formats for the meta.tsv file but haven't solved it.\r\n\r\nCongratulations again!",
    "88992": "[quote=Keaton;88986]\r\n\r\nCongratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.\r\n\r\nSo far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.\r\n\r\nThanks in advanced for any feedback.\r\n\r\n[/quote]\r\n\r\nWe used Torch framework from the beginning to the end.  As you pointed out, each framework has its pros and cons.  Some framework may have more development/support for a particular problem than others. I would suggest to use the one that you feel most comfortable with.  Visiting the forum for each of the framework may also help you make the decision.",
    "88986": "Congratulations and great research.  I am curious, did your team stick to only one neural network framework and/or library, from the beginning to the end of the competition, or did you ever change the frameworks and/or library you were using.  I am interested in learning about which frameworks and libraries are better suited toward specific convolutional neural network, deep learning research.\r\n\r\nSo far, I have heard of the Torch library, Theano, cxxnet, and Caffe.  Most of my experience with NN research so far is with Caffe.  I find that there are pros and cons to all of these.\r\n\r\nThanks in advanced for any feedback.",
    "88918": "[quote=Junonia;88899]\r\n\r\nFor your second question,  the single dropout following the convolutional layer seemed to provide enough regularization in our settings.  We experimented with more dropout between the fully connected layers (as well as additional dropout between convolutional layers),  we did not notice a difference in performance in this dataset (as well as in the plankton dataset) and sometimes it actually hurt a little bit.  However, this may not be true for other dataset or in other settings.\r\n[/quote]\r\n\r\nThis is a very interesting finding! Usually most of the parameters in a net like this are precisely between those two fully connected layers, so that is where regularization is needed the most. Then again, adding dropout anywhere else in the network will of course have a slight regularization effect on all of the other layers too, maybe that sufficed in this case.\r\n",
    "88898": "[quote=Julian de Wit;88892]\r\n\r\nHello and congrats.\r\nThank you for your clear report.\r\n\r\nI tried clipped MSE at one moment but did not see much change in score. \r\nDo you have an estimate how much it helped ? Did it also help in faster training for you ?\r\n\r\nThanks for pointing out to Nvidia NPP. That was something I really liked to have. \r\n\r\nRegards,\r\n  Julian.\r\n\r\n\r\n\r\n\r\n[/quote]\r\n\r\nJulian,  \r\n\r\nWe did not have a comprehensive evaluation of the effect of clipped MSE.  Since it did not appear to hurt score and had minimum effect on running time, we have kept using it.\r\n\r\nThanks\r\n\r\nJun ",
    "88895": "Nice work, congrats! I'm not sure I fully understand the architecture description for the cyclic pooling net. From the table, you seem to have 4 cyclic slice layers at multiple points in the net, which doesn't really make sense to me. Could you explain what you mean by this?\r\n\r\nAnother thing I noticed is the lack of dropout between the fully connected layers, and between the topmost hidden layer and the output layer. Did you find it beneficial to only have dropout between the convolutional and the dense part of the net? I would imagine this doesn't regularize things very much.",
    "88892": "Hello and congrats.\r\nThank you for your clear report.\r\n\r\nI tried clipped MSE at one moment but did not see much change in score. \r\nDo you have an estimate how much it helped ? Did it also help in faster training for you ?\r\n\r\nThanks for pointing out to Nvidia NPP. That was something I really liked to have. \r\n\r\nRegards,\r\n  Julian.\r\n\r\n\r\n",
    "103752": ""
  }
}