{
  "id": 15617,
  "title": "Team o_O Solution Summary",
  "url": "/competitions/diabetic-retinopathy-detection/writeups/o-o-team-o-o-solution-summary",
  "author_name": "",
  "post_date": "2015-07-31T08:14:08.113Z",
  "votes": 46,
  "comment_count": 30,
  "views": 9225,
  "content": "<p>First I'd like to thank team mate Stephan, the hosts and everyone \nelse who joined and congratulate the other winners! </p>\n\n<p>Below is a brief solution summary while things are fresh before we get\ntime to do a proper writeup and publish our code.</p>\n\n<h2>Team o_O Solution Summary</h2>\n\n<p>Neural networks trained using <a href=\"https://github.com/Lasagne/Lasagne\">lasagne</a>\nand <a href=\"https://github.com/dnouri/nolearn\">nolearn</a>. Both libraries are awesome, \nwell documented and easy to get started with.</p>\n\n<p>A lot of inspiration and some code was taken from winners of\nthe national data science bowl competition. Thanks!</p>\n\n<ul>\n<li><a href=\"https://github.com/benanne/kaggle-ndsb\">&#8779; Deep Sea &#8779;</a>: \ndata augmentation, RMS pooling, pseudo random augmentation averaging \n(used for feature extraction here)</li>\n<li><a href=\"https://www.kaggle.com/c/datasciencebowl/forums/t/13166/happy-lantern-festival-report-and-code\">Happy Lantern Festival</a>\nCapacity and coverage analysis of conv nets (it's built into nolearn)</li>\n</ul>\n\n<h3>Preprocessing</h3>\n\n<ul>\n<li>We selected the smallest rectangle that contained the entire eye.</li>\n<li>When this didn't work well (dark images) or the image was already\nalmost a square we defaulted to selecting the center square.</li>\n<li>Resized to 512, 256, 128 pixel squares.</li>\n<li>Before the data augmentation we scaled each channel to have zero mean \nand unit variance.</li>\n</ul>\n\n<h3>Augmentation</h3>\n\n<ul>\n<li>360 degrees rotation, translation, scaling, stretching (all uniform)</li>\n<li>Krizhevsky color augmentation (gaussian)</li>\n</ul>\n\n<p>These augmentations were always applied.</p>\n\n<h3>Network Configurations</h3>\n\n<pre><code>                            net A       |          net B\n            units   filter stride size  |  filter stride size\n 1 Input                           448  |     4           448\n 2 Conv        32     5       2    224  |     4     2     224\n 3 Conv        32     3            224  |     4           225 \n 4 MaxPool            3       2    111  |     3     2     112 \n 5 Conv        64     5       2     56  |     4     2      56\n 6 Conv        64     3             56  |     4            57\n 7 Conv        64     3             56  |     4            56\n 8 MaxPool            3       2     27  |     3     2      27\n 9 Conv       128     3             27  |     4            28\n10 Conv       128     3             27  |     4            27\n11 Conv       128     3             27  |     4            28\n12 MaxPool            3       2     13  |     3     2      13\n13 Conv       256     3             13  |     4            14\n14 Conv       256     3             13  |     4            13\n15 Conv       256     3             13  |     4            14\n16 MaxPool            3       2      6  |     3     2       6\n17 Conv       512     3              6  |     4             5\n18 Conv       512     3              6  |   n/a           n/a\n19 RMSPool            3       3      2  |     3     2       2\n20 Dropout\n21 Dense     1024\n22 Maxout    512\n23 Dropout\n24 Dense     1024\n25 Maxout    512\n</code></pre>\n\n<ul>\n<li>Leaky (0.01) rectifier units following each conv and dense layer.</li>\n<li>L2 weight decay factor 0.0005</li>\n<li>Mean squared error objective.</li>\n<li>Untied biases.</li>\n</ul>\n\n<h3>Training</h3>\n\n<ul>\n<li>We used 10% of the patients as validation set.</li>\n<li>128 px images -&gt; layers 1 - 11 and 20 to 25.</li>\n<li>256 px images -&gt; layers 1 - 15 and 20 to 25. Weights of layer 1 - 11 initialized \nwith weights from above.</li>\n<li>512 px images -&gt; all layers. Weights of layers 1 - 15 initialized with\nweights from above.</li>\n<li>Training started with resampling such that all classes were present in equal \nfractions. We then gradually decreased the balancing after each epoch to\narrive at final &quot;resampling weights&quot; of 1 for class 0 and 2 for the other \nclasses.</li>\n<li>Nesterov momentum (0.9) with fixed learning rate schedule over 250 epochs.\n\n<ul><li>epoch 0: 0.003</li>\n<li>epoch 150: 0.0003</li>\n<li>epoch 220: 0.00003</li>\n<li>For the nets for 256 and 128 pixels images we stopped training already \nafter 200 epochs.</li></ul></li>\n</ul>\n\n<p>Models were trained on a GTX 970 and a GTX 980Ti (highly recommended) with batch\nsizes small enough to fit into memory (usually 24 to 48 for the large networks\nand 128 for the smaller ones).</p>\n\n<h3>&quot;Per Patient&quot; Blend</h3>\n\n<p>Extracted mean and standard deviation of RMSPool layer for 50 pseudo random \naugmentations for three sets of weights (best validation score, best kappa,\nfinal weights) for net A and B.</p>\n\n<p>For each eye (or patient) used the following as input features for blending,</p>\n\n<pre><code>[this_eye_mean, other_eye_mean, this_eye_stddev, other_eye_stddev, left_eye_indicator]\n</code></pre>\n\n<p>We standardized all features to have zero mean and unit variance and used them to \ntrain a network of shape,</p>\n\n<pre><code>Input        8193\nDense          32\nMaxout         16\nDense          32\nMaxout         16\n</code></pre>\n\n<ul>\n<li>L1 regularization (2e-5) on the first layer and L2 regularization (0.005) \neverywhere.</li>\n<li><a href=\"http://arxiv.org/abs/1412.6980\">Adam Updates</a> with fixed learning rate\nschedule over 100 epochs.\n\n<ul><li>epoch  0: 5e-4</li>\n<li>epoch 60: 5e-5</li>\n<li>epoch 80: 5e-6</li>\n<li>epoch 90: 5e-7</li></ul></li>\n<li>Mean squared error objective.</li>\n<li>Batches of size 128, replace batch with probability\n\n<ul><li>0.2 with batch sampled such that classes are balanced</li>\n<li>0.5 with batch sampled uniformly from all images (shuffled)</li></ul></li>\n</ul>\n\n<p>The mean output of the six blend networks (conv nets A and B, 3 sets of\nweights each) was then thresholded at <code>[0.5, 1.5, 2.5, 3.5]</code> to get \ninteger levels for submission.</p>\n\n<h3>Notes</h3>\n\n<p>I just noticed that we achieved a slightly higher private LB score for a \nsubmission where we optimized the thresholds on the validation set to\nmaximize kappa. But this didn't work as well on the public LB and we \ndidn't pursue it any further or select it for scoring at the end.</p>\n\n<p>When using very leaky (0.33) rectifier units we managed to train the\nnetworks for 512 px images directly and the MSE and kappa score were comparable\nwith what we achieved with leaky rectifiers. However after blending the features \nfor each patients over pseudo random augmentations the validation and LB kappa \nwere never quite as good as what we obtained with leaky rectifier units and\ninitialization from smaller nets. We went back to using 0.01 leaky rectifiers \nafterwards and gave up on training the larger networks from scratch.</p>",
  "messages": [
    {
      "id": "87432",
      "postDate": "07/29/2015 14:34:53",
      "content": "<p>First I'd like to thank team mate Stephan, the hosts and everyone \nelse who joined and congratulate the other winners! </p>\n\n<p>Below is a brief solution summary while things are fresh before we get\ntime to do a proper writeup and publish our code.</p>\n\n<h2>Team o_O Solution Summary</h2>\n\n<p>Neural networks trained using <a href=\"https://github.com/Lasagne/Lasagne\">lasagne</a>\nand <a href=\"https://github.com/dnouri/nolearn\">nolearn</a>. Both libraries are awesome, \nwell documented and easy to get started with.</p>\n\n<p>A lot of inspiration and some code was taken from winners of\nthe national data science bowl competition. Thanks!</p>\n\n<ul>\n<li><a href=\"https://github.com/benanne/kaggle-ndsb\">&#8779; Deep Sea &#8779;</a>: \ndata augmentation, RMS pooling, pseudo random augmentation averaging \n(used for feature extraction here)</li>\n<li><a href=\"https://www.kaggle.com/c/datasciencebowl/forums/t/13166/happy-lantern-festival-report-and-code\">Happy Lantern Festival</a>\nCapacity and coverage analysis of conv nets (it's built into nolearn)</li>\n</ul>\n\n<h3>Preprocessing</h3>\n\n<ul>\n<li>We selected the smallest rectangle that contained the entire eye.</li>\n<li>When this didn't work well (dark images) or the image was already\nalmost a square we defaulted to selecting the center square.</li>\n<li>Resized to 512, 256, 128 pixel squares.</li>\n<li>Before the data augmentation we scaled each channel to have zero mean \nand unit variance.</li>\n</ul>\n\n<h3>Augmentation</h3>\n\n<ul>\n<li>360 degrees rotation, translation, scaling, stretching (all uniform)</li>\n<li>Krizhevsky color augmentation (gaussian)</li>\n</ul>\n\n<p>These augmentations were always applied.</p>\n\n<h3>Network Configurations</h3>\n\n<pre><code>                            net A       |          net B\n            units   filter stride size  |  filter stride size\n 1 Input                           448  |     4           448\n 2 Conv        32     5       2    224  |     4     2     224\n 3 Conv        32     3            224  |     4           225 \n 4 MaxPool            3       2    111  |     3     2     112 \n 5 Conv        64     5       2     56  |     4     2      56\n 6 Conv        64     3             56  |     4            57\n 7 Conv        64     3             56  |     4            56\n 8 MaxPool            3       2     27  |     3     2      27\n 9 Conv       128     3             27  |     4            28\n10 Conv       128     3             27  |     4            27\n11 Conv       128     3             27  |     4            28\n12 MaxPool            3       2     13  |     3     2      13\n13 Conv       256     3             13  |     4            14\n14 Conv       256     3             13  |     4            13\n15 Conv       256     3             13  |     4            14\n16 MaxPool            3       2      6  |     3     2       6\n17 Conv       512     3              6  |     4             5\n18 Conv       512     3              6  |   n/a           n/a\n19 RMSPool            3       3      2  |     3     2       2\n20 Dropout\n21 Dense     1024\n22 Maxout    512\n23 Dropout\n24 Dense     1024\n25 Maxout    512\n</code></pre>\n\n<ul>\n<li>Leaky (0.01) rectifier units following each conv and dense layer.</li>\n<li>L2 weight decay factor 0.0005</li>\n<li>Mean squared error objective.</li>\n<li>Untied biases.</li>\n</ul>\n\n<h3>Training</h3>\n\n<ul>\n<li>We used 10% of the patients as validation set.</li>\n<li>128 px images -&gt; layers 1 - 11 and 20 to 25.</li>\n<li>256 px images -&gt; layers 1 - 15 and 20 to 25. Weights of layer 1 - 11 initialized \nwith weights from above.</li>\n<li>512 px images -&gt; all layers. Weights of layers 1 - 15 initialized with\nweights from above.</li>\n<li>Training started with resampling such that all classes were present in equal \nfractions. We then gradually decreased the balancing after each epoch to\narrive at final &quot;resampling weights&quot; of 1 for class 0 and 2 for the other \nclasses.</li>\n<li>Nesterov momentum (0.9) with fixed learning rate schedule over 250 epochs.\n\n<ul><li>epoch 0: 0.003</li>\n<li>epoch 150: 0.0003</li>\n<li>epoch 220: 0.00003</li>\n<li>For the nets for 256 and 128 pixels images we stopped training already \nafter 200 epochs.</li></ul></li>\n</ul>\n\n<p>Models were trained on a GTX 970 and a GTX 980Ti (highly recommended) with batch\nsizes small enough to fit into memory (usually 24 to 48 for the large networks\nand 128 for the smaller ones).</p>\n\n<h3>&quot;Per Patient&quot; Blend</h3>\n\n<p>Extracted mean and standard deviation of RMSPool layer for 50 pseudo random \naugmentations for three sets of weights (best validation score, best kappa,\nfinal weights) for net A and B.</p>\n\n<p>For each eye (or patient) used the following as input features for blending,</p>\n\n<pre><code>[this_eye_mean, other_eye_mean, this_eye_stddev, other_eye_stddev, left_eye_indicator]\n</code></pre>\n\n<p>We standardized all features to have zero mean and unit variance and used them to \ntrain a network of shape,</p>\n\n<pre><code>Input        8193\nDense          32\nMaxout         16\nDense          32\nMaxout         16\n</code></pre>\n\n<ul>\n<li>L1 regularization (2e-5) on the first layer and L2 regularization (0.005) \neverywhere.</li>\n<li><a href=\"http://arxiv.org/abs/1412.6980\">Adam Updates</a> with fixed learning rate\nschedule over 100 epochs.\n\n<ul><li>epoch  0: 5e-4</li>\n<li>epoch 60: 5e-5</li>\n<li>epoch 80: 5e-6</li>\n<li>epoch 90: 5e-7</li></ul></li>\n<li>Mean squared error objective.</li>\n<li>Batches of size 128, replace batch with probability\n\n<ul><li>0.2 with batch sampled such that classes are balanced</li>\n<li>0.5 with batch sampled uniformly from all images (shuffled)</li></ul></li>\n</ul>\n\n<p>The mean output of the six blend networks (conv nets A and B, 3 sets of\nweights each) was then thresholded at <code>[0.5, 1.5, 2.5, 3.5]</code> to get \ninteger levels for submission.</p>\n\n<h3>Notes</h3>\n\n<p>I just noticed that we achieved a slightly higher private LB score for a \nsubmission where we optimized the thresholds on the validation set to\nmaximize kappa. But this didn't work as well on the public LB and we \ndidn't pursue it any further or select it for scoring at the end.</p>\n\n<p>When using very leaky (0.33) rectifier units we managed to train the\nnetworks for 512 px images directly and the MSE and kappa score were comparable\nwith what we achieved with leaky rectifiers. However after blending the features \nfor each patients over pseudo random augmentations the validation and LB kappa \nwere never quite as good as what we obtained with leaky rectifier units and\ninitialization from smaller nets. We went back to using 0.01 leaky rectifiers \nafterwards and gave up on training the larger networks from scratch.</p>",
      "rawMarkdown": "First I'd like to thank team mate Stephan, the hosts and everyone \r\nelse who joined and congratulate the other winners! \r\n\r\nBelow is a brief solution summary while things are fresh before we get\r\ntime to do a proper writeup and publish our code.\r\n\r\n## Team o_O Solution Summary\r\n\r\nNeural networks trained using [lasagne](https://github.com/Lasagne/Lasagne)\r\nand [nolearn](https://github.com/dnouri/nolearn). Both libraries are awesome, \r\nwell documented and easy to get started with.\r\n\r\nA lot of inspiration and some code was taken from winners of\r\nthe national data science bowl competition. Thanks!\r\n\r\n- [≋ Deep Sea ≋](https://github.com/benanne/kaggle-ndsb): \r\n  data augmentation, RMS pooling, pseudo random augmentation averaging \r\n  (used for feature extraction here)\r\n- [Happy Lantern Festival](https://www.kaggle.com/c/datasciencebowl/forums/t/13166/happy-lantern-festival-report-and-code)\r\n  Capacity and coverage analysis of conv nets (it's built into nolearn)\r\n\r\n### Preprocessing\r\n- We selected the smallest rectangle that contained the entire eye.\r\n- When this didn't work well (dark images) or the image was already\r\n  almost a square we defaulted to selecting the center square.\r\n- Resized to 512, 256, 128 pixel squares.\r\n- Before the data augmentation we scaled each channel to have zero mean \r\n  and unit variance.\r\n\r\n### Augmentation\r\n- 360 degrees rotation, translation, scaling, stretching (all uniform)\r\n- Krizhevsky color augmentation (gaussian)\r\n\r\nThese augmentations were always applied.\r\n\r\n### Network Configurations\r\n\r\n                                net A       |          net B\r\n                units   filter stride size  |  filter stride size\r\n     1 Input                           448  |     4           448\r\n     2 Conv        32     5       2    224  |     4     2     224\r\n     3 Conv        32     3            224  |     4           225 \r\n     4 MaxPool            3       2    111  |     3     2     112 \r\n     5 Conv        64     5       2     56  |     4     2      56\r\n     6 Conv        64     3             56  |     4            57\r\n     7 Conv        64     3             56  |     4            56\r\n     8 MaxPool            3       2     27  |     3     2      27\r\n     9 Conv       128     3             27  |     4            28\r\n    10 Conv       128     3             27  |     4            27\r\n    11 Conv       128     3             27  |     4            28\r\n    12 MaxPool            3       2     13  |     3     2      13\r\n    13 Conv       256     3             13  |     4            14\r\n    14 Conv       256     3             13  |     4            13\r\n    15 Conv       256     3             13  |     4            14\r\n    16 MaxPool            3       2      6  |     3     2       6\r\n    17 Conv       512     3              6  |     4             5\r\n    18 Conv       512     3              6  |   n/a           n/a\r\n    19 RMSPool            3       3      2  |     3     2       2\r\n    20 Dropout\r\n    21 Dense     1024\r\n    22 Maxout    512\r\n    23 Dropout\r\n    24 Dense     1024\r\n    25 Maxout    512\r\n\r\n- Leaky (0.01) rectifier units following each conv and dense layer.\r\n- L2 weight decay factor 0.0005\r\n- Mean squared error objective.\r\n- Untied biases.\r\n\r\n### Training\r\n- We used 10% of the patients as validation set.\r\n- 128 px images -> layers 1 - 11 and 20 to 25.\r\n- 256 px images -> layers 1 - 15 and 20 to 25. Weights of layer 1 - 11 initialized \r\n  with weights from above.\r\n- 512 px images -> all layers. Weights of layers 1 - 15 initialized with\r\n  weights from above.\r\n- Training started with resampling such that all classes were present in equal \r\n  fractions. We then gradually decreased the balancing after each epoch to\r\n  arrive at final \"resampling weights\" of 1 for class 0 and 2 for the other \r\n  classes.\r\n- Nesterov momentum (0.9) with fixed learning rate schedule over 250 epochs.\r\n    + epoch 0: 0.003\r\n    + epoch 150: 0.0003\r\n    + epoch 220: 0.00003\r\n    + For the nets for 256 and 128 pixels images we stopped training already \r\n      after 200 epochs.\r\n\r\nModels were trained on a GTX 970 and a GTX 980Ti (highly recommended) with batch\r\nsizes small enough to fit into memory (usually 24 to 48 for the large networks\r\nand 128 for the smaller ones).\r\n\r\n### \"Per Patient\" Blend\r\nExtracted mean and standard deviation of RMSPool layer for 50 pseudo random \r\naugmentations for three sets of weights (best validation score, best kappa,\r\nfinal weights) for net A and B.\r\n\r\nFor each eye (or patient) used the following as input features for blending,\r\n\r\n    [this_eye_mean, other_eye_mean, this_eye_stddev, other_eye_stddev, left_eye_indicator]\r\n\r\nWe standardized all features to have zero mean and unit variance and used them to \r\ntrain a network of shape,\r\n\r\n    Input        8193\r\n    Dense          32\r\n    Maxout         16\r\n    Dense          32\r\n    Maxout         16\r\n\r\n- L1 regularization (2e-5) on the first layer and L2 regularization (0.005) \r\n  everywhere.\r\n- [Adam Updates](http://arxiv.org/abs/1412.6980) with fixed learning rate\r\n  schedule over 100 epochs.\r\n  + epoch  0: 5e-4\r\n  + epoch 60: 5e-5\r\n  + epoch 80: 5e-6\r\n  + epoch 90: 5e-7\r\n- Mean squared error objective.\r\n- Batches of size 128, replace batch with probability\r\n  + 0.2 with batch sampled such that classes are balanced\r\n  + 0.5 with batch sampled uniformly from all images (shuffled)\r\n\r\nThe mean output of the six blend networks (conv nets A and B, 3 sets of\r\nweights each) was then thresholded at `[0.5, 1.5, 2.5, 3.5]` to get \r\ninteger levels for submission.\r\n\r\n### Notes\r\nI just noticed that we achieved a slightly higher private LB score for a \r\nsubmission where we optimized the thresholds on the validation set to\r\nmaximize kappa. But this didn't work as well on the public LB and we \r\ndidn't pursue it any further or select it for scoring at the end.\r\n\r\nWhen using very leaky (0.33) rectifier units we managed to train the\r\nnetworks for 512 px images directly and the MSE and kappa score were comparable\r\nwith what we achieved with leaky rectifiers. However after blending the features \r\nfor each patients over pseudo random augmentations the validation and LB kappa \r\nwere never quite as good as what we obtained with leaky rectifier units and\r\ninitialization from smaller nets. We went back to using 0.01 leaky rectifiers \r\nafterwards and gave up on training the larger networks from scratch.",
      "votes": null
    },
    {
      "id": "87445",
      "postDate": "07/29/2015 15:39:33",
      "content": "<p>Congrats! Very interesting use of RMS pooling for spatial pooling -- in the NDSB we used it for pooling across rotations, where it clearly reduced overfitting compared to mean pooling or max pooling. Did you observe the same effect using it for spatial pooling? Or did it improve things in a different way?</p>\n\n<p>Regarding very leaky rectifiers: in my experience it can help to make the network even deeper. My intuition is that this compensates for the fact that these units are slightly less nonlinear than regular rectifiers (or leaky rectifiers with a small leak rate). Of course, having more layers also eats into your computational budget, so it's not always an option.</p>",
      "rawMarkdown": "Congrats! Very interesting use of RMS pooling for spatial pooling -- in the NDSB we used it for pooling across rotations, where it clearly reduced overfitting compared to mean pooling or max pooling. Did you observe the same effect using it for spatial pooling? Or did it improve things in a different way?\r\n\r\nRegarding very leaky rectifiers: in my experience it can help to make the network even deeper. My intuition is that this compensates for the fact that these units are slightly less nonlinear than regular rectifiers (or leaky rectifiers with a small leak rate). Of course, having more layers also eats into your computational budget, so it's not always an option.",
      "votes": null
    },
    {
      "id": "87465",
      "postDate": "07/29/2015 18:02:45",
      "content": "<p>Hello congrats!\nA question... In the end (also due to time pressure) I also bootstapped new bigger nets with the first layer(s) of previous nets. Have you any idea if this results in less variation ?  </p>\n\n<p>Somehow the new nets did not add too much to the ensemble. Daniels Hammacks net which had a completely different architecture added a lot.</p>",
      "rawMarkdown": "Hello congrats!\r\nA question... In the end (also due to time pressure) I also bootstapped new bigger nets with the first layer(s) of previous nets. Have you any idea if this results in less variation ?  \r\n\r\nSomehow the new nets did not add too much to the ensemble. Daniels Hammacks net which had a completely different architecture added a lot.",
      "votes": null
    },
    {
      "id": "87510",
      "postDate": "07/29/2015 22:57:33",
      "content": "<p>Congratulations and thanks for the post!</p>\n\n<p>So you didn't use any &quot;unusual&quot; loss function to approximate the kappa score?! Kappa was explicitly used only in &quot;blending&quot; phase?</p>\n\n<p>And one more question: did you check what would be the score without &quot;blending&quot; phase?</p>",
      "rawMarkdown": "Congratulations and thanks for the post!\r\n\r\nSo you didn't use any \"unusual\" loss function to approximate the kappa score?! Kappa was explicitly used only in \"blending\" phase?\r\n\r\nAnd one more question: did you check what would be the score without \"blending\" phase?",
      "votes": null
    },
    {
      "id": "87525",
      "postDate": "07/30/2015 01:31:09",
      "content": "<p>Thanks guys.</p>\n\n<p>@sedielem At this point I can't say with confidence if it made any difference at all. Around 2 months ago I tried replacing all MaxPool layers with RMSPool layers but that didn't work very well. But when using RMS pooling only in the last pooling layer the results were similar to using only max pooling layers and from then on I just stuck with it. It would be interesting to run it once more without the RMS pooling layer but the entire process of training one network takes 3-4 days.  I'll let you know if I get around to it. I also tried starting with leaky rate 0.5 and then gradually decreasing it down to zero using a shared variable. During training this worked great, but when trying to extract features or make predictions I noticed I was unable to affect the results by changing the leaky rate. Not knowing what was going on at this point I didn't pursue it any further.</p>\n\n<p>One thing I wanted to look at but didn't get around to yet was how much variance there was in the trained filter weights within each layer for different training strategies and rectifiers. I was also wondering if it's possible to apply a kind of penalty that &quot;forces&quot; the filters within each conv layer to become more diverse. Hoping this could help speed up training or allow smaller networks with similar performance. I trained my first neural net around the beginning of this year so I'm not sure how much sense this all makes.</p>\n\n<p>@julian I didn't try to use the same weights to bootstrap larger nets with different architecture so I can't comment on that. But I can give you some numbers of how much (or little) we gained from ensembling</p>\n\n<pre><code>net    set of weights   blend iter  augment averages       public LB     private LB\nB              1            1              20                0.84803        0.83916  \nB              1            1              50                0.84963        0.83973\nA+B            1            1              50                0.85090        0.84291\nA+B            1           10              50                0.85201        0.84308 \nA+B            2            1              50                0.85425        0.84379\nA+B            2           10              50                0.85199        0.84425\nA+B            3            1              50                0.85339        0.84479\nA+B            2            1              50                0.85316        0.84674 *\n(*in the last row thresholds were optimized for kappa on validation set)\n</code></pre>\n\n<p>@Hrant I did try using softmax classification in the beginning but the results weren't encouraging. I wondered quite a bit about how one could implement something close to kappa in theano but I never figured it out. At the end of the training phase both networks had a kappa score of about 0.80 (using info from one eye only). When feeding the output of the RMSPool layer (without averaging over different augmentations or blending patient eyes) into the blend network using only the features for one eye kappa is 0.805. When doing the same but using features averaged over 50 pseudo random augmentations kappa is 0.812. Then adding the standard deviation of the features brings kappa to 0.816. These numbers are however subject to some fluctuations.</p>",
      "rawMarkdown": "Thanks guys.\r\n\r\n@sedielem At this point I can't say with confidence if it made any difference at all. Around 2 months ago I tried replacing all MaxPool layers with RMSPool layers but that didn't work very well. But when using RMS pooling only in the last pooling layer the results were similar to using only max pooling layers and from then on I just stuck with it. It would be interesting to run it once more without the RMS pooling layer but the entire process of training one network takes 3-4 days.  I'll let you know if I get around to it. I also tried starting with leaky rate 0.5 and then gradually decreasing it down to zero using a shared variable. During training this worked great, but when trying to extract features or make predictions I noticed I was unable to affect the results by changing the leaky rate. Not knowing what was going on at this point I didn't pursue it any further.\r\n\r\nOne thing I wanted to look at but didn't get around to yet was how much variance there was in the trained filter weights within each layer for different training strategies and rectifiers. I was also wondering if it's possible to apply a kind of penalty that \"forces\" the filters within each conv layer to become more diverse. Hoping this could help speed up training or allow smaller networks with similar performance. I trained my first neural net around the beginning of this year so I'm not sure how much sense this all makes.\r\n\r\n@julian I didn't try to use the same weights to bootstrap larger nets with different architecture so I can't comment on that. But I can give you some numbers of how much (or little) we gained from ensembling\r\n\r\n    net    set of weights   blend iter  augment averages       public LB     private LB\r\n    B              1            1              20                0.84803        0.83916  \r\n    B              1            1              50                0.84963        0.83973\r\n    A+B            1            1              50                0.85090        0.84291\r\n    A+B            1           10              50                0.85201        0.84308 \r\n    A+B            2            1              50                0.85425        0.84379\r\n    A+B            2           10              50                0.85199        0.84425\r\n    A+B            3            1              50                0.85339        0.84479\r\n    A+B            2            1              50                0.85316        0.84674 *\r\n    (*in the last row thresholds were optimized for kappa on validation set)\r\n\r\n@Hrant I did try using softmax classification in the beginning but the results weren't encouraging. I wondered quite a bit about how one could implement something close to kappa in theano but I never figured it out. At the end of the training phase both networks had a kappa score of about 0.80 (using info from one eye only). When feeding the output of the RMSPool layer (without averaging over different augmentations or blending patient eyes) into the blend network using only the features for one eye kappa is 0.805. When doing the same but using features averaged over 50 pseudo random augmentations kappa is 0.812. Then adding the standard deviation of the features brings kappa to 0.816. These numbers are however subject to some fluctuations.",
      "votes": null
    },
    {
      "id": "87558",
      "postDate": "07/30/2015 08:11:33",
      "content": "<p>I also tried RMS-pooling in the last layer only and had similar experiences (about the same score, no spectacular differences at a quick glance). I was planning on using some networks with RMS-pooling in an ensemble but other things came up.</p>\n\n<p>I think most of your gains, relative to mine and some others, came from the blending. I didn't have enough time to get something working and it seemed that it would be hard to regularise. </p>\n\n<p>Congratulations, btw!</p>",
      "rawMarkdown": "I also tried RMS-pooling in the last layer only and had similar experiences (about the same score, no spectacular differences at a quick glance). I was planning on using some networks with RMS-pooling in an ensemble but other things came up.\r\n\r\nI think most of your gains, relative to mine and some others, came from the blending. I didn't have enough time to get something working and it seemed that it would be hard to regularise. \r\n\r\nCongratulations, btw!",
      "votes": null
    },
    {
      "id": "87587",
      "postDate": "07/30/2015 13:15:49",
      "content": "<p>Congratulations!  Many thanks for the post! I have a question. Could you give me link with explanation &quot;Krizhevsky color augmentation (gaussian)&quot;?</p>",
      "rawMarkdown": "Congratulations!  Many thanks for the post! I have a question. Could you give me link with explanation \"Krizhevsky color augmentation (gaussian)\"?",
      "votes": null
    },
    {
      "id": "87592",
      "postDate": "07/30/2015 13:49:32",
      "content": "<p>@Jeffrey, I agree it's probably what made up for most of the difference in score. It took some time to get the get the learning rate, L1 and L2 factors and most importantly resampling strategy to work well together. What I quite liked about it was that I could iterate on it rapidly without having to wait days to see results because it only takes a few minutes to train the blending network. When we trained a better/different network for feature extraction at a later point most of the improvements made to the blending network would persist.</p>\n\n<p>@Andrey, I think it originally appeared in the <a href=\"http://www.cs.toronto.edu/~fritz/absps/imagenet.pdf\">Imagenet Paper</a> in section 4.1. I used standard deviation 0.5 for net A and standard deviation 0.25 for net B as opposed to 0.1 in the paper. I guess our numbers are rather high but the images in this dataset looked like they vary a lot in terms of color composition and I was trying to take this into account. I'm uncertain if adding the color augmentation helped in the end but I figured even if didn't improve the results there's a good chance it would help against overfitting so we kept it.</p>",
      "rawMarkdown": "Jeffrey, I agree it's probably what made up for most of the difference in score. It took some time to get the get the learning rate, L1 and L2 factors and most importantly resampling strategy to work well together. What I quite liked about it was that I could iterate on it rapidly without having to wait days to see results because it only takes a few minutes to train the blending network. When we trained a better/different network for feature extraction at a later point most of the improvements made to the blending network would persist.\r\n\r\n@Andrey, I think it originally appeared in the [Imagenet Paper](http://www.cs.toronto.edu/~fritz/absps/imagenet.pdf) in section 4.1. I used standard deviation 0.5 for net A and standard deviation 0.25 for net B as opposed to 0.1 in the paper. I guess our numbers are rather high but the images in this dataset looked like they vary a lot in terms of color composition and I was trying to take this into account. I'm uncertain if adding the color augmentation helped in the end but I figured even if didn't improve the results there's a good chance it would help against overfitting so we kept it.",
      "votes": null
    },
    {
      "id": "87597",
      "postDate": "07/30/2015 14:26:13",
      "content": "<p>@Mathis Antony, Congrats!\nWhat is feature pool you mentioned, maxout? </p>",
      "rawMarkdown": "Mathis Antony, Congrats!\r\nWhat is feature pool you mentioned, maxout?",
      "votes": null
    },
    {
      "id": "87599",
      "postDate": "07/30/2015 14:51:04",
      "content": "<p>@old-ufo thanks. Yes it's maxout, pooled over 2 units. I updated the summary.</p>",
      "rawMarkdown": "old-ufo thanks. Yes it's maxout, pooled over 2 units. I updated the summary.",
      "votes": null
    },
    {
      "id": "88056",
      "postDate": "08/03/2015 06:02:20",
      "content": "<p>Thanks for the explanations. Where can we get the source code?</p>",
      "rawMarkdown": "Thanks for the explanations. Where can we get the source code?",
      "votes": null
    },
    {
      "id": "88058",
      "postDate": "08/03/2015 06:21:39",
      "content": "<p>@Ali we haven't released it yet. As far as I understand we'll first submit it to kaggle and when we fulfill all the requirements release it on github.</p>",
      "rawMarkdown": "Ali we haven't released it yet. As far as I understand we'll first submit it to kaggle and when we fulfill all the requirements release it on github.",
      "votes": null
    },
    {
      "id": "105863",
      "postDate": "01/27/2016 03:45:24",
      "content": "<p>I got this error</p>\n\n<p>from lasagne.objectives import Objective\nImportError: cannot import name Objective</p>\n\n<p>Can't find the solutions, any suggestions please?</p>",
      "rawMarkdown": "I got this error\r\n\r\nfrom lasagne.objectives import Objective\r\nImportError: cannot import name Objective\r\n\r\nCan't find the solutions, any suggestions please?",
      "votes": null
    },
    {
      "id": "105865",
      "postDate": "01/27/2016 04:01:22",
      "content": "<p>Newer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.</p>",
      "rawMarkdown": "Newer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.",
      "votes": null
    },
    {
      "id": "105869",
      "postDate": "01/27/2016 04:45:23",
      "content": "<p>Thanks, however the requirement.txt will give</p>\n\n<p>Cloning <a href=\"https://github.com/benanne/Lasagne.git\">https://github.com/benanne/Lasagne.git</a> (to 9f591a5f3a192028df9947ba1e4903b3b46e8fe0) to ./src/lasagne-master\n  Could not find a tag or branch '9f591a5f3a192028df9947ba1e4903b3b46e8fe0', assuming commit.</p>\n\n<p>[quote=Mathis Antony;105865]</p>\n\n<p>Newer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Thanks, however the requirement.txt will give\r\n\r\nCloning https://github.com/benanne/Lasagne.git (to 9f591a5f3a192028df9947ba1e4903b3b46e8fe0) to ./src/lasagne-master\r\n  Could not find a tag or branch '9f591a5f3a192028df9947ba1e4903b3b46e8fe0', assuming commit.\r\n\r\n\r\n\r\n[quote=Mathis Antony;105865]\r\n\r\nNewer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "105870",
      "postDate": "01/27/2016 04:46:43",
      "content": "<p>That is the expected output you get when installing with pip.</p>",
      "rawMarkdown": "That is the expected output you get when installing with pip.",
      "votes": null
    },
    {
      "id": "105882",
      "postDate": "01/27/2016 06:39:09",
      "content": "<p>nn.py&quot;, line 177, in _create_iter_funcs\n    for input_layer in input_layers]\nAttributeError: 'module' object has no attribute 'Param'</p>\n\n<p>This error seems linked with theano?</p>",
      "rawMarkdown": "nn.py\", line 177, in _create_iter_funcs\r\n    for input_layer in input_layers]\r\nAttributeError: 'module' object has no attribute 'Param'\r\n\r\nThis error seems linked with theano?",
      "votes": null
    },
    {
      "id": "105883",
      "postDate": "01/27/2016 06:40:23",
      "content": "<p>The details are </p>\n\n<pre><code>    X_inputs = [theano.Param(input_layer.input_var, name=input_layer.name)\n                for input_layer in input_layers]\n</code></pre>",
      "rawMarkdown": "The details are \r\n\r\n        X_inputs = [theano.Param(input_layer.input_var, name=input_layer.name)\r\n                    for input_layer in input_layers]",
      "votes": null
    },
    {
      "id": "105892",
      "postDate": "01/27/2016 08:13:54",
      "content": "<p>Traceback (most recent call last):\n  File &quot;train_nn.py&quot;, line 40, in \n    main()\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 610, in <strong>call</strong>\n    return self.main(*args, **kwargs)\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 590, in main\n    rv = self.invoke(ctx)\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 782, in invoke\n    return ctx.invoke(self.callback, **ctx.params)\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 416, in invoke\n    return callback(*args, **kwargs)\n  File &quot;train_nn.py&quot;, line 31, in main\n    net.load_params_from(weights_from)\n  File &quot;/home/xusongcen/DRD/codes/no2/src/nolearn-master/nolearn/lasagne/base.py&quot;, line 487, in load_params_from\n    self.initialize()\n  File &quot;/home/xusongcen/DRD/codes/no2/nn.py&quot;, line 136, in initialize\n    self.y_tensor_type,\n  File &quot;/home/xusongcen/DRD/codes/no2/nn.py&quot;, line 177, in _create_iter_funcs\n    for input_layer in input_layers]\nAttributeError: 'module' object has no attribute 'Param'</p>",
      "rawMarkdown": "Traceback (most recent call last):\r\n  File \"train_nn.py\", line 40, in <module>\r\n    main()\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 610, in __call__\r\n    return self.main(*args, **kwargs)\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 590, in main\r\n    rv = self.invoke(ctx)\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 782, in invoke\r\n    return ctx.invoke(self.callback, **ctx.params)\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 416, in invoke\r\n    return callback(*args, **kwargs)\r\n  File \"train_nn.py\", line 31, in main\r\n    net.load_params_from(weights_from)\r\n  File \"/home/xusongcen/DRD/codes/no2/src/nolearn-master/nolearn/lasagne/base.py\", line 487, in load_params_from\r\n    self.initialize()\r\n  File \"/home/xusongcen/DRD/codes/no2/nn.py\", line 136, in initialize\r\n    self.y_tensor_type,\r\n  File \"/home/xusongcen/DRD/codes/no2/nn.py\", line 177, in _create_iter_funcs\r\n    for input_layer in input_layers]\r\nAttributeError: 'module' object has no attribute 'Param'",
      "votes": null
    },
    {
      "id": "106203",
      "postDate": "01/29/2016 02:35:30",
      "content": "<p>Hi there, when I checked the definition of the network A or B, I noticed that the last layer was defined as:</p>\n\n<p>(DenseLayer, {'num_units': 1}),</p>\n\n<p>does this mean the number of outputs of both networks A,B are just one? instead of 5-way classification ?</p>",
      "rawMarkdown": "Hi there, when I checked the definition of the network A or B, I noticed that the last layer was defined as:\r\n\r\n (DenseLayer, {'num_units': 1}),\r\n\r\ndoes this mean the number of outputs of both networks A,B are just one? instead of 5-way classification ?",
      "votes": null
    },
    {
      "id": "106210",
      "postDate": "01/29/2016 03:43:27",
      "content": "<p>Maybe it would work if you install the following theano commit instead:</p>\n\n<pre><code>pip install git+https://github.com/Theano/Theano.git@dfb2730348d05f6aadd116ce492e836a4c0ba6d6\n</code></pre>\n\n<p>It does regression instead of classification and the output is then thresholded at 0.5, 1.5, 2.5, 3.5 to get integer levels for predictions.</p>",
      "rawMarkdown": "Maybe it would work if you install the following theano commit instead:\r\n\r\n    pip install git+https://github.com/Theano/Theano.git@dfb2730348d05f6aadd116ce492e836a4c0ba6d6\r\n\r\nIt does regression instead of classification and the output is then thresholded at 0.5, 1.5, 2.5, 3.5 to get integer levels for predictions.",
      "votes": null
    },
    {
      "id": "127927",
      "postDate": "07/16/2016 09:13:10",
      "content": "<p>Hi!\n   Following your solution,I run the code &quot;python train_nn.py --cnf configs/c_512_5x5_32.py --weights_from weights/c_256_5x5_32/weights_final.pkl&quot;, but I get the warning like &quot;Loaded parameters to layer 'conv2ddnn1' (shape 32x3x5x5).\nCould not load parameters to layer 'conv2ddnn1' because shapes did not match: 32x224x224 vs 32x112x112.\nLoaded parameters to layer 'conv2ddnn2' (shape 32x32x3x3).\nCould not load parameters to layer 'conv2ddnn2' because shapes did not match: 32x224x224 vs 32x112x112.\n...&quot;\nIs there something wrong for the shape of the trainning net initialized from the pretrained weights?\nHow should I solve this problem?\nThanks!</p>",
      "rawMarkdown": "Hi!\r\n   Following your solution,I run the code \"python train_nn.py --cnf configs/c_512_5x5_32.py --weights_from weights/c_256_5x5_32/weights_final.pkl\", but I get the warning like \"Loaded parameters to layer 'conv2ddnn1' (shape 32x3x5x5).\r\nCould not load parameters to layer 'conv2ddnn1' because shapes did not match: 32x224x224 vs 32x112x112.\r\nLoaded parameters to layer 'conv2ddnn2' (shape 32x32x3x3).\r\nCould not load parameters to layer 'conv2ddnn2' because shapes did not match: 32x224x224 vs 32x112x112.\r\n...\"\r\nIs there something wrong for the shape of the trainning net initialized from the pretrained weights?\r\nHow should I solve this problem?\r\nThanks!",
      "votes": null
    },
    {
      "id": "135363",
      "postDate": "09/13/2016 12:52:38",
      "content": "<p>Congratulations Team o_O! @skying This is a shape mismatch between bias matrices because the biases are untied. You could possibly switch from untied biases to regular that would be identical in all 3 training phases. I am also curious how Team o_O handled it? Thanks!</p>",
      "rawMarkdown": "Congratulations Team o_O! @skying This is a shape mismatch between bias matrices because the biases are untied. You could possibly switch from untied biases to regular that would be identical in all 3 training phases. I am also curious how Team o_O handled it? Thanks!",
      "votes": null
    },
    {
      "id": "146846",
      "postDate": "11/27/2016 09:14:04",
      "content": "<p>@Hi mathis\nthanks a lot for sending your code\nhow can I  use that with corei7 CPU only?</p>",
      "rawMarkdown": "Hi mathis\r\nthanks a lot for sending your code\r\nhow can I  use that with corei7 CPU only?",
      "votes": null
    },
    {
      "id": "166535",
      "postDate": "03/10/2017 02:13:08",
      "content": "<p>yeah,I know.The description about the shape dismatch is about the \"b\" ,becasue in the code ,Conv parameters set the \"united_bias\" to \"True\",which means different values in each postion in each channel.So,the \"bias\" cannnot use the pretratined value,its initial value is set by \"init.Const(0.05\").I think the bias weight a little in the all huge parameters,so it does not care when it cannot be loaded.As usual,we set the bias to 0 in convolution networks.And thanks for you help!</p>",
      "rawMarkdown": "yeah,I know.The description about the shape dismatch is about the \"b\" ,becasue in the code ,Conv parameters set the \"united_bias\" to \"True\",which means different values in each postion in each channel.So,the \"bias\" cannnot use the pretratined value,its initial value is set by \"init.Const(0.05\").I think the bias weight a little in the all huge parameters,so it does not care when it cannot be loaded.As usual,we set the bias to 0 in convolution networks.And thanks for you help!",
      "votes": null
    },
    {
      "id": "168090",
      "postDate": "03/16/2017 07:56:19",
      "content": "<p>@Mathis Antony  HI,I am so excited to read your Solution Summary, may I ask a question? are you opening your source code?would you mind sending me your source code, I use another network to classify the diabetic retinopathy but get very low accuracy.thanks a lot</p>",
      "rawMarkdown": "Mathis Antony  HI,I am so excited to read your Solution Summary, may I ask a question? are you opening your source code?would you mind sending me your source code, I use another network to classify the diabetic retinopathy but get very low accuracy.thanks a lot",
      "votes": null
    },
    {
      "id": "176365",
      "postDate": "04/20/2017 07:28:55",
      "content": "<p>First I want to thank you for this post.\nWhen i try your solution, I got this error :</p>\n\n<p>The 'Objective' class is no longer supported, please use 'nolearn.lasagne.objective' or similar.</p>\n\n<p>Can't find the solutions, any suggestions please?</p>",
      "rawMarkdown": "First I want to thank you for this post.\nWhen i try your solution, I got this error :\n\nThe 'Objective' class is no longer supported, please use 'nolearn.lasagne.objective' or similar.\n\nCan't find the solutions, any suggestions please?",
      "votes": null
    },
    {
      "id": "183720",
      "postDate": "05/19/2017 02:06:36",
      "content": "<p>hi, @mathis ,thanks for your post. I am a little confuse about the followings:\n    0.2 with batch sampled such that classes are balanced\n    0.5 with batch sampled uniformly from all images (shuffled)\ncan you help me?</p>",
      "rawMarkdown": "hi, @mathis ,thanks for your post. I am a little confuse about the followings:\n    0.2 with batch sampled such that classes are balanced\n    0.5 with batch sampled uniformly from all images (shuffled)\ncan you help me?",
      "votes": null
    },
    {
      "id": "373398",
      "postDate": "08/21/2018 09:39:01",
      "content": "<p>First I would like to say thank you to Agost Biro. His answer on stackoverflow helped me alot!</p>\n\n<p>Secondly, I somehow figured out why Team o_O chose 448 instead of other numbers. This is because in their augmentation settings,  their zoom range is 1/1.15 to 1.15, and in fact 512/1.15=445.2 ~ 448. 448 is a round number of 445.2 and it can be divided by 2 for multiple times(448 = 2^6*7).</p>\n\n<hr>\n\n<h2>Original Post :</h2>\n\n<p>Hello! I am new to CNNs and recently I am studying your solution.  </p>\n\n<p>However, I have some trouble understanding your networks. </p>\n\n<p>What does the \"units\" mean?</p>\n\n<p>What does the \"size\" mean？ Does it mean batch size? If so why would you choose 448? </p>\n\n<p>And why would you /2 every time you increase(x2) units? </p>\n\n<p>Any answer would be appreciated.</p>",
      "rawMarkdown": "First I would like to say thank you to Agost Biro. His answer on stackoverflow helped me alot!\n\nSecondly, I somehow figured out why Team o_O chose 448 instead of other numbers. This is because in their augmentation settings,  their zoom range is 1/1.15 to 1.15, and in fact 512/1.15=445.2 ~ 448. 448 is a round number of 445.2 and it can be divided by 2 for multiple times(448 = 2^6*7).\n\n--------------------------------------------------------------------------\nOriginal Post :\n--------------------------------------------------------------------------\nHello! I am new to CNNs and recently I am studying your solution.  \n\nHowever, I have some trouble understanding your networks. \n\nWhat does the \"units\" mean?\n\nWhat does the \"size\" mean？ Does it mean batch size? If so why would you choose 448? \n\nAnd why would you /2 every time you increase(x2) units? \n\nAny answer would be appreciated.",
      "votes": null
    },
    {
      "id": "373487",
      "postDate": "08/21/2018 12:40:51",
      "content": "<p>The \"size\" means the vertical and horizontal dimensions of the activations. I've posted an answer to your question on StackOverflow with more details: <a href=\"https://stackoverflow.com/a/51948851/2650622\">https://stackoverflow.com/a/51948851/2650622</a></p>",
      "rawMarkdown": "The \"size\" means the vertical and horizontal dimensions of the activations. I've posted an answer to your question on StackOverflow with more details: https://stackoverflow.com/a/51948851/2650622",
      "votes": null
    },
    {
      "id": "600495",
      "postDate": "08/16/2019 07:23:04",
      "content": "<p>Is this solution open-sourced? If so, please mention the link.</p>",
      "rawMarkdown": "Is this solution open-sourced? If so, please mention the link.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 87445,
      "author_name": "sedielem",
      "author_url": "",
      "post_date": "07/29/2015 15:39:33",
      "content": "<p>Congrats! Very interesting use of RMS pooling for spatial pooling -- in the NDSB we used it for pooling across rotations, where it clearly reduced overfitting compared to mean pooling or max pooling. Did you observe the same effect using it for spatial pooling? Or did it improve things in a different way?</p>\n\n<p>Regarding very leaky rectifiers: in my experience it can help to make the network even deeper. My intuition is that this compensates for the fact that these units are slightly less nonlinear than regular rectifiers (or leaky rectifiers with a small leak rate). Of course, having more layers also eats into your computational budget, so it's not always an option.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87465,
      "author_name": "juliandewit",
      "author_url": "",
      "post_date": "07/29/2015 18:02:45",
      "content": "<p>Hello congrats!\nA question... In the end (also due to time pressure) I also bootstapped new bigger nets with the first layer(s) of previous nets. Have you any idea if this results in less variation ?  </p>\n\n<p>Somehow the new nets did not add too much to the ensemble. Daniels Hammacks net which had a completely different architecture added a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87510,
      "author_name": "hrantkhachatrian",
      "author_url": "",
      "post_date": "07/29/2015 22:57:33",
      "content": "<p>Congratulations and thanks for the post!</p>\n\n<p>So you didn't use any &quot;unusual&quot; loss function to approximate the kappa score?! Kappa was explicitly used only in &quot;blending&quot; phase?</p>\n\n<p>And one more question: did you check what would be the score without &quot;blending&quot; phase?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87525,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "07/30/2015 01:31:09",
      "content": "<p>Thanks guys.</p>\n\n<p>@sedielem At this point I can't say with confidence if it made any difference at all. Around 2 months ago I tried replacing all MaxPool layers with RMSPool layers but that didn't work very well. But when using RMS pooling only in the last pooling layer the results were similar to using only max pooling layers and from then on I just stuck with it. It would be interesting to run it once more without the RMS pooling layer but the entire process of training one network takes 3-4 days.  I'll let you know if I get around to it. I also tried starting with leaky rate 0.5 and then gradually decreasing it down to zero using a shared variable. During training this worked great, but when trying to extract features or make predictions I noticed I was unable to affect the results by changing the leaky rate. Not knowing what was going on at this point I didn't pursue it any further.</p>\n\n<p>One thing I wanted to look at but didn't get around to yet was how much variance there was in the trained filter weights within each layer for different training strategies and rectifiers. I was also wondering if it's possible to apply a kind of penalty that &quot;forces&quot; the filters within each conv layer to become more diverse. Hoping this could help speed up training or allow smaller networks with similar performance. I trained my first neural net around the beginning of this year so I'm not sure how much sense this all makes.</p>\n\n<p>@julian I didn't try to use the same weights to bootstrap larger nets with different architecture so I can't comment on that. But I can give you some numbers of how much (or little) we gained from ensembling</p>\n\n<pre><code>net    set of weights   blend iter  augment averages       public LB     private LB\nB              1            1              20                0.84803        0.83916  \nB              1            1              50                0.84963        0.83973\nA+B            1            1              50                0.85090        0.84291\nA+B            1           10              50                0.85201        0.84308 \nA+B            2            1              50                0.85425        0.84379\nA+B            2           10              50                0.85199        0.84425\nA+B            3            1              50                0.85339        0.84479\nA+B            2            1              50                0.85316        0.84674 *\n(*in the last row thresholds were optimized for kappa on validation set)\n</code></pre>\n\n<p>@Hrant I did try using softmax classification in the beginning but the results weren't encouraging. I wondered quite a bit about how one could implement something close to kappa in theano but I never figured it out. At the end of the training phase both networks had a kappa score of about 0.80 (using info from one eye only). When feeding the output of the RMSPool layer (without averaging over different augmentations or blending patient eyes) into the blend network using only the features for one eye kappa is 0.805. When doing the same but using features averaged over 50 pseudo random augmentations kappa is 0.812. Then adding the standard deviation of the features brings kappa to 0.816. These numbers are however subject to some fluctuations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87558,
      "author_name": "jeffreydf",
      "author_url": "",
      "post_date": "07/30/2015 08:11:33",
      "content": "<p>I also tried RMS-pooling in the last layer only and had similar experiences (about the same score, no spectacular differences at a quick glance). I was planning on using some networks with RMS-pooling in an ensemble but other things came up.</p>\n\n<p>I think most of your gains, relative to mine and some others, came from the blending. I didn't have enough time to get something working and it seemed that it would be hard to regularise. </p>\n\n<p>Congratulations, btw!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87587,
      "author_name": "vinokurov",
      "author_url": "",
      "post_date": "07/30/2015 13:15:49",
      "content": "<p>Congratulations!  Many thanks for the post! I have a question. Could you give me link with explanation &quot;Krizhevsky color augmentation (gaussian)&quot;?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87592,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "07/30/2015 13:49:32",
      "content": "<p>@Jeffrey, I agree it's probably what made up for most of the difference in score. It took some time to get the get the learning rate, L1 and L2 factors and most importantly resampling strategy to work well together. What I quite liked about it was that I could iterate on it rapidly without having to wait days to see results because it only takes a few minutes to train the blending network. When we trained a better/different network for feature extraction at a later point most of the improvements made to the blending network would persist.</p>\n\n<p>@Andrey, I think it originally appeared in the <a href=\"http://www.cs.toronto.edu/~fritz/absps/imagenet.pdf\">Imagenet Paper</a> in section 4.1. I used standard deviation 0.5 for net A and standard deviation 0.25 for net B as opposed to 0.1 in the paper. I guess our numbers are rather high but the images in this dataset looked like they vary a lot in terms of color composition and I was trying to take this into account. I'm uncertain if adding the color augmentation helped in the end but I figured even if didn't improve the results there's a good chance it would help against overfitting so we kept it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87597,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "07/30/2015 14:26:13",
      "content": "<p>@Mathis Antony, Congrats!\nWhat is feature pool you mentioned, maxout? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87599,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "07/30/2015 14:51:04",
      "content": "<p>@old-ufo thanks. Yes it's maxout, pooled over 2 units. I updated the summary.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88056,
      "author_name": "alishakiba",
      "author_url": "",
      "post_date": "08/03/2015 06:02:20",
      "content": "<p>Thanks for the explanations. Where can we get the source code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88058,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "08/03/2015 06:21:39",
      "content": "<p>@Ali we haven't released it yet. As far as I understand we'll first submit it to kaggle and when we fulfill all the requirements release it on github.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105863,
      "author_name": "peterxu1986",
      "author_url": "",
      "post_date": "01/27/2016 03:45:24",
      "content": "<p>I got this error</p>\n\n<p>from lasagne.objectives import Objective\nImportError: cannot import name Objective</p>\n\n<p>Can't find the solutions, any suggestions please?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105865,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "01/27/2016 04:01:22",
      "content": "<p>Newer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105869,
      "author_name": "peterxu1986",
      "author_url": "",
      "post_date": "01/27/2016 04:45:23",
      "content": "<p>Thanks, however the requirement.txt will give</p>\n\n<p>Cloning <a href=\"https://github.com/benanne/Lasagne.git\">https://github.com/benanne/Lasagne.git</a> (to 9f591a5f3a192028df9947ba1e4903b3b46e8fe0) to ./src/lasagne-master\n  Could not find a tag or branch '9f591a5f3a192028df9947ba1e4903b3b46e8fe0', assuming commit.</p>\n\n<p>[quote=Mathis Antony;105865]</p>\n\n<p>Newer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105870,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "01/27/2016 04:46:43",
      "content": "<p>That is the expected output you get when installing with pip.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105882,
      "author_name": "peterxu1986",
      "author_url": "",
      "post_date": "01/27/2016 06:39:09",
      "content": "<p>nn.py&quot;, line 177, in _create_iter_funcs\n    for input_layer in input_layers]\nAttributeError: 'module' object has no attribute 'Param'</p>\n\n<p>This error seems linked with theano?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105883,
      "author_name": "peterxu1986",
      "author_url": "",
      "post_date": "01/27/2016 06:40:23",
      "content": "<p>The details are </p>\n\n<pre><code>    X_inputs = [theano.Param(input_layer.input_var, name=input_layer.name)\n                for input_layer in input_layers]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105892,
      "author_name": "peterxu1986",
      "author_url": "",
      "post_date": "01/27/2016 08:13:54",
      "content": "<p>Traceback (most recent call last):\n  File &quot;train_nn.py&quot;, line 40, in \n    main()\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 610, in <strong>call</strong>\n    return self.main(*args, **kwargs)\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 590, in main\n    rv = self.invoke(ctx)\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 782, in invoke\n    return ctx.invoke(self.callback, **ctx.params)\n  File &quot;/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py&quot;, line 416, in invoke\n    return callback(*args, **kwargs)\n  File &quot;train_nn.py&quot;, line 31, in main\n    net.load_params_from(weights_from)\n  File &quot;/home/xusongcen/DRD/codes/no2/src/nolearn-master/nolearn/lasagne/base.py&quot;, line 487, in load_params_from\n    self.initialize()\n  File &quot;/home/xusongcen/DRD/codes/no2/nn.py&quot;, line 136, in initialize\n    self.y_tensor_type,\n  File &quot;/home/xusongcen/DRD/codes/no2/nn.py&quot;, line 177, in _create_iter_funcs\n    for input_layer in input_layers]\nAttributeError: 'module' object has no attribute 'Param'</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 106203,
      "author_name": "peterxu1986",
      "author_url": "",
      "post_date": "01/29/2016 02:35:30",
      "content": "<p>Hi there, when I checked the definition of the network A or B, I noticed that the last layer was defined as:</p>\n\n<p>(DenseLayer, {'num_units': 1}),</p>\n\n<p>does this mean the number of outputs of both networks A,B are just one? instead of 5-way classification ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 106210,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "01/29/2016 03:43:27",
      "content": "<p>Maybe it would work if you install the following theano commit instead:</p>\n\n<pre><code>pip install git+https://github.com/Theano/Theano.git@dfb2730348d05f6aadd116ce492e836a4c0ba6d6\n</code></pre>\n\n<p>It does regression instead of classification and the output is then thresholded at 0.5, 1.5, 2.5, 3.5 to get integer levels for predictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127927,
      "author_name": "liguo888",
      "author_url": "",
      "post_date": "07/16/2016 09:13:10",
      "content": "<p>Hi!\n   Following your solution,I run the code &quot;python train_nn.py --cnf configs/c_512_5x5_32.py --weights_from weights/c_256_5x5_32/weights_final.pkl&quot;, but I get the warning like &quot;Loaded parameters to layer 'conv2ddnn1' (shape 32x3x5x5).\nCould not load parameters to layer 'conv2ddnn1' because shapes did not match: 32x224x224 vs 32x112x112.\nLoaded parameters to layer 'conv2ddnn2' (shape 32x32x3x3).\nCould not load parameters to layer 'conv2ddnn2' because shapes did not match: 32x224x224 vs 32x112x112.\n...&quot;\nIs there something wrong for the shape of the trainning net initialized from the pretrained weights?\nHow should I solve this problem?\nThanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 135363,
      "author_name": "wglogowski",
      "author_url": "",
      "post_date": "09/13/2016 12:52:38",
      "content": "<p>Congratulations Team o_O! @skying This is a shape mismatch between bias matrices because the biases are untied. You could possibly switch from untied biases to regular that would be identical in all 3 training phases. I am also curious how Team o_O handled it? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 166535,
          "author_name": "liguo888",
          "author_url": "",
          "post_date": "03/10/2017 02:13:08",
          "content": "<p>yeah,I know.The description about the shape dismatch is about the \"b\" ,becasue in the code ,Conv parameters set the \"united_bias\" to \"True\",which means different values in each postion in each channel.So,the \"bias\" cannnot use the pretratined value,its initial value is set by \"init.Const(0.05\").I think the bias weight a little in the all huge parameters,so it does not care when it cannot be loaded.As usual,we set the bias to 0 in convolution networks.And thanks for you help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 146846,
      "author_name": "miladz6868",
      "author_url": "",
      "post_date": "11/27/2016 09:14:04",
      "content": "<p>@Hi mathis\nthanks a lot for sending your code\nhow can I  use that with corei7 CPU only?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 168090,
      "author_name": "yufengli",
      "author_url": "",
      "post_date": "03/16/2017 07:56:19",
      "content": "<p>@Mathis Antony  HI,I am so excited to read your Solution Summary, may I ask a question? are you opening your source code?would you mind sending me your source code, I use another network to classify the diabetic retinopathy but get very low accuracy.thanks a lot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 176365,
      "author_name": "madiskou",
      "author_url": "",
      "post_date": "04/20/2017 07:28:55",
      "content": "<p>First I want to thank you for this post.\nWhen i try your solution, I got this error :</p>\n\n<p>The 'Objective' class is no longer supported, please use 'nolearn.lasagne.objective' or similar.</p>\n\n<p>Can't find the solutions, any suggestions please?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 183720,
      "author_name": "w7256037",
      "author_url": "",
      "post_date": "05/19/2017 02:06:36",
      "content": "<p>hi, @mathis ,thanks for your post. I am a little confuse about the followings:\n    0.2 with batch sampled such that classes are balanced\n    0.5 with batch sampled uniformly from all images (shuffled)\ncan you help me?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 373398,
      "author_name": "fcterrydr",
      "author_url": "",
      "post_date": "08/21/2018 09:39:01",
      "content": "<p>First I would like to say thank you to Agost Biro. His answer on stackoverflow helped me alot!</p>\n\n<p>Secondly, I somehow figured out why Team o_O chose 448 instead of other numbers. This is because in their augmentation settings,  their zoom range is 1/1.15 to 1.15, and in fact 512/1.15=445.2 ~ 448. 448 is a round number of 445.2 and it can be divided by 2 for multiple times(448 = 2^6*7).</p>\n\n<hr>\n\n<h2>Original Post :</h2>\n\n<p>Hello! I am new to CNNs and recently I am studying your solution.  </p>\n\n<p>However, I have some trouble understanding your networks. </p>\n\n<p>What does the \"units\" mean?</p>\n\n<p>What does the \"size\" mean？ Does it mean batch size? If so why would you choose 448? </p>\n\n<p>And why would you /2 every time you increase(x2) units? </p>\n\n<p>Any answer would be appreciated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 373487,
          "author_name": "agostbiro",
          "author_url": "",
          "post_date": "08/21/2018 12:40:51",
          "content": "<p>The \"size\" means the vertical and horizontal dimensions of the activations. I've posted an answer to your question on StackOverflow with more details: <a href=\"https://stackoverflow.com/a/51948851/2650622\">https://stackoverflow.com/a/51948851/2650622</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 600495,
      "author_name": "rajireddya",
      "author_url": "",
      "post_date": "08/16/2019 07:23:04",
      "content": "<p>Is this solution open-sourced? If so, please mention the link.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "87432": "First I'd like to thank team mate Stephan, the hosts and everyone \r\nelse who joined and congratulate the other winners! \r\n\r\nBelow is a brief solution summary while things are fresh before we get\r\ntime to do a proper writeup and publish our code.\r\n\r\n## Team o_O Solution Summary\r\n\r\nNeural networks trained using [lasagne](https://github.com/Lasagne/Lasagne)\r\nand [nolearn](https://github.com/dnouri/nolearn). Both libraries are awesome, \r\nwell documented and easy to get started with.\r\n\r\nA lot of inspiration and some code was taken from winners of\r\nthe national data science bowl competition. Thanks!\r\n\r\n- [≋ Deep Sea ≋](https://github.com/benanne/kaggle-ndsb): \r\n  data augmentation, RMS pooling, pseudo random augmentation averaging \r\n  (used for feature extraction here)\r\n- [Happy Lantern Festival](https://www.kaggle.com/c/datasciencebowl/forums/t/13166/happy-lantern-festival-report-and-code)\r\n  Capacity and coverage analysis of conv nets (it's built into nolearn)\r\n\r\n### Preprocessing\r\n- We selected the smallest rectangle that contained the entire eye.\r\n- When this didn't work well (dark images) or the image was already\r\n  almost a square we defaulted to selecting the center square.\r\n- Resized to 512, 256, 128 pixel squares.\r\n- Before the data augmentation we scaled each channel to have zero mean \r\n  and unit variance.\r\n\r\n### Augmentation\r\n- 360 degrees rotation, translation, scaling, stretching (all uniform)\r\n- Krizhevsky color augmentation (gaussian)\r\n\r\nThese augmentations were always applied.\r\n\r\n### Network Configurations\r\n\r\n                                net A       |          net B\r\n                units   filter stride size  |  filter stride size\r\n     1 Input                           448  |     4           448\r\n     2 Conv        32     5       2    224  |     4     2     224\r\n     3 Conv        32     3            224  |     4           225 \r\n     4 MaxPool            3       2    111  |     3     2     112 \r\n     5 Conv        64     5       2     56  |     4     2      56\r\n     6 Conv        64     3             56  |     4            57\r\n     7 Conv        64     3             56  |     4            56\r\n     8 MaxPool            3       2     27  |     3     2      27\r\n     9 Conv       128     3             27  |     4            28\r\n    10 Conv       128     3             27  |     4            27\r\n    11 Conv       128     3             27  |     4            28\r\n    12 MaxPool            3       2     13  |     3     2      13\r\n    13 Conv       256     3             13  |     4            14\r\n    14 Conv       256     3             13  |     4            13\r\n    15 Conv       256     3             13  |     4            14\r\n    16 MaxPool            3       2      6  |     3     2       6\r\n    17 Conv       512     3              6  |     4             5\r\n    18 Conv       512     3              6  |   n/a           n/a\r\n    19 RMSPool            3       3      2  |     3     2       2\r\n    20 Dropout\r\n    21 Dense     1024\r\n    22 Maxout    512\r\n    23 Dropout\r\n    24 Dense     1024\r\n    25 Maxout    512\r\n\r\n- Leaky (0.01) rectifier units following each conv and dense layer.\r\n- L2 weight decay factor 0.0005\r\n- Mean squared error objective.\r\n- Untied biases.\r\n\r\n### Training\r\n- We used 10% of the patients as validation set.\r\n- 128 px images -> layers 1 - 11 and 20 to 25.\r\n- 256 px images -> layers 1 - 15 and 20 to 25. Weights of layer 1 - 11 initialized \r\n  with weights from above.\r\n- 512 px images -> all layers. Weights of layers 1 - 15 initialized with\r\n  weights from above.\r\n- Training started with resampling such that all classes were present in equal \r\n  fractions. We then gradually decreased the balancing after each epoch to\r\n  arrive at final \"resampling weights\" of 1 for class 0 and 2 for the other \r\n  classes.\r\n- Nesterov momentum (0.9) with fixed learning rate schedule over 250 epochs.\r\n    + epoch 0: 0.003\r\n    + epoch 150: 0.0003\r\n    + epoch 220: 0.00003\r\n    + For the nets for 256 and 128 pixels images we stopped training already \r\n      after 200 epochs.\r\n\r\nModels were trained on a GTX 970 and a GTX 980Ti (highly recommended) with batch\r\nsizes small enough to fit into memory (usually 24 to 48 for the large networks\r\nand 128 for the smaller ones).\r\n\r\n### \"Per Patient\" Blend\r\nExtracted mean and standard deviation of RMSPool layer for 50 pseudo random \r\naugmentations for three sets of weights (best validation score, best kappa,\r\nfinal weights) for net A and B.\r\n\r\nFor each eye (or patient) used the following as input features for blending,\r\n\r\n    [this_eye_mean, other_eye_mean, this_eye_stddev, other_eye_stddev, left_eye_indicator]\r\n\r\nWe standardized all features to have zero mean and unit variance and used them to \r\ntrain a network of shape,\r\n\r\n    Input        8193\r\n    Dense          32\r\n    Maxout         16\r\n    Dense          32\r\n    Maxout         16\r\n\r\n- L1 regularization (2e-5) on the first layer and L2 regularization (0.005) \r\n  everywhere.\r\n- [Adam Updates](http://arxiv.org/abs/1412.6980) with fixed learning rate\r\n  schedule over 100 epochs.\r\n  + epoch  0: 5e-4\r\n  + epoch 60: 5e-5\r\n  + epoch 80: 5e-6\r\n  + epoch 90: 5e-7\r\n- Mean squared error objective.\r\n- Batches of size 128, replace batch with probability\r\n  + 0.2 with batch sampled such that classes are balanced\r\n  + 0.5 with batch sampled uniformly from all images (shuffled)\r\n\r\nThe mean output of the six blend networks (conv nets A and B, 3 sets of\r\nweights each) was then thresholded at `[0.5, 1.5, 2.5, 3.5]` to get \r\ninteger levels for submission.\r\n\r\n### Notes\r\nI just noticed that we achieved a slightly higher private LB score for a \r\nsubmission where we optimized the thresholds on the validation set to\r\nmaximize kappa. But this didn't work as well on the public LB and we \r\ndidn't pursue it any further or select it for scoring at the end.\r\n\r\nWhen using very leaky (0.33) rectifier units we managed to train the\r\nnetworks for 512 px images directly and the MSE and kappa score were comparable\r\nwith what we achieved with leaky rectifiers. However after blending the features \r\nfor each patients over pseudo random augmentations the validation and LB kappa \r\nwere never quite as good as what we obtained with leaky rectifier units and\r\ninitialization from smaller nets. We went back to using 0.01 leaky rectifiers \r\nafterwards and gave up on training the larger networks from scratch.",
    "87445": "Congrats! Very interesting use of RMS pooling for spatial pooling -- in the NDSB we used it for pooling across rotations, where it clearly reduced overfitting compared to mean pooling or max pooling. Did you observe the same effect using it for spatial pooling? Or did it improve things in a different way?\r\n\r\nRegarding very leaky rectifiers: in my experience it can help to make the network even deeper. My intuition is that this compensates for the fact that these units are slightly less nonlinear than regular rectifiers (or leaky rectifiers with a small leak rate). Of course, having more layers also eats into your computational budget, so it's not always an option.",
    "87465": "Hello congrats!\r\nA question... In the end (also due to time pressure) I also bootstapped new bigger nets with the first layer(s) of previous nets. Have you any idea if this results in less variation ?  \r\n\r\nSomehow the new nets did not add too much to the ensemble. Daniels Hammacks net which had a completely different architecture added a lot.",
    "87510": "Congratulations and thanks for the post!\r\n\r\nSo you didn't use any \"unusual\" loss function to approximate the kappa score?! Kappa was explicitly used only in \"blending\" phase?\r\n\r\nAnd one more question: did you check what would be the score without \"blending\" phase?",
    "87525": "Thanks guys.\r\n\r\n@sedielem At this point I can't say with confidence if it made any difference at all. Around 2 months ago I tried replacing all MaxPool layers with RMSPool layers but that didn't work very well. But when using RMS pooling only in the last pooling layer the results were similar to using only max pooling layers and from then on I just stuck with it. It would be interesting to run it once more without the RMS pooling layer but the entire process of training one network takes 3-4 days.  I'll let you know if I get around to it. I also tried starting with leaky rate 0.5 and then gradually decreasing it down to zero using a shared variable. During training this worked great, but when trying to extract features or make predictions I noticed I was unable to affect the results by changing the leaky rate. Not knowing what was going on at this point I didn't pursue it any further.\r\n\r\nOne thing I wanted to look at but didn't get around to yet was how much variance there was in the trained filter weights within each layer for different training strategies and rectifiers. I was also wondering if it's possible to apply a kind of penalty that \"forces\" the filters within each conv layer to become more diverse. Hoping this could help speed up training or allow smaller networks with similar performance. I trained my first neural net around the beginning of this year so I'm not sure how much sense this all makes.\r\n\r\n@julian I didn't try to use the same weights to bootstrap larger nets with different architecture so I can't comment on that. But I can give you some numbers of how much (or little) we gained from ensembling\r\n\r\n    net    set of weights   blend iter  augment averages       public LB     private LB\r\n    B              1            1              20                0.84803        0.83916  \r\n    B              1            1              50                0.84963        0.83973\r\n    A+B            1            1              50                0.85090        0.84291\r\n    A+B            1           10              50                0.85201        0.84308 \r\n    A+B            2            1              50                0.85425        0.84379\r\n    A+B            2           10              50                0.85199        0.84425\r\n    A+B            3            1              50                0.85339        0.84479\r\n    A+B            2            1              50                0.85316        0.84674 *\r\n    (*in the last row thresholds were optimized for kappa on validation set)\r\n\r\n@Hrant I did try using softmax classification in the beginning but the results weren't encouraging. I wondered quite a bit about how one could implement something close to kappa in theano but I never figured it out. At the end of the training phase both networks had a kappa score of about 0.80 (using info from one eye only). When feeding the output of the RMSPool layer (without averaging over different augmentations or blending patient eyes) into the blend network using only the features for one eye kappa is 0.805. When doing the same but using features averaged over 50 pseudo random augmentations kappa is 0.812. Then adding the standard deviation of the features brings kappa to 0.816. These numbers are however subject to some fluctuations.",
    "87558": "I also tried RMS-pooling in the last layer only and had similar experiences (about the same score, no spectacular differences at a quick glance). I was planning on using some networks with RMS-pooling in an ensemble but other things came up.\r\n\r\nI think most of your gains, relative to mine and some others, came from the blending. I didn't have enough time to get something working and it seemed that it would be hard to regularise. \r\n\r\nCongratulations, btw!",
    "87587": "Congratulations!  Many thanks for the post! I have a question. Could you give me link with explanation \"Krizhevsky color augmentation (gaussian)\"?",
    "87592": "Jeffrey, I agree it's probably what made up for most of the difference in score. It took some time to get the get the learning rate, L1 and L2 factors and most importantly resampling strategy to work well together. What I quite liked about it was that I could iterate on it rapidly without having to wait days to see results because it only takes a few minutes to train the blending network. When we trained a better/different network for feature extraction at a later point most of the improvements made to the blending network would persist.\r\n\r\n@Andrey, I think it originally appeared in the [Imagenet Paper](http://www.cs.toronto.edu/~fritz/absps/imagenet.pdf) in section 4.1. I used standard deviation 0.5 for net A and standard deviation 0.25 for net B as opposed to 0.1 in the paper. I guess our numbers are rather high but the images in this dataset looked like they vary a lot in terms of color composition and I was trying to take this into account. I'm uncertain if adding the color augmentation helped in the end but I figured even if didn't improve the results there's a good chance it would help against overfitting so we kept it.",
    "87597": "Mathis Antony, Congrats!\r\nWhat is feature pool you mentioned, maxout?",
    "87599": "old-ufo thanks. Yes it's maxout, pooled over 2 units. I updated the summary.",
    "88056": "Thanks for the explanations. Where can we get the source code?",
    "88058": "Ali we haven't released it yet. As far as I understand we'll first submit it to kaggle and when we fulfill all the requirements release it on github.",
    "105863": "I got this error\r\n\r\nfrom lasagne.objectives import Objective\r\nImportError: cannot import name Objective\r\n\r\nCan't find the solutions, any suggestions please?",
    "105865": "Newer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.",
    "105869": "Thanks, however the requirement.txt will give\r\n\r\nCloning https://github.com/benanne/Lasagne.git (to 9f591a5f3a192028df9947ba1e4903b3b46e8fe0) to ./src/lasagne-master\r\n  Could not find a tag or branch '9f591a5f3a192028df9947ba1e4903b3b46e8fe0', assuming commit.\r\n\r\n\r\n\r\n[quote=Mathis Antony;105865]\r\n\r\nNewer versions of lasagne don't have the Objective class. You can use the lasagne version specified in requirements.txt , there is some more information in the README file.\r\n\r\n[/quote]",
    "105870": "That is the expected output you get when installing with pip.",
    "105882": "nn.py\", line 177, in _create_iter_funcs\r\n    for input_layer in input_layers]\r\nAttributeError: 'module' object has no attribute 'Param'\r\n\r\nThis error seems linked with theano?",
    "105883": "The details are \r\n\r\n        X_inputs = [theano.Param(input_layer.input_var, name=input_layer.name)\r\n                    for input_layer in input_layers]",
    "105892": "Traceback (most recent call last):\r\n  File \"train_nn.py\", line 40, in <module>\r\n    main()\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 610, in __call__\r\n    return self.main(*args, **kwargs)\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 590, in main\r\n    rv = self.invoke(ctx)\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 782, in invoke\r\n    return ctx.invoke(self.callback, **ctx.params)\r\n  File \"/home/liuchuanjian/anaconda/lib/python2.7/site-packages/click/core.py\", line 416, in invoke\r\n    return callback(*args, **kwargs)\r\n  File \"train_nn.py\", line 31, in main\r\n    net.load_params_from(weights_from)\r\n  File \"/home/xusongcen/DRD/codes/no2/src/nolearn-master/nolearn/lasagne/base.py\", line 487, in load_params_from\r\n    self.initialize()\r\n  File \"/home/xusongcen/DRD/codes/no2/nn.py\", line 136, in initialize\r\n    self.y_tensor_type,\r\n  File \"/home/xusongcen/DRD/codes/no2/nn.py\", line 177, in _create_iter_funcs\r\n    for input_layer in input_layers]\r\nAttributeError: 'module' object has no attribute 'Param'",
    "106203": "Hi there, when I checked the definition of the network A or B, I noticed that the last layer was defined as:\r\n\r\n (DenseLayer, {'num_units': 1}),\r\n\r\ndoes this mean the number of outputs of both networks A,B are just one? instead of 5-way classification ?",
    "106210": "Maybe it would work if you install the following theano commit instead:\r\n\r\n    pip install git+https://github.com/Theano/Theano.git@dfb2730348d05f6aadd116ce492e836a4c0ba6d6\r\n\r\nIt does regression instead of classification and the output is then thresholded at 0.5, 1.5, 2.5, 3.5 to get integer levels for predictions.",
    "127927": "Hi!\r\n   Following your solution,I run the code \"python train_nn.py --cnf configs/c_512_5x5_32.py --weights_from weights/c_256_5x5_32/weights_final.pkl\", but I get the warning like \"Loaded parameters to layer 'conv2ddnn1' (shape 32x3x5x5).\r\nCould not load parameters to layer 'conv2ddnn1' because shapes did not match: 32x224x224 vs 32x112x112.\r\nLoaded parameters to layer 'conv2ddnn2' (shape 32x32x3x3).\r\nCould not load parameters to layer 'conv2ddnn2' because shapes did not match: 32x224x224 vs 32x112x112.\r\n...\"\r\nIs there something wrong for the shape of the trainning net initialized from the pretrained weights?\r\nHow should I solve this problem?\r\nThanks!",
    "135363": "Congratulations Team o_O! @skying This is a shape mismatch between bias matrices because the biases are untied. You could possibly switch from untied biases to regular that would be identical in all 3 training phases. I am also curious how Team o_O handled it? Thanks!",
    "146846": "Hi mathis\r\nthanks a lot for sending your code\r\nhow can I  use that with corei7 CPU only?",
    "166535": "yeah,I know.The description about the shape dismatch is about the \"b\" ,becasue in the code ,Conv parameters set the \"united_bias\" to \"True\",which means different values in each postion in each channel.So,the \"bias\" cannnot use the pretratined value,its initial value is set by \"init.Const(0.05\").I think the bias weight a little in the all huge parameters,so it does not care when it cannot be loaded.As usual,we set the bias to 0 in convolution networks.And thanks for you help!",
    "168090": "Mathis Antony  HI,I am so excited to read your Solution Summary, may I ask a question? are you opening your source code?would you mind sending me your source code, I use another network to classify the diabetic retinopathy but get very low accuracy.thanks a lot",
    "176365": "First I want to thank you for this post.\nWhen i try your solution, I got this error :\n\nThe 'Objective' class is no longer supported, please use 'nolearn.lasagne.objective' or similar.\n\nCan't find the solutions, any suggestions please?",
    "183720": "hi, @mathis ,thanks for your post. I am a little confuse about the followings:\n    0.2 with batch sampled such that classes are balanced\n    0.5 with batch sampled uniformly from all images (shuffled)\ncan you help me?",
    "373398": "First I would like to say thank you to Agost Biro. His answer on stackoverflow helped me alot!\n\nSecondly, I somehow figured out why Team o_O chose 448 instead of other numbers. This is because in their augmentation settings,  their zoom range is 1/1.15 to 1.15, and in fact 512/1.15=445.2 ~ 448. 448 is a round number of 445.2 and it can be divided by 2 for multiple times(448 = 2^6*7).\n\n--------------------------------------------------------------------------\nOriginal Post :\n--------------------------------------------------------------------------\nHello! I am new to CNNs and recently I am studying your solution.  \n\nHowever, I have some trouble understanding your networks. \n\nWhat does the \"units\" mean?\n\nWhat does the \"size\" mean？ Does it mean batch size? If so why would you choose 448? \n\nAnd why would you /2 every time you increase(x2) units? \n\nAny answer would be appreciated.",
    "373487": "The \"size\" means the vertical and horizontal dimensions of the activations. I've posted an answer to your question on StackOverflow with more details: https://stackoverflow.com/a/51948851/2650622",
    "600495": "Is this solution open-sourced? If so, please mention the link."
  },
  "source": "meta"
}