{
  "id": 15801,
  "title": "Competition report (min-pooling) and thank you",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/15801",
  "author_name": "Benjamin Graham",
  "post_date": "2015-08-06T11:48:24.360000",
  "votes": 80,
  "comment_count": 9,
  "views": 7303,
  "content": "<p>Hello. Attached is my competition report. </p>\n\n<p>Many thanks to the competition organisers at Kaggle, EyePACS and the California Healthcare Foundation. Thank you also to everyone who posted interesting observations about the dataset on the forums.</p>\n\n<p>Some links:</p>\n\n<ul>\n<li><p>The SparseConvNet library (GPL):\n<a href=\"https://github.com/btgraham/SparseConvNet\">https://github.com/btgraham/SparseConvNet</a></p></li>\n<li><p>A Git branch for the competition with pre-/post-processing scripts, etc\n<a href=\"https://github.com/btgraham/SparseConvNet/tree/kaggle_Diabetic_Retinopathy_competition\">https://github.com/btgraham/SparseConvNet/tree/kaggle_Diabetic_Retinopathy_competition</a></p></li>\n<li><p>Fractional max-pooling\n<a href=\"http://arxiv.org/abs/1412.6071\">http://arxiv.org/abs/1412.6071</a></p></li>\n</ul>",
  "messages": [
    {
      "id": 3136033,
      "postDate": "2025-02-28T05:39:09.500Z",
      "content": "<p>whats the new github link? the old one doesnt open</p>",
      "rawMarkdown": "whats the new github link? the old one doesnt open",
      "votes": 1
    },
    {
      "id": 88655,
      "postDate": "2015-08-06T11:48:24.360Z",
      "content": "<p>Hello. Attached is my competition report. </p>\n\n<p>Many thanks to the competition organisers at Kaggle, EyePACS and the California Healthcare Foundation. Thank you also to everyone who posted interesting observations about the dataset on the forums.</p>\n\n<p>Some links:</p>\n\n<ul>\n<li><p>The SparseConvNet library (GPL):\n<a href=\"https://github.com/btgraham/SparseConvNet\">https://github.com/btgraham/SparseConvNet</a></p></li>\n<li><p>A Git branch for the competition with pre-/post-processing scripts, etc\n<a href=\"https://github.com/btgraham/SparseConvNet/tree/kaggle_Diabetic_Retinopathy_competition\">https://github.com/btgraham/SparseConvNet/tree/kaggle_Diabetic_Retinopathy_competition</a></p></li>\n<li><p>Fractional max-pooling\n<a href=\"http://arxiv.org/abs/1412.6071\">http://arxiv.org/abs/1412.6071</a></p></li>\n</ul>",
      "rawMarkdown": "Hello. Attached is my competition report. \r\n\r\nMany thanks to the competition organisers at Kaggle, EyePACS and the California Healthcare Foundation. Thank you also to everyone who posted interesting observations about the dataset on the forums.\r\n\r\nSome links:\r\n\r\n- The SparseConvNet library (GPL):\r\n  https://github.com/btgraham/SparseConvNet\r\n\r\n- A Git branch for the competition with pre-/post-processing scripts, etc\r\n  https://github.com/btgraham/SparseConvNet/tree/kaggle_Diabetic_Retinopathy_competition\r\n\r\n- Fractional max-pooling\r\n  http://arxiv.org/abs/1412.6071",
      "votes": 80
    },
    {
      "id": 370950,
      "postDate": "2018-08-15T18:03:10.057Z",
      "content": "<p>Hello Benjamin,</p>\n\n<p>I've been trying to reproduce your results <em>using your code as-is</em> and I haven't been able to complete the execution of it due to lack of memory in the GPU card. The very surprising thing is that I'm only using an small subset of this competition images (2000 images (1000 patients) for training, 400 for validation and 1000 for testing), and I am using a GTX 1080 Ti card. Images were scaled at 300 pixels (<a href=\"https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/Data/kaggleDiabeticRetinopathy/preprocessImages.py\">https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/Data/kaggleDiabeticRetinopathy/preprocessImages.py</a>).\nI'm using CUDA 9.0 with the four cuBLAS Patch Updates to date, cuDNN v7.2.1 for CUDA 9.0 on Ubuntu 16.04.</p>\n\n<p>One of the invocations to allocate memory in the GPU card (<code>cudaSafeCall(cudaMalloc((void**) &amp;d_vec, sizeof(t)*dAllocated));</code>) in <a href=\"https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/vectorCUDA.cu\">https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/vectorCUDA.cu</a> fails because it tries to allocate a chunk of memory larger than the available memory in the GPU card. This happens with one of the networks (<a href=\"https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/kaggleDiabetes1.cpp\">https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/kaggleDiabetes1.cpp</a>), and of course this error did not encourage me to train the other 2 networks. Of course I'm monitoring the memory usage of the card while running the code and it's in fact growing until it runs out.</p>\n\n<p>I would appreciate your help on this (I understand this may not be easy after 3 years...), I can provide whatever info in my development environment you need.</p>\n\n<p>I also take this opportunity to ask other people in the forum if anyone was able to successfully train the 3 networks from Benjamin's code.</p>\n\n<p>Thank you !\nPablo</p>",
      "rawMarkdown": "Hello Benjamin,\n\nI've been trying to reproduce your results *using your code as-is* and I haven't been able to complete the execution of it due to lack of memory in the GPU card. The very surprising thing is that I'm only using an small subset of this competition images (2000 images (1000 patients) for training, 400 for validation and 1000 for testing), and I am using a GTX 1080 Ti card. Images were scaled at 300 pixels (https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/Data/kaggleDiabeticRetinopathy/preprocessImages.py).\nI'm using CUDA 9.0 with the four cuBLAS Patch Updates to date, cuDNN v7.2.1 for CUDA 9.0 on Ubuntu 16.04.\n\nOne of the invocations to allocate memory in the GPU card (`cudaSafeCall(cudaMalloc((void**) &amp;d_vec, sizeof(t)*dAllocated));`) in https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/vectorCUDA.cu fails because it tries to allocate a chunk of memory larger than the available memory in the GPU card. This happens with one of the networks (https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/kaggleDiabetes1.cpp), and of course this error did not encourage me to train the other 2 networks. Of course I'm monitoring the memory usage of the card while running the code and it's in fact growing until it runs out.\n\nI would appreciate your help on this (I understand this may not be easy after 3 years...), I can provide whatever info in my development environment you need.\n\nI also take this opportunity to ask other people in the forum if anyone was able to successfully train the 3 networks from Benjamin's code.\n\nThank you !\nPablo",
      "votes": 1,
      "replies": [
        {
          "id": 812365,
          "postDate": "2020-04-18T17:34:03.477Z",
          "content": "<p>Hello <a href=\"/pablorios\">@pablorios</a>. The issue you have mentioned here generally relates to the scenario where the GPU can't handle the data you are currently feeding it. You may try to reduce the batch size a tad bit and give it another try. </p>",
          "rawMarkdown": "Hello @pablorios. The issue you have mentioned here generally relates to the scenario where the GPU can't handle the data you are currently feeding it. You may try to reduce the batch size a tad bit and give it another try. "
        }
      ]
    },
    {
      "id": 89062,
      "postDate": "2015-08-10T21:34:52.373Z",
      "content": "<p>Julian, Daniel: Hello. I am hoping to improve upon my competition solution so suggestions are welcome.</p>\n\n<p>Daniel:\n1) For this competition, I initialised the networks using the training data, at each layer scaling the Uniform[-1,+1] distribution so that the hidden units roughly had standard deviation 1.</p>\n\n<p>2) I found that ensembling did not add much. This surprised me as on things like CIFAR-10 and ImageNet, ensembling is massively useful.\n(I don't want to compare fractional max-pooling with VGG-style convnets on the basis of this competition as I didn't have enough time to be as thorough as I would have liked.)</p>\n\n<p>3) I don't think there is anything fundamentally wrong with very small batches. It is just a trade-off between training a smaller network for more epochs or a larger network for fewer.</p>\n\n<p>4) Up to a week. I used SparseConvNet because it's familiar and it implements fractional max-pooling. It is designed to handle sparse input data efficiently, which has not really relevant for this competition. For dense images, SparseConvNet cannot compete speed-wise with Caffe+cuDNN. But it is a work in progress, and I have high-hopes for optimizing it further.</p>",
      "rawMarkdown": "Julian, Daniel: Hello. I am hoping to improve upon my competition solution so suggestions are welcome.\r\n\r\nDaniel:\r\n1) For this competition, I initialised the networks using the training data, at each layer scaling the Uniform[-1,+1] distribution so that the hidden units roughly had standard deviation 1.\r\n\r\n2) I found that ensembling did not add much. This surprised me as on things like CIFAR-10 and ImageNet, ensembling is massively useful.\r\n(I don't want to compare fractional max-pooling with VGG-style convnets on the basis of this competition as I didn't have enough time to be as thorough as I would have liked.)\r\n\r\n3) I don't think there is anything fundamentally wrong with very small batches. It is just a trade-off between training a smaller network for more epochs or a larger network for fewer.\r\n\r\n4) Up to a week. I used SparseConvNet because it's familiar and it implements fractional max-pooling. It is designed to handle sparse input data efficiently, which has not really relevant for this competition. For dense images, SparseConvNet cannot compete speed-wise with Caffe+cuDNN. But it is a work in progress, and I have high-hopes for optimizing it further.\r\n\r\n",
      "votes": 5
    },
    {
      "id": 88851,
      "postDate": "2015-08-07T19:31:48.283Z",
      "content": "<p>Congrats Ben. </p>\n\n<p>Your approach was quite straightforward and fairly similar to ours. From what I can tell, it's more evidence that fractional max pooling is just better than regular max pooling. I have a couple of questions:</p>\n\n<p>1) How did you initialize your networks?</p>\n\n<p>2) How well do your networks perform before ensembling? Are the fractional max pooling ones better than the VGG style convnet?</p>\n\n<p>3) Is your software more memory-efficient than Theano? I have 12gb of GPU memory but I don't think I could have used Keras to train networks as big as yours without taking a huge hit on batch size!</p>\n\n<p>4) What were your training times like? I'm curious about seconds/image as well as how many epochs you trained for.</p>\n\n<p>Again, congrats. I wish now that I had tried images bigger than 256x256!</p>\n\n<p>Daniel</p>",
      "rawMarkdown": "Congrats Ben. \r\n\r\nYour approach was quite straightforward and fairly similar to ours. From what I can tell, it's more evidence that fractional max pooling is just better than regular max pooling. I have a couple of questions:\r\n\r\n1) How did you initialize your networks?\r\n\r\n2) How well do your networks perform before ensembling? Are the fractional max pooling ones better than the VGG style convnet?\r\n\r\n3) Is your software more memory-efficient than Theano? I have 12gb of GPU memory but I don't think I could have used Keras to train networks as big as yours without taking a huge hit on batch size!\r\n\r\n4) What were your training times like? I'm curious about seconds/image as well as how many epochs you trained for.\r\n\r\nAgain, congrats. I wish now that I had tried images bigger than 256x256!\r\n\r\nDaniel",
      "votes": 6
    },
    {
      "id": 88811,
      "postDate": "2015-08-07T12:42:27.280Z",
      "content": "<p>Congratulations Mr Graham.\nI really suffered from a bad case of fractional max pooling envy during the last part of the competition :P\nIn the last 2 weeks I got your software from github but somehow I could not get the net to score good. So I'll be studying where I did go wrong.</p>\n\n<p>That said.. I really enjoyed reading your papers and your code. You have great out of the box ideas. </p>\n\n<p>If we are all gonna Kumbaya ensembling perhaps we also have some nice additions.\nFirst I started out with a symptom localizer/counter but in the end a blackbox convnet worked better. However, the &quot;bloodspot counts&quot; still added about 1-2 kappa in our 2nd stage classiefier. We also used image size as a feature. Thinking that a score for a good image should weigh heavier than one for a bad image. Daniels net was also quite different with a big avg pooler in the end.</p>",
      "rawMarkdown": "Congratulations Mr Graham.\r\nI really suffered from a bad case of fractional max pooling envy during the last part of the competition :P\r\nIn the last 2 weeks I got your software from github but somehow I could not get the net to score good. So I'll be studying where I did go wrong.\r\n\r\nThat said.. I really enjoyed reading your papers and your code. You have great out of the box ideas. \r\n\r\nIf we are all gonna Kumbaya ensembling perhaps we also have some nice additions.\r\nFirst I started out with a symptom localizer/counter but in the end a blackbox convnet worked better. However, the \"bloodspot counts\" still added about 1-2 kappa in our 2nd stage classiefier. We also used image size as a feature. Thinking that a score for a good image should weigh heavier than one for a bad image. Daniels net was also quite different with a big avg pooler in the end.\r\n\r\n\r\n\r\n\r\n\r\n",
      "votes": 4
    },
    {
      "id": 88677,
      "postDate": "2015-08-06T14:54:49.887Z",
      "content": "<p>Congratulations Benjamin!</p>\n\n<p>I would be very interested to know how well an ensemble of the top x competitors would perform. It's a little tricky because of the discrete predictions but maybe a max voting ensemble of the top 10 competitors. I have attached my best submission: 0.84135 on public leaderboard,  0.82899 on private in case people want to fiddle around with it.</p>",
      "rawMarkdown": "Congratulations Benjamin!\r\n\r\nI would be very interested to know how well an ensemble of the top x competitors would perform. It's a little tricky because of the discrete predictions but maybe a max voting ensemble of the top 10 competitors. I have attached my best submission: 0.84135 on public leaderboard,\t0.82899 on private in case people want to fiddle around with it.",
      "votes": 4
    },
    {
      "id": 105230,
      "postDate": "2016-01-21T07:41:22.283Z",
      "content": "<p>Hi Graham,could you advice me on how to approach San Fransico Crime competition.\nI want to use pandas as data structure ,but I'm not strong in python. How do I preprocess the data.Do i have tokenize the text to vectors of zeros ones for the attributes?how do I treat dates.I'm more familiar with  numerical attributes.\nThanks alot\nUmar</p>",
      "rawMarkdown": "Hi Graham,could you advice me on how to approach San Fransico Crime competition.\r\nI want to use pandas as data structure ,but I'm not strong in python. How do I preprocess the data.Do i have tokenize the text to vectors of zeros ones for the attributes?how do I treat dates.I'm more familiar with  numerical attributes.\r\nThanks alot\r\nUmar",
      "replies": [
        {
          "id": 2691935,
          "postDate": "2024-03-11T15:08:01.497Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3136033,
      "author_name": "Fira",
      "author_url": "",
      "post_date": "2025-02-28T05:39:09.500000",
      "content": "<p>whats the new github link? the old one doesnt open</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 370950,
      "author_name": "Pablo Rios",
      "author_url": "",
      "post_date": "2018-08-15T18:03:10.057000",
      "content": "<p>Hello Benjamin,</p>\n\n<p>I've been trying to reproduce your results <em>using your code as-is</em> and I haven't been able to complete the execution of it due to lack of memory in the GPU card. The very surprising thing is that I'm only using an small subset of this competition images (2000 images (1000 patients) for training, 400 for validation and 1000 for testing), and I am using a GTX 1080 Ti card. Images were scaled at 300 pixels (<a href=\"https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/Data/kaggleDiabeticRetinopathy/preprocessImages.py\">https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/Data/kaggleDiabeticRetinopathy/preprocessImages.py</a>).\nI'm using CUDA 9.0 with the four cuBLAS Patch Updates to date, cuDNN v7.2.1 for CUDA 9.0 on Ubuntu 16.04.</p>\n\n<p>One of the invocations to allocate memory in the GPU card (<code>cudaSafeCall(cudaMalloc((void**) &amp;d_vec, sizeof(t)*dAllocated));</code>) in <a href=\"https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/vectorCUDA.cu\">https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/vectorCUDA.cu</a> fails because it tries to allocate a chunk of memory larger than the available memory in the GPU card. This happens with one of the networks (<a href=\"https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/kaggleDiabetes1.cpp\">https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/kaggleDiabetes1.cpp</a>), and of course this error did not encourage me to train the other 2 networks. Of course I'm monitoring the memory usage of the card while running the code and it's in fact growing until it runs out.</p>\n\n<p>I would appreciate your help on this (I understand this may not be easy after 3 years...), I can provide whatever info in my development environment you need.</p>\n\n<p>I also take this opportunity to ask other people in the forum if anyone was able to successfully train the 3 networks from Benjamin's code.</p>\n\n<p>Thank you !\nPablo</p>",
      "votes": 1,
      "replies": [
        {
          "id": 812365,
          "author_name": "Akib Shahriyar",
          "author_url": "",
          "post_date": "2020-04-18T17:34:03.477000",
          "content": "<p>Hello <a href=\"/pablorios\">@pablorios</a>. The issue you have mentioned here generally relates to the scenario where the GPU can't handle the data you are currently feeding it. You may try to reduce the batch size a tad bit and give it another try. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 89062,
      "author_name": "Benjamin Graham",
      "author_url": "",
      "post_date": "2015-08-10T21:34:52.373000",
      "content": "<p>Julian, Daniel: Hello. I am hoping to improve upon my competition solution so suggestions are welcome.</p>\n\n<p>Daniel:\n1) For this competition, I initialised the networks using the training data, at each layer scaling the Uniform[-1,+1] distribution so that the hidden units roughly had standard deviation 1.</p>\n\n<p>2) I found that ensembling did not add much. This surprised me as on things like CIFAR-10 and ImageNet, ensembling is massively useful.\n(I don't want to compare fractional max-pooling with VGG-style convnets on the basis of this competition as I didn't have enough time to be as thorough as I would have liked.)</p>\n\n<p>3) I don't think there is anything fundamentally wrong with very small batches. It is just a trade-off between training a smaller network for more epochs or a larger network for fewer.</p>\n\n<p>4) Up to a week. I used SparseConvNet because it's familiar and it implements fractional max-pooling. It is designed to handle sparse input data efficiently, which has not really relevant for this competition. For dense images, SparseConvNet cannot compete speed-wise with Caffe+cuDNN. But it is a work in progress, and I have high-hopes for optimizing it further.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 88851,
      "author_name": "dhammack",
      "author_url": "",
      "post_date": "2015-08-07T19:31:48.283000",
      "content": "<p>Congrats Ben. </p>\n\n<p>Your approach was quite straightforward and fairly similar to ours. From what I can tell, it's more evidence that fractional max pooling is just better than regular max pooling. I have a couple of questions:</p>\n\n<p>1) How did you initialize your networks?</p>\n\n<p>2) How well do your networks perform before ensembling? Are the fractional max pooling ones better than the VGG style convnet?</p>\n\n<p>3) Is your software more memory-efficient than Theano? I have 12gb of GPU memory but I don't think I could have used Keras to train networks as big as yours without taking a huge hit on batch size!</p>\n\n<p>4) What were your training times like? I'm curious about seconds/image as well as how many epochs you trained for.</p>\n\n<p>Again, congrats. I wish now that I had tried images bigger than 256x256!</p>\n\n<p>Daniel</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 88811,
      "author_name": "Julian de Wit",
      "author_url": "",
      "post_date": "2015-08-07T12:42:27.280000",
      "content": "<p>Congratulations Mr Graham.\nI really suffered from a bad case of fractional max pooling envy during the last part of the competition :P\nIn the last 2 weeks I got your software from github but somehow I could not get the net to score good. So I'll be studying where I did go wrong.</p>\n\n<p>That said.. I really enjoyed reading your papers and your code. You have great out of the box ideas. </p>\n\n<p>If we are all gonna Kumbaya ensembling perhaps we also have some nice additions.\nFirst I started out with a symptom localizer/counter but in the end a blackbox convnet worked better. However, the &quot;bloodspot counts&quot; still added about 1-2 kappa in our 2nd stage classiefier. We also used image size as a feature. Thinking that a score for a good image should weigh heavier than one for a bad image. Daniels net was also quite different with a big avg pooler in the end.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 88677,
      "author_name": "Jeffrey",
      "author_url": "",
      "post_date": "2015-08-06T14:54:49.887000",
      "content": "<p>Congratulations Benjamin!</p>\n\n<p>I would be very interested to know how well an ensemble of the top x competitors would perform. It's a little tricky because of the discrete predictions but maybe a max voting ensemble of the top 10 competitors. I have attached my best submission: 0.84135 on public leaderboard,  0.82899 on private in case people want to fiddle around with it.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 105230,
      "author_name": "AbdullahiUmar",
      "author_url": "",
      "post_date": "2016-01-21T07:41:22.283000",
      "content": "<p>Hi Graham,could you advice me on how to approach San Fransico Crime competition.\nI want to use pandas as data structure ,but I'm not strong in python. How do I preprocess the data.Do i have tokenize the text to vectors of zeros ones for the attributes?how do I treat dates.I'm more familiar with  numerical attributes.\nThanks alot\nUmar</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2691935,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-11T15:08:01.497000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3136033": "whats the new github link? the old one doesnt open",
    "88655": "Hello. Attached is my competition report. \r\n\r\nMany thanks to the competition organisers at Kaggle, EyePACS and the California Healthcare Foundation. Thank you also to everyone who posted interesting observations about the dataset on the forums.\r\n\r\nSome links:\r\n\r\n- The SparseConvNet library (GPL):\r\n  https://github.com/btgraham/SparseConvNet\r\n\r\n- A Git branch for the competition with pre-/post-processing scripts, etc\r\n  https://github.com/btgraham/SparseConvNet/tree/kaggle_Diabetic_Retinopathy_competition\r\n\r\n- Fractional max-pooling\r\n  http://arxiv.org/abs/1412.6071",
    "370950": "Hello Benjamin,\n\nI've been trying to reproduce your results *using your code as-is* and I haven't been able to complete the execution of it due to lack of memory in the GPU card. The very surprising thing is that I'm only using an small subset of this competition images (2000 images (1000 patients) for training, 400 for validation and 1000 for testing), and I am using a GTX 1080 Ti card. Images were scaled at 300 pixels (https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/Data/kaggleDiabeticRetinopathy/preprocessImages.py).\nI'm using CUDA 9.0 with the four cuBLAS Patch Updates to date, cuDNN v7.2.1 for CUDA 9.0 on Ubuntu 16.04.\n\nOne of the invocations to allocate memory in the GPU card (`cudaSafeCall(cudaMalloc((void**) &amp;d_vec, sizeof(t)*dAllocated));`) in https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/vectorCUDA.cu fails because it tries to allocate a chunk of memory larger than the available memory in the GPU card. This happens with one of the networks (https://github.com/btgraham/SparseConvNet/blob/kaggle_Diabetic_Retinopathy_competition/kaggleDiabetes1.cpp), and of course this error did not encourage me to train the other 2 networks. Of course I'm monitoring the memory usage of the card while running the code and it's in fact growing until it runs out.\n\nI would appreciate your help on this (I understand this may not be easy after 3 years...), I can provide whatever info in my development environment you need.\n\nI also take this opportunity to ask other people in the forum if anyone was able to successfully train the 3 networks from Benjamin's code.\n\nThank you !\nPablo",
    "89062": "Julian, Daniel: Hello. I am hoping to improve upon my competition solution so suggestions are welcome.\r\n\r\nDaniel:\r\n1) For this competition, I initialised the networks using the training data, at each layer scaling the Uniform[-1,+1] distribution so that the hidden units roughly had standard deviation 1.\r\n\r\n2) I found that ensembling did not add much. This surprised me as on things like CIFAR-10 and ImageNet, ensembling is massively useful.\r\n(I don't want to compare fractional max-pooling with VGG-style convnets on the basis of this competition as I didn't have enough time to be as thorough as I would have liked.)\r\n\r\n3) I don't think there is anything fundamentally wrong with very small batches. It is just a trade-off between training a smaller network for more epochs or a larger network for fewer.\r\n\r\n4) Up to a week. I used SparseConvNet because it's familiar and it implements fractional max-pooling. It is designed to handle sparse input data efficiently, which has not really relevant for this competition. For dense images, SparseConvNet cannot compete speed-wise with Caffe+cuDNN. But it is a work in progress, and I have high-hopes for optimizing it further.\r\n\r\n",
    "88851": "Congrats Ben. \r\n\r\nYour approach was quite straightforward and fairly similar to ours. From what I can tell, it's more evidence that fractional max pooling is just better than regular max pooling. I have a couple of questions:\r\n\r\n1) How did you initialize your networks?\r\n\r\n2) How well do your networks perform before ensembling? Are the fractional max pooling ones better than the VGG style convnet?\r\n\r\n3) Is your software more memory-efficient than Theano? I have 12gb of GPU memory but I don't think I could have used Keras to train networks as big as yours without taking a huge hit on batch size!\r\n\r\n4) What were your training times like? I'm curious about seconds/image as well as how many epochs you trained for.\r\n\r\nAgain, congrats. I wish now that I had tried images bigger than 256x256!\r\n\r\nDaniel",
    "88811": "Congratulations Mr Graham.\r\nI really suffered from a bad case of fractional max pooling envy during the last part of the competition :P\r\nIn the last 2 weeks I got your software from github but somehow I could not get the net to score good. So I'll be studying where I did go wrong.\r\n\r\nThat said.. I really enjoyed reading your papers and your code. You have great out of the box ideas. \r\n\r\nIf we are all gonna Kumbaya ensembling perhaps we also have some nice additions.\r\nFirst I started out with a symptom localizer/counter but in the end a blackbox convnet worked better. However, the \"bloodspot counts\" still added about 1-2 kappa in our 2nd stage classiefier. We also used image size as a feature. Thinking that a score for a good image should weigh heavier than one for a bad image. Daniels net was also quite different with a big avg pooler in the end.\r\n\r\n\r\n\r\n\r\n\r\n",
    "88677": "Congratulations Benjamin!\r\n\r\nI would be very interested to know how well an ensemble of the top x competitors would perform. It's a little tricky because of the discrete predictions but maybe a max voting ensemble of the top 10 competitors. I have attached my best submission: 0.84135 on public leaderboard,\t0.82899 on private in case people want to fiddle around with it.",
    "105230": "Hi Graham,could you advice me on how to approach San Fransico Crime competition.\r\nI want to use pandas as data structure ,but I'm not strong in python. How do I preprocess the data.Do i have tokenize the text to vectors of zeros ones for the attributes?how do I treat dates.I'm more familiar with  numerical attributes.\r\nThanks alot\r\nUmar"
  }
}