{
  "id": 18548,
  "title": "Keras Deep Learning tutorial (~0.0359)",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18548",
  "author_name": "Marko Jocic",
  "post_date": "2016-01-24T20:17:25.607000",
  "votes": 32,
  "comment_count": 121,
  "views": 43077,
  "content": "<p>Hi everyone,</p>\n\n<p>a Keras-based deep learning tutorial for ~0.0359 CRPS is available on:</p>\n\n<p><a href=\"https://github.com/jocicmarko/kaggle-dsb2-keras/\">https://github.com/jocicmarko/kaggle-dsb2-keras/</a></p>\n\n<p>Basically it is a conv net for linear regression task.\nThere wasn't much experimenting with net structure, hyper-parameters and image pre-processing, so these might be a good place to start.</p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
  "messages": [
    {
      "id": 105563,
      "postDate": "2016-01-24T20:17:25.607Z",
      "content": "<p>Hi everyone,</p>\n\n<p>a Keras-based deep learning tutorial for ~0.0359 CRPS is available on:</p>\n\n<p><a href=\"https://github.com/jocicmarko/kaggle-dsb2-keras/\">https://github.com/jocicmarko/kaggle-dsb2-keras/</a></p>\n\n<p>Basically it is a conv net for linear regression task.\nThere wasn't much experimenting with net structure, hyper-parameters and image pre-processing, so these might be a good place to start.</p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "Hi everyone,\r\n\r\na Keras-based deep learning tutorial for ~0.0359 CRPS is available on:\r\n\r\nhttps://github.com/jocicmarko/kaggle-dsb2-keras/\r\n\r\nBasically it is a conv net for linear regression task.\r\nThere wasn't much experimenting with net structure, hyper-parameters and image pre-processing, so these might be a good place to start.\r\n\r\nCheers,\r\n\r\nMarko",
      "votes": 32
    },
    {
      "id": 105662,
      "postDate": "2016-01-25T18:26:08.057Z",
      "content": "<p>@phunter</p>\n\n<p>From the top of my head, I think the main difference is that in our example the two models are trained basically at the same time (which doubles the memory usage), while MXnet example trains them one by one.</p>",
      "rawMarkdown": "@phunter\r\n\r\nFrom the top of my head, I think the main difference is that in our example the two models are trained basically at the same time (which doubles the memory usage), while MXnet example trains them one by one.",
      "votes": 6
    },
    {
      "id": 105608,
      "postDate": "2016-01-25T05:39:42.757Z",
      "content": "<p>@ertuka: yes, you should update your version of Keras, from the Keras repository:</p>\n\n<pre><code>pip install git+git://github.com/fchollet/keras.git --upgrade --no-deps\n</code></pre>\n\n<p>(may need to be prefixed with &quot;sudo&quot;)</p>",
      "rawMarkdown": "@ertuka: yes, you should update your version of Keras, from the Keras repository:\r\n\r\n    pip install git+git://github.com/fchollet/keras.git --upgrade --no-deps\r\n\r\n(may need to be prefixed with \"sudo\")",
      "votes": 3
    },
    {
      "id": 105658,
      "postDate": "2016-01-25T17:10:38.827Z",
      "content": "<p>[quote=fchollet;105655]</p>\n\n<p>This is a pretty ridiculous attempt at saying &quot;lol MXnet is faster than Keras&quot;. No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. </p>\n\n<p>Stay classy man...</p>\n\n<p>[/quote]</p>\n\n<p>Oh, no, I didn't mean MXnet was faster because of what what, why downvote me? @Marko Jocic said 128x128 was too large to fit in his GTX 770 video card, and I said I could run 128x128 with mxnet on GTX 960, an even slower video card, so @WD didn't have to worry about memory of running it. We ran on different cards different resolutions different CPUs for data augment etc, and there was no comparisons. </p>\n\n<p>Stay calm :-)</p>",
      "rawMarkdown": "[quote=fchollet;105655]\r\n\r\nThis is a pretty ridiculous attempt at saying \"lol MXnet is faster than Keras\". No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. \r\n\r\nStay classy man...\r\n\r\n[/quote]\r\n\r\nOh, no, I didn't mean MXnet was faster because of what what, why downvote me? @Marko Jocic said 128x128 was too large to fit in his GTX 770 video card, and I said I could run 128x128 with mxnet on GTX 960, an even slower video card, so @WD didn't have to worry about memory of running it. We ran on different cards different resolutions different CPUs for data augment etc, and there was no comparisons. \r\n\r\nStay calm :-)",
      "votes": 4
    },
    {
      "id": 105919,
      "postDate": "2016-01-27T12:46:11.080Z",
      "content": "<p>Hi!</p>\n\n<p>Thank you sharing your model!</p>\n\n<p>When I submit this model code, get the high errors(~0.077926).(same piotr happening)\nThis code run 150 epochs running.</p>\n\n<p>So I think this code is influented random state or other cause.</p>",
      "rawMarkdown": "Hi!\r\n\r\nThank you sharing your model!\r\n\r\nWhen I submit this model code, get the high errors(~0.077926).(same piotr happening)\r\nThis code run 150 epochs running.\r\n\r\nSo I think this code is influented random state or other cause.",
      "votes": 2
    },
    {
      "id": 105652,
      "postDate": "2016-01-25T16:38:58.690Z",
      "content": "<p>[quote=WD;105630]</p>\n\n<p>@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.</p>\n\n<ul>\n<li><p>Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?</p></li>\n<li><p>You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? </p></li>\n<li><p>Does the code while len(images) &lt; 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? </p></li>\n</ul>\n\n<p>many thanks, W</p>\n\n<p>[/quote]</p>\n\n<p>Keras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) </p>\n\n<p>Update: Seems like this post has caused some confusions. Please, the statement above didn't mean any comparisons between MXnet and Keras. It was a simple run of MXnet with this tutorial's network from my own hardware. @Marko and I had different hardware, different image resolution etc, there was no apple-to-apple comparison of these two toolkits. Keras and Mxnet are both good deep learning frameworks, and it is the same with caffe/theano/torch/Lasagne etc which may later on have good tutorials for this Kaggle competition too, so please select your favorite one or ones for winning 200k $. And sorry to @Marko and @WD for this topic divergence. </p>",
      "rawMarkdown": "[quote=WD;105630]\r\n\r\n@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.\r\n\r\n* Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?\r\n\r\n* You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? \r\n\r\n* Does the code while len(images) < 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? \r\n\r\nmany thanks, W\r\n \r\n\r\n[/quote]\r\n\r\nKeras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) \r\n\r\nUpdate: Seems like this post has caused some confusions. Please, the statement above didn't mean any comparisons between MXnet and Keras. It was a simple run of MXnet with this tutorial's network from my own hardware. @Marko and I had different hardware, different image resolution etc, there was no apple-to-apple comparison of these two toolkits. Keras and Mxnet are both good deep learning frameworks, and it is the same with caffe/theano/torch/Lasagne etc which may later on have good tutorials for this Kaggle competition too, so please select your favorite one or ones for winning 200k $. And sorry to @Marko and @WD for this topic divergence. ",
      "votes": 2
    },
    {
      "id": 105889,
      "postDate": "2016-01-27T07:52:35.513Z",
      "content": "<p>Hi!</p>\n\n<p>Thanks for this tutorial! I cloned github code and run</p>\n\n<pre><code>python data.py\npython train.py\npython submission.py\n</code></pre>\n\n<p>and this gave me    0.074981 on LB, I didn't make any changes in the code. Does anyone have idea what can goes wrong? I expect LB ~ 0.0359</p>",
      "rawMarkdown": "Hi!\r\n\r\nThanks for this tutorial! I cloned github code and run\r\n\r\n    python data.py\r\n    python train.py\r\n    python submission.py\r\n\r\nand this gave me \t0.074981 on LB, I didn't make any changes in the code. Does anyone have idea what can goes wrong? I expect LB ~ 0.0359",
      "votes": 3
    },
    {
      "id": 111308,
      "postDate": "2016-03-13T14:52:17.347Z",
      "content": "<p>you guys are now 1st just wondering what you did? I know you are not using the model in this tutorial. It also looks like you did not upload your model so no prize money :)</p>",
      "rawMarkdown": "you guys are now 1st just wondering what you did? I know you are not using the model in this tutorial. It also looks like you did not upload your model so no prize money :)",
      "votes": 1
    },
    {
      "id": 111178,
      "postDate": "2016-03-12T12:20:24.693Z",
      "content": "<p>I tried with vgg model on 192 x192 images and tweaking some parameters,  got 0.0101. </p>",
      "rawMarkdown": "I tried with vgg model on 192 x192 images and tweaking some parameters,  got 0.0101. ",
      "votes": 1
    },
    {
      "id": 110930,
      "postDate": "2016-03-09T18:18:53.763Z",
      "content": "<p>[quote=datapool;110927]</p>\n\n<p>Thanks for the feedback.\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper &quot;ImageNet Classification with Deep Convolutional Neural Networks&quot;.\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>For the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.</p>",
      "rawMarkdown": "[quote=datapool;110927]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nFor the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.\r\n",
      "votes": 1
    },
    {
      "id": 110921,
      "postDate": "2016-03-09T17:24:58.533Z",
      "content": "<p>[quote=datapool;110899]</p>\n\n<p>Hello Marko Jocic,</p>\n\n<p>Thanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?</p>\n\n<p>[/quote]</p>\n\n<p>Hi, yes we also had that approach, and it yielded somewhat smaller accuracy, but on the other side double faster training, so it's definitely OK for more model diversity.</p>",
      "rawMarkdown": "[quote=datapool;110899]\r\n\r\nHello Marko Jocic,\r\n\r\nThanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?\r\n\r\n[/quote]\r\n\r\nHi, yes we also had that approach, and it yielded somewhat smaller accuracy, but on the other side double faster training, so it's definitely OK for more model diversity.\r\n",
      "votes": 1
    },
    {
      "id": 110377,
      "postDate": "2016-03-04T23:09:58.893Z",
      "content": "<p>[quote=JasonMay;109956]</p>\n\n<p>My question is do you know where all the shape must be addressed to alter image size to 128X128?</p>\n\n<p>Thanks,\nJason</p>\n\n<p>[/quote]</p>\n\n<p>You would need a bigger/beter model for 128x128 images.</p>\n\n<p>[quote=Gongning Luo;110012]</p>\n\n<p>How big of memory(RAM), the parameter of GPU and CPU?</p>\n\n<p>[/quote]</p>\n\n<p>You need ~1GB of GPU RAM for this tutorial.</p>\n\n<p>[quote=Lawrence Chernin;110353]</p>\n\n<p>I created a public AWS image: ami-d42a59b4 that has all the needed libraries to run this code.\nRunning on g2.2xlarge. (sorry, you may also need  to install cython and h5py )</p>\n\n<p>[/quote]</p>\n\n<p>Would you care sharing the image with the rest of Kagglers?</p>\n\n<p>Thanks! Marko</p>",
      "rawMarkdown": "[quote=JasonMay;109956]\r\n\r\nMy question is do you know where all the shape must be addressed to alter image size to 128X128?\r\n\r\nThanks,\r\nJason\r\n\r\n[/quote]\r\n\r\nYou would need a bigger/beter model for 128x128 images.\r\n\r\n[quote=Gongning Luo;110012]\r\n\r\nHow big of memory(RAM), the parameter of GPU and CPU?\r\n\r\n[/quote]\r\n\r\nYou need ~1GB of GPU RAM for this tutorial.\r\n\r\n\r\n[quote=Lawrence Chernin;110353]\r\n\r\nI created a public AWS image: ami-d42a59b4 that has all the needed libraries to run this code.\r\nRunning on g2.2xlarge. (sorry, you may also need  to install cython and h5py )\r\n\r\n[/quote]\r\n\r\nWould you care sharing the image with the rest of Kagglers?\r\n\r\nThanks! Marko",
      "votes": 1
    },
    {
      "id": 106509,
      "postDate": "2016-02-01T14:25:10.947Z",
      "content": "<p>[quote=Ertuka;106465]</p>\n\n<p>1) Could it be weights_systole_best better to use than weights_systole weights?\n2) Could I be missing anything else  to preload the previous model /run in train.py?\n3) Is there any difference using get_model versus to_json/model_from_json? Any idea why the later is not working in this model?</p>\n\n<p>[/quote]</p>\n\n<p>1) You could try both. Both should eventually converge more or less the same.</p>\n\n<p>2) 3) That should be it. It might be a bug currently in Keras, maybe the custom activation function (in first layer) can't be saved/loaded with to_json/from_json. You can check that. However, get_model() should work the same.</p>\n\n<p>[quote=WDharvard;106483]</p>\n\n<p>what are reasonable root mean squared error (RMSE) to observe during training?</p>\n\n<p>[/quote]</p>\n\n<p>You should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (&gt;0.04).</p>",
      "rawMarkdown": "[quote=Ertuka;106465]\r\n\r\n1) Could it be weights_systole_best better to use than weights_systole weights?\r\n2) Could I be missing anything else  to preload the previous model /run in train.py?\r\n3) Is there any difference using get_model versus to_json/model_from_json? Any idea why the later is not working in this model?\r\n\r\n[/quote]\r\n\r\n1) You could try both. Both should eventually converge more or less the same.\r\n\r\n2) 3) That should be it. It might be a bug currently in Keras, maybe the custom activation function (in first layer) can't be saved/loaded with to_json/from_json. You can check that. However, get_model() should work the same.\r\n\r\n[quote=WDharvard;106483]\r\n\r\nwhat are reasonable root mean squared error (RMSE) to observe during training?\r\n\r\n[/quote]\r\n\r\nYou should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (>0.04).",
      "votes": 1
    },
    {
      "id": 106225,
      "postDate": "2016-01-29T09:16:56.503Z",
      "content": "<p>@Marko, @patruf @ Wei Wu,</p>\n\n<p>I was able to get 0.359 with Marko code, but I made changes in submission.py or in val_los.txt files - I don't remember precisely now. But the network trained with train.py is OK. There is only a problem with submission.py.</p>",
      "rawMarkdown": "@Marko, @patruf @ Wei Wu,\r\n\r\nI was able to get 0.359 with Marko code, but I made changes in submission.py or in val_los.txt files - I don't remember precisely now. But the network trained with train.py is OK. There is only a problem with submission.py.",
      "votes": 1
    },
    {
      "id": 106213,
      "postDate": "2016-01-29T04:25:56.280Z",
      "content": "<p>@Marko , @patruff</p>\n\n<p>I can confirm the same problem as @patruff.\nWhen I run the model training, at the end the test CRPS was about 0.029.\nAfter I run submission.py and submit the file, the LB score was about 0.08. (It was a fresh run there was no old files.)</p>\n\n<p>I am running python 2.7 under a windows environment.</p>",
      "rawMarkdown": "@Marko , @patruff\r\n\r\nI can confirm the same problem as @patruff.\r\nWhen I run the model training, at the end the test CRPS was about 0.029.\r\nAfter I run submission.py and submit the file, the LB score was about 0.08. (It was a fresh run there was no old files.)\r\n\r\nI am running python 2.7 under a windows environment.",
      "votes": 1
    },
    {
      "id": 106106,
      "postDate": "2016-01-28T14:36:23.987Z",
      "content": "<h2>Iteration 200/200</h2>\n\n<p>Augmenting images - rotations\n4265/4265 [==============================] - 33s <br>\nAugmenting images - shifts\n4265/4265 [==============================] - 8s <br>\nFitting systole model...\nTrain on 4265 samples, validate on 1066 samples\nEpoch 1/1\n4265/4265 [==============================] - 22s - loss: 14.9959 - val_loss: 19.2078\nFitting diastole model...\nTrain on 4265 samples, validate on 1066 samples\nEpoch 1/1\n4265/4265 [==============================] - 22s - loss: 25.9386 - val_loss: 36.3604\nEvaluating CRPS...\n4265/4265 [==============================] - 9s <br>\n4265/4265 [==============================] - 9s <br>\n1066/1066 [==============================] - 2s <br>\n1066/1066 [==============================] - 2s <br>\nCRPS(train) = 0.02606546647974614\nCRPS(test) = 0.03470680849898871\nSaving weights...</p>\n\n<p>and then I ran the submission.py file but I still ended up with ~0.07 on the submission. Do the above values for CRPS look correct?</p>",
      "rawMarkdown": "Iteration 200/200\r\n--------------------------------------------------\r\nAugmenting images - rotations\r\n4265/4265 [==============================] - 33s       \r\nAugmenting images - shifts\r\n4265/4265 [==============================] - 8s        \r\nFitting systole model...\r\nTrain on 4265 samples, validate on 1066 samples\r\nEpoch 1/1\r\n4265/4265 [==============================] - 22s - loss: 14.9959 - val_loss: 19.2078\r\nFitting diastole model...\r\nTrain on 4265 samples, validate on 1066 samples\r\nEpoch 1/1\r\n4265/4265 [==============================] - 22s - loss: 25.9386 - val_loss: 36.3604\r\nEvaluating CRPS...\r\n4265/4265 [==============================] - 9s     \r\n4265/4265 [==============================] - 9s     \r\n1066/1066 [==============================] - 2s     \r\n1066/1066 [==============================] - 2s     \r\nCRPS(train) = 0.02606546647974614\r\nCRPS(test) = 0.03470680849898871\r\nSaving weights...\r\n\r\nand then I ran the submission.py file but I still ended up with ~0.07 on the submission. Do the above values for CRPS look correct?",
      "votes": 1
    },
    {
      "id": 105970,
      "postDate": "2016-01-27T19:04:59.843Z",
      "content": "<p>I just pushed code changes to repo. It seems that fit_generator() function doesn't work as expected, which needs further investigation. So, for now I changed the code to manually augment the data (rotations + shifts) and to use fit() function (check utils.py).</p>\n\n<p>Also, please create the data again (<code>python data.py</code>), since there was a small bug too. Then try to train the model again.</p>\n\n<p>Once again I apologize for this inconvenience, the tutorial is a product of much copy-pasting (best type of coding, right) from our original project, and some mistakes happened.</p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "I just pushed code changes to repo. It seems that fit_generator() function doesn't work as expected, which needs further investigation. So, for now I changed the code to manually augment the data (rotations + shifts) and to use fit() function (check utils.py).\r\n\r\nAlso, please create the data again (```python data.py```), since there was a small bug too. Then try to train the model again.\r\n\r\nOnce again I apologize for this inconvenience, the tutorial is a product of much copy-pasting (best type of coding, right) from our original project, and some mistakes happened.\r\n\r\nCheers,\r\n\r\nMarko",
      "votes": 1
    },
    {
      "id": 105786,
      "postDate": "2016-01-26T18:23:09.480Z",
      "content": "<p>@Marko. I am not running Keras myself - but am intrigued by the approach. Could you help me to understand how you translate the point estimate on the volumes into a CDF? i understand that you use the following helper function (per below). I don't fully understand the function of sigma, how this is linked to hist_systole.history['loss'][-1] (and what the latter variable is, as hard to see without running the code)! </p>\n\n<p>At a high level, would love to get your views on the more fundamental advantages  / disadvantages between this approach and the approach of estimating 600 values! </p>\n\n<p>Many thanks - W</p>\n\n<blockquote>\n  <p>def real_to_cdf(y, sigma=1e-10):\n      &quot;&quot;&quot;\n      Utility function for creating CDF from real number and sigma (uncertainty measure).</p>\n\n<pre><code>:param y: array of real values\n:param sigma: uncertainty measure. The higher sigma, the more imprecise the prediction is, and vice versa.\n\nDefault value for sigma is 1e-10 to produce step function if needed.\n&quot;&quot;&quot;\ncdf = np.zeros((y.shape[0], 600))\nfor i in range(y.shape[0]):\n    cdf[i] = norm.cdf(np.linspace(0, 599, 600), y[i], sigma)\nreturn cdf\n</code></pre>\n</blockquote>",
      "rawMarkdown": "@Marko. I am not running Keras myself - but am intrigued by the approach. Could you help me to understand how you translate the point estimate on the volumes into a CDF? i understand that you use the following helper function (per below). I don't fully understand the function of sigma, how this is linked to hist_systole.history['loss'][-1] (and what the latter variable is, as hard to see without running the code)! \r\n\r\nAt a high level, would love to get your views on the more fundamental advantages  / disadvantages between this approach and the approach of estimating 600 values! \r\n\r\nMany thanks - W\r\n\r\n> def real_to_cdf(y, sigma=1e-10):\r\n>     \"\"\"\r\n>     Utility function for creating CDF from real number and sigma (uncertainty measure).\r\n> \r\n>     :param y: array of real values\r\n>     :param sigma: uncertainty measure. The higher sigma, the more imprecise the prediction is, and vice versa.\r\n> \r\n>     Default value for sigma is 1e-10 to produce step function if needed.\r\n>     \"\"\"\r\n>     cdf = np.zeros((y.shape[0], 600))\r\n>     for i in range(y.shape[0]):\r\n>         cdf[i] = norm.cdf(np.linspace(0, 599, 600), y[i], sigma)\r\n>     return cdf\r\n\r\n",
      "votes": 1
    },
    {
      "id": 105730,
      "postDate": "2016-01-26T09:41:20.810Z",
      "content": "<p>@Woolsey</p>\n\n<p>From what I can see, you are running it on CPU... You should use Theano flags when running the code, or specify .theanorc file.</p>\n\n<p>You should check these links for more detailed instructions:\n<a href=\"http://deeplearning.net/software/theano/tutorial/using_gpu.html\">http://deeplearning.net/software/theano/tutorial/using_gpu.html</a>\n<a href=\"http://deeplearning.net/software/theano/library/config.html\">http://deeplearning.net/software/theano/library/config.html</a></p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "@Woolsey\r\n\r\nFrom what I can see, you are running it on CPU... You should use Theano flags when running the code, or specify .theanorc file.\r\n\r\nYou should check these links for more detailed instructions:\r\nhttp://deeplearning.net/software/theano/tutorial/using_gpu.html\r\nhttp://deeplearning.net/software/theano/library/config.html\r\n\r\nCheers,\r\n\r\nMarko",
      "votes": 1
    },
    {
      "id": 105655,
      "postDate": "2016-01-25T16:48:26.953Z",
      "content": "<p>[quote=phunter;105652]</p>\n\n<p>Keras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) </p>\n\n<p>[/quote]</p>\n\n<p>This is a pretty ridiculous attempt at saying &quot;lol MXnet is faster than Keras&quot;. No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. </p>\n\n<p>Stay classy man...</p>",
      "rawMarkdown": "[quote=phunter;105652]\r\n\r\nKeras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) \r\n\r\n[/quote]\r\n\r\nThis is a pretty ridiculous attempt at saying \"lol MXnet is faster than Keras\". No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. \r\n\r\nStay classy man...\r\n\r\n\r\n"
    },
    {
      "id": 110376,
      "postDate": "2016-03-04T23:05:36.020Z",
      "content": "<p>[quote=Manuele Tamburrano;109680]</p>\n\n<p>Hi Marko, thank you for sharing your code.</p>\n\n<p>I see you do a 2d convolution on the 30 time series, but I don't find any correlation among different sax of a same study. So do you predict volumes independently for each sax and let the net to guess itself the &quot;level&quot; of a sax?\nI mean, the volume should be predicted by a combination of all sax, so a 3d convolution would seems the right choice here, is there some reason why you opted for the 2d one? (computational reasons aside)</p>\n\n<p>thank you again</p>\n\n<p>[/quote]</p>\n\n<p>You are completely right - in this tutorial the network seems to pick up stuff that helps it determine the volume even if the &quot;level&quot; of sax is not directly given. 3D convolutions are definitely a good idea (which we are using as well), just make sure you are sorting slices correctly ;).</p>",
      "rawMarkdown": "[quote=Manuele Tamburrano;109680]\r\n\r\nHi Marko, thank you for sharing your code.\r\n\r\nI see you do a 2d convolution on the 30 time series, but I don't find any correlation among different sax of a same study. So do you predict volumes independently for each sax and let the net to guess itself the \"level\" of a sax?\r\nI mean, the volume should be predicted by a combination of all sax, so a 3d convolution would seems the right choice here, is there some reason why you opted for the 2d one? (computational reasons aside)\r\n\r\nthank you again\r\n\r\n[/quote]\r\n\r\nYou are completely right - in this tutorial the network seems to pick up stuff that helps it determine the volume even if the \"level\" of sax is not directly given. 3D convolutions are definitely a good idea (which we are using as well), just make sure you are sorting slices correctly ;).\r\n",
      "votes": 2
    },
    {
      "id": 108848,
      "postDate": "2016-02-20T15:30:42.120Z",
      "content": "<p>[quote=PengPai;108816]</p>\n\n<p>@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?</p>\n\n<p>[/quote]</p>\n\n<p>Besides parameter tuning, I found that the most significant performance boosts lied in better pre-processing (i.e. segmentation, ROI extraction) and smart(er) usage of metadata provided in DICOM images. One should have much better insights of all the provided data in order to utilize them properly, instead of just feeding images to CNNs.</p>",
      "rawMarkdown": "[quote=PengPai;108816]\r\n\r\n@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?\r\n\r\n[/quote]\r\n\r\nBesides parameter tuning, I found that the most significant performance boosts lied in better pre-processing (i.e. segmentation, ROI extraction) and smart(er) usage of metadata provided in DICOM images. One should have much better insights of all the provided data in order to utilize them properly, instead of just feeding images to CNNs.",
      "votes": 2
    },
    {
      "id": 108816,
      "postDate": "2016-02-20T06:35:04.677Z",
      "content": "<p>@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?</p>",
      "rawMarkdown": "@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?",
      "votes": 2
    },
    {
      "id": 108319,
      "postDate": "2016-02-16T17:51:29.857Z",
      "content": "<p>[quote=jwjohnson314;108271]</p>\n\n<p>So I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.</p>\n\n<p>[/quote]</p>\n\n<p>When you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.</p>",
      "rawMarkdown": "[quote=jwjohnson314;108271]\r\n\r\nSo I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.\r\n\r\n[/quote]\r\n\r\nWhen you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.",
      "votes": 2
    },
    {
      "id": 108156,
      "postDate": "2016-02-15T22:20:42.160Z",
      "content": "<p>[quote=Marko Jocic;108154]</p>\n\n<p>RMSE objective was removed from Keras a week ago. You can either use MSE as your objective or revert Keras to 0.3.1 version in order to use RMSE. You can also write RMSE as a custom objective function if you wish to use bleeding edge version of Keras.</p>\n\n<p>[/quote]</p>\n\n<p>I guess it would help the spread of Keras if it retained more stability in its API..</p>\n\n<p>I used Keras in the EEG Kaggle competition and used the JZS3 recurrent layers which for these data were faster and better than LSTM (by a noticable factor) - after the competition I upgraded Keras and found that JZS3 was removed...</p>\n\n<p>I really enjoy using Keras - and also perfectly understand making breaking changes if they significantly improve the usability of the API, however removing things only for aesthetic reasons (?) really causes unpleasant surprises for users - and also makes them think twice about upgrading...</p>",
      "rawMarkdown": "[quote=Marko Jocic;108154]\r\n\r\nRMSE objective was removed from Keras a week ago. You can either use MSE as your objective or revert Keras to 0.3.1 version in order to use RMSE. You can also write RMSE as a custom objective function if you wish to use bleeding edge version of Keras.\r\n\r\n[/quote]\r\n\r\nI guess it would help the spread of Keras if it retained more stability in its API..\r\n\r\nI used Keras in the EEG Kaggle competition and used the JZS3 recurrent layers which for these data were faster and better than LSTM (by a noticable factor) - after the competition I upgraded Keras and found that JZS3 was removed...\r\n\r\nI really enjoy using Keras - and also perfectly understand making breaking changes if they significantly improve the usability of the API, however removing things only for aesthetic reasons (?) really causes unpleasant surprises for users - and also makes them think twice about upgrading...\r\n\r\n ",
      "votes": 2
    },
    {
      "id": 111320,
      "postDate": "2016-03-13T18:04:42.247Z",
      "content": "<p>[quote=Yuanfang Guan;111314]</p>\n\n<p>it is pretty clear to anyone that was original top 10 that to perform better than tencia/wash is almost mission impossible. and to reduce 20% of error on top of their initial submissions literally  means a direct labeling of test set.  </p>\n\n<p>[/quote]</p>\n\n<p>One could say the same for your submissions as well, as you got more than 100% improvement in score since the first phase. So let's drop the ball for the manual labeling part.</p>\n\n<p>What Alexander said - bear in mind that current leaderboard is on calculated on 1% of test data, and my guess is that our model &quot;got lucky&quot; on those few studies.  Anyways, I expect much different standings tomorrow night. Good luck.</p>\n\n<p>Marko</p>",
      "rawMarkdown": "[quote=Yuanfang Guan;111314]\r\n\r\n it is pretty clear to anyone that was original top 10 that to perform better than tencia/wash is almost mission impossible. and to reduce 20% of error on top of their initial submissions literally  means a direct labeling of test set.  \r\n\r\n[/quote]\r\n\r\nOne could say the same for your submissions as well, as you got more than 100% improvement in score since the first phase. So let's drop the ball for the manual labeling part.\r\n\r\nWhat Alexander said - bear in mind that current leaderboard is on calculated on 1% of test data, and my guess is that our model \"got lucky\" on those few studies.  Anyways, I expect much different standings tomorrow night. Good luck.\r\n\r\nMarko"
    },
    {
      "id": 111319,
      "postDate": "2016-03-13T17:49:38.333Z",
      "content": "<p>@Yuanfang Guan The current leaderboard positions is highly unstable, I would not count on it even for the remote hints of the final private standing. At best it is &quot;directionally correct&quot;</p>\n\n<p>@DavidGbodiOdaibo Unfortunately due to a technical glitch we were 1 minute late to upload the model. \nThe model is loosely based on the tutorial with number of enhancements ;) It does not use any hand labeling</p>",
      "rawMarkdown": "@Yuanfang Guan The current leaderboard positions is highly unstable, I would not count on it even for the remote hints of the final private standing. At best it is \"directionally correct\"\r\n\r\n@DavidGbodiOdaibo Unfortunately due to a technical glitch we were 1 minute late to upload the model. \r\nThe model is loosely based on the tutorial with number of enhancements ;) It does not use any hand labeling\r\n\r\n"
    },
    {
      "id": 106450,
      "postDate": "2016-01-31T22:02:02.863Z",
      "content": "<p>Thanks, Just to report, I trained vgg on 224 x224 images and scored &lt; 0.03, but it takes 20 min per epochs and would need 30 -40 epochs before overfit . I also tried with 3d convolution, but so far do not have any good result. Since theano does not support multi-gpu, I am moving to torch </p>",
      "rawMarkdown": "Thanks, Just to report, I trained vgg on 224 x224 images and scored < 0.03, but it takes 20 min per epochs and would need 30 -40 epochs before overfit . I also tried with 3d convolution, but so far do not have any good result. Since theano does not support multi-gpu, I am moving to torch ",
      "votes": 2
    },
    {
      "id": 106076,
      "postDate": "2016-01-28T10:27:41.383Z",
      "content": "<p>@Marko, after yours fix I was able to reproduce LB ~ 0.0359. Thanks for sharing!</p>",
      "rawMarkdown": "@Marko, after yours fix I was able to reproduce LB ~ 0.0359. Thanks for sharing!",
      "votes": 2
    },
    {
      "id": 105938,
      "postDate": "2016-01-27T14:30:28.673Z",
      "content": "<p>@ tereka</p>\n\n<p>Thank you for reporting this. Since already two people experienced the same thing, I will investigate what is going on. At least, now we know there certainly is some sort of random element to the code, and I wish to offer my apologies for not seeing this earlier. It is unfortunate, since some people already said they could reproduce similar results with this model (even phunter on MXnet).</p>\n\n<p>If anyone figures out what causes this randomness and shares it with the rest of us, I guess everyone would be thankful and I would gladly update the tutorial accordingly to be more stable.</p>",
      "rawMarkdown": "@ tereka\r\n\r\nThank you for reporting this. Since already two people experienced the same thing, I will investigate what is going on. At least, now we know there certainly is some sort of random element to the code, and I wish to offer my apologies for not seeing this earlier. It is unfortunate, since some people already said they could reproduce similar results with this model (even phunter on MXnet).\r\n\r\nIf anyone figures out what causes this randomness and shares it with the rest of us, I guess everyone would be thankful and I would gladly update the tutorial accordingly to be more stable.",
      "votes": 2
    },
    {
      "id": 105698,
      "postDate": "2016-01-25T22:28:59.810Z",
      "content": "<p>@Woolsey</p>\n\n<p>Please upgrade Theano to latest version:</p>\n\n<pre><code>pip install --upgrade --no-deps git+git://github.com/Theano/Theano.git\n</code></pre>",
      "rawMarkdown": "@Woolsey\r\n\r\nPlease upgrade Theano to latest version:\r\n\r\n\r\n    pip install --upgrade --no-deps git+git://github.com/Theano/Theano.git\r\n\r\n",
      "votes": 2
    },
    {
      "id": 105637,
      "postDate": "2016-01-25T14:12:22.717Z",
      "content": "<p>@WD</p>\n\n<ul>\n<li>I didn't use difference between frames intentionally (as in MXNet tutorial) - just to show that other approaches (in this case - TV denoising) might work fine as well. Definitely a good place for trying different stuff.</li>\n<li>Indeed, that might be true - good pre-processing could really make a difference. Well, I didn't experiment much with hyper-parameters and network itself, but I noticed, for example, that using (5,5) convolutions in first layers yielded roughly same results as using (3,3). More or less it was the same with number of convolutional filters... In this example, the biggest improvement in score was switching from classification (600 sigmoid outputs) to linear regression (only 1 output). However, I do recommend trying out different parameters, because today I changed Adam learning rate a bit, and got slightly better results. Also, this example could maybe use a bit more regularization, since it slightly overfits. Ideally, one should do cross-validation for determining these hyper-parameters, but on my hardware that would be very time consuming.</li>\n</ul>",
      "rawMarkdown": "@WD\r\n\r\n - I didn't use difference between frames intentionally (as in MXNet tutorial) - just to show that other approaches (in this case - TV denoising) might work fine as well. Definitely a good place for trying different stuff.\r\n - Indeed, that might be true - good pre-processing could really make a difference. Well, I didn't experiment much with hyper-parameters and network itself, but I noticed, for example, that using (5,5) convolutions in first layers yielded roughly same results as using (3,3). More or less it was the same with number of convolutional filters... In this example, the biggest improvement in score was switching from classification (600 sigmoid outputs) to linear regression (only 1 output). However, I do recommend trying out different parameters, because today I changed Adam learning rate a bit, and got slightly better results. Also, this example could maybe use a bit more regularization, since it slightly overfits. Ideally, one should do cross-validation for determining these hyper-parameters, but on my hardware that would be very time consuming.",
      "votes": 2
    },
    {
      "id": 105632,
      "postDate": "2016-01-25T13:08:14.227Z",
      "content": "<p>@WD</p>\n\n<ul>\n<li>At first I was training on 128*128 images, but my poor hardware couldn't handle it. But yeah, I would expect improvement on the score with larger images.</li>\n<li>Maybe I could've said that better :). What I meant is - ~100 secs systole model + ~100 secs diastole model + ~60 secs for CRPS evaluation = ~260 secs, which is something below 5 minutes for one iteration. And then you <em>only</em> need 149 more iterations like this.</li>\n<li>Yes, in order to have same input shape for all samples.</li>\n</ul>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "@WD\r\n\r\n\r\n - At first I was training on 128*128 images, but my poor hardware couldn't handle it. But yeah, I would expect improvement on the score with larger images.\r\n - Maybe I could've said that better :). What I meant is - ~100 secs systole model + ~100 secs diastole model + ~60 secs for CRPS evaluation = ~260 secs, which is something below 5 minutes for one iteration. And then you *only* need 149 more iterations like this.\r\n - Yes, in order to have same input shape for all samples.\r\n\r\nCheers,\r\n\r\nMarko",
      "votes": 2
    },
    {
      "id": 105893,
      "postDate": "2016-01-27T08:56:13.660Z",
      "content": "<p>@piotr</p>\n\n<p>Did you let it run for all 150 iterations? 0.074981 is pretty high error, and model should get better then that in first couple of iterations...</p>\n\n<p>Also, our current score (~0.030) uses this same model, but with a bit tweaked parameters, and our previous score (~0.0359) was from the model in the tutorial.</p>",
      "rawMarkdown": "@piotr\r\n\r\nDid you let it run for all 150 iterations? 0.074981 is pretty high error, and model should get better then that in first couple of iterations...\r\n\r\nAlso, our current score (~0.030) uses this same model, but with a bit tweaked parameters, and our previous score (~0.0359) was from the model in the tutorial.",
      "votes": -1
    },
    {
      "id": 111366,
      "postDate": "2016-03-14T05:32:37.273Z",
      "content": "<p>@Marko would love to see your final architecture regardless of how well your team ends up doing!</p>",
      "rawMarkdown": "@Marko would love to see your final architecture regardless of how well your team ends up doing!"
    },
    {
      "id": 111324,
      "postDate": "2016-03-13T18:21:10.657Z",
      "content": "<p>[quote=Yuanfang Guan;111322]</p>\n\n<p>i didn't point to you why are you so defensive?</p>\n\n<p>i meant a 20% error reduction over wash's method in the complete final test set. \neveryone improved almost 100% over their own initial submission due to the selection of three easy cases.</p>\n\n<p>at least we uploaded our models which are binary reproducible. </p>\n\n<p>i hope the final winning model will be open to public scrutiny.  </p>\n\n<p>i don't really count on getting anything from this competition anyway, so i can assure you i am now dropping the topic completely now.</p>\n\n<p>have a nice day.</p>\n\n<p>[/quote]</p>\n\n<p>Please accept my sincere apologies, I totally misunderstood you there.</p>\n\n<p>It's also too bad that we didn't manage to upload our model on time (literally 30 seconds late), but the competition was really fun. </p>\n\n<p>I do hope as well that the winning model will be open for public - it could lead to much improvement and thus better usage in clinical purposes.</p>",
      "rawMarkdown": "[quote=Yuanfang Guan;111322]\r\n\r\ni didn't point to you why are you so defensive?\r\n\r\ni meant a 20% error reduction over wash's method in the complete final test set. \r\neveryone improved almost 100% over their own initial submission due to the selection of three easy cases.\r\n\r\nat least we uploaded our models which are binary reproducible. \r\n\r\ni hope the final winning model will be open to public scrutiny.  \r\n\r\ni don't really count on getting anything from this competition anyway, so i can assure you i am now dropping the topic completely now.\r\n\r\nhave a nice day.\r\n\r\n[/quote]\r\n\r\nPlease accept my sincere apologies, I totally misunderstood you there.\r\n\r\nIt's also too bad that we didn't manage to upload our model on time (literally 30 seconds late), but the competition was really fun. \r\n\r\nI do hope as well that the winning model will be open for public - it could lead to much improvement and thus better usage in clinical purposes."
    },
    {
      "id": 111255,
      "postDate": "2016-03-13T06:20:39.603Z",
      "content": "<p>[quote=Marko Jocic;111237]</p>\n\n<p>[quote=datapool;111183]</p>\n\n<p>After disabling all my normalization layers its didn't improve training speed.</p>\n\n<p>[/quote]</p>\n\n<p>Do you have cuDNN installed?</p>\n\n<p>[/quote]\nYes, installed this cudnn-7.0-linux-x64-v3.0-prod.tgz.</p>",
      "rawMarkdown": "[quote=Marko Jocic;111237]\r\n\r\n[quote=datapool;111183]\r\n\r\nAfter disabling all my normalization layers its didn't improve training speed.\r\n\r\n[/quote]\r\n\r\nDo you have cuDNN installed?\r\n\r\n[/quote]\r\nYes, installed this cudnn-7.0-linux-x64-v3.0-prod.tgz."
    },
    {
      "id": 111237,
      "postDate": "2016-03-12T23:56:52.340Z",
      "content": "<p>[quote=datapool;111183]</p>\n\n<p>After disabling all my normalization layers its didn't improve training speed.</p>\n\n<p>[/quote]</p>\n\n<p>Do you have cuDNN installed?</p>",
      "rawMarkdown": "[quote=datapool;111183]\r\n\r\nAfter disabling all my normalization layers its didn't improve training speed.\r\n\r\n[/quote]\r\n\r\nDo you have cuDNN installed?"
    },
    {
      "id": 111183,
      "postDate": "2016-03-12T15:27:17.877Z",
      "content": "<p>[quote=Marko Jocic;110930]</p>\n\n<p>[quote=datapool;110927]</p>\n\n<p>Thanks for the feedback.\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper &quot;ImageNet Classification with Deep Convolutional Neural Networks&quot;.\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>For the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.</p>\n\n<p>[/quote]\nAfter disabling all my normalization layers its didn't improve training speed.</p>",
      "rawMarkdown": "[quote=Marko Jocic;110930]\r\n\r\n[quote=datapool;110927]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nFor the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.\r\n\r\n\r\n[/quote]\r\nAfter disabling all my normalization layers its didn't improve training speed."
    },
    {
      "id": 111180,
      "postDate": "2016-03-12T15:16:09.913Z",
      "content": "<p>[quote=SecondPlan;111178]</p>\n\n<p>I tried with vgg model on 192 x192 images and tweaking some parameters,  got 0.0101. </p>\n\n<p>[/quote]\nI did 3D model similar to Slow Fusion with input of 24x40x40. Couldn't fit anything larger on 4GB of K520. I crop an ROI using very basic image processing.</p>",
      "rawMarkdown": "[quote=SecondPlan;111178]\r\n\r\nI tried with vgg model on 192 x192 images and tweaking some parameters,  got 0.0101. \r\n\r\n[/quote]\r\nI did 3D model similar to Slow Fusion with input of 24x40x40. Couldn't fit anything larger on 4GB of K520. I crop an ROI using very basic image processing."
    },
    {
      "id": 111177,
      "postDate": "2016-03-12T12:16:19.220Z",
      "content": "<p>Does anyone know whats the score of this method trained on train+validate and scored on the new test data? Just wanted to see if i was able to improve by using this as benchmark. I am at 0.014539. </p>",
      "rawMarkdown": "Does anyone know whats the score of this method trained on train+validate and scored on the new test data? Just wanted to see if i was able to improve by using this as benchmark. I am at 0.014539. "
    },
    {
      "id": 110933,
      "postDate": "2016-03-09T18:46:03.013Z",
      "content": "<p>[quote=Marko Jocic;110930]</p>\n\n<p>[quote=datapool;110927]</p>\n\n<p>Thanks for the feedback.\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper &quot;ImageNet Classification with Deep Convolutional Neural Networks&quot;.\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>For the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.</p>\n\n<p>[/quote]\nOops, i didn't realize normalization layer can make it so slow. I tough depth of my layer is the reason. Anyway good to learn by mistake. I will experiment after competition is over.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "[quote=Marko Jocic;110930]\r\n\r\n[quote=datapool;110927]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nFor the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.\r\n\r\n\r\n[/quote]\r\nOops, i didn't realize normalization layer can make it so slow. I tough depth of my layer is the reason. Anyway good to learn by mistake. I will experiment after competition is over.\r\n\r\nThanks"
    },
    {
      "id": 110927,
      "postDate": "2016-03-09T17:46:03.143Z",
      "content": "<p>[quote=Marko Jocic;110921]</p>\n\n<p>[quote=datapool;110899]</p>\n\n<p>Hello Marko Jocic,</p>\n\n<p>Thanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?</p>\n\n<p>[/quote]</p>\n\n<p>Hi, yes we also had that approach, and it yielded somewhat smaller accuracy, but on the other side double faster training, so it's definitely OK for more model diversity.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for the feedback.\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper &quot;ImageNet Classification with Deep Convolutional Neural Networks&quot;.\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "[quote=Marko Jocic;110921]\r\n\r\n[quote=datapool;110899]\r\n\r\nHello Marko Jocic,\r\n\r\nThanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?\r\n\r\n[/quote]\r\n\r\nHi, yes we also had that approach, and it yielded somewhat smaller accuracy, but on the other side double faster training, so it's definitely OK for more model diversity.\r\n\r\n\r\n[/quote]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks"
    },
    {
      "id": 110899,
      "postDate": "2016-03-09T11:26:12.790Z",
      "content": "<p>Hello Marko Jocic,</p>\n\n<p>Thanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?</p>\n\n<p>[quote=Marko Jocic;109600]</p>\n\n<p>[quote=Alexander Popov;109573]</p>\n\n<p>Thank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you for sharing that. Good luck in your further work!</p>\n\n<p>Cheers, Marko</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Hello Marko Jocic,\r\n\r\nThanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?\r\n\r\n[quote=Marko Jocic;109600]\r\n\r\n[quote=Alexander Popov;109573]\r\n\r\nThank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.\r\n\r\n\r\n[/quote]\r\n\r\nThank you for sharing that. Good luck in your further work!\r\n\r\nCheers, Marko\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 110839,
      "postDate": "2016-03-08T18:10:46.037Z",
      "content": "<p>Thanks a lot for this tutorial. \ndata.py is working fine. I am getting the following error while running train.py.  Any idea what's the problem?</p>\n\n<p>Thanks</p>\n\n<hr>\n\n<p>(myproject) khushhall@tunga:~/myproject/kaggle-dsb2-keras$ python train.py </p>\n\n<p>Using Theano backend.</p>\n\n<p>/home/khushhall/myproject/local/lib/python2.7/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\n  warnings.warn(&quot;downsample module has been moved to the pool module.&quot;)</p>\n\n<p>Loading and compiling models...</p>\n\n<p>Loading training data...</p>\n\n<p>Pre-processing images...</p>\n\n<p>Traceback (most recent call last):</p>\n\n<p>File &quot;train.py&quot;, line 150, in \n    train()</p>\n\n<p>File &quot;train.py&quot;, line 65, in train</p>\n\n<pre><code>X_train, y_train, X_test, y_test = split_data(X, y, split_ratio=0.2)\n</code></pre>\n\n<p>File &quot;train.py&quot;, line 42, in split_data</p>\n\n<pre><code>X_test = X[:split, :, :, :]\n</code></pre>\n\n<p>IndexError: too many indices for array</p>",
      "rawMarkdown": "Thanks a lot for this tutorial. \r\ndata.py is working fine. I am getting the following error while running train.py.  Any idea what's the problem?\r\n\r\nThanks\r\n\r\n------------------------------------------------------------------\r\n\r\n\r\n(myproject) khushhall@tunga:~/myproject/kaggle-dsb2-keras$ python train.py \r\n\r\nUsing Theano backend.\r\n\r\n/home/khushhall/myproject/local/lib/python2.7/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n  warnings.warn(\"downsample module has been moved to the pool module.\")\r\n\r\nLoading and compiling models...\r\n\r\nLoading training data...\r\n\r\nPre-processing images...\r\n\r\nTraceback (most recent call last):\r\n\r\n  File \"train.py\", line 150, in <module>\r\n    train()\r\n\r\n  File \"train.py\", line 65, in train\r\n\r\n    X_train, y_train, X_test, y_test = split_data(X, y, split_ratio=0.2)\r\n\r\n  File \"train.py\", line 42, in split_data\r\n\r\n    X_test = X[:split, :, :, :]\r\n\r\nIndexError: too many indices for array\r\n",
      "replies": [
        {
          "id": 226512,
          "postDate": "2017-10-02T14:43:35.657Z",
          "content": "<p>A little late, but from an earlier post it appears that this feature has become depricated.</p>\n\n<pre><code>DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\n</code></pre>\n\n<p>To fix this, you can change the (split) into int(split).</p>\n\n<pre><code>X_test = X[:int(split), :, :, :]\ny_test = y[:int(split), :]\nX_train = X[int(split):, :, :, :]\ny_train = y[int(split):, :]\n</code></pre>",
          "rawMarkdown": "A little late, but from an earlier post it appears that this feature has become depricated.\n\n    DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\n\nTo fix this, you can change the (split) into int(split).\n\n    X_test = X[:int(split), :, :, :]\n    y_test = y[:int(split), :]\n    X_train = X[int(split):, :, :, :]\n    y_train = y[int(split):, :]"
        }
      ]
    },
    {
      "id": 110393,
      "postDate": "2016-03-05T01:13:58.620Z",
      "content": "<p>It is a publicly available image. More details of how to find it in AWS are attached. The smallest gpu on AWS is the g2.2xlarge which has 15GB ram. I see the process is using round 3GB ram (RES).</p>",
      "rawMarkdown": "It is a publicly available image. More details of how to find it in AWS are attached. The smallest gpu on AWS is the g2.2xlarge which has 15GB ram. I see the process is using round 3GB ram (RES).\r\n"
    },
    {
      "id": 110353,
      "postDate": "2016-03-04T20:10:08.610Z",
      "content": "<p>I created a public AWS image: ami-d42a59b4 that has all the needed libraries to run this code.\nRunning on g2.2xlarge. (sorry, you may also need  to install cython and h5py )</p>",
      "rawMarkdown": "I created a public AWS image: ami-d42a59b4 that has all the needed libraries to run this code.\r\nRunning on g2.2xlarge. (sorry, you may also need  to install cython and h5py )"
    },
    {
      "id": 110012,
      "postDate": "2016-03-02T02:11:29.730Z",
      "content": "<p>@Marko Jocic\nWhen I use the source code, some errors occurs as following:</p>\n\n<p><em>File &quot;train.py&quot;, line 18, in load_train_data\n  X=X.astype\n     MemoryError</em></p>\n\n<p>I am a beginner in the area. I want to know the requirement of hardware to achive running the code.\nHow big of memory(RAM), the parameter of GPU and CPU\nI also want to know How can I merge with you and to follow you in your team?</p>\n\n<p>Thank you very much!</p>",
      "rawMarkdown": "@Marko Jocic\r\nWhen I use the source code, some errors occurs as following:\r\n      \r\n*File \"train.py\", line 18, in load_train_data\r\n  X=X.astype<np.float32>\r\n     MemoryError*\r\n\r\nI am a beginner in the area. I want to know the requirement of hardware to achive running the code.\r\nHow big of memory(RAM), the parameter of GPU and CPU\r\nI also want to know How can I merge with you and to follow you in your team?\r\n\r\nThank you very much!\r\n"
    },
    {
      "id": 109956,
      "postDate": "2016-03-01T18:11:28.397Z",
      "content": "<p>[quote=Marko Jocic;105632]</p>\n\n<p>@WD</p>\n\n<ul>\n<li>At first I was training on 128*128 images, but my poor hardware couldn't handle it. But yeah, I would expect improvement on the score with larger images.</li>\n</ul>\n\n<p>Cheers,</p>\n\n<p>Marko</p>\n\n<p>[/quote]</p>\n\n<p>Hi Marko,</p>\n\n<p>Thanks for the tutorial.  It has saved my group, which is taking on the DSB as our Graduate Capstone project.  We were able to achieve ~ .033 CRPS at about 3 minutes per iteration.</p>\n\n<p>I am interested in trying different image sizes (96X96 or 128X128), but think I am missing step.</p>\n\n<p>I changed img_shape = (64,64) in the data.py file, but when I ran train.py, I got a significantly worse CRPS in the first few iterations.</p>\n\n<p>I tried also altering model.py to input_shape=(30,128,128) but received an error message.</p>\n\n<p>My question is do you know where all the shape must be addressed to alter image size to 128X128?</p>\n\n<p>Thanks,\nJason</p>",
      "rawMarkdown": "[quote=Marko Jocic;105632]\r\n\r\n@WD\r\n\r\n\r\n - At first I was training on 128*128 images, but my poor hardware couldn't handle it. But yeah, I would expect improvement on the score with larger images.\r\n\r\nCheers,\r\n\r\nMarko\r\n\r\n[/quote]\r\n\r\nHi Marko,\r\n\r\nThanks for the tutorial.  It has saved my group, which is taking on the DSB as our Graduate Capstone project.  We were able to achieve ~ .033 CRPS at about 3 minutes per iteration.\r\n\r\nI am interested in trying different image sizes (96X96 or 128X128), but think I am missing step.\r\n\r\nI changed img_shape = (64,64) in the data.py file, but when I ran train.py, I got a significantly worse CRPS in the first few iterations.\r\n\r\nI tried also altering model.py to input_shape=(30,128,128) but received an error message.\r\n\r\nMy question is do you know where all the shape must be addressed to alter image size to 128X128?\r\n\r\nThanks,\r\nJason\r\n"
    },
    {
      "id": 109680,
      "postDate": "2016-02-29T11:41:35.107Z",
      "content": "<p>Hi Marko, thank you for sharing your code.</p>\n\n<p>I see you do a 2d convolution on the 30 time series, but I don't find any correlation among different sax of a same study. So do you predict volumes independently for each sax and let the net to guess itself the &quot;level&quot; of a sax?\nI mean, the volume should be predicted by a combination of all sax, so a 3d convolution would seems the right choice here, is there some reason why you opted for the 2d one? (computational reasons aside)</p>\n\n<p>thank you again</p>",
      "rawMarkdown": "Hi Marko, thank you for sharing your code.\r\n\r\nI see you do a 2d convolution on the 30 time series, but I don't find any correlation among different sax of a same study. So do you predict volumes independently for each sax and let the net to guess itself the \"level\" of a sax?\r\nI mean, the volume should be predicted by a combination of all sax, so a 3d convolution would seems the right choice here, is there some reason why you opted for the 2d one? (computational reasons aside)\r\n\r\nthank you again"
    },
    {
      "id": 109600,
      "postDate": "2016-02-28T12:16:04.993Z",
      "content": "<p>[quote=Alexander Popov;109573]</p>\n\n<p>Thank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you for sharing that. Good luck in your further work!</p>\n\n<p>Cheers, Marko</p>",
      "rawMarkdown": "[quote=Alexander Popov;109573]\r\n\r\nThank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.\r\n\r\n\r\n[/quote]\r\n\r\nThank you for sharing that. Good luck in your further work!\r\n\r\nCheers, Marko\r\n"
    },
    {
      "id": 109573,
      "postDate": "2016-02-27T23:18:23.453Z",
      "content": "<p>Thank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.</p>",
      "rawMarkdown": "Thank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.\r\n"
    },
    {
      "id": 109505,
      "postDate": "2016-02-26T21:27:41.637Z",
      "content": "<p>[quote=Andrew Beam;109502]</p>\n\n<p>It looks like the space represented by a pixel can vary quite a bit from image to image. Have you found normalizing the image spacing to be crucial to success?</p>\n\n<p>[/quote]</p>\n\n<p>It depends on the type of analysis you do. However, in our case we found it definitely helps</p>",
      "rawMarkdown": "[quote=Andrew Beam;109502]\r\n\r\nIt looks like the space represented by a pixel can vary quite a bit from image to image. Have you found normalizing the image spacing to be crucial to success?\r\n\r\n[/quote]\r\n\r\nIt depends on the type of analysis you do. However, in our case we found it definitely helps"
    },
    {
      "id": 109502,
      "postDate": "2016-02-26T20:57:22.987Z",
      "content": "<p>Hi Marko, great tutorial. I was wondering if you would be willing to comment on how useful you have found the pixel spacing and slice location to be? It looks like the space represented by a pixel can vary quite a bit from image to image. Have you found normalizing the image spacing to be crucial to success?</p>",
      "rawMarkdown": "Hi Marko, great tutorial. I was wondering if you would be willing to comment on how useful you have found the pixel spacing and slice location to be? It looks like the space represented by a pixel can vary quite a bit from image to image. Have you found normalizing the image spacing to be crucial to success?"
    },
    {
      "id": 109471,
      "postDate": "2016-02-26T14:14:35.097Z",
      "content": "<p>[quote=Marko Jocic;109466]</p>\n\n<p>[quote=DirkWillemWonnink;109400]</p>\n\n<p>Hi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   </p>\n\n<p>[/quote]</p>\n\n<p>As far as I can remember, <code>imresize</code> (from <code>scipy.misc</code>) scales images to [0,255], and that function is used in <code>crop_resize()</code> in the tutorial.\nIf you manage to figure out what causes the black images, please let me know and/or solve it through a PR on the tutorial.</p>\n\n<p>Thank you,</p>\n\n<p>Marko</p>\n\n<p>[/quote]</p>\n\n<p>Indeed I read now that imsize is converting to uint8 and automatically scaling it to 0 .. 255. For my test I switched to opencv for the rescaling as part of some additional functions. So this 'hidden conversion' was not executed anymore. It seems you can save some conversions in your tutorial model. :-) (ranging to 0 ..1 and convert to uint8) . Thanks for this clarification!</p>\n\n<p>Gr. Dirk Willem</p>",
      "rawMarkdown": "[quote=Marko Jocic;109466]\r\n\r\n[quote=DirkWillemWonnink;109400]\r\n\r\nHi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   \r\n\r\n[/quote]\r\n\r\nAs far as I can remember, ```imresize``` (from ```scipy.misc```) scales images to [0,255], and that function is used in ```crop_resize()``` in the tutorial.\r\nIf you manage to figure out what causes the black images, please let me know and/or solve it through a PR on the tutorial.\r\n\r\nThank you,\r\n\r\nMarko\r\n\r\n\r\n[/quote]\r\n\r\nIndeed I read now that imsize is converting to uint8 and automatically scaling it to 0 .. 255. For my test I switched to opencv for the rescaling as part of some additional functions. So this 'hidden conversion' was not executed anymore. It seems you can save some conversions in your tutorial model. :-) (ranging to 0 ..1 and convert to uint8) . Thanks for this clarification!\r\n\r\nGr. Dirk Willem"
    },
    {
      "id": 109466,
      "postDate": "2016-02-26T12:35:28.223Z",
      "content": "<p>[quote=DirkWillemWonnink;109400]</p>\n\n<p>Hi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   </p>\n\n<p>[/quote]</p>\n\n<p>As far as I can remember, <code>imresize</code> (from <code>scipy.misc</code>) scales images to [0,255], and that function is used in <code>crop_resize()</code> in the tutorial.\nIf you manage to figure out what causes the black images, please let me know and/or solve it through a PR on the tutorial.</p>\n\n<p>Thank you,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "[quote=DirkWillemWonnink;109400]\r\n\r\nHi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   \r\n\r\n[/quote]\r\n\r\nAs far as I can remember, ```imresize``` (from ```scipy.misc```) scales images to [0,255], and that function is used in ```crop_resize()``` in the tutorial.\r\nIf you manage to figure out what causes the black images, please let me know and/or solve it through a PR on the tutorial.\r\n\r\nThank you,\r\n\r\nMarko\r\n"
    },
    {
      "id": 109400,
      "postDate": "2016-02-25T18:43:43.627Z",
      "content": "<p>Hi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   </p>",
      "rawMarkdown": "Hi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   "
    },
    {
      "id": 109351,
      "postDate": "2016-02-25T11:57:13.520Z",
      "content": "<p>[quote=Alex Risman;109284]</p>\n\n<p>Hey Marko, this is awesome, thanks so much! Quick question: what is the purpose of the zero padding layers in the model?</p>\n\n<p>[/quote]</p>\n\n<p>Convolutional layers with border model &quot;valid&quot; and filter size greater than 1x1 reduce the size of feature maps, so I just zero pad these feature maps with one row/column on all sides to maintain the size of feature maps.</p>",
      "rawMarkdown": "[quote=Alex Risman;109284]\r\n\r\nHey Marko, this is awesome, thanks so much! Quick question: what is the purpose of the zero padding layers in the model?\r\n\r\n[/quote]\r\n\r\nConvolutional layers with border model \"valid\" and filter size greater than 1x1 reduce the size of feature maps, so I just zero pad these feature maps with one row/column on all sides to maintain the size of feature maps.\r\n"
    },
    {
      "id": 109284,
      "postDate": "2016-02-24T19:47:14.067Z",
      "content": "<p>Hey Marko, this is awesome, thanks so much! Quick question: what is the purpose of the zero padding layers in the model?</p>",
      "rawMarkdown": "Hey Marko, this is awesome, thanks so much! Quick question: what is the purpose of the zero padding layers in the model?"
    },
    {
      "id": 108849,
      "postDate": "2016-02-20T15:38:49.410Z",
      "content": "<p>[quote=John Gunawan;108673]</p>\n\n<p>Hi guys, </p>\n\n<p>I'm pretty new to all of this, so I apologize if my question is very easy or just bad but I got stuck running data.py.  I'm getting the following error: </p>\n\n<p>ValueError: need more than 1 value to unpack</p>\n\n<p>Can anyone advise on what I should do to troubleshoot this? </p>\n\n<p>Any help is appreciated!  Thanks! </p>\n\n<p>[/quote]\n@John: probably your train.csv is not the original anymore. Had the same issue.</p>",
      "rawMarkdown": "[quote=John Gunawan;108673]\r\n\r\nHi guys, \r\n\r\nI'm pretty new to all of this, so I apologize if my question is very easy or just bad but I got stuck running data.py.  I'm getting the following error: \r\n\r\nValueError: need more than 1 value to unpack\r\n\r\nCan anyone advise on what I should do to troubleshoot this? \r\n\r\nAny help is appreciated!  Thanks! \r\n\r\n[/quote]\r\n@John: probably your train.csv is not the original anymore. Had the same issue.\r\n"
    },
    {
      "id": 108809,
      "postDate": "2016-02-20T05:38:30.553Z",
      "content": "<p>[quote=EIGSI;106532]</p>\n\n<p>I can't understand why the performance of the diastole network is consistently worse than that of systole? It is the same input data, same model why the difference?\nI noticed this on mxnet example too, which is using the frame differences. In terms of CRPS there is always a .01 difference and in terms of RMSE here the difference seems to be around 10.</p>\n\n<p>[quote=Marko Jocic;106509]</p>\n\n<p>You should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (&gt;0.04).</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>Maybe this is because diastole volume has more variations than systole volume (std dev: 59 vs 43).</p>",
      "rawMarkdown": "[quote=EIGSI;106532]\r\n\r\nI can't understand why the performance of the diastole network is consistently worse than that of systole? It is the same input data, same model why the difference?\r\nI noticed this on mxnet example too, which is using the frame differences. In terms of CRPS there is always a .01 difference and in terms of RMSE here the difference seems to be around 10.\r\n\r\n[quote=Marko Jocic;106509]\r\n\r\n \r\n\r\nYou should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (>0.04).\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\nMaybe this is because diastole volume has more variations than systole volume (std dev: 59 vs 43)."
    },
    {
      "id": 108673,
      "postDate": "2016-02-19T02:50:11.457Z",
      "content": "<p>Hi guys, </p>\n\n<p>I'm pretty new to all of this, so I apologize if my question is very easy or just bad but I got stuck running data.py.  I'm getting the following error: </p>\n\n<p>ValueError: need more than 1 value to unpack</p>\n\n<p>Can anyone advise on what I should do to troubleshoot this? </p>\n\n<p>Any help is appreciated!  Thanks! </p>",
      "rawMarkdown": "Hi guys, \r\n\r\nI'm pretty new to all of this, so I apologize if my question is very easy or just bad but I got stuck running data.py.  I'm getting the following error: \r\n\r\nValueError: need more than 1 value to unpack\r\n\r\nCan anyone advise on what I should do to troubleshoot this? \r\n\r\nAny help is appreciated!  Thanks! "
    },
    {
      "id": 108573,
      "postDate": "2016-02-18T12:01:51.733Z",
      "content": "<p>[quote=Maksim Korolev;108544]</p>\n\n<p>I also recorded error at every step, here are the graphs if anybody is interested. Seems like it overfit but easier to tell on the CRPS graph. Interested to hear others' thoughts.</p>\n\n<p>[/quote]</p>\n\n<p>Yes it is obvious it overfits, but rotations and shifts should reduce overfitting.\nAlso, maybe try adding more regularization to the model.</p>",
      "rawMarkdown": "[quote=Maksim Korolev;108544]\r\n\r\nI also recorded error at every step, here are the graphs if anybody is interested. Seems like it overfit but easier to tell on the CRPS graph. Interested to hear others' thoughts.\r\n\r\n[/quote]\r\n\r\nYes it is obvious it overfits, but rotations and shifts should reduce overfitting.\r\nAlso, maybe try adding more regularization to the model.\r\n"
    },
    {
      "id": 108544,
      "postDate": "2016-02-18T05:44:10.380Z",
      "content": "<p>Thank you Marko! I used your code, but I removed the denoising and rotation and was still able to get around the ~0.035 in 100 iterations (older hardware). I also recorded error at every step, here are the graphs if anybody is interested. Seems like it overfit but easier to tell on the CRPS graph. Interested to hear others' thoughts.</p>",
      "rawMarkdown": "Thank you Marko! I used your code, but I removed the denoising and rotation and was still able to get around the ~0.035 in 100 iterations (older hardware). I also recorded error at every step, here are the graphs if anybody is interested. Seems like it overfit but easier to tell on the CRPS graph. Interested to hear others' thoughts."
    },
    {
      "id": 108490,
      "postDate": "2016-02-17T19:59:55.747Z",
      "content": "<p>I just updated the code to use custom RMSE loss function.</p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "I just updated the code to use custom RMSE loss function.\r\n\r\nCheers,\r\n\r\nMarko"
    },
    {
      "id": 108344,
      "postDate": "2016-02-16T20:03:08.147Z",
      "content": "<p>[quote=Guilherme Goto Escudero;108319]</p>\n\n<p>[quote=jwjohnson314;108271]</p>\n\n<p>So I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.</p>\n\n<p>[/quote]</p>\n\n<p>When you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks Guilherme - I missed that!</p>",
      "rawMarkdown": "[quote=Guilherme Goto Escudero;108319]\r\n\r\n[quote=jwjohnson314;108271]\r\n\r\nSo I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.\r\n\r\n[/quote]\r\n\r\nWhen you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.\r\n\r\n[/quote]\r\n\r\nThanks Guilherme - I missed that!\r\n"
    },
    {
      "id": 108333,
      "postDate": "2016-02-16T19:10:35.863Z",
      "content": "<p>[quote=Guilherme Goto Escudero;108319]</p>\n\n<p>When you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.</p>\n\n<p>[/quote]</p>\n\n<p>Exactly!</p>",
      "rawMarkdown": "[quote=Guilherme Goto Escudero;108319]\r\n\r\nWhen you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.\r\n\r\n[/quote]\r\n\r\nExactly!"
    },
    {
      "id": 108271,
      "postDate": "2016-02-16T14:55:24.753Z",
      "content": "<p>So I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.</p>",
      "rawMarkdown": "So I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date."
    },
    {
      "id": 108154,
      "postDate": "2016-02-15T21:33:56.507Z",
      "content": "<p>[quote=kaldes;108131]</p>\n\n<p>Exception: Invalid objective: rmse</p>\n\n<p>Tried upgrading Theano, no luck. </p>\n\n<p>Any suggestions?</p>\n\n<p>[/quote]</p>\n\n<p>RMSE objective was removed from Keras a week ago. You can either use MSE as your objective or revert Keras to 0.3.1 version in order to use RMSE. You can also write RMSE as a custom objective function if you wish to use bleeding edge version of Keras.</p>\n\n<p>Hope that helps,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "[quote=kaldes;108131]\r\n\r\nException: Invalid objective: rmse\r\n\r\nTried upgrading Theano, no luck. \r\n\r\nAny suggestions?\r\n\r\n[/quote]\r\n\r\nRMSE objective was removed from Keras a week ago. You can either use MSE as your objective or revert Keras to 0.3.1 version in order to use RMSE. You can also write RMSE as a custom objective function if you wish to use bleeding edge version of Keras.\r\n\r\nHope that helps,\r\n\r\nMarko\r\n"
    },
    {
      "id": 108135,
      "postDate": "2016-02-15T19:20:51.487Z",
      "content": "<p>Its working now. Issue was with the  version of MSVS</p>\n\n<p>[quote=Woolsey;108127]</p>\n\n<p>Mohd, </p>\n\n<p>Sorry I can't really answer your question as I gave up on using convnet for now (no nvidia).</p>\n\n<p>However if the improvement is below x10, I'd guess you have a software problem, either within the code or because of an installation issue.</p>\n\n<p>Good luck...</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "\r\nIts working now. Issue was with the  version of MSVS\r\n\r\n[quote=Woolsey;108127]\r\n\r\n\r\n\r\nMohd, \r\n\r\nSorry I can't really answer your question as I gave up on using convnet for now (no nvidia).\r\n\r\nHowever if the improvement is below x10, I'd guess you have a software problem, either within the code or because of an installation issue.\r\n\r\nGood luck...\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 108133,
      "postDate": "2016-02-15T19:17:56.473Z",
      "content": "<p>use loss='mse'\n[quote=kaldes;108131]</p>\n\n<p>Just joined the contest and very excited to see the collaborative energy in the forum! Thank you guys! Eager to try DL with Keras on this problem. </p>\n\n<p>Got through prepping the data files without error. </p>\n\n<h1>Trying to train, I am getting the following error (Mac Book Pro, OS X El Capitano)</h1>\n\n<p>Python 3.5.1 |Anaconda 2.4.1 (x86_64)| (default, Dec  7 2015, 11:24:55) \n[GCC 4.2.1 (Apple Inc. build 5577)] on darwin\nType &quot;help&quot;, &quot;copyright&quot;, &quot;credits&quot; or &quot;license&quot; for more information.</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>runfile('/Users/aards/Desktop/temp/train.py', wdir='/Users/aards/Desktop/temp')\n      Using Theano backend.\n      /Users/aards/anaconda/lib/python3.5/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\n        warnings.warn(&quot;downsample module has been moved to the pool module.&quot;)\n      Loading and compiling models...\n      Traceback (most recent call last):\n        File &quot;&quot;, line 1, in \n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py&quot;, line 699, in runfile\n          execfile(filename, namespace)\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py&quot;, line 88, in execfile\n          exec(compile(open(filename, 'rb').read(), filename, 'exec'), namespace)\n        File &quot;/Users/aards/Desktop/temp/train.py&quot;, line 147, in \n          train()\n        File &quot;/Users/aards/Desktop/temp/train.py&quot;, line 52, in train\n          model_systole = get_model()\n        File &quot;/Users/aards/Desktop/temp/model.py&quot;, line 52, in get_model\n          model.compile(optimizer=adam, loss='rmse')\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/keras/models.py&quot;, line 460, in compile\n          self.loss = objectives.get(loss)\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/keras/objectives.py&quot;, line 62, in get\n          return get_from_module(identifier, globals(), 'objective')\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/keras/utils/generic_utils.py&quot;, line 14, in get_from_module\n          str(identifier))</p>\n      \n      <h1>Exception: Invalid objective: rmse</h1>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>Tried upgrading Theano, no luck. </p>\n\n<p>Any suggestions?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "use loss='mse'\r\n[quote=kaldes;108131]\r\n\r\nJust joined the contest and very excited to see the collaborative energy in the forum! Thank you guys! Eager to try DL with Keras on this problem. \r\n\r\nGot through prepping the data files without error. \r\n\r\nTrying to train, I am getting the following error (Mac Book Pro, OS X El Capitano)\r\n======\r\nPython 3.5.1 |Anaconda 2.4.1 (x86_64)| (default, Dec  7 2015, 11:24:55) \r\n[GCC 4.2.1 (Apple Inc. build 5577)] on darwin\r\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\r\n>>> runfile('/Users/aards/Desktop/temp/train.py', wdir='/Users/aards/Desktop/temp')\r\nUsing Theano backend.\r\n/Users/aards/anaconda/lib/python3.5/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n  warnings.warn(\"downsample module has been moved to the pool module.\")\r\nLoading and compiling models...\r\nTraceback (most recent call last):\r\n  File \"<stdin>\", line 1, in <module>\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 699, in runfile\r\n    execfile(filename, namespace)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 88, in execfile\r\n    exec(compile(open(filename, 'rb').read(), filename, 'exec'), namespace)\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 147, in <module>\r\n    train()\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 52, in train\r\n    model_systole = get_model()\r\n  File \"/Users/aards/Desktop/temp/model.py\", line 52, in get_model\r\n    model.compile(optimizer=adam, loss='rmse')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/models.py\", line 460, in compile\r\n    self.loss = objectives.get(loss)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/objectives.py\", line 62, in get\r\n    return get_from_module(identifier, globals(), 'objective')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/utils/generic_utils.py\", line 14, in get_from_module\r\n    str(identifier))\r\nException: Invalid objective: rmse\r\n=========\r\n\r\nTried upgrading Theano, no luck. \r\n\r\nAny suggestions?\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 108131,
      "postDate": "2016-02-15T18:52:51.933Z",
      "content": "<p>Just joined the contest and very excited to see the collaborative energy in the forum! Thank you guys! Eager to try DL with Keras on this problem. </p>\n\n<p>Got through prepping the data files without error. </p>\n\n<h1>Trying to train, I am getting the following error (Mac Book Pro, OS X El Capitano)</h1>\n\n<p>Python 3.5.1 |Anaconda 2.4.1 (x86_64)| (default, Dec  7 2015, 11:24:55) \n[GCC 4.2.1 (Apple Inc. build 5577)] on darwin\nType &quot;help&quot;, &quot;copyright&quot;, &quot;credits&quot; or &quot;license&quot; for more information.</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>runfile('/Users/aards/Desktop/temp/train.py', wdir='/Users/aards/Desktop/temp')\n      Using Theano backend.\n      /Users/aards/anaconda/lib/python3.5/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\n        warnings.warn(&quot;downsample module has been moved to the pool module.&quot;)\n      Loading and compiling models...\n      Traceback (most recent call last):\n        File &quot;&quot;, line 1, in \n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py&quot;, line 699, in runfile\n          execfile(filename, namespace)\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py&quot;, line 88, in execfile\n          exec(compile(open(filename, 'rb').read(), filename, 'exec'), namespace)\n        File &quot;/Users/aards/Desktop/temp/train.py&quot;, line 147, in \n          train()\n        File &quot;/Users/aards/Desktop/temp/train.py&quot;, line 52, in train\n          model_systole = get_model()\n        File &quot;/Users/aards/Desktop/temp/model.py&quot;, line 52, in get_model\n          model.compile(optimizer=adam, loss='rmse')\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/keras/models.py&quot;, line 460, in compile\n          self.loss = objectives.get(loss)\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/keras/objectives.py&quot;, line 62, in get\n          return get_from_module(identifier, globals(), 'objective')\n        File &quot;/Users/aards/anaconda/lib/python3.5/site-packages/keras/utils/generic_utils.py&quot;, line 14, in get_from_module\n          str(identifier))</p>\n      \n      <h1>Exception: Invalid objective: rmse</h1>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>Tried upgrading Theano, no luck. </p>\n\n<p>Any suggestions?</p>",
      "rawMarkdown": "Just joined the contest and very excited to see the collaborative energy in the forum! Thank you guys! Eager to try DL with Keras on this problem. \r\n\r\nGot through prepping the data files without error. \r\n\r\nTrying to train, I am getting the following error (Mac Book Pro, OS X El Capitano)\r\n======\r\nPython 3.5.1 |Anaconda 2.4.1 (x86_64)| (default, Dec  7 2015, 11:24:55) \r\n[GCC 4.2.1 (Apple Inc. build 5577)] on darwin\r\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\r\n>>> runfile('/Users/aards/Desktop/temp/train.py', wdir='/Users/aards/Desktop/temp')\r\nUsing Theano backend.\r\n/Users/aards/anaconda/lib/python3.5/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n  warnings.warn(\"downsample module has been moved to the pool module.\")\r\nLoading and compiling models...\r\nTraceback (most recent call last):\r\n  File \"<stdin>\", line 1, in <module>\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 699, in runfile\r\n    execfile(filename, namespace)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 88, in execfile\r\n    exec(compile(open(filename, 'rb').read(), filename, 'exec'), namespace)\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 147, in <module>\r\n    train()\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 52, in train\r\n    model_systole = get_model()\r\n  File \"/Users/aards/Desktop/temp/model.py\", line 52, in get_model\r\n    model.compile(optimizer=adam, loss='rmse')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/models.py\", line 460, in compile\r\n    self.loss = objectives.get(loss)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/objectives.py\", line 62, in get\r\n    return get_from_module(identifier, globals(), 'objective')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/utils/generic_utils.py\", line 14, in get_from_module\r\n    str(identifier))\r\nException: Invalid objective: rmse\r\n=========\r\n\r\nTried upgrading Theano, no luck. \r\n\r\nAny suggestions?\r\n"
    },
    {
      "id": 108127,
      "postDate": "2016-02-15T18:28:14.073Z",
      "content": "<p>Mohd, </p>\n\n<p>Sorry I can't really answer your question as I gave up on using convnet for now (no nvidia).</p>\n\n<p>However if the improvement is below x10, I'd guess you have a software problem, either within the code or because of an installation issue.</p>\n\n<p>Good luck...</p>",
      "rawMarkdown": "Mohd, \r\n\r\nSorry I can't really answer your question as I gave up on using convnet for now (no nvidia).\r\n\r\nHowever if the improvement is below x10, I'd guess you have a software problem, either within the code or because of an installation issue.\r\n\r\nGood luck..."
    },
    {
      "id": 108056,
      "postDate": "2016-02-15T05:42:28.177Z",
      "content": "<p>Hi Woolsey,</p>\n\n<p>I am currently facing the same issue (iteration 1 took around 10 hrs). I have applied the fix suggested by MJ, am using &quot;Nvidia Grid K520 GPU&quot;. How much time does iteration 1 take after applying the fix suggested by MJ. \nWhat was your total training time?</p>\n\n<p>[quote=Woolsey;105719]</p>\n\n<p>One more question please. Iteration 1 took 6-7 hours on my system (XPS8100, no decent GPU). Does that sound right given this poor hardware, or I should suspect some problem with the software? </p>\n\n<pre><code>   runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\n    Using Theano backend.\n    Loading and compiling models...\n    Loading training data...\n    Pre-processing images...\n    5331/5331 [==============================] - 934s   C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\theano\\tensor\\signal\\downsample.py:5: UserWarning: downsample module has been moved to the pool module.\n      warnings.warn(&quot;downsample module has been moved to the pool module.&quot;)\n    C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py:39: DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\n      X_test = X[:split, :, :, :]\n\n\n--------------------------------------------------\nTraining...\n--------------------------------------------------\n--------------------------------------------------\nIteration 1/150\n--------------------------------------------------\nEpoch 1/1\n4290/4265 [==============================] - 5542s - loss: 34.4657 - val_loss: 42.3799\nFitting diastole model...\nEpoch 1/1\n4290/4265 [==============================] - 5523s - loss: 55.8717 - val_loss: 86.3558\nEvaluating CRPS...\n4265/4265 [==============================] - 1766s     \n4265/4265 [==============================] - 1763s     \n1066/1066 [==============================] - 441s     \n1066/1066 [==============================] - 439s     \nCRPS(train) = 0.0785559420969\nCRPS(test) = 0.0743980740534\nSaving weights...\n--------------------------------------------------\nIteration 2/150\n--------------------------------------------------\n</code></pre>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Hi Woolsey,\r\n\r\nI am currently facing the same issue (iteration 1 took around 10 hrs). I have applied the fix suggested by MJ, am using \"Nvidia Grid K520 GPU\". How much time does iteration 1 take after applying the fix suggested by MJ. \r\nWhat was your total training time?\r\n\r\n[quote=Woolsey;105719]\r\n\r\nOne more question please. Iteration 1 took 6-7 hours on my system (XPS8100, no decent GPU). Does that sound right given this poor hardware, or I should suspect some problem with the software? \r\n\r\n\r\n \r\n\r\n       runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\r\n        Using Theano backend.\r\n        Loading and compiling models...\r\n        Loading training data...\r\n        Pre-processing images...\r\n        5331/5331 [==============================] - 934s   C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\theano\\tensor\\signal\\downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n          warnings.warn(\"downsample module has been moved to the pool module.\")\r\n        C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py:39: DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\r\n          X_test = X[:split, :, :, :]\r\n\r\n    \r\n    --------------------------------------------------\r\n    Training...\r\n    --------------------------------------------------\r\n    --------------------------------------------------\r\n    Iteration 1/150\r\n    --------------------------------------------------\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5542s - loss: 34.4657 - val_loss: 42.3799\r\n    Fitting diastole model...\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5523s - loss: 55.8717 - val_loss: 86.3558\r\n    Evaluating CRPS...\r\n    4265/4265 [==============================] - 1766s     \r\n    4265/4265 [==============================] - 1763s     \r\n    1066/1066 [==============================] - 441s     \r\n    1066/1066 [==============================] - 439s     \r\n    CRPS(train) = 0.0785559420969\r\n    CRPS(test) = 0.0743980740534\r\n    Saving weights...\r\n    --------------------------------------------------\r\n    Iteration 2/150\r\n    --------------------------------------------------\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 108055,
      "postDate": "2016-02-15T05:32:50.787Z",
      "rawMarkdown": ""
    },
    {
      "id": 107753,
      "postDate": "2016-02-12T08:54:25.113Z",
      "content": "<p>check how the y_train.npy file is created and what's inside </p>",
      "rawMarkdown": "check how the y_train.npy file is created and what's inside "
    },
    {
      "id": 107752,
      "postDate": "2016-02-12T08:54:13.250Z",
      "content": "<p>[quote=Danny Malter;107718]</p>\n\n<p>Can somebody please explain the difference between y_train[:0] and y_train[:,1].  I realize it is the first and second column of y_train, but why is the first column considered the systole data and second column considered diastole.</p>\n\n<p>[/quote]</p>\n\n<p>If you check <code>train.csv</code> file provided by Kaggle, they put like that there.</p>",
      "rawMarkdown": "[quote=Danny Malter;107718]\r\n\r\nCan somebody please explain the difference between y_train[:0] and y_train[:,1].  I realize it is the first and second column of y_train, but why is the first column considered the systole data and second column considered diastole.\r\n\r\n[/quote]\r\n\r\nIf you check ```train.csv``` file provided by Kaggle, they put like that there.\r\n"
    },
    {
      "id": 107718,
      "postDate": "2016-02-12T04:42:14.517Z",
      "content": "<p>Can somebody please explain the difference between y_train[:0] and y_train[:,1].  I realize it is the first and second column of y_train, but why is the first column considered the systole data and second column considered diastole.</p>\n\n<pre><code>        print('Fitting systole model...')\n        hist_systole = model_systole.fit(X_train_aug, y_train[:, 0], \n            shuffle=True, nb_epoch=epochs_per_iter,\n        batch_size=batch_size, validation_data=(X_test, y_test[:, 0]))\n\n        print('Fitting diastole model...')\n        hist_diastole = model_diastole.fit(X_train_aug, y_train[:, 1], \n            shuffle=True, nb_epoch=epochs_per_iter,\n        batch_size=batch_size, validation_data=(X_test, y_test[:, 1]))\n</code></pre>",
      "rawMarkdown": "Can somebody please explain the difference between y_train[:0] and y_train[:,1].  I realize it is the first and second column of y_train, but why is the first column considered the systole data and second column considered diastole.\r\n\r\n            print('Fitting systole model...')\r\n            hist_systole = model_systole.fit(X_train_aug, y_train[:, 0], \r\n                shuffle=True, nb_epoch=epochs_per_iter,\r\n            batch_size=batch_size, validation_data=(X_test, y_test[:, 0]))\r\n\r\n            print('Fitting diastole model...')\r\n            hist_diastole = model_diastole.fit(X_train_aug, y_train[:, 1], \r\n                shuffle=True, nb_epoch=epochs_per_iter,\r\n            batch_size=batch_size, validation_data=(X_test, y_test[:, 1]))"
    },
    {
      "id": 106929,
      "postDate": "2016-02-05T00:38:53.833Z",
      "content": "<p>[quote=hassiktir;106920]</p>\n\n<p>Hi, I had a general keras question: what is the difference between using nb_epoch vs looping of .fit() ?\nObviously if there is something like data augmentation in the loop, there will be a new random augmentation of the data before each fit but excluding this type of stuff, are they equivalent?\ni.e.\nis </p>\n\n<pre><code>for x in range(5):\n    model.fit(X,y)\n</code></pre>\n\n<p>&quot;equivalent to&quot; model.fit(X,y,nb_epoch=5)</p>\n\n<p>[/quote]</p>\n\n<p>Yes, it is basically the same thing. As you noticed, I used nb_epoch=1 in the tutorial just because of the data augmentation for each epoch.</p>",
      "rawMarkdown": "[quote=hassiktir;106920]\r\n\r\nHi, I had a general keras question: what is the difference between using nb_epoch vs looping of .fit() ?\r\nObviously if there is something like data augmentation in the loop, there will be a new random augmentation of the data before each fit but excluding this type of stuff, are they equivalent?\r\ni.e.\r\nis \r\n\r\n    for x in range(5):\r\n        model.fit(X,y)\r\n\r\n\"equivalent to\" model.fit(X,y,nb_epoch=5)\r\n\r\n\r\n[/quote]\r\n\r\nYes, it is basically the same thing. As you noticed, I used nb_epoch=1 in the tutorial just because of the data augmentation for each epoch.\r\n"
    },
    {
      "id": 106927,
      "postDate": "2016-02-04T23:46:20.150Z",
      "content": "<p>@Phunter - would you be able to post the ported model in Mxnet Tutoral discussion forum? I am especially intersted how you ported the linear regression model to Mxnet, as i have failed to do so succesfully</p>",
      "rawMarkdown": "@Phunter - would you be able to post the ported model in Mxnet Tutoral discussion forum? I am especially intersted how you ported the linear regression model to Mxnet, as i have failed to do so succesfully"
    },
    {
      "id": 106924,
      "postDate": "2016-02-04T23:30:40.060Z",
      "content": "<p>@WD since it is Keras topic and not a good place of discussing MXnet due to some reasons, please refer to MXnet's official IO document page <a href=\"http://mxnet.readthedocs.org/en/latest/python/io.html\">http://mxnet.readthedocs.org/en/latest/python/io.html</a> there image IO function has augment operations.</p>",
      "rawMarkdown": "@WD since it is Keras topic and not a good place of discussing MXnet due to some reasons, please refer to MXnet's official IO document page http://mxnet.readthedocs.org/en/latest/python/io.html there image IO function has augment operations."
    },
    {
      "id": 106920,
      "postDate": "2016-02-04T22:21:06.227Z",
      "content": "<p>Hi, I had a general keras question: what is the difference between using nb_epoch vs looping of .fit() ?\nObviously if there is something like data augmentation in the loop, there will be a new random augmentation of the data before each fit but excluding this type of stuff, are they equivalent?\ni.e.\nis </p>\n\n<pre><code>for x in range(5):\n    model.fit(X,y)\n</code></pre>\n\n<p>&quot;equivalent to&quot; model.fit(X,y,nb_epoch=5)</p>",
      "rawMarkdown": "Hi, I had a general keras question: what is the difference between using nb_epoch vs looping of .fit() ?\r\nObviously if there is something like data augmentation in the loop, there will be a new random augmentation of the data before each fit but excluding this type of stuff, are they equivalent?\r\ni.e.\r\nis \r\n\r\n    for x in range(5):\r\n        model.fit(X,y)\r\n\r\n\"equivalent to\" model.fit(X,y,nb_epoch=5)\r\n"
    },
    {
      "id": 106889,
      "postDate": "2016-02-04T18:05:08.013Z",
      "content": "<p>@phunter. i have been trying to port this linear regression script to mxnet, but have had trouble to get the model to reach good results. My systole CRPS values are 0.0550, rather than the 0.0350 values that i would get on the logistic approach on the MxNet tutoral. I am likely doing something wrong. Would you be willing to post your model? Potentially I am making mistakes in porting some of the Keras features (border_mode, etc) to MxNet. I am more than happy to potentially then help to upload this model to the MxNet depository to help others as well! Many thanks in advance! </p>",
      "rawMarkdown": "@phunter. i have been trying to port this linear regression script to mxnet, but have had trouble to get the model to reach good results. My systole CRPS values are 0.0550, rather than the 0.0350 values that i would get on the logistic approach on the MxNet tutoral. I am likely doing something wrong. Would you be willing to post your model? Potentially I am making mistakes in porting some of the Keras features (border_mode, etc) to MxNet. I am more than happy to potentially then help to upload this model to the MxNet depository to help others as well! Many thanks in advance! "
    },
    {
      "id": 106824,
      "postDate": "2016-02-04T03:40:46.443Z",
      "content": "<p>[quote=phunter;105652]</p>\n\n<p>[quote=WD;105630]</p>\n\n<p>@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.</p>\n\n<ul>\n<li><p>Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?</p></li>\n<li><p>You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? </p></li>\n<li><p>Does the code while len(images) &lt; 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? </p></li>\n</ul>\n\n<p>many thanks, W</p>\n\n<p>[/quote]</p>\n\n<p>Keras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) </p>\n\n<p>Update: Seems like this post has caused some confusions. Please, the statement above didn't mean any comparisons between MXnet and Keras. It was a simple run of MXnet with this tutorial's network from my own hardware. @Marko and I had different hardware, different image resolution etc, there was no apple-to-apple comparison of these two toolkits. Keras and Mxnet are both good deep learning frameworks, and it is the same with caffe/theano/torch/Lasagne etc which may later on have good tutorials for this Kaggle competition too, so please select your favorite one or ones for winning 200k $. And sorry to @Marko and @WD for this topic divergence. </p>\n\n<p>[/quote]</p>\n\n<p>@phunter  I also went through the MXNET tutorial for NDSB2. I noticed that it did not involve data augmentation, nor evaluation of validation. Perhaps due to its CSV loading fashion instead of reading from images directly? Do you have any idea to implement both augmentation and validation based on the MXNET tutorial?</p>",
      "rawMarkdown": "[quote=phunter;105652]\r\n\r\n[quote=WD;105630]\r\n\r\n@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.\r\n\r\n* Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?\r\n\r\n* You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? \r\n\r\n* Does the code while len(images) < 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? \r\n\r\nmany thanks, W\r\n \r\n\r\n[/quote]\r\n\r\nKeras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) \r\n\r\nUpdate: Seems like this post has caused some confusions. Please, the statement above didn't mean any comparisons between MXnet and Keras. It was a simple run of MXnet with this tutorial's network from my own hardware. @Marko and I had different hardware, different image resolution etc, there was no apple-to-apple comparison of these two toolkits. Keras and Mxnet are both good deep learning frameworks, and it is the same with caffe/theano/torch/Lasagne etc which may later on have good tutorials for this Kaggle competition too, so please select your favorite one or ones for winning 200k $. And sorry to @Marko and @WD for this topic divergence. \r\n\r\n[/quote]\r\n\r\n@phunter  I also went through the MXNET tutorial for NDSB2. I noticed that it did not involve data augmentation, nor evaluation of validation. Perhaps due to its CSV loading fashion instead of reading from images directly? Do you have any idea to implement both augmentation and validation based on the MXNET tutorial?\r\n"
    },
    {
      "id": 106603,
      "postDate": "2016-02-02T12:08:45.193Z",
      "content": "<p>@EIGSI. That is strange indeed. Will think about this more. Woudl you be able to check if your predictions (linear regression) are normally distributed around the actual values? I have a weird bug that they all seem to be larger (rather than normaly distributed...) </p>",
      "rawMarkdown": "@EIGSI. That is strange indeed. Will think about this more. Woudl you be able to check if your predictions (linear regression) are normally distributed around the actual values? I have a weird bug that they all seem to be larger (rather than normaly distributed...) "
    },
    {
      "id": 106532,
      "postDate": "2016-02-01T21:20:40.357Z",
      "content": "<p>I can't understand why the performance of the diastole network is consistently worse than that of systole? It is the same input data, same model why the difference?\nI noticed this on mxnet example too, which is using the frame differences. In terms of CRPS there is always a .01 difference and in terms of RMSE here the difference seems to be around 10.</p>\n\n<p>[quote=Marko Jocic;106509]</p>\n\n<p>You should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (&gt;0.04).</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I can't understand why the performance of the diastole network is consistently worse than that of systole? It is the same input data, same model why the difference?\r\nI noticed this on mxnet example too, which is using the frame differences. In terms of CRPS there is always a .01 difference and in terms of RMSE here the difference seems to be around 10.\r\n\r\n[quote=Marko Jocic;106509]\r\n\r\n \r\n\r\nYou should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (>0.04).\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 106483,
      "postDate": "2016-02-01T08:56:18.113Z",
      "content": "<p>I am moving my model from logistic regression based to linear regression based (with RSME metric). Just as a quick question - what are reasonable root mean squared error (RMSE) to observe during training? just order of magnitude? That will help me to check and compare with previous CRPS scores and whether i have implemented the model correctly. Currently my model has TRAIN-RSME around 25 and Validation-RSME around 95 - which seems high. Also - the predictions tend to be in nearly alll cases higher than the actual values for both diastole and systole - which seems odd. please let me know.</p>",
      "rawMarkdown": "I am moving my model from logistic regression based to linear regression based (with RSME metric). Just as a quick question - what are reasonable root mean squared error (RMSE) to observe during training? just order of magnitude? That will help me to check and compare with previous CRPS scores and whether i have implemented the model correctly. Currently my model has TRAIN-RSME around 25 and Validation-RSME around 95 - which seems high. Also - the predictions tend to be in nearly alll cases higher than the actual values for both diastole and systole - which seems odd. please let me know."
    },
    {
      "id": 106465,
      "postDate": "2016-02-01T03:17:08.880Z",
      "content": "<p>Thanks a lot for the tutorial. Since I am using a laptop I am trying to run iterations in chunks by saving model/weights, and reloading in later runs. I have added:</p>\n\n<blockquote>\n<pre><code>model_systole.load_weights('weights_systole.hdf5')\nmodel_diastole.load_weights('weights_diastole.hdf5')\n</code></pre>\n  \n  <p>To train,py, and model_from_json failed to load some reason after model_systole.to_json so utilized get_model like submission.py.  Also read val_loss saved from the last run.  Couple questions I have\n  1) Could it be weights_systole_best better to use than weights_systole weights?\n  2) Could I be missing anything else  to preload the previous model /run in train.py?\n  3) Is there any difference using get_model versus to_json/model_from_json? Any idea why the later is not working in this model?</p>\n</blockquote>",
      "rawMarkdown": "Thanks a lot for the tutorial. Since I am using a laptop I am trying to run iterations in chunks by saving model/weights, and reloading in later runs. I have added:\r\n>     model_systole.load_weights('weights_systole.hdf5')\r\n>     model_diastole.load_weights('weights_diastole.hdf5')\r\nTo train,py, and model_from_json failed to load some reason after model_systole.to_json so utilized get_model like submission.py.  Also read val_loss saved from the last run.  Couple questions I have\r\n1) Could it be weights_systole_best better to use than weights_systole weights?\r\n2) Could I be missing anything else  to preload the previous model /run in train.py?\r\n3) Is there any difference using get_model versus to_json/model_from_json? Any idea why the later is not working in this model?\r\n"
    },
    {
      "id": 106324,
      "postDate": "2016-01-30T00:27:05.387Z",
      "content": "<p>Yes, the recent changes have fixed the problem completely. Thanks so much for looking into this!</p>",
      "rawMarkdown": " Yes, the recent changes have fixed the problem completely. Thanks so much for looking into this!"
    },
    {
      "id": 106317,
      "postDate": "2016-01-29T22:50:23.433Z",
      "content": "<p>Thank you for confirming that it now works. Glad you had good result with it ;).</p>",
      "rawMarkdown": "Thank you for confirming that it now works. Glad you had good result with it ;)."
    },
    {
      "id": 106316,
      "postDate": "2016-01-29T22:48:08.543Z",
      "content": "<p>@Marko that was the problem.</p>\n\n<p>I was about to find the problem. I changed the submission.py to load the models and train data and calculate the CRPS on the train data to debug the problem. I did find that loading the best weights gave me  ~0.08 and loading the last weights gave me &lt;0.03.\nHowever, I did not get a chance to discover the bug when I saw your post.\nThanks.</p>",
      "rawMarkdown": "@Marko that was the problem.\r\n\r\nI was about to find the problem. I changed the submission.py to load the models and train data and calculate the CRPS on the train data to debug the problem. I did find that loading the best weights gave me  ~0.08 and loading the last weights gave me <0.03.\r\nHowever, I did not get a chance to discover the bug when I saw your post.\r\nThanks."
    },
    {
      "id": 106240,
      "postDate": "2016-01-29T10:49:08.637Z",
      "content": "<p>During the training, diastole weights from the best iteration are not saved properly.\nSystole weights were saved instead of diastole weights. So I just changed:</p>\n\n<p><code>model_systole.save_weights('weights_diastole_best.hdf5', overwrite=True)</code></p>\n\n<p>to</p>\n\n<p><code>model_diastole.save_weights('weights_diastole_best.hdf5', overwrite=True)</code></p>\n\n<p>I also just pushed this change to repo, so you can update it.\nSo there are two options, either retrain the model (which I recommend), or in submission.py change the weights loaded (weights_diastole.hdf5 instead of weights_diastole_best.hdf5).  Second option might not give the best result since it uses the weights from the last iteration, instead of weights from the best iteration.</p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "rawMarkdown": "During the training, diastole weights from the best iteration are not saved properly.\r\nSystole weights were saved instead of diastole weights. So I just changed:\r\n\r\n``` model_systole.save_weights('weights_diastole_best.hdf5', overwrite=True)```\r\n\r\nto\r\n\r\n``` model_diastole.save_weights('weights_diastole_best.hdf5', overwrite=True)```\r\n\r\nI also just pushed this change to repo, so you can update it.\r\nSo there are two options, either retrain the model (which I recommend), or in submission.py change the weights loaded (weights_diastole.hdf5 instead of weights_diastole_best.hdf5).  Second option might not give the best result since it uses the weights from the last iteration, instead of weights from the best iteration.\r\n\r\nCheers,\r\n\r\nMarko"
    },
    {
      "id": 106239,
      "postDate": "2016-01-29T10:35:00.850Z",
      "content": "<p>It seems that the problem was in val_loss.txt file, mine looks like:</p>\n\n<p>26.8219207778</p>\n\n<p>43.2115525287</p>\n\n<p>How about yours?</p>",
      "rawMarkdown": "It seems that the problem was in val_loss.txt file, mine looks like:\r\n\r\n26.8219207778\r\n\r\n43.2115525287\r\n\r\nHow about yours?"
    },
    {
      "id": 106235,
      "postDate": "2016-01-29T10:02:06.053Z",
      "content": "<p>[quote=piotr;106225]</p>\n\n<p>@Marko, @patruf @ Wei Wu,</p>\n\n<p>I was able to get 0.359 with Marko code, but I made changes in submission.py or in val_los.txt files - I don't remember precisely now. But the network trained with train.py is OK. There is only a problem with submission.py.</p>\n\n<p>[/quote]</p>\n\n<p>Would you bother to check or recall what did you change? I'm getting blind here trying to see what could cause the problem...</p>",
      "rawMarkdown": "[quote=piotr;106225]\r\n\r\n@Marko, @patruf @ Wei Wu,\r\n\r\nI was able to get 0.359 with Marko code, but I made changes in submission.py or in val_los.txt files - I don't remember precisely now. But the network trained with train.py is OK. There is only a problem with submission.py.\r\n\r\n[/quote]\r\n\r\nWould you bother to check or recall what did you change? I'm getting blind here trying to see what could cause the problem..."
    },
    {
      "id": 106226,
      "postDate": "2016-01-29T09:17:53.857Z",
      "content": "<p>There is probably a problem with submission generation. Will look into it now.</p>",
      "rawMarkdown": "There is probably a problem with submission generation. Will look into it now."
    },
    {
      "id": 106199,
      "postDate": "2016-01-29T01:45:58.937Z",
      "content": "<p>Nope, got 0.084... just now, and the warning was just &quot;downsample module has been moved to the pool module&quot; but I don't think that should change anything.</p>",
      "rawMarkdown": " Nope, got 0.084... just now, and the warning was just \"downsample module has been moved to the pool module\" but I don't think that should change anything."
    },
    {
      "id": 106150,
      "postDate": "2016-01-28T20:15:42.907Z",
      "content": "<p>I am using it right now to generate a submission. I did have one question; are the original image files necessary to keep after the .npy files have been generated? The .npy files are much smaller in size and would fit much better on the SSD connected directly to the motherboard, whereas the original dataset image files need to reside on an external drive. Thanks very much for this! I was looking to start using Keras for this and other image recognition challenges!</p>",
      "rawMarkdown": "I am using it right now to generate a submission. I did have one question; are the original image files necessary to keep after the .npy files have been generated? The .npy files are much smaller in size and would fit much better on the SSD connected directly to the motherboard, whereas the original dataset image files need to reside on an external drive. Thanks very much for this! I was looking to start using Keras for this and other image recognition challenges!"
    },
    {
      "id": 106119,
      "postDate": "2016-01-28T16:42:41.573Z",
      "content": "<p>I tried to check on it but didn't have time this morning. I decided to start over and will check the results tonight. The only thing I changed was in the submission.py file the fo.writerow(fi.next()) to __ next __() for python 3. Also, I get a Keras warning when I run, something similar to warning: dropout module has been moved to pool, but I didn't think it was important since everything else seemed to work. Thanks for all your responses, I will figure it out. I really like the model, I think everything is well documented and understandable.</p>",
      "rawMarkdown": " I tried to check on it but didn't have time this morning. I decided to start over and will check the results tonight. The only thing I changed was in the submission.py file the fo.writerow(fi.next()) to __ next __() for python 3. Also, I get a Keras warning when I run, something similar to warning: dropout module has been moved to pool, but I didn't think it was important since everything else seemed to work. Thanks for all your responses, I will figure it out. I really like the model, I think everything is well documented and understandable."
    },
    {
      "id": 106117,
      "postDate": "2016-01-28T15:58:39.373Z",
      "content": "<p>@patruff</p>\n\n<p>Yes, they do. Since you got CRPS(test) = ~0.0347, you should expect somewhat similar result on submission (probably just a bit higher though).\nSince the fix, I had multiple people confirming that the code now works properly (I also checked myself to make sure). From the top of my head, I can't think of anything right now that would produce such a bad result. Hm, maybe submission.csv file wasn't overwritten properly, could you check?</p>",
      "rawMarkdown": "@patruff\r\n\r\nYes, they do. Since you got CRPS(test) = ~0.0347, you should expect somewhat similar result on submission (probably just a bit higher though).\r\nSince the fix, I had multiple people confirming that the code now works properly (I also checked myself to make sure). From the top of my head, I can't think of anything right now that would produce such a bad result. Hm, maybe submission.csv file wasn't overwritten properly, could you check?"
    },
    {
      "id": 106077,
      "postDate": "2016-01-28T10:30:35.460Z",
      "content": "<p>@piotr</p>\n\n<p>You are welcome. Thanks for your understanding and patience!</p>",
      "rawMarkdown": "@piotr\r\n\r\nYou are welcome. Thanks for your understanding and patience!"
    },
    {
      "id": 106047,
      "postDate": "2016-01-28T03:14:27.973Z",
      "content": "<p>Maybe it's related: I had my PR just merged which fixed a bug in kera's ImageDataGenerator <a href=\"https://github.com/fchollet/keras/issues/1551\">https://github.com/fchollet/keras/issues/1551</a></p>",
      "rawMarkdown": "Maybe it's related: I had my PR just merged which fixed a bug in kera's ImageDataGenerator https://github.com/fchollet/keras/issues/1551"
    },
    {
      "id": 106031,
      "postDate": "2016-01-28T00:00:08.487Z",
      "content": "<p>For anyone interested and using older graphics cards (like myself, using cuda 3.0) I made a docker image that uses tensorflow gpu + keras with python3.  Images are a bit big and would love feedback to make them smaller but need to put the dockerfiles somewhere.  </p>\n\n<p>docker pull grahama/tf:keras</p>\n\n<p>Also should probably note to run it, best to use <a href=\"https://github.com/tensorflow/tensorflow/blob/master/tensorflow/tools/docker/docker_run_gpu.sh\">https://github.com/tensorflow/tensorflow/blob/master/tensorflow/tools/docker/docker_run_gpu.sh</a>\nsince official nvidia-docker thing was not working for me.\nso for instance to run a container with data/ and these python files in /root/</p>\n\n<p>docker_run_gpu -v /root/:/root/ grahama/tf:keras</p>\n\n<p>then in container:</p>\n\n<p>python3 data.py &amp;&amp; python3 train.py</p>\n\n<p>One last thing, \nawesome that fchollet is here, love (and have contributed to) keras.  Would anyone be able to explain the significance of batch_size in relation to loss and val_loss?  I thought I had read it was 'just' a speed thing (and that smaller was possibly better <a href=\"https://github.com/fchollet/keras/issues/68#issuecomment-95413935\">https://github.com/fchollet/keras/issues/68#issuecomment-95413935</a> ) but maybe I am thinking of a different batch_size since it seems like small batch_size for this script is faster than larger batch_size but makes accuracy poor. </p>",
      "rawMarkdown": "For anyone interested and using older graphics cards (like myself, using cuda 3.0) I made a docker image that uses tensorflow gpu + keras with python3.  Images are a bit big and would love feedback to make them smaller but need to put the dockerfiles somewhere.  \r\n\r\ndocker pull grahama/tf:keras\r\n\r\nAlso should probably note to run it, best to use https://github.com/tensorflow/tensorflow/blob/master/tensorflow/tools/docker/docker_run_gpu.sh\r\nsince official nvidia-docker thing was not working for me.\r\nso for instance to run a container with data/ and these python files in /root/\r\n\r\ndocker_run_gpu -v /root/:/root/ grahama/tf:keras\r\n\r\nthen in container:\r\n\r\npython3 data.py && python3 train.py\r\n\r\n\r\nOne last thing, \r\nawesome that fchollet is here, love (and have contributed to) keras.  Would anyone be able to explain the significance of batch_size in relation to loss and val_loss?  I thought I had read it was 'just' a speed thing (and that smaller was possibly better https://github.com/fchollet/keras/issues/68#issuecomment-95413935 ) but maybe I am thinking of a different batch_size since it seems like small batch_size for this script is faster than larger batch_size but makes accuracy poor. "
    },
    {
      "id": 105980,
      "postDate": "2016-01-27T19:42:39.967Z",
      "content": "<p>Marko, no apology necessary. Thanks for fixing things so quickly. I'm going to try it out again later tonight.</p>",
      "rawMarkdown": " Marko, no apology necessary. Thanks for fixing things so quickly. I'm going to try it out again later tonight."
    },
    {
      "id": 105939,
      "postDate": "2016-01-27T14:46:45.823Z",
      "content": "<p>Make that a third person. 0.08... for me without changing anything except I ran it using python3. Thanks for looking into it Marko.</p>",
      "rawMarkdown": " Make that a third person. 0.08... for me without changing anything except I ran it using python3. Thanks for looking into it Marko."
    },
    {
      "id": 105810,
      "postDate": "2016-01-26T21:07:00.347Z",
      "content": "<p>Thanks Marko for sharing this tutorial and your explanations!</p>",
      "rawMarkdown": " Thanks Marko for sharing this tutorial and your explanations!"
    },
    {
      "id": 105808,
      "postDate": "2016-01-26T20:50:40.137Z",
      "content": "<p>Thx again Marko.</p>",
      "rawMarkdown": "Thx again Marko."
    },
    {
      "id": 105793,
      "postDate": "2016-01-26T19:14:18.280Z",
      "content": "<p>@WD</p>\n\n<p>Ok, so basically every model is imperfect/imprecise - i.e. sometimes it gets things right, sometimes it doesn't. We usually represent this imprecision by value of loss function - the lower it is the better our model. So, in a tutorial I use RMS error as a loss function, which means if value of loss function is, say, 20 - that is an indicator that my model usually misses the true value by ~20ml. I wanted to incorporate this uncertainty in the calculation of CDF - ideally, if loss was 0 (model predicts everything right) the CDF would be a step function, but if it is &gt;0 the CDF has that sigmoid-like part around the predicted volume. The larger the uncertainty, the wider is that sigmoid-like area. Because of this, using RMSE value for sigma seemed very natural.</p>\n\n<p>hist_systole.history['loss'][-1] is a Keras specific part of code, since function fit() returns a History object, which has a current value of loss function both on train and test split ('loss' and 'val_loss', respectively). Since hist_systole.history['loss'] returns an array, I just pick up the last value with [-1].</p>\n\n<p>Also, the fundamental advantage of predicting one real value to 600 [0,1] values is simple because it is more natural for this problem! (well, to me at least). If the real volume is N and you predict volume N+1, its no big deal, RMSE is just 1. But imagine you have 600 [0,1] values, even if the real volume corresponds to N-th output and if you predict it is on N+1-th output, the error would be same as if you predicted it is on N+10-th output. That is you are using softmax at output (1 of N classification). I assume it would be better to use sigmoids, but still I found using linear regression as the most appealing approach.</p>",
      "rawMarkdown": "@WD\r\n\r\nOk, so basically every model is imperfect/imprecise - i.e. sometimes it gets things right, sometimes it doesn't. We usually represent this imprecision by value of loss function - the lower it is the better our model. So, in a tutorial I use RMS error as a loss function, which means if value of loss function is, say, 20 - that is an indicator that my model usually misses the true value by ~20ml. I wanted to incorporate this uncertainty in the calculation of CDF - ideally, if loss was 0 (model predicts everything right) the CDF would be a step function, but if it is >0 the CDF has that sigmoid-like part around the predicted volume. The larger the uncertainty, the wider is that sigmoid-like area. Because of this, using RMSE value for sigma seemed very natural.\r\n\r\nhist_systole.history['loss'][-1] is a Keras specific part of code, since function fit() returns a History object, which has a current value of loss function both on train and test split ('loss' and 'val_loss', respectively). Since hist_systole.history['loss'] returns an array, I just pick up the last value with [-1].\r\n\r\nAlso, the fundamental advantage of predicting one real value to 600 [0,1] values is simple because it is more natural for this problem! (well, to me at least). If the real volume is N and you predict volume N+1, its no big deal, RMSE is just 1. But imagine you have 600 [0,1] values, even if the real volume corresponds to N-th output and if you predict it is on N+1-th output, the error would be same as if you predicted it is on N+10-th output. That is you are using softmax at output (1 of N classification). I assume it would be better to use sigmoids, but still I found using linear regression as the most appealing approach.\r\n"
    },
    {
      "id": 105738,
      "postDate": "2016-01-26T10:49:14.913Z",
      "content": "<p>To use gpu add this to the  import section of your python script</p>\n\n<p>import os</p>\n\n<p>os.environ['THEANO_FLAGS'] = 'device=gpu'</p>",
      "rawMarkdown": "To use gpu add this to the  import section of your python script\r\n\r\nimport os\r\n\r\nos.environ['THEANO_FLAGS'] = 'device=gpu'"
    },
    {
      "id": 105719,
      "postDate": "2016-01-26T04:26:20.383Z",
      "content": "<p>One more question please. Iteration 1 took 6-7 hours on my system (XPS8100, no decent GPU). Does that sound right given this poor hardware, or I should suspect some problem with the software? </p>\n\n<pre><code>   runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\n    Using Theano backend.\n    Loading and compiling models...\n    Loading training data...\n    Pre-processing images...\n    5331/5331 [==============================] - 934s   C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\theano\\tensor\\signal\\downsample.py:5: UserWarning: downsample module has been moved to the pool module.\n      warnings.warn(&quot;downsample module has been moved to the pool module.&quot;)\n    C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py:39: DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\n      X_test = X[:split, :, :, :]\n\n\n--------------------------------------------------\nTraining...\n--------------------------------------------------\n--------------------------------------------------\nIteration 1/150\n--------------------------------------------------\nEpoch 1/1\n4290/4265 [==============================] - 5542s - loss: 34.4657 - val_loss: 42.3799\nFitting diastole model...\nEpoch 1/1\n4290/4265 [==============================] - 5523s - loss: 55.8717 - val_loss: 86.3558\nEvaluating CRPS...\n4265/4265 [==============================] - 1766s     \n4265/4265 [==============================] - 1763s     \n1066/1066 [==============================] - 441s     \n1066/1066 [==============================] - 439s     \nCRPS(train) = 0.0785559420969\nCRPS(test) = 0.0743980740534\nSaving weights...\n--------------------------------------------------\nIteration 2/150\n--------------------------------------------------\n</code></pre>",
      "rawMarkdown": "One more question please. Iteration 1 took 6-7 hours on my system (XPS8100, no decent GPU). Does that sound right given this poor hardware, or I should suspect some problem with the software? \r\n\r\n\r\n \r\n\r\n       runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\r\n        Using Theano backend.\r\n        Loading and compiling models...\r\n        Loading training data...\r\n        Pre-processing images...\r\n        5331/5331 [==============================] - 934s   C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\theano\\tensor\\signal\\downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n          warnings.warn(\"downsample module has been moved to the pool module.\")\r\n        C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py:39: DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\r\n          X_test = X[:split, :, :, :]\r\n\r\n    \r\n    --------------------------------------------------\r\n    Training...\r\n    --------------------------------------------------\r\n    --------------------------------------------------\r\n    Iteration 1/150\r\n    --------------------------------------------------\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5542s - loss: 34.4657 - val_loss: 42.3799\r\n    Fitting diastole model...\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5523s - loss: 55.8717 - val_loss: 86.3558\r\n    Evaluating CRPS...\r\n    4265/4265 [==============================] - 1766s     \r\n    4265/4265 [==============================] - 1763s     \r\n    1066/1066 [==============================] - 441s     \r\n    1066/1066 [==============================] - 439s     \r\n    CRPS(train) = 0.0785559420969\r\n    CRPS(test) = 0.0743980740534\r\n    Saving weights...\r\n    --------------------------------------------------\r\n    Iteration 2/150\r\n    --------------------------------------------------\r\n"
    },
    {
      "id": 105705,
      "postDate": "2016-01-25T23:48:36.403Z",
      "content": "<p>Many thanks! </p>\n\n<p>PS: for windows users, if git gives you some trouble, this <a href=\"https://github.com/notanumber/gitst2/issues/10\">link</a> might help:  </p>",
      "rawMarkdown": "Many thanks! \r\n\r\nPS: for windows users, if git gives you some trouble, this [link][1] might help:  \r\n\r\n[1]: https://github.com/notanumber/gitst2/issues/10"
    },
    {
      "id": 105697,
      "postDate": "2016-01-25T22:27:11.490Z",
      "content": "<p>Many thanks for this keras tutorial! </p>\n\n<p>Data.py ran well (Anaconda/Spyder, windows 64). Keras installation seems ok, but Train.py won't work. Any idea how to fix the following? </p>\n\n<pre><code>&gt;&gt;&gt; runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\nUsing Theano backend.\nLoading and compiling models...\nTraceback (most recent call last):\n  File &quot;&lt;stdin&gt;&quot;, line 1, in &lt;module&gt;\n  File &quot;C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\spyderlib\\widgets\\externalshell\\sitecustomize.py&quot;, line 699, in runfile\n    execfile(filename, namespace)\n</code></pre>\n\n<p>(more in Woolsey_error.txt attached)</p>\n\n<pre><code>  File &quot;C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\keras\\backend\\theano_backend.py&quot;, line 463, in relu\n    x = T.nnet.relu(x, alpha)\nAttributeError: 'module' object has no attribute 'relu'\n&gt;&gt;&gt; \n</code></pre>",
      "rawMarkdown": "Many thanks for this keras tutorial! \r\n\r\nData.py ran well (Anaconda/Spyder, windows 64). Keras installation seems ok, but Train.py won't work. Any idea how to fix the following? \r\n\r\n    >>> runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\r\n    Using Theano backend.\r\n    Loading and compiling models...\r\n    Traceback (most recent call last):\r\n      File \"<stdin>\", line 1, in <module>\r\n      File \"C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\spyderlib\\widgets\\externalshell\\sitecustomize.py\", line 699, in runfile\r\n        execfile(filename, namespace)\r\n(more in Woolsey_error.txt attached)\r\n\r\n      File \"C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\keras\\backend\\theano_backend.py\", line 463, in relu\r\n        x = T.nnet.relu(x, alpha)\r\n    AttributeError: 'module' object has no attribute 'relu'\r\n    >>> "
    },
    {
      "id": 105684,
      "postDate": "2016-01-25T21:01:43.110Z",
      "content": "<p>It can be hard to see that there is no need to train 2 separate models, I had to convince myself that this was the case by answering a few questions.</p>\n\n<p>Assuming there was no systolic volume in the competition just diastolic </p>\n\n<ol>\n<li>Will training a model to predict patient 1 &#8211; 500 labels for patient give you the diastolic distribution? YES\nAssuming there was no diastolic volume in the competition just systolic </li>\n<li>Will training a model to predict patient 1 &#8211; 500 labels for patient give you the systolic distribution? YES</li>\n</ol>\n\n<p><strong>a.ha mo.ment</strong></p>\n\n<ol>\n<li>Is there any difference in the labeling and training procedure for the 2 models above? NO (same label same data)</li>\n</ol>\n\n<p>Training a network to predict patient learns both the systolic and diastolic volumes simultaneously :) </p>\n\n<p><strong>Send the check in the mail&#8230;</strong></p>",
      "rawMarkdown": "It can be hard to see that there is no need to train 2 separate models, I had to convince myself that this was the case by answering a few questions.\r\n\r\nAssuming there was no systolic volume in the competition just diastolic \r\n\r\n 1. Will training a model to predict patient 1 – 500 labels for patient give you the diastolic distribution? YES\r\nAssuming there was no diastolic volume in the competition just systolic \r\n 2. Will training a model to predict patient 1 – 500 labels for patient give you the systolic distribution? YES\r\n\r\n\r\n**a.ha mo.ment**\r\n\r\n 1. Is there any difference in the labeling and training procedure for the 2 models above? NO (same label same data)\r\n\r\nTraining a network to predict patient learns both the systolic and diastolic volumes simultaneously :) \r\n\r\n**Send the check in the mail…**\r\n"
    },
    {
      "id": 105679,
      "postDate": "2016-01-25T19:49:02.583Z",
      "content": "<p>cuDNN v3 gave me 5x speedup per epoch... using lasagne I had to jump through hoops to get it installed on my aws gpu because NVidia makes you register before downloading it, makes no sense!..I am still waiting on my registration approval.  A little trick I tried to avoid training 2 models was instead of predicting systole and diastole vols, Just predict patient 0-500 labels. Train a single model then translate the patient predictions to systole and diastole volumes. This got me a score of 0.040xxx. Which was as good as I got with 2 separate models for systole and diastole. I am still not convinced this is not a viable strategy. What is the difference btw a network trained for systole and one trained for diastole? Absolutely nothing. The same features are learnt by the net;  </p>",
      "rawMarkdown": "cuDNN v3 gave me 5x speedup per epoch... using lasagne I had to jump through hoops to get it installed on my aws gpu because NVidia makes you register before downloading it, makes no sense!..I am still waiting on my registration approval.  A little trick I tried to avoid training 2 models was instead of predicting systole and diastole vols, Just predict patient 0-500 labels. Train a single model then translate the patient predictions to systole and diastole volumes. This got me a score of 0.040xxx. Which was as good as I got with 2 separate models for systole and diastole. I am still not convinced this is not a viable strategy. What is the difference btw a network trained for systole and one trained for diastole? Absolutely nothing. The same features are learnt by the net;  "
    },
    {
      "id": 105675,
      "postDate": "2016-01-25T19:16:46.843Z",
      "content": "<p>[quote=phunter;105673]</p>\n\n<p>I guess so, haven't had chance to run Keras on my machine but will try. Maybe Keras can try training the model one by one?</p>\n\n<p>[/quote]</p>\n\n<p>Of course, people can modify the code for themselves and train the models one by one.</p>",
      "rawMarkdown": "[quote=phunter;105673]\r\n\r\nI guess so, haven't had chance to run Keras on my machine but will try. Maybe Keras can try training the model one by one?\r\n\r\n[/quote]\r\n\r\n\r\nOf course, people can modify the code for themselves and train the models one by one."
    },
    {
      "id": 105673,
      "postDate": "2016-01-25T19:13:29.217Z",
      "content": "<p>[quote=Marko Jocic;105662]</p>\n\n<p>@phunter</p>\n\n<p>From the top of my head, I think the main difference is that in our example the two models are trained basically at the same time (which doubles the memory usage), while MXnet example trains them one by one.</p>\n\n<p>[/quote]\nI guess so, haven't had chance to run Keras on my machine but will try. Maybe Keras can try training the model one by one?</p>",
      "rawMarkdown": "[quote=Marko Jocic;105662]\r\n\r\n@phunter\r\n\r\nFrom the top of my head, I think the main difference is that in our example the two models are trained basically at the same time (which doubles the memory usage), while MXnet example trains them one by one.\r\n\r\n[/quote]\r\nI guess so, haven't had chance to run Keras on my machine but will try. Maybe Keras can try training the model one by one?"
    },
    {
      "id": 105635,
      "postDate": "2016-01-25T13:52:49.587Z",
      "content": "<p>@Marko - many thanks. and thanks for the well-documented code. </p>\n\n<ul>\n<li><p>I noticed that you are training on the frames themselves (rather than the difference between frames - as explored by others). It might be interesting by folks to see if this would make a difference</p></li>\n<li><p>Marko - out of curiosity - and from your experience - what do you think the margin of improvement might be if one changes the convolution window sizes, the learning rates, the weight decay and other parameters? Do you feel that this might shave off only a little bit (if at all), or do you think that parameter optimization can make a big difference? So far - my experience is that the real gains lie in preprocessing, and that most network configurations give roughly similar results - but would love to get your expertise and perspective!</p></li>\n</ul>",
      "rawMarkdown": "@Marko - many thanks. and thanks for the well-documented code. \r\n\r\n* I noticed that you are training on the frames themselves (rather than the difference between frames - as explored by others). It might be interesting by folks to see if this would make a difference\r\n\r\n* Marko - out of curiosity - and from your experience - what do you think the margin of improvement might be if one changes the convolution window sizes, the learning rates, the weight decay and other parameters? Do you feel that this might shave off only a little bit (if at all), or do you think that parameter optimization can make a big difference? So far - my experience is that the real gains lie in preprocessing, and that most network configurations give roughly similar results - but would love to get your expertise and perspective!"
    },
    {
      "id": 105630,
      "postDate": "2016-01-25T13:01:31.857Z",
      "content": "<p>@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.</p>\n\n<ul>\n<li><p>Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?</p></li>\n<li><p>You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? </p></li>\n<li><p>Does the code while len(images) &lt; 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? </p></li>\n</ul>\n\n<p>many thanks, W</p>",
      "rawMarkdown": "@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.\r\n\r\n* Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?\r\n\r\n* You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? \r\n\r\n* Does the code while len(images) < 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? \r\n\r\nmany thanks, W\r\n "
    },
    {
      "id": 105610,
      "postDate": "2016-01-25T05:59:59.780Z",
      "content": "<p>@fchollet: Thanks, upgrading to Keras 0.3.1 from 0.3.0 fixed the border_mode issue. </p>",
      "rawMarkdown": "@fchollet: Thanks, upgrading to Keras 0.3.1 from 0.3.0 fixed the border_mode issue. "
    },
    {
      "id": 105604,
      "postDate": "2016-01-25T04:14:39.930Z",
      "content": "<p>Thanks posting keras tutorial. Is keras needs to be updated? Mine version is only a month old, using python 2.7.10, I've tried train.py but it fails as:</p>\n\n<pre><code>X = self.get_input(train)\n</code></pre>\n\n<p>File &quot;/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py&quot;, line 102, in get_input\n    return self.previous.get_output(train=train)\n  File &quot;/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py&quot;, line 512, in get_output\n    X = self.get_input(train)\n  File &quot;/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py&quot;, line 102, in get_input\n    return self.previous.get_output(train=train)\n  File &quot;/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/convolutional.py&quot;, line 215, in get_output\n    dim_ordering=self.dim_ordering)\n  File &quot;/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/backend/theano_backend.py&quot;, line 543, in conv2d\n    border_mode=(pad_x, pad_y))\n  File &quot;/home/esorar/theano/t_env/lib/python2.7/site-packages/theano/sandbox/cuda/dnn.py&quot;, line 1192, in dnn_conv\n    conv_mode=conv_mode, precision=precision)(img.shape,\n  File &quot;/home/esorar/theano/t_env/lib/python2.7/site-packages/theano/sandbox/cuda/dnn.py&quot;, line 260, in <strong>init</strong>\n    border_mode = tuple(map(int, border_mode))\nTypeError: int() argument must be a string or a number, not 'TensorVariable'</p>",
      "rawMarkdown": "Thanks posting keras tutorial. Is keras needs to be updated? Mine version is only a month old, using python 2.7.10, I've tried train.py but it fails as:\r\n\r\n    X = self.get_input(train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py\", line 102, in get_input\r\n    return self.previous.get_output(train=train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py\", line 512, in get_output\r\n    X = self.get_input(train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py\", line 102, in get_input\r\n    return self.previous.get_output(train=train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/convolutional.py\", line 215, in get_output\r\n    dim_ordering=self.dim_ordering)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/backend/theano_backend.py\", line 543, in conv2d\r\n    border_mode=(pad_x, pad_y))\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/theano/sandbox/cuda/dnn.py\", line 1192, in dnn_conv\r\n    conv_mode=conv_mode, precision=precision)(img.shape,\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/theano/sandbox/cuda/dnn.py\", line 260, in __init__\r\n    border_mode = tuple(map(int, border_mode))\r\nTypeError: int() argument must be a string or a number, not 'TensorVariable'"
    },
    {
      "id": 111322,
      "postDate": "2016-03-13T18:13:48.620Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 111314,
      "postDate": "2016-03-13T16:13:36.947Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 105662,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-01-25T18:26:08.057000",
      "content": "<p>@phunter</p>\n\n<p>From the top of my head, I think the main difference is that in our example the two models are trained basically at the same time (which doubles the memory usage), while MXnet example trains them one by one.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 105608,
      "author_name": "fchollet",
      "author_url": "",
      "post_date": "2016-01-25T05:39:42.757000",
      "content": "<p>@ertuka: yes, you should update your version of Keras, from the Keras repository:</p>\n\n<pre><code>pip install git+git://github.com/fchollet/keras.git --upgrade --no-deps\n</code></pre>\n\n<p>(may need to be prefixed with &quot;sudo&quot;)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 105658,
      "author_name": "phunter",
      "author_url": "",
      "post_date": "2016-01-25T17:10:38.827000",
      "content": "<p>[quote=fchollet;105655]</p>\n\n<p>This is a pretty ridiculous attempt at saying &quot;lol MXnet is faster than Keras&quot;. No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. </p>\n\n<p>Stay classy man...</p>\n\n<p>[/quote]</p>\n\n<p>Oh, no, I didn't mean MXnet was faster because of what what, why downvote me? @Marko Jocic said 128x128 was too large to fit in his GTX 770 video card, and I said I could run 128x128 with mxnet on GTX 960, an even slower video card, so @WD didn't have to worry about memory of running it. We ran on different cards different resolutions different CPUs for data augment etc, and there was no comparisons. </p>\n\n<p>Stay calm :-)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 105919,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "2016-01-27T12:46:11.080000",
      "content": "<p>Hi!</p>\n\n<p>Thank you sharing your model!</p>\n\n<p>When I submit this model code, get the high errors(~0.077926).(same piotr happening)\nThis code run 150 epochs running.</p>\n\n<p>So I think this code is influented random state or other cause.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 105652,
      "author_name": "phunter",
      "author_url": "",
      "post_date": "2016-01-25T16:38:58.690000",
      "content": "<p>[quote=WD;105630]</p>\n\n<p>@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.</p>\n\n<ul>\n<li><p>Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?</p></li>\n<li><p>You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? </p></li>\n<li><p>Does the code while len(images) &lt; 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? </p></li>\n</ul>\n\n<p>many thanks, W</p>\n\n<p>[/quote]</p>\n\n<p>Keras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) </p>\n\n<p>Update: Seems like this post has caused some confusions. Please, the statement above didn't mean any comparisons between MXnet and Keras. It was a simple run of MXnet with this tutorial's network from my own hardware. @Marko and I had different hardware, different image resolution etc, there was no apple-to-apple comparison of these two toolkits. Keras and Mxnet are both good deep learning frameworks, and it is the same with caffe/theano/torch/Lasagne etc which may later on have good tutorials for this Kaggle competition too, so please select your favorite one or ones for winning 200k $. And sorry to @Marko and @WD for this topic divergence. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 105889,
      "author_name": "Piotr",
      "author_url": "",
      "post_date": "2016-01-27T07:52:35.513000",
      "content": "<p>Hi!</p>\n\n<p>Thanks for this tutorial! I cloned github code and run</p>\n\n<pre><code>python data.py\npython train.py\npython submission.py\n</code></pre>\n\n<p>and this gave me    0.074981 on LB, I didn't make any changes in the code. Does anyone have idea what can goes wrong? I expect LB ~ 0.0359</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 111308,
      "author_name": "DavidGbodiOdaibo",
      "author_url": "",
      "post_date": "2016-03-13T14:52:17.347000",
      "content": "<p>you guys are now 1st just wondering what you did? I know you are not using the model in this tutorial. It also looks like you did not upload your model so no prize money :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111178,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-03-12T12:20:24.693000",
      "content": "<p>I tried with vgg model on 192 x192 images and tweaking some parameters,  got 0.0101. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110930,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-03-09T18:18:53.763000",
      "content": "<p>[quote=datapool;110927]</p>\n\n<p>Thanks for the feedback.\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper &quot;ImageNet Classification with Deep Convolutional Neural Networks&quot;.\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>For the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110921,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-03-09T17:24:58.533000",
      "content": "<p>[quote=datapool;110899]</p>\n\n<p>Hello Marko Jocic,</p>\n\n<p>Thanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?</p>\n\n<p>[/quote]</p>\n\n<p>Hi, yes we also had that approach, and it yielded somewhat smaller accuracy, but on the other side double faster training, so it's definitely OK for more model diversity.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 110377,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-03-04T23:09:58.893000",
      "content": "<p>[quote=JasonMay;109956]</p>\n\n<p>My question is do you know where all the shape must be addressed to alter image size to 128X128?</p>\n\n<p>Thanks,\nJason</p>\n\n<p>[/quote]</p>\n\n<p>You would need a bigger/beter model for 128x128 images.</p>\n\n<p>[quote=Gongning Luo;110012]</p>\n\n<p>How big of memory(RAM), the parameter of GPU and CPU?</p>\n\n<p>[/quote]</p>\n\n<p>You need ~1GB of GPU RAM for this tutorial.</p>\n\n<p>[quote=Lawrence Chernin;110353]</p>\n\n<p>I created a public AWS image: ami-d42a59b4 that has all the needed libraries to run this code.\nRunning on g2.2xlarge. (sorry, you may also need  to install cython and h5py )</p>\n\n<p>[/quote]</p>\n\n<p>Would you care sharing the image with the rest of Kagglers?</p>\n\n<p>Thanks! Marko</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 106509,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-02-01T14:25:10.947000",
      "content": "<p>[quote=Ertuka;106465]</p>\n\n<p>1) Could it be weights_systole_best better to use than weights_systole weights?\n2) Could I be missing anything else  to preload the previous model /run in train.py?\n3) Is there any difference using get_model versus to_json/model_from_json? Any idea why the later is not working in this model?</p>\n\n<p>[/quote]</p>\n\n<p>1) You could try both. Both should eventually converge more or less the same.</p>\n\n<p>2) 3) That should be it. It might be a bug currently in Keras, maybe the custom activation function (in first layer) can't be saved/loaded with to_json/from_json. You can check that. However, get_model() should work the same.</p>\n\n<p>[quote=WDharvard;106483]</p>\n\n<p>what are reasonable root mean squared error (RMSE) to observe during training?</p>\n\n<p>[/quote]</p>\n\n<p>You should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (&gt;0.04).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 106225,
      "author_name": "Piotr",
      "author_url": "",
      "post_date": "2016-01-29T09:16:56.503000",
      "content": "<p>@Marko, @patruf @ Wei Wu,</p>\n\n<p>I was able to get 0.359 with Marko code, but I made changes in submission.py or in val_los.txt files - I don't remember precisely now. But the network trained with train.py is OK. There is only a problem with submission.py.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 106213,
      "author_name": "Wei Wu",
      "author_url": "",
      "post_date": "2016-01-29T04:25:56.280000",
      "content": "<p>@Marko , @patruff</p>\n\n<p>I can confirm the same problem as @patruff.\nWhen I run the model training, at the end the test CRPS was about 0.029.\nAfter I run submission.py and submit the file, the LB score was about 0.08. (It was a fresh run there was no old files.)</p>\n\n<p>I am running python 2.7 under a windows environment.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 106106,
      "author_name": "patruff",
      "author_url": "",
      "post_date": "2016-01-28T14:36:23.987000",
      "content": "<h2>Iteration 200/200</h2>\n\n<p>Augmenting images - rotations\n4265/4265 [==============================] - 33s <br>\nAugmenting images - shifts\n4265/4265 [==============================] - 8s <br>\nFitting systole model...\nTrain on 4265 samples, validate on 1066 samples\nEpoch 1/1\n4265/4265 [==============================] - 22s - loss: 14.9959 - val_loss: 19.2078\nFitting diastole model...\nTrain on 4265 samples, validate on 1066 samples\nEpoch 1/1\n4265/4265 [==============================] - 22s - loss: 25.9386 - val_loss: 36.3604\nEvaluating CRPS...\n4265/4265 [==============================] - 9s <br>\n4265/4265 [==============================] - 9s <br>\n1066/1066 [==============================] - 2s <br>\n1066/1066 [==============================] - 2s <br>\nCRPS(train) = 0.02606546647974614\nCRPS(test) = 0.03470680849898871\nSaving weights...</p>\n\n<p>and then I ran the submission.py file but I still ended up with ~0.07 on the submission. Do the above values for CRPS look correct?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 105970,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-01-27T19:04:59.843000",
      "content": "<p>I just pushed code changes to repo. It seems that fit_generator() function doesn't work as expected, which needs further investigation. So, for now I changed the code to manually augment the data (rotations + shifts) and to use fit() function (check utils.py).</p>\n\n<p>Also, please create the data again (<code>python data.py</code>), since there was a small bug too. Then try to train the model again.</p>\n\n<p>Once again I apologize for this inconvenience, the tutorial is a product of much copy-pasting (best type of coding, right) from our original project, and some mistakes happened.</p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 105786,
      "author_name": "WD",
      "author_url": "",
      "post_date": "2016-01-26T18:23:09.480000",
      "content": "<p>@Marko. I am not running Keras myself - but am intrigued by the approach. Could you help me to understand how you translate the point estimate on the volumes into a CDF? i understand that you use the following helper function (per below). I don't fully understand the function of sigma, how this is linked to hist_systole.history['loss'][-1] (and what the latter variable is, as hard to see without running the code)! </p>\n\n<p>At a high level, would love to get your views on the more fundamental advantages  / disadvantages between this approach and the approach of estimating 600 values! </p>\n\n<p>Many thanks - W</p>\n\n<blockquote>\n  <p>def real_to_cdf(y, sigma=1e-10):\n      &quot;&quot;&quot;\n      Utility function for creating CDF from real number and sigma (uncertainty measure).</p>\n\n<pre><code>:param y: array of real values\n:param sigma: uncertainty measure. The higher sigma, the more imprecise the prediction is, and vice versa.\n\nDefault value for sigma is 1e-10 to produce step function if needed.\n&quot;&quot;&quot;\ncdf = np.zeros((y.shape[0], 600))\nfor i in range(y.shape[0]):\n    cdf[i] = norm.cdf(np.linspace(0, 599, 600), y[i], sigma)\nreturn cdf\n</code></pre>\n</blockquote>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 105730,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-01-26T09:41:20.810000",
      "content": "<p>@Woolsey</p>\n\n<p>From what I can see, you are running it on CPU... You should use Theano flags when running the code, or specify .theanorc file.</p>\n\n<p>You should check these links for more detailed instructions:\n<a href=\"http://deeplearning.net/software/theano/tutorial/using_gpu.html\">http://deeplearning.net/software/theano/tutorial/using_gpu.html</a>\n<a href=\"http://deeplearning.net/software/theano/library/config.html\">http://deeplearning.net/software/theano/library/config.html</a></p>\n\n<p>Cheers,</p>\n\n<p>Marko</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 105655,
      "author_name": "fchollet",
      "author_url": "",
      "post_date": "2016-01-25T16:48:26.953000",
      "content": "<p>[quote=phunter;105652]</p>\n\n<p>Keras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) </p>\n\n<p>[/quote]</p>\n\n<p>This is a pretty ridiculous attempt at saying &quot;lol MXnet is faster than Keras&quot;. No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. </p>\n\n<p>Stay classy man...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110376,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-03-04T23:05:36.020000",
      "content": "<p>[quote=Manuele Tamburrano;109680]</p>\n\n<p>Hi Marko, thank you for sharing your code.</p>\n\n<p>I see you do a 2d convolution on the 30 time series, but I don't find any correlation among different sax of a same study. So do you predict volumes independently for each sax and let the net to guess itself the &quot;level&quot; of a sax?\nI mean, the volume should be predicted by a combination of all sax, so a 3d convolution would seems the right choice here, is there some reason why you opted for the 2d one? (computational reasons aside)</p>\n\n<p>thank you again</p>\n\n<p>[/quote]</p>\n\n<p>You are completely right - in this tutorial the network seems to pick up stuff that helps it determine the volume even if the &quot;level&quot; of sax is not directly given. 3D convolutions are definitely a good idea (which we are using as well), just make sure you are sorting slices correctly ;).</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 108848,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-02-20T15:30:42.120000",
      "content": "<p>[quote=PengPai;108816]</p>\n\n<p>@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?</p>\n\n<p>[/quote]</p>\n\n<p>Besides parameter tuning, I found that the most significant performance boosts lied in better pre-processing (i.e. segmentation, ROI extraction) and smart(er) usage of metadata provided in DICOM images. One should have much better insights of all the provided data in order to utilize them properly, instead of just feeding images to CNNs.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 108816,
      "author_name": "PengPai",
      "author_url": "",
      "post_date": "2016-02-20T06:35:04.677000",
      "content": "<p>@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 108319,
      "author_name": "Guilherme Goto Escudero",
      "author_url": "",
      "post_date": "2016-02-16T17:51:29.857000",
      "content": "<p>[quote=jwjohnson314;108271]</p>\n\n<p>So I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.</p>\n\n<p>[/quote]</p>\n\n<p>When you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 108156,
      "author_name": "rmldj",
      "author_url": "",
      "post_date": "2016-02-15T22:20:42.160000",
      "content": "<p>[quote=Marko Jocic;108154]</p>\n\n<p>RMSE objective was removed from Keras a week ago. You can either use MSE as your objective or revert Keras to 0.3.1 version in order to use RMSE. You can also write RMSE as a custom objective function if you wish to use bleeding edge version of Keras.</p>\n\n<p>[/quote]</p>\n\n<p>I guess it would help the spread of Keras if it retained more stability in its API..</p>\n\n<p>I used Keras in the EEG Kaggle competition and used the JZS3 recurrent layers which for these data were faster and better than LSTM (by a noticable factor) - after the competition I upgraded Keras and found that JZS3 was removed...</p>\n\n<p>I really enjoy using Keras - and also perfectly understand making breaking changes if they significantly improve the usability of the API, however removing things only for aesthetic reasons (?) really causes unpleasant surprises for users - and also makes them think twice about upgrading...</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 111320,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-03-13T18:04:42.247000",
      "content": "<p>[quote=Yuanfang Guan;111314]</p>\n\n<p>it is pretty clear to anyone that was original top 10 that to perform better than tencia/wash is almost mission impossible. and to reduce 20% of error on top of their initial submissions literally  means a direct labeling of test set.  </p>\n\n<p>[/quote]</p>\n\n<p>One could say the same for your submissions as well, as you got more than 100% improvement in score since the first phase. So let's drop the ball for the manual labeling part.</p>\n\n<p>What Alexander said - bear in mind that current leaderboard is on calculated on 1% of test data, and my guess is that our model &quot;got lucky&quot; on those few studies.  Anyways, I expect much different standings tomorrow night. Good luck.</p>\n\n<p>Marko</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111319,
      "author_name": "==>(AL)<==",
      "author_url": "",
      "post_date": "2016-03-13T17:49:38.333000",
      "content": "<p>@Yuanfang Guan The current leaderboard positions is highly unstable, I would not count on it even for the remote hints of the final private standing. At best it is &quot;directionally correct&quot;</p>\n\n<p>@DavidGbodiOdaibo Unfortunately due to a technical glitch we were 1 minute late to upload the model. \nThe model is loosely based on the tutorial with number of enhancements ;) It does not use any hand labeling</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106450,
      "author_name": "SecondPlan",
      "author_url": "",
      "post_date": "2016-01-31T22:02:02.863000",
      "content": "<p>Thanks, Just to report, I trained vgg on 224 x224 images and scored &lt; 0.03, but it takes 20 min per epochs and would need 30 -40 epochs before overfit . I also tried with 3d convolution, but so far do not have any good result. Since theano does not support multi-gpu, I am moving to torch </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 106076,
      "author_name": "Piotr",
      "author_url": "",
      "post_date": "2016-01-28T10:27:41.383000",
      "content": "<p>@Marko, after yours fix I was able to reproduce LB ~ 0.0359. Thanks for sharing!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 105938,
      "author_name": "Marko Jocic",
      "author_url": "",
      "post_date": "2016-01-27T14:30:28.673000",
      "content": "<p>@ tereka</p>\n\n<p>Thank you for reporting this. Since already two people experienced the same thing, I will investigate what is going on. At least, now we know there certainly is some sort of random element to the code, and I wish to offer my apologies for not seeing this earlier. It is unfortunate, since some people already said they could reproduce similar results with this model (even phunter on MXnet).</p>\n\n<p>If anyone figures out what causes this randomness and shares it with the rest of us, I guess everyone would be thankful and I would gladly update the tutorial accordingly to be more stable.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 105698,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T22:28:59.810000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 105637,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T14:12:22.717000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 105632,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T13:08:14.227000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 105893,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-27T08:56:13.660000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 111366,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-14T05:32:37.273000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111324,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-13T18:21:10.657000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111255,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-13T06:20:39.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111237,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-12T23:56:52.340000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111183,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-12T15:27:17.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111180,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-12T15:16:09.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111177,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-12T12:16:19.220000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110933,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-09T18:46:03.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110927,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-09T17:46:03.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110899,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-09T11:26:12.790000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110839,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-08T18:10:46.037000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 226512,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-10-02T14:43:35.657000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 110393,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-05T01:13:58.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110353,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-04T20:10:08.610000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 110012,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-02T02:11:29.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109956,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-01T18:11:28.397000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109680,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-29T11:41:35.107000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109600,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-28T12:16:04.993000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109573,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-27T23:18:23.453000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109505,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-26T21:27:41.637000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109502,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-26T20:57:22.987000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109471,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-26T14:14:35.097000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109466,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-26T12:35:28.223000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109400,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-25T18:43:43.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109351,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-25T11:57:13.520000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 109284,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-24T19:47:14.067000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108849,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-20T15:38:49.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108809,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-20T05:38:30.553000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108673,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-19T02:50:11.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108573,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-18T12:01:51.733000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108544,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-18T05:44:10.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108490,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-17T19:59:55.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108344,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-16T20:03:08.147000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108333,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-16T19:10:35.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108271,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-16T14:55:24.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108154,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-15T21:33:56.507000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108135,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-15T19:20:51.487000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108133,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-15T19:17:56.473000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108131,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-15T18:52:51.933000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108127,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-15T18:28:14.073000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108056,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-15T05:42:28.177000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 108055,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-15T05:32:50.787000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 107753,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-12T08:54:25.113000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 107752,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-12T08:54:13.250000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 107718,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-12T04:42:14.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106929,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-05T00:38:53.833000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106927,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-04T23:46:20.150000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106924,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-04T23:30:40.060000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106920,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-04T22:21:06.227000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106889,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-04T18:05:08.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106824,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-04T03:40:46.443000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106603,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-02T12:08:45.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106532,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-01T21:20:40.357000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106483,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-01T08:56:18.113000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106465,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-02-01T03:17:08.880000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106324,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-30T00:27:05.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106317,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-29T22:50:23.433000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106316,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-29T22:48:08.543000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106240,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-29T10:49:08.637000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106239,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-29T10:35:00.850000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106235,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-29T10:02:06.053000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106226,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-29T09:17:53.857000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106199,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-29T01:45:58.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106150,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-28T20:15:42.907000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106119,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-28T16:42:41.573000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106117,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-28T15:58:39.373000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106077,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-28T10:30:35.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106047,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-28T03:14:27.973000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 106031,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-28T00:00:08.487000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105980,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-27T19:42:39.967000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105939,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-27T14:46:45.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105810,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-26T21:07:00.347000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105808,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-26T20:50:40.137000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105793,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-26T19:14:18.280000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105738,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-26T10:49:14.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105719,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-26T04:26:20.383000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105705,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T23:48:36.403000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105697,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T22:27:11.490000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105684,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T21:01:43.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105679,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T19:49:02.583000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105675,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T19:16:46.843000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105673,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T19:13:29.217000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105635,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T13:52:49.587000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105630,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T13:01:31.857000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105610,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T05:59:59.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 105604,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-01-25T04:14:39.930000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111322,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-13T18:13:48.620000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 111314,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-13T16:13:36.947000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "105563": "Hi everyone,\r\n\r\na Keras-based deep learning tutorial for ~0.0359 CRPS is available on:\r\n\r\nhttps://github.com/jocicmarko/kaggle-dsb2-keras/\r\n\r\nBasically it is a conv net for linear regression task.\r\nThere wasn't much experimenting with net structure, hyper-parameters and image pre-processing, so these might be a good place to start.\r\n\r\nCheers,\r\n\r\nMarko",
    "105662": "@phunter\r\n\r\nFrom the top of my head, I think the main difference is that in our example the two models are trained basically at the same time (which doubles the memory usage), while MXnet example trains them one by one.",
    "105608": "@ertuka: yes, you should update your version of Keras, from the Keras repository:\r\n\r\n    pip install git+git://github.com/fchollet/keras.git --upgrade --no-deps\r\n\r\n(may need to be prefixed with \"sudo\")",
    "105658": "[quote=fchollet;105655]\r\n\r\nThis is a pretty ridiculous attempt at saying \"lol MXnet is faster than Keras\". No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. \r\n\r\nStay classy man...\r\n\r\n[/quote]\r\n\r\nOh, no, I didn't mean MXnet was faster because of what what, why downvote me? @Marko Jocic said 128x128 was too large to fit in his GTX 770 video card, and I said I could run 128x128 with mxnet on GTX 960, an even slower video card, so @WD didn't have to worry about memory of running it. We ran on different cards different resolutions different CPUs for data augment etc, and there was no comparisons. \r\n\r\nStay calm :-)",
    "105919": "Hi!\r\n\r\nThank you sharing your model!\r\n\r\nWhen I submit this model code, get the high errors(~0.077926).(same piotr happening)\r\nThis code run 150 epochs running.\r\n\r\nSo I think this code is influented random state or other cause.",
    "105652": "[quote=WD;105630]\r\n\r\n@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.\r\n\r\n* Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?\r\n\r\n* You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? \r\n\r\n* Does the code while len(images) < 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? \r\n\r\nmany thanks, W\r\n \r\n\r\n[/quote]\r\n\r\nKeras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) \r\n\r\nUpdate: Seems like this post has caused some confusions. Please, the statement above didn't mean any comparisons between MXnet and Keras. It was a simple run of MXnet with this tutorial's network from my own hardware. @Marko and I had different hardware, different image resolution etc, there was no apple-to-apple comparison of these two toolkits. Keras and Mxnet are both good deep learning frameworks, and it is the same with caffe/theano/torch/Lasagne etc which may later on have good tutorials for this Kaggle competition too, so please select your favorite one or ones for winning 200k $. And sorry to @Marko and @WD for this topic divergence. ",
    "105889": "Hi!\r\n\r\nThanks for this tutorial! I cloned github code and run\r\n\r\n    python data.py\r\n    python train.py\r\n    python submission.py\r\n\r\nand this gave me \t0.074981 on LB, I didn't make any changes in the code. Does anyone have idea what can goes wrong? I expect LB ~ 0.0359",
    "111308": "you guys are now 1st just wondering what you did? I know you are not using the model in this tutorial. It also looks like you did not upload your model so no prize money :)",
    "111178": "I tried with vgg model on 192 x192 images and tweaking some parameters,  got 0.0101. ",
    "110930": "[quote=datapool;110927]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nFor the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.\r\n",
    "110921": "[quote=datapool;110899]\r\n\r\nHello Marko Jocic,\r\n\r\nThanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?\r\n\r\n[/quote]\r\n\r\nHi, yes we also had that approach, and it yielded somewhat smaller accuracy, but on the other side double faster training, so it's definitely OK for more model diversity.\r\n",
    "110377": "[quote=JasonMay;109956]\r\n\r\nMy question is do you know where all the shape must be addressed to alter image size to 128X128?\r\n\r\nThanks,\r\nJason\r\n\r\n[/quote]\r\n\r\nYou would need a bigger/beter model for 128x128 images.\r\n\r\n[quote=Gongning Luo;110012]\r\n\r\nHow big of memory(RAM), the parameter of GPU and CPU?\r\n\r\n[/quote]\r\n\r\nYou need ~1GB of GPU RAM for this tutorial.\r\n\r\n\r\n[quote=Lawrence Chernin;110353]\r\n\r\nI created a public AWS image: ami-d42a59b4 that has all the needed libraries to run this code.\r\nRunning on g2.2xlarge. (sorry, you may also need  to install cython and h5py )\r\n\r\n[/quote]\r\n\r\nWould you care sharing the image with the rest of Kagglers?\r\n\r\nThanks! Marko",
    "106509": "[quote=Ertuka;106465]\r\n\r\n1) Could it be weights_systole_best better to use than weights_systole weights?\r\n2) Could I be missing anything else  to preload the previous model /run in train.py?\r\n3) Is there any difference using get_model versus to_json/model_from_json? Any idea why the later is not working in this model?\r\n\r\n[/quote]\r\n\r\n1) You could try both. Both should eventually converge more or less the same.\r\n\r\n2) 3) That should be it. It might be a bug currently in Keras, maybe the custom activation function (in first layer) can't be saved/loaded with to_json/from_json. You can check that. However, get_model() should work the same.\r\n\r\n[quote=WDharvard;106483]\r\n\r\nwhat are reasonable root mean squared error (RMSE) to observe during training?\r\n\r\n[/quote]\r\n\r\nYou should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (>0.04).",
    "106225": "@Marko, @patruf @ Wei Wu,\r\n\r\nI was able to get 0.359 with Marko code, but I made changes in submission.py or in val_los.txt files - I don't remember precisely now. But the network trained with train.py is OK. There is only a problem with submission.py.",
    "106213": "@Marko , @patruff\r\n\r\nI can confirm the same problem as @patruff.\r\nWhen I run the model training, at the end the test CRPS was about 0.029.\r\nAfter I run submission.py and submit the file, the LB score was about 0.08. (It was a fresh run there was no old files.)\r\n\r\nI am running python 2.7 under a windows environment.",
    "106106": "Iteration 200/200\r\n--------------------------------------------------\r\nAugmenting images - rotations\r\n4265/4265 [==============================] - 33s       \r\nAugmenting images - shifts\r\n4265/4265 [==============================] - 8s        \r\nFitting systole model...\r\nTrain on 4265 samples, validate on 1066 samples\r\nEpoch 1/1\r\n4265/4265 [==============================] - 22s - loss: 14.9959 - val_loss: 19.2078\r\nFitting diastole model...\r\nTrain on 4265 samples, validate on 1066 samples\r\nEpoch 1/1\r\n4265/4265 [==============================] - 22s - loss: 25.9386 - val_loss: 36.3604\r\nEvaluating CRPS...\r\n4265/4265 [==============================] - 9s     \r\n4265/4265 [==============================] - 9s     \r\n1066/1066 [==============================] - 2s     \r\n1066/1066 [==============================] - 2s     \r\nCRPS(train) = 0.02606546647974614\r\nCRPS(test) = 0.03470680849898871\r\nSaving weights...\r\n\r\nand then I ran the submission.py file but I still ended up with ~0.07 on the submission. Do the above values for CRPS look correct?",
    "105970": "I just pushed code changes to repo. It seems that fit_generator() function doesn't work as expected, which needs further investigation. So, for now I changed the code to manually augment the data (rotations + shifts) and to use fit() function (check utils.py).\r\n\r\nAlso, please create the data again (```python data.py```), since there was a small bug too. Then try to train the model again.\r\n\r\nOnce again I apologize for this inconvenience, the tutorial is a product of much copy-pasting (best type of coding, right) from our original project, and some mistakes happened.\r\n\r\nCheers,\r\n\r\nMarko",
    "105786": "@Marko. I am not running Keras myself - but am intrigued by the approach. Could you help me to understand how you translate the point estimate on the volumes into a CDF? i understand that you use the following helper function (per below). I don't fully understand the function of sigma, how this is linked to hist_systole.history['loss'][-1] (and what the latter variable is, as hard to see without running the code)! \r\n\r\nAt a high level, would love to get your views on the more fundamental advantages  / disadvantages between this approach and the approach of estimating 600 values! \r\n\r\nMany thanks - W\r\n\r\n> def real_to_cdf(y, sigma=1e-10):\r\n>     \"\"\"\r\n>     Utility function for creating CDF from real number and sigma (uncertainty measure).\r\n> \r\n>     :param y: array of real values\r\n>     :param sigma: uncertainty measure. The higher sigma, the more imprecise the prediction is, and vice versa.\r\n> \r\n>     Default value for sigma is 1e-10 to produce step function if needed.\r\n>     \"\"\"\r\n>     cdf = np.zeros((y.shape[0], 600))\r\n>     for i in range(y.shape[0]):\r\n>         cdf[i] = norm.cdf(np.linspace(0, 599, 600), y[i], sigma)\r\n>     return cdf\r\n\r\n",
    "105730": "@Woolsey\r\n\r\nFrom what I can see, you are running it on CPU... You should use Theano flags when running the code, or specify .theanorc file.\r\n\r\nYou should check these links for more detailed instructions:\r\nhttp://deeplearning.net/software/theano/tutorial/using_gpu.html\r\nhttp://deeplearning.net/software/theano/library/config.html\r\n\r\nCheers,\r\n\r\nMarko",
    "105655": "[quote=phunter;105652]\r\n\r\nKeras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) \r\n\r\n[/quote]\r\n\r\nThis is a pretty ridiculous attempt at saying \"lol MXnet is faster than Keras\". No it's not, you are comparing two different models. The performance difference is largely due to the data augmentation process. Keras runs cuDNN under the hood, making it as fast as everything else on the market. \r\n\r\nStay classy man...\r\n\r\n\r\n",
    "110376": "[quote=Manuele Tamburrano;109680]\r\n\r\nHi Marko, thank you for sharing your code.\r\n\r\nI see you do a 2d convolution on the 30 time series, but I don't find any correlation among different sax of a same study. So do you predict volumes independently for each sax and let the net to guess itself the \"level\" of a sax?\r\nI mean, the volume should be predicted by a combination of all sax, so a 3d convolution would seems the right choice here, is there some reason why you opted for the 2d one? (computational reasons aside)\r\n\r\nthank you again\r\n\r\n[/quote]\r\n\r\nYou are completely right - in this tutorial the network seems to pick up stuff that helps it determine the volume even if the \"level\" of sax is not directly given. 3D convolutions are definitely a good idea (which we are using as well), just make sure you are sorting slices correctly ;).\r\n",
    "108848": "[quote=PengPai;108816]\r\n\r\n@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?\r\n\r\n[/quote]\r\n\r\nBesides parameter tuning, I found that the most significant performance boosts lied in better pre-processing (i.e. segmentation, ROI extraction) and smart(er) usage of metadata provided in DICOM images. One should have much better insights of all the provided data in order to utilize them properly, instead of just feeding images to CNNs.",
    "108816": "@Marko, thank you for sharing such a good tutorial for keras fans. Could you please share your insights or suggestions if we would like to boost our performance besides tuning parameters ?",
    "108319": "[quote=jwjohnson314;108271]\r\n\r\nSo I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.\r\n\r\n[/quote]\r\n\r\nWhen you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.",
    "108156": "[quote=Marko Jocic;108154]\r\n\r\nRMSE objective was removed from Keras a week ago. You can either use MSE as your objective or revert Keras to 0.3.1 version in order to use RMSE. You can also write RMSE as a custom objective function if you wish to use bleeding edge version of Keras.\r\n\r\n[/quote]\r\n\r\nI guess it would help the spread of Keras if it retained more stability in its API..\r\n\r\nI used Keras in the EEG Kaggle competition and used the JZS3 recurrent layers which for these data were faster and better than LSTM (by a noticable factor) - after the competition I upgraded Keras and found that JZS3 was removed...\r\n\r\nI really enjoy using Keras - and also perfectly understand making breaking changes if they significantly improve the usability of the API, however removing things only for aesthetic reasons (?) really causes unpleasant surprises for users - and also makes them think twice about upgrading...\r\n\r\n ",
    "111320": "[quote=Yuanfang Guan;111314]\r\n\r\n it is pretty clear to anyone that was original top 10 that to perform better than tencia/wash is almost mission impossible. and to reduce 20% of error on top of their initial submissions literally  means a direct labeling of test set.  \r\n\r\n[/quote]\r\n\r\nOne could say the same for your submissions as well, as you got more than 100% improvement in score since the first phase. So let's drop the ball for the manual labeling part.\r\n\r\nWhat Alexander said - bear in mind that current leaderboard is on calculated on 1% of test data, and my guess is that our model \"got lucky\" on those few studies.  Anyways, I expect much different standings tomorrow night. Good luck.\r\n\r\nMarko",
    "111319": "@Yuanfang Guan The current leaderboard positions is highly unstable, I would not count on it even for the remote hints of the final private standing. At best it is \"directionally correct\"\r\n\r\n@DavidGbodiOdaibo Unfortunately due to a technical glitch we were 1 minute late to upload the model. \r\nThe model is loosely based on the tutorial with number of enhancements ;) It does not use any hand labeling\r\n\r\n",
    "106450": "Thanks, Just to report, I trained vgg on 224 x224 images and scored < 0.03, but it takes 20 min per epochs and would need 30 -40 epochs before overfit . I also tried with 3d convolution, but so far do not have any good result. Since theano does not support multi-gpu, I am moving to torch ",
    "106076": "@Marko, after yours fix I was able to reproduce LB ~ 0.0359. Thanks for sharing!",
    "105938": "@ tereka\r\n\r\nThank you for reporting this. Since already two people experienced the same thing, I will investigate what is going on. At least, now we know there certainly is some sort of random element to the code, and I wish to offer my apologies for not seeing this earlier. It is unfortunate, since some people already said they could reproduce similar results with this model (even phunter on MXnet).\r\n\r\nIf anyone figures out what causes this randomness and shares it with the rest of us, I guess everyone would be thankful and I would gladly update the tutorial accordingly to be more stable.",
    "105698": "@Woolsey\r\n\r\nPlease upgrade Theano to latest version:\r\n\r\n\r\n    pip install --upgrade --no-deps git+git://github.com/Theano/Theano.git\r\n\r\n",
    "105637": "@WD\r\n\r\n - I didn't use difference between frames intentionally (as in MXNet tutorial) - just to show that other approaches (in this case - TV denoising) might work fine as well. Definitely a good place for trying different stuff.\r\n - Indeed, that might be true - good pre-processing could really make a difference. Well, I didn't experiment much with hyper-parameters and network itself, but I noticed, for example, that using (5,5) convolutions in first layers yielded roughly same results as using (3,3). More or less it was the same with number of convolutional filters... In this example, the biggest improvement in score was switching from classification (600 sigmoid outputs) to linear regression (only 1 output). However, I do recommend trying out different parameters, because today I changed Adam learning rate a bit, and got slightly better results. Also, this example could maybe use a bit more regularization, since it slightly overfits. Ideally, one should do cross-validation for determining these hyper-parameters, but on my hardware that would be very time consuming.",
    "105632": "@WD\r\n\r\n\r\n - At first I was training on 128*128 images, but my poor hardware couldn't handle it. But yeah, I would expect improvement on the score with larger images.\r\n - Maybe I could've said that better :). What I meant is - ~100 secs systole model + ~100 secs diastole model + ~60 secs for CRPS evaluation = ~260 secs, which is something below 5 minutes for one iteration. And then you *only* need 149 more iterations like this.\r\n - Yes, in order to have same input shape for all samples.\r\n\r\nCheers,\r\n\r\nMarko",
    "105893": "@piotr\r\n\r\nDid you let it run for all 150 iterations? 0.074981 is pretty high error, and model should get better then that in first couple of iterations...\r\n\r\nAlso, our current score (~0.030) uses this same model, but with a bit tweaked parameters, and our previous score (~0.0359) was from the model in the tutorial.",
    "111366": "@Marko would love to see your final architecture regardless of how well your team ends up doing!",
    "111324": "[quote=Yuanfang Guan;111322]\r\n\r\ni didn't point to you why are you so defensive?\r\n\r\ni meant a 20% error reduction over wash's method in the complete final test set. \r\neveryone improved almost 100% over their own initial submission due to the selection of three easy cases.\r\n\r\nat least we uploaded our models which are binary reproducible. \r\n\r\ni hope the final winning model will be open to public scrutiny.  \r\n\r\ni don't really count on getting anything from this competition anyway, so i can assure you i am now dropping the topic completely now.\r\n\r\nhave a nice day.\r\n\r\n[/quote]\r\n\r\nPlease accept my sincere apologies, I totally misunderstood you there.\r\n\r\nIt's also too bad that we didn't manage to upload our model on time (literally 30 seconds late), but the competition was really fun. \r\n\r\nI do hope as well that the winning model will be open for public - it could lead to much improvement and thus better usage in clinical purposes.",
    "111255": "[quote=Marko Jocic;111237]\r\n\r\n[quote=datapool;111183]\r\n\r\nAfter disabling all my normalization layers its didn't improve training speed.\r\n\r\n[/quote]\r\n\r\nDo you have cuDNN installed?\r\n\r\n[/quote]\r\nYes, installed this cudnn-7.0-linux-x64-v3.0-prod.tgz.",
    "111237": "[quote=datapool;111183]\r\n\r\nAfter disabling all my normalization layers its didn't improve training speed.\r\n\r\n[/quote]\r\n\r\nDo you have cuDNN installed?",
    "111183": "[quote=Marko Jocic;110930]\r\n\r\n[quote=datapool;110927]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nFor the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.\r\n\r\n\r\n[/quote]\r\nAfter disabling all my normalization layers its didn't improve training speed.",
    "111180": "[quote=SecondPlan;111178]\r\n\r\nI tried with vgg model on 192 x192 images and tweaking some parameters,  got 0.0101. \r\n\r\n[/quote]\r\nI did 3D model similar to Slow Fusion with input of 24x40x40. Couldn't fit anything larger on 4GB of K520. I crop an ROI using very basic image processing.",
    "111177": "Does anyone know whats the score of this method trained on train+validate and scored on the new test data? Just wanted to see if i was able to improve by using this as benchmark. I am at 0.014539. ",
    "110933": "[quote=Marko Jocic;110930]\r\n\r\n[quote=datapool;110927]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nFor the exact reason you just mentioned :). Although it made the network to converge faster in the first few iterations, it made the training (per iteration) much slower. So I decided to leave it out for practical reasons.\r\n\r\n\r\n[/quote]\r\nOops, i didn't realize normalization layer can make it so slow. I tough depth of my layer is the reason. Anyway good to learn by mistake. I will experiment after competition is over.\r\n\r\nThanks",
    "110927": "[quote=Marko Jocic;110921]\r\n\r\n[quote=datapool;110899]\r\n\r\nHello Marko Jocic,\r\n\r\nThanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?\r\n\r\n[/quote]\r\n\r\nHi, yes we also had that approach, and it yielded somewhat smaller accuracy, but on the other side double faster training, so it's definitely OK for more model diversity.\r\n\r\n\r\n[/quote]\r\n\r\nThanks for the feedback.\r\nI have noticed that most CNN implementation code on web based on Theano (keras/lasagne) doesn't have a local response normalization layer. This is the layer discussed in section 3.3 in famous paper \"ImageNet Classification with Deep Convolutional Neural Networks\".\r\nThis can be added by model.add(BatchNormalization(epsilon=1e-06, mode=0, axis=1, momentum=0.9)) in Keras. I added a couple of these layers to be as close to what i read in papers as possible.\r\nMy model takes 21 minutes per iteration so i didn't have time to experiment and see if its useful to have this layer or not. Any thoughts on why you didn't use this layer.\r\n\r\nThanks",
    "110899": "Hello Marko Jocic,\r\n\r\nThanks for sharing your code. I used it as a starting point and developed on top of it. Now that compitition code is frozen, i was wondering if it would have been beter to train one model for systole and diastole? So the final layer would be Dense(2). Would that be less/more accurate than having two seperate models?\r\n\r\n[quote=Marko Jocic;109600]\r\n\r\n[quote=Alexander Popov;109573]\r\n\r\nThank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.\r\n\r\n\r\n[/quote]\r\n\r\nThank you for sharing that. Good luck in your further work!\r\n\r\nCheers, Marko\r\n\r\n\r\n[/quote]\r\n",
    "110839": "Thanks a lot for this tutorial. \r\ndata.py is working fine. I am getting the following error while running train.py.  Any idea what's the problem?\r\n\r\nThanks\r\n\r\n------------------------------------------------------------------\r\n\r\n\r\n(myproject) khushhall@tunga:~/myproject/kaggle-dsb2-keras$ python train.py \r\n\r\nUsing Theano backend.\r\n\r\n/home/khushhall/myproject/local/lib/python2.7/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n  warnings.warn(\"downsample module has been moved to the pool module.\")\r\n\r\nLoading and compiling models...\r\n\r\nLoading training data...\r\n\r\nPre-processing images...\r\n\r\nTraceback (most recent call last):\r\n\r\n  File \"train.py\", line 150, in <module>\r\n    train()\r\n\r\n  File \"train.py\", line 65, in train\r\n\r\n    X_train, y_train, X_test, y_test = split_data(X, y, split_ratio=0.2)\r\n\r\n  File \"train.py\", line 42, in split_data\r\n\r\n    X_test = X[:split, :, :, :]\r\n\r\nIndexError: too many indices for array\r\n",
    "110393": "It is a publicly available image. More details of how to find it in AWS are attached. The smallest gpu on AWS is the g2.2xlarge which has 15GB ram. I see the process is using round 3GB ram (RES).\r\n",
    "110353": "I created a public AWS image: ami-d42a59b4 that has all the needed libraries to run this code.\r\nRunning on g2.2xlarge. (sorry, you may also need  to install cython and h5py )",
    "110012": "@Marko Jocic\r\nWhen I use the source code, some errors occurs as following:\r\n      \r\n*File \"train.py\", line 18, in load_train_data\r\n  X=X.astype<np.float32>\r\n     MemoryError*\r\n\r\nI am a beginner in the area. I want to know the requirement of hardware to achive running the code.\r\nHow big of memory(RAM), the parameter of GPU and CPU\r\nI also want to know How can I merge with you and to follow you in your team?\r\n\r\nThank you very much!\r\n",
    "109956": "[quote=Marko Jocic;105632]\r\n\r\n@WD\r\n\r\n\r\n - At first I was training on 128*128 images, but my poor hardware couldn't handle it. But yeah, I would expect improvement on the score with larger images.\r\n\r\nCheers,\r\n\r\nMarko\r\n\r\n[/quote]\r\n\r\nHi Marko,\r\n\r\nThanks for the tutorial.  It has saved my group, which is taking on the DSB as our Graduate Capstone project.  We were able to achieve ~ .033 CRPS at about 3 minutes per iteration.\r\n\r\nI am interested in trying different image sizes (96X96 or 128X128), but think I am missing step.\r\n\r\nI changed img_shape = (64,64) in the data.py file, but when I ran train.py, I got a significantly worse CRPS in the first few iterations.\r\n\r\nI tried also altering model.py to input_shape=(30,128,128) but received an error message.\r\n\r\nMy question is do you know where all the shape must be addressed to alter image size to 128X128?\r\n\r\nThanks,\r\nJason\r\n",
    "109680": "Hi Marko, thank you for sharing your code.\r\n\r\nI see you do a 2d convolution on the 30 time series, but I don't find any correlation among different sax of a same study. So do you predict volumes independently for each sax and let the net to guess itself the \"level\" of a sax?\r\nI mean, the volume should be predicted by a combination of all sax, so a 3d convolution would seems the right choice here, is there some reason why you opted for the 2d one? (computational reasons aside)\r\n\r\nthank you again",
    "109600": "[quote=Alexander Popov;109573]\r\n\r\nThank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.\r\n\r\n\r\n[/quote]\r\n\r\nThank you for sharing that. Good luck in your further work!\r\n\r\nCheers, Marko\r\n",
    "109573": "Thank you for sharing very useful tutorial,  Marko. Just tried  200 iterations and it produced LB score 0.031665 which was  very close to the best CPRS on validation data set.  Each iteration took about 150 sec in my setup.\r\n",
    "109505": "[quote=Andrew Beam;109502]\r\n\r\nIt looks like the space represented by a pixel can vary quite a bit from image to image. Have you found normalizing the image spacing to be crucial to success?\r\n\r\n[/quote]\r\n\r\nIt depends on the type of analysis you do. However, in our case we found it definitely helps",
    "109502": "Hi Marko, great tutorial. I was wondering if you would be willing to comment on how useful you have found the pixel spacing and slice location to be? It looks like the space represented by a pixel can vary quite a bit from image to image. Have you found normalizing the image spacing to be crucial to success?",
    "109471": "[quote=Marko Jocic;109466]\r\n\r\n[quote=DirkWillemWonnink;109400]\r\n\r\nHi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   \r\n\r\n[/quote]\r\n\r\nAs far as I can remember, ```imresize``` (from ```scipy.misc```) scales images to [0,255], and that function is used in ```crop_resize()``` in the tutorial.\r\nIf you manage to figure out what causes the black images, please let me know and/or solve it through a PR on the tutorial.\r\n\r\nThank you,\r\n\r\nMarko\r\n\r\n\r\n[/quote]\r\n\r\nIndeed I read now that imsize is converting to uint8 and automatically scaling it to 0 .. 255. For my test I switched to opencv for the rescaling as part of some additional functions. So this 'hidden conversion' was not executed anymore. It seems you can save some conversions in your tutorial model. :-) (ranging to 0 ..1 and convert to uint8) . Thanks for this clarification!\r\n\r\nGr. Dirk Willem",
    "109466": "[quote=DirkWillemWonnink;109400]\r\n\r\nHi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   \r\n\r\n[/quote]\r\n\r\nAs far as I can remember, ```imresize``` (from ```scipy.misc```) scales images to [0,255], and that function is used in ```crop_resize()``` in the tutorial.\r\nIf you manage to figure out what causes the black images, please let me know and/or solve it through a PR on the tutorial.\r\n\r\nThank you,\r\n\r\nMarko\r\n",
    "109400": "Hi Marko, while testing some adjustments I ran into almost totally black images. Not sure what changed it but the issue is related to the uint8 conversion in the end of data.py. Earlier the images are changed to float-type between 0 and 1. How is this conversion meant to work? (I would expect it needs a (float) number between 0 and 255?) . Thanks!   ",
    "109351": "[quote=Alex Risman;109284]\r\n\r\nHey Marko, this is awesome, thanks so much! Quick question: what is the purpose of the zero padding layers in the model?\r\n\r\n[/quote]\r\n\r\nConvolutional layers with border model \"valid\" and filter size greater than 1x1 reduce the size of feature maps, so I just zero pad these feature maps with one row/column on all sides to maintain the size of feature maps.\r\n",
    "109284": "Hey Marko, this is awesome, thanks so much! Quick question: what is the purpose of the zero padding layers in the model?",
    "108849": "[quote=John Gunawan;108673]\r\n\r\nHi guys, \r\n\r\nI'm pretty new to all of this, so I apologize if my question is very easy or just bad but I got stuck running data.py.  I'm getting the following error: \r\n\r\nValueError: need more than 1 value to unpack\r\n\r\nCan anyone advise on what I should do to troubleshoot this? \r\n\r\nAny help is appreciated!  Thanks! \r\n\r\n[/quote]\r\n@John: probably your train.csv is not the original anymore. Had the same issue.\r\n",
    "108809": "[quote=EIGSI;106532]\r\n\r\nI can't understand why the performance of the diastole network is consistently worse than that of systole? It is the same input data, same model why the difference?\r\nI noticed this on mxnet example too, which is using the frame differences. In terms of CRPS there is always a .01 difference and in terms of RMSE here the difference seems to be around 10.\r\n\r\n[quote=Marko Jocic;106509]\r\n\r\n \r\n\r\nYou should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (>0.04).\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\nMaybe this is because diastole volume has more variations than systole volume (std dev: 59 vs 43).",
    "108673": "Hi guys, \r\n\r\nI'm pretty new to all of this, so I apologize if my question is very easy or just bad but I got stuck running data.py.  I'm getting the following error: \r\n\r\nValueError: need more than 1 value to unpack\r\n\r\nCan anyone advise on what I should do to troubleshoot this? \r\n\r\nAny help is appreciated!  Thanks! ",
    "108573": "[quote=Maksim Korolev;108544]\r\n\r\nI also recorded error at every step, here are the graphs if anybody is interested. Seems like it overfit but easier to tell on the CRPS graph. Interested to hear others' thoughts.\r\n\r\n[/quote]\r\n\r\nYes it is obvious it overfits, but rotations and shifts should reduce overfitting.\r\nAlso, maybe try adding more regularization to the model.\r\n",
    "108544": "Thank you Marko! I used your code, but I removed the denoising and rotation and was still able to get around the ~0.035 in 100 iterations (older hardware). I also recorded error at every step, here are the graphs if anybody is interested. Seems like it overfit but easier to tell on the CRPS graph. Interested to hear others' thoughts.",
    "108490": "I just updated the code to use custom RMSE loss function.\r\n\r\nCheers,\r\n\r\nMarko",
    "108344": "[quote=Guilherme Goto Escudero;108319]\r\n\r\n[quote=jwjohnson314;108271]\r\n\r\nSo I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.\r\n\r\n[/quote]\r\n\r\nWhen you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.\r\n\r\n[/quote]\r\n\r\nThanks Guilherme - I missed that!\r\n",
    "108333": "[quote=Guilherme Goto Escudero;108319]\r\n\r\nWhen you construct the cdf from the value you receive from your network, you use the error from the net as sigma. Since you are not using rmse anymore, sigma is exploding. You should change that to a fixed value, or take the square root of the error.\r\n\r\n[/quote]\r\n\r\nExactly!",
    "108271": "So I cloned from github yesterday, changed rmse to mse, and ran with no errors. CRPS (train) lands at 0.135934288308, and CRPS (test) at 0.188673192829. The test CRPS is about what ends up on the LB. Not sure why the numbers are so bad. My installs (Keras and Theano) are up-to-date.",
    "108154": "[quote=kaldes;108131]\r\n\r\nException: Invalid objective: rmse\r\n\r\nTried upgrading Theano, no luck. \r\n\r\nAny suggestions?\r\n\r\n[/quote]\r\n\r\nRMSE objective was removed from Keras a week ago. You can either use MSE as your objective or revert Keras to 0.3.1 version in order to use RMSE. You can also write RMSE as a custom objective function if you wish to use bleeding edge version of Keras.\r\n\r\nHope that helps,\r\n\r\nMarko\r\n",
    "108135": "\r\nIts working now. Issue was with the  version of MSVS\r\n\r\n[quote=Woolsey;108127]\r\n\r\n\r\n\r\nMohd, \r\n\r\nSorry I can't really answer your question as I gave up on using convnet for now (no nvidia).\r\n\r\nHowever if the improvement is below x10, I'd guess you have a software problem, either within the code or because of an installation issue.\r\n\r\nGood luck...\r\n\r\n[/quote]\r\n",
    "108133": "use loss='mse'\r\n[quote=kaldes;108131]\r\n\r\nJust joined the contest and very excited to see the collaborative energy in the forum! Thank you guys! Eager to try DL with Keras on this problem. \r\n\r\nGot through prepping the data files without error. \r\n\r\nTrying to train, I am getting the following error (Mac Book Pro, OS X El Capitano)\r\n======\r\nPython 3.5.1 |Anaconda 2.4.1 (x86_64)| (default, Dec  7 2015, 11:24:55) \r\n[GCC 4.2.1 (Apple Inc. build 5577)] on darwin\r\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\r\n>>> runfile('/Users/aards/Desktop/temp/train.py', wdir='/Users/aards/Desktop/temp')\r\nUsing Theano backend.\r\n/Users/aards/anaconda/lib/python3.5/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n  warnings.warn(\"downsample module has been moved to the pool module.\")\r\nLoading and compiling models...\r\nTraceback (most recent call last):\r\n  File \"<stdin>\", line 1, in <module>\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 699, in runfile\r\n    execfile(filename, namespace)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 88, in execfile\r\n    exec(compile(open(filename, 'rb').read(), filename, 'exec'), namespace)\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 147, in <module>\r\n    train()\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 52, in train\r\n    model_systole = get_model()\r\n  File \"/Users/aards/Desktop/temp/model.py\", line 52, in get_model\r\n    model.compile(optimizer=adam, loss='rmse')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/models.py\", line 460, in compile\r\n    self.loss = objectives.get(loss)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/objectives.py\", line 62, in get\r\n    return get_from_module(identifier, globals(), 'objective')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/utils/generic_utils.py\", line 14, in get_from_module\r\n    str(identifier))\r\nException: Invalid objective: rmse\r\n=========\r\n\r\nTried upgrading Theano, no luck. \r\n\r\nAny suggestions?\r\n\r\n\r\n[/quote]\r\n",
    "108131": "Just joined the contest and very excited to see the collaborative energy in the forum! Thank you guys! Eager to try DL with Keras on this problem. \r\n\r\nGot through prepping the data files without error. \r\n\r\nTrying to train, I am getting the following error (Mac Book Pro, OS X El Capitano)\r\n======\r\nPython 3.5.1 |Anaconda 2.4.1 (x86_64)| (default, Dec  7 2015, 11:24:55) \r\n[GCC 4.2.1 (Apple Inc. build 5577)] on darwin\r\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\r\n>>> runfile('/Users/aards/Desktop/temp/train.py', wdir='/Users/aards/Desktop/temp')\r\nUsing Theano backend.\r\n/Users/aards/anaconda/lib/python3.5/site-packages/theano/tensor/signal/downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n  warnings.warn(\"downsample module has been moved to the pool module.\")\r\nLoading and compiling models...\r\nTraceback (most recent call last):\r\n  File \"<stdin>\", line 1, in <module>\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 699, in runfile\r\n    execfile(filename, namespace)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/spyderlib/widgets/externalshell/sitecustomize.py\", line 88, in execfile\r\n    exec(compile(open(filename, 'rb').read(), filename, 'exec'), namespace)\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 147, in <module>\r\n    train()\r\n  File \"/Users/aards/Desktop/temp/train.py\", line 52, in train\r\n    model_systole = get_model()\r\n  File \"/Users/aards/Desktop/temp/model.py\", line 52, in get_model\r\n    model.compile(optimizer=adam, loss='rmse')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/models.py\", line 460, in compile\r\n    self.loss = objectives.get(loss)\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/objectives.py\", line 62, in get\r\n    return get_from_module(identifier, globals(), 'objective')\r\n  File \"/Users/aards/anaconda/lib/python3.5/site-packages/keras/utils/generic_utils.py\", line 14, in get_from_module\r\n    str(identifier))\r\nException: Invalid objective: rmse\r\n=========\r\n\r\nTried upgrading Theano, no luck. \r\n\r\nAny suggestions?\r\n",
    "108127": "Mohd, \r\n\r\nSorry I can't really answer your question as I gave up on using convnet for now (no nvidia).\r\n\r\nHowever if the improvement is below x10, I'd guess you have a software problem, either within the code or because of an installation issue.\r\n\r\nGood luck...",
    "108056": "Hi Woolsey,\r\n\r\nI am currently facing the same issue (iteration 1 took around 10 hrs). I have applied the fix suggested by MJ, am using \"Nvidia Grid K520 GPU\". How much time does iteration 1 take after applying the fix suggested by MJ. \r\nWhat was your total training time?\r\n\r\n[quote=Woolsey;105719]\r\n\r\nOne more question please. Iteration 1 took 6-7 hours on my system (XPS8100, no decent GPU). Does that sound right given this poor hardware, or I should suspect some problem with the software? \r\n\r\n\r\n \r\n\r\n       runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\r\n        Using Theano backend.\r\n        Loading and compiling models...\r\n        Loading training data...\r\n        Pre-processing images...\r\n        5331/5331 [==============================] - 934s   C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\theano\\tensor\\signal\\downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n          warnings.warn(\"downsample module has been moved to the pool module.\")\r\n        C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py:39: DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\r\n          X_test = X[:split, :, :, :]\r\n\r\n    \r\n    --------------------------------------------------\r\n    Training...\r\n    --------------------------------------------------\r\n    --------------------------------------------------\r\n    Iteration 1/150\r\n    --------------------------------------------------\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5542s - loss: 34.4657 - val_loss: 42.3799\r\n    Fitting diastole model...\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5523s - loss: 55.8717 - val_loss: 86.3558\r\n    Evaluating CRPS...\r\n    4265/4265 [==============================] - 1766s     \r\n    4265/4265 [==============================] - 1763s     \r\n    1066/1066 [==============================] - 441s     \r\n    1066/1066 [==============================] - 439s     \r\n    CRPS(train) = 0.0785559420969\r\n    CRPS(test) = 0.0743980740534\r\n    Saving weights...\r\n    --------------------------------------------------\r\n    Iteration 2/150\r\n    --------------------------------------------------\r\n\r\n\r\n[/quote]\r\n",
    "108055": "",
    "107753": "check how the y_train.npy file is created and what's inside ",
    "107752": "[quote=Danny Malter;107718]\r\n\r\nCan somebody please explain the difference between y_train[:0] and y_train[:,1].  I realize it is the first and second column of y_train, but why is the first column considered the systole data and second column considered diastole.\r\n\r\n[/quote]\r\n\r\nIf you check ```train.csv``` file provided by Kaggle, they put like that there.\r\n",
    "107718": "Can somebody please explain the difference between y_train[:0] and y_train[:,1].  I realize it is the first and second column of y_train, but why is the first column considered the systole data and second column considered diastole.\r\n\r\n            print('Fitting systole model...')\r\n            hist_systole = model_systole.fit(X_train_aug, y_train[:, 0], \r\n                shuffle=True, nb_epoch=epochs_per_iter,\r\n            batch_size=batch_size, validation_data=(X_test, y_test[:, 0]))\r\n\r\n            print('Fitting diastole model...')\r\n            hist_diastole = model_diastole.fit(X_train_aug, y_train[:, 1], \r\n                shuffle=True, nb_epoch=epochs_per_iter,\r\n            batch_size=batch_size, validation_data=(X_test, y_test[:, 1]))",
    "106929": "[quote=hassiktir;106920]\r\n\r\nHi, I had a general keras question: what is the difference between using nb_epoch vs looping of .fit() ?\r\nObviously if there is something like data augmentation in the loop, there will be a new random augmentation of the data before each fit but excluding this type of stuff, are they equivalent?\r\ni.e.\r\nis \r\n\r\n    for x in range(5):\r\n        model.fit(X,y)\r\n\r\n\"equivalent to\" model.fit(X,y,nb_epoch=5)\r\n\r\n\r\n[/quote]\r\n\r\nYes, it is basically the same thing. As you noticed, I used nb_epoch=1 in the tutorial just because of the data augmentation for each epoch.\r\n",
    "106927": "@Phunter - would you be able to post the ported model in Mxnet Tutoral discussion forum? I am especially intersted how you ported the linear regression model to Mxnet, as i have failed to do so succesfully",
    "106924": "@WD since it is Keras topic and not a good place of discussing MXnet due to some reasons, please refer to MXnet's official IO document page http://mxnet.readthedocs.org/en/latest/python/io.html there image IO function has augment operations.",
    "106920": "Hi, I had a general keras question: what is the difference between using nb_epoch vs looping of .fit() ?\r\nObviously if there is something like data augmentation in the loop, there will be a new random augmentation of the data before each fit but excluding this type of stuff, are they equivalent?\r\ni.e.\r\nis \r\n\r\n    for x in range(5):\r\n        model.fit(X,y)\r\n\r\n\"equivalent to\" model.fit(X,y,nb_epoch=5)\r\n",
    "106889": "@phunter. i have been trying to port this linear regression script to mxnet, but have had trouble to get the model to reach good results. My systole CRPS values are 0.0550, rather than the 0.0350 values that i would get on the logistic approach on the MxNet tutoral. I am likely doing something wrong. Would you be willing to post your model? Potentially I am making mistakes in porting some of the Keras features (border_mode, etc) to MxNet. I am more than happy to potentially then help to upload this model to the MxNet depository to help others as well! Many thanks in advance! ",
    "106824": "[quote=phunter;105652]\r\n\r\n[quote=WD;105630]\r\n\r\n@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.\r\n\r\n* Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?\r\n\r\n* You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? \r\n\r\n* Does the code while len(images) < 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? \r\n\r\nmany thanks, W\r\n \r\n\r\n[/quote]\r\n\r\nKeras tutorial network looks promising: I ported it to MXnet and got my current rank with 0.030x score without data augment (well, I should do it). More details: I train mxnet 128x128 on my poor GTX 960. It takes about 54 seconds per epoch and 800 MB memory. Just a benchmark :-) \r\n\r\nUpdate: Seems like this post has caused some confusions. Please, the statement above didn't mean any comparisons between MXnet and Keras. It was a simple run of MXnet with this tutorial's network from my own hardware. @Marko and I had different hardware, different image resolution etc, there was no apple-to-apple comparison of these two toolkits. Keras and Mxnet are both good deep learning frameworks, and it is the same with caffe/theano/torch/Lasagne etc which may later on have good tutorials for this Kaggle competition too, so please select your favorite one or ones for winning 200k $. And sorry to @Marko and @WD for this topic divergence. \r\n\r\n[/quote]\r\n\r\n@phunter  I also went through the MXNET tutorial for NDSB2. I noticed that it did not involve data augmentation, nor evaluation of validation. Perhaps due to its CSV loading fashion instead of reading from images directly? Do you have any idea to implement both augmentation and validation based on the MXNET tutorial?\r\n",
    "106603": "@EIGSI. That is strange indeed. Will think about this more. Woudl you be able to check if your predictions (linear regression) are normally distributed around the actual values? I have a weird bug that they all seem to be larger (rather than normaly distributed...) ",
    "106532": "I can't understand why the performance of the diastole network is consistently worse than that of systole? It is the same input data, same model why the difference?\r\nI noticed this on mxnet example too, which is using the frame differences. In terms of CRPS there is always a .01 difference and in terms of RMSE here the difference seems to be around 10.\r\n\r\n[quote=Marko Jocic;106509]\r\n\r\n \r\n\r\nYou should expect at least ~20 for systole model and at least ~35 for diastole model. Anything higher would lead to a bad result (>0.04).\r\n\r\n[/quote]\r\n",
    "106483": "I am moving my model from logistic regression based to linear regression based (with RSME metric). Just as a quick question - what are reasonable root mean squared error (RMSE) to observe during training? just order of magnitude? That will help me to check and compare with previous CRPS scores and whether i have implemented the model correctly. Currently my model has TRAIN-RSME around 25 and Validation-RSME around 95 - which seems high. Also - the predictions tend to be in nearly alll cases higher than the actual values for both diastole and systole - which seems odd. please let me know.",
    "106465": "Thanks a lot for the tutorial. Since I am using a laptop I am trying to run iterations in chunks by saving model/weights, and reloading in later runs. I have added:\r\n>     model_systole.load_weights('weights_systole.hdf5')\r\n>     model_diastole.load_weights('weights_diastole.hdf5')\r\nTo train,py, and model_from_json failed to load some reason after model_systole.to_json so utilized get_model like submission.py.  Also read val_loss saved from the last run.  Couple questions I have\r\n1) Could it be weights_systole_best better to use than weights_systole weights?\r\n2) Could I be missing anything else  to preload the previous model /run in train.py?\r\n3) Is there any difference using get_model versus to_json/model_from_json? Any idea why the later is not working in this model?\r\n",
    "106324": " Yes, the recent changes have fixed the problem completely. Thanks so much for looking into this!",
    "106317": "Thank you for confirming that it now works. Glad you had good result with it ;).",
    "106316": "@Marko that was the problem.\r\n\r\nI was about to find the problem. I changed the submission.py to load the models and train data and calculate the CRPS on the train data to debug the problem. I did find that loading the best weights gave me  ~0.08 and loading the last weights gave me <0.03.\r\nHowever, I did not get a chance to discover the bug when I saw your post.\r\nThanks.",
    "106240": "During the training, diastole weights from the best iteration are not saved properly.\r\nSystole weights were saved instead of diastole weights. So I just changed:\r\n\r\n``` model_systole.save_weights('weights_diastole_best.hdf5', overwrite=True)```\r\n\r\nto\r\n\r\n``` model_diastole.save_weights('weights_diastole_best.hdf5', overwrite=True)```\r\n\r\nI also just pushed this change to repo, so you can update it.\r\nSo there are two options, either retrain the model (which I recommend), or in submission.py change the weights loaded (weights_diastole.hdf5 instead of weights_diastole_best.hdf5).  Second option might not give the best result since it uses the weights from the last iteration, instead of weights from the best iteration.\r\n\r\nCheers,\r\n\r\nMarko",
    "106239": "It seems that the problem was in val_loss.txt file, mine looks like:\r\n\r\n26.8219207778\r\n\r\n43.2115525287\r\n\r\nHow about yours?",
    "106235": "[quote=piotr;106225]\r\n\r\n@Marko, @patruf @ Wei Wu,\r\n\r\nI was able to get 0.359 with Marko code, but I made changes in submission.py or in val_los.txt files - I don't remember precisely now. But the network trained with train.py is OK. There is only a problem with submission.py.\r\n\r\n[/quote]\r\n\r\nWould you bother to check or recall what did you change? I'm getting blind here trying to see what could cause the problem...",
    "106226": "There is probably a problem with submission generation. Will look into it now.",
    "106199": " Nope, got 0.084... just now, and the warning was just \"downsample module has been moved to the pool module\" but I don't think that should change anything.",
    "106150": "I am using it right now to generate a submission. I did have one question; are the original image files necessary to keep after the .npy files have been generated? The .npy files are much smaller in size and would fit much better on the SSD connected directly to the motherboard, whereas the original dataset image files need to reside on an external drive. Thanks very much for this! I was looking to start using Keras for this and other image recognition challenges!",
    "106119": " I tried to check on it but didn't have time this morning. I decided to start over and will check the results tonight. The only thing I changed was in the submission.py file the fo.writerow(fi.next()) to __ next __() for python 3. Also, I get a Keras warning when I run, something similar to warning: dropout module has been moved to pool, but I didn't think it was important since everything else seemed to work. Thanks for all your responses, I will figure it out. I really like the model, I think everything is well documented and understandable.",
    "106117": "@patruff\r\n\r\nYes, they do. Since you got CRPS(test) = ~0.0347, you should expect somewhat similar result on submission (probably just a bit higher though).\r\nSince the fix, I had multiple people confirming that the code now works properly (I also checked myself to make sure). From the top of my head, I can't think of anything right now that would produce such a bad result. Hm, maybe submission.csv file wasn't overwritten properly, could you check?",
    "106077": "@piotr\r\n\r\nYou are welcome. Thanks for your understanding and patience!",
    "106047": "Maybe it's related: I had my PR just merged which fixed a bug in kera's ImageDataGenerator https://github.com/fchollet/keras/issues/1551",
    "106031": "For anyone interested and using older graphics cards (like myself, using cuda 3.0) I made a docker image that uses tensorflow gpu + keras with python3.  Images are a bit big and would love feedback to make them smaller but need to put the dockerfiles somewhere.  \r\n\r\ndocker pull grahama/tf:keras\r\n\r\nAlso should probably note to run it, best to use https://github.com/tensorflow/tensorflow/blob/master/tensorflow/tools/docker/docker_run_gpu.sh\r\nsince official nvidia-docker thing was not working for me.\r\nso for instance to run a container with data/ and these python files in /root/\r\n\r\ndocker_run_gpu -v /root/:/root/ grahama/tf:keras\r\n\r\nthen in container:\r\n\r\npython3 data.py && python3 train.py\r\n\r\n\r\nOne last thing, \r\nawesome that fchollet is here, love (and have contributed to) keras.  Would anyone be able to explain the significance of batch_size in relation to loss and val_loss?  I thought I had read it was 'just' a speed thing (and that smaller was possibly better https://github.com/fchollet/keras/issues/68#issuecomment-95413935 ) but maybe I am thinking of a different batch_size since it seems like small batch_size for this script is faster than larger batch_size but makes accuracy poor. ",
    "105980": " Marko, no apology necessary. Thanks for fixing things so quickly. I'm going to try it out again later tonight.",
    "105939": " Make that a third person. 0.08... for me without changing anything except I ran it using python3. Thanks for looking into it Marko.",
    "105810": " Thanks Marko for sharing this tutorial and your explanations!",
    "105808": "Thx again Marko.",
    "105793": "@WD\r\n\r\nOk, so basically every model is imperfect/imprecise - i.e. sometimes it gets things right, sometimes it doesn't. We usually represent this imprecision by value of loss function - the lower it is the better our model. So, in a tutorial I use RMS error as a loss function, which means if value of loss function is, say, 20 - that is an indicator that my model usually misses the true value by ~20ml. I wanted to incorporate this uncertainty in the calculation of CDF - ideally, if loss was 0 (model predicts everything right) the CDF would be a step function, but if it is >0 the CDF has that sigmoid-like part around the predicted volume. The larger the uncertainty, the wider is that sigmoid-like area. Because of this, using RMSE value for sigma seemed very natural.\r\n\r\nhist_systole.history['loss'][-1] is a Keras specific part of code, since function fit() returns a History object, which has a current value of loss function both on train and test split ('loss' and 'val_loss', respectively). Since hist_systole.history['loss'] returns an array, I just pick up the last value with [-1].\r\n\r\nAlso, the fundamental advantage of predicting one real value to 600 [0,1] values is simple because it is more natural for this problem! (well, to me at least). If the real volume is N and you predict volume N+1, its no big deal, RMSE is just 1. But imagine you have 600 [0,1] values, even if the real volume corresponds to N-th output and if you predict it is on N+1-th output, the error would be same as if you predicted it is on N+10-th output. That is you are using softmax at output (1 of N classification). I assume it would be better to use sigmoids, but still I found using linear regression as the most appealing approach.\r\n",
    "105738": "To use gpu add this to the  import section of your python script\r\n\r\nimport os\r\n\r\nos.environ['THEANO_FLAGS'] = 'device=gpu'",
    "105719": "One more question please. Iteration 1 took 6-7 hours on my system (XPS8100, no decent GPU). Does that sound right given this poor hardware, or I should suspect some problem with the software? \r\n\r\n\r\n \r\n\r\n       runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\r\n        Using Theano backend.\r\n        Loading and compiling models...\r\n        Loading training data...\r\n        Pre-processing images...\r\n        5331/5331 [==============================] - 934s   C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\theano\\tensor\\signal\\downsample.py:5: UserWarning: downsample module has been moved to the pool module.\r\n          warnings.warn(\"downsample module has been moved to the pool module.\")\r\n        C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py:39: DeprecationWarning: using a non-integer number instead of an integer will result in an error in the future\r\n          X_test = X[:split, :, :, :]\r\n\r\n    \r\n    --------------------------------------------------\r\n    Training...\r\n    --------------------------------------------------\r\n    --------------------------------------------------\r\n    Iteration 1/150\r\n    --------------------------------------------------\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5542s - loss: 34.4657 - val_loss: 42.3799\r\n    Fitting diastole model...\r\n    Epoch 1/1\r\n    4290/4265 [==============================] - 5523s - loss: 55.8717 - val_loss: 86.3558\r\n    Evaluating CRPS...\r\n    4265/4265 [==============================] - 1766s     \r\n    4265/4265 [==============================] - 1763s     \r\n    1066/1066 [==============================] - 441s     \r\n    1066/1066 [==============================] - 439s     \r\n    CRPS(train) = 0.0785559420969\r\n    CRPS(test) = 0.0743980740534\r\n    Saving weights...\r\n    --------------------------------------------------\r\n    Iteration 2/150\r\n    --------------------------------------------------\r\n",
    "105705": "Many thanks! \r\n\r\nPS: for windows users, if git gives you some trouble, this [link][1] might help:  \r\n\r\n[1]: https://github.com/notanumber/gitst2/issues/10",
    "105697": "Many thanks for this keras tutorial! \r\n\r\nData.py ran well (Anaconda/Spyder, windows 64). Keras installation seems ok, but Train.py won't work. Any idea how to fix the following? \r\n\r\n    >>> runfile('C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master/train.py', wdir='C:/Users/Woolsey/Desktop/scibowl/kaggle-dsb2-keras-master')\r\n    Using Theano backend.\r\n    Loading and compiling models...\r\n    Traceback (most recent call last):\r\n      File \"<stdin>\", line 1, in <module>\r\n      File \"C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\spyderlib\\widgets\\externalshell\\sitecustomize.py\", line 699, in runfile\r\n        execfile(filename, namespace)\r\n(more in Woolsey_error.txt attached)\r\n\r\n      File \"C:\\Users\\Woolsey\\Anaconda\\lib\\site-packages\\keras\\backend\\theano_backend.py\", line 463, in relu\r\n        x = T.nnet.relu(x, alpha)\r\n    AttributeError: 'module' object has no attribute 'relu'\r\n    >>> ",
    "105684": "It can be hard to see that there is no need to train 2 separate models, I had to convince myself that this was the case by answering a few questions.\r\n\r\nAssuming there was no systolic volume in the competition just diastolic \r\n\r\n 1. Will training a model to predict patient 1 – 500 labels for patient give you the diastolic distribution? YES\r\nAssuming there was no diastolic volume in the competition just systolic \r\n 2. Will training a model to predict patient 1 – 500 labels for patient give you the systolic distribution? YES\r\n\r\n\r\n**a.ha mo.ment**\r\n\r\n 1. Is there any difference in the labeling and training procedure for the 2 models above? NO (same label same data)\r\n\r\nTraining a network to predict patient learns both the systolic and diastolic volumes simultaneously :) \r\n\r\n**Send the check in the mail…**\r\n",
    "105679": "cuDNN v3 gave me 5x speedup per epoch... using lasagne I had to jump through hoops to get it installed on my aws gpu because NVidia makes you register before downloading it, makes no sense!..I am still waiting on my registration approval.  A little trick I tried to avoid training 2 models was instead of predicting systole and diastole vols, Just predict patient 0-500 labels. Train a single model then translate the patient predictions to systole and diastole volumes. This got me a score of 0.040xxx. Which was as good as I got with 2 separate models for systole and diastole. I am still not convinced this is not a viable strategy. What is the difference btw a network trained for systole and one trained for diastole? Absolutely nothing. The same features are learnt by the net;  ",
    "105675": "[quote=phunter;105673]\r\n\r\nI guess so, haven't had chance to run Keras on my machine but will try. Maybe Keras can try training the model one by one?\r\n\r\n[/quote]\r\n\r\n\r\nOf course, people can modify the code for themselves and train the models one by one.",
    "105673": "[quote=Marko Jocic;105662]\r\n\r\n@phunter\r\n\r\nFrom the top of my head, I think the main difference is that in our example the two models are trained basically at the same time (which doubles the memory usage), while MXnet example trains them one by one.\r\n\r\n[/quote]\r\nI guess so, haven't had chance to run Keras on my machine but will try. Maybe Keras can try training the model one by one?",
    "105635": "@Marko - many thanks. and thanks for the well-documented code. \r\n\r\n* I noticed that you are training on the frames themselves (rather than the difference between frames - as explored by others). It might be interesting by folks to see if this would make a difference\r\n\r\n* Marko - out of curiosity - and from your experience - what do you think the margin of improvement might be if one changes the convolution window sizes, the learning rates, the weight decay and other parameters? Do you feel that this might shave off only a little bit (if at all), or do you think that parameter optimization can make a big difference? So far - my experience is that the real gains lie in preprocessing, and that most network configurations give roughly similar results - but would love to get your expertise and perspective!",
    "105630": "@Marko. This looks great. I had some questions after scanning the code - would love to get your perspective and thoughts.\r\n\r\n* Would you expect better results if training with 96*96 images? or do you expect that the difference will be marginal?\r\n\r\n* You noted that the time needed for each epoch is ~100 seconds, and whole iteration (training both models + CRPS evaluation) takes around 5 minutes. However - if it runs it for 150 epochs at 100 seconds - would that not take a lot longer than 5 minutes? \r\n\r\n* Does the code while len(images) < 30: images.append(images[x]) ; x += 1 copy in additinal images for those missing images in those folders with less than 30 images? \r\n\r\nmany thanks, W\r\n ",
    "105610": "@fchollet: Thanks, upgrading to Keras 0.3.1 from 0.3.0 fixed the border_mode issue. ",
    "105604": "Thanks posting keras tutorial. Is keras needs to be updated? Mine version is only a month old, using python 2.7.10, I've tried train.py but it fails as:\r\n\r\n    X = self.get_input(train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py\", line 102, in get_input\r\n    return self.previous.get_output(train=train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py\", line 512, in get_output\r\n    X = self.get_input(train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/core.py\", line 102, in get_input\r\n    return self.previous.get_output(train=train)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/layers/convolutional.py\", line 215, in get_output\r\n    dim_ordering=self.dim_ordering)\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/keras/backend/theano_backend.py\", line 543, in conv2d\r\n    border_mode=(pad_x, pad_y))\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/theano/sandbox/cuda/dnn.py\", line 1192, in dnn_conv\r\n    conv_mode=conv_mode, precision=precision)(img.shape,\r\n  File \"/home/esorar/theano/t_env/lib/python2.7/site-packages/theano/sandbox/cuda/dnn.py\", line 260, in __init__\r\n    border_mode = tuple(map(int, border_mode))\r\nTypeError: int() argument must be a string or a number, not 'TensorVariable'",
    "111322": "",
    "111314": ""
  }
}