{
  "id": 21277,
  "title": "Keras : Can I run a model training with small epoch in a loop",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/21277",
  "author_name": "",
  "post_date": "2016-05-28T08:04:46.367Z",
  "votes": null,
  "comment_count": 6,
  "views": 2423,
  "content": "<p>Hi, I am new to Data Science, and very new to Image processing, actually this will be my first attempt. Hence my question may be naive. (Also kindly do not blast me for no submissions. I am a genuine person and trying to learn from Kaggle. I am trying, but it is taking humongous amount of time to train these models with my limited resources). I want to train a Keras model with a small epoch size( Say 5). Then I want to put this training in a loop. The loop will run for 10 iterations. I will state the reason as to why I am doing this. I don't have access to powerful hardware. A simple epoch on my machine takes 3 Hrs to run.I was trying to run for 50 epochs. 3 days into the training, Windows crashed and I lost every thing, time and whatever progress the model made. My sample code would be like this: ============================================== </p>\n\n<h1>for i in range(10): print(i) model.fit(X_train, Y_train, batch_size=batch_size, nb_epoch=5, shuffle=True, verbose=1, validation_data=(X_valid, Y_valid), callbacks=callbacks) print(&quot;Saving model....&quot;) save_model(model) </h1>\n\n<p>My idea is to run the model for small epochs, save the model, re load, recompile and continue with training for the next round with the same data. I believe this is equivalent to running the model for 50 epochs. From what I read about CNN and Keras, I understand this is possible as these models can be trained incrementally. Just wanted the advice of the experts as to whether this will work fine, or my understanding is wrong and I am making a fundamental mistake. Apologies for the long post. Thanks and regards Sayantan Raha</p>",
  "messages": [
    {
      "id": "121639",
      "postDate": "05/28/2016 08:04:46",
      "content": "<p>Hi, I am new to Data Science, and very new to Image processing, actually this will be my first attempt. Hence my question may be naive. (Also kindly do not blast me for no submissions. I am a genuine person and trying to learn from Kaggle. I am trying, but it is taking humongous amount of time to train these models with my limited resources). I want to train a Keras model with a small epoch size( Say 5). Then I want to put this training in a loop. The loop will run for 10 iterations. I will state the reason as to why I am doing this. I don't have access to powerful hardware. A simple epoch on my machine takes 3 Hrs to run.I was trying to run for 50 epochs. 3 days into the training, Windows crashed and I lost every thing, time and whatever progress the model made. My sample code would be like this: ============================================== </p>\n\n<h1>for i in range(10): print(i) model.fit(X_train, Y_train, batch_size=batch_size, nb_epoch=5, shuffle=True, verbose=1, validation_data=(X_valid, Y_valid), callbacks=callbacks) print(&quot;Saving model....&quot;) save_model(model) </h1>\n\n<p>My idea is to run the model for small epochs, save the model, re load, recompile and continue with training for the next round with the same data. I believe this is equivalent to running the model for 50 epochs. From what I read about CNN and Keras, I understand this is possible as these models can be trained incrementally. Just wanted the advice of the experts as to whether this will work fine, or my understanding is wrong and I am making a fundamental mistake. Apologies for the long post. Thanks and regards Sayantan Raha</p>",
      "rawMarkdown": "Hi, I am new to Data Science, and very new to Image processing, actually this will be my first attempt. Hence my question may be naive. (Also kindly do not blast me for no submissions. I am a genuine person and trying to learn from Kaggle. I am trying, but it is taking humongous amount of time to train these models with my limited resources). I want to train a Keras model with a small epoch size( Say 5). Then I want to put this training in a loop. The loop will run for 10 iterations. I will state the reason as to why I am doing this. I don't have access to powerful hardware. A simple epoch on my machine takes 3 Hrs to run.I was trying to run for 50 epochs. 3 days into the training, Windows crashed and I lost every thing, time and whatever progress the model made. My sample code would be like this: ============================================== \r\nfor i in range(10): print(i) model.fit(X_train, Y_train, batch_size=batch_size, nb_epoch=5, shuffle=True, verbose=1, validation_data=(X_valid, Y_valid), callbacks=callbacks) print(\"Saving model....\") save_model(model) \r\n=========================================================== \r\nMy idea is to run the model for small epochs, save the model, re load, recompile and continue with training for the next round with the same data. I believe this is equivalent to running the model for 50 epochs. From what I read about CNN and Keras, I understand this is possible as these models can be trained incrementally. Just wanted the advice of the experts as to whether this will work fine, or my understanding is wrong and I am making a fundamental mistake. Apologies for the long post. Thanks and regards Sayantan Raha",
      "votes": null
    },
    {
      "id": "121645",
      "postDate": "05/28/2016 09:58:12",
      "content": "<p>1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.</p>\n\n<p>2, You could save the model to disk, reload the model from disk and continue the training phase afterwards.</p>\n\n<p>3, As you used &quot;fit&quot; function, did you load all the image into memory? The resized images might be too small and even a human could not identify correctly. You may want to check the &quot;fit_generator&quot; function instead.</p>",
      "rawMarkdown": "1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.\r\n\r\n2, You could save the model to disk, reload the model from disk and continue the training phase afterwards.\r\n\r\n3, As you used \"fit\" function, did you load all the image into memory? The resized images might be too small and even a human could not identify correctly. You may want to check the \"fit_generator\" function instead.",
      "votes": null
    },
    {
      "id": "121651",
      "postDate": "05/28/2016 12:01:38",
      "content": "<p>for keras you can also use the callbacks and checkpoint to save your weights to disk periodically like below,</p>\n\n<p>model = VGG_16()</p>\n\n<p>checkpointer = ModelCheckpoint(filepath=&quot;kerasweightsvgg16.pkl&quot;, verbose=1, save_best_only=True)</p>\n\n<p>model.fit_generator(.......,   callbacks=[checkpointer])</p>",
      "rawMarkdown": "for keras you can also use the callbacks and checkpoint to save your weights to disk periodically like below,\r\n\r\nmodel = VGG_16()\r\n \r\ncheckpointer = ModelCheckpoint(filepath=\"kerasweightsvgg16.pkl\", verbose=1, save_best_only=True)\r\n\r\nmodel.fit_generator(.......,   callbacks=[checkpointer])",
      "votes": null
    },
    {
      "id": "121661",
      "postDate": "05/28/2016 14:50:10",
      "content": "<p>Thanks ChaseWind and DavidGbodiOdaibo .</p>\n\n<p>I did resize the picture to 64x64, which as you said is really difficult for a human being to see. I tried with 224 and 128. But they were even slower. </p>\n\n<p>I am loading the entire train set in memory, but training in batches. I will check out the fit_generator and see how I can optimize the entire cycle. Thanks once more. </p>\n\n<p>Regards\nSayantan Raha</p>",
      "rawMarkdown": "Thanks ChaseWind and DavidGbodiOdaibo .\r\n\r\nI did resize the picture to 64x64, which as you said is really difficult for a human being to see. I tried with 224 and 128. But they were even slower. \r\n\r\nI am loading the entire train set in memory, but training in batches. I will check out the fit_generator and see how I can optimize the entire cycle. Thanks once more. \r\n\r\n\r\nRegards\r\nSayantan Raha",
      "votes": null
    },
    {
      "id": "121671",
      "postDate": "05/28/2016 15:49:15",
      "content": "<p>[quote=ChaseWind;121645]\n1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.\n[/quote]</p>\n\n<p>The on-demand price is $0.65/hour. I found it hard to get spot pricing that was much cheaper and reliable.</p>\n\n<p>I ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.</p>\n\n<p>The upcoming 1080 looks like a very cost effective option.</p>",
      "rawMarkdown": "[quote=ChaseWind;121645]\r\n1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.\r\n[/quote]\r\n\r\nThe on-demand price is $0.65/hour. I found it hard to get spot pricing that was much cheaper and reliable.\r\n\r\nI ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.\r\n\r\nThe upcoming 1080 looks like a very cost effective option.",
      "votes": null
    },
    {
      "id": "126854",
      "postDate": "07/12/2016 17:35:50",
      "content": "<p>[quote=gauss256;121671]</p>\n\n<p>I ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.</p>\n\n<p>The upcoming 1080 looks like a very cost effective option.</p>\n\n<p>[/quote]</p>\n\n<p>That sounds good. What's the configuration of your machine?</p>",
      "rawMarkdown": "[quote=gauss256;121671]\r\n\r\nI ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.\r\n\r\nThe upcoming 1080 looks like a very cost effective option.\r\n\r\n[/quote]\r\n\r\nThat sounds good. What's the configuration of your machine?",
      "votes": null
    },
    {
      "id": "127355",
      "postDate": "07/14/2016 00:53:26",
      "content": "<p>Ubuntu 14.04, 32GB RAM, 512GB SSD, Nvidia 980 Ti, Intel i7</p>\n\n<p>I'm very happy with it. :)</p>",
      "rawMarkdown": "Ubuntu 14.04, 32GB RAM, 512GB SSD, Nvidia 980 Ti, Intel i7\r\n\r\nI'm very happy with it. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 121645,
      "author_name": "xingyang",
      "author_url": "",
      "post_date": "05/28/2016 09:58:12",
      "content": "<p>1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.</p>\n\n<p>2, You could save the model to disk, reload the model from disk and continue the training phase afterwards.</p>\n\n<p>3, As you used &quot;fit&quot; function, did you load all the image into memory? The resized images might be too small and even a human could not identify correctly. You may want to check the &quot;fit_generator&quot; function instead.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121651,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "05/28/2016 12:01:38",
      "content": "<p>for keras you can also use the callbacks and checkpoint to save your weights to disk periodically like below,</p>\n\n<p>model = VGG_16()</p>\n\n<p>checkpointer = ModelCheckpoint(filepath=&quot;kerasweightsvgg16.pkl&quot;, verbose=1, save_best_only=True)</p>\n\n<p>model.fit_generator(.......,   callbacks=[checkpointer])</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121661,
      "author_name": "sayantan",
      "author_url": "",
      "post_date": "05/28/2016 14:50:10",
      "content": "<p>Thanks ChaseWind and DavidGbodiOdaibo .</p>\n\n<p>I did resize the picture to 64x64, which as you said is really difficult for a human being to see. I tried with 224 and 128. But they were even slower. </p>\n\n<p>I am loading the entire train set in memory, but training in batches. I will check out the fit_generator and see how I can optimize the entire cycle. Thanks once more. </p>\n\n<p>Regards\nSayantan Raha</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121671,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "05/28/2016 15:49:15",
      "content": "<p>[quote=ChaseWind;121645]\n1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.\n[/quote]</p>\n\n<p>The on-demand price is $0.65/hour. I found it hard to get spot pricing that was much cheaper and reliable.</p>\n\n<p>I ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.</p>\n\n<p>The upcoming 1080 looks like a very cost effective option.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126854,
      "author_name": "akashtandon",
      "author_url": "",
      "post_date": "07/12/2016 17:35:50",
      "content": "<p>[quote=gauss256;121671]</p>\n\n<p>I ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.</p>\n\n<p>The upcoming 1080 looks like a very cost effective option.</p>\n\n<p>[/quote]</p>\n\n<p>That sounds good. What's the configuration of your machine?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127355,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "07/14/2016 00:53:26",
      "content": "<p>Ubuntu 14.04, 32GB RAM, 512GB SSD, Nvidia 980 Ti, Intel i7</p>\n\n<p>I'm very happy with it. :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121639": "Hi, I am new to Data Science, and very new to Image processing, actually this will be my first attempt. Hence my question may be naive. (Also kindly do not blast me for no submissions. I am a genuine person and trying to learn from Kaggle. I am trying, but it is taking humongous amount of time to train these models with my limited resources). I want to train a Keras model with a small epoch size( Say 5). Then I want to put this training in a loop. The loop will run for 10 iterations. I will state the reason as to why I am doing this. I don't have access to powerful hardware. A simple epoch on my machine takes 3 Hrs to run.I was trying to run for 50 epochs. 3 days into the training, Windows crashed and I lost every thing, time and whatever progress the model made. My sample code would be like this: ============================================== \r\nfor i in range(10): print(i) model.fit(X_train, Y_train, batch_size=batch_size, nb_epoch=5, shuffle=True, verbose=1, validation_data=(X_valid, Y_valid), callbacks=callbacks) print(\"Saving model....\") save_model(model) \r\n=========================================================== \r\nMy idea is to run the model for small epochs, save the model, re load, recompile and continue with training for the next round with the same data. I believe this is equivalent to running the model for 50 epochs. From what I read about CNN and Keras, I understand this is possible as these models can be trained incrementally. Just wanted the advice of the experts as to whether this will work fine, or my understanding is wrong and I am making a fundamental mistake. Apologies for the long post. Thanks and regards Sayantan Raha",
    "121645": "1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.\r\n\r\n2, You could save the model to disk, reload the model from disk and continue the training phase afterwards.\r\n\r\n3, As you used \"fit\" function, did you load all the image into memory? The resized images might be too small and even a human could not identify correctly. You may want to check the \"fit_generator\" function instead.",
    "121651": "for keras you can also use the callbacks and checkpoint to save your weights to disk periodically like below,\r\n\r\nmodel = VGG_16()\r\n \r\ncheckpointer = ModelCheckpoint(filepath=\"kerasweightsvgg16.pkl\", verbose=1, save_best_only=True)\r\n\r\nmodel.fit_generator(.......,   callbacks=[checkpointer])",
    "121661": "Thanks ChaseWind and DavidGbodiOdaibo .\r\n\r\nI did resize the picture to 64x64, which as you said is really difficult for a human being to see. I tried with 224 and 128. But they were even slower. \r\n\r\nI am loading the entire train set in memory, but training in batches. I will check out the fit_generator and see how I can optimize the entire cycle. Thanks once more. \r\n\r\n\r\nRegards\r\nSayantan Raha",
    "121671": "[quote=ChaseWind;121645]\r\n1, The g2.2xlarge instance on AWS is sufficient for the computational needs of this competition. In most cases, it costs less than $0.2 per hour. Compared with the GPU on laptops, the g2.2xlarge should be faster.\r\n[/quote]\r\n\r\nThe on-demand price is $0.65/hour. I found it hard to get spot pricing that was much cheaper and reliable.\r\n\r\nI ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.\r\n\r\nThe upcoming 1080 looks like a very cost effective option.",
    "126854": "[quote=gauss256;121671]\r\n\r\nI ended up buying a 980 Ti for my own machine and it turned out to be 2.5x faster than AWS.\r\n\r\nThe upcoming 1080 looks like a very cost effective option.\r\n\r\n[/quote]\r\n\r\nThat sounds good. What's the configuration of your machine?",
    "127355": "Ubuntu 14.04, 32GB RAM, 512GB SSD, Nvidia 980 Ti, Intel i7\r\n\r\nI'm very happy with it. :)"
  },
  "source": "meta"
}