{
  "id": 71044,
  "title": "Problems committing kernels (training gets canceled half-way)",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71044",
  "author_name": "",
  "post_date": "2018-11-09T14:30:57.053284100Z",
  "votes": 9,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'm running into problems trying to commit kernels in order to submit my results. When trying to commit my kernel, the console always dies (I have to refresh the page at some later time), and for some reason the training cell of my notebook is always terminated half-way.</p>\n\n<p><strong>More details:</strong></p>\n\n<ol>\n<li>In the editor, I can run the complete notebook without running out of memory / time or anything. It takes around 30 minutes.</li>\n<li>Previously, I committed a smaller version of the model which ran fine and the results were accepted by Kaggle, so I'm sure my code is correct</li>\n<li>When trying to commit, the UI is terrible. The console always stops outputting after a while and all I can do is refresh the page. I've tried waiting for hours but the UI never recovers from this by itself.</li>\n<li>I can check the cell outputs in the final kernel preview. For some reason, my Keras <code>fit()</code> call always ends halfway through the second epoch (see below)</li>\n<li>After canceling the training, the notebook then continues it's course and produces a final submission file.</li>\n<li>Retrying did not help. The total run time seems to hover around 2400 seconds</li>\n<li>The run is always marked as successful, but in the log there is a warning (<code>[NbConvertApp] WARNING | Timeout waiting for IOPub output</code>).</li>\n</ol>\n\n<p>Here is a typical output of my <code>fit</code> cell when run in the commit:</p>\n\n<p><code>\nTrain on 1175509 samples, validate on 130613 samples\nEpoch 1/10\n1175509/1175509 [==============================] - 355s 302us/step - loss: 0.1170 - acc: 0.9543 - val_loss: 0.1049 - val_acc: 0.9574\nEpoch 2/10\n 351744/1175509 [=======&amp;gt;......................] - ETA: 3:57 - loss: 0.1040 - acc: 0.9590\n</code></p>\n\n<p>Does anyone else have this problem? Is there a time limit for single notebook cells?</p>",
  "messages": [
    {
      "id": "418244",
      "postDate": "11/09/2018 14:30:57",
      "content": "<p>I'm running into problems trying to commit kernels in order to submit my results. When trying to commit my kernel, the console always dies (I have to refresh the page at some later time), and for some reason the training cell of my notebook is always terminated half-way.</p>\n\n<p><strong>More details:</strong></p>\n\n<ol>\n<li>In the editor, I can run the complete notebook without running out of memory / time or anything. It takes around 30 minutes.</li>\n<li>Previously, I committed a smaller version of the model which ran fine and the results were accepted by Kaggle, so I'm sure my code is correct</li>\n<li>When trying to commit, the UI is terrible. The console always stops outputting after a while and all I can do is refresh the page. I've tried waiting for hours but the UI never recovers from this by itself.</li>\n<li>I can check the cell outputs in the final kernel preview. For some reason, my Keras <code>fit()</code> call always ends halfway through the second epoch (see below)</li>\n<li>After canceling the training, the notebook then continues it's course and produces a final submission file.</li>\n<li>Retrying did not help. The total run time seems to hover around 2400 seconds</li>\n<li>The run is always marked as successful, but in the log there is a warning (<code>[NbConvertApp] WARNING | Timeout waiting for IOPub output</code>).</li>\n</ol>\n\n<p>Here is a typical output of my <code>fit</code> cell when run in the commit:</p>\n\n<p><code>\nTrain on 1175509 samples, validate on 130613 samples\nEpoch 1/10\n1175509/1175509 [==============================] - 355s 302us/step - loss: 0.1170 - acc: 0.9543 - val_loss: 0.1049 - val_acc: 0.9574\nEpoch 2/10\n 351744/1175509 [=======&amp;gt;......................] - ETA: 3:57 - loss: 0.1040 - acc: 0.9590\n</code></p>\n\n<p>Does anyone else have this problem? Is there a time limit for single notebook cells?</p>",
      "rawMarkdown": "I'm running into problems trying to commit kernels in order to submit my results. When trying to commit my kernel, the console always dies (I have to refresh the page at some later time), and for some reason the training cell of my notebook is always terminated half-way.\n\n**More details:**\n\n 1. In the editor, I can run the complete notebook without running out of memory / time or anything. It takes around 30 minutes.\n 2. Previously, I committed a smaller version of the model which ran fine and the results were accepted by Kaggle, so I'm sure my code is correct\n 3. When trying to commit, the UI is terrible. The console always stops outputting after a while and all I can do is refresh the page. I've tried waiting for hours but the UI never recovers from this by itself.\n 4. I can check the cell outputs in the final kernel preview. For some reason, my Keras `fit()` call always ends halfway through the second epoch (see below)\n 5. After canceling the training, the notebook then continues it's course and produces a final submission file.\n 6. Retrying did not help. The total run time seems to hover around 2400 seconds\n 7. The run is always marked as successful, but in the log there is a warning (`[NbConvertApp] WARNING | Timeout waiting for IOPub output`).\n\nHere is a typical output of my `fit` cell when run in the commit:\n\n```\nTrain on 1175509 samples, validate on 130613 samples\nEpoch 1/10\n1175509/1175509 [==============================] - 355s 302us/step - loss: 0.1170 - acc: 0.9543 - val_loss: 0.1049 - val_acc: 0.9574\nEpoch 2/10\n 351744/1175509 [=======&gt;......................] - ETA: 3:57 - loss: 0.1040 - acc: 0.9590\n```\n\nDoes anyone else have this problem? Is there a time limit for single notebook cells?",
      "votes": null
    },
    {
      "id": "418440",
      "postDate": "11/09/2018 22:40:23",
      "content": "<p>Yes I found the same problem... stops half way through.</p>",
      "rawMarkdown": "Yes I found the same problem... stops half way through.",
      "votes": null
    },
    {
      "id": "418911",
      "postDate": "11/10/2018 21:10:00",
      "content": "<p>same problem here. A easy work around is to put each epoch of training in its own block. </p>",
      "rawMarkdown": "same problem here. A easy work around is to put each epoch of training in its own block.",
      "votes": null
    },
    {
      "id": "418984",
      "postDate": "11/11/2018 03:01:32",
      "content": "<p>I have similar problems, I committed a kernel that should take less than 1 hour, 20 hours later it is still running with no results. </p>",
      "rawMarkdown": "I have similar problems, I committed a kernel that should take less than 1 hour, 20 hours later it is still running with no results.",
      "votes": null
    },
    {
      "id": "419639",
      "postDate": "11/12/2018 10:50:30",
      "content": "<p>I just started a new Kernel and did not run into this problem. Then I added a <code>ModelCheckpoint</code> callback to be able to get the best model of the training afterwards. Suddenly, I ran into the same issue again. This may be related to the problem. It saves the model after each epoch (or well, whenever it improves). Maybe this causes the kernel to cancel it? I will try disabling it and report back.</p>",
      "rawMarkdown": "I just started a new Kernel and did not run into this problem. Then I added a `ModelCheckpoint` callback to be able to get the best model of the training afterwards. Suddenly, I ran into the same issue again. This may be related to the problem. It saves the model after each epoch (or well, whenever it improves). Maybe this causes the kernel to cancel it? I will try disabling it and report back.",
      "votes": null
    },
    {
      "id": "419743",
      "postDate": "11/12/2018 13:55:00",
      "content": "<p>Never mind, that was not the problem. Even when not saving the best model, some cells are canceled, some are not. Here are two more examples (in the same kernel in the same commit!)</p>\n\n<p>```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 208s 167us/step - loss: 0.1391 - acc: 0.9459 - val_loss: 0.1071 - val_acc: 0.9585</p>\n\n<p>F1 Score - epoch: 1 - score: 0.649526 </p>\n\n<p>Epoch 2/4\n1240815/1240815 [==============================] - 201s 162us/step - loss: 0.1045 - acc: 0.9588 - val_loss: 0.1055 - val_acc: 0.9587</p>\n\n<p>F1 Score - epoch: 2 - score: 0.646368 </p>\n\n<p>Epoch 3/4\n 576512/1240815 [============&gt;.................] - ETA: 1:46 - loss: 0.0964 - acc: 0.9616\n```</p>\n\n<p>... Some other stuff ...</p>\n\n<p>```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 157s 127us/step - loss: 0.1620 - acc: 0.9372 - val_loss: 0.1080 - val_acc: 0.9583</p>\n\n<p>F1 Score - epoch: 1 - score: 0.650368 </p>\n\n<p>Epoch 2/4\n1240815/1240815 [==============================] - 152s 122us/step - loss: 0.1074 - acc: 0.9577 - val_loss: 0.1049 - val_acc: 0.9590</p>\n\n<p>F1 Score - epoch: 2 - score: 0.653796 </p>\n\n<p>Epoch 3/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0994 - acc: 0.9607 - val_loss: 0.1051 - val_acc: 0.9587</p>\n\n<p>F1 Score - epoch: 3 - score: 0.649134 </p>\n\n<p>Epoch 4/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0933 - acc: 0.9630 - val_loss: 0.1037 - val_acc: 0.9598</p>\n\n<p>F1 Score - epoch: 4 - score: 0.665111 \n```</p>\n\n<p>So the first cell stops after ~500 seconds, the second one finishes even though it needs 600 seconds to complete!</p>",
      "rawMarkdown": "Never mind, that was not the problem. Even when not saving the best model, some cells are canceled, some are not. Here are two more examples (in the same kernel in the same commit!)\n\n```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 208s 167us/step - loss: 0.1391 - acc: 0.9459 - val_loss: 0.1071 - val_acc: 0.9585\n\n F1 Score - epoch: 1 - score: 0.649526 \n\nEpoch 2/4\n1240815/1240815 [==============================] - 201s 162us/step - loss: 0.1045 - acc: 0.9588 - val_loss: 0.1055 - val_acc: 0.9587\n\n F1 Score - epoch: 2 - score: 0.646368 \n\nEpoch 3/4\n 576512/1240815 [============&gt;.................] - ETA: 1:46 - loss: 0.0964 - acc: 0.9616\n```\n\n... Some other stuff ...\n\n```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 157s 127us/step - loss: 0.1620 - acc: 0.9372 - val_loss: 0.1080 - val_acc: 0.9583\n\n F1 Score - epoch: 1 - score: 0.650368 \n\nEpoch 2/4\n1240815/1240815 [==============================] - 152s 122us/step - loss: 0.1074 - acc: 0.9577 - val_loss: 0.1049 - val_acc: 0.9590\n\n F1 Score - epoch: 2 - score: 0.653796 \n\nEpoch 3/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0994 - acc: 0.9607 - val_loss: 0.1051 - val_acc: 0.9587\n\n F1 Score - epoch: 3 - score: 0.649134 \n\nEpoch 4/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0933 - acc: 0.9630 - val_loss: 0.1037 - val_acc: 0.9598\n\n F1 Score - epoch: 4 - score: 0.665111 \n```\n\nSo the first cell stops after ~500 seconds, the second one finishes even though it needs 600 seconds to complete!",
      "votes": null
    },
    {
      "id": "419837",
      "postDate": "11/12/2018 16:08:43",
      "content": "<p>This seems to be more common than I thought. Check out this public kernel: <a href=\"https://www.kaggle.com/danofer/different-embeddings-with-attention-fork\">https://www.kaggle.com/danofer/different-embeddings-with-attention-fork</a>\nCell 13 is also canceled before completion. Cell 18 should have the same run time but runs fine.</p>\n\n<p>So it may be that this only happens for the first long-running cell and is fine for all later ones. A work-around may be to run the training twice in two separate cells and hope the second time it finishes.</p>",
      "rawMarkdown": "This seems to be more common than I thought. Check out this public kernel: https://www.kaggle.com/danofer/different-embeddings-with-attention-fork\nCell 13 is also canceled before completion. Cell 18 should have the same run time but runs fine.\n\nSo it may be that this only happens for the first long-running cell and is fine for all later ones. A work-around may be to run the training twice in two separate cells and hope the second time it finishes.",
      "votes": null
    },
    {
      "id": "421959",
      "postDate": "11/15/2018 15:55:53",
      "content": "<p>Hi Max,\nI ran into the same problem of commit stopping halfway. I tried your work-around, but in vain. Did you get any other alternative for this?</p>",
      "rawMarkdown": "Hi Max,\nI ran into the same problem of commit stopping halfway. I tried your work-around, but in vain. Did you get any other alternative for this?",
      "votes": null
    },
    {
      "id": "422692",
      "postDate": "11/16/2018 16:32:21",
      "content": "<p>Hi Aafrin, I have not found a work-around, except to split the training into several cells. I have also contacted one of the organizers, who said they will look into it, and then didn't follow up... Maybe if we all contact them and make some ruckus we will be heard :)</p>",
      "rawMarkdown": "Hi Aafrin, I have not found a work-around, except to split the training into several cells. I have also contacted one of the organizers, who said they will look into it, and then didn't follow up... Maybe if we all contact them and make some ruckus we will be heard :)",
      "votes": null
    },
    {
      "id": "422968",
      "postDate": "11/17/2018 06:37:34",
      "content": "<p>just convert into a script.  Works perfectly.</p>",
      "rawMarkdown": "just convert into a script.  Works perfectly.",
      "votes": null
    },
    {
      "id": "428503",
      "postDate": "11/27/2018 11:47:39",
      "content": "<p>I believe I found a reason / workaround for this: When using Keras, simply pass <code>verbose=2</code> to your <code>model.fit()</code> call. The progress bars that get logged by default cause quite a lot of traffic, and it seems this is what causes the cell to be canceled. By setting verbose to 2, you only get one line of log per epoch of training.</p>",
      "rawMarkdown": "I believe I found a reason / workaround for this: When using Keras, simply pass `verbose=2` to your `model.fit()` call. The progress bars that get logged by default cause quite a lot of traffic, and it seems this is what causes the cell to be canceled. By setting verbose to 2, you only get one line of log per epoch of training.",
      "votes": null
    },
    {
      "id": "454130",
      "postDate": "01/11/2019 06:57:36",
      "content": "<p>after I set verbose = 0 in the model.fit(), the problem was solved.</p>",
      "rawMarkdown": "after I set verbose = 0 in the model.fit(), the problem was solved.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 418440,
      "author_name": "venkykrishna",
      "author_url": "",
      "post_date": "11/09/2018 22:40:23",
      "content": "<p>Yes I found the same problem... stops half way through.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 418911,
      "author_name": "sijunhe9248",
      "author_url": "",
      "post_date": "11/10/2018 21:10:00",
      "content": "<p>same problem here. A easy work around is to put each epoch of training in its own block. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 418984,
      "author_name": "agentili",
      "author_url": "",
      "post_date": "11/11/2018 03:01:32",
      "content": "<p>I have similar problems, I committed a kernel that should take less than 1 hour, 20 hours later it is still running with no results. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419639,
      "author_name": "mschumacher",
      "author_url": "",
      "post_date": "11/12/2018 10:50:30",
      "content": "<p>I just started a new Kernel and did not run into this problem. Then I added a <code>ModelCheckpoint</code> callback to be able to get the best model of the training afterwards. Suddenly, I ran into the same issue again. This may be related to the problem. It saves the model after each epoch (or well, whenever it improves). Maybe this causes the kernel to cancel it? I will try disabling it and report back.</p>",
      "votes": null,
      "replies": [
        {
          "id": 419743,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "11/12/2018 13:55:00",
          "content": "<p>Never mind, that was not the problem. Even when not saving the best model, some cells are canceled, some are not. Here are two more examples (in the same kernel in the same commit!)</p>\n\n<p>```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 208s 167us/step - loss: 0.1391 - acc: 0.9459 - val_loss: 0.1071 - val_acc: 0.9585</p>\n\n<p>F1 Score - epoch: 1 - score: 0.649526 </p>\n\n<p>Epoch 2/4\n1240815/1240815 [==============================] - 201s 162us/step - loss: 0.1045 - acc: 0.9588 - val_loss: 0.1055 - val_acc: 0.9587</p>\n\n<p>F1 Score - epoch: 2 - score: 0.646368 </p>\n\n<p>Epoch 3/4\n 576512/1240815 [============&gt;.................] - ETA: 1:46 - loss: 0.0964 - acc: 0.9616\n```</p>\n\n<p>... Some other stuff ...</p>\n\n<p>```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 157s 127us/step - loss: 0.1620 - acc: 0.9372 - val_loss: 0.1080 - val_acc: 0.9583</p>\n\n<p>F1 Score - epoch: 1 - score: 0.650368 </p>\n\n<p>Epoch 2/4\n1240815/1240815 [==============================] - 152s 122us/step - loss: 0.1074 - acc: 0.9577 - val_loss: 0.1049 - val_acc: 0.9590</p>\n\n<p>F1 Score - epoch: 2 - score: 0.653796 </p>\n\n<p>Epoch 3/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0994 - acc: 0.9607 - val_loss: 0.1051 - val_acc: 0.9587</p>\n\n<p>F1 Score - epoch: 3 - score: 0.649134 </p>\n\n<p>Epoch 4/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0933 - acc: 0.9630 - val_loss: 0.1037 - val_acc: 0.9598</p>\n\n<p>F1 Score - epoch: 4 - score: 0.665111 \n```</p>\n\n<p>So the first cell stops after ~500 seconds, the second one finishes even though it needs 600 seconds to complete!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419837,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "11/12/2018 16:08:43",
          "content": "<p>This seems to be more common than I thought. Check out this public kernel: <a href=\"https://www.kaggle.com/danofer/different-embeddings-with-attention-fork\">https://www.kaggle.com/danofer/different-embeddings-with-attention-fork</a>\nCell 13 is also canceled before completion. Cell 18 should have the same run time but runs fine.</p>\n\n<p>So it may be that this only happens for the first long-running cell and is fine for all later ones. A work-around may be to run the training twice in two separate cells and hope the second time it finishes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 421959,
          "author_name": "aafrin",
          "author_url": "",
          "post_date": "11/15/2018 15:55:53",
          "content": "<p>Hi Max,\nI ran into the same problem of commit stopping halfway. I tried your work-around, but in vain. Did you get any other alternative for this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422692,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "11/16/2018 16:32:21",
          "content": "<p>Hi Aafrin, I have not found a work-around, except to split the training into several cells. I have also contacted one of the organizers, who said they will look into it, and then didn't follow up... Maybe if we all contact them and make some ruckus we will be heard :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422968,
          "author_name": "sijunhe9248",
          "author_url": "",
          "post_date": "11/17/2018 06:37:34",
          "content": "<p>just convert into a script.  Works perfectly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 428503,
      "author_name": "mschumacher",
      "author_url": "",
      "post_date": "11/27/2018 11:47:39",
      "content": "<p>I believe I found a reason / workaround for this: When using Keras, simply pass <code>verbose=2</code> to your <code>model.fit()</code> call. The progress bars that get logged by default cause quite a lot of traffic, and it seems this is what causes the cell to be canceled. By setting verbose to 2, you only get one line of log per epoch of training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454130,
      "author_name": "yyhhlancelot",
      "author_url": "",
      "post_date": "01/11/2019 06:57:36",
      "content": "<p>after I set verbose = 0 in the model.fit(), the problem was solved.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "418244": "I'm running into problems trying to commit kernels in order to submit my results. When trying to commit my kernel, the console always dies (I have to refresh the page at some later time), and for some reason the training cell of my notebook is always terminated half-way.\n\n**More details:**\n\n 1. In the editor, I can run the complete notebook without running out of memory / time or anything. It takes around 30 minutes.\n 2. Previously, I committed a smaller version of the model which ran fine and the results were accepted by Kaggle, so I'm sure my code is correct\n 3. When trying to commit, the UI is terrible. The console always stops outputting after a while and all I can do is refresh the page. I've tried waiting for hours but the UI never recovers from this by itself.\n 4. I can check the cell outputs in the final kernel preview. For some reason, my Keras `fit()` call always ends halfway through the second epoch (see below)\n 5. After canceling the training, the notebook then continues it's course and produces a final submission file.\n 6. Retrying did not help. The total run time seems to hover around 2400 seconds\n 7. The run is always marked as successful, but in the log there is a warning (`[NbConvertApp] WARNING | Timeout waiting for IOPub output`).\n\nHere is a typical output of my `fit` cell when run in the commit:\n\n```\nTrain on 1175509 samples, validate on 130613 samples\nEpoch 1/10\n1175509/1175509 [==============================] - 355s 302us/step - loss: 0.1170 - acc: 0.9543 - val_loss: 0.1049 - val_acc: 0.9574\nEpoch 2/10\n 351744/1175509 [=======&gt;......................] - ETA: 3:57 - loss: 0.1040 - acc: 0.9590\n```\n\nDoes anyone else have this problem? Is there a time limit for single notebook cells?",
    "418440": "Yes I found the same problem... stops half way through.",
    "418911": "same problem here. A easy work around is to put each epoch of training in its own block.",
    "418984": "I have similar problems, I committed a kernel that should take less than 1 hour, 20 hours later it is still running with no results.",
    "419639": "I just started a new Kernel and did not run into this problem. Then I added a `ModelCheckpoint` callback to be able to get the best model of the training afterwards. Suddenly, I ran into the same issue again. This may be related to the problem. It saves the model after each epoch (or well, whenever it improves). Maybe this causes the kernel to cancel it? I will try disabling it and report back.",
    "419743": "Never mind, that was not the problem. Even when not saving the best model, some cells are canceled, some are not. Here are two more examples (in the same kernel in the same commit!)\n\n```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 208s 167us/step - loss: 0.1391 - acc: 0.9459 - val_loss: 0.1071 - val_acc: 0.9585\n\n F1 Score - epoch: 1 - score: 0.649526 \n\nEpoch 2/4\n1240815/1240815 [==============================] - 201s 162us/step - loss: 0.1045 - acc: 0.9588 - val_loss: 0.1055 - val_acc: 0.9587\n\n F1 Score - epoch: 2 - score: 0.646368 \n\nEpoch 3/4\n 576512/1240815 [============&gt;.................] - ETA: 1:46 - loss: 0.0964 - acc: 0.9616\n```\n\n... Some other stuff ...\n\n```\nTrain on 1240815 samples, validate on 65307 samples\nEpoch 1/4\n1240815/1240815 [==============================] - 157s 127us/step - loss: 0.1620 - acc: 0.9372 - val_loss: 0.1080 - val_acc: 0.9583\n\n F1 Score - epoch: 1 - score: 0.650368 \n\nEpoch 2/4\n1240815/1240815 [==============================] - 152s 122us/step - loss: 0.1074 - acc: 0.9577 - val_loss: 0.1049 - val_acc: 0.9590\n\n F1 Score - epoch: 2 - score: 0.653796 \n\nEpoch 3/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0994 - acc: 0.9607 - val_loss: 0.1051 - val_acc: 0.9587\n\n F1 Score - epoch: 3 - score: 0.649134 \n\nEpoch 4/4\n1240815/1240815 [==============================] - 151s 122us/step - loss: 0.0933 - acc: 0.9630 - val_loss: 0.1037 - val_acc: 0.9598\n\n F1 Score - epoch: 4 - score: 0.665111 \n```\n\nSo the first cell stops after ~500 seconds, the second one finishes even though it needs 600 seconds to complete!",
    "419837": "This seems to be more common than I thought. Check out this public kernel: https://www.kaggle.com/danofer/different-embeddings-with-attention-fork\nCell 13 is also canceled before completion. Cell 18 should have the same run time but runs fine.\n\nSo it may be that this only happens for the first long-running cell and is fine for all later ones. A work-around may be to run the training twice in two separate cells and hope the second time it finishes.",
    "421959": "Hi Max,\nI ran into the same problem of commit stopping halfway. I tried your work-around, but in vain. Did you get any other alternative for this?",
    "422692": "Hi Aafrin, I have not found a work-around, except to split the training into several cells. I have also contacted one of the organizers, who said they will look into it, and then didn't follow up... Maybe if we all contact them and make some ruckus we will be heard :)",
    "422968": "just convert into a script.  Works perfectly.",
    "428503": "I believe I found a reason / workaround for this: When using Keras, simply pass `verbose=2` to your `model.fit()` call. The progress bars that get logged by default cause quite a lot of traffic, and it seems this is what causes the cell to be canceled. By setting verbose to 2, you only get one line of log per epoch of training.",
    "454130": "after I set verbose = 0 in the model.fit(), the problem was solved."
  },
  "source": "meta"
}